AI Safety Benchmark Lead & Evaluation Architect
ALICE
Alice is seeking a seasoned researcher-leader to helm AI safety benchmarks and evaluations. You will own the taxonomy, harness, and release cadence, coordinating researchers and freelancers while reporting to the CTO office. You will influence the public research agenda and shape quarterly plans with cross-functional leads. You will collaborate with top AI labs and universities, ensure rigorous evaluation standards, and mentor a growing team of researchers. #J-18808-Ljbffr Alice
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Safety Benchmark Lead & Evaluation Architect in New York, NY vacancy
- Rise Data Labs seeks a Media & Information AI Training Specialist to craft realistic media tasks and rubrics for... ..., editing, writing, or video production experience and comfort evaluating professional-grade work to benchmark AI performance. #J-18808-Ljbffr BAM VenturesSuggested
- Mercor is partnering with a leading AI research organization to engage experienced sales engineers for a project focused on evaluating how well AI systems perform real-world technical sales work. You will define what excellent work looks like by designing task-specific...Suggested
- Mercor is partnering with a leading AI research organization to engage experienced UI/UX and product designers for a project that evaluates how well AI systems perform real-world digital product design work. You will define what excellent work looks like by designing task...Suggested
- ...high-quality academic assessment content for an AI research initiative. You will write and verify rigorous... ...choice questions across core psychology domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. You will be...SuggestedRemote job
- ...high-quality academic assessment content for an AI research initiative. You will write and verify rigorous... ...questions across core mathematics domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. You will be assigned...SuggestedRemote job
- ...high-quality academic assessment content for an AI research initiative. You will write and verify rigorous... ...choice questions across core psychology domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. You will be...Remote job10 hours per week
- A leading AI research firm in New York, NY, is looking for a specialist to oversee the red-teaming and adversarial evaluation pipeline for their models. The ideal candidate should possess a graduate... ...have a deep understanding of LLM safety and adversarial techniques. A...
- Mercor seeks experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous topics. Responsibilities...
- ...A tech-focused research lab is seeking Property Manager Experts to design and evaluate property management challenges for training AI models. This remote role allows you to leverage your expertise in real estate operations while making a significant impact on AI solutions...Remote work
$110k - $120k
...transforming strategies into leading-edge tech platforms, at scale... ...including engineering, data, AI, cybersecurity, product and technology... ...and Google Cloud Platform Evaluate scalability, resilience,... ...Collaborate with client architects, engineers, product leaders and...Full timeSummer workInternshipWork at officeLocal area$190k - $199k
...SS&C is a leading provider of mission-critical, AI-powered technology and services empowering financial services and healthcare organizations to work... ...outcomes, operational priorities and regulatory requirements. Evaluate opportunities based on value, feasibility and customer...Ongoing contractFull timeWorldwide$270k - $330k
...cybersecurity SaaS company, focused on shaping AI strategy and designing large-scale... ..., production-grade systems. Evaluate and define AI governance, safety, and compliance frameworks... ...intelligence domains. ~ Experience leading technical strategy and influencing across...Full timeWork at office$225k - $304.2k
...make an impact? Applied AI Engineer/Architect West Monroe is seeking... ...AI Engineer/Architect to lead the design and delivery of... ...strategies, workflow orchestration, evaluation frameworks, observability,... ...methodologies using benchmark datasets, regression testing...Full timeLocal areaImmediate startFlexible hours- Senior Agentic AI Engineer Are you looking to join an innovative, market-leading company where you can truly elevate your career... ..., agentic workflows, quality evaluators, constraint validation, and... ..., product-oriented solutions. Architect large-scale graph, semantic, and...Work at officeRemote workFlexible hours
- Mercor is building high-fidelity simulated work environments to evaluate and improve AI agents on real marketing work. You will help design the briefs, artifacts, the judgment calls, and the errors a weaker practitioner would miss. Tasks span diagnosing, prioritizing, planning...
- Mercor is seeking senior litigation professionals to build evaluation tasks for AI systems operating in civil litigation and dispute resolution contexts. The workflows are calibrated to case complexity, evidentiary stakes, and procedural scope of major commercial litigation...
$80 - $90 per hour
AI Architect -26-02164 Hybrid/Onsite in NYC 2yrs Duration TempW2orC2C NTT DATA strives... ...-native scale‑up) Experience building evaluation systems that proved the agent was reliable... ...and connectivity. We are one of the leading providers of digital and AI infrastructure...Hourly payTemporary workRemote workFlexible hours- Mercor is seeking senior insurance professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise insurance and risk contexts. The workflows align with underwriting complexity, regulatory stakes, and large carrier scales. Contributors design...
- Position: Generative AI Architect Location: Allentown PA(Remote) Generative AI Architect The... ...objectives. This role involves leading the architecture and deployment of advanced... ...preprocessing, feature engineering, and model evaluation techniques. Knowledge of API...Remote job
- Mercor is building high-fidelity simulated work environments to evaluate AI agents in real marketing work. You will design and pressure-test Field Marketing environments, including briefs, artifacts, and the decision criteria that separate strong responses from plausible...
$80 per hour
Ethos is seeking experienced curriculum coordinators to train an AI language model on professional document, spreadsheet, and slide deck tasks. The role focuses on creating, evaluating, and refining AI-generated curricula across core workflows in a K-12 setting. Flexible...Remote jobImmediate startFlexible hours- ...including CAD/CAE, PLM, MBSE, digital twins, and AI-powered engineering tools. Establish a... ..., and semantic retrieval frameworks. Lead architecture for AI-enabled engineering use... ...vendors and research organizations to evaluate emerging AI and graph technologies. Governance...
- Overview The AI Solutions Architect defines the technical direction for AI, ML, and data‑driven capabilities across the enterprise. The role... ...pipelines, training workflows, inference layers, and governance. Evaluate emerging AI and cloud capabilities and guide technology...
- Senior AI Architect BMO is establishing a dedicated AI Engineering function, and this role is... ...adjudicate Microsoft AI boundary questions, and evaluate preview features against bank governance... ...and federation with the AI Registry. Lead assessment of emerging AI capabilities,...
- Hybrid/Onsite in NYC 2yrs Duration Job Description AI Architect --- Principal/Staff Engineer, Agentic AI Systems Why: Client's AIOps team... ...(or equivalent AI-native scale-up) Experience building evaluation systems that proved the agent was reliable enough to run with...
- Obsidian in New York builds high‑fidelity simulated work environments to evaluate AI agents on real marketing tasks. You will design creator and influencer scenarios, including briefs, artifacts, and the judging criteria that separate strong responses from plausible but...
$55 per hour
A leading technology firm is seeking an Evaluation Scenario Writer - QA to ensure the quality of evaluation scenarios for AI projects. This part-time, flexible role involves reviewing tests, spotting inconsistencies, and collaborating on automation. Ideal for those with...Remote jobPart timeFlexible hours- Senior AI Architect - Finance Technology AI Enablement We are looking for a Senior AI Architect who will serve as the top AI/ML technical... ...high-volume financial data processing and analytics pipelines. Evaluate and select tools, platforms, and frameworks (cloud-native and...
- ...possible. Join us and help the world’s leading organizationsunlock the value of technology... ...seeking a highly experienced Agentic AI Architect to lead the design and implementation of... ...PerformanceScalabilityReliabilityCost optimization Continuously evaluate emerging AI technologies, tools, and...
- ...on behalf of our client for a Generative AI Architect . Role: Our Client - the world’s first... ...thinking . This individual will lead the design, development, and deployment... ...cost optimization in deployed AI systems. Evaluate and integrate emerging AI models, frameworks...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Safety Benchmark Lead & Evaluation Architect. Be the first to apply!

