Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Safety Benchmark Lead & Evaluation Architect

ALICE

Alice is seeking a seasoned researcher-leader to helm AI safety benchmarks and evaluations. You will own the taxonomy, harness, and release cadence, coordinating researchers and freelancers while reporting to the CTO office. You will influence the public research agenda and shape quarterly plans with cross-functional leads. You will collaborate with top AI labs and universities, ensure rigorous evaluation standards, and mentor a growing team of researchers. #J-18808-Ljbffr Alice

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Safety Benchmark Lead & Evaluation Architect in New York, NY vacancy
  • Rise Data Labs seeks a Media & Information AI Training Specialist to craft realistic media tasks and rubrics for...  ..., editing, writing, or video production experience and comfort evaluating professional-grade work to benchmark AI performance. #J-18808-Ljbffr BAM Ventures
    Suggested

    BAM Ventures

    New York, NY
    2 days ago
  • Mercor is partnering with a leading AI research organization to engage experienced sales engineers for a project focused on evaluating how well AI systems perform real-world technical sales work. You will define what excellent work looks like by designing task-specific... 
    Suggested

    Mercor

    New York, NY
    3 days ago
  • Mercor is partnering with a leading AI research organization to engage experienced UI/UX and product designers for a project that evaluates how well AI systems perform real-world digital product design work. You will define what excellent work looks like by designing task... 
    Suggested

    Mercor

    New York, NY
    2 days ago
  •  ...high-quality academic assessment content for an AI research initiative. You will write and verify rigorous...  ...choice questions across core psychology domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. You will be... 
    Suggested
    Remote job

    Mercor

    New York, NY
    4 days ago
  •  ...high-quality academic assessment content for an AI research initiative. You will write and verify rigorous...  ...questions across core mathematics domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. You will be assigned... 
    Suggested
    Remote job

    Obsidian

    New York, NY
    22 hours ago
  •  ...high-quality academic assessment content for an AI research initiative. You will write and verify rigorous...  ...choice questions across core psychology domains, evaluate solution quality, and help establish gold-standard benchmarks used to advance AI capabilities. You will be... 
    Remote job
    10 hours per week

    Obsidian

    New York, NY
    1 day ago
  • A leading AI research firm in New York, NY, is looking for a specialist to oversee the red-teaming and adversarial evaluation pipeline for their models. The ideal candidate should possess a graduate...  ...have a deep understanding of LLM safety and adversarial techniques. A... 

    Reflection AI

    New York, NY
    3 days ago
  • Mercor seeks experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous topics. Responsibilities... 

    Mercor

    New York, NY
    2 days ago
  •  ...A tech-focused research lab is seeking Property Manager Experts to design and evaluate property management challenges for training AI models. This remote role allows you to leverage your expertise in real estate operations while making a significant impact on AI solutions... 
    Remote work

    AfterQuery Experts

    New York, NY
    5 days ago
  • $110k - $120k

     ...transforming strategies into leading-edge tech platforms, at scale...  ...including engineering, data, AI, cybersecurity, product and technology...  ...and Google Cloud Platform Evaluate scalability, resilience,...  ...Collaborate with client architects, engineers, product leaders and... 
    Full time
    Summer work
    Internship
    Work at office
    Local area

    Boston Consulting Group

    New York, NY
    2 days ago
  • $190k - $199k

     ...SS&C is a leading provider of mission-critical, AI-powered technology and services empowering financial services and healthcare organizations to work...  ...outcomes, operational priorities and regulatory requirements. Evaluate opportunities based on value, feasibility and customer... 
    Ongoing contract
    Full time
    Worldwide

    SS&C Technologies

    New York, NY
    2 days ago
  • $270k - $330k

     ...cybersecurity SaaS company, focused on shaping AI strategy and designing large-scale...  ..., production-grade systems. Evaluate and define AI governance, safety, and compliance frameworks...  ...intelligence domains. ~ Experience leading technical strategy and influencing across... 
    Full time
    Work at office

    Clera

    New York, NY
    4 days ago
  • $225k - $304.2k

     ...make an impact? Applied AI Engineer/Architect West Monroe is seeking...  ...AI Engineer/Architect to lead the design and delivery of...  ...strategies, workflow orchestration, evaluation frameworks, observability,...  ...methodologies using benchmark datasets, regression testing... 
    Full time
    Local area
    Immediate start
    Flexible hours

    West Monroe

    New York, NY
    5 days ago
  • Senior Agentic AI Engineer Are you looking to join an innovative, market-leading company where you can truly elevate your career...  ..., agentic workflows, quality evaluators, constraint validation, and...  ..., product-oriented solutions. Architect large-scale graph, semantic, and... 
    Work at office
    Remote work
    Flexible hours

    Kinaxis

    New York, NY
    2 days ago
  • Mercor is building high-fidelity simulated work environments to evaluate and improve AI agents on real marketing work. You will help design the briefs, artifacts, the judgment calls, and the errors a weaker practitioner would miss. Tasks span diagnosing, prioritizing, planning... 

    Mercor

    New York, NY
    4 days ago
  • Mercor is seeking senior litigation professionals to build evaluation tasks for AI systems operating in civil litigation and dispute resolution contexts. The workflows are calibrated to case complexity, evidentiary stakes, and procedural scope of major commercial litigation... 

    Mercor

    New York, NY
    4 days ago
  • $80 - $90 per hour

    AI Architect -26-02164 Hybrid/Onsite in NYC 2yrs Duration TempW2orC2C NTT DATA strives...  ...-native scale‑up) Experience building evaluation systems that proved the agent was reliable...  ...and connectivity. We are one of the leading providers of digital and AI infrastructure... 
    Hourly pay
    Temporary work
    Remote work
    Flexible hours

    NTT DATA North America

    New York, NY
    2 days ago
  • Mercor is seeking senior insurance professionals to build evaluation tasks for AI systems operating in Fortune 500 enterprise insurance and risk contexts. The workflows align with underwriting complexity, regulatory stakes, and large carrier scales. Contributors design... 

    Mercor Inc

    New York, NY
    4 days ago
  • Position: Generative AI Architect Location: Allentown PA(Remote) Generative AI Architect The...  ...objectives. This role involves leading the architecture and deployment of advanced...  ...preprocessing, feature engineering, and model evaluation techniques. Knowledge of API... 
    Remote job

    Georgia IT Inc

    New York, NY
    1 day ago
  • Mercor is building high-fidelity simulated work environments to evaluate AI agents in real marketing work. You will design and pressure-test Field Marketing environments, including briefs, artifacts, and the decision criteria that separate strong responses from plausible... 

    Mercor

    New York, NY
    3 days ago
  • $80 per hour

    Ethos is seeking experienced curriculum coordinators to train an AI language model on professional document, spreadsheet, and slide deck tasks. The role focuses on creating, evaluating, and refining AI-generated curricula across core workflows in a K-12 setting. Flexible... 
    Remote job
    Immediate start
    Flexible hours

    Ethos

    New York, NY
    1 day ago
  •  ...including CAD/CAE, PLM, MBSE, digital twins, and AI-powered engineering tools. Establish a...  ..., and semantic retrieval frameworks. Lead architecture for AI-enabled engineering use...  ...vendors and research organizations to evaluate emerging AI and graph technologies. Governance... 

    PRI Technology

    New York, NY
    2 days ago
  • Overview The AI Solutions Architect defines the technical direction for AI, ML, and data‑driven capabilities across the enterprise. The role...  ...pipelines, training workflows, inference layers, and governance. Evaluate emerging AI and cloud capabilities and guide technology... 

    Saransh Inc

    New York, NY
    2 days ago
  • Senior AI Architect BMO is establishing a dedicated AI Engineering function, and this role is...  ...adjudicate Microsoft AI boundary questions, and evaluate preview features against bank governance...  ...and federation with the AI Registry. Lead assessment of emerging AI capabilities,... 

    Bmo

    New York, NY
    3 days ago
  • Hybrid/Onsite in NYC 2yrs Duration Job Description AI Architect --- Principal/Staff Engineer, Agentic AI Systems Why: Client's AIOps team...  ...(or equivalent AI-native scale-up) Experience building evaluation systems that proved the agent was reliable enough to run with... 

    Shree Narayani Networking Solutions LLC

    New York, NY
    2 days ago
  • Obsidian in New York builds high‑fidelity simulated work environments to evaluate AI agents on real marketing tasks. You will design creator and influencer scenarios, including briefs, artifacts, and the judging criteria that separate strong responses from plausible but... 

    Obsidian

    New York, NY
    4 days ago
  • $55 per hour

    A leading technology firm is seeking an Evaluation Scenario Writer - QA to ensure the quality of evaluation scenarios for AI projects. This part-time, flexible role involves reviewing tests, spotting inconsistencies, and collaborating on automation. Ideal for those with... 
    Remote job
    Part time
    Flexible hours

    Mindrift

    New York, NY
    5 days ago
  • Senior AI Architect - Finance Technology AI Enablement We are looking for a Senior AI Architect who will serve as the top AI/ML technical...  ...high-volume financial data processing and analytics pipelines. Evaluate and select tools, platforms, and frameworks (cloud-native and... 

    RIT Solutions

    New York, NY
    1 day ago
  •  ...possible. Join us and help the world’s leading organizationsunlock the value of technology...  ...seeking a highly experienced Agentic AI Architect to lead the design and implementation of...  ...PerformanceScalabilityReliabilityCost optimization Continuously evaluate emerging AI technologies, tools, and... 

    Capgemini

    New York, NY
    2 days ago
  •  ...on behalf of our client for a Generative AI Architect . Role: Our Client - the world’s first...  ...thinking . This individual will lead the design, development, and deployment...  ...cost optimization in deployed AI systems. Evaluate and integrate emerging AI models, frameworks... 

    Arka Innovate

    New York, NY
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Safety Benchmark Lead & Evaluation Architect. Be the first to apply!