Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Safety Scientist: Agent Robustness & Risk Evaluation

Scale

Scale Labs seeks a Research Scientist focused on Agent Robustness in a multi-location setting. You will tackle fundamental challenges in building safe AI agents and aligning them with humans, including benchmarking methods and mitigation strategies. Ideal candidates have 3+ years in ML research, a publications track record in ML, and experience with RLHF, DPO, or GRPO, plus strong cross-functional communication. Some familiarity with agent evaluation tooling is a plus. #J-18808-Ljbffr Scale

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Safety Scientist: Agent Robustness & Risk Evaluation in San Francisco, CA vacancy
  • Scale Labs in San Francisco is seeking a Research Scientist focused on Agent Robustness to advance safe, aligned AI. You will tackle fundamental challenges in evaluating agent capabilities, safety, and risk, and help design benchmarks and protocols. You will design harnesses... 
    Risk

    Scale AI, Inc.

    San Francisco, CA
    5 days ago
  • Scale Labs is seeking a Research Scientist focused on Agent Robustness to advance safe and aligned AI agents. You will contribute to evaluating risks, building testing harnesses, and prototyping mitigation strategies across agent interactions and environments. The role... 
    Risk

    United States Digital Space LLC

    San Francisco, CA
    3 days ago
  • A leading AI research organization in San Francisco is seeking a Research Scientist to work on agent robustness. The role involves conducting research on AI safety, designing tests for AI agents, and developing...  ...mitigations for potential risks. Candidates should have a... 
    Risk
    Full time

    Scale AI

    San Francisco, CA
    5 days ago
  • $197.4k - $246.75k

    Scale Labs, Research Scientist — Agent Robustness As the leading data and evaluation partner for frontier AI companies, Scale plays an integral...  ...scientific decisions about AI risks and capabilities. Our...  ...focus on how they relate to safety, risk factors, and methodologies... 
    Risk
    Full time

    Scale AI

    San Francisco, CA
    5 days ago
  • $60 - $70 per hour

     ...technical talent with leading AI research labs. Headquartered in...  ...Dorsey . Position: AI Safety Practitioner Type: Contract...  ...Role Responsibilities Evaluate AI-generated responses for safety...  ...reviewing complex, high-risk, or ambiguous content. Application... 
    Risk
    Contract work
    Summer work
    Remote work

    Mercor

    San Francisco, CA
    4 days ago
  •  ...seeking exceptional research engineers to push the boundaries of frontier AI safety, shaping empirical understanding of risk and owning end-to-end threads within this effort. You’ll design evaluations of frontier models against real threat models, develop datasets,... 
    Risk

    United States Digital Space LLC

    San Francisco, CA
    2 days ago
  • Scale Labs seeks a Research Scientist focused on Frontier Risk Evaluations to design evaluation measures, harnesses and datasets for measuring risks posed by frontier AI systems. You will build harnesses to test models, collaborate with government agencies to scope evaluations... 
    Risk

    Scale

    San Francisco, CA
    3 days ago
  • $216k - $270k

    Scale AI, Inc. is looking for a Research Scientist specializing in Frontier Risk Evaluations to develop measures for assessing risks of advanced AI systems. In this role, you...  ...publish reports to inform policymakers about AI safety. Ideal candidates have a Ph.D., experience... 
    Risk

    Scale AI, Inc.

    San Francisco, CA
    5 days ago
  • An innovative tech company in New York is seeking a Research Scientist focused on Frontier Risk Evaluations. The ideal candidate will contribute to designing and creating evaluation measures for assessing AI risks and will have strong experience in machine learning and... 
    Risk

    Scale AI

    San Francisco, CA
    5 days ago
  • $35 per hour

     ...believe the safest AI is the one that’s already...  ...AI models and agents: jailbreaks, prompt...  ...and flag systemic risks Apply structure: follow...  ...AI systems Evaluation coverage expands: more...  ...trust the safety of their AI because...  ...making AI systems more robust, safe, and trustworthy... 
    Risk
    Remote job

    Obsidian

    San Francisco, CA
    3 days ago
  • $215k - $230k

     ...our trajectory. The AI Engineering Team is chartered...  ...mission is to build robust pipelines, high-...  ...deployed with speed, safety, and scale . We manage...  ...also deeply involved in evaluating and integrating cutting...  ...edge tools in the LLM and agent space — including open... 
    Local area
    Remote work

    Crypto Pro Network

    San Francisco, CA
    4 days ago
  • OpenAI is seeking a researcher to advance frontier evaluations and environments for safe AGI/ASI. You will help design north star model...  ...products. Collaborate with researchers, engineers, product and safety teams to decide what to measure, how to measure it, and how to... 

    Neura Market

    San Francisco, CA
    2 days ago
  • $218.5k - $288k

     ...OpportunityAs an Applied Scientist specializing in Small Language Models and AI Training, you will lead...  ...Design, implement, and evaluate model training...  ...to improve performance, robustness, and generalization of...  ...applications.Ensure AI safety, fairness, and alignment... 
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    6 hours ago
  • $177k - $197k

     ...As an Innovation AI Developer at Kirkland...  ...leaders, data scientists, and...  ...efficiency, quality, risk reduction), and architect...  ...solutions. Develop AI agents that automate...  ...’s AI ecosystem. Evaluation & Model...  ...measure accuracy, robustness, safety, and alignment to... 
    Risk
    Contract work
    Worldwide
    Flexible hours

    Kirkland & Ellis

    San Francisco, CA
    4 days ago
  •  ...important part of the Safety Systems org at...  ...Framework. Frontier AI models have the...  ...severe risks. To ensure that AI...  ...capability assessment, evaluations, and internal red...  ...'re hiring a Data Scientist to help build, evaluate...  ...ability to design robust measurement under... 
    Risk
    Shift work

    OpenAI

    San Francisco, CA
    3 days ago
  • $207k - $285k

     ...a Technical Program Manager in San Francisco to lead initiatives that ensure the safety and robustness of its AI models. The role involves collaborating with diverse teams to turn risks into actionable plans. Ideal candidates will have experience in technical program... 
    Risk

    OpenAI

    San Francisco, CA
    5 days ago
  • $154.39k - $247.02k

     ...society’s most critical safety and justice issues with...  ...where you matter.AI Infrastructure Engineer...  ...applications that use LLMs, AI agents, RAG workflows,...  ...workflows, model APIs, evaluations, and AI safety considerations...  ....Communicate technical risks, tradeoffs, support... 
    Risk
    Work experience placement

    Axon

    San Francisco, CA
    3 days ago
  • Welo Data is looking for a Data Labeling Associate in San Francisco to evaluate AI systems' handling of Arabic nuances. The role requires professional-level proficiency in Arabic and 2 years of AI safety experience. Responsibilities include critiquing Arabic AI outputs... 

    Welo Data

    San Francisco, CA
    1 day ago
  • Welo Data in San Francisco seeks a full-time AI Evaluator with professional proficiency in Portuguese (Portugal) and experience in Generative AI safety. The role involves critiquing AI outputs, identifying biases, and refining evaluation frameworks. Candidates should possess... 
    Full time

    Welo Data

    San Francisco, CA
    3 days ago
  • $211k - $290.5k

     ...Compass — Faire’s user facing AI bet within the Discovery...  ...As a Senior Applied AI/ML Scientist on the Compass team, you...  ...for this product — driving agent quality through data, evaluation, and modeling, while...  ...product bets into sequenced, de-risked tactical plans — what to... 
    Risk
    Work experience placement
    Work at office
    Local area
    Immediate start
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    2 days ago
  • $215.2k - $245.6k

    Lead AI Engineer (Gen AI Platform Services: Agentic AI, Agent Guardrails, Agent Evaluation, Agent Memory) Overview At Capital One, we are creating responsible and reliable...  ...cross-functional team of engineers, research scientists, technical program managers, and product... 
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    1 day ago
  • Reflection is seeking a highly skilled safety researcher to own red-teaming and adversarial evaluation for our open-weight models. You will probe for failure...  ...as a gatekeeper for model releases, balancing risk with ambitious AI capabilities. #J-18808-Ljbffr Visa Hunt
    Risk

    Visa Hunt

    San Francisco, CA
    3 days ago
  • Mercor is building a remote AI red team to stress test conversational models...  ...human data generation, structured evaluation, and clear reporting of risk. You will work across projects with...  ...guidelines and wellness resources in a safety-focused environment. You will review... 
    Risk
    Remote job

    Obsidian

    San Francisco, CA
    3 days ago
  • Variance is building disruptive AI agents to aid fraud and Risk & Compliance teams, enabling faster investigations and more reliable decisions. You design, ship, and iterate production-grade agents that operate in enterprise risk workflows, from pilot to deployment, with... 
    Risk

    Doist

    San Francisco, CA
    2 days ago
  • $200.8k - $251k

    Scale AI in San Francisco is seeking an AI Researcher focused on advancing intelligent agents through data strategy and research publications. This role requires expertise in...  ...range of $200,800 - $251,000, along with equity and robust benefits. #J-18808-Ljbffr Scale AI

    Scale AI

    San Francisco, CA
    5 days ago
  • $204.44k - $324.99k

     ...and configure Agentforce agents, topics, instructions,...  ...templates and grounded AI experiences using...  ...Implement agent testing, evaluation, observability, and guardrails...  ..., hallucination risk, data access, human-in-...  ...insurance, 401(k) plans, and a robust suite of personal well-... 
    Risk
    H1b
    Local area

    KPMG

    San Francisco, CA
    1 day ago
  •  ...Staff AI Agent Engineer San Francisco Bay Area About Us: Liberate builds AI agents to automate manual tasks for the $2.7T insurance...  ...Ability to reason about quality, failure modes, and operational risk ~ Comfort working directly with customers on technically... 
    Risk
    Work at office
    2 days per week

    Liberate

    San Francisco, CA
    1 day ago
  • $40 - $50 per hour

     ...rental properties with occasional opportunities to conduct property evaluations. Average pay ranges between $40-50 per showing. Key...  ...regardless of vacancies. We combine modern data science for efficient risk modeling with tech-powered operations to deliver consistent,... 
    Risk
    Hourly pay
    Local area
    Flexible hours

    Doorstead Inc

    San Francisco, CA
    6 days ago
  • $93.75k - $133.2k

     ...Special Agent Position Overview: The position advertised has...  ...crimes that threaten public safety. Your transition from a specialized...  ...Qualifications and Evaluations GL-09 All applicants must have...  ...vulnerabilities, considering risks, and choosing the best outcome... 
    Risk
    Permanent employment
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start

    ClearanceJobs

    San Francisco, CA
    4 days ago
  • $133k - $391.5k

     ...selves. TheSenior Manager, Engineering - AI Agents will lead the development and delivery...  ...Communicate engineering progress, risks, and technical decisions clearly to stakeholders...  ...in software development workflows Evaluate emerging developer AI tools and... 
    Risk
    Full time
    Temporary work
    Work at office
    Worldwide
    3 days per week

    Five9

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Safety Scientist: Agent Robustness & Risk Evaluation. Be the first to apply!