Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Autonomy Safety Evaluations Research Engineer

$315k

Anthropic

We are looking for Research Engineers to build “gold standard” evaluations for catastrophic risks, in order to understand what AI Safety Level (ASL) to assign to models. Research leads on this team collaborate with engineers in one of our focus areas: CBRN, Cyber, Autonomy (this list may expand over time). This will have major implications for the way we train, deploy, and secure our models, as detailed in our Responsible Scaling Policy (RSP). The policy defines a series of capability thresholds – AI Safety Levels (ASLs) – that represent increasing risks – crossing an ASL threshold would trigger a commitment to more stringent safety, security, and operational measures, intended to handle the increased level of risk. Please note: We are currently only hiring for the Autonomous Replication and Adaption (Autonomy) threats workstream. We will also be prioritizing candidates who can start ASAP and can be based in either our San Francisco or London office. Responsibilities: Research Engineers will be responsible for designing and running the evaluations needed to measure dangerous capabilities in models, and determine when we cross an ASL threshold. You’ll lead projects with world class experts in fields like biosecurity, autonomous replication, cybersecurity, and national security, and experiment with new evals, in order to measure how risky AI systems are. Done well, this will inform decisions at the highest levels of the company. You may be a good fit if you: Have an ML-focused background and engineering and research skills (e.g. experience in Python) Have experience managing research programs of dozens of technical and non-technical experts Are driven to find solutions to ambiguously scoped problems Design and run experiments and iterate quickly to solve machine learning problems Thrive in collaborative environment (we love pair programming!) Have experience training, working with, and prompting models For all workstreams, experience designing and building evaluations would be valuable, but is definitely not essential. For National Security threats workstreams, we will particularly value experience working on confidential or sensitive projects and demonstrated integrity, responsibility, and trustworthiness. We will also value domain specific knowledge, although it is not necessary. For ARA threats workstreams, we would value experience with language model agents, although this is not essential. Sample Projects: ARA risks – building infrastructure and tooling for testing for these capabilities, and iterating with external ARA experts to scope possible tasks. This will involve building custom “testing environments” and new infrastructure.

  • Not currently hiring for) CBRN risks – working with external experts in the field of biosecurity to design clear and repeatable CBRN evaluations, based on a summary of dangerous biological capabilities. Using our post training infrastructure to prepare new generations of models for routine evaluations.
  • Not currently hiring for) Cyber risks – working with external cyber experts to co-design a set of clear and repeatable cyber evaluations. This is likely to involve building custom environments or additions onto existing tooling and infrastructure, or locating specialized datasets.
Deadline to apply: None. Applications will be reviewed on a rolling basis. The expected salary range for this position is: Annual Salary:

$315,000—$510,000 USD

Logistics Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. US visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate; operations roles are especially difficult to support. But if we make you an offer, we will make every effort to get you into the United States, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Compensation and Benefits* Anthropic’s compensation package consists of three elements: salary, equity, and benefits. We are committed to pay fairness and aim for these three elements collectively to be highly competitive with market rates. Equity - For eligible roles, equity will be a major component of the total compensation. We aim to offer higher-than-average equity compensation for a company of our size, and communicate equity amounts at the time of offer issuance. US Benefits - The following benefits are for our US-based employees: Optional equity donation matching. Comprehensive health, dental, and vision insurance for you and all your dependents. 401(k) plan with 4% matching. 22 weeks of paid parental leave. Unlimited PTO – most staff take between 4-6 weeks each year, sometimes more! Stipends for education, home office improvements, commuting, and wellness. Fertility benefits via Carrot. Daily lunches and snacks in our office. Relocation support for those moving to the Bay Area. UK Benefits - The following benefits are for our UK-based employees: Optional equity donation matching. Private health, dental, and vision insurance for you and your dependents. Pension contribution (matching 4% of your salary). 21 weeks of paid parental leave. Unlimited PTO – most staff take between 4-6 weeks each year, sometimes more! Health cash plan. Life insurance and income protection. Daily lunches and snacks in our office. #J-18808-Ljbffr Anthropic

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Autonomy Safety Evaluations Research Engineer in New York, NY vacancy
  • $270k - $380k

     ...customers. Cohere is a team of researchers, engineers, designers, and more, who...  ...Research Engineer in our Safety team, you will play a key role...  ...in both model training and evaluation. You will own the cohesive...  .... You will have a lot of autonomy and need to be opinionated... 
    Suggested
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  •  ...Research Engineer (Senior Staff+), Safeguards LabsSan Francisco, CA | New York City, NYAbout AnthropicAnthropic...  ..., chartered to investigate novel safety methods that protect Claude and the...  ...classifiers and detection systems, and evaluate their effectiveness.Develop and iterate... 
    Suggested
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    2 days ago
  • $224k - $356.5k

     ...As part of this mission, our AI Safety & Security Engineering team builds and evaluates AI-powered tooling that finds, validates...  .... We are seeking a Security Research Engineer to advance our Validate...  ...a creative engineer who enjoys autonomy and shares our passion for... 
    Suggested
    Full time
    Immediate start
    Remote work

    Nvidia

    New York, NY
    3 days ago
  • $350k

     ...Research Engineer, Production Model Post-TrainingSan Francisco, CA | New York City, NY | Seattle...  ...enhance their capabilities, alignment, and safety. As a Research Engineer on our Post-...  ...with training, fine-tuning, or evaluating large language modelsCan balance research... 
    Suggested
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    2 days ago
  •  ...the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their...  ...Montreal, Seoul, Germany and Paris. Join us! Why this role? Evaluation is critical to making progress in scaling intelligence. As... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Home office

    cohere

    New York, NY
    5 days ago
  • $175k - $225k

     ...markets around the world. We value autonomy and the ability to quickly...  ...with traders and quantitative researchers to implement, refine and deploy alpha signals, evaluate and maintain trading tools,...  ...time.About the RoleAs a Research Engineer, you will be an integral... 
    Temporary work
    Flexible hours

    DRW

    New York, NY
    1 day ago
  •  ...ICML, and ACL—and see that research deployed to a user base of over...  ...AI Agents Applied Research/Engineering Lead in our The Digital Team...  ...planning, tool use, and safety; building production systems...  ...batching, prompt governance, and evaluation frameworks.Implement privacy... 

    JP Morgan Chase

    New York, NY
    6 hours ago
  •  ...Research EngineerAs a Research Engineer on our team, you will work on real production use cases of LLMs and other...  ...role, with a high level of autonomy and responsibility. You will be expected...  ...and techniques to adopt––including evaluating when to use open-source models, proprietary... 
    Full time
    Work at office

    Forus

    New York, NY
    2 days ago
  • $150k - $196k

    The Lead Research Engineer is a key technical contributor and small engineering team lead on multiple programs aligned with division and/or group objectives. Mission autonomy is the ability for autonomous agents or teams to conduct strategic decision-making and coordinated... 
    Temporary work
    Summer work

    Scientific Systems Company, Inc.

    New York, NY
    5 days ago
  •  ...Basis is a nonprofit applied AI research organization with two...  ...first. About the Role Research Engineers in Operations at Basis build...  ...Progress with a high degree of autonomy and under uncertainty. Have...  ...materials for both hiring evaluation and recruitment-related research... 
    Full time
    Contract work
    Work at office

    Basis Research Institute

    New York, NY
    1 day ago
  • $300k - $405k

     ...Research Engineer, Cybersecurity RL (Reinforcement Learning)San Francisco...  ...with significant impact on the autonomy, coding, and reasoning...  ...conducting experiments and evaluations, delivering your work into production...  ...in this work. Your safety matters to us. To protect yourself... 
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    2 days ago
  • $124k - $335k

     ...products and solutions.Those in software engineering at PwC will focus on developing...  ...solve complex problems. As you increase in autonomy, you apply sound judgment, recognising when...  ...collaborating closely with team members. We evaluate these factors thoughtfully to establish... 
    Full time
    H1b

    PwC

    New York, NY
    3 days ago
  •  ...NYAbout the OpportunityWe are seeking a disciplined and creative Research Engineer to join an elite systematic trading group. In this role, you...  ...deployment.Technological Advancement: Continuously evaluate emerging technologies to enhance the team’s existing infrastructure... 
    Immediate start

    Objective Paradigm

    New York, NY
    4 days ago
  • $165k - $260k

    Senior LLM Research Engineer - Artificial Intelligence Location New York Business Area Engineering and CTO Ref...  ...adaptation of LLMs to financial domains, dialogue interfaces, evaluation of LLMs, model safety and responsible AI.What's in it for you:Collaborate... 
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    1 day ago
  • $174k - $252k

    Apply research ideas to high-impact problems by prototyping, curating...  ...ambiguous problems.Train, evaluate, and iterate on deep neural models...  ...objectives. Influence engineering best practices by championing...  ...scientific discovery, ensuring safety and ethics are always our highest... 
    Full time
    Work at office

    Google

    New York, NY
    3 days ago
  • $197.3k - $313.7k

     ...looking for talented software and platform engineers to embed in our AI team to bridge the...  ...engineering skills directly enable world-class research and products used by millions?At...  ...production services, like APIs, UIs, agentic evaluators, and more.Roll out scalable cloud... 
    Full time

    Salesforce

    New York, NY
    6 hours ago
  • $110.7k - $379.2k

    Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare...  ...training team, you will design, train, evaluate, and align the models that reason about...  ...how each affects reasoning quality, safety, latency, cost, and reliability. •... 
    Local area
    Visa sponsorship

    Deloitte

    New York, NY
    6 hours ago
  • $200k - $300k

     ...collaborative coding environment empowers talented engineers to make significant contributions and...  ...seeking highly motivated and skilled Research Engineers who will work very closely...  ...our predictions into profitable trades.Evaluating sim vs. live differences to improve the... 
    Temporary work
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    1 day ago
  • $175k - $250k

     ...organizations.What you’ll doAs a Machine Learning Engineer - Applied Scientist you will play a...  .... You will manage all aspects of the research process including methodology selection,...  ...testing, prototyping, and performance evaluation. You will apply, adapt, and extend existing... 
    Work experience placement

    Point72

    New York, NY
    2 days ago
  •  ...AI EngineerAs an AI Engineer on our research team, you'll work on our hardest research and AI infra...  ...orchestration, prompt optimizations, evaluation framework, and AI infraEnable agent/workflow...  ...your domain. Drive initiatives with autonomy and accountability. Think deeply,... 
    Work at office

    Amperos Health, Inc

    New York, NY
    2 days ago
  • $350k

     ...ML/Research Engineer, SafeguardsSan Francisco, CA | New York City, NYAbout AnthropicAnthropic...  ...to automatically source representative evaluations to iterate onBuild systems to monitor...  ...across contextsEvaluate and improve the safety of agentic products—developing both... 
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    New York, NY
    2 days ago
  •  ...Research Engineer/Scientist (Reinforcement Learning)Percepta's mission is to transform critical institutions with applied AI. We care that...  ...simulation environments and data pipelines to training and evaluation frameworks.Conduct in-the-wild evaluations at scale that drive... 

    Percepta

    New York, NY
    2 days ago
  •  ...computers truly come alive.Responsibilities:Own evaluation pipelines — design, build, and automate...  ...models, not slide decks — partner with research and infra to prototype, train, and...  ...Qualifications:Expert-level PyTorch.Proven software engineer who loves ML; comfortable writing... 
    Full time
    Contract work
    Shift work

    SESAME

    New York, NY
    2 days ago
  • $216k - $270k

     ...when it matters most, combining rigorous evaluation with full-stack deployment so our...  ...AI in production, paired with applied ML research, design, and evaluation to ensure these...  ...RoleAs a Staff Machine Learning Research Engineer, you will operate across the full breadth... 
    Full time

    Scale AI

    New York, NY
    6 hours ago
  • $264.8k - $331k

     ...complex agents in enterprises around the world. The Enterprise ML Research Lab works on the front lines of this AI revolution. We are...  ...for the same role. This allows us to ensure a fair and thorough evaluation of all applicants.About Us:At Scale, our mission is to develop... 
    Full time

    Scale AI

    New York, NY
    1 day ago
  •  ...in markets around the world. We value autonomy and the ability to quickly pivot to capture...  ...to challenge consensus. As a  Research Engineer , you will be an integral member of a...  ...prototype to production deployment Evaluate new technology and improve our technology... 
    Immediate start

    DRW

    New York, NY
    more than 2 months ago
  • $315k

    As a Research Engineer or Research Scientist in Applied Finetuning, you will directly train the models we launch to the public via Claude....  ...implement new algorithms, run experiments on data mixes, design evaluations, and improve our production model training pipeline. This... 
    Work at office
    Home office
    Visa sponsorship
    Relocation package

    Anthropic

    New York, NY
    5 days ago
  • $165k - $260k

    A leading financial technology company in New York seeks a Senior NLP Engineer to work on innovative AI-driven solutions. This role involves designing, training, and evaluating NLP models, collaborating across teams, and publishing findings. Ideal candidates will have... 

    Bloomberg

    New York, NY
    4 days ago
  • Attendance Works is seeking a Director of Evaluation and Research to lead measurement of impact across programs and initiatives, guiding data-informed decision making for continuous improvement. The role collaborates with the CEO, VPs, and directors to shape the organization... 

    Attendance Works

    New York, NY
    5 days ago
  • $264.8k - $331k

     ...RoleAs a Senior/Staff Machine Learning Engineer (MLE) on the General Agents team, you’ll...  ...lifecycle—from model and system design to evaluation, deployment, and iteration—bridging...  ...in ambiguous problem spaces, balancing research-driven approaches with pragmatic product... 
    Full time

    Scale AI

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Autonomy Safety Evaluations Research Engineer. Be the first to apply!