Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Research Engineer, Agent Evals & Post-training

Jobleads-US

Staff Research Engineer, Agent Evals & Post-training

Locations: New York, NY • Los Angeles, CA • London • Austin, TX • San Francisco, CA • Seattle, WA
Industry: AI
Work environment: In office
Contract: Full-time

About Nominal

Our mission is to accelerate how the world engineers new hardware. Nominal's connected test and operations platform powers the world's most advanced hardware programs and its most ambitious startups, from spacecraft, racecars, and autonomous vehicles to next-generation defense and energy programs. Our customers include Anduril, Shield AI, Hermeus, Albedo, Shinkei, and Pratt Miller Motorsports, as well as U.S. Navy and U.S. Air Force programs. Now we're expanding across the entire hardware lifecycle, building the foundation, AI-native applications, and agents that accelerate innovators' work and change what's possible to build.

About Hardware Intelligence

We're the team behind Nominal's agents, AI-native applications, and MCP, and its forward-leaning AI bets. Our mission is to unlock the bottlenecks of the hardware lifecycle with AI. Our agents reason over physical reality, from high-rate telemetry and test campaigns to designs and simulations, where real test results are the ground truth their work is checked against. We believe opinionated AI, built for the real work of hardware programs, will change how the world engineers.

The Role:

As a Staff Research Engineer supporting agent evals & post-training, you'll define how Nominal measures its agents, for customers and for ourselves, and build the path from evals to post-trained models when hardware needs them. Evals come first; post-training follows when air-gapped deployment or cost makes it the right investment.

What You'll Do:

  • Build eval suites for our agents, the MCP tool layer, and our internal company agent, grounded in real hardware tasks.
  • Invent new benchmarks for agents working over physical engineering data, where none exist today.
  • Benchmark our agents against frontier agents using our MCP, and show where and why ours win.
  • Build the eval infrastructure: datasets, model-based and human graders, and regression gates in CI.
  • Turn evals into reward signals and training data, and lead post-training (fine-tuning, RL, distillation) when we need models we can run anywhere, including air-gapped environments.
  • Set Nominal's strategy for measurement and model improvement, working with hardware experts on what "correct" means.

What you'll bring:

  • 8+ years in ML engineering or research, including evals or post-training work you led in production.
  • Statistical rigor: you design evals that don't fool you, and you know when a difference is real.
  • Deep experience evaluating LLM or agent systems, including model-graded evals and their limits.
  • Hands‑on post‑training experience: fine‑tuning, RL from feedback, or distillation on real tasks.
  • The judgment to know when to measure, when to train, and when a better prompt or tool is the answer.
  • A track record of setting technical direction across a team and raising the bar for the engineers around you.
  • You build with modern AI coding agents (Claude Code, Cursor, Codex) every day, and stay curious and open to better ways of working. The tools keep changing, and so do we.

Nice to have:

  • You've built evals or post-training at a frontier lab or an AI-native company.
  • You've published benchmarks or eval methods that others use.
  • You've trained or served open-weight models in restricted, on‑prem, or air‑gapped environments.
  • You've worked in test, reliability, or verification engineering for physical systems.

Benefits/Perks

  • 100% coverage of medical, dental, and vision insurance
  • Unlimited PTO and sick leave
  • Free lunch, snacks, and coffee
  • Professional Development Stipend
  • Annual company retreat

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, or national origin.

ITAR Requirements

To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff Research Engineer, Agent Evals & Post-training in San Francisco, CA vacancy
  •  ...Nominal seeks a Staff Research Engineer to shape how we measure agent evals and lead post-training efforts. You will build eval suites grounded in hardware tasks and drive the data-to-model path, including RL, distillation, and model fine-tuning, across air-gapped environments... 
    Training

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $264.8k - $331k

     ...capabilities.About the General Agents TeamThe General...  ...the RoleAs a Senior/Staff Machine Learning Engineer (MLE) on the General...  ...spaces, balancing research-driven approaches...  ...on each job posting reflects the minimum...  ...relevant education or training. Scale employees in... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  •  ...growing group of committed researchers, engineers, policy experts, and business...  ...the role Large teams of agents are starting to solve problems...  ...combination of education, training, and/or experience Required...  ...: Currently, we expect all staff to be in one of our offices... 
    Training
    Visa sponsorship

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $200k - $300k

     ...role You’ll own the agent and the experience....  ...This is a hands-on engineering role with a lot of product...  ...code, the prompts, the evals, the integrations, and...  ...This is a senior to staff level role. We’re looking...  ...make ship decisions. Post-training or fine-tuning experience... 
    Training
    Remote job
    Full time

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $200k - $250k

     ...building consumer-grade agents that run privately on...  ...Without a serious evals function, releases are...  ...team knows whether each training run or prompt tweak helped...  ...tools that speed up research iterations and make...  ...care about), product engineers (tracking how real users... 
    Training
    Full time
    Relocation package

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $265k - $331k

     ...build production AI agents that automate complex...  ...architectures, and research papers emerge every...  ...one of the hardest engineering challenges.As a Staff Frontier Agent Engineer...  ...on each job posting reflects the minimum...  ...relevant education or training. Scale employees in... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    8 hours ago
  •  ...Primarily On-site We are seeking an Research Engineer – RL Infrastructure & Agent Environments to build the...  ...supporting infrastructure used to train and assess long-horizon enterprise...  ...behind realistic agent environments, post-training systems, and reliable evaluation... 
    Training

    MaxIT Consulting - Max Corporate Group

    San Francisco, CA
    22 days ago
  • $155k - $269k

     ...: You will… Be part of a multidisciplinary team of Research Scientists and Engineers building the content backbone of a best-in-class multi-sensor...  ...meets the diversity, realism, and coverage demands of training and closed-loop evaluation. Qualifications: ~... 
    Training
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    11 days ago
  •  ...Model is building automated ML research engineering. Existing frontier models...  ...the lack of high-quality RL training environments. Our first step...  ...to push the frontier of post-training on large language models...  ..., and methodologies for RL agents. Profile and optimize... 
    Training
    Visa sponsorship
    Relocation package

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $220k - $290k

    Senior / Staff Machine Learning Research Engineer Who We Are:Calico (Calico Life Sciences LLC) is an Alphabet-founded research and development company...  ...acceleratorsDive deep into model architectures to optimize training and inference performance, implementing advanced... 
    Training

    Calico Life Sciences

    South San Francisco, CA
    3 days ago
  • $249.12k - $367.92k

     ...enables innovation across Plaid.As a Staff Machine Learning Engineer, you will lead the technical...  ...repeatable pipelines that translate research into production impact. You will also...  ...objectives, architecture design, distributed training, serving infrastructure, monitoring,... 
    Training
    Work experience placement
    Work at office
    Local area
    Immediate start

    Plaid Financial

    San Francisco, CA
    17 hours ago
  • $150k - $250k

     ...global social organizations.We research and deploy technologies that...  ...ForAt Distyl, Research Engineers build the bridge between frontier...  ...and run post-training workflows that improve the behavior...  ...reward modeling, synthetic data, evals, or related post-training techniquesStrong... 
    Training
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    1 day ago
  •  ...teams: a security-focused product where an agent registers as a governed identity in a...  ...what ships as code. You'll build and run evals against replicas of real customer environments...  ...get better real-task data into eval and training Company stage: Pre-seed. Work type: In... 
    Training
    Full time
    Contract work
    H1b
    Visa sponsorship

    Thomas Talent Network

    San Francisco, CA
    22 days ago
  • $200k - $300k

     ...Job Description Senior/Staff Applied AI Engineer, Agent Harness Company: Init Intelligence...  ...You will build and run evals against replicas of real...  ...data into evaluation and training What You Bring ~4+...  ...environments, and scaling them AI research published at top... 
    Training
    Full time
    Contract work
    H1b
    Work at office
    Visa sponsorship
    Monday to Friday

    Transparent Search Group

    San Francisco, CA
    17 days ago
  •  ...The Role As a Research Engineer, you'll build the systems that let us post-train models continuously: the training pipelines,...  ...and the RL stack that improves agents over time. When a training run...  ...versioning training environments and evals. Develop the tooling that... 
    Training

    ThirdLayer, Inc.

    San Francisco, CA
    4 days ago
  • $175k - $250k

     ...Research Engineer About Scorecard We’re a small, nimble team backed...  ...on creating simulations for agents interacting with multiple users...  ...memory. Design and run evals . We need to validate that...  ...build; your simusers run inside training environments where their... 
    Training
    Work at office

    Kindredventures

    San Francisco, CA
    2 days ago
  •  ...timescales. We are making agents that do that...  ...What you'll do The research that matters most to us...  ...end, building agents, training models, and designing the...  ...benchmarks, environments, and evals that tell us whether it...  ...about current PhDs and post-docs in continual learning... 
    Training

    Orin Labs

    San Francisco, CA
    2 days ago
  • $231k - $340k

     ...getting started. Role Overview Post-training is how Harvey turns expert feedback and agent traces into models that are...  ...work. We are looking for a research engineer who can help scale that loop: defining...  ...findings into training data, evals, or harness changes. Work... 
    Training

    Jobleads-US

    San Francisco, CA
    2 days ago
  • $205k - $275k

     ...Xaira Therapeutics seeks an experienced AI research engineer to develop and deploy machine learning models for drug discovery and protein...  ...with modern deep learning toolchains and large-scale model training Strong understanding of deep learning fundamentals and complex... 
    Training

    Xaira Therapeutics

    South San Francisco, CA
    3 days ago
  •  ...Hillclimb is a research-focused startup aiming to improve agent capabilities toward recursive self-improvement. Our team designs environments and evals that push frontier models to ideate, experiment...  ...validating environments before training, with a collaborative frontier/... 
    Training

    Jobleads-US

    San Francisco, CA
    1 day ago
  •  ...encouraged to apply. Mission Design, train, ship, iterate on, and innovate on...  ...The Path’s AI Therapist. Combine research, data science, and engineering to create models, orchestration, and...  ...generic benchmarks. Can look at evals, transcripts, and metrics and quickly... 
    Training

    The Path

    San Francisco, CA
    4 days ago
  • $147k - $210k

     ...and energy systems modeling and optimizationAt Google, research-focused Software Engineers are embedded throughout the company, allowing them to setup...  ...-related skills, experience, and relevant education or training. US: $147000 - $210000 (USD) + 15% bonus target + equity... 
    Training

    Google

    San Francisco, CA
    2 days ago
  • $197.3k - $313.7k

     ...CRM, where humans with agents drive customer success...  ...software and platform engineers to embed in our AI...  ...directly enable world-class research and products used by...  ...and opt out options.Posting StatementSalesforce is...  ...promotion, benefits, training, assessment of job performance... 
    Training
    Full time

    Salesforce

    San Francisco, CA
    17 hours ago
  • $193.3k - $261.5k

     ...this team, you will lead research and development...  ...to enable intelligent agents that perform complex reasoning...  ...collaboration with engineering teams to bring...  ...transformer architecture, training/inference lifecycles,...  ...employees, supervisors, and staff; adhere to standards... 
    Training
    Internship
    Local area
    Flexible hours

    AmazonWebServices

    San Francisco, CA
    4 days ago
  • $228.6k - $342.8k

     ...The MissionDatabricks agents are only as good as...  ...are hiring a Senior Staff Applied AI Engineer to own context retrieval...  .... Stand up offline evals (nDCG, MRR, ****@*****.***,...  ...engineers, partner with Research, Product, and...  ...similar venues.Experience training or fine-tuning embedding... 
    Training
    Work at office
    Local area
    Immediate start
    Worldwide

    DataBricks

    San Francisco, CA
    4 days ago
  •  ...teams: a security-focused product where an agent registers as a governed identity in a...  ...is a backend- and systems-heavy founding engineering role where \"full-stack\" means owning the...  ...and deterministic execution (validation, evals, release gates); build isolated virtual environments... 
    Full time
    H1b
    Work at office
    Visa sponsorship
    Monday to Friday

    Jobleads-US

    San Francisco, CA
    5 days ago
  •  ...Requirements Strong general software engineering skills Thorough knowledge of the deep learning literature Experience with pre- and post-training of LLMs Ability to come up with and evaluate research ideas Experience working with large distributed systems... 
    Training

    Magic Inc

    San Francisco, CA
    2 days ago
  •  ...conversation for every patient. An AI agent that knows who you are,...  ...About the Role As a Research Engineer, you’ll be responsible for building...  ...milestones and roadmaps Train, fine-tune, validate, and...  ...~ Prior experience post-training and deploying LLMs in... 
    Training
    Work at office
    Flexible hours

    Assort Health

    San Francisco, CA
    2 days ago
  •  ...United States Digital Space LLC is seeking a generalist researcher-engineer to design and run large experiments on agent teams and to build scalable systems the teams rely on. You'll work at the research-engineering boundary, turning vague questions into runnable experiments... 

    Jobleads-US

    San Francisco, CA
    4 days ago
  •  ...Proximal is building the research systems needed to identify what...  ...early team has built coding agents and RL infrastructure at companies...  ...the role As a research engineer, you will work on open-ended...  ...correctness Modify our RL training stack to support end-to-end training... 
    Training

    Proximal LLC

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Research Engineer, Agent Evals & Post-training. Be the first to apply!