Staff Research Engineer, Agent Evals & Post-training
Jobleads-US
Staff Research Engineer, Agent Evals & Post-training
Locations: New York, NY • Los Angeles, CA • London • Austin, TX • San Francisco, CA • Seattle, WA
Industry: AI
Work environment: In office
Contract: Full-time
About Nominal
Our mission is to accelerate how the world engineers new hardware. Nominal's connected test and operations platform powers the world's most advanced hardware programs and its most ambitious startups, from spacecraft, racecars, and autonomous vehicles to next-generation defense and energy programs. Our customers include Anduril, Shield AI, Hermeus, Albedo, Shinkei, and Pratt Miller Motorsports, as well as U.S. Navy and U.S. Air Force programs. Now we're expanding across the entire hardware lifecycle, building the foundation, AI-native applications, and agents that accelerate innovators' work and change what's possible to build.
About Hardware Intelligence
We're the team behind Nominal's agents, AI-native applications, and MCP, and its forward-leaning AI bets. Our mission is to unlock the bottlenecks of the hardware lifecycle with AI. Our agents reason over physical reality, from high-rate telemetry and test campaigns to designs and simulations, where real test results are the ground truth their work is checked against. We believe opinionated AI, built for the real work of hardware programs, will change how the world engineers.
The Role:
As a Staff Research Engineer supporting agent evals & post-training, you'll define how Nominal measures its agents, for customers and for ourselves, and build the path from evals to post-trained models when hardware needs them. Evals come first; post-training follows when air-gapped deployment or cost makes it the right investment.
What You'll Do:
- Build eval suites for our agents, the MCP tool layer, and our internal company agent, grounded in real hardware tasks.
- Invent new benchmarks for agents working over physical engineering data, where none exist today.
- Benchmark our agents against frontier agents using our MCP, and show where and why ours win.
- Build the eval infrastructure: datasets, model-based and human graders, and regression gates in CI.
- Turn evals into reward signals and training data, and lead post-training (fine-tuning, RL, distillation) when we need models we can run anywhere, including air-gapped environments.
- Set Nominal's strategy for measurement and model improvement, working with hardware experts on what "correct" means.
What you'll bring:
- 8+ years in ML engineering or research, including evals or post-training work you led in production.
- Statistical rigor: you design evals that don't fool you, and you know when a difference is real.
- Deep experience evaluating LLM or agent systems, including model-graded evals and their limits.
- Hands‑on post‑training experience: fine‑tuning, RL from feedback, or distillation on real tasks.
- The judgment to know when to measure, when to train, and when a better prompt or tool is the answer.
- A track record of setting technical direction across a team and raising the bar for the engineers around you.
- You build with modern AI coding agents (Claude Code, Cursor, Codex) every day, and stay curious and open to better ways of working. The tools keep changing, and so do we.
Nice to have:
- You've built evals or post-training at a frontier lab or an AI-native company.
- You've published benchmarks or eval methods that others use.
- You've trained or served open-weight models in restricted, on‑prem, or air‑gapped environments.
- You've worked in test, reliability, or verification engineering for physical systems.
Benefits/Perks
- 100% coverage of medical, dental, and vision insurance
- Unlimited PTO and sick leave
- Free lunch, snacks, and coffee
- Professional Development Stipend
- Annual company retreat
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, or national origin.
ITAR Requirements
To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here.
#J-18808-Ljbffr Jobleads-US- ...Nominal seeks a Staff Research Engineer to shape how we measure agent evals and lead post-training efforts. You will build eval suites grounded in hardware tasks and drive the data-to-model path, including RL, distillation, and model fine-tuning, across air-gapped environments...Training
$264.8k - $331k
...capabilities.About the General Agents TeamThe General... ...the RoleAs a Senior/Staff Machine Learning Engineer (MLE) on the General... ...spaces, balancing research-driven approaches... ...on each job posting reflects the minimum... ...relevant education or training. Scale employees in...TrainingFull time- ...growing group of committed researchers, engineers, policy experts, and business... ...the role Large teams of agents are starting to solve problems... ...combination of education, training, and/or experience Required... ...: Currently, we expect all staff to be in one of our offices...TrainingVisa sponsorship
$200k - $300k
...role You’ll own the agent and the experience.... ...This is a hands-on engineering role with a lot of product... ...code, the prompts, the evals, the integrations, and... ...This is a senior to staff level role. We’re looking... ...make ship decisions. Post-training or fine-tuning experience...TrainingRemote jobFull time$200k - $250k
...building consumer-grade agents that run privately on... ...Without a serious evals function, releases are... ...team knows whether each training run or prompt tweak helped... ...tools that speed up research iterations and make... ...care about), product engineers (tracking how real users...TrainingFull timeRelocation package$265k - $331k
...build production AI agents that automate complex... ...architectures, and research papers emerge every... ...one of the hardest engineering challenges.As a Staff Frontier Agent Engineer... ...on each job posting reflects the minimum... ...relevant education or training. Scale employees in...TrainingFull time- ...Primarily On-site We are seeking an Research Engineer – RL Infrastructure & Agent Environments to build the... ...supporting infrastructure used to train and assess long-horizon enterprise... ...behind realistic agent environments, post-training systems, and reliable evaluation...Training
$155k - $269k
...: You will… Be part of a multidisciplinary team of Research Scientists and Engineers building the content backbone of a best-in-class multi-sensor... ...meets the diversity, realism, and coverage demands of training and closed-loop evaluation. Qualifications: ~...TrainingFull timeWork at officeWork from homeFlexible hours- ...Model is building automated ML research engineering. Existing frontier models... ...the lack of high-quality RL training environments. Our first step... ...to push the frontier of post-training on large language models... ..., and methodologies for RL agents. Profile and optimize...TrainingVisa sponsorshipRelocation package
$220k - $290k
Senior / Staff Machine Learning Research Engineer Who We Are:Calico (Calico Life Sciences LLC) is an Alphabet-founded research and development company... ...acceleratorsDive deep into model architectures to optimize training and inference performance, implementing advanced...Training$249.12k - $367.92k
...enables innovation across Plaid.As a Staff Machine Learning Engineer, you will lead the technical... ...repeatable pipelines that translate research into production impact. You will also... ...objectives, architecture design, distributed training, serving infrastructure, monitoring,...TrainingWork experience placementWork at officeLocal areaImmediate start$150k - $250k
...global social organizations.We research and deploy technologies that... ...ForAt Distyl, Research Engineers build the bridge between frontier... ...and run post-training workflows that improve the behavior... ...reward modeling, synthetic data, evals, or related post-training techniquesStrong...TrainingWork at office3 days per week- ...teams: a security-focused product where an agent registers as a governed identity in a... ...what ships as code. You'll build and run evals against replicas of real customer environments... ...get better real-task data into eval and training Company stage: Pre-seed. Work type: In...TrainingFull timeContract workH1bVisa sponsorship
$200k - $300k
...Job Description Senior/Staff Applied AI Engineer, Agent Harness Company: Init Intelligence... ...You will build and run evals against replicas of real... ...data into evaluation and training What You Bring ~4+... ...environments, and scaling them AI research published at top...TrainingFull timeContract workH1bWork at officeVisa sponsorshipMonday to Friday- ...The Role As a Research Engineer, you'll build the systems that let us post-train models continuously: the training pipelines,... ...and the RL stack that improves agents over time. When a training run... ...versioning training environments and evals. Develop the tooling that...Training
$175k - $250k
...Research Engineer About Scorecard We’re a small, nimble team backed... ...on creating simulations for agents interacting with multiple users... ...memory. Design and run evals . We need to validate that... ...build; your simusers run inside training environments where their...TrainingWork at office- ...timescales. We are making agents that do that... ...What you'll do The research that matters most to us... ...end, building agents, training models, and designing the... ...benchmarks, environments, and evals that tell us whether it... ...about current PhDs and post-docs in continual learning...Training
$231k - $340k
...getting started. Role Overview Post-training is how Harvey turns expert feedback and agent traces into models that are... ...work. We are looking for a research engineer who can help scale that loop: defining... ...findings into training data, evals, or harness changes. Work...Training$205k - $275k
...Xaira Therapeutics seeks an experienced AI research engineer to develop and deploy machine learning models for drug discovery and protein... ...with modern deep learning toolchains and large-scale model training Strong understanding of deep learning fundamentals and complex...Training- ...Hillclimb is a research-focused startup aiming to improve agent capabilities toward recursive self-improvement. Our team designs environments and evals that push frontier models to ideate, experiment... ...validating environments before training, with a collaborative frontier/...Training
- ...encouraged to apply. Mission Design, train, ship, iterate on, and innovate on... ...The Path’s AI Therapist. Combine research, data science, and engineering to create models, orchestration, and... ...generic benchmarks. Can look at evals, transcripts, and metrics and quickly...Training
$147k - $210k
...and energy systems modeling and optimizationAt Google, research-focused Software Engineers are embedded throughout the company, allowing them to setup... ...-related skills, experience, and relevant education or training. US: $147000 - $210000 (USD) + 15% bonus target + equity...Training$197.3k - $313.7k
...CRM, where humans with agents drive customer success... ...software and platform engineers to embed in our AI... ...directly enable world-class research and products used by... ...and opt out options.Posting StatementSalesforce is... ...promotion, benefits, training, assessment of job performance...TrainingFull time$193.3k - $261.5k
...this team, you will lead research and development... ...to enable intelligent agents that perform complex reasoning... ...collaboration with engineering teams to bring... ...transformer architecture, training/inference lifecycles,... ...employees, supervisors, and staff; adhere to standards...TrainingInternshipLocal areaFlexible hours$228.6k - $342.8k
...The MissionDatabricks agents are only as good as... ...are hiring a Senior Staff Applied AI Engineer to own context retrieval... .... Stand up offline evals (nDCG, MRR, ****@*****.***,... ...engineers, partner with Research, Product, and... ...similar venues.Experience training or fine-tuning embedding...TrainingWork at officeLocal areaImmediate startWorldwide- ...teams: a security-focused product where an agent registers as a governed identity in a... ...is a backend- and systems-heavy founding engineering role where \"full-stack\" means owning the... ...and deterministic execution (validation, evals, release gates); build isolated virtual environments...Full timeH1bWork at officeVisa sponsorshipMonday to Friday
- ...Requirements Strong general software engineering skills Thorough knowledge of the deep learning literature Experience with pre- and post-training of LLMs Ability to come up with and evaluate research ideas Experience working with large distributed systems...Training
- ...conversation for every patient. An AI agent that knows who you are,... ...About the Role As a Research Engineer, you’ll be responsible for building... ...milestones and roadmaps Train, fine-tune, validate, and... ...~ Prior experience post-training and deploying LLMs in...TrainingWork at officeFlexible hours
- ...United States Digital Space LLC is seeking a generalist researcher-engineer to design and run large experiments on agent teams and to build scalable systems the teams rely on. You'll work at the research-engineering boundary, turning vague questions into runnable experiments...
- ...Proximal is building the research systems needed to identify what... ...early team has built coding agents and RL infrastructure at companies... ...the role As a research engineer, you will work on open-ended... ...correctness Modify our RL training stack to support end-to-end training...Training
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Research Engineer, Agent Evals & Post-training. Be the first to apply!
- assistant engineer San Francisco, CA
- senior staff systems engineer San Francisco, CA
- technology administrator San Francisco, CA
- engineering aide San Francisco, CA
- senior staff engineer San Francisco, CA
- staff design engineer San Francisco, CA
- software engineer staff San Francisco, CA
- staff engineer San Francisco, CA
- staff security engineer San Francisco, CA
- deep learning research engineer San Francisco, CA




