Software Engineer, AI Evaluation
Nuna
Chronic disease isn't managed in a clinic. It is managed at home, in relationships, in the everyday. What's on the dinner table, what gets said, and who notices when someone's struggling. For the 130 million Americans managing a chronic condition, the healthcare system has offered the same answer for decades: a 15-minute doctor's visit, a pamphlet, and a portal login they'll never use.
At Nuna, we are building an AI health coach that shows up like a person who actually has time: available at 3am, infinitely patient, and never behind a waiting room. We use motivational interviewing to help patients and their families see themselves clearly, design experiments that fit their real lives, and navigate a system that has not historically been on their side. We are building from the ground up around a simple belief: patients don't want to be healthy; they want their lives back.
We're not competing with other health apps. We're competing with the moment a person gives up on getting better. If that's a problem you want to work on, we'd like to talk.
Your team
We are a small, interdisciplinary team - engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health coach. Our products are only as good as the care and science behind them, and your piece is how we know the coach is safe and working. You'll own the evaluation system for the team building the coach: a data scientist partners with you on the science, clinicians and designers supply the ground truth, and the engineers shipping the agents depend on the signal you produce to decide what ships.
The role
This is a net-new, build-first role for someone who wants to own how we evaluate our AI agents end to end. You'll build the harnesses, datasets, judges, and release gates that tell us whether the coach is safe and good, and you'll own both that infrastructure and the evals that run on it. This is not a test-execution role - you write the code and own the system, rather than running tests someone else designed. You'll make the day-to-day calls on standards, methods, and trade-offs, often with incomplete information and the freedom to define the right answer yourself. Comfort with ambiguity is part of the job.
What you'll do
Build testing harnesses and evaluation infrastructure for our agentic products and our internal agentic tooling
Own our evals end to end - both the architecture and the content - with support from data science and clinical partners
Make every agentic deployment run through the testing apparatus before it ships, and own the release gates that keep unsafe or low-quality behavior from reaching patients
Build the ground truth, judges, and metrics, and validate that the evaluation itself can be trusted: calibration to human labels, reliability, and honest confidence on every number, in partnership with our data scientist
Build functional tooling for labeling and review workflows, so clinicians, coaches, and designers can author and review evaluation scenarios without an engineer in the loop
Help close the loop from evaluation results to model and prompt refinement, working toward systems that iterate safely with less human hand-holding
What we're looking for
Significant experience building and shipping reliable production systems and tooling
Deep understanding of how to evaluate AI systems - LLM-as-judge, red-teaming and adversarial testing, synthetic scenario generation, and multi-turn and agentic evaluation - and a clear sense of how evals themselves fail. You've deployed evals and automated AI tooling in production, not just prototyped them
A testing mindset applied to building the measurement system, not running tests against a spec: adversarial instinct, coverage thinking, regression discipline, and documentation others can build on
You use AI in your daily work and build tools that make the people around you more effective
Enough fluency in statistics and experimental design to partner with a data scientist on calibration and reliability
Can design the workflow and build a functional UI for non-engineers like clinicians and labelers
A genuine interest in improving healthcare alongside an interdisciplinary team, with the judgment to tell a launch-blocking issue from a nice-to-have
Bonus Points
Experience in healthcare or another regulated, high-trust domain, and familiarity with the regulatory landscape
Hands-on experience with the eval tooling ecosystem (LangSmith, Braintrust, DeepEval, Ragas, Promptfoo, or similar)
Red-teaming or AI safety experience - prompt injection, jailbreaks, adversarial and stress testing
Experience with automated, eval-driven model or prompt optimization
You've built in an early-stage or fast-moving environment
Nuna is an Equal Employment Opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, disability, genetics and/or veteran status.
#LI-LM1
- ...About the Team The Applied AI team works across research, engineering, product, and design to bring OpenAI’s technology to the world. We seek to learn... ..., you’ll lead development of the systems we use to evaluate the quality of our AI models and products. You’ll help...SuggestedFull timeWork at officeRelocation package
$175k - $215k
...dynamics, and state-of-the-art Generative AI to create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge: How... ...virtual world is "real"? We are looking for a Software Engineer to build the metrics and pipelines that...SuggestedFull timeRemote work$170k - $216k
...powers the Waymo Driver. Our software allows the Waymo Driver to perceive... ...sensors, enabling software engineers like you to develop multi-... ...-critical automation and evaluation frameworks that establish the... ...of experience in industrial AI applications involving the creation...SuggestedFull timeRemote work- ...mission is to organize human intelligence to power the AI economy. We partner with leading AI labs and enterprises... ...or London offices. About the Role As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of building the data...SuggestedFull timeWork at officeRelocation package
- ...Software Engineer, Agent Evaluation and Quality Engineering · Full-time · San Francisco; New York Our mission is to automate coding. The first... ...You'll Work On Designing and building best-in-class AI evaluation system: curated datasets, offline replay, scorers...SuggestedFull timeWork at office
$238k - $302k
...+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large... ...looking for quantitatively-minded engineers to research and propose new ways... ...in a heavily quantitative software engineering area ~ Experience navigating...Full timeRemote work$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering... ..., we build AI systems using Evaluation-Driven Development—an... ...in production.AI Evaluation Engineers focus on designing and implementing... ...We Require2+ years of software engineering experience Strong...Work at office3 days per week$120k - $170k
...adventure? Loft Orbital is looking for a Software Engineer to join our Ground Software Solutions... ...this role is intentionally wide as we evaluate individuals based on their unique... ...observation, IoT connectivity, on-orbit AI, national security missions, and more. Leveraging...Full timeTemporary workWork at officeRelocation packageFlexible hours$200k
...Member of Technical Staff to manage our internal evaluations platform, essential for improving AI model performance. You will be responsible for designing... ...in measurements. The ideal candidate has strong software engineering fundamentals and experience with machine learning...- Kindredventures is seeking research engineers to build a central evaluation framework and scalable evaluation pipelines for models and datasets. You will design benchmarks, implement baselines, and create dashboards that translate results into actionable insights for the...
$193.4k - $290k
...By combining frontier agentic AI, an enterprise-grade platform... ...— from leadership to engineers — and work together to solve... ...Role Overview As a Backend Software Engineer on the Product Engineering... ...s most sensitive matters ~ Evaluating LLMs across a 10k+ leaf taxonomy...Full timeRelocation package- ...are looking for a self-starter full stack engineer who can help us rapidly prototype and... ...researchers, such as visualization for our evaluation of models. You should be comfortable being... ...forward. About OpenAI OpenAI is an AI research and deployment company dedicated...Full timeWork at officeRelocation package
- ...About HappyRobot HappyRobot is the AI-native operating system for the real economy—... ...The Role At HappyRobot , Full Stack Engineers own meaningful parts of the product and infrastructure... ...of your personal data for the purpose of evaluating and selecting you as a candidate for the...Full timeShift work
$202k - $237k
...computing and make it accessible to software developers of all skill... ...accelerate the progress of AI applications out into the real... ...are seeking a Backend Software Engineer to join our team focused on... ...Opportunity Employer. Candidates are evaluated without regard to age, race,...Full timeWork at officeFlexible hours- ...Minerals Mariana Minerals is a software-first, vertically integrated... ...powering modern energy, AI, and defense technologies. We... ...Senior Full Stack Software Engineer to lead critical technical initiatives... ...Conduct thorough technical evaluations of third-party software...Full time
- HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning environments... ...use cases, creating simulators, reward functions, and evaluation harnesses to measure model performance. The role blends...
- ...Francisco seeks a bio safety researcher to design and run capability evaluations for biology-focused models, build and curate datasets for safety classifiers, and iterate on those classifiers with ML engineers. You will work at the intersection of applied ML and biosecurity...
$208k - $312k
...team behind Next.js, v0, and AI SDK, we create products that... ...exceptional developer experience.Now, software is entering a new era, and... ...for a Senior (IC4) software engineer with a strong security... ...deployed full-time with v0, and is evaluated as much on shipped product...Full timeWork from homeWorldwideFlexible hours$172.5k - $260.1k
...DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents drive... ...ahead!What you will be doingAs a Senior Software Engineer on the Vulnerability Management team, you... ...tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout...Permanent employmentFull time$123.7k - $254.67k
...career you love? It’s Possible.At Pinterest, AI isn't just a feature, it's a powerful... ...Pinterest is seeking an experienced Security Engineer to build and implement detection and... ...outputs.Strong track record of critical evaluation and verification of AI-assisted work (e.g...Work at officeLocal areaRemote workRelocationRelocation package- ...critical inference for the world's most dynamic AI companies, like Cursor, Notion,... ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE... ...deployments within partner environments. Evaluate and implement emerging virtualization...Full timeFlexible hours
- ...Beacon Software is a permanent capital holding company which acquires... ...below. Senior Software Engineer, Platform Own the architecture... ...move faster. This is an AI-first orchestration role:... ...Tune prompts, guardrails, and evaluation harnesses for accuracy, tone,...Permanent employmentFull time
$200k - $250k
...compute capabilities to the world’s biggest AI Labs at industry-defining speeds. Our... ...cloud provider, is looking for a Software Engineer, Infrastructure Platform to build the foundational... ...growth Technical Leadership Evaluate build vs. buy decisions for platform...Full timeLocal area- ...help businesses build better, more human customer experiences with AI. We are primarily an in-person company based in San Francisco,... ...doesn't precisely match the job description. We strive to evaluate all applicants consistently without regard to race, color, religion...Full timeFlexible hours
- ...About JazzX AI: Vision: Enterprises operating on institutional... ..., AI-powered enterprise software companies. SAIGroup’s... ...Overview As a Software Engineer, AI Platform, you will help design... ...workflows, retrieval systems, evaluation pipelines, and production-...Full timeWorldwide
- ...computing and make it accessible to software developers of all skill... ...accelerate the progress of AI applications out into the real... ...Anyscale is looking for a Software Engineer to join the Infrastructure... ...Employer. Candidates are evaluated without regard to age, race,...Full time
- ...As one of our core engineers, you’ll be critical in creating data infrastructure that will... ...model labs, and we serve all of the frontier AI labs. We are based in San Francisco,... ...pipelines. Any background in AI research, LLM evaluation, or human‑in‑the‑loop systems...Full time
$185k - $325k
...end-to-end. By combining frontier agentic AI, an enterprise-grade platform, and deep... ...close to our customers — from leadership to engineers — and work together to solve real... ...capability every product team depends on. Evaluation Infrastructure. Build the shared eval tooling...Full time$170.11k - $237k
...computing and make it accessible to software developers of all skill... ...accelerate the progress of AI applications out into the real... ...is actively seeking talented engineers to join our team and contribute... ...Employer. Candidates are evaluated without regard to age, race,...Full timeWork experience placementWork at officeFlexible hours- Skyrocket Ventures is looking for engineers who are proficient in Python and willing... ...building and improving simulation and evaluation platforms for AI agents in San Francisco. Successful... ...will have 5-10 years of experience in software engineering, particularly in backend...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, AI Evaluation. Be the first to apply!
- senior robotics software engineer San Francisco, CA
- software system engineer San Francisco, CA
- part time software developer San Francisco, CA
- fall software engineering internship San Francisco, CA
- security software engineer San Francisco, CA
- intel software engineer San Francisco, CA
- software developer fintech San Francisco, CA
- new graduate software engineer San Francisco, CA
- software development engineer aws San Francisco, CA
- information technology software engineer San Francisco, CA


