Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, AI Evaluation

Full-time

Nuna

Chronic disease isn't managed in a clinic. It is managed at home, in relationships, in the everyday. What's on the dinner table, what gets said, and who notices when someone's struggling. For the 130 million Americans managing a chronic condition, the healthcare system has offered the same answer for decades: a 15-minute doctor's visit, a pamphlet, and a portal login they'll never use.

At Nuna, we are building an AI health coach that shows up like a person who actually has time: available at 3am, infinitely patient, and never behind a waiting room. We use motivational interviewing to help patients and their families see themselves clearly, design experiments that fit their real lives, and navigate a system that has not historically been on their side. We are building from the ground up around a simple belief: patients don't want to be healthy; they want their lives back.

We're not competing with other health apps. We're competing with the moment a person gives up on getting better. If that's a problem you want to work on, we'd like to talk.

Your team

We are a small, interdisciplinary team - engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health coach. Our products are only as good as the care and science behind them, and your piece is how we know the coach is safe and working. You'll own the evaluation system for the team building the coach: a data scientist partners with you on the science, clinicians and designers supply the ground truth, and the engineers shipping the agents depend on the signal you produce to decide what ships.

The role

This is a net-new, build-first role for someone who wants to own how we evaluate our AI agents end to end. You'll build the harnesses, datasets, judges, and release gates that tell us whether the coach is safe and good, and you'll own both that infrastructure and the evals that run on it. This is not a test-execution role - you write the code and own the system, rather than running tests someone else designed. You'll make the day-to-day calls on standards, methods, and trade-offs, often with incomplete information and the freedom to define the right answer yourself. Comfort with ambiguity is part of the job.

What you'll do

  • Build testing harnesses and evaluation infrastructure for our agentic products and our internal agentic tooling

  • Own our evals end to end - both the architecture and the content - with support from data science and clinical partners

  • Make every agentic deployment run through the testing apparatus before it ships, and own the release gates that keep unsafe or low-quality behavior from reaching patients

  • Build the ground truth, judges, and metrics, and validate that the evaluation itself can be trusted: calibration to human labels, reliability, and honest confidence on every number, in partnership with our data scientist

  • Build functional tooling for labeling and review workflows, so clinicians, coaches, and designers can author and review evaluation scenarios without an engineer in the loop

  • Help close the loop from evaluation results to model and prompt refinement, working toward systems that iterate safely with less human hand-holding

What we're looking for

  • Significant experience building and shipping reliable production systems and tooling

  • Deep understanding of how to evaluate AI systems - LLM-as-judge, red-teaming and adversarial testing, synthetic scenario generation, and multi-turn and agentic evaluation - and a clear sense of how evals themselves fail. You've deployed evals and automated AI tooling in production, not just prototyped them

  • A testing mindset applied to building the measurement system, not running tests against a spec: adversarial instinct, coverage thinking, regression discipline, and documentation others can build on

  • You use AI in your daily work and build tools that make the people around you more effective

  • Enough fluency in statistics and experimental design to partner with a data scientist on calibration and reliability

  • Can design the workflow and build a functional UI for non-engineers like clinicians and labelers

  • A genuine interest in improving healthcare alongside an interdisciplinary team, with the judgment to tell a launch-blocking issue from a nice-to-have

Bonus Points

  • Experience in healthcare or another regulated, high-trust domain, and familiarity with the regulatory landscape

  • Hands-on experience with the eval tooling ecosystem (LangSmith, Braintrust, DeepEval, Ragas, Promptfoo, or similar)

  • Red-teaming or AI safety experience - prompt injection, jailbreaks, adversarial and stress testing

  • Experience with automated, eval-driven model or prompt optimization

  • You've built in an early-stage or fast-moving environment

Nuna is an Equal Employment Opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, disability, genetics and/or veteran status.

#LI-LM1

Vacancy posted 8 days ago
Similar jobs that could be interesting for youBased on the Software Engineer, AI Evaluation in San Francisco, CA vacancy
  •  ...About the Team The Applied AI team works across research, engineering, product, and design to bring OpenAI’s technology to the world. We seek to learn...  ..., you’ll lead development of the systems we use to evaluate the quality of our AI models and products. You’ll help... 
    Suggested
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    16 hours ago
  • $175k - $215k

     ...dynamics, and state-of-the-art Generative AI to create a training ground for the Waymo Driver. The Simulator Evaluation team faces the ultimate data challenge: How...  ...virtual world is "real"? We are looking for a Software Engineer to build the metrics and pipelines that... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    16 hours ago
  • $170k - $216k

     ...powers the Waymo Driver. Our software allows the Waymo Driver to perceive...  ...sensors, enabling software engineers like you to develop multi-...  ...-critical automation and evaluation frameworks that establish the...  ...of experience in industrial AI applications involving the creation... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    16 hours ago
  •  ...mission is to organize human intelligence to power the AI economy. We partner with leading AI labs and enterprises...  ...or London offices. About the Role As a Senior Software Engineer (AI Data & Evaluation) at Mercor, you will be at the core of building the data... 
    Suggested
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    16 hours ago
  •  ...Software Engineer, Agent Evaluation and Quality Engineering · Full-time · San Francisco; New York Our mission is to automate coding. The first...  ...You'll Work On Designing and building best-in-class AI evaluation system: curated datasets, offline replay, scorers... 
    Suggested
    Full time
    Work at office

    Anysphere

    San Francisco, CA
    3 days ago
  • $238k - $302k

     ...+ U.S. states. The Large Model Evaluation team is at the nexus of Waymo’s AI ambition . With advancements in Large...  ...looking for quantitatively-minded engineers to research and propose new ways...  ...in a heavily quantitative software engineering area ~ Experience navigating... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    16 hours ago
  • $150k - $250k

    About Distyl AI Distyl is an applied AI technology company partnering...  ..., we build AI systems using Evaluation-Driven Development—an...  ...in production.AI Evaluation Engineers focus on designing and implementing...  ...We Require2+ years of software engineering experience Strong... 
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    4 days ago
  • $120k - $170k

     ...adventure?   Loft Orbital is looking for a Software Engineer to join our Ground Software Solutions...  ...this role is intentionally wide as we evaluate individuals based on their unique...  ...observation, IoT connectivity, on-orbit AI, national security missions, and more. Leveraging... 
    Full time
    Temporary work
    Work at office
    Relocation package
    Flexible hours

    Loft Orbital Solutions

    San Francisco, CA
    16 hours ago
  • $200k

     ...Member of Technical Staff to manage our internal evaluations platform, essential for improving AI model performance. You will be responsible for designing...  ...in measurements. The ideal candidate has strong software engineering fundamentals and experience with machine learning... 

    Dormont Manufacturing Co

    San Francisco, CA
    1 day ago
  • Kindredventures is seeking research engineers to build a central evaluation framework and scalable evaluation pipelines for models and datasets. You will design benchmarks, implement baselines, and create dashboards that translate results into actionable insights for the... 

    Kindredventures

    San Francisco, CA
    16 hours ago
  • $193.4k - $290k

     ...By combining frontier agentic AI, an enterprise-grade platform...  ...— from leadership to engineers — and work together to solve...  ...Role Overview As a Backend Software Engineer on the Product Engineering...  ...s most sensitive matters ~ Evaluating LLMs across a 10k+ leaf taxonomy... 
    Full time
    Relocation package

    Harvey

    San Francisco, CA
    16 hours ago
  •  ...are looking for a self-starter full stack engineer who can help us rapidly prototype and...  ...researchers, such as visualization for our evaluation of models. You should be comfortable being...  ...forward. About OpenAI OpenAI is an AI research and deployment company dedicated... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    16 hours ago
  •  ...About HappyRobot HappyRobot is the AI-native operating system for the real economy—...  ...The Role At HappyRobot , Full Stack Engineers own meaningful parts of the product and infrastructure...  ...of your personal data for the purpose of evaluating and selecting you as a candidate for the... 
    Full time
    Shift work

    HappyRobot

    San Francisco, CA
    16 hours ago
  • $202k - $237k

     ...computing and make it accessible to software developers of all skill...  ...accelerate the progress of AI applications out into the real...  ...are seeking a Backend Software Engineer to join our team focused on...  ...Opportunity Employer. Candidates are evaluated without regard to age, race,... 
    Full time
    Work at office
    Flexible hours

    Anyscale

    San Francisco, CA
    16 hours ago
  •  ...Minerals Mariana Minerals is a software-first, vertically integrated...  ...powering modern energy, AI, and defense technologies. We...  ...Senior Full Stack Software Engineer to lead critical technical initiatives...  ...Conduct thorough technical evaluations of third-party software... 
    Full time

    Mariana Minerals

    San Francisco, CA
    16 hours ago
  • HeyMilo AI is hiring a Research Engineer to join our Applied AI team in San Francisco. You’ll design and build reinforcement learning environments...  ...use cases, creating simulators, reward functions, and evaluation harnesses to measure model performance. The role blends... 

    HeyMilo AI

    San Francisco, CA
    4 days ago
  •  ...Francisco seeks a bio safety researcher to design and run capability evaluations for biology-focused models, build and curate datasets for safety classifiers, and iterate on those classifiers with ML engineers. You will work at the intersection of applied ML and biosecurity... 

    Anthropic

    San Francisco, CA
    3 days ago
  • $208k - $312k

     ...team behind Next.js, v0, and AI SDK, we create products that...  ...exceptional developer experience.Now, software is entering a new era, and...  ...for a Senior (IC4) software engineer with a strong security...  ...deployed full-time with v0, and is evaluated as much on shipped product... 
    Full time
    Work from home
    Worldwide
    Flexible hours

    Vercel

    San Francisco, CA
    2 days ago
  • $172.5k - $260.1k

     ...DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents drive...  ...ahead!What you will be doingAs a Senior Software Engineer on the Vulnerability Management team, you...  ...tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout... 
    Permanent employment
    Full time

    Salesforce

    San Francisco, CA
    3 days ago
  • $123.7k - $254.67k

     ...career you love? It’s Possible.At Pinterest, AI isn't just a feature, it's a powerful...  ...Pinterest is seeking an experienced Security Engineer to build and implement detection and...  ...outputs.Strong track record of critical evaluation and verification of AI-assisted work (e.g... 
    Work at office
    Local area
    Remote work
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    1 day ago
  •  ...critical inference for the world's most dynamic AI companies, like Cursor, Notion,...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE...  ...deployments within partner environments. Evaluate and implement emerging virtualization... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    16 hours ago
  •  ...Beacon Software is a permanent capital holding company which acquires...  ...below. Senior Software Engineer, Platform Own the architecture...  ...move faster. This is an AI-first orchestration role:...  ...Tune prompts, guardrails, and evaluation harnesses for accuracy, tone,... 
    Permanent employment
    Full time

    Beacon Software

    San Francisco, CA
    16 hours ago
  • $200k - $250k

     ...compute capabilities to the world’s biggest AI Labs at industry-defining speeds. Our...  ...cloud provider, is looking for a Software Engineer, Infrastructure Platform to build the foundational...  ...growth Technical Leadership Evaluate build vs. buy decisions for platform... 
    Full time
    Local area

    Fluidstack

    San Francisco, CA
    16 hours ago
  •  ...help businesses build better, more human customer experiences with AI. We are primarily an in-person company based in San Francisco,...  ...doesn't precisely match the job description. We strive to evaluate all applicants consistently without regard to race, color, religion... 
    Full time
    Flexible hours

    Sierra

    San Francisco, CA
    16 hours ago
  •  ...About JazzX AI:    Vision: Enterprises operating on institutional...  ..., AI-powered enterprise software companies. SAIGroup’s...  ...Overview As a Software Engineer, AI Platform, you will help design...  ...workflows, retrieval systems, evaluation pipelines, and production-... 
    Full time
    Worldwide

    Jazzx Ai

    San Francisco, CA
    16 hours ago
  •  ...computing and make it accessible to software developers of all skill...  ...accelerate the progress of AI applications out into the real...  ...Anyscale is looking for a Software Engineer to join the Infrastructure...  ...Employer. Candidates are evaluated without regard to age, race,... 
    Full time

    Anyscale

    San Francisco, CA
    16 hours ago
  •  ...As one of our core engineers, you’ll be critical in creating data infrastructure that will...  ...model labs, and we serve all of the frontier AI labs. We are based in San Francisco,...  ...pipelines. Any background in AI research, LLM evaluation, or human‑in‑the‑loop systems... 
    Full time

    AfterQuery

    San Francisco, CA
    16 hours ago
  • $185k - $325k

     ...end-to-end. By combining frontier agentic AI, an enterprise-grade platform, and deep...  ...close to our customers — from leadership to engineers — and work together to solve real...  ...capability every product team depends on. Evaluation Infrastructure. Build the shared eval tooling... 
    Full time

    Harvey

    San Francisco, CA
    16 hours ago
  • $170.11k - $237k

     ...computing and make it accessible to software developers of all skill...  ...accelerate the progress of AI applications out into the real...  ...is actively seeking talented engineers to join our team and contribute...  ...Employer. Candidates are evaluated without regard to age, race,... 
    Full time
    Work experience placement
    Work at office
    Flexible hours

    Anyscale

    San Francisco, CA
    16 hours ago
  • Skyrocket Ventures is looking for engineers who are proficient in Python and willing...  ...building and improving simulation and evaluation platforms for AI agents in San Francisco. Successful...  ...will have 5-10 years of experience in software engineering, particularly in backend... 

    Skyrocket Ventures

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, AI Evaluation. Be the first to apply!