Lead, Build Agent Evaluation & Model Benchmarking
ServiceNow
ServiceNow in Santa Clara, CA is seeking a leader for the Build Agent evaluation framework. You will own eval strategy, roadmap, telemetry, and cross-team quality standards, guiding an 8‑engineer team to deliver scalable, data‑driven improvements.
You will benchmark models, decide on model support, and communicate results to stakeholders. Strong experience with evaluation systems for LLMs and cross‑team accountability is required.
#J-18808-Ljbffr- NVIDIA Corporation in Santa Clara is seeking a Senior Research Manager to lead world-model evaluation and benchmarking across the NVIDIA Physical AI portfolio. This role will build the team and research agenda for evaluating world models through closed-system evaluations...Suggested
- NVIDIA is seeking a Senior Research Manager to lead world-model evaluation and benchmarking for Physical AI. This position involves developing evaluation methods and driving model improvement through rigorous scientific standards. The ideal candidate will have a PhD in...Suggested
- NVIDIA Gruppe is seeking a Senior Research Manager to lead world-model evaluation and benchmarking efforts in Santa Clara, California. The successful candidate... ...that defines useful world models for Physical AI and build a team of research scientists focused on innovative...Suggested
- ...that offer MaaS solutions to businesses and researchers, spanning model serving, tool calling, and analytics. This full-time role emphasizes designing evaluation systems for LLM agents, building benchmarks and pipelines, analyzing traces and feedback, and collaborating...SuggestedFull time
- ...Apple is seeking a Senior SDET to lead automated model evaluation across the Intelligence Platform. You will own eval automation, build LLM-as-a-judge in pipelines, and catch regressions before they reach live use. You will design, implement, and maintain evaluation...Suggested
$272k - $431.25k
At NVIDIA, we’re not just building the future, we’re generating it! Our world model team is pushing the boundaries of multimodal AI, robotics... ...looking for a Senior Research Manager to lead world-model evaluation and benchmarking across NVIDIA’s Physical AI model...Full time$224k - $356.5k
...Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play... ...you'll be doing:Define and build evaluation methodologies for... ...LLMs, RAG systems, agents, and vision/multimodal models... ...improving evaluation frameworks, benchmarks, or ML infrastructure used...Full time- NVIDIA seeks a highly analytical Product Evaluations Lead to own evaluation, measurement, and go-forward... ...of GenAI software including Nemotron models for strategic enterprise ISV partners. You will drive experiments, benchmarks, and translate results into insights and prioritized...
$152k - $241.5k
...Senior Software Engineer for Agent Simulation & Evaluation! Today, NVIDIA is tapping... ..., driven by reasoning models, coding agents, parallel subagents... ....What you will be doing:Building agentic systems that... ...critical LPU workflows such as benchmarking, workload characterization...Full time$184.7k - $324.8k
Research Scientist / Engineer, Foundation Model Evaluation Cupertino, California, United States Software and Services We build frontier foundation models that power intelligent... ...and interests. Responsibilities Benchmark Design & Development: Design and implement...Relocation$128k - $256k
...operate Large Language Model (LLM) service... ...mainland China. We are building full-stack, end-to... ..., and intelligent agent systems. Beyond... ...year. Design evaluation systems for LLM-based... .... Build benchmarks and automated judging... ...great people. We lead with curiosity, humility...Temporary workInternshipLocal area$166.5k - $291.4k
...moment inspired Fred to build a company that could... ...teamBuild Agent is ServiceNow's AI coding... ...metadata types. This role leads the team responsible for the Build Agent evaluation framework, model support, and telemetry... ....Model Support & Benchmarking: Structured evaluation...Work at officeImmediate startRemote workFlexible hours- Join Sekai as we build an AI-driven consumer platform reminiscent of TikTok for mini-apps. We're looking... ...for a technical leader to design and own our agent runtime and orchestration layer, build workflows, and develop evaluation loops. This role offers the chance to work...Remote job
- ...Applied Machine Learning Ark team, working on large language model service platforms that span text and multimodal capabilities. You will help build end-to-end MaaS solutions, data pipelines, and evaluation systems across US and international markets. The internship...Internship
- ...NVIDIA is seeking a Senior Software Engineer for Agent Architecture and Evaluation to build agentic AI systems and analyze inference workloads in realistic settings. You will work with teams advancing evaluation across software and hardware to improve accuracy and efficiency...
$184k - $287.5k
...are seeking a senior vision language model engineer to design and build agentic data and training workflows... ...our researchers to develop and evaluate prototypes of our latest models, such... ...novel algorithms and search pipelines, benchmarking, and integrating prototypes in...Full time$184k - $287.5k
...and the Relational Foundation Model team is helping lead that transformation. We are building a unified foundation model that... ...models: you will design, build, and evaluate novel Transformer and graph... ...that moves beyond single-table benchmarks, we would love to hear from youWhat...Full time- ...Senior Machine Learning Scientist in San Jose. The role focuses on building AI solutions to enhance customer interactions across various... ...have a PhD in a quantitative field and experience deploying ML models in production. The position offers a diverse set of...
- ...Institute of Foundation Models We are a dedicated research lab for building, understanding, using,... ...post-training, and evaluation benchmarks. The role combines cutting... ...reasoning and agent development. Key Responsibilities... ...publication record in leading AI conferences such as...
- NVIDIA Corporation in Santa Clara, CA seeks a Product Evaluations Lead to own GenAI software evaluation, measurement, and go-forward analysis... ...enterprise ISV partners. You will drive experiments, benchmarks, and translate results into actionable insights guiding product...
- NVIDIA is seeking a highly analytical Product Evaluations Lead to own the evaluation, measurement, and go-forward analysis of our Generative... ...enterprise ISV partners. You will drive experiments and benchmarks, translating results into insights that guide product, research...
- NVIDIA seeks a highly analytical Product Evaluations Lead to own evaluation, measurement, and go-... ...partners. You will drive experiments, benchmarks, and translate results into clear... ...with GenAI product strategy, influencing model release criteria, product direction, and...
- ...ServiceNow is seeking a Software Engineering Manager for Build Agent to lead an 8-engineer team focused on evaluation infrastructure, model benchmarking, and telemetry. The role drives strategy, cross-team collaboration, and scalable processes to ensure robust evaluation...
- Apple Inc. is seeking a Research Scientist/Engineer to design evaluation systems for foundation models powering Apple products. You will work hands‑on across evaluation design, experimentation, and cross‑team collaboration to drive model improvement and product quality...
$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours$231.44k - $282.88k
...an experienced leader to head benchmarking and performance... ...team in creating benchmarks for evaluating CPU and accelerator performance... ...methodology for complex workloads. Build and deploy efficient tools and... ...from analytical to cycle‑based models; and from bare‑metal...- ...About the Institute of Foundation Models We are a dedicated research lab for building, understanding, using, and risk... ..., experimentation, and evaluation workflows. ~ This role balances... ...automated model evaluation or benchmarking systems. ~ Knowledge of cost...Visa sponsorship
$244.14k - $413.16k
XPENG is a leading smart technology company at the... ...learning systems that allow agents to plan over long... ..., large language models, and real-world autonomous... ...data. You will help build the learning systems... ...platform.Evaluation and benchmarking systems for agent capabilities...Full time- ...NVIDIA is seeking a Senior Software Engineer for Agent Simulation & Evaluation to help build agentic systems and provide performance analytics across GPU and LPU platforms. The role emphasizes Python tooling, AI agents, and collaboration with inference, hardware, runtime...
$160.5k - $240.7k
...developersoptimizeand deploy machine learning models on edge and mobile hardware.... ...and generative AI models Build tooling to analyze, profile,... ...quantization pipelines and evaluation harnesses to scale model... ...pipelines and model benchmarking at scale is a plus Level...Work experience placementImmediate startWork from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead, Build Agent Evaluation & Model Benchmarking. Be the first to apply!
- commissioning agent Santa Clara, CA
- state farm agent Santa Clara, CA
- executive protection agent Santa Clara, CA
- cruise agent Santa Clara, CA
- airport agent Santa Clara, CA
- agent Santa Clara, CA
- import export agent Santa Clara, CA
- remote chat agent Santa Clara, CA
- tsa agent Santa Clara, CA
- entry level special agent


