Tech Lead, Agent Eval Platform
ServiceNow
Company Description Who we are Moveworks: the Agentic AI Assistant platform that empowers the entire workforce. Our platform enables employees to converse with all of their business systems through natural language to quickly find answers and automate tasks. Powered by the world's most advanced LLMs, our proprietary models, and a sophisticated Agentic AI platform, we're transforming how work gets done by allowing AI to take initiative, streamline complex workflows, and continuously learn and adapt. Moveworks is trusted by over 5.5 million employees at more than 350 of the world’s largest companies, including 10% of the Fortune 500, to automate everyday tasks and streamline business operations. Recognized on the Forbes Cloud 100 and AI 50 lists, Moveworks was also named one of Fast Company’s 2025 Most Innovative Companies and Inc’s Best in Business, in the Best in Innovation category. Moveworks was also recognized at Microsoft’s 2025 Partner of the Year and in 2024, received the AI Breakthrough Award. In December 2025, Moveworks was acquired by ServiceNow, marking a pivotal milestone in our journey to create a single front door to work for all business systems. By combining ServiceNow’s leading workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we deliver the AI platform for every person and every workflow. Built to go beyond basic summaries to deliver meaningful business impact. Together, our AI acts across enterprise systems to turn conversations into completed work. By joining our team, you’ll be at the forefront of the AI transformation, backed by the global scale of ServiceNow and the agility of a high-growth company. We are looking for world-class talent to help us extend agentic AI to every employee across every corner of the business.Come join us! ServiceNow: it all started in sunny San Diego, California in 2004 when a visionary engineer, Fred Luddy, saw the potential to transform how we work. Fast forward to today — ServiceNow stands as a global market leader, bringing innovative AI-enhanced technology to over 8,100 customers, including 85% of the Fortune 500®. Our intelligent cloud-based platform seamlessly connects people, systems, and processes to empower organizations to find smarter, faster, and better ways to work. But this is just the beginning of our journey. Join us as we pursue our purpose to make the world work better for everyone. Job Description The Role Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better? That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing. This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments. What you get to do in this role: We're hiring across three areas. You'll anchor on one and touch the others; which one is a conversation we have with you, not a slot we drop you into. Eval orchestration at scale The runtime that executes multi-turn agent scenarios end-to-end — stand up the environment and user simulator, drive the useragentworld loop, collect transcripts, traces, and final state, run validators and scoring, tear down Scheduling, retries, high-concurrency execution, and run isolation at production dataset sizes Versioned specs, datasets, and reports, with run-to-run comparison as a first‑class operation Consolidating evals that run today as one‑off workflows onto a single orchestration service — one source of truth, one place to schedule and retry Establishing a reliability floor and an SLO for the harness itself Getting to self‑serve, so any team runs an eval without bespoke integration Agent observability and tracing Leading the move to OpenTelemetry‑native observability for the agent platform, replacing the parallel per‑service logging, correlation, and redaction mechanisms in use today The span data model for agent trajectories — prompts, tool calls, plan updates, outcomes — so a trajectory is queryable, not reconstructed by hand from log files Trace context propagation across async boundaries and sessions that stay alive for minutes or hours Making full prompts and completions survive the pipeline intact, and keeping eval traffic from contaminating its own data Fault attribution and cross‑run diffing: which component actually broke, and what changed since the last green run The debug surface support and harness engineers use, and the tracing contract with the team that builds the agent Stateful simulation The simulation environment itself: stateful fakes of the enterprise systems agents call — ITSM, HR, knowledge bases, inventory — backed by a real datastore that persists changes during a run, so a created ticket is visible to a later read Per‑run data injection and programmatic setup/teardown so every run is hermetic and repeatable LLM‑driven user simulators for open‑ended personas, and scripted state‑machine simulators for deterministic flows Contract‑testing mocks against real API schemas in CI, so simulation fidelity can't quietly drift as vendor APIs change Ahead of us: isolated sandbox environments reproducing the config, identity, search content, and permissions an agent actually reads — provisioned from an identical baseline and torn down every run And across all three: laying the foundation for using eval signal to optimize the agent, not just measure it. Qualifications To be successful in this role you have: Experience in at least 3 of these: Distributed systems: idempotency, delivery guarantees, isolation, and — unusually central here — determinism and reproducibility Orchestration and workflow runtimes: DAG execution, scheduling, retries, backfills, high‑concurrency job systems (Temporal, Airflow, Argo, or something you built yourself) Observability internals as a builder, not just a user: OpenTelemetry SDKs and collectors, semantic conventions, span context propagation, high‑cardinality trace data Concurrent and async programming: Python asyncio, Go concurrency, structured cancellation Data‑intensive pipelines: high‑volume ingest, schema evolution, sampling and retention trade‑offs gRPC/protobuf service and interface design Required: Ability to tech lead other engineers and the end to enddelivery of a project. Good communicationand soft skills. 10+ years building production backend or infrastructure systems Strong in Python or Go (ideally both) Experience designing and operating systems that handle real traffic at scale Comfort making a non‑deterministic system measurable. You don't need an ML background — but you should find it interesting to turn fuzzy agent behavior into a signal engineers are willing to gate releases on Comfort with ambiguity; these are novel problems without textbook solutions Additional Information Work Personas We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third‑party service. Equal Opportunity Employer ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. Accommodations We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact View email address on click.appcast.io for assistance. Export Control Regulations For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. From Fortune. ©2025 Fortune Media IP Limited. All rights reserved. Used under license. #J-18808-Ljbffr ServiceNow
- ServiceNow is seeking an experienced engineer to own the evaluation framework for AI agents within Moveworks. You will design rubrics and judges, orchestrate large-scale multi-turn scenarios, and build stateful simulators that enable reliable, reusable evals across enterprise...Platform
- ...Moveworks: the Agentic AI Assistant platform that empowers the entire... .... By combining ServiceNow’s leading workflow automation with Moveworks... ...The Role Moveworks' AI agents don't just generate text — they... ...configuration next to the dataset — so eval authors express intent, rather...PlatformFull timeWork at officeRemote workFlexible hours
$193.93k - $352.29k
...building a universal autonomy platform: self-driving for all roads and... ..., T. Rowe Price, and other leading investors. About the Team... ...build the platform that lets AI agents operate autonomously inside Nuro... ...of experiments, including the eval and confidence machinery required...PlatformImmediate startFlexible hours$258k - $387k
...around us, that’s why we’re building a universal autonomy platform: self-driving for all roads and all rides.Founded in 20... ..., Google, Softbank, Fidelity, T. Rowe Price, and other leading investors.About the RoleThe Eval Platform team owns the simulation and evaluation and...PlatformImmediate startFlexible hours$213k - $263k
...ride‑hail service and can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver has provided over ten... ...new cities and countries, each with their unique challenges. The Eval Authoring APIs team’s mission is to establish the foundational...PlatformFull timeRemote work$238k - $302k
...can also be applied to a range of vehicle platforms and product use cases. The Waymo Driver... ...be accountable for include: Clustering: Lead the creation of a dynamic event clustering... ...skills Excited about autonomous driving, Sim+Eval, eager to learn We prefer: 10+ years of...PlatformFull timeRemote work$193.93k - $291.15k
...why we’re building a universal autonomy platform: self-driving for all roads and all rides... ...Softbank, Fidelity, T. Rowe Price, and other leading investors.About the RoleAs the Technical... ...a proven track record of acting as a Tech Lead on complex software projects.BS, MS,...PlatformImmediate startRemote workFlexible hours- Join our team to play a pivotal role in mitigating tech risks and upholding operational excellence, driving innovation in risk management. As a Tech Risk & Controls Lead in Core Foundational Platforms , you will be responsible for identifying and mitigating compliance and...Platform
$207k - $300k
...Temporal orchestration, automated node re-bootstrapping, and self-service remediation.Own architectural and design decisions for the platform, collaborating with leadership and cross-functional teams to prioritize efforts, resolve roadblocks, and manage technical debt....Platform$138k - $225k
...the business needs of the team. We’re hiring a Data Foundations Lead to architect and scale the core data foundations that enable... ...Finance, Engineering, and Finance Technology and deliver durable platforms. Preferred Qualifications Finance domain fluency: Experience...PlatformFor contractorsWork at officeFlexible hours$170k - $275k
...relentless work. The Role As a Software Engineer on the Agent Harnessing team, you will build the core architecture that allows... ...communication layers, and execution environments—designed as a platform that Scout AI engineers can extend and depend on....PlatformFull timeRelocation package- ...accommodate family commitments. Dana is Applied's agentic platform for physical AI industries: apps, agents, and data on one governed platform. You will own Dana... ...agents: shipped an agent product, built harnesses or eval loops, or founded in the space Fluency in the current...PlatformFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift
$262k - $364k
...context working memory, and multi-agent collaboration for complex, long-running loops.Lead the development of low-latency... ...a team of Software Engineers and Tech Leads, driving resource allocation... ...Experience in agentic AI systems and platforms.Preferred qualifications:Master’s...Platform$155k - $185k
...OpportunityWe are looking for a Senior AI Agent & LLM Engineer who combines strong... ...experiences, evaluation systems, and the shared platform required to operate them reliably at... ...meetings transcribed, Otter.ai is the world’s leading tool for meeting transcription,...PlatformPermanent employment$207k - $300k
Lead the technical strategy and architecture of next-generation control planes and software infrastructure for highly-available ML... ...computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware...PlatformWorldwide$207k - $301k
...end to end stack and analysis tools.Partner with product area leads to understand model optimization use cases, drive cross functional... ...power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware...PlatformWorldwide$174k - $252k
..., deploy, monitor, and optimize enterprise-grade generative AI agents.Architect and implement intuitive, AI-centric workflows that simplify... ...the power of Google's next generation agentic development platform to empower anyone to build enterprise grade agents. We work...Platform- ...once it's out there. We're looking for a Tech Lead to own fleet deployment, testing... ...site Build and evolve the fleet management platform — the tooling, dashboards, and APIs used... ...Docker, and edge-appropriate orchestration/agent frameworks) for packaging and deploying...PlatformRemote work
- ...converge. The Role We are looking for a Tech Lead Manager who thinks like a product builder... ...large volumes of social data, run AI agents on that data in real time, and deliver insights... ...AI-powered or agentic software platforms at scale. Tech Stack Familiarity: Hands-...PlatformShift work
$120k - $220k
...5, NewsBreak is the Content Intelligence platform shaping the future content economy. With... ...visit About the Role We're building the agent platform that powers NewsBreak's next-... ...weekly, works closely with product, and takes eval and observability seriously. Our codebase...PlatformFull timeLocal areaWork from home$193.93k - $291.15k
...components and network software. Design and lead development of scalable, fault-tolerant... ...and next-generation autonomous vehicle platforms. Lead cross-functional technical... ...experience and a proven track record as a Tech Lead on complex software projects. ~ BS...PlatformFull time$207k - $300k
...HTML, CSS or equivalent.3 years of experience with Google Cloud Platform.3 years of experience with large scale software design and... ...algorithms.3 years of experience in a technical leadership role leading project teams and setting technical direction.3 years of experience...Platform- Moveworks is seeking a Tech Lead for its Agentic AI Product in Mountain View, California. In this role, you'll leverage cutting-edge Machine Learning technologies to drive automation and innovation within enterprise uses. You will lead a team, design product features, and...Platform
- JPMorgan Chase & Co. seeks a Tech Risk & Controls Lead to mitigate technology risks and uphold operational excellence within Core Foundational Platforms. You will identify and mitigate compliance and risk issues while guiding technology owners and regulators toward robust...Platform
$180k - $260k
Clockwork.io in Palo Alto is seeking a Tech Lead to architect and develop a high-performance network monitoring platform. This role demands strong programming skills in languages such as C++, Go, or Python and significant experience with distributed systems and networking...Platform$160k - $200k
...Staff Engineer to design and ship the production multi-agent systems at the core of LeanData’s new platform of autonomous agents for go-to-market teams.This is a... ...dataEvaluate everything you ship: build the eval cases, rubrics, and regression tests that prove a change...PlatformWork at office$180k - $260k
...About the Role We are looking for a passionate and experienced Tech Lead - Frontend / Full Stack to join our growing engineering team.... ...and deliver high-impact solutions. Continuously improve platform performance, usability, and developer productivity by adopting...Platform$193.93k - $291.15k
...why we’re building a universal autonomy platform: self-driving for all roads and all rides... ...Softbank, Fidelity, T. Rowe Price, and other leading investors.About the RoleThe mandate of... ...a key member of the Prediction and Smart Agents team, you will focus on building state-of...PlatformImmediate startFlexible hours- ...hiring founding researchers who influence agenda, culture, and infrastructure. You'll access real-world data at scale and work with ML platform, product, and operations teams to shape research that translates into real-world impact. We seek exceptional researchers with a...Platform
$217k - $271.5k
...Intelligent Content Management. Our platform enables organizations to fuel... ...2005, Box simplifies work for leading global organizations,... ...infrastructure that powers AI agents across our entire product ecosystem... ..., and quality.Champion an eval-driven development culture — define...PlatformLive inWork at officeImmediate startShift work3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Tech Lead, Agent Eval Platform. Be the first to apply!
- technical leader Mountain View, CA
- technical lead Mountain View, CA
- state farm agent Mountain View, CA
- tsa agent Mountain View, CA
- operations agent Mountain View, CA
- import export agent Mountain View, CA
- commissioning agent Mountain View, CA
- remote chat agent Mountain View, CA
- agent Mountain View, CA
- executive protection agent Mountain View, CA




