Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff ML Engineer, Agent Training & Environments

$250k - $280k
Full-time

Labelbox

Shape the Future of AI


At Labelbox, we're building the critical infrastructure that powers breakthrough AI models at leading research labs and enterprises. Since 2018, we've been pioneering data-centric approaches that are fundamental to AI development, and our work becomes even more essential as AI capabilities expand exponentially.

About Labelbox


We're the only company offering three integrated solutions for frontier AI development:


  1. Enterprise Platform & Tools : Advanced annotation tools, workflow automation, and quality control systems that enable teams to produce high-quality training data at scale

  2. Frontier Data Labeling Service : Specialized data labeling through Alignerr, leveraging subject matter experts for next-generation AI models

  3. Expert Marketplace : Connecting AI teams with highly skilled annotators and domain experts for flexible scaling

Why Join Us



  • High-Impact Environment : We operate like an early-stage startup, focusing on impact over process. You'll take on expanded responsibilities quickly, with career growth directly tied to your contributions.

  • Technical Excellence : Work at the cutting edge of AI development, collaborating with industry leaders and shaping the future of artificial intelligence.

  • Innovation at Speed : We celebrate those who take ownership, move fast, and deliver impact. Our environment rewards high agency and rapid execution.

  • Continuous Growth : Every role requires continuous learning and evolution. You'll be surrounded by curious minds solving complex problems at the frontier of AI.

  • Clear Ownership : You'll know exactly what you're responsible for and have the autonomy to execute. We empower people to drive results through clear ownership and metrics.

Role Overview


Labelbox is the RL data factory for advancing frontier agent capabilities. We build the data, environments, and evaluations that frontier labs use to train and judge their agents.

This role sits where training meets infrastructure. You will run the experiments and build the systems that run them: environments agents act in, verifiers that decide whether they succeeded, and the fine-tuning pipelines that turn that signal into a better model. We're looking for someone who does both halves — the engineering throughput of a strong platform engineer, and real depth in post-training agents.

The bar is high: engineers with strong judgment who set technical direction, turn prototypes into reliable systems fast, and are at the frontier of agent-first engineering practice.

 

What you'll work on



  • RL environments for agentic tasks: task definitions, tool surfaces, state and reset semantics, reward design — and the harness that runs thousands of them in parallel.

  • Verifiers and graders: programmatic checks, LLM judges, rubric pipelines, View email address on jobs.jobcopilot.com scoring. Deciding what "the agent succeeded" means, and making that judgment trustworthy at scale.

  • Fine-tuning pipelines that turn evaluation signals into measurable agent improvements — SFT and RL, from data collection through training to checkpoint evaluation.

  • Eval systems that run millions of agent trajectories to measure model and product quality.

  • Training and serving infrastructure that scales to the throughput frontier labs need: multi-launcher orchestration, long-running job fault tolerance, cost accounting.

What we're looking for


As an engineer


  • A 3+ year track record of shipping systems that customers and other engineers still rely on.

  • Exceptional throughput, without the quality tax. You ship a lot, you review a lot, and the v1 you ship becomes the foundation the rest of the team builds on.

  • Strong system and API design judgment. Hard architecture calls land with you: you make them, defend them under pressure, and update fast when someone else is right.

  • You ship production code with coding agents daily. You know where they break and what it takes to make them reliable, and you use that to move the whole team faster.

  • You build the substrate other people's work runs on — tooling, CI, harnesses, libraries — and you treat that as the job, not a distraction from it.

  • You move fast in ambiguous, startup-pace environments, with influence over authority.

  • Deep proficiency in Python, and comfort across the rest of the stack.

As an RL post-training practitioner


  • You have fine-tuned models for agentic tasks and made them measurably better. SFT plus at least one RL method (GRPO, PPO, DPO, or similar) in production.

  • You have built environments agents operate in, and you know why reward and task design is where most of the difficulty actually lives.

  • You have designed verifiers or graders for open-ended work, and you know how they get gamed.

  • You debug training runs forensically and methodically.

  • You reason about compute-economics. You know what an experiment costs, when a run is not worth finishing, and how to get the same signal for a tenth of the spend.

  • You write up what you learned so it changes what the team does next.

 

Nice to have



  • Experience with agent harnesses and coding agents as subjects of training and evaluation.

  • Multi-tenancy and isolation for untrusted agent execution: sandboxing, egress control, credential handling.

  • Background in production distributed systems, ML infrastructure, or data systems at scale.

  • Experience working directly with frontier labs or other highly technical customers.

 

Our Technology Stack


Our engineering team works with a modern tech stack designed for scalability, performance, and developer efficiency:


  • Frontend: React.js with Redux, TypeScript

  • Backend: Node.js, TypeScript, Python, some Java & Kotlin

  • APIs: GraphQL

  • Cloud & Infrastructure: Google Cloud Platform (GCP), Kubernetes

  • Databases: MySQL, Spanner, PostgreSQL

  • Queueing / Streaming: Kafka, PubSub

Labelbox strives to ensure pay parity across the organization and discuss compensation transparently. The expected annual base salary range for United States-based candidates   is below. This range is not inclusive of any potential equity packages or additional benefits. Exact compensation varies based on a variety of factors, including skills and competencies, experience, and geographical location.

Annual base salary range

$250,000 - $280,000 USD

Life at Labelbox



  • Location : Join our dedicated tech hub in San Francisco

  • Work Style : Hybrid model with 3 days per week in office, combining collaboration and flexibility

  • Environment : Fast-paced and high-intensity, perfect for ambitious individuals who thrive on ownership and quick decision-making

  • Growth : Career advancement opportunities directly tied to your impact

  • Vision : Be part of building the foundation for humanity's most transformative technology

Our Vision


We believe data will remain crucial in achieving artificial general intelligence. As AI models become more sophisticated, the need for high-quality, specialized training data will only grow. Join us in developing new products and services that enable the next generation of AI breakthroughs.

Labelbox is backed by leading investors including SoftBank, Andreessen Horowitz, B Capital, Gradient Ventures, Databricks Ventures, and Kleiner Perkins. Our customers include Fortune 500 enterprises and leading AI labs.

Your Personal Data Privacy : Any personal information you provide Labelbox as a part of your application will be processed in accordance with Labelbox’s Job Applicant Privacy notice .

Any emails from Labelbox team members will originate from a @labelbox.com email address. If you encounter anything that raises suspicions during your interactions, we encourage you to exercise caution and suspend or discontinue communications.

Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the Staff ML Engineer, Agent Training & Environments in Remote vacancy
  • Staff Machine Learning Engineer, Agent Memory & Reasoning (University) Location Employment Type Full time Location...  ...least the next six months: no model training, no fine tuning, no deep GPU or CUDA...  ...fundamentals to go with your ML and agent experience. Practical fluency... 
    Training
    Full time
    Remote work
    Shift work

    Wand Inc.

    Brooklyn, NY
    4 days ago
  •  ...deploying real world systems for demanding environments, the Stack team is dedicated to...  ...driving environment. We're looking for a Staff or Senior Engineer to lead the development, maintenance,...  ...Extensive experience architecting, training, and deploying deep learning models... 
    Training
    Remote job
    Full time

    Stack Av

    Remote
    13 hours ago
  • $251k - $310k

     ...Conduct comprehensive experimentation to train and deploy state-of-the-art Multimodal LLMs...  ...and Radar.. Partner effectively with engineering and research teams across Waymo to deploy...  ...(e.g., xprof), and debugging of ML models. You have: PhD or Masters... 
    Training
    Full time
    Temporary work
    Remote work

    Waymo

    New York, NY
    13 hours ago
  •  ...customers. Spara is building AI agents to handle this explosion of...  ...of Triplemint, where he ran training, coaching, and operations...  ...We are seeking a Staff ML Engineer with a passion for building...  ...a lightweight, early-stage environment and helping to define solutions... 
    Training
    Full time
    Work at office
    Work from home
    Flexible hours
    3 days per week

    Spara

    New York, NY
    13 hours ago
  • $220k - $247k

     ...Do As a  Senior Staff Machine Learning Engineer , you will operate...  ...design of large-scale ML systems and shared...  ...capabilities across agents, text, image, audio,...  ...scalable ML platforms (training, evaluation,...  ...high-growth startup environments   Location This... 
    Training
    Full time
    Work at office
    Immediate start
    Flexible hours
    3 days per week

    Typeface

    Remote
    13 hours ago
  • $249.6k - $299.5k

     ...data, contributing to Torc's data workflow for training and validation across the complete AV stack. As Staff ML Engineer, you will lead this team, setting its...  ...the integration of the framework in a cloud environment and automate the pipeline to scale target verification... 
    Training
    Full time
    Work experience placement
    Immediate start
    Relocation

    Torc Robotics

    Remote
    13 hours ago
  •  ...best parts of a startup environment - small team, high...  ...spans cutting-edge LLMs / ML, large-scale data systems...  ...a team of exceptional engineers, analysts, and investors...  ...We're looking for a Staff ML Engineer to lead the...  ...spanning data pipelines, training workflows, inference... 
    Training
    Full time
    Work at office
    Local area

    Versant Limited

    Remote
    13 hours ago
  • $207k - $300k

     ...generated content through advanced context engineering and agentic feedback loops to identify...  ...(AIGC) to fulfill user interests. As a Staff Machine Learning Engineer focusing on Search...  ..., experience, and relevant education or training. US: $207000 - $300000 (USD) + 20% bonus... 
    Training

    Google

    Mountain View, CA
    1 day ago
  • $189.3k - $320.7k

     ...behavior across real-world scenarios.As a Staff ML Engineer on the Prometheus team within the...  ...of autonomous vehicle development—from training and validation to testing and safety....  ...Experience deploying ML models into production environments and understanding end-to-end... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $189k - $300k

     ...works on and delivers ML models to the product...  ...foundation model pre-training and fine-tuning with data...  ...-impact team of AI/ML engineers, data scientists and...  ...autonomous vehicles. As a Staff AI/ML Engineer in the...  ...into production environments and understanding end-... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    1 day ago
  •  ...-connected world.Role OverviewAs our Staff Software Engineer, ML infra Engineer for Search & Discovery...  ...structured and unstructured data needed to train complex ML models and efficiently...  ...committed to providing a safe work environment for its employees and its consumers.If... 
    Training
    Temporary work

    Coupang

    Mountain View, CA
    2 days ago
  • $281k - $356k

     ...machine learning models to deliver training and evaluation data for...  ...for researchers and software engineers who are passionate about developing...  ...in Python and standard ML frameworks (e.g., JAX, TensorFlow...  ..., or complex simulation environments ~ Deep understanding of state... 
    Training
    Full time

    Waymo

    Remote
    13 hours ago
  • $251k - $310k

     ...simulations of realistic environments for testing, training, and validation of the Waymo...  ...of machine learning (ML) engineers, software engineers, and ML...  ...world, encompassing realistic agents, roads, traffic systems,...  ...you will report to a Senior Staff Engineering Manager... 
    Training
    Full time
    Remote work

    Waymo

    Remote
    13 hours ago
  • $252k - $315k

     ...build production AI agents that automate complex...  ...one of the hardest engineering challenges.As a Staff Frontier Agent Engineer...  ....Unlike traditional ML roles that focus on...  ...in high-stakes environments.Collaborate with infrastructure...  ...education or training. Scale employees in... 
    Training
    Full time

    Scale AI

    New York, NY
    4 days ago
  • $151k - $177.5k

     ...battery from the inside out today. We engineer and manufacture ground-breaking battery...  ...or other industrial data Experience training, validating, and deploying custom models...  ...systems architecture, or controls-adjacent environments Experience with sequence modeling,... 
    Training
    Full time

    Sila

    Remote
    13 hours ago
  • $141k - $249k

     ...... - Build standardized distributed training frameworks for research and production,...  ...techniques. - Work with researchers and ML engineers on best-practices for optimal resource...  ...- Experience in Bazel in a monorepo environment, and integrating third party packages into... 
    Training
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    Pittsburgh, PA
    more than 2 months ago
  • $189.3k - $290.7k

     ...behavior across real-world scenarios.As a Staff ML Infra Engineer, you will drive the development of...  ...enable rapid dataset generation, training, evaluation, and iteration of our most...  ...providing an inclusive workplace creates an environment in which our employees can thrive and... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    13 hours ago
  • Wand is seeking a Staff Machine Learning Engineer to join University, a new team focused on agent memory, evolution, and reasoning. You will design and build memory systems, ensure secure, scalable deployment in the cloud, and work with a world-class team to push the frontier... 
    Remote job

    Wand Inc.

    Brooklyn, NY
    4 days ago
  •  ...Perplexity is seeking experienced ML engineers to design, build, and optimize the recommendation...  ...Experience with large-scale ranking and training infrastructure (multi-stage retrieval...  ...something and hope. They must have armies of agents and workers who can constantly work in... 
    Training
    Full time

    Perplexity®️

    San Francisco, CA
    13 hours ago
  • $171.7k - $303.9k

     ...the world!The Data Labeling Engineering team designs, builds, and operates...  ..., data engineering, and AI/ML, defining the strategies,...  ...controls that create reliable training data at scale. Our tools and...  ...inclusive workplace creates an environment in which our employees can thrive... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...at Boulevard. We’re looking for a Staff ML Engineer to help shape the next generation of our...  ...and feature engineering to model training, deployment, monitoring and continue improvement...  ...learning systems in production environments. ~ Strong programming skills in Python... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    Boulevard

    Remote
    a month ago
  •  ...seeking a Member of Technical Staff to advance state-of-the-art models...  ...’ll work across research and engineering to push performance and scale...  ...engineers, Python and ML framework expertise, and experience...  ...with large-scale distributed training, aiming to push frontier AI capabilities... 
    Training
    Remote work

    Cohere

    San Francisco, CA
    13 hours ago
  • $85 per hour

     ...Evaluate and improve frontier AI coding agents by completing realistic machine learning engineering tasks and assessing model outputs...  ...that reflect production ML workflows, helping a leading AI research...  ...implementations involving model training, inference systems, MLOps, and... 
    Training
    Hourly pay
    Remote work

    SaidGig

    United States
    28 days ago
  • ServiceNow is seeking a Senior Staff engineer to shape the next generation of agentic systems. You will own cross-cutting architecture, tool...  ...and React/TypeScript frontend. Strong capability in multi-agent coordination, context management, and real-time execution is essential... 
    Remote job

    Worky

    Santa Clara, CA
    1 day ago
  • $189.3k - $320.7k

     ...behavior across real-world scenarios.As a Staff AI/ML Future Sensing Engineer in the Embodied AI organization,...  ...efforts spanning data curation, training, validation, performance...  ...preparing ML models for production environments and understanding end-to-end deployment... 
    Training
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    2 days ago
  • $218.8k - $335.3k

    Job DescriptionStaff AI/ML Engineer, AV ML Infra We’re General Motors...  ...optimizes large-scale ML training and inference across cloud and...  ...commercialization.Position Overview: As a Staff AI/ML Engineer, you will be...  ...workplace creates an environment in which our employees can... 
    Training
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Austin, TX
    2 days ago
  • $262k - $361k

     ...energy, AI, software, engineering, and product to build...  ...development of production ML/AI systems. You will...  ...mentoring senior and staff-level engineers, establishing...  ...in production environments. Establish scalable...  ...experience building, training, and deploying large-scale... 
    Training
    Full time
    Remote work
    Flexible hours

    Tapestry, Inc.

    Remote
    13 hours ago
  • $238k - $302k

     ...simulation across 15+ U.S. states. The DUE ML Core team will build and operate scalable...  ...machine learning models to deliver training and evaluation data for hundreds of metrics...  ...are looking for researchers and software engineers who are passionate about developing... 
    Training
    Full time
    Remote work

    Waymo

    Remote
    13 hours ago
  • $251k - $310k

     ...simulations of realistic environments for testing and training the Waymo Driver. Our...  ...collaborative group of software engineers, machine learning (ML) engineers, and data...  ..., including realistic agents (vehicles, pedestrians,...  ...will report to a Sr Staff TLM. You will:... 
    Training
    Full time
    Remote work

    Waymo

    Remote
    13 hours ago
  • $229k - $317k

     ...and efficient navigation in complex environments.   As an engineer in the ODIN team, you will develop advanced...  ...architectures and sophisticated training techniques, leveraging all the inputs...  ...we have at Zoox. Drive end-to-end ML solutions from research to production... 
    Training
    Full time

    Zoox

    Remote
    13 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff ML Engineer, Agent Training & Environments. Be the first to apply!