Senior / Staff ML Ops Engineer
$184k - $272kWaabi
Job Description
Job Description
Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI. With a world-class team, we're unlocking the next era of autonomous transportation with technology that's powering commercial autonomous trucks and robotaxis. Waabi is backed by and partners with world leaders in AI, automotive, logistics, and deep tech.
With offices in Toronto, San Francisco, Dallas, and Pittsburgh, Waabi is growing quickly and looking for diverse, innovative and collaborative candidates who want to impact the world in a positive way. To learn more visit:
You will..
- Build and evolve our training infrastructure on Kubernetes with Infrastructure — GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, and the operators and workflow engines that keep long-running training reliable.
- Shape the developer-facing surface — CLIs, SDKs, job submission, templates, paved paths — designed with the teams who'll use them. Make the common case one command and keep the uncommon case possible.
- Shorten the inner loop. Time to first training run, edit-to-signal latency, local iteration before a job hits the cluster, fast failure over slow mystery. Measure it, publish it, drive it down.
- Evangelize best-in-class tooling and frameworks. Track what the ecosystem is shipping, evaluate honestly, and make the case with working prototypes and migration paths — or say plainly when a shiny thing isn't worth the switching cost.
- Strengthen the data and artifact layer. Dataset versioning, sharding, and high-throughput loading of large multimodal sensor data, so jobs saturate GPUs instead of waiting on I/O.
- Turn one-off Python into durable tooling — tested, documented, observable libraries, CLIs, and services with sane defaults, and deletions where they're overdue.
- Make experiments legible, with the teams who live in them: experiment hygiene, dashboards researchers trust, a real model registry, and lineage from dataset to checkpoint to simulation result.
- Ship CI/CD for models alongside autonomy and simulation, so a model change is validated the same way a code change is.
- Build observability across the ML stack — utilization, throughput, failure modes, queue times, cost per experiment. When a job fails at 3am on node 47, the researcher should find out why without you.
- Treat docs, onboarding, and support as product surface — golden-path guides, a new researcher productive on day two, office hours that turn repeat questions into shipped fixes.
- Drive adoption, not just availability. Prototype with real users, watch them work, iterate. A tool nobody adopts didn't ship.
- Make the platform boringly reliable — fewer failures, faster recovery, and none of the manual steps that quietly cost a team days.
- Build guardrails that don't feel like walls, with Security, IT, and Infrastructure: access controls, data handling, and cost governance that hold up in an IP-sensitive environment while staying self-serve.
Qualifications:
- 5+ years of software or infrastructure engineering, including tools or platforms used by other engineers and operating ML or data-intensive production systems.
- Hands-on Kubernetes expertise — GPU scheduling, autoscaling, Helm or equivalent, networking fundamentals, and the ability to debug a cluster under load rather than restart it.
- Excellent Python, and a track record of designing APIs and CLIs other people enjoy using.
Practical AWS depth: object storage at scale, IAM, GPU compute, networking, cost management, and infrastructure as code (Terraform, Pulumi, or similar). - Distributed training in PyTorch (DDP, FSDP, or similar), plus experiment tracking and model registry tooling — from the perspective of someone who made them pleasant for others to use.
- Fluency with containers, CI/CD, and modern build systems, including large monorepos.
- The ability to influence without authority: evaluate a framework on its merits, pilot it credibly, and persuade skeptical senior engineers to change how they work.
- A collaborative default — you'd rather co-own a system than draw a boundary around your part of it.
- User empathy: you'd rather fix the third-most-interesting problem blocking ten people than the most interesting one blocking nobody.
- Strong product instincts, strong writing, and comfort operating autonomously in ambiguous territory.
- Passionate about self-driving technologies and frontier AI, and about what a small, world-class team can do with the right infrastructure.
Bonus/nice to have:
- Internal developer platform, research platform, or DevEx work — with a story about a tool whose adoption you grew from zero.
- Large-scale distributed GPU training: hundreds to thousands of accelerators, NCCL, high-performance cluster networking, collective communication tuning.
- High-throughput loading of LiDAR or camera data, and formats such as Parquet or WebDataset.
- Workflow and scheduling systems — Argo Workflows, Ray, Flyte, Kubeflow, or Slurm.
- Build-system depth (Bazel or similar), including remote caching in a monorepo.
- Simulation infrastructure or large-scale batch evaluation pipelines.
- Background in ML, robotics, or autonomous systems infrastructure.
- Security- and IP-sensitive production environments.
- Open-source contributions to ML infrastructure or developer tools.
The US yearly salary range for this role is: $184,000 - $272,000 USD in addition to competitive perks & benefits. Waabi (US) Inc.’s yearly salary ranges are determined based on several factors in accordance with the Company’s compensation practices. The salary base range is reflective of the minimum and maximum target for new hire salaries for the position across all US locations. Note: The Company provides additional compensation for employees in this role, including equity incentive awards and an annual performance bonus.
Perks/Benefits:
- Competitive compensation and equity awards.
- Health and Wellness benefits encompassing Medical, Dental and Vision coverage (for full-time employees only).
- Unlimited Vacation.
- Flexible hours and Work from Home support.
- Daily drinks, snacks and catered meals (when in office).
- Regularly scheduled team building activities and social events both on-site, off-site & virtually.
- As we grow, this list continues to evolve!
Waabi is a technology start-up building technologies to transform the way the world moves. Join our talented team to be a part of the future and to make an impact!
Waabi is an equal opportunity employer. We celebrate diversity and are committed to creating a supportive, inclusive, and accessible workplace for all our employees. We seek applicants of all backgrounds and identities, across race, color, ethnicity, national origin or ancestry, age, citizenship, religion, sex, sexual orientation, gender identity or expression, military or veteran status, marital status, pregnancy or parental status, caregiver status, disability, or any other characteristic protected by law. We make workplace accommodations for qualified individuals with disabilities as required by applicable law. If reasonable accommodation is needed to participate in the job application or interview process please let our recruiting team know.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
$228.96k - $315.36k
...and stop fraud before it happens. The team owns the full ML lifecycle—from feature pipelines and model training to production... ...as we grow to support hundreds of customers.As a Senior Machine Learning Engineer, you will own the development of high-performance feature...SeniorWork experience placementWork at officeLocal area$147k - $268.4k
...never been made means doing what's never been done. If you're an engineer, scientist, or builder who thrives on problems no one has... ...of patients across the globe. What You'll Be Doing As an ML Ops Engineer, you build and operate the platforms that run the end-...SuggestedRemote workFlexible hours2 days per week$250k - $350k
...in production, paired with applied ML research, design, and evaluation to... ...agentic AI looks like.About the RoleAs a Staff Machine Learning Research Engineer, you will operate across the full... ...capabilitiesSet AI/ML technical direction, mentor senior and staff-track engineers and...SeniorFull time$118k - $169k
...platforms, tools, and processes that take our models from ideas to production models, serving predictions in real time. The Sr. ML Ops Engineer will partner with our Data Science, Data Product Management, Product Engineering, and Data Platform teams to create and support...SeniorHourly payWork experience placementWork at officeImmediate startVisa sponsorshipWork visaFlexible hours- ...governance, maintain auditability, and deliver reliable outcomes at scale. About the Role We are looking for a visionary Senior ML Engineer who will bridge the gap between high-level architecture and hands-on execution, specifically focusing on simplifying...SeniorFull timeShift work
$200k - $300k
...configurations, and manipulation scenarios out of the box. At Chef, we're building that model: the Food Foundation Model. As a Senior ML Engineer, Foundation Models, you will work at the frontier of large-scale robot learning: training and fine-tuning the Food Foundation...SeniorFull timeFlexible hours- ...that will be used by 1000s of developers and enterprise users.ML performance, quality, and systems acumen-ship: Experience in tuning... ...$149,400 - $195,050 Qualifications 7+ years of software engineering experience building and operating enterprise systems, APIs, or...SeniorWork at office
- ...RDQ127R59 Summary As a Senior Applied ML Engineer on the Applied AI team at Databricks, you will use machine learning, scheduling, and optimization algorithms to maximize the efficiency and performance of our infrastructure. Your work will span the entire stack—from...SeniorFull time
$170.1k - $258.3k
...approaches to model export, kernel development, and performance engineering so that every cycle on our accelerators translates into better... ...kernels and custom libraries that sit at the heart of our on‑vehicle ML inference for ADAS and autonomous driving. We own making core...SeniorFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$128.7k - $261.3k
...export, kernel development, and performance engineering so that every cycle on our accelerators... ...of automated driving. The RoleAs a Senior Compiler Engineer on the AI Kernels & Compilers... ...path fast, reliable, and effortless for ML engineers across the AV organization to compile...SeniorFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$147k - $268.4k
...never been made means doing what's never been done. If you're an engineer, scientist, or builder who thrives on problems no one has... ...patients across the globe. What You'll Be Doing As an ML Ops Engineer, you build and operate the platforms that run the end-...Full timeRemote workFlexible hours2 days per week$130.2k - $195.3k
...Senior ML Platform Engineer We're on a mission to unleash the power of content… you in? We've got the brands, we've got the stars, we've got the... ...globe. Specific projects will include designing the ML Ops platform and Agentic AI layer that will enable the Data Science...Senior- ...Senior Client Engineer SAN FRANCISCO, CA ENGINEERING FULL-TIME What will you be doing? Training machine learning models over billions of data points. Quantifying predictive uncertainty using probabilistic and Bayesian methods. Creating models that quickly generalize...SeniorFull timeWork experience placement
$166k - $210.25k
RDQ127R59SummaryAs a Senior Applied ML Engineer on the Applied AI team at Databricks, you will use machine learning, scheduling, and optimization algorithms to maximize the efficiency and performance of our infrastructure. Your work will span the entire stack—from cluster...SeniorLocal areaWorldwide$141k - $249k
...collaborative candidates who want to impact the world in a positive way. You will... Collaborate closely with autonomy and algorithm engineers to scale safe self-driving systems using an AI-first approach. Expand the model deployment pipeline to new GPUs and embedded...SeniorWork at officeWork from homeFlexible hours$300k
...Senior Technical Leader Grindr is an AI-native platform powering how millions of gay... ...term relationship. We're doubling down on ML as the future of Grindr, and in these early... ...across teams, collaborating with engineering, data science and product teams to turn bold...SeniorCasual workWork at officeImmediate startWorldwideFlexible hours- ...transformers and spatial models run efficiently on both cloud and edge compute resources. Learn more at About the Role As an ML / DevOps Engineer, you will play a pivotal role in advancing our infrastructure, scaling enterprise deployment workflows, and refining...Work at office
- ...About the role Our client is a well-funded AI startup building production-grade ML infrastructure used by enterprise customers. They are looking for a Senior AI/ML Engineer to own model training pipelines, evaluation systems, and inference serving at scale. Full-time...SeniorFull time
$298k - $368k
...from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently... ...In this hybrid role, you'll report to a Senior Staff Technical Lead Manager. You will:... ...for onboard requirements. Mentor ML engineers and foster an engineering...SeniorFull timeRemote workShift work$148.5k - $223.9k
...efforts. Job Category Software Engineering Job Details About Salesforce... ...you are the future of Salesforce. Senior Member of Technical Staff - Senior Machine Learning Engineering... ...managing automated, production-grade ML pipelines. Software Engineering...SeniorFull time- ...Job Description - Senior AI/ML Engineer Location: Melbourne (hybrid, 3+ days a week in our Melbourne office) About Artificial Analysis Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand...SeniorFull timeWork at office3 days per week
$240k - $270k
...to reshape how people build in Sigma, discover insights, and make smarter decisions—fast. That’s where you come in. As an AI/ML Engineer, you’ll join a growing team focused on building the AI foundation that will power Sigma for the future. Your work will become an...SeniorFull timeWork at officeFlexible hours- ...either side. The role This is the first Machine Learning Engineer role at Metriport. We have access to the richest clinical... ...them in production and keeping them honest. This is applied ML on messy, high-dimensional, real-world healthcare data - not a research...SeniorWork at officeWork from homeRelocationFlexible hours
- ...Qualifications: Expert-level PyTorch. Proven software engineer who loves ML; comfortable writing production code across the stack. Hands... ...happiness. Deep knowledge of the ML lifecycle: dataset ops, training pipelines, eval frameworks, deployment, and monitoring...Full timeContract workFlexible hoursShift work
$177k - $218k
...investments. You will partner with senior leadership, including VPs and... ...is the equivalent of a Senior Staff-level.Success looks like: a... ...on Design, Research, Writing, Ops and Front-end Development who... ...with our partners in Product, Engineering, Data, and Marketing to design...SeniorFull timeWork at officeLocal area2 days per week3 days per week$190k - $270k
...Role We're hiring a machine learning engineer to work on our Computational Design / CAD... ...collaborate effectively with our AI team (deep ML expertise is not required; we have that... ...and level.) Level: Open to a range of seniority; final leveling determined during the...SeniorFull time$200k - $260k
...applications — serving speech-to-text and text-to-speech models with best-in-class latency and reliability. We're looking for a Senior ML Engineer to drive the model serving layer for voice workloads. You'll work hands-on with inference engines like TRT-LLM and SGLang to...SeniorFull time- ...here. You'll be responsible for our core ML detection platform, and the systems you build... .... Raise the technical bar. As a senior voice on a lean team, your instincts and... ...You Are ~4+ years of machine learning engineering experience , with meaningful time shipping...SeniorFull timeWork at officeRelocation package
$244k - $320k
...notifications, our AI-powered personalization engine delivers bespoke experiences that drive... ...across thousands of brands. As a Senior Machine Learning Engineer, you will play a... ..., scaling, and operating production-grade ML systems that drive real-time personalization...SeniorFull time$204k - $259k
...collaborate with research teams at Alphabet. We have access to millions of miles of driving data from a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently and continuously learning from large scale real-world data, to (2) develop models...SeniorFull timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior / Staff ML Ops Engineer. Be the first to apply!
- staff data engineer San Francisco, CA
- senior staff engineer San Francisco, CA
- senior staff systems engineer San Francisco, CA
- engineering aide San Francisco, CA
- software engineer staff San Francisco, CA
- staff design engineer San Francisco, CA
- assistant engineer San Francisco, CA
- staff security engineer San Francisco, CA
- technology administrator San Francisco, CA
- staff engineer San Francisco, CA



