Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer, ML Infra, Autonomy

$206.5k - $258.1k
Full-time

Rivian

About Rivian

Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. 

As a company, we constantly challenge what’s possible, never simply accepting what has always been done. We reframe old problems, seek new solutions and operate comfortably in areas that are unknown. Our backgrounds are diverse, but our team shares a love of the outdoors and a desire to protect it for future generations.


Role Summary

Rivian Autonomy is building an ML Infrastructure team to give hundreds of ML engineers a training platform they can trust at fleet scale. We are seeking a Staff Software Engineer to help design and build the platform that trains and evaluates our autonomous driving models: the control plane that schedules jobs across accelerator clusters, the storage and I/O layer that feeds them, the observability that tells us where every GPU-hour goes, and the training framework layer that lets ML engineers write model code once and run it on any silicon we operate.

Autonomy at Rivian runs a training fleet of thousands of GPUs on Kubernetes over petabyte-scale sensor data, with additional accelerator types arriving in the coming months. The fleet is fully allocated, so the next step change comes from goodput — how much useful training each GPU-hour delivers. Raising it, through scheduling that keeps large gang jobs fed, a storage and I/O layer that keeps up with the accelerators, and per-workload baselines that make every optimization measurable, is the heart of this role.

The control plane and training fleet already exist and serve every ML engineer in Rivian Autonomy; the scheduling, storage/I/O, observability and framework layers on top of them are largely still to be built. This is a role for someone who wants to set technical direction rather than maintain it.

The platform spans four areas — job scheduling and multi-tenant cluster management, training data storage and I/O, observability and workload optimization, and the training framework layer. You will lead one or two of these areas end-to-end and contribute across the rest; we do not expect one person to be an expert in all four. You will partner closely with the model training teams who are the platform’s customers, with the Cloud Infrastructure team that owns the clusters underneath, and with the Data Infrastructure team that produces the datasets the platform serves.

As an early member of the team, you will help define its technical direction, operating model and future hiring.


Responsibilities

Job scheduling and multi-tenant cluster management

  • Evolve our control plane into the single entry point for every accelerator cluster we operate, across cloud providers and silicon types: a user asks for N accelerators of a given type, not for a specific cluster.
  • Design scheduling that maximizes fleet goodput rather than queue order: gang scheduling for jobs of hundreds of nodes, multi-factor priority and fair sharing across teams, quota borrowing with enforceable reclaim, and execution-time-aware backfill.
  • Treat availability as an engineered system: continuous node health checking with automatic cordon, drain and replace, and automatic classification of every failed job (user, out-of-memory, communication, hardware, platform).

Training data storage and I/O

  • Design the storage tiering between object storage, shared or node-local caches and memory, and the request patterns that keep thousands of concurrent readers from overwhelming the object store.
  • Kill the small-file problem for good: shard formats and indexing that serve both fleet-scale shuffled random access and sequential scans, source data stored once and referenced everywhere, and the evaluation and adoption of a training-native storage format.
  • Make “GPU-hours lost to input wait” a first-class metric and drive it down on real jobs.

Observability and workload optimization

  • Build the monitoring and profiling toolchain at every layer a job touches — storage, node, GPU and interconnect, scheduler, and per-job metrics surfaced to the job owner — so that platform and users see the same picture.
  • Establish a workload taxonomy and per-type execution baselines (step time, utilization, communication fraction, input wait, checkpoint cost) that make regressions detectable and every optimization quantifiable.
  • Lay the groundwork for AI-assisted triage and optimization: metrics, logs and scheduler state accessible enough that an agent can be the first responder for failed jobs and propose improvements measured against the baselines.

Training framework and stack currency

  • Help build a thin, opinionated training framework layer over open-source distributed training libraries: one API for users, per-accelerator backends underneath, with golden images, validated launch recipes and a model-zoo CI that gates every release.
  • Run coordinated upgrade programs across the ML stack (distributed compute framework, Kubernetes operators, queueing, experiment tracking) on a cluster that is never idle.

Technical leadership

  • Define the platform roadmap with the team lead — build-versus-buy decisions, boundaries with the cluster, data and model teams, and prioritization by measured impact on goodput and cost.
  • Work directly with ML engineers to find where the platform slows them down or fails, and turn that into improvements to the scheduler, the data path, the tooling and the documentation.
  • Lead architecture across organizational boundaries, communicate recommendations to engineering leadership, and mentor the engineers building and operating the platform.

Qualifications

Required

  • 5+ years of software engineering experience, or equivalent demonstrated impact, with substantial distributed-systems work in production.
  • Staff-level technical leadership: you identify the problems worth solving, shape strategy across teams, make pragmatic trade-offs, and drive ambiguous initiatives from evidence to production.
  • Hands-on experience building or operating large-scale compute or ML infrastructure — and the failure modes that only show up at scale: hung collectives, retry storms, stranded capacity, silent input-bound jobs.
  • Deep expertise in at least one of the following areas, with working familiarity with the others:
  • Cluster scheduling and multi-tenancy — Kubernetes-based scheduling and resource management (Kueue, Volcano, Slurm, YARN or equivalent), including quota, fairness and preemption design.
  • Training data storage and I/O — object-store request behavior, caching tiers, shard formats, shuffle-versus-locality trade-offs, and measuring whether a job is input-bound.
  • Distributed training frameworks and accelerators — PyTorch distributed, Ray Train, JAX or equivalent, on GPU, TPU or Trainium: how process groups, collectives, sharding and checkpointing actually behave.
  • Observability and performance engineering — profiling and monitoring distributed workloads from the kernel and GPU up to the scheduler, and turning measurements into optimizations.
  • Strong programming skills in Python and experience with at least one additional relevant language such as Go, Rust or C++.
  • Experience operating on a major cloud provider (AWS preferred) and on Kubernetes.
  • Strong communication and developer empathy, with a track record of building platforms that engineers adopt and trust; self-directed in ambiguous problem spaces.

Bonus Points

  • Deep understanding of operating systems — memory management, I/O and network stack, scheduling, kernel-level debugging — as applied to debugging and optimizing distributed workloads.
  • Experience with large-model training and cross-GPU communication: collective communication, tensor/pipeline parallelism, NCCL performance at scale.
  • Experience with Ray and KubeRay internals, Kueue or Kubernetes scheduler extensions, or maintaining a patch set on top of an upstream project.
  • Experience with GPU profiling and performance tooling (torch profiler, Nsight Systems and Compute, DCGM, NCCL telemetry) or their counterparts on other accelerators.
  • Familiarity with training-native or columnar data formats (Lance, WebDataset, Parquet row groups) and with GPU-side video decode in the input path.
  • Experience with experiment tracking and model registry platforms (MLflow or equivalent) at scale.
  • Background in applying LLM agents to infrastructure operations, triage or performance tuning.

Pay Disclosure

The salary range for this role is $206,500-$258,100 for San Francisco Bay Area based applicants. This is the lowest to highest salary we in good faith believe we would pay for this role at the time of this posting. An employee’s position within the salary range will be based on several factors including, but not limited to, specific competencies, relevant education, qualifications, certifications, experience, skills, geographic location, shift, and organizational needs.

We offer a comprehensive package of benefits for full-time and part-time employees, their spouse or domestic partner, and children up to age 26, including but not limited to paid vacation, paid sick leave, and a competitive portfolio of insurance benefits including life, medical, dental, vision, short-term disability insurance, and long-term disability insurance to eligible employees. You may also have the opportunity to participate in Rivian’s 401(k) Plan and Employee Stock Purchase Program if you meet certain eligibility requirements. Full-time employee coverage is effective on their first day of employment. Part-time employee coverage is effective the first of the month following 90 days of employment. More information about benefits is available at rivianbenefits.com.

Equal Opportunity

Rivian is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, sex, sexual orientation, gender, gender expression, gender identity, genetic information or characteristics, physical or mental disability, marital/domestic partner status, age, military/veteran status, medical condition, or any other characteristic protected by law.

Rivian is committed to ensuring that our hiring process is accessible for persons with disabilities. If you have a disability or limitation, such as those covered by the Americans with Disabilities Act, that requires accommodations to assist you in the search and application process, please email us at View email address on us.fitly.work .

Candidate Data Privacy and Technology

Rivian may collect, use and disclose your personal information or personal data (within the meaning of the applicable data protection laws) when you apply for employment and/or participate in our recruitment processes (“Candidate Personal Data”). This data includes contact, demographic, communications, educational, professional, employment, social media/website, network/device, recruiting system usage/interaction, security and preference information. Rivian may use your Candidate Personal Data for the purposes of (i) tracking interactions with our recruiting system; (ii) carrying out, analyzing and improving our application and recruitment process, including assessing you and your application and conducting employment, background and reference checks; (iii) establishing an employment relationship or entering into an employment contract with you; (iv) complying with our legal, regulatory and corporate governance obligations; (v) recordkeeping; (vi) ensuring network and information security and preventing fraud; and (vii) as otherwise required or permitted by applicable law.

Rivian may share your Candidate Personal Data with (i) internal personnel who have a need to know such information in order to perform their duties, including individuals on our People Team, Finance, Legal, and the team(s) with the position(s) for which you are applying; (ii) Rivian affiliates; and (iii) Rivian’s service providers, including providers of background checks, staffing services, and cloud services.

Rivian may transfer or store internationally your Candidate Personal Data, including to or in the United States, Canada, the United Kingdom, and the European Union and in the cloud, and this data may be subject to the laws and accessible to the courts, law enforcement and national security authorities of such jurisdictions. 

How We Use AI in Our Hiring Process: To ensure transparency, we want candidates to know that Rivian uses iCIMS Talent Cloud Artificial Intelligence (TCAI) and AI-enabled tools to assist with screening, reviewing, organizing and highlighting profiles and applications that match the key requirements for each role.

AI does not make hiring decisions: Qualified candidate applications are reviewed by a member of our team, and all decisions throughout the process are made by humans. We use AI to support efficiency and consistency, not to replace human judgment. We are committed to a fair, thoughtful, and equitable experience for every candidate.

Participation in AI profile matching is entirely voluntary. If you prefer that your profile not be used in this process, you can opt out at any time. Opting out means your profile will be excluded from automated matching and will not be surfaced for additional roles through this system. Your current application remains active and will not be affected in any way.

Please note that we are currently not accepting applications from third party application services.

Vacancy posted 7 days ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer, ML Infra, Autonomy in California vacancy
  •  ...and we operate with real autonomy and almost no bureaucracy....  ...looking for "generalist" engineers who can operate at a Staff level: You're confident in...  ...originally designed. Use AI/ML to turn freeform doctors'...  .... Identifying the biggest infra cost opportunities and shipping... 
    Suggested
    Work at office
    Remote work
    Work from home
    Relocation
    Flexible hours

    Metriport Inc

    San Francisco, CA
    3 days ago
  •  ...systems. Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon...  ...is both exciting and crucial.  As a Staff Engineer, you will lead the technical delivery...  ...control through optimization and applied ML/RL  ~ Background in collaborative... 
    Suggested
    Full time
    Temporary work
    Part time
    Work experience placement
    Work at office
    Worldwide

    Shield AI

    San Diego, CA
    26 days ago
  • $170k - $220k

     ...Job Description Job Description Title:   Staff Software Engineer, Neuronavigation and Autonomy Functional Area: Software Engineering Department: R&...  ...runtime connecting live sensor and robot data to CV and ML inference: low-latency data flow, concurrency, and... 
    Suggested

    Magnus Medical

    Burlingame, CA
    14 days ago
  • $260k - $300k

     ...releases, we’d like to hear from you. About this role As a Staff Backend/Infra Engineer at Concurrence, you'll build the core services behind our...  ...-latency system design Experience with production AI or ML systems Benefits (available to Full-Time Employees)... 
    Suggested
    Full time
    Flexible hours

    Concurrence

    San Francisco, CA
    15 days ago
  • $253k - $336k

     ...is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion,...  ...problem you would own. ABOUT THE JOB Staff Robotics Engineers lead the delivery of vehicle...  ...delivery of a multi-year, multi-stakeholder software roadmap that spans across multiple... 
    Suggested
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Costa Mesa, CA
    2 days ago
  • $166.5k - $291.4k

     ...It all started when engineer Fred Luddy wrote code that...  ...set. This is not an ML research role; you do not...  ...traditional full-stack staff work, where systems...  ...tools are exposed, and how autonomy is delegated. Own the...  ...You Bring 7+ years of software engineering,with a... 
    Contract work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    SmartRecruiters

    Santa Clara, CA
    1 day ago
  • $144.7k - $221.4k

     ...verification. Partnering with Autonomy, Simulation, Systems, and...  ...results into clear feedback for engineering and leadership, and help accelerate...  ...autonomous driving software performance atinterfaces across...  ...develop new statistical and ML methods to quantify performance... 
    Full time
    Work at office
    Local area
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  •  ...operating systems, and autonomy. Eighteen of the top 20...  ...research effort by building ML tools, infrastructure,...  ..., we encourage all engineers to take ownership over...  ...generation self-driving software Help scale end-to-end...  ...at scale) ML infra fluency - has used and/... 
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    a month ago
  • Lead Data Engineer Make Your Mark: We're looking for a Lead Data Engineer...  ...and drive design reviews for ML pipelines, model registries,...  ...efficiency. Collaborate with software engineers to integrate machine...  ...or experience | High degree of autonomy and exercises independent... 

    BlackLine

    Los Angeles, CA
    3 days ago
  • $178k - $201k

    Staff Software Engineer, Applied AI Flock Freight is looking for a Staff Software Engineer, Applied AI...  ...AI within corporate systems with high autonomy and ownership. Copilot Experiences & Integrations...  ..., with at least 1 year in applied AI/ML or intelligent automation in production... 
    Full time
    Work at office
    Local area
    Work from home
    Relocation

    Flock Freight - Encinitas, CA, US

    Encinitas, CA
    3 days ago
  • $300k - $350k

     ...the best product managers, software, and hardware talent at innovative...  ...help them hire. Senior / Staff Software Engineer Location - San Mateo, CA (On...  ...AI Research, Agentic AI, ML Infrastructure, AI Research...  ...team operates with exceptional autonomy and iteration speed.... 
    H1b
    Work at office
    Remote work
    Flexible hours

    Recruiting from Scratch

    San Mateo, CA
    3 days ago
  • $215k - $270k

     ...POSITION Saildrone is seeking a Staff Machine Learning Engineer to join our team. Reporting...  ...systems that enable autonomy and real-time intelligence...  ...takes ownership of production ML systems in mission-critical...  ...than disrupt the overall software stack stability. Technical... 
    Local area
    Relocation package
    Flexible hours
    3 days per week

    Saildrone Inc

    Alameda, CA
    3 days ago
  •  ...civilians with intelligent systems. Its products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon simulation and...  ...Qualifications: ~ BS/MS/PhD in Computer Science, Aerospace Engineering, Electrical Engineering, Robotics, Mechanical Engineering,... 
    Full time
    Temporary work
    Part time
    Work at office
    Worldwide

    Shield AI

    San Mateo, CA
    28 days ago
  • $193.93k - $352.29k

     ...why we're building a universal autonomy platform: self-driving for...  ...About the Role The Autonomy ML Infrastructure team is...  ...compression.  Work with autonomy engineers to optimize, validate, and...  ...Write robust, high quality software to increase our confidence in... 
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    16 days ago
  • Mach Industries is recruiting an Autonomy Software Engineer to design, build, and deploy software powering perception, localization, navigation, planning...  ...product lines. You may come from estimation, perception, ML, embedded systems, or planning but will work across the... 

    Mach Industries

    Huntington Beach, CA
    1 day ago
  • $177k - $265.6k

     .... We're the front-end design engine, defining what comes next. From...  ...challenges in survivability, autonomy, and battle management. We...  ...is seeking a highly motivated Staff Software Engineer to join our team onsite...  ...work closely with other AI/ML and software engineers to... 
    Relocation package
    Flexible hours
    Shift work

    Northrop Grumman

    El Segundo, CA
    3 days ago
  • $193.93k - $352.29k

     ...re building a universal autonomy platform: self-driving...  ...autonomously inside Nuro's own engineering organization, under the...  ...You ~5+ years of software engineering experience...  ...practical experience. Staff-level candidates should...  .... ~ Experience with ML training or research... 
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    a month ago
  • $125k - $160k

     ...Mach Industries is building an AI‑forward autonomy stack for contested environments where...  ...unavailable, degraded, or denied. As an Autonomy Software Engineer, you will design, build, and deploy the...  ...with Python for tooling, analysis, and ML workflows. Build and integrate... 
    Permanent employment
    Work experience placement
    Work at office
    Local area
    Night shift

    Mach Industries

    Huntington Beach, CA
    1 day ago
  • $160k - $200k

    Senior Software Engineer Have you ever worked somewhere that challenges you every day? A place where...  ...Zendar: Zendar builds a radar-centric autonomy stack which makes any vehicle - from cars...  ...data collection requirements from our ML team by adapting the current platform with... 
    Work at office
    Flexible hours

    Zendar

    Berkeley, CA
    3 days ago
  • $206.5k - $258.1k

     ...generations.  Role Summary The Autonomy org at Rivian is seeking a Staff Software Engineer, Data Ops to join the Data team...  ...AWS Cloud Platform and Data/Dev/ML Ops practices....  ...reliable, scalable, and distributed infra using microservice architecture.... 
    Full time
    Contract work
    Temporary work
    Part time
    Local area
    Shift work

    Rivian

    California
    15 days ago
  • $193.93k - $352.29k

     ...why we're building a universal autonomy platform: self-driving for all...  ...We are looking for a Senior/Staff Software Engineer to serve as a technical leader for Nuro's ML Data engine. You will sit at the...  ...across autonomy teams and data infra teams to build effective ML data... 
    Immediate start
    Flexible hours
    Shift work

    Nuro

    San Francisco, CA
    11 days ago
  •  ...into the physical world. As a Software Engineer, you will set the technical...  ...work closely with engineers, ML researchers, product managers...  ...offs across multiple domains (infra, ML, product) Collaborate across...  .... We operate with high autonomy, collaborate closely across functions... 
    Immediate start

    Siemens

    Santa Clara, CA
    1 day ago
  • $260k - $290k

     ...building together. About The Role VSCO is hiring a Senior Staff Engineer to own Reflex, our production GPU image-processing engine, and...  ..., color management, LUTs, film emulation On-device or cloud ML in a creative-tools context (segmentation, style transfer)... 
    Temporary work
    Local area
    Worldwide
    Flexible hours

    vsco39

    San Francisco, CA
    1 day ago
  • $170.5k - $193.5k

     ...delivering best-in-class financing and software products for sustainable solutions, from...  .... We are seeking a highly skilled Staff Software Engineer to bring technical excellence, initiative...  ..., and secure solutions to support AI/ML-powered features across Web, Mobile, SMS... 

    Socket

    San Francisco, CA
    3 days ago
  • $206k - $258k

     ...diverse, but our team shares a love of the outdoors and a desire to protect it for future generations. Role Summary As a Software Engineer specializing in safety-critical self-driving middleware, you will play a vital role in the design, development, and deployment... 
    Full time
    Contract work
    Local area

    Rivian

    Palo Alto, CA
    2 days ago
  • $176.1k - $308.2k

     ...Company Description It all started when engineer Fred Luddy wrote code that automated a...  ...card, will be considered. As a Staff Software Engineer, you will independently lead...  ...fundamentals with hands-on expertise in AI/ML integration, cloud architecture, and... 
    Permanent employment
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours
    2 days per week

    ServiceNow

    Santa Clara, CA
    1 day ago
  •  ...The Staff Software Engineer will lead the design and delivery of complex AI-native and cloud-native systems while providing technical mentorship...  ...platform performance and ensure the successful integration of AI/ML services within the ServiceNow ecosystem. Requirements... 
    Flexible hours

    engineeringjobs.net, Inc.

    Santa Clara, CA
    2 days ago
  •  ...they already use to move and process events at scale. As an engineer, you'll own delivery of significant pieces of this product — not...  ...or streaming data systems. You don't need a background in ML research or model training — this role is about building and operating... 
    Live in

    Confluent

    Mountain View, CA
    2 days ago
  •  ...love to meet you. About the Role Plenful is hiring a Staff Software Engineer to lead the design and development of systems that power our...  ...of our entire platform. You'll work closely with DevOps, ML/AI, Customer Success, and UX teams, and mentor engineers while... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful Inc.

    San Francisco, CA
    2 days ago
  • $75k - $100k

     ...we’re reimagining how software gets built. Our vision...  ...breakthroughs in AI, systems engineering, and product design....  ...efforts with high autonomy and ownership What You’...  ...distributed systems, cloud infra, and high-performance services...  ..., developer tools, or ML systems Even if you don... 
    Full time
    Worldwide

    Emergent

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer, ML Infra, Autonomy. Be the first to apply!