Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer, ML Training and Inference Infrastructure

$228k - $285k
Full-time

Rivian

About Rivian

Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. 

As a company, we constantly challenge what’s possible, never simply accepting what has always been done. We reframe old problems, seek new solutions and operate comfortably in areas that are unknown. Our backgrounds are diverse, but our team shares a love of the outdoors and a desire to protect it for future generations. 


Role Summary

As a Staff Software Engineer, ML training and inference infrastructure , you will be a member of the Perception team at Rivian, which develops advanced machine learning algorithms that directly impact safety critical self-driving features of our category defining vehicles.

We are looking for candidates with deep knowledge and strong enthusiasm towards establishing a state-of-art ML infrastructure for training and inference of large autonomous driving models; and optimizing the training and inference performance.


Responsibilities

  • Optimize the performance of Deep Learning training workload on NVIDIA GPU systems on a large scale
  • Optimize the latency of model inference and model pre- and post-processing on onboard systems
  • Design, train, and deploy large deep learning models that can leverage the vast amount of labeled and unlabeled data

Qualifications

  • PhD in CS/CE/EE, or equivalent, in industry experience
  • Deep knowledge of PyTorch
  • Knowledge of model training framework (e.g. PyTorch Lightning, ray, etc.)
  • In-depth knowledge of transformer architecture and ways to accelerate the training and inference of transformer models
  • Experience of performing large scale distributed training of models
  • A track record of profiling models and doing detective work to improve model training and inference speed

Preferred Skill Requirements:

  • Experience with CUDA or Triton language for writing custom ops
  • Knowledge of Nvidia TensorRT
  • Knowledge of NCCL
  • Experience with edge computing systems
  • A track record of efficiently solving complex problems collaboratively on larger teams

Pay Disclosure

Salary Range for California Based Applicants: $228,000.00 - $285,000.00 (actual compensation will be determined based on experience, location, and other factors permitted by law). 

Benefits Summary : Rivian provides robust medical/Rx, dental and vision insurance packages for full-time employees, their spouse or domestic partner, and children up to age 26. Coverage is effective on the first day of employment 

Equal Opportunity

Rivian is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, sex, sexual orientation, gender, gender expression, gender identity, genetic information or characteristics, physical or mental disability, marital/domestic partner status, age, military/veteran status, medical condition, or any other characteristic protected by law.

Rivian is committed to ensuring that our hiring process is accessible for persons with disabilities. If you have a disability or limitation, such as those covered by the Americans with Disabilities Act, that requires accommodations to assist you in the search and application process, please email us at  View email address on ev.careers .

Candidate Data Privacy and Technology

Rivian may collect, use and disclose your personal information or personal data (within the meaning of the applicable data protection laws) when you apply for employment and/or participate in our recruitment processes (“Candidate Personal Data”). This data includes contact, demographic, communications, educational, professional, employment, social media/website, network/device, recruiting system usage/interaction, security and preference information. Rivian may use your Candidate Personal Data for the purposes of (i) tracking interactions with our recruiting system; (ii) carrying out, analyzing and improving our application and recruitment process, including assessing you and your application and conducting employment, background and reference checks; (iii) establishing an employment relationship or entering into an employment contract with you; (iv) complying with our legal, regulatory and corporate governance obligations; (v) recordkeeping; (vi) ensuring network and information security and preventing fraud; and (vii) as otherwise required or permitted by applicable law.

Rivian may share your Candidate Personal Data with (i) internal personnel who have a need to know such information in order to perform their duties, including individuals on our People Team, Finance, Legal, and the team(s) with the position(s) for which you are applying; (ii) Rivian affiliates; and (iii) Rivian’s service providers, including providers of background checks, staffing services, and cloud services.

Rivian may transfer or store internationally your Candidate Personal Data, including to or in the United States, Canada, the United Kingdom, and the European Union and in the cloud, and this data may be subject to the laws and accessible to the courts, law enforcement and national security authorities of such jurisdictions. 

How We Use AI in Our Hiring Process: To ensure transparency, we want candidates to know that Rivian uses iCIMS Talent Cloud Artificial Intelligence (TCAI) and AI-enabled tools to assist with screening, reviewing, organizing and highlighting profiles and applications that match the key requirements for each role.

AI does not make hiring decisions: Qualified candidate applications are reviewed by a member of our team, and all decisions throughout the process are made by humans. We use AI to support efficiency and consistency, not to replace human judgment. We are committed to a fair, thoughtful, and equitable experience for every candidate.

Participation in AI profile matching is entirely voluntary. If you prefer that your profile not be used in this process, you can opt out at any time. Opting out means your profile will be excluded from automated matching and will not be surfaced for additional roles through this system. Your current application remains active and will not be affected in any way.

Please note that we are currently not accepting applications from third party application services.

Vacancy posted 24 days ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer, ML Training and Inference Infrastructure in California vacancy
  • $228k - $285k

     ...shares a love of the outdoors and a desire to protect it for future generations. Role Summary As a Staff Software Engineer, ML training and inference infrastructure, you will be a member of the Perception team at Rivian, which develops advanced machine learning... 
    Training
    Full time
    Contract work
    Local area

    Rivian

    Palo Alto, CA
    13 hours ago
  • $190.9k - $232.8k

     ...1285About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture...  ..., distributed inference infrastructure - orchestrate across...  ...understanding of ML inference internals: attention...  ...certifications and training, and specific work location... 
    Training
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    8 hours ago
  • $198k - $326k

     ...LinkedIn's AI model training, feature engineering and serving with...  ...infra, compute software, and hardware to...  ....Model Training Infrastructure: As an engineer...  ...GPU based inference for a large variety...  ...hundreds of new ML models per quarter...  ...scale.As a Sr. Staff Software Engineer... 
    Training
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Sunnyvale, CA
    4 days ago
  •  ...leader in AI cloud infrastructure serving tens of thousands...  ...groundbreaking AI training and inference possible.The Lambda Infrastructure Engineering organization forges...  ...seeking a seasoned Staff Storage Software Engineer with deep...  ...in HPC, AI/ML infrastructure, or... 
    Training
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $245.4k - $429.45k

     ...our recruiting process here.Sr. Staff Software Engineer, Product ML InfrastructureMillions of...  ...with AI.Pinterest’s Product ML Infrastructure (PMLI) team enables fast, safe,...  ...products. We build unified data, training, feature, and inference infrastructure; this role will... 
    Training
    Work at office
    Local area
    Remote work
    Relocation
    Relocation package

    Pinterest

    Palo Alto, CA
    1 day ago
  • $190k - $270k

    Staff Software Engineer - AI Research InfrastructureP-1215At Databricks...  ...ranging from post-training open source LLMs to...  ..., AI Research Infrastructure, you will be developing...  ...‑scale training and inference experiment workloads...  ...scientists, ML engineers, and platform... 
    Training
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • $193.93k - $352.29k

     ...fungible is the infrastructure that decides...  ...inside Nuro's own engineering organization,...  ...against the training pipelines...  ...You5+ years of software engineering experience...  ...experience. Staff-level...  ...under the hood at inference. Attention and...  ...Experience with ML training or research... 
    Training
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    3 days ago
  •  ...physical world. Critical infrastructure is constrained by labor...  ...the Role At Watney, ML Infrastructure engineers turn data collected from a...  ...systems will require larger training runs with more data, expanded...  ...ll Do Own training and inference infrastructure Build... 
    Training

    Watney

    San Francisco, CA
    a month ago
  • $210k - $300k

     ...economical commerce. We're training robot AGI to power a...  ...team of the world's best engineers and operators. If you are...  ...We're looking for a Software Engineer to join our ML Infrastructure team. In this role, you'...  ...build the training and inference systems that power our general... 
    Training
    Local area
    Flexible hours

    Nimble Robotics

    San Francisco, CA
    21 days ago
  • $207k - $300k

     ...aspects of machine learning infrastructure and modeling. Change infrastructure (training and serving) to support...  ...testing, and launching software products.5 years of experience leading ML design and optimizing ML...  ...’s degree or PhD in Engineering, Computer Science, or a... 
    Training
    Immediate start

    Google

    Mountain View, CA
    8 hours ago
  • $262k - $364k

     ...architecture of highly complex ML pipelines and...  ...senior and mid-level engineers, fostering a culture of...  ...years of experience in software development.7 years of...  ..., and working with ML infrastructure (e.g., model...  ...relevant education or training. US: $262000 - $364000... 
    Training

    Google

    Mountain View, CA
    1 day ago
  •  ...deliver industry-leading training and inference speeds; over 10 times faster...  ...AI.We’re looking for a Software Engineer focused on Observability to...  ...logging, tracing, and alerting infrastructure that enables fast...  ...performance computing, AI/ML systems, or inference platformsHardware... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $254k - $350k

     ...Remote (United States)Software - Software...  ...this mission. The ML Platform team at...  ...techniques in distributed training, quantization,...  ...strong software engineers and act as a...  ...more about our ML Infrastructure, here is one of our...  ...ML Training OR Inference performance optimization... 
    Training
    Full time
    Remote work

    Zoox

    San Mateo, CA
    5 days ago
  • $235.7k - $277k

     ...a separate system to run inference or build an agent, our customers...  ...events at scale.As an engineer, you'll own delivery of...  .... This is production infrastructure serving live inference, so...  ...don't need a background in ML research or model training — this role is about... 
    Training
    Live in

    Confluent

    Mountain View, CA
    2 days ago
  • $245k - $307k

     ...for a Lead Data Engineer to design, build,...  ...performance data infrastructure while collaborating...  ...reviews for ML pipelines, model...  ...Collaborate with software engineers to integrate...  ...maintain APIs for model inference. Infrastructure...  ...and manage training infrastructure including... 
    Training
    Full time
    Temporary work
    Work at office
    Shift work
    3 days per week

    BlackLine

    Pleasanton, CA
    2 days ago
  • $229k - $343k

     ...We’re looking for a Staff Software Engineer to join Snap Inc on...  ...for large-scale batch inference and low-latency...  ...storage patterns for training-scale reads and low-...  ...microservices, cloud infrastructure and/or platform architectureProficiency...  ...building or scaling ML Infrastructure... 
    Training
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    2 days ago
  • $192k - $260k

     ...best data and AI infrastructure platform so our customers...  ...frontier AI model inference for open source...  ...role, no prior ML or AI experience is...  ...We’re looking for engineers who have owned...  ...runtimes at scale.As a Staff Engineer, you’ll...  ...certifications and training, and specific work... 
    Training
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    8 hours ago
  •  ...deliver industry-leading training and inference speeds; over 10 times faster...  ..., and optimize cloud infrastructure supporting CI workloads, balancing...  ...developer velocity and engineering productivity.Identify...  ...professional experience in software engineering, infrastructure... 
    Training

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $165.2k - $223.6k

     ...builds AWS Neuron, the software development kit...  ...and Trainium ML accelerators. This...  ...enabling unparalleled ML inference and training performance....  ...software boundary, our engineers build systematic infrastructure, innovate new...  ...supervisors, and staff; adhere to standards... 
    Training
    Full time
    Work experience placement
    Internship
    Local area
    Flexible hours

    Annapurna Labs (U.S.) Inc.

    Cupertino, CA
    1 hour ago
  • $262k - $364k

     ...predictions.Optimize ML inference and resolve system-level...  ...of experience in software development.7 years of...  ...with industry-scale ML infrastructure (e.g., model...  ...Master’s degree or PhD in Engineering, Computer Science, or...  ...relevant education or training. US: $262000 - $36400... 
    Training

    Google

    Mountain View, CA
    1 day ago
  • $227.2k - $324.5k

    About the Role:As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine...  ...world-class machine learning inference platforms. These platforms power...  ...(e.g. Feast), ElastiCache, model training orchestration, etc.Understanding... 
    Training
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    8 hours ago
  • $193.93k - $352.29k

     ...looking for a Senior/Staff Software Engineer to serve as a...  ...technical leader for Nuro’s ML Data engine. You...  ...Machine Learning, and Infrastructure, acting as an architect...  ...into high-value training signals for autonomy...  ...compute embeddings or run inference at scale, manage... 
    Training
    Immediate start
    Flexible hours
    Shift work

    Nuro

    Mountain View, CA
    8 hours ago
  • $205k - $250k

     ...Backend Engineer At 3Y Health, we are developing an AI business...  ..., secure, and scalable infrastructure. Ideally, you've worked on...  ...external AI APIs, managing ML inference pipelines, or supporting data infrastructure for model training. You're not just a code contributor... 
    Training
    Work experience placement
    Private practice
    Work at office

    3Y

    San Francisco, CA
    4 days ago
  •  ...is a financial infrastructure platform for...  ...to innovate in ML Platform at Stripe...  ...enable ML engineers and data scientists...  ...spans ML training infrastructure...  ...products. As a Staff Engineer, you'...  ...latency model inference, large-scale...  ...operating great software systems.Who... 
    Training
    Flexible hours

    Stripe

    San Francisco, CA
    2 days ago
  •  ...builds the shared infrastructure that helps DoorDash...  ...high-throughput batch inference, and fine-tuning on...  ...and inference engines, fine-tuning and training pipelines, GPU autoscaling...  ...bar for — ML engineers, product...  ...industry experience in software engineeringDeep backend... 
    Training
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    2 days ago
  • $278.53k - $345.04k

     ...shared experiences for everyone.ML Platform @ Roblox today...  ...ML use cases and billions of inferences per day across Discovery, Safety, Engine, and much more. As an Infrastructure Engineer on the ML Platform...  ...lifecycle such as model serving, training, model CI/CD, and GPU... 
    Training
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    3 days ago
  • $180k - $300k

     ...But a large portion of training compute is wasted...  ...using far less compute at inference time, substantially reducing...  ...research and data engineering necessary to solve...  ...looking for an experienced Infrastructure Engineer to join as a...  ...~ Experience building ML/DL infrastructure and/... 
    Training
    Work at office
    Work from home
    Relocation package

    DatologyAI

    San Mateo, CA
    1 day ago
  • $175k - $287k

     ...one of the largest privately managed compute infrastructures in the world outside the public cloud providers. As a Staff Software Engineer on the Compute Infrastructure team, you...  ...pipelines processing petabytes daily, and the AI/ML infrastructure driving recommendations,... 
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    2 days ago
  • $193.93k - $352.29k

     ...investors.About the RoleThe Autonomy ML Infrastructure team is responsible for building & improving...  ...model compression. Work with autonomy engineers to optimize, validate, and deploy...  ...framework, FTL.Write robust, high quality software to increase our confidence in our... 
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    3 days ago
  •  ...looking for a systems-minded engineer who lives at the intersection of large-scale model inference, distributed systems, and...  ...This role focuses on post-training and inference infrastructure, with particular emphasis...  ...-engineering modern ML infrastructure, reasoning... 
    Training

    AMD

    San Jose, CA
    8 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer, ML Training and Inference Infrastructure. Be the first to apply!