Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten

A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating with cross-functional teams to adapt models efficiently. Ideal candidates have over 5 years of experience in ML infrastructure and a strong background in distributed training frameworks. Competitive compensation and benefits package is offered.#J-18808-Ljbffr

Vacancy posted 3 hours ago
Similar jobs that could be interesting for youBased on the Senior ML Training Systems Engineer - Distributed GPU Infra in San Francisco, CA vacancy
  • $227.2k - $417k

     ...Role:As a Software Engineer on the ML Infrastructure...  ...ML model serving systems that support Deep...  ...and a mentor to senior engineers, fostering...  ...and low latency distributed systems using ScalaBuild...  ...of our infra. Lead large scale...  ...ElastiCache, model training orchestration, etc... 
    Training
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    1 day ago
  • $250k

     ...next-generation GPU platform designed for AI training, experimentation...  ...is looking for a Senior / Staff Site Reliability Engineer to support and...  ...with platform, ML, and infrastructure...  ...across distributed compute environments...  ...infrastructure systems Improve CI/CD... 
    Senior
    Training
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $250k

    A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines. The role involves building scalable infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include... 
    Senior

    Acceler8 Talent

    San Francisco, CA
    3 hours ago
  • $117.2k - $313.7k

     ...opportunities for Lead software engineers who want their lines...  .../frameworks in distributed filesystems in an ever...  ...that improve system scalability, robustness...  ...Experience with Big-Data/ML and S3 Hands-on experience...  ...promotion, benefits, training, assessment of job performance... 
    Senior
    Training
    Full time
    Immediate start
    Remote work

    Salesforce

    San Francisco, CA
    2 days ago
  •  ...—like the da Vinci surgical system and Ion—have transformed how...  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics,...  ...alongside research, SW/ HW/ ML engineering, regulatory, controls... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    1 day ago
  • $205.9k - $407.5k

     ...are looking to bring on a Senior Principal ML GPU Architect to lead the ML GPU...  ...with the Director of ML Engineering who is responsible for our...  ...step function changes in training and inference speed/scale...  ...inference code for large, distributed training/inference with FP... 
    Senior
    Training
    Temporary work

    Adobe

    San Francisco, CA
    4 hours ago
  • $166k - $225k

     ...improve their business. Founded by engineers — and customer obsessed — we leap...  ...be building the next generation distributed data storage and processing systems that can outperform specialized...  ...experience, relevant certifications and training, and specific work location.... 
    Senior
    Training
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  •  ...and help build the platform engineers turn to to ship AI products....  ...building the global operating system for distributed, heterogeneous AI hardware....  ...engineers to lead our GPU Networking efforts, making RDMA...  ...~ Exposure to a variety of ML startups, offering unparalleled... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    23 hours ago
  • $179k - $218k

     ...bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to...  ...Telemetry: Leverage AI/ML methodologies to analyze...  ...before they impact customer training runs.Technical Sparing...  ...Cause Analysis (RCA) on systemic issues that span the... 
    Senior
    Training
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...for an exceptional Staff Software Engineer to help define and build the distributed backend that powers our AI agent...  ...patterns, and build the foundational systems that will support millions of...  ...systems. Experience with LLMOps or ML infrastructure is highly desirable... 
    Senior
    Immediate start
    Remote work
    Flexible hours

    Workato

    San Francisco, CA
    3 hours ago
  • $148.5k - $223.9k

     ...Salesforce is seeking a senior engineering candidate to join the...  ...proactively design systems that prevent them,...  ....Understanding of AI/ML concepts applied to operations...  ....Strong knowledge of distributed systems and Linux/...  ...promotion, benefits, training, assessment of job... 
    Senior
    Training
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    23 hours ago
  • $160k - $194k

     ...The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in...  ...Demonstrated track record for Senior or Staff level software...  ...network architectures and AI/ML landscape is a big plus...  ...conditions, experience and training, licensure and... 
    Training
    Full time
    Remote work
    Flexible hours

    Ad Verse

    San Francisco, CA
    23 hours ago
  •  ...building production-grade ML infrastructure used by enterprise...  .... They are looking for a Senior AI/ML Engineer to own model training pipelines, evaluation systems, and inference serving at...  ...~ Experience with distributed training, GPU optimization, or inference... 
    Senior
    Training
    Full time

    Clera

    San Francisco, CA
    23 hours ago
  • $300k

     ...full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer...  ...of this GPU-powered infrastructure...  ...workloads, automate systems at petascale, and...  ...Collaborate with ML, networking, and platform...  ...engineering, distributed systems, or... 
    Senior
    Training
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  •  ...Scale AI is hiring for a senior software engineer to design, build, and scale full‑stack systems powering our GenAI data engine. You will work across front-end, back-end...  ..., and Temporal. You’ll collaborate with ML teams and product to deliver high-quality, scalable... 
    Senior

    Scale

    San Francisco, CA
    4 hours ago
  • $165k - $200k

     ...with us at Crusoe.About This RoleAs a Senior Systems Engineer, you’ll play a key role in building...  ...broader infrastructure needsMentoring and training Service Desk team members, developing...  ...off-hours syncs to support distributed operationsWhat You’ll Bring to the Team... 
    Senior
    Training
    Temporary work

    Crusoe

    San Francisco, CA
    23 hours ago
  • $190k - $230k

     ...About This RoleWe’re seeking a Senior Systems Engineer to play a key role in...  ...teams through code reviews, training, and technical enablement programsEvaluating...  ..., including 3+ years in AI/ML or AI application...  ...designing scalable, distributed systems in cloud environments... 
    Senior
    Training
    Temporary work

    Crusoe

    San Francisco, CA
    23 hours ago
  • $160k - $225k

     ...improve their business.Training and customizing...  ...for large-scale GPU training and fine...  ...the world.As a Senior Software Engineer for AI Runtime,...  ...and scaling the systems that make large-scale...  ...and capacity, distributed training performance...  ...computing, or ML systems.Experience... 
    Senior
    Training
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $200.8k - $251k

     ...member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    3 hours ago
  • $160k - $225k

     ...Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in...  ...instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and...  ...5+ years of experience in distributed systems, proficiency with GPU training... 
    Senior
    Training

    Cacheflow

    San Francisco, CA
    3 hours ago
  • $110 per hour

     ...Position: MLOps Engineer (JAX, PyTorch, Pallas...  ...performance in MLOps , training infrastructure, and ML framework-level...  ...to MLOps and ML systems problems . Evaluate...  ...pipeline design, distributed systems reasoning, and...  ...optimizing custom GPU kernels using Pallas... 
    Training
    Remote job
    Contract work
    Summer work
    Weekday work

    Mercor

    San Francisco, CA
    2 days ago
  • $219k - $315k

     ...experienced software engineer to work on large-scale...  ...vehicle. These are data and GPU intensive workloads...  ...of keeping production systems running with high...  ...behaviorsImprove the ML training pipelines supporting the...  ...optimizing large-scale distributed systems for cost and efficiencyExperience... 
    Senior
    Training
    Full time
    Temporary work
    Relocation package

    Zoox

    San Francisco, CA
    2 days ago
  • $229.9k - $262.4k

     ...Overview Senior Lead Software Engineer, Distributed Systems (Golang + Python on Kubernetes) Do you love building and pioneering in the technology space?...  ...committed to pioneering and responsibly implementing AI/ML across Capital One . We achieve this by building... 
    Senior
    Full time
    Part time
    Internship
    Local area

    Capital One

    San Francisco, CA
    a month ago
  • $255k - $405k

     ...Lambda is the #1 GPU Cloud for ML/AI teams training, fine-tuning and inferencing AI models, where engineers can easily, securely and affordably...  ...portfolio includes on-prem GPU systems, hosted GPUs across public...  ...for large-scale, distributed systems. Familiarity with... 
    Senior
    Training
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    23 hours ago
  • $220k - $320k

     ...techniques into production systems, we'd love to meet...  ...net Inference.net trains and hosts...  ...ten‑person team of engineers who work in‑person...  ...CUDA kernels and GPU utilization across...  ...Collaborate with applied ML engineers to...  ...) Experience with distributed inference and... 
    Senior
    Training
    Work at office

    SOLANA FOUNDATION

    San Francisco, CA
    3 days ago
  •  ...Paradigm is seeking a Senior Software Engineer in San Francisco, California. This role requires 7+ years of...  ..., particularly in managing and deploying distributed services. Candidates should be skilled in debugging complex systems and should possess familiarity with tools... 
    Senior

    Paradigm

    San Francisco, CA
    11 hours ago
  • $250k - $300k

     ...experience in storage engineering, including operating distributed storage at multi-...  ...performance storage for GPU or HPC clusters...  ...build production-grade systems and tooling We...  ...have familiarity with ML/AI storage patterns...  ...wide throughput for training and inference workloads... 
    Training
    Full time
    Remote work

    Together AI

    San Francisco, CA
    2 days ago
  • $120k - $200k

     ...Senior Infrastructure EngineerAt Bland.com...  ...Senior Infrastructure Engineer at Bland, you'll...  ...'re architecting distributed systems that handle real-time...  ...processing, scale ML inference, and...  ...our AI models, from training pipelines to real-...  ...model serving, or GPU computing.... 
    Senior
    Training
    Work at office
    Night shift

    Bland AI

    San Francisco, CA
    2 days ago
  • $204k - $300k

     ...science and electrical engineering, such as AI/ML, algorithms, digital...  ...& analytics, distributed systems, cloud, edge & mobile...  ...Will AccomplishAs a senior research leader in the...  ...Desired)• Experience with GPU/CPU architectures,...  ...education or training. Your recruiter can... 
    Senior
    Training
    Full time
    Local area
    Worldwide
    Flexible hours

    Dolby

    San Francisco, CA
    3 days ago
  • $225k

     ...next-generation GPU platforms designed...  ...for large-scale AI training, experimentation,...  ...is looking for a Senior Network Engineer to design, deploy...  ...issues in distributed HPC environments...  ...infrastructure, platform, and ML engineering teams...  ...) Strong Linux systems knowledge and... 
    Senior
    Training
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior ML Training Systems Engineer - Distributed GPU Infra. Be the first to apply!