Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Performance Engineer — Scale Distributed Training & Throughput

applied

Applied Intuition, Inc. is seeking a performance engineer to accelerate large-scale ML workloads in the data center. You will own profiling, optimization, and cost-efficiency for distributed training and large offline inferences.

You will work across accelerators, ML frameworks, and data infra, partnering with teams to land improvements that shorten time-to-result and reduce data processing costs. Collaboration and technical excellence are valued.

#J-18808-Ljbffr
Vacancy posted 21 hours ago
Similar jobs that could be interesting for youBased on the ML Performance Engineer — Scale Distributed Training & Throughput in Sunnyvale, CA vacancy
  •  ...Intuition in Sunnyvale is seeking a Performance Engineer to accelerate large-scale ML workloads in data centers. You will optimize distributed training across many nodes and improve batch...  ...-scale sensor logs, targeting throughput and cost-per-data processed. You... 
    Training
    Performance

    Applied Intuition

    Sunnyvale, CA
    21 hours ago
  •  ...Applied Intuition, Inc. in Sunnyvale, CA, is seeking a performance engineer to optimize large-scale ML workloads in the datacenter. This role focuses on distributed training across many nodes and high-throughput batch inference over petabytes of real-world autonomy logs... 
    Training
    Performance

    NLP PEOPLE

    Sunnyvale, CA
    21 hours ago
  • $153.2k - $234.1k

     ...transportation on a global scale.Role Overview:Are...  ...machine learning engineer working on our...  ...the safety and performance of the car, rather...  ...vehicles.As a Senior ML Infra Engineer,...  ...machine learning model training and evaluation...  ...large-scale distributed systems/applications... 
    Training
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  •  ...Intuition, Inc. is a Silicon Valley leader powering the future of physical AI. We seek a Performance Engineer to optimize large-scale ML workloads, focusing on distributed training, batch inference, and cost-effective data processing. You will own profiling across the stack... 
    Training
    Performance

    Decisive Point

    Sunnyvale, CA
    21 hours ago
  • $189.3k - $290.7k

     ...transportation on a global scale. Role:Are you...  ...scenarios.As a Staff ML Infra Engineer, you will drive the development...  ...dataset generation, training, evaluation, and...  ...pipelines that are performant, easy to use, and exceptionally...  ...building large-scale distributed systems, applications... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    16 hours ago
  •  ...is building an AI training and model post-training...  ...for large-scale training, RL experiments...  ...intersection of distributed systems, GPU performance, and ML framework...  ...Python and PyTorch engineering skills, hands-on experience...  ...to optimize throughput and memory across... 
    Training
    Performance

    Nebius B.V.

    Palo Alto, CA
    4 days ago
  •  ...in Mountain View is seeking a Staff / Principal ML Training Systems Engineer to lead the performance of large-scale multimodal training systems. This role involves...  ...will have proven experience in boosting distributed training performance and hands-on skills with modern... 
    Training
    Performance

    Rhoda AI

    Mountain View, CA
    16 hours ago
  • $150k

     ...edge foundation model training, alongside world-...  ...data scientists, and engineers, tackling the most fundamental...  ...global hub for high-performance computing in deep...  .... The Role The Distributed ML Engineer will play a...  ...methodologies, and large-scale machine learning... 
    Training
    Performance
    Full time
    Work experience placement
    Visa sponsorship

    Institute Of Foundation Models

    Sunnyvale, CA
    16 hours ago
  •  ...Principal Machine Learning Engineer to join our Models...  ...by the challenge of distributed training of large models on a...  ...generative AI at scale.THE PERSON:The ideal...  ...end training pipeline performance.Optimize the distributed...  ...EXPERIENCE:Experience with ML/DL frameworks such as... 
    Training
    Performance

    AMD

    San Jose, CA
    4 days ago
  • $189k - $300k

     ...transportation on a global scale. The Data Scaling team...  ...works on and delivers ML models to the product...  ...directly impacting AV product performance through smart use of...  ...foundation model pre-training and fine-tuning with...  ...high-impact team of AI/ML engineers, data scientists and... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  •  ...Santa Clara is seeking a skilled software engineer with over 8 years of experience to...  ...applications. The role involves designing distributed training strategies, collaborating with ML researchers, and developing tools for performance enhancement. Ideal candidates have a... 
    Training
    Performance

    Odyssey

    Santa Clara, CA
    16 hours ago
  •  ...seeking a Machine Learning Engineer in Palo Alto,...  ...will play a key role in scaling our Ray and PyTorch-based...  ...extensive experience with distributed systems, strong proficiency...  ...the efficiency and performance of our data processing and model training pipelines. #J-18808-Ljbffr... 
    Training
    Performance

    Orbifold AI

    Palo Alto, CA
    4 days ago
  • $165.2k - $223.6k

     ...Software Development Engineer to own the...  ...developing high-performance compute kernels,...  ...operate at frontier scale with large distributed models.This is a...  ...for a custom ML accelerator architecture...  ...latency and throughput improvements...  ...transformer architecture, training/inference... 
    Training
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $122.6k - $185k

     ...cloud for AI training and inference?...  ...continuous price performance improvements...  ...scalability in AI/ML and HPC...  ...those at cloud scale? If yes, then...  ...AWS Hardware Engineering team creates server...  ..., high throughput systems and knowledge...  ...of complex distributed systemsAmazon... 
    Training
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    16 hours ago
  • $250k - $350k

     ...Staff level Inference Engineers to accelerate the performance of Pika's AI-driven...  ...experiences at scale.You will design and...  ...computing kernels and distributed workloads using...  ...production.Improve Training Efficiency: (Bonus)...  ...HaveExperience with high-throughput video or real-time... 
    Training
    Performance
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    3 days ago
  • $195k - $230k

     ...challenges at scale.Together, we reached...  ...Learning Engineer to help evolve...  ...iterate on high-throughput, low-latency recommendation...  ...from offline training online...  ...drift, and system performance in production....  ...scale data and ML systems (e.g., Spark, distributed training, real-... 
    Training
    Performance
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    2 days ago
  • $155.42k - $395.9k

     ...DescriptionAbout the Team:The ML Compute Platform is...  ...supports the training and deployment of...  ...with a focus on performance, availability, concurrency...  ...a Senior Software Engineer to join our team and help us scale our platform for...  ...background in distributed systems, infrastructure... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...infrastructure company in California seeks a Member of Technical Staff — Training to design and optimize large-scale distributed training systems for frontier AI models. Candidates should have 5+ years of experience in ML systems and be proficient in Python along with another... 
    Training

    RadixArk

    Palo Alto, CA
    3 days ago
  • $165.2k - $223.6k

     ...forefront of maximizing performance for AWS's custom ML accelerators....  ...software boundary, our engineers craft high-...  ...unparalleled ML inference and training performance.As part...  ...computing, and distributed architectures,...  ...patterns, reliability and scaling) of new and... 
    Training
    Performance
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $207k - $300k

     ...through advanced context engineering and agentic feedback...  ...to improve the performance of the AIGC stack.Resolve...  ...at massive scale, and extend well beyond...  ...information retrieval, distributed computing, large-scale...  ...relevant education or training. US: $207000 - $30000... 
    Training
    Performance

    Google

    Mountain View, CA
    1 day ago
  • $124k - $250k

     ...of our software engineering infra team, you'll...  ...team builds a high-performance, high availability, globally distributed ecosystem...  ...infrastructure with high throughput and low latency...  ...delivery, including training, serving, and...  ...and maintain large-scale distributed... 
    Training
    Performance

    AppLovin

    Palo Alto, CA
    2 days ago
  •  ...We’re a team of engineers, clinicians,...  ...helps care teams perform with greater precision...  ...on large-scale image data....  ...and implement AI/ML approaches to extract...  ..., annotate, train, and test on...  ...Experience with distributed training and cloud...  ...inference, GPU/throughput optimization (e... 
    Training
    Performance
    Work at office
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    4 days ago
  • $296.3k

     ...seeking a Principal AI Engineer to lead the...  ...that powers large-scale training and cloud...  ...accelerating training throughput, scaling multi-modal...  ...across distributed training, training...  ...optimize core AI/ML platform infrastructure...  ...scalability, and performance across the AI/ML... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $189.4k - $300.6k

     ...transportation on a global scale. Are you passionate...  ...development. We engineer high-performance tools that identify top...  ...with data-intensive ML teams to drive rapid...  ...back to data selection, training, and launch decisions...  ..., Node.js, high-throughput data streamingData/Infra... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    16 hours ago
  •  ...OverviewAs our Staff Software Engineer, ML infra Engineer for...  ...data needed to train complex ML models and...  ...pipelinesDevelop and scale data infrastructure that...  ...structures, algorithms, performance complexity, and implications...  ...services with high throughput and low... 
    Training
    Performance
    Temporary work

    Coupang

    Mountain View, CA
    2 days ago
  • $174.9k - $261.3k

     ...world!The Data Labeling Engineering team designs, builds...  ...engineering, and AI/ML, defining the...  ...that create reliable training data at scale. Our tools and platform...  ...test scalable, high‑performance user experiences and...  ...building robust distributed platforms and applications... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...We are looking for a performance engineer who specializes in making large-scale machine learning workloads...  ...role is focused on distributed training runs spanning many nodes, and high-throughput batch inference sweeping...  ...intersection of accelerators, ML frameworks, and large‑... 
    Training
    Performance
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Decisive Point

    Sunnyvale, CA
    21 hours ago
  • $184.7k - $324.8k

     ...vision and machine learning engineers building real-time 3D...  ...understanding. This includes training and optimizing deep...  ...products Experience with large-scale distributed training and model...  ...or experience integrating ML models into performance-critical systems Strong... 
    Training
    Performance
    Relocation

    Apple

    Sunnyvale, CA
    4 days ago
  •  ...Technologies is seeking a Machine Learning Data Engineer to build and operate large-scale data systems powering AI training and evaluation pipelines. The role combines...  ...ingestion, transformation, lineage, and high-throughput delivery of data to training jobs across... 
    Training
    Remote work

    Bright Vision Technologies

    Santa Clara, CA
    21 hours ago
  •  ...building an AI training and model post-training...  ...makes large-scale training and RL...  ...intersection of distributed systems, GPU performance, model training...  ...and production engineering. Your responsibilities...  ..., and training throughput. Diagnose...  ...training, large-scale ML systems, or GPU... 
    Training
    Performance

    Nebius B.V.

    Palo Alto, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Performance Engineer — Scale Distributed Training & Throughput. Be the first to apply!