Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Performance Engineer — Scale Distributed Training & Throughput

applied

Applied Intuition, Inc. is seeking a performance engineer to accelerate large-scale ML workloads in the data center. You will own profiling, optimization, and cost-efficiency for distributed training and large offline inferences. You will work across accelerators, ML frameworks, and data infra, partnering with teams to land improvements that shorten time-to-result and reduce data processing costs. Collaboration and technical excellence are valued. #J-18808-Ljbffr applied

Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the ML Performance Engineer — Scale Distributed Training & Throughput in Sunnyvale, CA vacancy
  •  ...Intuition in Sunnyvale is seeking a Performance Engineer to accelerate large-scale ML workloads in data centers. You will optimize distributed training across many nodes and improve batch...  ...-scale sensor logs, targeting throughput and cost-per-data processed. You will... 
    Training
    Performance

    Applied Intuition

    Sunnyvale, CA
    2 days ago
  • Applied Intuition, Inc. in Sunnyvale, CA, is seeking a performance engineer to optimize large-scale ML workloads in the datacenter. This role focuses on distributed training across many nodes and high-throughput batch inference over petabytes of real-world autonomy logs... 
    Training
    Performance

    NLP PEOPLE

    Sunnyvale, CA
    1 day ago
  • $153.2k - $234.1k

     ...transportation on a global scale.Role Overview:Are...  ...machine learning engineer working on our...  ...the safety and performance of the car, rather...  ...vehicles.As a Senior ML Infra Engineer,...  ...machine learning model training and evaluation...  ...large-scale distributed systems/applications... 
    Training
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $189.3k - $290.7k

     ...transportation on a global scale. Role:Are you...  ...scenarios.As a Staff ML Infra Engineer, you will drive the development...  ...dataset generation, training, evaluation, and...  ...pipelines that are performant, easy to use, and exceptionally...  ...building large-scale distributed systems, applications... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...Intuition, Inc. is a Silicon Valley leader powering the future of physical AI. We seek a Performance Engineer to optimize large-scale ML workloads, focusing on distributed training, batch inference, and cost-effective data processing. You will own profiling across the stack... 
    Training
    Performance

    Decisive Point

    Sunnyvale, CA
    1 day ago
  •  ...is building an AI training and model post-training...  ...for large-scale training, RL experiments...  ...intersection of distributed systems, GPU performance, and ML framework...  ...Python and PyTorch engineering skills, hands-on experience...  ...to optimize throughput and memory across... 
    Training
    Performance

    Nebius B.V.

    Palo Alto, CA
    1 day ago
  •  ...in Mountain View is seeking a Staff / Principal ML Training Systems Engineer to lead the performance of large-scale multimodal training systems. This role involves...  ...will have proven experience in boosting distributed training performance and hands-on skills with modern... 
    Training
    Performance

    Rhoda AI

    Mountain View, CA
    2 days ago
  • $150k

     ...edge foundation model training, alongside world-...  ...data scientists, and engineers, tackling the most fundamental...  ...global hub for high-performance computing in deep...  .... The Role The Distributed ML Engineer will play a...  ...methodologies, and large-scale machine learning... 
    Training
    Performance
    Full time
    Work experience placement
    Visa sponsorship

    Institute Of Foundation Models

    Sunnyvale, CA
    11 hours ago
  •  ...Principal Machine Learning Engineer to join our Models...  ...by the challenge of distributed training of large models on a...  ...generative AI at scale.THE PERSON:The ideal...  ...end training pipeline performance.Optimize the distributed...  ...EXPERIENCE:Experience with ML/DL frameworks such as... 
    Training
    Performance

    AMD

    San Jose, CA
    1 day ago
  • $189k - $300k

     ...transportation on a global scale. The Data Scaling team...  ...works on and delivers ML models to the product...  ...directly impacting AV product performance through smart use of...  ...foundation model pre-training and fine-tuning with...  ...high-impact team of AI/ML engineers, data scientists and... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  •  ...Santa Clara is seeking a skilled software engineer with over 8 years of experience to...  ...applications. The role involves designing distributed training strategies, collaborating with ML researchers, and developing tools for performance enhancement. Ideal candidates have a... 
    Training
    Performance

    Odyssey

    Santa Clara, CA
    2 days ago
  •  ...seeking a Machine Learning Engineer in Palo Alto,...  ...will play a key role in scaling our Ray and PyTorch-based...  ...extensive experience with distributed systems, strong proficiency...  ...the efficiency and performance of our data processing and model training pipelines. #J-18808-Ljbffr... 
    Training
    Performance

    Orbifold AI

    Palo Alto, CA
    1 day ago
  • $165.2k - $223.6k

     ...Software Development Engineer to own the...  ...developing high-performance compute kernels,...  ...operate at frontier scale with large distributed models.This is a...  ...for a custom ML accelerator architecture...  ...latency and throughput improvements...  ...transformer architecture, training/inference... 
    Training
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $122.6k - $185k

     ...cloud for AI training and inference?...  ...continuous price performance improvements...  ...scalability in AI/ML and HPC...  ...those at cloud scale? If yes, then...  ...AWS Hardware Engineering team creates server...  ..., high throughput systems and knowledge...  ...of complex distributed systemsAmazon... 
    Training
    Performance
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  •  ...company seeking an expert ML engineer to scale large models and build robust data pipelines for training and evaluation. You will design...  ...leverage Spark and Ray for distributed processing in a high-growth...  ...across iterations to push performance ceilings. #J-18808-Ljbffr... 
    Training
    Performance

    SpaceXAI

    Palo Alto, CA
    4 days ago
  • $250k - $350k

     ...Staff level Inference Engineers to accelerate the performance of Pika's AI-driven...  ...experiences at scale.You will design and...  ...computing kernels and distributed workloads using...  ...production.Improve Training Efficiency: (Bonus)...  ...HaveExperience with high-throughput video or real-time... 
    Training
    Performance
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    11 hours ago
  • $195k - $230k

     ...challenges at scale.Together, we reached...  ...Learning Engineer to help evolve...  ...iterate on high-throughput, low-latency recommendation...  ...from offline training online...  ...drift, and system performance in production....  ...scale data and ML systems (e.g., Spark, distributed training, real-... 
    Training
    Performance
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    7 hours ago
  •  ...infrastructure company in California seeks a Member of Technical Staff — Training to design and optimize large-scale distributed training systems for frontier AI models. Candidates should have 5+ years of experience in ML systems and be proficient in Python along with another... 
    Training

    RadixArk

    Palo Alto, CA
    11 hours ago
  • $165.2k - $223.6k

     ...forefront of maximizing performance for AWS's custom ML accelerators....  ...software boundary, our engineers craft high-...  ...unparalleled ML inference and training performance.As part...  ...computing, and distributed architectures,...  ...patterns, reliability and scaling) of new and... 
    Training
    Performance
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $155.42k - $395.9k

     ...DescriptionAbout the Team:The ML Compute Platform is...  ...supports the training and deployment of...  ...with a focus on performance, availability, concurrency...  ...a Senior Software Engineer to join our team and help us scale our platform for...  ...background in distributed systems, infrastructure... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  • $207k - $300k

     ...through advanced context engineering and agentic feedback...  ...to improve the performance of the AIGC stack.Resolve...  ...at massive scale, and extend well beyond...  ...information retrieval, distributed computing, large-scale...  ...relevant education or training. US: $207000 - $30000... 
    Training
    Performance

    Google

    Mountain View, CA
    3 days ago
  • $124k - $250k

     ...of our software engineering infra team, you'll...  ...team builds a high-performance, high availability, globally distributed ecosystem...  ...infrastructure with high throughput and low latency...  ...delivery, including training, serving, and...  ...and maintain large-scale distributed... 
    Training
    Performance

    AppLovin

    Palo Alto, CA
    4 days ago
  •  ...We’re a team of engineers, clinicians,...  ...helps care teams perform with greater precision...  ...on large-scale image data....  ...and implement AI/ML approaches to extract...  ..., annotate, train, and test on...  ...Experience with distributed training and cloud...  ...inference, GPU/throughput optimization (e... 
    Training
    Performance
    Work at office
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    1 day ago
  • $296.3k

     ...seeking a Principal AI Engineer to lead the...  ...that powers large-scale training and cloud...  ...accelerating training throughput, scaling multi-modal...  ...across distributed training, training...  ...optimize core AI/ML platform infrastructure...  ...scalability, and performance across the AI/ML... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  •  ...OverviewAs our Staff Software Engineer, ML infra Engineer for...  ...data needed to train complex ML models and...  ...pipelinesDevelop and scale data infrastructure that...  ...structures, algorithms, performance complexity, and implications...  ...services with high throughput and low... 
    Training
    Performance
    Temporary work

    Coupang

    Mountain View, CA
    4 days ago
  • $184.7k - $324.8k

     ...and as a research engineer on our team, you...  ...across the full ML lifecycle — including training infrastructure, performance optimization, data...  ...colleagues to design, scale, and harden the...  ...achieve quality, throughput, computational...  ...Experience with distributed training, large-scale... 
    Training
    Performance
    Relocation

    Apple

    Cupertino, CA
    2 days ago
  • $174.9k - $261.3k

     ...world!The Data Labeling Engineering team designs, builds...  ...engineering, and AI/ML, defining the...  ...that create reliable training data at scale. Our tools and platform...  ...test scalable, high‑performance user experiences and...  ...building robust distributed platforms and applications... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  • $213k - $263k

     ...states. The ML Optimization team...  ...looking for engineers with ML software...  ..., high-performance ML runtime and...  ...compute and large-scale, offboard data...  ...with the high-throughput, highly concurrent...  ...concurrent distributed backend systems...  ..., relevant training and education,... 
    Training
    Performance
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $250k - $350k

     ...the world's leading ML systems engineers, including leaders...  ...powering our large-scale training, inference, and reinforcement...  ...stack to maximize performance, scalability,...  ...systems Design distributed runtimes and...  ...communication for maximum throughput and end-to-end... 
    Training
    Performance
    Visa sponsorship

    Periodic Labs

    Menlo Park, CA
    4 days ago
  • $150k - $350k

     ...AI to assist engineers in RTL design,...  ...are seeking an ML Systems Engineer...  ...optimize the performance and efficiency...  ...clusters for training and inference...  ...limits of LLM throughput and latency. Your...  ...with large‑scale ML systems, GPU...  ...batching strategies, distributed inference, or... 
    Training
    Performance

    ChipAgents

    San Jose, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Performance Engineer — Scale Distributed Training & Throughput. Be the first to apply!