ML Performance Engineer: Scale Training & Throughput
Applied Intuition
Applied Intuition in Sunnyvale is seeking a Performance Engineer to accelerate large-scale ML workloads in data centers. You will optimize distributed training across many nodes and improve batch inference on petabyte-scale sensor logs, targeting throughput and cost-per-data processed.
You will own profiling across the stack, from data loading to kernel execution, and work with accelerators, ML frameworks, and data infrastructure to close performance gaps and drive faster iterations.
#J-18808-Ljbffr- ...Applied Intuition, Inc. in Sunnyvale, CA, is seeking a performance engineer to optimize large-scale ML workloads in the datacenter. This role focuses on distributed training across many nodes and high-throughput batch inference over petabytes of real-world autonomy logs...TrainingPerformance
- ...Applied Intuition, Inc. is seeking a performance engineer to accelerate large-scale ML workloads in the data center. You will own profiling, optimization, and cost-efficiency for distributed training and large offline inferences. You will work across accelerators, ML...TrainingPerformance
$189k - $300k
...transportation on a global scale. The Data Scaling team... ...works on and delivers ML models to the product... ...directly impacting AV product performance through smart use of... ...foundation model pre-training and fine-tuning with... ...high-impact team of AI/ML engineers, data scientists and...TrainingPerformanceFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$153.2k - $234.1k
...transportation on a global scale.Role Overview:Are you... ...every machine learning engineer working on our cutting-... ...the safety and performance of the car, rather than... ...driverless vehicles.As a Senior ML Infra Engineer, you... ...machine learning model training and evaluation workflows...TrainingPerformanceFull timeWork at officeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours- ...Intuition, Inc. is a Silicon Valley leader powering the future of physical AI. We seek a Performance Engineer to optimize large-scale ML workloads, focusing on distributed training, batch inference, and cost-effective data processing. You will own profiling across the stack...TrainingPerformance
$189.3k - $290.7k
...transportation on a global scale. Role:Are you passionate about... ...-world scenarios.As a Staff ML Infra Engineer, you will drive the... ...enable rapid dataset generation, training, evaluation, and iteration of... ...training pipelines that are performant, easy to use, and exceptionally...TrainingPerformanceFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours- Rhoda AI in Mountain View is seeking a Staff / Principal ML Training Systems Engineer to lead the performance of large-scale multimodal training systems. This role involves improving training efficiency and collaborating closely with research teams to accelerate model iteration...TrainingPerformance
- ...Principal Machine Learning Engineer to join our Models and... ...of distributed training of large models on a large... ...training generative AI at scale.THE PERSON:The ideal... ...-end training pipeline performance.Optimize the distributed... ...EXPERIENCE:Experience with ML/DL frameworks such as...TrainingPerformance
$189.4k - $300.6k
...transportation on a global scale. Are you passionate... ...development. We engineer high-performance tools that identify top... ...with data-intensive ML teams to drive rapid... ...back to data selection, training, and launch decisions... ..., Node.js, high-throughput data streamingData/Infra...TrainingPerformanceFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours- ...OverviewAs our Staff Software Engineer, ML infra Engineer for... ...data needed to train complex ML models and... ...pipelinesDevelop and scale data infrastructure that... ...structures, algorithms, performance complexity, and implications... ...services with high throughput and low...TrainingPerformanceTemporary work
- ...AI is seeking a Machine Learning Engineer in Palo Alto, California, to build... .... You will play a key role in scaling our Ray and PyTorch-based systems... ...directly influence the efficiency and performance of our data processing and model training pipelines. #J-18808-Ljbffr...TrainingPerformance
$193.3k - $261.5k
...Software Development Engineer to own the design... ..., developing high-performance compute kernels,... ...operate at frontier scale with large... ...kernels for a custom ML accelerator architecture... ...drive latency and throughput improvements from... ...architecture, training/inference lifecycles...TrainingPerformanceLocal areaFlexible hours- ...Technologies is seeking a Machine Learning Data Engineer to build and operate large-scale data systems powering AI training and evaluation pipelines. The role combines... ...ingestion, transformation, lineage, and high-throughput delivery of data to training jobs across...TrainingRemote work
$122.6k - $185k
...the cloud for AI training and inference? Want... ...continuous price performance improvements in the... ...scalability in AI/ML and HPC workloads.... ...to those at cloud scale? If yes, then come... ...The AWS Hardware Engineering team creates server... ...concurrent, high throughput systems and knowledge...TrainingPerformanceLocal areaFlexible hours$117.7k - $221.4k
...value data from large-scale real-world sensor streams... ...scenarios, prepare training-ready data, and support... ...approach that first performs the cheapest reusable... ...model reflects how Cola engineers think: build durable intermediate... ...compute utilization, throughput, storage, and query...TrainingPerformanceFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours$250k - $350k
...Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven... ...leading user experiences at scale.You will design and... ...into production.Improve Training Efficiency: (Bonus) Contribute... ...HaveExperience with high-throughput video or real-time streaming...TrainingPerformanceWork at office3 days per week$195k - $230k
...meaningful challenges at scale.Together, we... ...Machine Learning Engineer to help evolve our... ...and iterate on high-throughput, low-latency... ...systems from offline training online inference A... ...drift, and system performance in production.AI &... ...large-scale data and ML systems (e.g.,...TrainingPerformanceFull timeLocal areaWork from home$213k - $263k
...U.S. states. The ML Optimization team... ...We are looking for engineers with ML software &... ...an efficient, high-performance ML runtime and... ...compute and large-scale, offboard data center... ...systems with the high-throughput, highly concurrent... ..., relevant training and education, and...TrainingPerformanceFull timeRemote work$155.42k - $395.9k
...DescriptionAbout the Team:The ML Compute Platform is part... ...platform supports the training and deployment of state-... ...learning models with a focus on performance, availability,... ...looking for a Senior Software Engineer to join our team and help us scale our platform for...TrainingPerformanceFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$165.2k - $223.6k
...the forefront of maximizing performance for AWS's custom ML accelerators. Working at... ...hardware-software boundary, our engineers craft high-performance... ...ML inference and training performance.As part of the... ...patterns, reliability and scaling) of new and existing systems...TrainingPerformanceInternshipLocal areaWork from homeFlexible hours- ...is building next-generation generalist robots and is seeking a Staff / Principal ML Training Systems Engineer to own training systems performance end-to-end. You will optimize large-scale multimodal training, define parallelism strategies, and drive efficiency across GPUs...TrainingPerformance
$207k - $300k
...content through advanced context engineering and agentic feedback loops... ...and feedback to improve the performance of the AIGC stack.Resolve... ...handle information at massive scale, and extend well beyond web... ..., and relevant education or training. US: $207000 - $300000 (USD)...TrainingPerformance$124k - $250k
...member of our software engineering infra team, you'll... ...The team builds a high-performance, high availability, globally... ...with high throughput and low latency to our... ...model delivery, including training, serving, and optimizations... ..., and maintain large-scale distributed systemsCollaborate...TrainingPerformance- ...worldwide.We’re a team of engineers, clinicians, and... ...helps care teams perform with greater... ...learning models on large-scale image data.... ...and implement AI/ML approaches to extract... ...store, annotate, train, and test on large... ...time inference, GPU/throughput optimization (e.g....TrainingPerformanceWork at officeLocal areaWorldwideFlexible hours
$296.3k
...seeking a Principal AI Engineer to lead the design and... ...that powers large-scale training and cloud inference. This... ...accelerating training throughput, scaling multi-modal models... ...and optimize core AI/ML platform... ...reliability, scalability, and performance across the AI/ML platform...TrainingPerformanceFull timeLocal areaRemote workWork from homeFlexible hours$281k - $356k
...learning models to deliver training and evaluation data for... ...and software engineers who are passionate about... ...driver to improve the performance of our technology stack... ...and fine-tuning large-scale generative models to produce... ...in Python and standard ML frameworks (e.g., JAX,...TrainingPerformanceFull time$90.1k - $191.8k
...the world!The Data Labeling Engineering team designs, builds, and... ...engineering, data engineering, and ML, defining labeling... ...controls that create reliable training data at scale.Our team builds the mission... ..., and test scalable, high‑performance user experiences and services...TrainingPerformanceFull timeWork experience placementLocal areaRemote workWork from homeRelocation packageFlexible hours$174.9k - $261.3k
...the world!The Data Labeling Engineering team designs, builds, and... ..., data engineering, and AI/ML, defining the strategies, tooling... ...that create reliable training data at scale. Our tools and platform are... ...implement, and test scalable, high‑performance user experiences and...TrainingPerformanceFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...We are looking for a performance engineer who specializes in making large-scale machine learning workloads... ...is focused on distributed training runs spanning many nodes, and high-throughput batch inference sweeping... ...intersection of accelerators, ML frameworks, and large‑...TrainingPerformanceFull timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift
$184.7k - $324.8k
...vision and machine learning engineers building real-time 3D... ...understanding. This includes training and optimizing deep learning... ...products Experience with large-scale distributed training and... ...or experience integrating ML models into performance-critical systems Strong communication...TrainingPerformanceRelocation
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Performance Engineer: Scale Training & Throughput. Be the first to apply!
- ai ml engineer Sunnyvale, CA
- senior ml engineer Sunnyvale, CA
- computer vision machine learning engineer Sunnyvale, CA
- machine learning engineer Sunnyvale, CA
- machine learning ai engineer Sunnyvale, CA
- machine learning software engineer Sunnyvale, CA
- senior performance tester Sunnyvale, CA
- lead performance test engineer Sunnyvale, CA
- performance windows Sunnyvale, CA
- performance specialist Sunnyvale, CA

