Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Performance Engineer - Offboard Training & Inference

Decisive Point

Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving machine on the planet. Applied Intuition services the automotive, defense, trucking, construction, mining and agriculture industries in three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the top 20 global automakers, as well as the United States military and its allies, trust the company’s solutions to deliver physical intelligence. Applied Intuition is headquartered in Sunnyvale, California, with offices in Washington, D.C.; San Diego; Ft. Walton Beach, Florida; Ann Arbor, Michigan; London; Stuttgart; Munich; Stockholm; Bangalore; Seoul; and Tokyo. Learn more at applied.co.

We are an in-office company, and our expectation is that full-time employees primarily work from their Applied Intuition office 5 days a week. However, we also recognize the importance of flexibility and trust our employees to manage their schedules responsibly. This may include occasional remote work, starting the day with morning meetings from home before heading to the office, or leaving earlier when needed to accommodate family commitments. This in-office expectation does not apply to contractor positions

About the Role

We are looking for a performance engineer who specializes in making large-scale machine learning workloads fast and cost-efficient in the datacenter. This role is focused on distributed training runs spanning many nodes, and high-throughput batch inference sweeping petabytes of real‑world autonomy logs for auto‑labeling, data mining, ground‑truth generation, and evaluation.

The optimization target here is not tail latency on a vehicle - it is throughput, cluster goodput, and cost per unit of data processed. A training run that wastes 30% of its GPU‑hours on stalled data loaders, or an offline inference sweep that takes a week instead of a day, directly slows down how fast the whole company can iterate. You will own the gap between what our fleet of accelerators is theoretically capable of and what our workloads actually achieve: profiling across the stack, finding where the compute and the wall‑clock time actually go, and closing the difference.

You will work at the intersection of accelerators, ML frameworks, and large‑scale data infrastructure, partnering with the teams who own each layer to land wins that show up in training time‑to‑result and offline processing cost. At Applied, we encourage all engineers to take ownership over technical and product decisions, closely interact with users to collect feedback, and contribute to a thoughtful, dynamic team culture.

At Applied, you will:

  • Profile and optimize distributed training end to end - data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing
  • Optimize large-scale offline and batch inference over petabyte‑scale sensor logs: batching and scheduling strategies, quantization and low‑precision execution, graph optimization, and accelerator saturation across long‑running sweeps
  • Establish roofline and performance models for our workloads, quantify the gap between achieved and theoretical performance, and stack‑rank optimization opportunities by impact and effort
  • Improve multi‑node scaling efficiency: sharding and parallelism strategies, collective communication, interconnect utilization, and memory‑bandwidth and kernel‑fusion bottlenecks
  • Drive cluster goodput - reduce GPU idle time from input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery on long‑running jobs
  • Build the benchmarking, observability, and regression‑detection tooling that keeps performance from silently degrading as models and code evolve
  • Collaborate with engineers across functions to solve complex data and compute problems at scale
  • Contribute to a team culture that values effective collaboration, technical excellence, and innovation

We're looking for someone who has:

  • Hands‑on ML performance engineering experience: profiling, roofline analysis, throughput optimization, and root‑cause investigation in production systems
  • Experience with distributed multi‑node training at scale (FSDP, DeepSpeed, Megatron, NCCL, or equivalent), including diagnosing scaling inefficiency as node count grows
  • Deep familiarity with GPU or accelerator performance concepts - memory bandwidth, kernel launch overhead, occupancy, quantization, collective communication
  • Experience with high‑throughput or batch inference systems (NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar)
  • Fluency in Python and proficiency in C++ or another systems language
  • Excellent debugging, analytical, and problem‑solving skills
  • A deep understanding of machine learning foundations, and the ability to develop technical solutions for problems with no established playbook

Nice to have:

  • GPU kernel development experience: CUDA, Triton, CUTLASS, or hand‑tuned attention implementations
  • Experience with profiling toolchains such as Nsight Systems/Compute, PyTorch Profiler, or perf
  • Experience with GPU scheduling and orchestration on Kubernetes, Slurm, or Ray, including multi‑tenant cluster utilization
  • Experience with fault tolerance and elastic training for long‑running jobs - checkpointing strategy, straggler mitigation, preemption recovery
  • Familiarity with autonomy or robotics data (ROS, OpenCV, multi‑sensor log formats)

Applied Intuition is an equal opportunity employer and federal contractor or subcontractor. Consequently, the parties agree that, as applicable, they will abide by the requirements of 41 CFR 60‑1.4(a), 41 CFR 60‑300.5(a) and 41 CFR 60‑741.5(a) and that these laws are incorporated herein by reference. These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities, and prohibit discrimination against all individuals based on their race, color, religion, sex, sexual orientation, gender identity or national origin. These regulations require that covered prime contractors and subcontractors take affirmative action to employ and advance in employment individuals without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status or disability. The parties also agree that, as applicable, they will abide by the requirements of Executive Order 13496 (29 CFR Part 471, Appendix A to Subpart A), relating to the notice of employee rights under federal labor laws.

FOR US-BASED ROLES:

Applied Intuition is committed to providing an accessible and inclusive application and interview experience to applicants who are disabled veterans and other applicants with disabilities or medical conditions. Reasonable accommodations are available, requesting an accommodation will not affect your candidacy in any way, and you are not required to disclose the nature of your disability or medical condition in order to make a request.
If you require an accommodation please contact View email address on click.appcast.io. We will work with you!

#J-18808-Ljbffr
Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the Machine Learning Performance Engineer - Offboard Training & Inference in Sunnyvale, CA vacancy
  •  ...intelligence to every moving machine on the planet....  ...; Seoul; and Tokyo. Learn more at applied.co....  ...We are looking for a performance engineer who specializes in making...  ...on distributed training runs spanning many nodes...  ...high-throughput batch inference sweeping petabytes of... 
    Training
    Performance
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    8 hours ago
  • $119.25k - $150.85k

     ...GM AV, the Model Deployment & Inference Solutions team deploys machine learning models from training frameworks (e.g., PyTorch)...  ...Escalade IQ, and we’re hiring engineers to help deliver the next generation...  ...on deployment workflows, performance investigations, model-... 
    Training
    Performance
    Full time
    Internship
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $189.3k - $320.7k

     ...developing and deploying machine learning solutions that...  ...scenarios.As a Staff ML Engineer on the Prometheus team...  ...and deploying offboard machine learning solutions...  ...vehicle development—from training and validation to...  ...payouts based on company performance, job level, and... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $117.7k - $221.4k

     ...the intersection of machine learning, data infrastructure,...  ...important scenarios, prepare training-ready data, and...  ...approach that first performs the cheapest reusable...  ...model reflects how Cola engineers think: build durable...  ..., featurization, and inference foundations that... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $150k

     ...edge foundation model training, alongside world-...  ...scientists, and engineers, tackling the most...  ...global hub for high-performance computing in deep learning, driving impactful...  ...performance for the machine learning software stacks...  ...at training and inference, and support the team... 
    Training
    Performance
    Full time
    Work experience placement
    Visa sponsorship

    Institute Of Foundation Models

    Sunnyvale, CA
    1 day ago
  •  ...excellence, constantly learning and evolving as...  ...As an  ML Engineer within the Application...  ...of core driving performance and feature-level...  ...choices, training methodologies, and...  ...for success as a Machine Learning Engineer...  ...or driver intent inference. Experience integrating... 
    Training
    Performance
    Full time
    Work at office
    Work from home

    Wayve

    Sunnyvale, CA
    1 day ago
  • $250k - $350k

     ...RoleWe are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In...  ...models into production.Improve Training Efficiency: (Bonus) Contribute to...  ...attention acceleration, and deep learning compiler stacks.GPU & Parallelism... 
    Training
    Performance
    Work at office
    3 days per week

    Pika

    Palo Alto, CA
    3 days ago
  • $195k - $230k

     ...the RoleWe are looking for a Senior Machine Learning Engineer to help evolve our large-scale...  ...metrics.Own systems from offline training online inference A/B experimentation metric analysis...  ...quality, model drift, and system performance in production.AI & LLM ApplicationsApply... 
    Training
    Performance
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    2 days ago
  • $174.72k - $295.68k

     ...cutting-edge R&D in AI, machine learning, and smart...  ...PTQ, QAT, on-vehicle inference and related fields.Key...  ...numerical consistency with training models, and productionize...  ...team to establish performance estimates and prove model...  ...and software engineering skills.Ability to work... 
    Training
    Performance
    Full time

    XPENG Motors

    Santa Clara, CA
    1 day ago
  • $230k - $265k

     ...industry-veteran scientists and engineers. As a Senior Machine Learning Engineer, you’ll bring your...  ...and implementation of training, fine-tuning, post-training, and inference strategies for large language...  ...improvements in model performance, robustness, observability,... 
    Training
    Performance
    Permanent employment

    Otter.ai

    Mountain View, CA
    2 days ago
  • $184k - $287.5k

    Intelligent machines powered by Artificial Intelligence...  ...that can learn, reason, and...  ...Machine Learning Engineers with a background...  ...Development: Design, train, and optimize innovative...  ..., continuous performance instrumentation, and...  ...optimization for real-time inference on embedded or... 
    Training
    Performance
    Full time
    Worldwide
    Night shift

    Nvidia

    Santa Clara, CA
    3 days ago
  • $300k

     ...edge foundation model training, alongside world-...  ...scientists, and engineers, tackling the most...  ...global hub for high-performance computing in deep learning, driving impactful...  ...and/or distributed inference optimization team...  ...Experience with large-scale machine learning workloads... 
    Training
    Performance
    Full time
    Flexible hours

    Institute Of Foundation Models

    Sunnyvale, CA
    1 day ago
  • $224k - $356.5k

     ...are looking for outstanding Machine Learning Engineers to join our Physical AI...  ...ensuring our AI agents are trained on the most diverse and rigorous...  ...of ML software, including performance optimization, testing, and...  ...the performance during inference/training.Familiarity with... 
    Training
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $193.3k - $261.5k

    The Product: AWS Machine Learning accelerators are at the forefront of...  ...delivers best-in-class ML inference performance at the lowest cost in...  ...deliver the best-in-class ML training performance with the most...  ...disciplines including silicon engineering, hardware design and... 
    Training
    Performance
    Internship
    Local area
    Work from home
    Relocation
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  •  ...Applied Intuition, Inc. is seeking a performance engineer to accelerate large-scale ML workloads in the data center. You will...  ..., optimization, and cost-efficiency for distributed training and large offline inferences. You will work across accelerators, ML frameworks, and... 
    Training
    Performance

    Applied PC

    Sunnyvale, CA
    2 days ago
  •  ..., Inc. is a Silicon Valley leader powering the future of physical AI. We seek a Performance Engineer to optimize large-scale ML workloads, focusing on distributed training, batch inference, and cost-effective data processing. You will own profiling across the stack, identify... 
    Training
    Performance

    Decisive Point

    Sunnyvale, CA
    3 days ago
  • $179.4k - $303.6k

     ...cutting-edge R&D in AI, machine learning, and smart...  ...strong Machine Learning Engineer / Computer Vision Engineer...  ...data preparation, model training, evaluation,...  ...traffic sign detection performance across diverse real-world...  ...TensorRT / quantization / inference acceleration.Work... 
    Training
    Performance
    Full time

    XPENG Motors

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...NVIDIA is looking for a Machine Learning (ML) Engineer to join the GPU accelerated Apache Spark team...  ...for ETL, SQL, and ML/DL model training and inference pipelines, spanning many domains and...  ...implement machine learning solutions for performance prediction and optimization of GPU... 
    Training
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...work helps care teams perform with greater...  ...Opportunity:As a Staff Machine Learning Engineer, you will be...  ...to store, annotate, train, and test on large pathology...  ...deployment for real-time inference, GPU/throughput... 
    Training
    Performance
    Work at office
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    4 days ago
  • $296.3k

     ...are seeking a Principal AI Engineer to lead the design and...  ...infrastructure that powers large-scale training and cloud inference. This includes...  ...reliability, scalability, and performance across the AI/ML platform....  ...realizing your ambitions. Learn how GM supports a... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $184.7k - $324.8k

     ...California, United States Machine Learning and AI Video is at the...  ...products, and as a research engineer on our team, you will...  ...lifecycle — including training infrastructure, performance optimization, data, and...  ...performance optimization for inference, ML model development... 
    Training
    Performance
    Relocation

    Apple

    Cupertino, CA
    10 hours ago
  • $160k - $200k

    Santa Clara, CAData Engineering - Machine Learning and Data Engineer /Full-time /HybridPlusAI is a Physical AI company pioneering...  ...petabytes of data while ensuring optimal performance for both training and inference phases. You will build robust pipelines for managing... 
    Training
    Performance
    Full time

    Plus.ai

    Santa Clara, CA
    2 days ago
  • $201.3k - $352.3k

     ...DescriptionIt all started when engineer Fred Luddy wrote code...  ...robustness, performance, safety, and real-world...  ...systems. Formal grounding in machine learning fundamentals — modeling, training, evaluation, and the...  ...to LLM fine-tuning or inference optimization in production... 
    Training
    Performance
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    2 days ago
  •  ...hiring Computer Vision / Machine Learning Software Engineers to build compute-constrained...  ...detection Optimize performance, accuracy, and speed of...  ...for dataset management, training, and deployment Participate...  ...(both training and inference) ~ Knowledge of model optimization... 
    Training
    Performance
    Full time
    Visa sponsorship

    Corvus Robotics

    Mountain View, CA
    1 day ago
  • $111k - $231.25k

     ...Machine Learning Engineer II (Finance) Yahoo Mail is the ultimate consumer inbox...  ..., maintainable, and high-performing code that operates...  ...on experience developing, training, and deploying machine learning...  ...vLLM, TensorRT-LLM, Triton Inference Server, or ONNX Runtime. Modern... 
    Training
    Performance
    Work at office
    Flexible hours

    Yahoo Holdings Inc. in

    Mountain View, CA
    2 days ago
  • $170.6k - $261.3k

     ...a global scale.As a Senior Machine Learning Engineer on the State Estimation and...  ...safety. What You'll DoDesign, train, and evaluate ML perception...  ...to improve model performance against those metrics. Analyze...  ...Implement efficient training and inference pipelines, including model... 
    Training
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    19 hours ago
  • $2,000 per month

     ...Machine Learning Research EngineerCupertino, CAEtched is building...  ...into maximally performant instruction sequences...  ...Implement model-specific inference-time acceleration...  ...distributed inference/training...  ...Cupertino, and greatly value engineering skills. We do not have... 
    Training
    Performance

    ETCHED LLC

    Cupertino, CA
    2 days ago
  • $235.03k - $352.29k

     ...RoleWe are looking for a Staff Machine Learning Engineer to be a technical leader on...  ...at the Frontier: Design, train, and productionize state-of...  ...: data, training, onboard inference, closed-loop and open-loop...  ...also eligible for an annual performance bonus, equity, and a... 
    Training
    Performance
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    2 days ago
  • $184.7k - $324.8k

     ...is seeking an experienced Machine Learning Engineer to build, operate, and scale...  ...to ship reliable, high-performance ML-powered features across...  ...sampling and collecting data for training, labeling via human...  ...Qualifications Familiarity with inference optimization techniques... 
    Training
    Performance
    Work experience placement
    Relocation

    Apple Inc.

    Cupertino, CA
    8 hours ago
  • $213k - $263k

     ...are looking for engineers with ML software...  ...Waymo onboard ML inference engine for Waymo...  ...from efficient deep learning models, model compression...  ...ML workload performance at the hardware...  ...onboard and offboard deployment....  ...with designing, training and debugging deep... 
    Training
    Performance
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Performance Engineer - Offboard Training & Inference. Be the first to apply!