Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Performance Engineer

Applied Intuition

Performance Engineer

Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving machine on the planet. Applied Intuition services the automotive, defense, trucking, construction, mining and agriculture industries in three core areas: tools and infrastructure, operating systems, and autonomy. Eighteen of the top 20 global automakers, as well as the United States military and its allies, trust the company's solutions to deliver physical intelligence. Applied Intuition is headquartered in Sunnyvale, California, with offices in Washington, D.C.; San Diego; Ft. Walton Beach, Florida; Ann Arbor, Michigan; London; Stuttgart; Munich; Stockholm; Bangalore; Seoul; and Tokyo.

We are an in-office company, and our expectation is that full-time employees primarily work from their Applied Intuition office 5 days a week. However, we also recognize the importance of flexibility and trust our employees to manage their schedules responsibly. This may include occasional remote work, starting the day with morning meetings from home before heading to the office, or leaving earlier when needed to accommodate family commitments. This in-office expectation does not apply to contractor positions.

About the Role

We are looking for a performance engineer who specializes in making large-scale machine learning workloads fast and cost-efficient in the datacenter. This role is focused on distributed training runs spanning many nodes, and high-throughput batch inference sweeping petabytes of real-world autonomy logs for auto-labeling, data mining, ground-truth generation, and evaluation.

The optimization target here is not tail latency on a vehicle - it is throughput, cluster goodput, and cost per unit of data processed. A training run that wastes 30% of its GPU-hours on stalled data loaders, or an offline inference sweep that takes a week instead of a day, directly slows down how fast the whole company can iterate. You will own the gap between what our fleet of accelerators is theoretically capable of and what our workloads actually achieve: profiling across the stack, finding where the compute and the wall-clock time actually go, and closing the difference.

You will work at the intersection of accelerators, ML frameworks, and large-scale data infrastructure, partnering with the teams who own each layer to land wins that show up in training time-to-result and offline processing cost. At Applied, we encourage all engineers to take ownership over technical and product decisions, closely interact with users to collect feedback, and contribute to a thoughtful, dynamic team culture.

At Applied, you will:

  • Profile and optimize distributed training end to end - data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing
  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs: batching and scheduling strategies, quantization and low-precision execution, graph optimization, and accelerator saturation across long-running sweeps
  • Establish roofline and performance models for our workloads, quantify the gap between achieved and theoretical performance, and stack-rank optimization opportunities by impact and effort
  • Improve multi-node scaling efficiency: sharding and parallelism strategies, collective communication, interconnect utilization, and memory-bandwidth and kernel-fusion bottlenecks
  • Drive cluster goodput - reduce GPU idle time from input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery on long-running jobs
  • Build the benchmarking, observability, and regression-detection tooling that keeps performance from silently degrading as models and code evolve
  • Collaborate with engineers across functions to solve complex data and compute problems at scale
  • Contribute to a team culture that values effective collaboration, technical excellence, and innovation
We're looking for someone who has:
  • Hands-on ML performance engineering experience: profiling, roofline analysis, throughput optimization, and root-cause investigation in production systems
  • Experience with distributed multi-node training at scale (FSDP, DeepSpeed, Megatron, NCCL, or equivalent), including diagnosing scaling inefficiency as node count grows
  • Deep familiarity with GPU or accelerator performance concepts - memory bandwidth, kernel launch overhead, occupancy, quantization, collective communication
  • Experience with high-throughput or batch inference systems (NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar)
  • Fluency in Python and proficiency in C++ or another systems language
  • Excellent debugging, analytical, and problem-solving skills
  • A deep understanding of machine learning foundations, and the ability to develop technical solutions for problems with no established playbook

Nice to have:

  • GPU kernel development experience: CUDA, Triton, CUTLASS, or hand-tuned attention implementations
  • Experience with profiling toolchains such as Nsight Systems/Compute, PyTorch Profiler, or perf
  • Experience with GPU scheduling and orchestration on Kubernetes, Slurm, or Ray, including multi-tenant cluster utilization
  • Experience with fault tolerance and elastic training for long-running jobs - checkpointing strategy, straggler mitigation, preemption recovery
  • Familiarity with autonomy or robotics data (ROS, OpenCV, multi-sensor log formats)

Don't meet every single requirement? If you're excited about this role but your past experience doesn't align perfectly with every qualification in the job description, we encourage you to apply anyway. You may be just the right candidate for this or other roles.

Applied Intuition is an equal opportunity employer and federal contractor or subcontractor. Consequently, the parties agree that, as applicable, they will abide by the requirements of 41 CFR 60-1.4(a), 41 CFR 60-300.5(a) and 41 CFR 60-741.5(a) and that these laws are incorporated herein by reference. These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities, and prohibit discrimination against all individuals based on their race, color, religion, sex, sexual orientation, gender identity or national origin. These regulations require that covered prime contractors and subcontractors take affirmative action to employ and advance in employment individuals without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status or disability. The parties also agree that, as applicable, they will abide by the requirements of Executive Order 13496 (29 CFR Part 471, Appendix A to Subpart A), relating to the notice of employee rights under federal labor laws.

FOR US-BASED ROLES: Applied Intuition is committed to providing an accessible and inclusive application and interview experience to applicants who are disabled veterans and other applicants with disabilities or medical conditions. Reasonable accommodations are available, requesting an accommodation will not affect your candidacy in any way, and you are not required to disclose the nature of your disability or medical condition in order to make a request. If you require an accommodation please contact View email address on click.appcast.io. We will work with you!

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI Performance Engineer in Sunnyvale, CA vacancy
  • $100k

    Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI...  ...seeking a senior High Speed Interconnect / Signal Integrity Engineer to design and validate high-bandwidth links for large-... 
    Performance
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

    We are now looking for a Senior Agentic AI Software Engineer! Today, NVIDIA is tapping into the unlimited potential of AI to define the next...  ...with building AI agentsFamiliarity with how to evaluate performance of AI agents Experience using coding agents like Codex and... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build...  ...TVM, MLIR)Strong experience in GPU kernel development and performance optimizations (especially using CUDA C/C++, cuTile, Triton,... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...advance your career. THE ROLEWe are hiring AI Engineers to build recursive self-improvement...  ...sits at the intersection of AI systems, performance engineering, hardware-aware optimization... 
    Performance

    AMD

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will have a pivotal role in enhancing efficiency for our researchers by implementing progressions throughout the entire stack... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

    We are now looking for a Senior AI Frameworks Engineer (C++/Python)! NVIDIA's high-performance computing platforms are powering the AI revolution across many applications and industries. Within our software stack, CUTLASS stands out as a popular open-source ecosystem dedicated... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...recently, GPU deep learning ignited modern AI — the next era of computing — with the...  ...company”.NVIDIA is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress in deep...  ...engine end to end, with a focus on performance, fast builds, low memory use, and Ahead-of... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...breakthroughs in gaming, computer graphics, high-performance computing, and artificial intelligence....  ...powers everything from generative AI to autonomous systems, and we continue...  ..., and tools that enable researchers and engineers to develop the next generation of AI/ML... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $120k - $220k

     ...news and information powered by advanced AI, recommendation systems, and adtech.Recognized...  ...sharper every cycle.We're hiring the engineer who owns this agent end-to-end.What You’...  ...creative critic loop — Extend our LLM critic + performance-feedback regeneration from images-only to... 
    Performance
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    4 days ago
  • $152k - $241.5k

    We’re currently seeking a Senior AI Developer Technology Engineer, Financial Sector!Would you like to help shape the future of financial AI and data...  ...system bottlenecks to achieve the best possible performance of computer hardware? Could you be thrilled about an opportunity... 
    Performance
    Full time
    Work experience placement
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...looking for a Forward Deployed Research Engineer to build, evaluate, and deploy cutting-edge...  ...modeling.Diagnose and optimize AI system performance across models, data, infrastructure, orchestration... 
    Performance

    AMD

    Santa Clara, CA
    23 hours ago
  • $176k - $276k

    NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on...  ...systems specialist to architect, develop and bring up large scale performance platforms.What you will be doing:Design, implement and... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...tapping into the unlimited potential of AI to define the next era of computing. An era...  ...for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for...  ...compiler must deliver leading inference performance, fast build time, reduced memory footprints... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $94k - $141k

    AI EngineerLead I - ML EngineeringWho We Are:Born digital, UST transforms lives through...  ....com.You Are:UST is searching for an AI Engineer to build and deliver production-grade...  ...and tasks typically associated with the performance of the position. Other relevant essential... 
    Performance
    Full time
    Temporary work
    Part time
    Work at office
    Local area
    Flexible hours

    UST Global

    Santa Clara, CA
    2 days ago
  • $144k - $236k

     ...this role is hybrid, meaning it will be performed both from home and from a LinkedIn office...  ...the business needs of the team. The Video AI team sits at the heart of our LinkedIn’s...  ...billion members. As a Senior AI Software Engineer, you will own end-to-end machine learning... 
    Performance
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    1 day ago
  • $150k - $210k

     ...measurable business value.About the Position - Embodied AI EngineerWe are seeking an exceptional Embodied AI Engineer to build the foundation of LG's vision of theZero...  ...a combined bank of paid sick and vacation time. Performance based Short-Term Incentives (varies by role).... 
    Performance
    Full time
    Temporary work
    For contractors
    Local area
    Immediate start

    LG Electronics

    Santa Clara, CA
    23 hours ago
  • $195.2k - $275.58k

    Job Details:Job Description: The Software and AI (SAI) organization is seeking a highly skilled Software Development Engineer to contribute to the development and...  ...oneDNN, a complex, cross‑platform, open‑source performance library for deep learning applications ().oneDNN... 
    Performance
    Full time
    Local area
    Immediate start
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Intel

    Santa Clara, CA
    23 hours ago
  • $152k - $241.5k

    We are looking for a software engineer with a strong background in parallel processing and GPU architecture to push the limits of performance at the intersection of AI, high-performance computing, and financial markets. In this role, you will dive deep into parallel algorithms... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...amazing people. As part of Nvidia's applied AI team for chip design, you will have the...  ...at the intersection of research, engineering, and product development, transforming innovative...  ...ensuring their seamless and efficient performance. If you're passionate about the latest... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...career. THE ROLEWe are hiring Applied AI Engineers to work directly with hardware and software...  ..., verification, simulation, firmware, performance debugging, routing, issue triage, and... 
    Performance

    AMD

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA is the industry leader in high performance computing, gaming and AI. Our GPUs and SOCs give outstanding performance and efficiency, revolutionizing...  ...the Blackwell generation alone! Now we're hiring the engineer who will lead the rebuild of that toolchain around AI.We... 
    Performance
    Full time
    Immediate start

    Nvidia

    Santa Clara, CA
    3 days ago
  • $170.5k - $315.49k

    Job Details:Job Description: We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures.In this role, you will dive deep into the inference stack and redefine peak performance... 
    Performance
    Full time
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    4 days ago
  • Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This...  ...The RoleYou'll own model quality and performance for Cerebras' inference offerings. You will...  ...do this on a loop."You'll sit between engineering, product, and customer-facing teams. What... 
    Performance

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $244.14k - $413.16k

     ...forefront of innovation, integrating advanced AI and autonomous driving technologies into...  ...looking for a hands-on Senior Staff AI Engineer to build and scale production-grade AI...  ...high bar for reliability, evaluation, and performance.Job Responsibilities:Lead the technical design... 
    Performance
    Full time

    XPENG Motors

    Santa Clara, CA
    23 hours ago
  • $152k - $241.5k

     ...software is built in the age of Generative AI? Join NVIDIA’s TensorRT team to help lead...  ...swarms of AI agents to produce high-performance, high-quality, modern C++ software at an...  ...scale.If you are a systems-thinking C++ engineer who wants to help scale out an agentic development... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $106k - $159k

    Lead AI EngineerLead II - ML EngineeringWho We Are:Born digital, UST transforms lives...  ...com.You Are:UST is searching for a Lead AI Engineer to architect, lead, and deliver...  ...and tasks typically associated with the performance of the position. Other relevant essential... 
    Performance
    Full time
    Temporary work
    Part time
    Work at office
    Local area
    Flexible hours

    UST Global

    Santa Clara, CA
    2 days ago
  • $200k - $322k

     ...recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a...  ...choice to join us today.Design-for-X Engineering at NVIDIA works on groundbreaking innovations...  ...and SW dev teams, while monitoring performance, automating deployments and maintaining... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $171k

    Sr. Staff ll, AI Engineer - Search (L7-2)We exist to wow our customers. We know we’re doing the right thing when we hear our customers...  ...improvements in user engagement, operational efficiency, or system performance.Deep technical understanding of modern AI stacks, including:... 
    Performance
    Temporary work
    Flexible hours

    Coupang

    Mountain View, CA
    23 hours ago
  • $204k - $337k

     ...this role is hybrid, meaning it will be performed both from home and from a LinkedIn office...  ...New York, NY. Team Overview: The Video AI team sits at the heart of our LinkedIn’s...  ...community for all.Responsibilities: As a Senior Engineering Manager you will lead a team of 15-20... 
    Performance
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    2 days ago
  • $189.3k - $290.7k

    Job DescriptionAs an AI Engineer on the team, you will build and deploy applied AI/ML solutions that directly support simulation workflows...  ...3D reconstruction models, and excel at building robust, high-performance inference pipelines. This role is not focused on training... 
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Performance Engineer. Be the first to apply!