Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Inference Engineer - High-Performance GPU Systems

Perplexity

Perplexity is seeking an engineer to join our inference stack. You will work on GPU-accelerated ML inference, supporting transformer models, caching, and low-latency serving. You\'ll collaborate across Rust, Python, CUDA, CuTe DSL, and deploy production distributed systems under real load with a focus on performance and reliability. You will contribute to a stack built around PyTorch, CUDA kernels, and modern ML tooling, delivering scalable inference in a fast-paced environment. #J-18808-Ljbffr Perplexity

Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the AI Inference Engineer - High-Performance GPU Systems in California, MO vacancy
  • $184k - $287.5k

    Join to apply for the Senior High Performance AI Engineer role at NVIDIA NVIDIA has...  .... An era in which our GPU acts as the brains of computers...  ...build and optimize agentic AI systems for the CUDA ecosystem. Co...  ...distributed training, and inference/serving—and with model/... 
    Performance

    NVIDIA

    California, MO
    9 hours ago
  • AI Systems Engineer - Codex Core Agents About The Team The Codex Core Agents...  ...harness, model interaction, inference, sandboxed execution, orchestration...  ...reliability, and the performance envelope around tokens, latency...  ..., inference/runtime stack, GPU fleet, and product surface.... 
    Performance

    OpenAI

    California, MO
    3 days ago
  • $251.1k - $385k

    It all started when engineer Fred Luddy wrote code that...  ...Today, ServiceNow is the AI control tower for...  ...multi‑model orchestration systems. Drive the enterprise...  ...infrastructure, including GPU/TPU‑accelerated and...  ...Build, scale, andretaina high‑performing global organization of... 
    Performance
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    California, MO
    9 hours ago
  •  ...Principal Technical Support Engineer to lead customer...  ...platform bring-up, and performance optimization for next-generation AI networking...  ...This role focuses on high-speed Ethernet fabrics...  ...based network operating systems, switch ASICs, and GPU cluster networking, requiring... 
    Performance

    Yoh, A Day & Zimmermann Company

    California, MO
    4 days ago
  • $184k - $287.5k

    A leading tech company is seeking a Senior Performance Engineer in California to enhance AI system performance and datacenter applications. The role requires extensive experience in accelerated computing, deep learning frameworks, and cloud/container architecture. Applicants... 
    Performance

    NVIDIA

    California, MO
    9 hours ago
  • OpenAI is seeking an AI Systems Engineer for Codex Core Agents to design and build the core agent harness and execution loop. You will enable...  ...configuration layers, focusing on reliability, safety, and performance. This role involves collaboration with research,... 
    Performance

    OpenAI

    California, MO
    3 days ago
  •  ...to lead the development of next-generation custom AI silicon optimized for edge deployments. You will shape architecture for high-performance, energy-efficient systems supporting on-device intelligence and ML inference. You will collaborate with internal teams and external... 
    Performance

    OpenAI

    California, MO
    3 days ago
  • $124k - $195.5k

    The AI revolution is not powered...  ...higher-throughput inference makes AI...  ...scientists and engineers explore more possibilities...  .... At NVIDIA, performance is not a...  ...many downstream systems and be repeated...  ...capabilities and GPU architectures...  ...employer. As we highly value diversity... 
    Performance

    NVIDIA

    California, MO
    9 hours ago
  • Rosendin Electric is seeking a Systems Engineer to provide technical engineering support and ensure compliance with project requirements and...  ...requires a minimum of 2 years' experience as a Systems Engineer, a high school diploma or equivalent, and preferred RCDD credential.... 
    For contractors

    Rosendin

    California, MO
    9 hours ago
  • ABOUT RETELL AI Retell AI is using first-principles...  ...hiring an Applied AI Engineer to work directly with enterprise...  ...architect scalable AI systems for production. Solve...  ..., reliability, and performance. Work directly in...  ...exceptional builders. High Ownership: Small teams,... 
    Performance
    Full time
    H1b

    Retell AI

    California, MO
    3 days ago
  • $123.24k - $180k

    Overview of Role As a Sr./Principal AI Engineer within TSMC's Artificial Intelligence for...  ...startup, building foundational systems from the ground up and tackling challenges...  ...learning engineering, or related fields in high-performance environments. 5+ years of hands-on... 
    Performance
    Work at office

    TSMC - Taiwan Semiconductor Manufacturing Company Limited

    California, MO
    1 day ago
  • A technology solutions company is seeking an Imaging Software Engineer to develop next-generation AR/VR imaging systems. The engineer will design high-performance imaging software using MATLAB, Python, C++, and CUDA. Responsibilities include automating workflows and collaborating... 
    Performance
    Remote job
    Hourly pay
    Contract work

    Tailored Management

    California, MO
    22 hours ago
  •  ...Oracle Cloud Infrastructure is seeking a Principal Systems Software Engineer to build low-level systems software and platform-management for next-generation GPU infrastructure. You will own capabilities spanning BMC software, platform management, firmware lifecycle,... 

    Jobleads-US

    California, MO
    2 days ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving...  .... One person, one GPU. If you'd like to build the...  ...advanced GPU and Networking systems. Document and update data center...  ...Experience with High Performance Compute GPU systems (air or water... 
    Performance
    Local area
    Flexible hours
    Shift work
    Day shift

    Neura Market

    California, MO
    2 days ago
  •  ...DDN is seeking a highly experienced Senior Staff Engineer specializing in AI Data Path & Storage...  ...of advanced storage systems with next-generation AI inference pipelines. This role...  ...design and deliver high-performance data movement...  ...data movement across GPU, memory, and... 
    Performance

    DDN

    California, MO
    22 hours ago
  • $125.5k - $200.7k

     ...Opportunity Overview The Physical AI Field Applications Engineer is a sales‑focused,...  ...sites. Help design systems using UR/MiR hardware, sensors...  ...AI. Troubleshoot robotic performance issues during pre‑ and post...  ...perception, control, and AI inference design on UR/MiR platforms... 
    Performance
    Local area
    Worldwide
    Relocation
    Flexible hours

    Universal Robots

    California, MO
    9 hours ago
  • $216k - $295k

     ...world’s best scientists, engineers, and business...  ...to build talent‑dense, high‑impact teams that are aligned...  ...transformative potential of AI. About The Role As the...  ...development of a high‑performing technical recruiting...  ...the capabilities of AI systems and seek to safely deploy... 
    Performance
    Work at office
    Work from home
    Relocation package

    SupportFinity™

    California, MO
    4 days ago
  •  ...looking for a senior backend engineer to help design and operate high-volume, distributed backend systems . This role focuses on data...  ...pragmatic decisions around performance, reliability, and cost. Core...  ...Kubernetes experience is a plus). ML inference or classification pipelines... 
    Performance

    Ranwalk

    California, MO
    4 days ago
  • $152k - $241.5k

     ...NVIDIA’s accelerated inference software stack....  ...boundaries of inference performance. Benchmark state‑...  ...generation of AI models and...  ...optimization, especially for GPU‑based applications...  ...processor and system‑level performance...  ...opportunity employer. As we highly value diversity in... 
    Performance

    NVIDIA

    California, MO
    9 hours ago
  • Cohere is seeking engineers to design and develop cutting-edge multimodal AI systems across text, speech, and vision. You will push the boundaries of model capabilities and work with leading compute infrastructure to explore novel ideas. Join a team that values rapid experimentation... 
    Remote job

    Cohere

    California, MO
    3 days ago
  • $50.24 - $55.82 per hour

     ...next‑generation memory for datacenter and AI workloads. In this internship you will...  ...and simulation platform — a multi‑agent system that turns large volumes of external and...  ...seeking a software‑centric AI Agentic Systems Engineer Intern to help design, build, and... 
    Full time
    Internship
    Local area
    Immediate start

    1040 Micron Semiconductor Prds

    California, MO
    3 days ago
  • $152k - $218.5k

     ...Software Engineer, CUDA-Q page is loaded## Software...  ...multi-processor systems. We are looking for...  ...record of building performant and robust...  ...MLIR* Proficiency in GPU- and/or FPGA-programmingNVIDIA...  ...to be one of high technology's most desirable...  ....NVIDIA uses AI tools in its... 
    Performance
    Remote work

    NVIDIA

    California, MO
    22 hours ago
  • $150k - $180k

    Skydance is seeking a Senior Gameplay AI Engineer to design, implement, and own core AI systems for NPCs. You will partner with Design, Animation, and Audio to deliver responsive, believable characters in a production environment. The role requires 7+ years in gameplay/... 
    Remote work

    Sky Dance

    California, MO
    1 day ago
  • $140k - $180k

     ...seeking a Senior Software Engineer specializing in AI and Machine Learning...  ...models and agentic systems that power...  ...pipelines. Optimize model inference for latency and...  ...caching, and other performance techniques. Develop...  ...training loops, autograd, GPU acceleration/CUDA, and... 
    Performance
    Remote work
    Afternoon shift
    Early shift

    Prosum

    California, MO
    1 day ago
  • $160k - $225k

     ...world's best data and AI infrastructure...  ...platform for large-scale GPU training and fine-...  ...a Senior Software Engineer for AI Runtime, you...  ...and scaling the systems that make large-scale...  ...training performance, fault tolerance, and...  ...delivering scalable, high-throughput, and resilient... 
    Performance
    Local area
    Worldwide

    Databricks

    California, MO
    3 days ago
  • Databricks is seeking a staff software engineer focusing on GenAI performance and kernel development. You will own high-performance GPU kernels for the GenAI inference stack, optimize for various...  ...collaborate with ML researchers, systems engineers, and product teams to push... 
    Performance

    Databricks

    California, MO
    3 days ago
  •  ...EnCharge AI is a leader in advanced AI hardware and software systems for edge-to-cloud computing. EnCharge...  ...class solutions. The high-performance architecture is...  ...looking for an Embedded SW Engineer to develop the...  ...Firmware used to deploy inference jobs on EAI processors... 
    Performance

    Polluxa, Inc.

    California, MO
    1 day ago
  • NVIDIA is seeking a Backend Compiler Engineer to enhance a GPU compiler backend, focusing on C++ improvements, new passes, and robust optimizations...  ...with global teams to refine architecture and deliver high-performance code generation for graphics and compute workloads. The... 
    Performance

    NVIDIA

    California, MO
    3 days ago
  • Micron Technology, Inc's Cloud Memory Business Unit invites an AI-focused intern to design, build, and productize an internal agentic...  ...-generation and test-planning capability while collaborating with senior memory and systems #J-18808-Ljbffr Micron Technology, Inc
    Internship

    Micron Technology, Inc

    California, MO
    2 days ago
  •  ...Oracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve the low-level systems software and platform-management capabilities that power next-generation GPU infrastructure. This is a hands-on systems engineering role... 

    Ll Oefentherapie

    California, MO
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Inference Engineer - High-Performance GPU Systems. Be the first to apply!