Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Platform Engineering Manager

NVIDIA AI

NVIDIA is seeking a highly capable Engineering Manager to lead the next generation of LLM/VLM inference software. You will architect and guide a team of engineers, interfacing with researchers and GPU architects to deliver production-grade software that sets the standard for AI performance. You’ll drive design and execution across distributed teams, focus on cutting-edge CUDA-enabled tooling, and ensure scalable, user-friendly APIs for AI developers. #J-18808-Ljbffr NVIDIA AI

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the LLM Inference Platform Engineering Manager in Santa Clara, CA vacancy
  • A leading technology company in California is seeking an Engineering Manager to lead the development of cutting-edge LLM/VLM technologies. In this hands-on leadership role, you will manage a team responsible for optimizing runtime and frameworks, while collaborating with... 
    Suggested

    NVIDIA Corporation

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

    NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of...  ...of large-scale models for LLM, multimodal, and generative AI...  ...optimization across CPU and GPU platforms.Proven ability to mentor... 
    Suggested
    Full time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...NVIDIA’s open-source benchmarking platform, AIPerf, is the growing standard for assessing LLM serving performance across various inference frameworks. Hyperscalers, cloud...  ...and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA’s Dynamo... 
    Suggested
    Full time
    Local area
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $207k - $301k

    People Management and Talent Development: Lead, mentor,...  ...team of systems and ML engineers. Drive a culture of...  ...strategy for enhancing the LLM serving stack,...  ...learning infrastructure, AI platforms, or high-performance...  ...Distributed Cloud (DSC) AI Inference Platform team operates... 
    Suggested

    Google

    Sunnyvale, CA
    1 day ago
  • $224k - $356.5k

     ...the AI revolution—we're accelerating it. We are accelerating LLM inference across the stack and across all open source LLM frameworks...  ...our team. We're seeking a highly skilled and driven Engineering Manager to take the lead in accelerating the next generation of LLM... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $245k - $325k

     ...SambaNova Systems is seeking a Director of Software Engineering to lead the SambaStack platform engineering team. This role involves ensuring the delivery of reliable AI inference services while managing a high-performing group of engineers. Candidates should have over... 

    Jobleads-US

    San Jose, CA
    2 days ago
  • $224k - $356.5k

    NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for observing, debugging, and optimizing...  ...visibility into model behavior, inference performance, reliability, and cost...  ...signals across large scale LLM and VLM deployments. It will help... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • NVIDIA in Santa Clara, CA seeks a senior software engineer focusing on GPU computing and ML inference to optimize LLM workloads on edge AI hardware. You will track open-source inference frameworks, map architectures to NVIDIA GPUs, and report performance metrics. With 1... 

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $157.2k - $254.1k

     ...most comprehensive AI security platform. Organizations are...  ...delivers model security, posture management, AI red teaming, and runtime...  ...Principal Machine Learning Inference Engineer, you will serve as a technical...  ...Demonstrated expertise with modern LLM inference engines (e.g.,... 
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    3 days ago
  •  ...ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team...  ...AI inference workloads on AMD GPU platforms. You will contribute to optimizing latency...  ...throughput, and cost efficiency for LLM and multimodal model serving in... 

    AMD

    San Jose, CA
    18 hours ago
  • $272k - $425.5k

    A leading technology firm in Santa Clara seeks a Principal Software Engineer for Large-Scale LLM Memory and Storage Systems. The role involves designing a unified memory layer for large-scale inference, integrating with LLM serving engines, and mentoring engineers. Candidates... 
    Remote job

    NVIDIA Corporation

    Santa Clara, CA
    8 hours ago
  • $240k - $260k

     ...Job Description Job Description AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform manages and governs human and non-human access to all...  ...training on Ray + H100s, the multi-engine LLM inference mesh (vLLM, SGLang, NVIDIA Triton... 

    Saviynt

    Milpitas, CA
    a month ago
  • $139k - $204k

    CoreWeave is seeking a Senior Engineer to lead designs and enhance engineering standards within their Kubernetes-native inference platform. Responsibilities include driving architecture, defining SLIs/SLOs, and mentoring engineers, with 3-8 years of experience preferred... 
    Remote job
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    1 day ago
  • $165k - $242k

    A cloud service provider is seeking a Senior Software Engineer II for their Inference team in Sunnyvale, California. In this role, you'll lead design reviews, implement optimizations, and improve service reliability. The ideal candidate has extensive experience with distributed... 

    CoreWeave

    Sunnyvale, CA
    8 hours ago
  • $201.3k - $352.3k

     ...DescriptionIt all started when engineer Fred Luddy wrote code...  .... Our ServiceNow AI platform brings together any AI...  ..., and product managers with a dual mission. We...  ...window efficiency, and inference costs.Lead a High-Performing...  ...ROUGE, BLEU, G-Eval, LLM-as-a-judge patterns,... 
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours
    Shift work

    ServiceNow

    Santa Clara, CA
    2 days ago
  • $201k - $402k

     ...OpportunityWe are looking for a Senior Engineering Manager to lead the design, development, and...  ...scaling of a next-generation Kubernetes platform powering enterprise environments. This...  ...of GPU scheduling, training, and inference systemsExecution & Operational ExcellenceTrack... 
    Work at office
    Local area
    Remote work
    Relocation package
    3 days per week

    Nutanix

    San Jose, CA
    4 days ago
  • $216k - $345k

     ...set strategy and drive execution across a portfolio of platforms, manage a team of experienced engineers and technical program managers, oversee vendor...  ...the crowd:Experience with agentic marketing workflows, LLM automation, audience segmentation, or AI-assisted campaign... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based...  ...inference. About the Role We're hiring a Software Engineer to help contribute to projects on our Inference Platform team. Our team primarily owns the orchestration... 
    Full time

    Cerebras Systems

    Sunnyvale, CA
    18 hours ago
  • Zoomcar is seeking a Principal Software Development Engineer in Sunnyvale to architect and implement functions for monitoring LLM requests and filter for prompt injection attacks. You will collaborate across teams to ensure secure and efficient system designs while adhering... 

    Zoomcar

    Sunnyvale, CA
    2 days ago
  • $152k - $241.5k

     ...computing company”.We are looking for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning &...  ...advancement of AI, our DLC has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices,... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $280k - $350k

     ...of top researchers and engineers, building the world’s top...  ..., optimizing realtime inference, and creating best-in-...  ...as a Staff / Principal Platform Engineer and take end-to...  ...for our TTS and LLM Router.Drive engineering...  ...performance of services.Manage pipelines to ensure smooth... 
    Full time
    Work at office
    Relocation

    Inworld AI

    Mountain View, CA
    18 hours ago
  •  ...looking for a systems-minded engineer who lives at the intersection of large-scale model inference, distributed systems, and performance...  ..., KV cache lifecycle management, and efficient offloading mechanisms...  ...and deeply understand modern LLM inference frameworks,... 

    AMD

    San Jose, CA
    1 day ago
  • $224k - $356.5k

     ...NVIDIA is the platform upon which every new AI-powered application...  ...deeply technical software manager to lead production AI inference for NVIDIA Inference...  ...combining optimized inference engines, model profiles/recipes,...  ...shipping production‑ready LLM NIMs, including planning,... 

    Socket.dev

    Santa Clara, CA
    2 days ago
  •  ...allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale...  ...inference.About The RoleWe're hiring a Principal Engineer for our Inference Cloud Platform. This team owns the cloud layer behind our Inference Service... 

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $262k - $364k

     ...Lead and coach a distributed engineering team, fostering innovation while...  ...deliverables.Large-Scale and LLM Architecture: Design and...  ...transport across various NICs and platforms, tuning the communication...  ...staying ahead of training and inference advancements.Collaborative System... 
    Remote work
    Worldwide

    Google

    Sunnyvale, CA
    3 days ago
  • $245k - $295k

     ...Crusoe. About the Role We are seeking a Senior Manager, Infrastructure Platform Engineering to lead a team building core systems that turn large-...  ...challenges of GPU clusters, AI training, and inference workloads Working knowledge of platform security and... 
    Temporary work
    Immediate start

    Crusoe

    Sunnyvale, CA
    a month ago
  • $296.3k - $374.8k

     ...received.Meet the TeamThe AI Platforms and Enablement team is the engine behind Cisco's internal...  ..., agents, harnesses, inference efficiency, and enterprise...  ...internal AI infrastructure, managing significant platform...  ...or operating production LLM, agentic AI, model-serving... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    3 days ago
  • Everpure is seeking a talented engineer in Santa Clara, CA to lead the Flashblade//EXA platform engineering team. This role involves building foundational architecture...  ...software execution, collaborate with product management, and manage vendor integration. Successful... 

    Everpure

    Santa Clara, CA
    1 day ago
  • $228.1k - $342.8k

     ...leading technology company in Cupertino seeks a Technical Leader to oversee storage infrastructure. This role involves managing a team of engineers, collaborating across departments, and focusing on innovation in distributed systems. Ideal candidates will have 10+ years... 

    Apple

    Cupertino, CA
    4 days ago
  • Google is seeking a Software Engineering Manager to lead multiple teams across locations, guiding technical strategy and project delivery while...  ..., with a focus on scalable systems across Google's platforms. In this role, you will manage engineers across teams, oversee... 

    Google

    Sunnyvale, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Platform Engineering Manager. Be the first to apply!