Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior GPU Cloud Production Engineer — Kubernetes & Reliability

$272k - $431.25k

NVIDIA

NVIDIA is seeking a Principal Software Engineer to define architecture and lead operations of large-scale GPU clusters. The role requires strong expertise in Kubernetes, Linux, and infrastructure automation. Candidates should have over 15 years of experience with distributed systems, proven leadership in technical initiatives, and a degree in Computer Science. A competitive salary between $272,000 and $431,250 based on experience and location will be offered. #J-18808-Ljbffr NVIDIA

Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the Senior GPU Cloud Production Engineer — Kubernetes & Reliability in Santa Clara, CA vacancy
  • $184k - $356.5k

    NVIDIA seeks a Senior Software Engineer to build and operate a large-scale GPU infrastructure for AI workloads. You'll automate large-scale GPU...  ...tools, and improve workflows within a production engineering team focused on Kubernetes. With 8+ years of experience required,... 
    Senior

    NVIDIA

    Santa Clara, CA
    3 days ago
  •  ...Energy Systems based in Sunnyvale, California, seeks a highly skilled engineer experienced in Kubernetes. You will architect and operate scalable Kubernetes clusters, ensuring performance and reliability across global datacenters. The ideal candidate has over 7 years in... 
    Senior

    Crusoe Energy Systems

    Sunnyvale, CA
    16 hours ago
  • A leading cloud service provider in California is seeking a Senior Platform Engineer II to champion reliability and administer multi-tenant Kubernetes platforms. The ideal candidate will have over 5 years of experience in Kubernetes, a strong background in resilient application... 
    Senior

    CoreWeave

    Sunnyvale, CA
    16 hours ago
  •  ...Clara is seeking a skilled software engineer to develop a cloud-native stack for managing data infrastructures...  ...and shipping services around Kubernetes and engaging with cutting-edge technologies...  ...and contribute to advances in AI, GPU technology, and high-performance... 
    Senior

    NVIDIA

    Santa Clara, CA
    16 hours ago
  • $184k - $356.5k

    NVIDIA is seeking a Senior Software Engineer to build automation for GPU clusters in Santa Clara, California. The ideal...  ...have 8+ years of experience in production infrastructure and strong...  ...automation and improving infrastructure reliability. The position offers a base... 
    Senior

    2100 NVIDIA USA

    Santa Clara, CA
    3 days ago
  •  ...California is seeking a hands-on Site Reliability Engineer (SRE) to design and manage their next-generation...  ...of experience, expertise in Linux and cloud technologies, and a passion for solving...  ...and scale, managing large multi-cloud GPU clusters, and ensuring robust security... 
    Senior

    Luma AI

    Palo Alto, CA
    16 hours ago
  • Crusoe Energy Systems, located in Sunnyvale, California, is seeking a Principal Engineer to lead the reliability and scalability of its cloud infrastructure. The role involves ensuring operational excellence across compute, storage, and networking systems, while mentoring... 
    Senior

    Crusoe Energy Systems

    Sunnyvale, CA
    16 hours ago
  • Production engineering is a field that involves crafting, building, and...  ...along with open‑source cloud‑enabling technologies such as Kubernetes, containers, and...  ...storage architectures are reliable, scalable, and efficient...  ...internal and external‑facing GPU cloud services meet... 
    Senior
    Flexible hours

    NVIDIA

    Santa Clara, CA
    16 hours ago
  • A technology services company is seeking a Senior Site Reliability Engineer / DevOps Engineer in Sunnyvale, CA. The ideal candidate will have over...  ...years of experience in DevOps, expertise in Docker and Kubernetes, and proficiency with Terraform or Ansible. Responsibilities... 
    Senior

    Donato Technologies Inc

    Sunnyvale, CA
    1 day ago
  • NVIDIA CFR is seeking a hands‑on senior engineer to own the lifecycle and automation of the Kubernetes platform that supports GNI...  ...network systems. You will provide production support for platform‑hosted...  ...and Bangalore teams to deliver reliable, scalable Kubernetes solutions... 
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $165k - $242k

    A cloud service provider is seeking a Senior Software Engineer II for their Inference team in Sunnyvale, California. In this...  ..., and improve service reliability. The ideal candidate has extensive...  ..., strong Python/Go skills, and Kubernetes expertise. This position offers... 
    Senior

    CoreWeave

    Sunnyvale, CA
    4 days ago
  • $156k - $190k

     ...in Sunnyvale, CA, is seeking a Staff Cloud Support Engineer to provide technical leadership in cloud...  ...will lead incident responses, design reliability architecture, and mentor team members...  ...experience and expertise in Linux, Kubernetes, and networking. We offer a... 
    Senior

    Crusoe Energy Systems

    Sunnyvale, CA
    16 hours ago
  • $204k - $247k

     ...Systems in Sunnyvale, California is seeking a skilled engineer specialized in Kubernetes to build and operate our next-generation platform. The...  ...architecting scalable Kubernetes clusters and ensuring their reliability, while also mentoring team members. Applicants should... 

    Crusoe Energy Systems

    Sunnyvale, CA
    16 hours ago
  • NVIDIA is seeking a Senior Systems Software Engineer to improve performance, reliability, and scalability of a system routing AI workloads across distributed GPU fleets. You’ll work on a polyglot platform with open source components, contributing to control plane and edge... 
    Senior

    NVIDIA

    Santa Clara, CA
    16 hours ago
  • NVIDIA is hiring engineers to scale up its AI...  ...platform that automates GPU asset provisioning...  ...management across cloud providers....  ...industry-leading reliability, availability, and...  ...experience on large-scale production systems. BS in...  ...systems such as Kubernetes, SLURM. Understanding... 
    Senior

    NVIDIA

    Santa Clara, CA
    16 hours ago
  • $184k - $287.5k

    The DGX Cloud organization at NVIDIA brings...  ...of innovative engineers dedicated to solving...  ...for an outstanding Senior Systems Software Engineer...  ...such as Kubernetes and containers, and...  ...the stack - from GPU operator and device...  ...large scale, ensuring reliability and efficiency.... 
    Senior
    Worldwide

    NVIDIA

    Santa Clara, CA
    16 hours ago
  •  ...over 10 times faster than GPU‑based hyperscale cloud inference services. This order...  ...powered by the Wafer‑Scale Engine (WSE). This team will help deliver world‑class, ultra‑reliable inference infrastructure...  ...familiarity with our current stack, production pain points, and high‑... 
    Shift work

    Cerebras

    Sunnyvale, CA
    16 hours ago
  • $108k - $172.5k

     ...computing. An era in which our GPU acts as the brains of...  ...We are seeking a motivated Senior HPC Support Engineer - Compute/GPU (DGX Platform...  ...of AI hardware and software products. As a primary point of contact...  ...(i.e. Docker, Kubernetes). You'd have nurtured a deep... 
    Senior
    Work experience placement

    Nvidia Corporation

    Santa Clara, CA
    16 hours ago
  • Clover is seeking a Sr. DevOps / Cloud Engineer for Billing Infrastructure in Sunnyvale, CA....  ...closely with finance, compliance, and product teams to deliver scalable, resilient cloud...  ...backgrounds, experience with Kubernetes, Terraform, and IaC, and a proven track... 
    Senior

    Clover

    Sunnyvale, CA
    16 hours ago
  • $300 per month

     ...center construction, and cloud services. If you...  ...Role: At Crusoe, the Production Engineering team plays a critical...  ...ensuring the performance, reliability, and scalability of...  ...support large‑scale GPU clusters and latency‑...  ...platforms such as Kubernetes or Docker Strong incident... 
    Senior
    Temporary work

    Crusoe Energy Systems

    Sunnyvale, CA
    16 hours ago
  • $272k - $431.25k

    NVIDIA DGX Cloud is scaling GPU infrastructure across internal...  ...Software Engineers to help shape the...  ...technical direction for production engineering, Kubernetes-based operations, automation, and reliability across large-scale...  ...This role is for senior technical leaders... 

    NVIDIA

    Santa Clara, CA
    16 hours ago
  • $165k - $242k

    CoreWeave in Sunnyvale, California is looking for a Senior Engineer to enhance their GPU performance testing platform. The role involves designing solutions for infrastructure testing, developing Kubernetes operators, and creating scalable back-end services. The ideal... 
    Senior

    CoreWeave

    Sunnyvale, CA
    16 hours ago
  •  ...is seeking an experienced Senior DevOps Engineer in Santa Clara, California....  ...should have strong expertise in cloud infrastructure and CI/CD...  ...in technologies like Kubernetes, AWS, Docker, and Terraform...  ...significantly to scalable and reliable deployments within a high-performing... 
    Senior

    KlearNow Ltd.

    Santa Clara, CA
    16 hours ago
  • $300 per month

     ...intelligence. We’re crafting the engine that powers a world...  ..., transformative cloud infrastructure. About...  ...Role: We are seeking a Production Engineer to play a critical...  ...in transitioning to Kubernetes and optimizing Crusoe's...  ...hardware experience and GPU troubleshooting... 
    Temporary work

    Crusoe Energy Systems

    Sunnyvale, CA
    16 hours ago
  • $159k - $231k

    Senior Robotics Automation Engineer, Platforms Infrastructure This is a specialized role...  ...engineering and robotics product development. Experience deploying...  ...scale, efficiency, reliability and velocity. Our customers...  ...include Googlers, Google Cloud customers, and billions of... 
    Senior
    Contract work
    Worldwide

    Google

    Sunnyvale, CA
    16 hours ago
  • $236k - $330k

     ...development and processing of engineering hardware must be performed on...  ...teams, and robotics product development. Experience in robotic...  ...unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google... 
    Senior
    Contract work
    Remote work
    Worldwide
    Flexible hours

    Google

    Sunnyvale, CA
    16 hours ago
  • $145k - $165k

    A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    4 days ago
  • Fortinet is seeking a talented Site Reliability Engineer to join our engineering team in the United States. This hands‑on role focuses on building, maintaining, and troubleshooting cloud service clusters, infrastructure, and monitoring systems to ensure high availability... 
    Senior

    Fortinet

    Sunnyvale, CA
    16 hours ago
  •  ...in Sunnyvale, California, is seeking a Staff Software Infrastructure Engineer. This critical role involves managing cloud infrastructure, developing automation tools, and transitioning to Kubernetes while ensuring efficient server provisioning and scaling operations. The... 

    Crusoe Energy Systems

    Sunnyvale, CA
    16 hours ago
  • $152k - $287.5k

    NVIDIA is seeking a Senior AI Infrastructure Engineer for our DGX Cloud group. This role involves designing, building, and maintaining large-scale production systems, ensuring GPU cloud services perform with maximum reliability. The ideal candidate will have a strong background... 
    Senior

    NVIDIA

    Santa Clara, CA
    16 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior GPU Cloud Production Engineer — Kubernetes & Reliability. Be the first to apply!