Senior GPU Cloud Production Engineer — Kubernetes & Reliability
$272k - $431.25kNVIDIA
NVIDIA is seeking a Principal Software Engineer to define architecture and lead operations of large-scale GPU clusters. The role requires strong expertise in Kubernetes, Linux, and infrastructure automation. Candidates should have over 15 years of experience with distributed systems, proven leadership in technical initiatives, and a degree in Computer Science. A competitive salary between $272,000 and $431,250 based on experience and location will be offered. #J-18808-Ljbffr NVIDIA
$184k - $356.5k
NVIDIA seeks a Senior Software Engineer to build and operate a large-scale GPU infrastructure for AI workloads. You'll automate large-scale GPU... ...tools, and improve workflows within a production engineering team focused on Kubernetes. With 8+ years of experience required,...Senior- ...Energy Systems based in Sunnyvale, California, seeks a highly skilled engineer experienced in Kubernetes. You will architect and operate scalable Kubernetes clusters, ensuring performance and reliability across global datacenters. The ideal candidate has over 7 years in...Senior
- A leading cloud service provider in California is seeking a Senior Platform Engineer II to champion reliability and administer multi-tenant Kubernetes platforms. The ideal candidate will have over 5 years of experience in Kubernetes, a strong background in resilient application...Senior
- ...Clara is seeking a skilled software engineer to develop a cloud-native stack for managing data infrastructures... ...and shipping services around Kubernetes and engaging with cutting-edge technologies... ...and contribute to advances in AI, GPU technology, and high-performance...Senior
$184k - $356.5k
NVIDIA is seeking a Senior Software Engineer to build automation for GPU clusters in Santa Clara, California. The ideal... ...have 8+ years of experience in production infrastructure and strong... ...automation and improving infrastructure reliability. The position offers a base...Senior- ...California is seeking a hands-on Site Reliability Engineer (SRE) to design and manage their next-generation... ...of experience, expertise in Linux and cloud technologies, and a passion for solving... ...and scale, managing large multi-cloud GPU clusters, and ensuring robust security...Senior
- Crusoe Energy Systems, located in Sunnyvale, California, is seeking a Principal Engineer to lead the reliability and scalability of its cloud infrastructure. The role involves ensuring operational excellence across compute, storage, and networking systems, while mentoring...Senior
- Production engineering is a field that involves crafting, building, and... ...along with open‑source cloud‑enabling technologies such as Kubernetes, containers, and... ...storage architectures are reliable, scalable, and efficient... ...internal and external‑facing GPU cloud services meet...SeniorFlexible hours
- A technology services company is seeking a Senior Site Reliability Engineer / DevOps Engineer in Sunnyvale, CA. The ideal candidate will have over... ...years of experience in DevOps, expertise in Docker and Kubernetes, and proficiency with Terraform or Ansible. Responsibilities...Senior
- NVIDIA CFR is seeking a hands‑on senior engineer to own the lifecycle and automation of the Kubernetes platform that supports GNI... ...network systems. You will provide production support for platform‑hosted... ...and Bangalore teams to deliver reliable, scalable Kubernetes solutions...Senior
$165k - $242k
A cloud service provider is seeking a Senior Software Engineer II for their Inference team in Sunnyvale, California. In this... ..., and improve service reliability. The ideal candidate has extensive... ..., strong Python/Go skills, and Kubernetes expertise. This position offers...Senior$156k - $190k
...in Sunnyvale, CA, is seeking a Staff Cloud Support Engineer to provide technical leadership in cloud... ...will lead incident responses, design reliability architecture, and mentor team members... ...experience and expertise in Linux, Kubernetes, and networking. We offer a...Senior$204k - $247k
...Systems in Sunnyvale, California is seeking a skilled engineer specialized in Kubernetes to build and operate our next-generation platform. The... ...architecting scalable Kubernetes clusters and ensuring their reliability, while also mentoring team members. Applicants should...- NVIDIA is seeking a Senior Systems Software Engineer to improve performance, reliability, and scalability of a system routing AI workloads across distributed GPU fleets. You’ll work on a polyglot platform with open source components, contributing to control plane and edge...Senior
- NVIDIA is hiring engineers to scale up its AI... ...platform that automates GPU asset provisioning... ...management across cloud providers.... ...industry-leading reliability, availability, and... ...experience on large-scale production systems. BS in... ...systems such as Kubernetes, SLURM. Understanding...Senior
$184k - $287.5k
The DGX Cloud organization at NVIDIA brings... ...of innovative engineers dedicated to solving... ...for an outstanding Senior Systems Software Engineer... ...such as Kubernetes and containers, and... ...the stack - from GPU operator and device... ...large scale, ensuring reliability and efficiency....SeniorWorldwide- ...over 10 times faster than GPU‑based hyperscale cloud inference services. This order... ...powered by the Wafer‑Scale Engine (WSE). This team will help deliver world‑class, ultra‑reliable inference infrastructure... ...familiarity with our current stack, production pain points, and high‑...Shift work
$108k - $172.5k
...computing. An era in which our GPU acts as the brains of... ...We are seeking a motivated Senior HPC Support Engineer - Compute/GPU (DGX Platform... ...of AI hardware and software products. As a primary point of contact... ...(i.e. Docker, Kubernetes). You'd have nurtured a deep...SeniorWork experience placement- Clover is seeking a Sr. DevOps / Cloud Engineer for Billing Infrastructure in Sunnyvale, CA.... ...closely with finance, compliance, and product teams to deliver scalable, resilient cloud... ...backgrounds, experience with Kubernetes, Terraform, and IaC, and a proven track...Senior
$300 per month
...center construction, and cloud services. If you... ...Role: At Crusoe, the Production Engineering team plays a critical... ...ensuring the performance, reliability, and scalability of... ...support large‑scale GPU clusters and latency‑... ...platforms such as Kubernetes or Docker Strong incident...SeniorTemporary work$272k - $431.25k
NVIDIA DGX Cloud is scaling GPU infrastructure across internal... ...Software Engineers to help shape the... ...technical direction for production engineering, Kubernetes-based operations, automation, and reliability across large-scale... ...This role is for senior technical leaders...$165k - $242k
CoreWeave in Sunnyvale, California is looking for a Senior Engineer to enhance their GPU performance testing platform. The role involves designing solutions for infrastructure testing, developing Kubernetes operators, and creating scalable back-end services. The ideal...Senior- ...is seeking an experienced Senior DevOps Engineer in Santa Clara, California.... ...should have strong expertise in cloud infrastructure and CI/CD... ...in technologies like Kubernetes, AWS, Docker, and Terraform... ...significantly to scalable and reliable deployments within a high-performing...Senior
$300 per month
...intelligence. We’re crafting the engine that powers a world... ..., transformative cloud infrastructure. About... ...Role: We are seeking a Production Engineer to play a critical... ...in transitioning to Kubernetes and optimizing Crusoe's... ...hardware experience and GPU troubleshooting...Temporary work$159k - $231k
Senior Robotics Automation Engineer, Platforms Infrastructure This is a specialized role... ...engineering and robotics product development. Experience deploying... ...scale, efficiency, reliability and velocity. Our customers... ...include Googlers, Google Cloud customers, and billions of...SeniorContract workWorldwide$236k - $330k
...development and processing of engineering hardware must be performed on... ...teams, and robotics product development. Experience in robotic... ...unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google...SeniorContract workRemote workWorldwideFlexible hours$145k - $165k
A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key...Senior- Fortinet is seeking a talented Site Reliability Engineer to join our engineering team in the United States. This hands‑on role focuses on building, maintaining, and troubleshooting cloud service clusters, infrastructure, and monitoring systems to ensure high availability...Senior
- ...in Sunnyvale, California, is seeking a Staff Software Infrastructure Engineer. This critical role involves managing cloud infrastructure, developing automation tools, and transitioning to Kubernetes while ensuring efficient server provisioning and scaling operations. The...
$152k - $287.5k
NVIDIA is seeking a Senior AI Infrastructure Engineer for our DGX Cloud group. This role involves designing, building, and maintaining large-scale production systems, ensuring GPU cloud services perform with maximum reliability. The ideal candidate will have a strong background...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior GPU Cloud Production Engineer — Kubernetes & Reliability. Be the first to apply!
- principal cloud engineer Santa Clara, CA
- cloud engineer Santa Clara, CA
- senior devops cloud engineer Santa Clara, CA
- senior aws cloud engineer Santa Clara, CA
- senior cloud data engineer Santa Clara, CA
- cloud network engineer Santa Clara, CA
- informatica cloud developer Santa Clara, CA
- cloud operations engineer Santa Clara, CA
- google cloud architect Santa Clara, CA
- aws cloud architect Santa Clara, CA
