Senior Platform and EngOps Engineer - Cluster Operations (Santa Clara)
$176k - $276kNvidia
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team of innovative engineers who develop and maintain software facilitating GPU communication, driving groundbreaking solutions in High Performance Computing and Deep Learning. We're looking for highly motivated EngOps and Platform Engineers to boost execution efficiency while managing and maintaining large GPU clusters interconnected via NVLink and InfiniBand.What you will be doing:Develop automated tools to efficiently deploy, provision, and maintain extensive GPU clusters interconnected via NVLink and InfiniBandImplement modern DevOps tools to automate software updates, perform maintenance tasks, and monitor cluster availability, ensuring seamless operations.Take ownership of daily cluster failures and issues, troubleshooting them promptly to maintain optimal cluster availability and performance.Manage the rollout and rollback of cluster software and firmware updates, ensuring smooth transitions and minimal disruptions.Collaborate effectively with dynamic Engineering and Product Teams across multiple time zones to align cluster operations with evolving project requirements.What we need to see:BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.8+ years of hands-on experience in deploying and administrating clusters, servers, switches, and related infrastructure.Automation expert with hands on skills in Ansible, Python and Shell Scripting.Deep understanding of operating systems, computer networks, and high-performance applications.Proven ability to work effectively with developers and test engineers across different teams and time zones.Proficient with Linux fundamentals.Ways to stand out from the crowd:Familiarity with resource scheduling managers, preferably Slurm.Direct experience with industry standard alerting tools and emergency response practices.Hands-on experience with GPU-focused hardware and software, such as DGX systems and Compute Clusters.Proficiency in crafting and implementing a robust metrics collection and alerting infrastructure.Proficiency in designing large scale networking technologies and the associated challenges.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 24, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$184k - $287.5k
...Cloud is building and operating large-scale GPU... ...We are looking for Senior Software Engineers to help build the automation... ...that make GPU clusters reliable, scalable,... ...up work.Partner with platform, storage, networking... ...SummaryLocation: US, CA, Santa Clara; US, RemoteType:...PlatformOperationsSeniorFull timePart time$332k
...centers running AI training clusters to autonomous vehicles operating in safety-critical... ...Reliability, you will set the engineering standard for how NVIDIA's... ..., Embedded, and emerging platforms. Ensure every product... ...law.SummaryLocation: US, CA, Santa ClaraType: Full time...PlatformOperationsSeniorFull timePart timeWork at office$272k - $431.25k
...NVIDIA is seeking a Senior MLOps Engineering Manager to join our Autonomous Driving organization in Santa Clara, CA. This role offers an outstanding... ...build, development, and operation of large‑scale, end‑to‑end... ..., CI/CD, and data platforms.What We Need to See:Bachelor...PlatformOperationsSeniorFull timePart time$168k - $258.75k
...infrastructure? At NVIDIA, we seek a Senior MGX Ecosystem Engineer to play a crucial role in... ...you will influence future platform architectures and... ...management, engineering, operations, manufacturing, and business... ...law.SummaryLocation: US, CA, Santa ClaraType: Full time...PlatformOperationsSeniorFull timePart time$184k - $287.5k
...NVIDIA is seeking an NCX Senior Engineer to join our DSX team,... ...from NVIDIA's AI platform across varied environments... ...latency, cost, and operational risk.Implement and... ...configuration of GPU‑accelerated clusters and NCP building... ...: US, CA, Santa Clara; US, Remote; US, WA,...PlatformSeniorFull timePart timeRemote work$242.35k - $363k
...uncompromising performance.As a Senior Distinguished Engineer, you will be one of... ...and global network operators. Here’s why top... ...large-scale GPU clusters and AI fabrics.•... ...next-gen switching platforms serving AI clusters... ...collaboration in Santa Clara—where silicon,...PlatformSeniorPermanent employmentPart timeInternshipWork from home- ...Senior Director, CTIO Engineering TechnologistsInfrastructure Solutions... ...continuity, and deliver the platforms, applications, and... ...Austin, Texas or Santa Clara, California.What... ...the strategic and operational objectives of their... ...compute) clusters, AI compute, AI Datacenter...PlatformSeniorPart time
$168k - $264.5k
...architecture, design, marketing, operations, and productization,... ....We are hiring a Senior Silicon Power and Thermal Controllers Engineer to design, implement,... ...where silicon, firmware, platform, and software meet, and... ...: US, CA, Santa ClaraType: Full time...PlatformOperationsSeniorFull timePart timeShift work$224k - $356.5k
...are looking for a Senior Software Architect... ...designing server platforms and has added understanding... ...with world class engineering teams, product management, Operations and Customer... ...of cloud and cluster level deployment and... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin;...PlatformOperationsSeniorFull timePart timeRemote workShift work$208k - $327.75k
...NVIDIA Enterprise Platforms Group is seeking a Senior System Architect to... ...automation, and AI/ML operations. This person can... ...native-platform and cluster reference builds.Architect... ...Architects, Engineering, Solution Architecture... ...: US, CA, Santa Clara; US, MA, Westford;...PlatformOperationsSeniorFull timePart time$184k - $287.5k
...forward‑thinking engineers tackling some of... ...searching for a Senior Systems Software... ...stack, including GPU operators, device plugins,... ...and major cloud platforms. You’ll own hard... ...operation at hyperscale cluster sizes, doing in... ...: US, CA, Santa Clara; US, WA, SeattleType...PlatformOperationsSeniorFull timePart timeRemote work$184k - $287.5k
...building an innovative Data Platform that employs advanced... ..., actionable cluster detection, effective... ...power of Data, AI, and Operations Research to help deliver... ...-specific feature engineering for real-time cloud gaming... ...: US, CA, Santa Clara; US, CA, RemoteType:...PlatformOperationsSeniorFull timePart time$184k - $287.5k
...provisioning and management. As a Senior Software Engineer - Datacenter Systems, you... ...large-scale GPU clusters connected through NVLink and... ...management tools, making cluster operations and visibility more... ....SummaryLocation: US, CA, Santa Clara; US, CA, Remote; US, AZ, RemoteType...OperationsSeniorFull timePart timeRemote work$134.39k - $201.3k
...highly integrated component platforms with high-speed Silicon photonics... .... These components are made operational with highly functional... ...firmware, and control loops. An engineer in this team will be involved... ...to employment.#LI-NF1SummaryLocation: US-CA - Santa ClaraType:...PlatformOperationsSeniorPermanent employmentPart timeInternshipWork from home$240k - $379.5k
...organization builds and operates the AI infrastructure... ....Partner with engineering, product, operations,... ...systems, or large-scale platform operations.Strong communication... ...AI/ML platforms, GPU clusters, or large-scale cloud... ...: US, CA, Santa Clara; US, WA, SeattleType:...PlatformOperationsSeniorFull timePart time$184k - $287.5k
...We are looking for a Senior Solutions Architect specializing... ...expert uniting engineering, field teams, and... ...high-performance clusters and establish performance... ...engineering and operations teams.Ways to stand out... ...SummaryLocation: US, CA, Santa Clara; US, TX, AustinType:...OperationsSeniorFull timePart time$272k - $431.25k
...NVIDIA's Object Storage Platform team builds and operates the company's... ...enabling researchers and engineers to reliably store... ...data closer to GPU clusters, minimizing idle... ...engineering organization to senior leadership,... ...: US, CA, Santa Clara; US, CA, RemoteType...PlatformOperationsSeniorFull timePart time$272k - $431.25k
...a Principal Software Engineer to join our CSP Engagements... ...deploy, monitor, and operate these systems reliably... ...in system software, platform firmware, or large-... ...management software, cluster management, or system-... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR...PlatformOperationsFull timePart timeRemote workShift work$177.82k - $266.4k
...Marvell silicon into deployable platforms by working shoulder-to-shoulder with customer engineering teams from early architecture... ....What You Can ExpectAs Senior Principal Engineer for Signal... ...Board/System HW, Software, and Operations teams to deliver fully integrated...PlatformOperationsSeniorPart timeInternshipWork from home$184k - $287.5k
...We are looking for a Senior Software Engineer to lead the bring-up... ...across NVIDIA GPU platforms at the largest scales... ...capabilities that keep large clusters productive. This is... ...for an engineer who operates at the intersection... ...: US, CA, Santa Clara; US, TX, Austin; US,...PlatformOperationsSeniorFull timePart timeRemote work$168k - $258.75k
...'s deep learning platforms are at the forefront... ...teams (Software Engineering, Production... ...and Data Center Operations) and their leadership... ...operational workflows, and cluster/capacity bring-up... ..., and align senior multi-functional... ...: US, CA, Santa Clara; US, WA, SeattleType...PlatformOperationsSeniorFull timePart timeRemote work$248k - $396.75k
...product leaders, and platform teams across NVIDIA and... ...product, sales, engineering, architecture, marketing... ...telemetry, and fleet operations.Build a robust opportunity... ..., and large-scale AI clusters.Foundation knowledge... ...: US, CA, Santa Clara; US, WA, SeattleType:...PlatformOperationsSeniorFull timePart time$184k - $287.5k
...OEM/SI partners, Operations, NVIDIA customer... ...building trust with senior audiences, and... ...Computer Science, Engineering, or a similar technical... ...systems, AI platforms, networking,... ...large-scale GPU clusters, HPC systems, networking... ...: US, CA, Santa Clara; US, TX, Remote;...PlatformOperationsSeniorFull timePart timeCasual workRemote work$176k - $276k
...deploy, integrate, and operate the Kubernetes-based platform and shared services used... ...this platform, including cluster provisioning and upgrades... ...looking for a hands-on senior engineer to own the lifecycle and... ...SummaryLocation: US, CA, Santa Clara; US, IL, Remote; US, WA,...PlatformOperationsSeniorFull timePart timeRemote workWeekend work$174.72k - $295.68k
...and data visualization. Support business operations such as autonomous driving, smart... ...Responsible for building a data management platform covering the entire process from data collection... ...higher in Computer Science, Software Engineering, Artificial Intelligence, or related...PlatformOperationsSeniorFull timePart timeOverseas$224k - $356.5k
...work.The data center platforms like GB200 NVL72 by NVIDIA... ...technical leader to engineer and propel innovation... ...deployments, and field operations.What You'll Be Doing:... ...across rack-level or cluster-level deployments.Background... ...: US, CA, Santa Clara; US, OR, HillsboroType...PlatformOperationsSeniorFull timePart time$272k - $431.25k
...Principal Software Engineers to help shape the technical... ..., Kubernetes-based operations, automation, and... ...large-scale GPU clusters.This role is for senior technical leaders... ...engineers and influence platform, infrastructure,... ...: US, CA, Santa Clara; US, RemoteType: Full...PlatformOperationsFull timePart time$184k - $287.5k
...are seeking a self‑motivated senior engineer for the Aerial Omniverse... ...‑class GPU and ray‑tracing platforms.What you'll be doing:As a member... ...ray‑tracing engine that operates at two time scales. At the... ...law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full...PlatformSeniorFull timePart time$184k - $287.5k
...We are seeking for an expert Senior Compiler Engineer to join our Compute Compiler Team, with a focus... ...ecosystem health with NVIDIA’s platform goalsAct as a technical ambassador for... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, TX, Remote; US,...PlatformSeniorFull timePart timeRemote work$184k - $287.5k
...NVIDIA's accelerated computing platform has revolutionized HPC and AI, and we have built... .... This role will be part of an engineering team developing, scaling, and optimizing... ...characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, NY, New YorkType: Full time...PlatformSeniorFull timePart time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Platform and EngOps Engineer - Cluster Operations (Santa Clara). Be the first to apply!
- platform developer Santa Clara, CA
- platform engineer Santa Clara, CA
- senior lighting artist Santa Clara, CA
- senior hvac project manager Santa Clara, CA
- senior technical product manager Santa Clara, CA
- senior medical science liaison Santa Clara, CA
- senior accountant remote Santa Clara, CA
- senior app developer Santa Clara, CA
- senior marketing account manager Santa Clara, CA
- senior robotics software engineer Santa Clara, CA




