Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior DevOps Engineer, Cloud Simulation Infrastructure

$184k - $287.5k

NVIDIA

We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline for SimReady assets! This role is critical to our product strategy, enabling us to transition from local, workstation-driven validation to high-scale, automated cloud validation on NVIDIA Cloud Functions (NVCF). You will be responsible for deploying a robust, multi-GPU pipeline that supports structural validation, AI-driven runtime behavioral testing, and automated asset remediation.What you'll be doing:Deployment: Deploy full Isaac Sim runtimes within GPU-aware NVCF containers. Manage container packaging, GPU initialization, and runtime utilities for physics, sensor, and rendering validation.Deploy Structural Validation: Deploy services to validate USD structure and compliance without runtime overhead.Deploy Runtime Validation: Architect scalable execution layers to conduct runtime behavior-based testing (e.g., drop/grasp tests). Deploy rule-based systems or AI based systems for automated pass/fail grading.Deploy Automated Remediation: Develop an AI-based pipeline that intercepts failures, triggers automated asset fixes, and re-validates results to ensure quality standards.Cloud Infrastructure Ownership: Scale execution from single-workstation validation to massive, multi-GPU cloud environments. Optimize for performance, addressing function-to-function networking, gRPC bottlenecks, and in-cluster proxy behavior.Artifact & Evidence Pipeline: Automate the generation of verification videos, thumbnails, feature-level reports, and validation metadata. Ensure all assets are traceable and linked to quality gates.Observability & CI/CD: Establish robust CI/CD, cluster verification, and monitoring pipelines. Implement logging, metrics, and tracing to ensure services are observable, debuggable, and production-ready.Operational Reliability: Implement atomic update semantics and safe failure handling to ensure validation processes never corrupt the primary asset library.What we need to see:BS or MS degree in Computer Science, Computer Engineering, or related field (or equivalent experience).8+ years of professional experience working on DevOps and/or cloud simulation.Extensive experience in production-grade DevOps, SRE, or Infrastructure Engineering, with a focus on GPU-backed cloud services.Proven expertise in container orchestration (Kubernetes/Docker) and CI/CD pipeline development.Experience with automated testing frameworks, preferably involving AI/ML inference, computer vision, or rule-based validation.Proficiency in Python and systems scripting for test orchestration and pipeline automation.Strong ability to design and maintain distributed job lifecycle services (submit/poll/fetch/cancel) and handle asynchronous failure states.Ability to diagnose and solve distributed network bottlenecks, including gRPC and function-to-function communication.Ways to stand out from the crowd:Direct experience deploying services on NVCF (NVIDIA Cloud Functions) or DGX Cloud.Deep familiarity with Isaac Sim, Omniverse, USD, or Sensor RTX workflows.Background in robotics simulation, physical AI, or large-scale content creation pipelines.Experience building "self-healing" or automated remediation workflows.Experience with cluster verification frameworks, stress testing, and deployment validation at scale.We value different paths to technical excellence and welcome candidates who bring strong judgment, curiosity, and a collaborative approach. Come build the future of autonomous vehicle simulation with us!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 21, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior DevOps Engineer, Cloud Simulation Infrastructure in Santa Clara, CA vacancy
  • $152k - $241.5k

     ...Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's...  ...hardware-in-the-loop (HIL) infrastructure to support the robust...  ..., Docker, Kubernetes) in cloud-native or hybrid environments...  ...of autonomous vehicle simulation with us!Your base salary... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $140k - $210k

     ...make a difference at Fiserv.Job TitleSr. DevOps / Cloud Engineer, Billing InfrastructureAbout CloverClover...  ...grow with confidence.What does a successful Senior DevOps / Cloud Engineer do at Clover?The Billing platform infrastructure supports Clover SaaS and App Market... 
    Senior
    Full time
    Worldwide

    Fiserv

    Sunnyvale, CA
    14 hours ago
  • $184k - $287.5k

    We are looking for a senior systems software engineer to improve the operation and user experience of distributed system infrastructure using AI. We are passionate about the opportunity to...  ...Docker and Kubernetes.Experience with DevOps tools such as Ansible, Terraform, and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $176k - $276k

    Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform...  ...environments.We are looking for a hands-on senior engineer to own the lifecycle and automation of the... 
    Senior
    Full time
    Remote work
    Weekend work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $159k - $231k

     ...with EDA vendors to enhance simulation capabilities and methodologies...  ...’s degree in Electrical Engineering, Computer Engineering, Physics...  ...Integrity Engineer within Platforms Infrastructure Engineering, you will play a...  ...include Googlers, Google Cloud customers, and billions of... 
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $210.6k - $305.1k

     ...TeamCisco is building next-generation infrastructure software for enterprise AI factories, neocloud providers, sovereign cloud environments, and modern data center fabrics...  ..., and scale.We are seeking a hands-on Senior Software Engineering Manager to lead a team working at the... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    CISCO Systems

    San Jose, CA
    2 days ago
  • $184k - $287.5k

     ...advanced large language model workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking, analysis, and...  ..., validation, and debugging of large-scale AI clusters, infrastructure, and end-to-end workloads, setting the standard for how... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $174k - $252k

     ...connectivity solution for hybrid/multi cloud, involving data plane and control plane...  ...years of experience developing large-scale infrastructure, distributed systems or networking, or...  ...Balancing, etc.).Google's software engineers develop the next-generation technologies... 
    Senior

    Google

    Sunnyvale, CA
    3 days ago
  • $163k - $237k

     ...qualifications:Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science...  ...for all usage scenarios.The AI and Infrastructure team is redefining what’s possible. We...  ...Our customers include Googlers, Google Cloud customers, and billions of Google users... 
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    13 hours ago
  • $174k - $253k

     ...years of experience developing large-scale infrastructure, distributed systems or networks, or...  ...technologies. Google's software engineers develop the next-generation technologies...  ...and fastest experience possible. Google Cloud accelerates every organization’s ability... 
    Senior

    Google

    Sunnyvale, CA
    2 days ago
  • $174k - $253k

     ...years of experience developing large-scale infrastructure, distributed systems or networks, or...  ...technologies. Google's software engineers develop the next-generation technologies...  ...Our customers include Googlers, Google Cloud customers, and billions of Google users... 
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    3 days ago
  • $150k - $218k

     ...report generation. Measure, simulate, analyze, and improve the thermal...  ...the test fixture for better engineering efficiency and data accuracy...  ...designs powerful computing infrastructures with custom-built machines....  ...include Googlers, Google Cloud customers, and billions of Google... 
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    4 days ago
  • $159k - $231k

     ...Robotics initiatives and roadmap, engineering execution with business...  ...service level availability simulation (SLAM), and proprioception pipelines...  ...to join our Physical infrastructure Robotics team. This team is...  ...customers include Googlers, Google Cloud customers, and billions of... 
    Senior
    Contract work
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago
  • $236k - $330k

     ...Robotics initiatives and roadmap, engineering execution with business...  ...liquid-cooled AI hardware infrastructure.Oversee the end-to-end deployment...  ...with world models and simulation.Understanding of recent advancements...  ...include Googlers, Google Cloud customers, and billions of... 
    Senior
    Contract work
    Remote work
    Worldwide
    Flexible hours

    Google

    Sunnyvale, CA
    4 days ago
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud Lepton Team means contributing to the...  ...well as developing scalable AI infrastructure services globally. We are...  ...an AI infrastructure software engineer to join our team. You'll be instrumental...  ...AI in production.As a senior DGX Cloud AI Infrastructure... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    14 hours ago
  • $144k - $198k

     ...testing pipelines for the GNC, Vehicle Simulation, Vehicle Management and Integrated...  ...existing, High-Performance Computing (HPC) infrastructure to support large-scale Monte Carlo...  ...need:BS/Advanced Degree in Aerospace Engineering, Computer Science, Electrical/Computer... 
    Senior
    Local area

    Archer Aviation

    San Jose, CA
    3 days ago
  • An innovative AI solutions company is seeking a Senior DevOps Engineer to architect and maintain the core infrastructure supporting cutting-edge AI applications. The role involves designing scalable environments, collaborating with teams for seamless deployments, and championing... 
    Senior
    Full time
    Remote work
    Flexible hours

    New Code Inc

    Palo Alto, CA
    3 days ago
  •  ...Santa Clara is seeking a Sr Site Reliability Engineer. The candidate will be responsible for maintaining highly reliable cloud infrastructure and will lead cross-functional...  ...at least 5 years of experience in SRE or DevOps and a solid background in cloud services,... 
    Senior

    Palo Alto Networks

    Santa Clara, CA
    3 days ago
  • $174k - $253k

     ...years of experience developing large-scale infrastructure, distributed systems or networks, or...  ...technologies. Google's software engineers develop the next-generation technologies...  ...Our customers include Googlers, Googler Cloud customers, and billions of Google users... 
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $280k - $380k

     ...disciplines.About the TeamOur DevOps/SRE team runs an active-active, multi-cloud platform on AWS and GCP...  ...and automation, engineering systems that perform under...  ..., and turning complex infrastructure into reliable, well‑...  ...Reliability Engineering) Senior Software Engineer to... 
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    4 days ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers...  ...s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $178k - $321k

     ...runtime, and the harness: a resilient cloud platform, the agentic runtime and evaluation...  ...setting, and the governed data and AI infrastructure everything else depends on. We hire on...  ...implement it. This is a two-person engineering team: you deploy, debug, and hotfix your... 
    Senior

    OKX

    San Jose, CA
    1 day ago
  • $159k - $230k

     ...reproduction.Design network, power and cooling infrastructure for end-to-end system testing.Drive...  ...of hardware/software designers, test engineers, on project planning within hardware...  ...manage the needs of multiple lab users.As a Senior Hardware Performance Test Engineer, you... 
    Senior

    Google

    Sunnyvale, CA
    2 days ago
  • $174k - $253k

     ...development code. Review code developed by other engineers and provide feedback to ensure best...  ...enhance software solutions.The AI and Infrastructure team is redefining what’s possible. We...  ...Our customers include Googlers, Google Cloud customers, and billions of Google users... 
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $174k - $299k

     ...tech, and hyper-connected world. Job Overview: As a Senior Staff Backend Engineer - Cloud Infrastructure, this role will focus on Cloud Platform Infrastructure...  ...Role Responsibilities: Strong experience on Cloud DevOps tools and coding (CI/CD, IaC, Terraform, Python, Cloud... 
    Senior
    Temporary work
    Work experience placement
    Flexible hours

    Coupang

    Mountain View, CA
    2 days ago
  •  ...wellbeing underline all aspects of Publicis Global Delivery.#WeArePGDOverviewWe are looking for a Senior DevOps Engineer with 5-7 years of experience in cloud infrastructure, automation, and CI/CD processes. Strong AWS, Terraform, and Kubernetes experience is required.... 
    Senior

    Publicis Media

    San Jose, CA
    2 days ago
  • $184k - $287.5k

    Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This...  ...an AI infrastructure software engineer to join our team. You'll be...  ...availability of AI systems.As a senior DGX Cloud AI Infrastructure software... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Looking for a Senior DevOps Engineer having good Handson with Linux Infrastructure experience, automating, and supporting enterprise-scale Linux environments across cloud and on-premises infrastructure. Candidate should have proven expertise in Kubernetes, Docker... 
    Senior

    Prophecy Technologies

    Sunnyvale, CA
    5 days ago
  • $300 per month

     ...Staff Software Infrastructure Engineer Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically...  ...across energy, manufacturing, data center construction, and cloud services. If you want to do the most meaningful work of your... 
    Senior
    Temporary work

    crusoe

    Sunnyvale, CA
    2 days ago
  • $256k - $414k

     ...NOW is the global leader in cloud gaming, dedicated to making...  ...scale.We are looking for a Senior Manager to lead the design,...  ...networking for GPU-based cloud infrastructure. This role is critical to enabling...  ...Science or a related engineering field (or equivalent experience... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior DevOps Engineer, Cloud Simulation Infrastructure. Be the first to apply!