Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Operations Engineering

$152k - $241.5k

NVIDIA

As a member of the Hardware Infrastructure EDA Compute team, you will optimize, scale, and support workload scheduling systems that directly impact design velocity and infrastructure efficiency. Beyond day-to-day operations, the role drives improvements in observability, service reliability, and automation, ensuring the EDA compute environment remains resilient, measurable, and aligned with long-term engineering demands.in a large-scale, multi-site environment supporting EDA and other compute-intensive workloadsAnalyze scheduler and infrastructure performance data to identify systemic bottlenecks and drive measurable improvements in utilization, throughput, and turnaround timeLead problem solving across scheduler, OS, and workload layers, ensuring timely resolution of service-impacting issuesIdentify recurring operational challenges and implement targeted automation or process improvements to reduce manual effort and prevent repeat incidentsHelp define and track reliable metrics and SLOs for service performance and reliability, partnering with customers to ensure expectations are realistic and measurableContribute to operational standards, documentation, and best practices to improve consistency across sitesPartner directly with customer teams to clarify requirements, translate technical tradeoffs, and drive issues to closureWhat we need to see:Bachelor’s degree in Computer Science or related field, or equivalent experienceMinimum 5+ years of experience operating and supporting large-scale Linux-based compute infrastructureStrong hands-on experience supporting and tuning job scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency in Linux systems administration (CentOS/RHEL)Strong problem solving skills and the ability to independently analyze complex system behavior under loadClear and effective communication skills, including the ability to articulate technical tradeoffs and reliability metrics to engineering stakeholdersWays to stand out from the crowd:Experience implementing reliability engineering practices within HPC scheduling environmentsDeep knowledge of job scheduling systems (LSF, Slurm, etc.) configuration tuning, scheduler internals, and advanced troubleshooting techniquesExperience building or enhancing observability systems, including metrics collection, monitoring pipelines, alerting strategies, and performance dashboardsBackground with container technologies such as Docker, Singularity, or Podman in HPC environmentsExperience influencing adoption of new infrastructure standards across multiple teams or sitesNVIDIA offers highly competitive salaries and a comprehensive benefits package. We are seeking creative and independent engineers with real passion for technology!#The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 24, 2026.As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.US, NC, DurhamType: Full time

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Operations Engineering in Santa Clara, CA vacancy
  • $124k - $195.5k

    As an HPC Operations Engineer at NVIDIA, you will play a pivotal role in ensuring the flawless operation of our high-performance computing (HPC)...  ...operational tasksSolid understanding of workload schedulers such as LSF, Slurm, or similar systemsStrong grasp of network computing... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $152k - $241.5k

     ...intelligence.We’re looking for a Senior SRE to join our Compute...  ...and implementation to operation and continuous...  ...integrate cleanly with HPC schedulers, storage, and...  ...clusters using Slurm, LSF or Kubernetes clusters,...  ...or Ruby.Mentored other engineers and influenced technical... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...ready to scale. You will work closely with engineering and AI teams, help shape how data moves...  ...more reliable, efficient, and easier to operate.Owning capacity planning, data lifecycle...  ...building or scaling storage for AI/ML or HPC workloads, including hybrid or multi cloud... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $189k - $301k

     ...the Role We are seeking a Senior Staff Engineer to build and optimize the EDA...  ...version control; set up, deploy, operate, and provide technical...  ...from design organizations LSF compute farm buildout and optimization...  ..., design infrastructure, or HPC for semiconductor design  ~... 
    Senior
    Contract work
    Work at office
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    19 days ago
  • $300 per month

     ...company built from the ground up, we own and operate each layer of the stack — from electrons...  ...At Crusoe Energy Systems, our Production Engineering (PE) team plays a mission-critical role...  ..., latency-sensitive workloads for AI and HPC use cases. This role directly supports... 
    Senior
    Temporary work

    Crusoe

    Sunnyvale, CA
    a month ago
  • $160k - $211k

     ...by Lattice OS, an AI-powered operating system that turns thousands of...  ....About the TeamThe Thermal Engineering team's work is essential to support...  ...the RoleWe are looking for a Senior Thermal Engineer in Costa...  ...of expeditionary AI/HPC data centers—air-transportable... 
    Senior
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Mountain View, CA
    5 hours ago
  • $255k - $340k

     ...design, and customer-scale AI infrastructure.We're looking for a Senior HPC Systems Architect with extensive experience designing,...  ...implementation teams.Provide technical leadership and mentoring to engineering teams, fostering best practices in HPC architecture.You8+... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $176k - $276k

    Production engineering is a field that involves crafting, building, and maintaining large-scale...  ...and ensure low-latency data access for HPC and AI/ML workloads.Storage Production...  ...a mindset focused on automating storage operations, improving data access efficiency, and optimizing... 
    Senior
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center...  ...NCCL, NVSHMEM, OpenSHMEM, UCX, UCC).Deep understanding of operating systems, computer and system architecture.Solid in... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Role Name: Senior Python Test and Automation Engineer Work Location: San Jose, California (Onsite)...  ...managing teams. Exposure to network operating systems, preferably SONiC very nice...  ...have. Familiarity with RDMA and HPC networks. Understanding of... 
    Senior

    Aita Consulting Services, Inc

    San Jose, CA
    3 days ago
  •  ...AI cloud, join us.What You’ll DoBuild and operate monitoring and alerting for cluster...  ...proactivelyRemotely deploy and configure large-scale HPC clusters for AI workloads using...  ...and feed clear requirements back to other engineering teams on simplification, stability, and operational... 
    Senior
    Work at office
    Local area
    Remote work
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $184k - $287.5k

    NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks.... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...next-gen distributed storage services for HPC workloads, optimizing both performance...  ...infrastructure environments, to automate operational monitoring and alerting, and to enable...  ...degree in Computer Science, Electrical Engineering or related field or equivalent experience... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $108k - $172.5k

     ...lasting impact on the world.We are seeking a highly motivated Senior HPC Support Engineer focussing on InfiniBand and NVLink technology, passionate...  ...for sophisticated installations, maintenance, or operations for a broad scope of groundbreaking networking products.... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $184k - $287.5k

     ...imagination and intelligence. Make the choice, join our diverse team today!We are looking for an outstanding hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of large-scale GPU compute clusters. Be a key player to enable the most... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ..., and tools that enable researchers and engineers to develop the next generation of AI/ML...  ...transformation. We are looking for a strong AI & HPC Observability Engineer to build and...  ...-grade coding, and a passion for operational excellence.What You Will Be Doing:Design... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $255k - $340k

     ...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and...  ..., at a scale few teams in the industry operate at, this is that team.What You’ll DoServe...  ...lead for integrating OEM and white-label HPC AI/ML, general purpose compute, storage,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    5 hours ago
  • $184k - $287.5k

    NVIDIA Math Libraries team is looking for a senior engineer to join our development efforts in the area of kernel generation for AI and HPC, specifically targeting matrix operations, JITing and fusions. Around the world, leading commercial and academic organizations are... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $176k - $276k

    NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building...  ...of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $179k - $218k

     ...infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens —...  ...Silicon Reality" must be bridged. We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive... 
    Senior
    Temporary work

    Crusoe

    Sunnyvale, CA
    a month ago
  • $100k - $138k

     ..., Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing...  ...community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.Job Summary:Supermicro... 
    Senior
    Work at office
    Worldwide

    Supermicro

    San Jose, CA
    1 day ago
  • $144.8k - $261.45k

    The OpportunityDetection & Automation Engineering operates Adobe's detection and response backbone, closing the gap between detection and containment.As a Senior Security Automation Engineer, you'll design and build production-grade automation, orchestration, and agentic... 
    Senior
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    3 days ago
  • $131.01k - $196.3k

     ...enabling timely delivery of current and next-generation products.Engineers on this team work close to real hardware, firmware, lab...  ...patterns, cable types, optical modules, switches, and customer-like operating conditions.Debug link-up, link-flap, signal integrity,... 
    Senior
    Permanent employment
    Full time
    Internship
    Work from home

    Marvell

    Santa Clara, CA
    3 days ago
  • $140k - $224.25k

     ...seeking a qualified Sr. Software Automation and Tools Development Engineer to join our GPU SWQA team. The successful candidate will...  ...cases, as well as an in-depth understanding of Windows and Linux operating systems. Comprehensive knowledge of system architecture is... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $200k - $322k

    NVIDIA is looking for a skilled and motivated Senior Infrastructure Engineer to join our dynamic team. You will contribute to innovative solutions and optimize our operations using brand new technology to agentize workflows to building agents for different domains of infrastructure... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $196k - $310.5k

     ...inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.We are seeking an Applied AI Engineer to lead end-to-end solution development — spanning data generation, model training, orchestration, and agentic automation — for... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224.9k - $309.3k

     ...support our extraordinary team who build great products and contribute to our growth, we’re looking to add a Sr Director, Operations Engineering located in San Jose, CA. Reporting to the VP, Global Operations the Sr Director, Operations Engineeringis responsible for leading... 
    Senior
    Full time
    Flexible hours

    Flextronics

    San Jose, CA
    4 days ago
  • $139k - $204k

     ...CRWV) in March 2025. Learn more at  About the Role Production Engineering ensures CoreWeave's cloud delivers world-class reliability, performance, and operational excellence. We are hiring a  Senior Production Engineer to take direct, hands-on ownership of critical... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    22 days ago
  • $141.3k - $360.7k

     ...easier for people to stream. We do this with state-of-the-art technology and engineering, keeping the customer at the center of everything we do.We're looking for a hands-on, systems-oriented Senior Software Engineer in Test (Sr. SDET) to join our Browse and Discovery Team... 
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $152k - $349k

    HPC/AI Principal Product ManagerThis role has been designed as 'Hybrid' with a requirement...  ...and drives the end to end strategy and operational product roadmap for one or more complex...  ...across the product lifecycle. (ie. Engineering: product development, Supply Chain: SKUs... 
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start
    2 days per week

    Hewlett Packard Enterprise

    San Jose, CA
    17 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Operations Engineering. Be the first to apply!