Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior HPC and LSF Operations Engineer

$152k - $241.5k

NVIDIA

As a member of the Hardware Infrastructure EDA Compute team, you will optimize, scale, and support workload scheduling systems that directly impact design velocity and infrastructure efficiency. Success in this role requires both operational precision along with developing and supporting forward-looking resource management solutions that address evolving compute demands. Beyond day-to-day operations, the role drives improvements in observability, service reliability, and automation, ensuring the EDA compute environment remains resilient, measurable, and aligned with long-term engineering demands.What you'll be doing:Manage, scale, and optimize job scheduling systems (LSF, Slurm, etc.) in a large-scale, multi-site environment supporting EDA and other compute-intensive workloadsAnalyze scheduler and infrastructure performance data to identify systemic bottlenecks and drive measurable improvements in utilization, throughput, and turnaround timeLead problem solving across scheduler, OS, and workload layers, ensuring timely resolution of service-impacting issuesIdentify recurring operational challenges and implement targeted automation or process improvements to reduce manual effort and prevent repeat incidentsHelp define and track reliable metrics and SLOs for service performance and reliability, partnering with customers to ensure expectations are realistic and measurableContribute to operational standards, documentation, and best practices to improve consistency across sitesPartner directly with customer teams to clarify requirements, translate technical tradeoffs, and drive issues to closureWhat we need to see:Bachelor’s degree in Computer Science or related field, or equivalent experienceMinimum 5+ years of experience operating and supporting large-scale Linux-based compute infrastructureStrong hands-on experience supporting and tuning job scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency in Linux systems administration (CentOS/RHEL)Strong problem solving skills and the ability to independently analyze complex system behavior under loadClear and effective communication skills, including the ability to articulate technical tradeoffs and reliability metrics to engineering stakeholdersWays to stand out from the crowd:Experience implementing reliability engineering practices within HPC scheduling environmentsDeep knowledge of job scheduling systems (LSF, Slurm, etc.) configuration tuning, scheduler internals, and advanced troubleshooting techniquesExperience building or enhancing observability systems, including metrics collection, monitoring pipelines, alerting strategies, and performance dashboardsBackground with container technologies such as Docker, Singularity, or Podman in HPC environmentsExperience influencing adoption of new infrastructure standards across multiple teams or sitesNVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most forward-thinking and hardworking people in the world on our team and our collaborative talent continues to drive NVIDIA's growth. We are seeking creative and independent engineers with real passion for technology!#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 24, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, DurhamType: Full time

Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the Senior HPC and LSF Operations Engineer in Santa Clara, CA vacancy
  • $124k - $195.5k

    As an HPC Operations Engineer at NVIDIA, you will play a pivotal role in ensuring the flawless operation of our high-performance computing (HPC)...  ...operational tasksSolid understanding of workload schedulers such as LSF, Slurm, or similar systemsStrong grasp of network computing... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...environment runs millions of cores across federated LSF cells, and every simulation, synthesis run...  ...LSF platform, and we are looking for an engineer who knows LSF at the level of its...  ...Engineering, or equivalent experience.8+ years in HPC or large-scale batch compute, with 5+... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $255k - $340k

     ...home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and...  ...product introduction (NPI), and fleet-scale operations — all engineered for gigawatt-scale AI...  ...system integration validation for new HPC AI/ML, general purpose compute, storage,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $152k - $241.5k

     ...world.We are seeking a highly skilled and experienced HPC Cluster Engineer to design, deploy, and operate GPU Compute Clusters for EDA (Electronic Design...  ...HPC job schedulers and orchestrators, such as Slurm, LSF, PBS or K8s. Applied experience with AI/HPC workflows... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $134k - $167.5k

     ...we invite you to bring your talents to Zscaler and help shape the future of cybersecurity.RoleWe are looking for a Senior IT Security Operations Engineer, Corporate Infrastructure to join our team. This is a Remote within the United States (with a hybrid preference for... 
    Senior
    Full time
    Work at office
    Local area
    Remote work

    Zscaler

    San Jose, CA
    3 days ago
  • $136k - $218.5k

    NVIDIA is looking for a Senior CPU Tooling and Design Automation Engineer! Do you want to help drive the development of CPU technology for architectures used...  ...AI) / deep learning (DL), high-performance computing (HPC), cloud service providers (CSP), gaming, virtual reality... 
    Senior
    Full time
    Work experience placement
    Night shift

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $272k - $431.25k

     ...ready to scale. You will work closely with engineering and AI teams, help shape how data moves...  ...more reliable, efficient, and easier to operate.Owning capacity planning, data lifecycle...  ...building or scaling storage for AI/ML or HPC workloads, including hybrid or multi cloud... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • Everpure in Santa Clara, CA, is seeking a strategic product marketer to lead GTM for Scalable AI & HPC, anchored by the FlashBlade//EXA product line. Define core narrative and positioning that differentiates in a fast-growth market. You’ll partner with Product Management... 
    Senior

    Everpure

    Santa Clara, CA
    11 hours ago
  •  ...Architect to bridge design and deployment of large-scale AI and HPC GPU infrastructure in Santa Clara, CA. You will drive end-to-...  ...solution integration with strategic customers and guide business and engineering teams based on customer feedback. Responsibilities include... 
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $189k - $301k

     ...the Role We are seeking a Senior Staff Engineer to build and optimize the EDA...  ...version control; set up, deploy, operate, and provide technical...  ...from design organizations LSF compute farm buildout and optimization...  ..., design infrastructure, or HPC for semiconductor design ~... 
    Senior
    Contract work
    Work at office
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    3 days ago
  •  ...NVIDIA Corporation is seeking a Senior Software Architect to enhance communication libraries crucial for scaling Deep Learning and HPC applications. You will investigate performance bottlenecks and design innovative communication technologies. The ideal candidate holds... 
    Senior
    Remote job

    Jobleads-US

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...foundation for its EDA compute farm, and we need an automation engineer to own it end to end. You are joining at the point where this is...  ...you'll be doing:Designing and owning the configuration schema for LSF cell deployment, so that a policy change is written once,... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $176k - $276k

    Production engineering is a field that involves crafting, building, and maintaining large-scale...  ...and ensure low-latency data access for HPC and AI/ML workloads.Storage Production...  ...a mindset focused on automating storage operations, improving data access efficiency, and optimizing... 
    Senior
    Full time
    Flexible hours

    Nvidia

    Santa Clara, CA
    11 hours ago
  • Data Direct Networks is seeking a Senior Technical Product Manager to lead the strategy, roadmap...  ...platform. You will partner with engineering, marketing, sales and customer success to deliver scalable solutions for AI, HPC and enterprise workloads. The role requires... 
    Senior

    Data Direct Networks

    Santa Clara, CA
    11 hours ago
  • Schlumberger is seeking a highly skilled HPC Engineer with deep expertise in discrete optimization, operations research, and quantum computing technologies. The role focuses on solving large-scale problems in scheduling, logistics, and resource allocation in data-driven... 
    Senior

    Schlumberger

    Sunnyvale, CA
    1 day ago
  • DDN is seeking a Senior Platform Product Manager to lead the strategy, roadmap, and execution for the Infinia...  ...hardware and storage media evolve to support AI, HPC, and enterprise workloads at scale, working across engineering, marketing, sales, and customer success. This... 
    Senior

    DDN

    Santa Clara, CA
    11 hours ago
  • $356.5k

     ...NVIDIA Gruppe is seeking a Senior Software Architect in Santa Clara, California. This role involves co-designing next-generation data...  ...developing scalable communications software to enhance Deep Learning and HPC applications. Candidates should have extensive experience in... 
    Senior

    Jobleads-US

    Santa Clara, CA
    1 day ago
  • $255k - $340k

     ...design, and customer-scale AI infrastructure.We're looking for a Senior HPC Systems Architect with extensive experience designing,...  ...implementation teams.Provide technical leadership and mentoring to engineering teams, fostering best practices in HPC architecture.You8+... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks.... 
    Senior

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center...  ...NCCL, NVSHMEM, OpenSHMEM, UCX, UCC).Deep understanding of operating systems, computer and system architecture.Solid in... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    11 hours ago
  • NVIDIA is seeking a Senior Software Engineer in Westford, Massachusetts to improve their HPC infrastructure. The role includes designing scalable systems and supporting multi-cloud environments. The ideal candidate will have 10+ years of experience, strong software development... 
    Senior

    NVIDIA

    Santa Clara, CA
    11 hours ago
  •  ...to design and development of accelerated and distributed Python APIs for numerical computing. Python dominates AI, data science and HPC, with NumPy, SciPy, TensorFlow and PyTorch. Join our team to develop GPU-accelerated Python libraries, optimize performance, and enable... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    1 hour ago
  • $196k - $310.5k

     ...impact on the world.We are now looking for a Senior Platform Engineer, Design Automation. You will join a small team that builds and operates the software platform used by NVIDIA chip-...  ....Experience with Docker, Kubernetes, HPC schedulers, or large compute farms.Background... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks.... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $184k - $287.5k

     ...next-gen distributed storage services for HPC workloads, optimizing both performance...  ...infrastructure environments, to automate operational monitoring and alerting, and to enable...  ...degree in Computer Science, Electrical Engineering or related field or equivalent experience... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...imagination and intelligence. Make the choice, join our diverse team today!We are looking for an outstanding hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of large-scale GPU compute clusters. Be a key player to enable the most... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $108k - $172.5k

     ...lasting impact on the world.We are seeking a highly motivated Senior HPC Support Engineer focussing on InfiniBand and NVLink technology, passionate...  ...for sophisticated installations, maintenance, or operations for a broad scope of groundbreaking networking products.... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...with 8-14 years of hands-on design and implementation expertise in HPC environments. You will work on enterprise-scale productivity...  ...solutions that scale across teams and geographies. Mentoring fellow engineers and collaborating with FAEs and customers are key aspects of the... 
    Senior

    Synopsys

    Sunnyvale, CA
    2 days ago
  • $152k - $241.5k

     ..., and tools that enable researchers and engineers to develop the next generation of AI/ML...  ...transformation. We are looking for a strong AI & HPC Observability Engineer to build and...  ...-grade coding, and a passion for operational excellence.What You Will Be Doing:Design... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $184k - $287.5k

    NVIDIA Math Libraries team is looking for a senior engineer to join our development efforts in the area of kernel generation for AI and HPC, specifically targeting matrix operations, JITing and fusions. Around the world, leading commercial and academic organizations are... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    11 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior HPC and LSF Operations Engineer. Be the first to apply!