Senior HPC and LSF Operations Engineer
$152k - $241.5kNVIDIA
As a member of the Hardware Infrastructure EDA Compute team, you will optimize, scale, and support workload scheduling systems that directly impact design velocity and infrastructure efficiency. Success in this role requires both operational precision along with developing and supporting forward-looking resource management solutions that address evolving compute demands. Beyond day-to-day operations, the role drives improvements in observability, service reliability, and automation, ensuring the EDA compute environment remains resilient, measurable, and aligned with long-term engineering demands.What you'll be doing:Manage, scale, and optimize job scheduling systems (LSF, Slurm, etc.) in a large-scale, multi-site environment supporting EDA and other compute-intensive workloadsAnalyze scheduler and infrastructure performance data to identify systemic bottlenecks and drive measurable improvements in utilization, throughput, and turnaround timeLead problem solving across scheduler, OS, and workload layers, ensuring timely resolution of service-impacting issuesIdentify recurring operational challenges and implement targeted automation or process improvements to reduce manual effort and prevent repeat incidentsHelp define and track reliable metrics and SLOs for service performance and reliability, partnering with customers to ensure expectations are realistic and measurableContribute to operational standards, documentation, and best practices to improve consistency across sitesPartner directly with customer teams to clarify requirements, translate technical tradeoffs, and drive issues to closureWhat we need to see:Bachelor’s degree in Computer Science or related field, or equivalent experienceMinimum 5+ years of experience operating and supporting large-scale Linux-based compute infrastructureStrong hands-on experience supporting and tuning job scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency in Linux systems administration (CentOS/RHEL)Strong problem solving skills and the ability to independently analyze complex system behavior under loadClear and effective communication skills, including the ability to articulate technical tradeoffs and reliability metrics to engineering stakeholdersWays to stand out from the crowd:Experience implementing reliability engineering practices within HPC scheduling environmentsDeep knowledge of job scheduling systems (LSF, Slurm, etc.) configuration tuning, scheduler internals, and advanced troubleshooting techniquesExperience building or enhancing observability systems, including metrics collection, monitoring pipelines, alerting strategies, and performance dashboardsBackground with container technologies such as Docker, Singularity, or Podman in HPC environmentsExperience influencing adoption of new infrastructure standards across multiple teams or sitesNVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most forward-thinking and hardworking people in the world on our team and our collaborative talent continues to drive NVIDIA's growth. We are seeking creative and independent engineers with real passion for technology!#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 9, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, DurhamType: Full time
$152k - $241.5k
...Success in this role requires both operational precision along with... ..., and aligned with long-term engineering demands.What you'll be doing:... ...optimize job scheduling systems (LSF, Slurm, etc.) in a large-... ...systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency...SeniorFull time$124k - $195.5k
As an HPC Operations Engineer at NVIDIA, you will play a pivotal role in ensuring the flawless operation of our high-performance computing (HPC)... ...operational tasksSolid understanding of workload schedulers such as LSF, Slurm, or similar systemsStrong grasp of network computing...SuggestedFull time$296k - $395k
...home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and... ...product introduction (NPI), and fleet-scale operations — all engineered for gigawatt-scale AI... ...system integration validation for new HPC AI/ML, general purpose compute, storage,...SeniorWork at officeLocal areaWork from homeFlexible hours$184k - $287.5k
...environment runs millions of cores across federated LSF cells, and every simulation, synthesis run... ...LSF platform, and we are looking for an engineer who knows LSF at the level of its... ...Engineering, or equivalent experience.8+ years in HPC or large-scale batch compute, with 5+...SeniorFull timeRemote work$300 per month
...company built from the ground up, we own and operate each layer of the stack — from electrons... ...Crusoe Energy Systems, our Production Engineering (PE) team plays a mission-critical role in... ..., latency-sensitive workloads for AI and HPC use cases. This role directly supports...SeniorTemporary work$272k - $431.25k
...ready to scale. You will work closely with engineering and AI teams, help shape how data moves... ...more reliable, efficient, and easier to operate.Owning capacity planning, data lifecycle... ...building or scaling storage for AI/ML or HPC workloads, including hybrid or multi cloud...SeniorFull time$205k - $260k
...place-and-route, and timing-closure workflows for engineering throughput and cloud cost efficiency. Own cloud HPC compute strategy, including auto-scaling, job... ...development cycle. ~ Experience building and operating data lakes and delivering role-tailored dashboards...SeniorFull time$300 per month
...company built from the ground up, we own and operate each layer of the stack — from electrons... ..., and our Compute-focused Production Engineers are the backbone of that mission. This role... ..., security, and scale for modern AI and HPC workloads. What You'll Be Working On...SeniorTemporary work- ...Role name: Senior Test and Automation Engineer Work Location San Jose, California (Onsite) Type of... ...managing teams. Exposure to network operating systems, preferably SONiC very nice... .... Familiarity with RDMA and HPC networks. Understanding of RoCE...SeniorContract work
$152k - $241.5k
...intelligence.We’re looking for a Senior SRE to join our Compute... ...and implementation to operation and continuous... ...integrate cleanly with HPC schedulers, storage, and... ...clusters using Slurm, LSF or Kubernetes clusters,... ...or Ruby.Mentored other engineers and influenced technical...SeniorFull time$176k - $276k
Production engineering is a field that involves crafting, building, and maintaining large-scale... ...and ensure low-latency data access for HPC and AI/ML workloads.Storage Production... ...a mindset focused on automating storage operations, improving data access efficiency, and optimizing...SeniorFull timeFlexible hours$184k - $287.5k
...foundation for its EDA compute farm, and we need an automation engineer to own it end to end. You are joining at the point where this is... ...you'll be doing:Designing and owning the configuration schema for LSF cell deployment, so that a policy change is written once,...SeniorFull time$189k - $301k
...the Role We are seeking a Senior Staff Engineer to build and optimize the EDA... ...version control; set up, deploy, operate, and provide technical... ...from design organizations LSF compute farm buildout and optimization... ..., design infrastructure, or HPC for semiconductor design ~...SeniorContract workWork at officeFlexible hours- ...NVIDIA in Santa Clara, CA is seeking a senior software developer to advance testing and automation for DriveOS. You will design test... .... Applicants should have a bachelor’s or master’s in engineering or CS, 5+ years in software development (Python or C++), experience...Senior
$98.9k - $228.7k
...you will be responsible for the reliability, scalability, and operational excellence of Zoom's government and military-facing products.... ...for how the team operates. This is a hands-on, on-call engineering role: you will own the systems you build, respond to production...SeniorPermanent employmentFull timeWork at officeRemote work$300 per month
...company built from the ground up, we own and operate each layer of the stack — from electrons... ...cloud platform — and Production Engineering sits at the heart of that mission. As a Production... ...that supports demanding AI and HPC workloads. You’ll partner closely with...SeniorTemporary work$255k - $340k
...design, and customer-scale AI infrastructure.We're looking for a Senior HPC Systems Architect with extensive experience designing,... ...implementation teams.Provide technical leadership and mentoring to engineering teams, fostering best practices in HPC architecture.You8+...SeniorWork at officeLocal areaWork from homeFlexible hours$100k - $138k
..., Cloud Computing, Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We are the #5 fastest growing... ...community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us.Job Summary:Supermicro...SeniorWork at officeWorldwide$184.7k - $324.8k
...Sr. ML Production Model Automation Engineer, Siri SpeechJoin the team redefining what a deeply personal and integrated assistant can be... ...staging promotion, production rollout and deprecation.Design and operate agent-based automation pipelines for ML models where agents own...SeniorRelocation$184k - $287.5k
...NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center... ...NCCL, NVSHMEM, OpenSHMEM, UCX, UCC).Deep understanding of operating systems, computer and system architecture.Solid in...SeniorFull timeRemote work$100k - $138k
...IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide... ..., passionate, and committed engineers, technologists, and business leaders... ...individual will collaborate with senior management of global production operations to establish quality-focused teams...SeniorContract workWork at officeWorldwideShift work$184k - $287.5k
...next-gen distributed storage services for HPC workloads, optimizing both performance... ...infrastructure environments, to automate operational monitoring and alerting, and to enable... ...degree in Computer Science, Electrical Engineering or related field or equivalent experience...SeniorFull time$184k - $287.5k
NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks....SeniorFull time$179k - $218k
...infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens —... ...Silicon Reality" must be bridged. We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive...SeniorTemporary work$104k - $152k
...Enterprise IT, Hadoop/ Big Data, Hyperscale, HPC and IoT/Embedded customers worldwide. We... ...talented, passionate, and committed engineers, technologists, and business leaders to... ...require skill sets to break down complex operations with a process oriented mindset, analyze...SeniorWork experience placementWorldwide$108k - $172.5k
...lasting impact on the world.We are seeking a highly motivated Senior HPC Support Engineer focussing on InfiniBand and NVLink technology, passionate... ...for sophisticated installations, maintenance, or operations for a broad scope of groundbreaking networking products....SeniorFull timeWork experience placement$184k - $287.5k
...imagination and intelligence. Make the choice, join our diverse team today!We are looking for an outstanding hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of large-scale GPU compute clusters. Be a key player to enable the most...SeniorFull timeRemote work$155k - $195k
...Senior Automation EngineerJoin Therma's dynamic team as Senior Automation Engineer and play a key role in leading the technical efforts in the design and implementation of... ...Functional Design Specifications and Sequence of Operations (SOO).Program and script PLC/DCS systems....Senior$131.01k - $196.3k
...enabling timely delivery of current and next-generation products.Engineers on this team work close to real hardware, firmware, lab... ...patterns, cable types, optical modules, switches, and customer-like operating conditions.Debug link-up, link-flap, signal integrity,...SeniorPermanent employmentFull timeInternshipWork from home$144.8k - $261.45k
The OpportunityDetection & Automation Engineering operates Adobe's detection and response backbone, closing the gap between detection and containment.As a Senior Security Automation Engineer, you'll design and build production-grade automation, orchestration, and agentic...SeniorFull timeTemporary workLocal areaWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior HPC and LSF Operations Engineer. Be the first to apply!
- senior security operations engineer Santa Clara, CA
- post production engineer Santa Clara, CA
- network operations center engineer Santa Clara, CA
- production operations engineer Santa Clara, CA
- data center operations engineer Santa Clara, CA
- application operations engineer Santa Clara, CA
- senior production engineer Santa Clara, CA
- operations engineer Santa Clara, CA
- security operations center engineer Santa Clara, CA
- senior technical service engineer Santa Clara, CA



