Senior HPC and LSF Operations Engineer
$152k - $241.5kNVIDIA
As a member of the Hardware Infrastructure EDA Compute team, you will optimize, scale, and support workload scheduling systems that directly impact design velocity and infrastructure efficiency. Success in this role requires both operational precision along with developing and supporting forward-looking resource management solutions that address evolving compute demands. Beyond day-to-day operations, the role drives improvements in observability, service reliability, and automation, ensuring the EDA compute environment remains resilient, measurable, and aligned with long-term engineering demands.What you'll be doing:Manage, scale, and optimize job scheduling systems (LSF, Slurm, etc.) in a large-scale, multi-site environment supporting EDA and other compute-intensive workloadsAnalyze scheduler and infrastructure performance data to identify systemic bottlenecks and drive measurable improvements in utilization, throughput, and turnaround timeLead problem solving across scheduler, OS, and workload layers, ensuring timely resolution of service-impacting issuesIdentify recurring operational challenges and implement targeted automation or process improvements to reduce manual effort and prevent repeat incidentsHelp define and track reliable metrics and SLOs for service performance and reliability, partnering with customers to ensure expectations are realistic and measurableContribute to operational standards, documentation, and best practices to improve consistency across sitesPartner directly with customer teams to clarify requirements, translate technical tradeoffs, and drive issues to closureWhat we need to see:Bachelor’s degree in Computer Science or related field, or equivalent experienceMinimum 5+ years of experience operating and supporting large-scale Linux-based compute infrastructureStrong hands-on experience supporting and tuning job scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency in Linux systems administration (CentOS/RHEL)Strong problem solving skills and the ability to independently analyze complex system behavior under loadClear and effective communication skills, including the ability to articulate technical tradeoffs and reliability metrics to engineering stakeholdersWays to stand out from the crowd:Experience implementing reliability engineering practices within HPC scheduling environmentsDeep knowledge of job scheduling systems (LSF, Slurm, etc.) configuration tuning, scheduler internals, and advanced troubleshooting techniquesExperience building or enhancing observability systems, including metrics collection, monitoring pipelines, alerting strategies, and performance dashboardsBackground with container technologies such as Docker, Singularity, or Podman in HPC environmentsExperience influencing adoption of new infrastructure standards across multiple teams or sitesNVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most forward-thinking and hardworking people in the world on our team and our collaborative talent continues to drive NVIDIA's growth. We are seeking creative and independent engineers with real passion for technology!#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 24, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, DurhamType: Full time
$124k - $195.5k
As an HPC Operations Engineer at NVIDIA, you will play a pivotal role in ensuring the flawless operation of our high-performance computing (HPC)... ...operational tasksSolid understanding of workload schedulers such as LSF, Slurm, or similar systemsStrong grasp of network computing...SuggestedFull time$152k - $241.5k
...world.We are seeking a highly skilled and experienced HPC Cluster Engineer to design, deploy, and operate GPU Compute Clusters for EDA (Electronic Design... ...HPC job schedulers and orchestrators, such as Slurm, LSF, PBS or K8s. Applied experience with AI/HPC workflows...SeniorFull time$152k - $241.5k
...intelligence.We’re looking for a Senior SRE to join our Compute... ...and implementation to operation and continuous... ...integrate cleanly with HPC schedulers, storage, and... ...clusters using Slurm, LSF or Kubernetes clusters,... ...or Ruby.Mentored other engineers and influenced technical...SeniorFull time$136k - $218.5k
NVIDIA is looking for a Senior CPU Tooling and Design Automation Engineer! Do you want to help drive the development of CPU technology for architectures used... ...AI) / deep learning (DL), high-performance computing (HPC), cloud service providers (CSP), gaming, virtual reality...SeniorFull timeWork experience placementNight shift$166k - $249k
General Information Job Title Digital Verification, Principal Engineer - 15149 HPC IP Job ID 15149 City Austin State/Province Texas Date Posted 05-Feb-2026 Job Category Engineering Job Subcategory ASIC Digital Design Hire Type...SuggestedRemote work$145k - $199.1k
Senior GenAI & High Performance Computing (HPC) Delivery EngineerDell Technologies has delivered HPC solutions for 25+ years, including support for Bright... ...GenAI & High Performance Computing (HPC) Delivery Engineer on our Service Delivery Team in Austin, Texas or Remote...SeniorRemote work- Job DescriptionThe Senior Director, Physical Design & Backend Engineering is accountable for end-to-end backend execution across HPC SoC and MCU programs, including implementation, signoff... ..., and area (PPA) targets. The role operates at the intersection of execution,...Senior
$184k - $287.5k
...NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co-design next-gen data center... ...NCCL, NVSHMEM, OpenSHMEM, UCX, UCC).Deep understanding of operating systems, computer and system architecture.Solid in...SeniorFull timeRemote work$106.8k - $194.8k
...diverse teams and take your career wherever you want it to go. Join EY and help to build a better working world. WAF Operations Solution Engineer PRACTICE DESCRIPTION: As a WAF Operations Solution Engineer, you will be responsible for implementing and...SeniorSummer holidayFlexible hours$184k - $287.5k
...next-gen distributed storage services for HPC workloads, optimizing both performance... ...infrastructure environments, to automate operational monitoring and alerting, and to enable... ...degree in Computer Science, Electrical Engineering or related field or equivalent experience...SeniorFull time$120k - $202.5k
Who We Are Looking For State Street's Cyber Data & Analytics (CyberDNA) team is seeking a Sr.Platform Operations Engineer to help shape the next generation of cybersecurity data, analytics, and AI-powered platforms. Partnering closely with Global Cyber Security, Infrastructure...SeniorFull timeTemporary workFlexible hours- ...supercomputers across industry, academia, and national labs.The AMD HPC & Sovereign AI applications team seeks a strong, experienced,... ..., bringing together developers, customers, and product engineers to deliver programs on-time. You delight in winning business by...Senior
$152k - $241.5k
...and see how you can make a lasting impact on the world.We are looking for a Senior Software Engineer to join our mission to continue improving our HPC infrastructure. Our team builds and operates sophisticated infrastructure to enable business critical services and AI...SeniorFull time- Senior Systems Development Engineer - HPCJoin us as a Senior Systems Development Engineer on our High Performance Computing (HPC) Solutions Engineering team in Austin, Texas to do the best work of... ...in the locations where we operate. Dell will not tolerate discrimination...SeniorWork experience placementStart working todayFlexible hours
$174k - $225k
...assumptions and reimagining how work gets done. Engineers define intent, author precise... ...member — right now at MyWellatDell.comThermal Senior Principal EngineerThe ISG Thermal Engineering... ...AI, Cloud, High Performance Computing (HPC), Edge computing devices and enterprise...Senior- Senior ServiceNow Developer (Modernization, Automation & AI)Be a part of a team that’s ensuring Dell Technologies' product integrity and customer satisfaction. Our IT Software Engineer team turns business requirements into technology solutions by designing, coding and...Senior
$106.7k - $133.4k
POSITION SUMMARY:The Senior Lab Automation Engineer is an experienced engineer with expertise in the development and implementation of complex laboratory... ...with external contractors.Experience in programming and operating liquid handlers. (Tecan, Hamilton, etc)Experience working...SeniorFor contractorsWork at officeLocal areaImmediate startWorldwide$107.1k - $160.7k
...the world’s leading integrated design practice. Our architects, engineers, interior designers, consultants, sustainability specialists,... ...Join us and design your place with Stantec.Your OpportunityThe Senior Automation Engineer for BAS/BMS/PLC systems, guides the technical...SeniorFull timeTemporary workPart timeCasual workLocal areaFlexible hours- Title: Senior Building Automation Systems (BAS) EngineerLocation: Austin, TXSalary: $13... ...in 1987, this organization is a global engineering firm specializing in building automation... ...SCADA systems to ensure reliable, 24/7 operations across complex environments. As we continue...Senior
$196k - $310.5k
...inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.We are seeking an Applied AI Engineer to lead end-to-end solution development — spanning data generation, model training, orchestration, and agentic automation — for...SeniorFull time$176k - $333.5k
...one else can solve. Make the choice to join us today. We are looking for a Senior Software Engineer to join our mission to continue improving our HPC infrastructure. Our team builds and operates sophisticated infrastructure to enable business critical services and AI applications...Senior- ...Job Description Job Description Senior analog/mixed-signal design engineer focusing on automation of high-performance analog-to-digital and digital-to-analog converters. The successful candidate in this role will be a member of the Automation Compiler Team, focused...Senior
- ...programming technologies to enable key mathematical operations on GPUDesign GPU computational software libraries for AI, HPC applicationsAid management in planning, and... ...development teams and other internal engineering teamsPREFERRED EXPERIENCE: 10+ years professional...
- ...AI and beyond. Together, we advance your career. THE ROLE: AMD is looking for a Enterprise AI/HPC GPU architect to join our Datacenter System Architecture and Engineering team to develop world-class products around Instinct GPUs. In this role you will be engaged with...
$114k - $148.2k
Operations is at the heart of Amazon’s business. We are known for our speed, accuracy, and... ...every day. The Reliability & Maintenance Engineering (RME) team are the business partners that... ...our journey! About the Role: As a Senior Automation Engineer, you will play a crucial...SeniorRemote workWorldwideFlexible hoursShift workNight shift- ...administrators and application teams—automating complex operational activities such as patching, release management, automated... ...at scale. We’re just getting started. We’re looking for a senior, high‑impact engineer with modern cloud development experience and a passion...SeniorFull timeLocal areaWork from homeRelocation package
- ...actively seeking passionate, collaborative, energetic, and forward-thinking individuals to join our team.We are seeking a Senior Software Engineer - Test Automation & Infrastructure to develop and scale automated test systems supporting the validation of complex RF and...SeniorPermanent employmentFull timeContract workWork experience placementLocal area
$184k - $287.5k
NVIDIA is seeking a Senior System Architect: Heterogeneous EDA Systems... ...resource waste. We need an engineer to develop and build an automated... ...Cause Analysis) pipelines for HPC or cloud-scale environments.... ...resource managers (Slurm, LSF, or Kubernetes) and how they manage...SeniorFull time$104k - $164k
...can’t wait to meet you.WHAT YOU'LL DOWe're looking for a Senior Lead Systems Engineer who's passionate about building the next generation of enterprise... ...teams to build reliable, scalable solutions that reduce operational overhead and improve the employee experience.This role...SeniorWork at officeLocal area- Job Description:About the Role:We are looking for a Senior SRE to join our Platform Engineering team where you’ll own the reliability, scalability, and operational excellence of our workflow orchestration platforms - primarily Apache Airflow and Broadcom Automic/UC4. This...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior HPC and LSF Operations Engineer. Be the first to apply!
- security operations center engineer Austin, TX
- remote operation drilling engineer Austin, TX
- data center operations engineer Austin, TX
- senior security operations engineer Austin, TX
- production operations engineer Austin, TX
- network operations center engineer Austin, TX
- operations quality engineer Austin, TX
- application operations engineer Austin, TX
- senior production engineer Austin, TX
- post production engineer Austin, TX


