Senior HPC and LSF Operations Engineer
$152k - $241.5kNVIDIA
As a member of the Hardware Infrastructure EDA Compute team, you will optimize, scale, and support workload scheduling systems that directly impact design velocity and infrastructure efficiency. Success in this role requires both operational precision along with developing and supporting forward-looking resource management solutions that address evolving compute demands. Beyond day-to-day operations, the role drives improvements in observability, service reliability, and automation, ensuring the EDA compute environment remains resilient, measurable, and aligned with long-term engineering demands.What you'll be doing:Manage, scale, and optimize job scheduling systems (LSF, Slurm, etc.) in a large-scale, multi-site environment supporting EDA and other compute-intensive workloadsAnalyze scheduler and infrastructure performance data to identify systemic bottlenecks and drive measurable improvements in utilization, throughput, and turnaround timeLead problem solving across scheduler, OS, and workload layers, ensuring timely resolution of service-impacting issuesIdentify recurring operational challenges and implement targeted automation or process improvements to reduce manual effort and prevent repeat incidentsHelp define and track reliable metrics and SLOs for service performance and reliability, partnering with customers to ensure expectations are realistic and measurableContribute to operational standards, documentation, and best practices to improve consistency across sitesPartner directly with customer teams to clarify requirements, translate technical tradeoffs, and drive issues to closureWhat we need to see:Bachelor’s degree in Computer Science or related field, or equivalent experienceMinimum 5+ years of experience operating and supporting large-scale Linux-based compute infrastructureStrong hands-on experience supporting and tuning job scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency in Linux systems administration (CentOS/RHEL)Strong problem solving skills and the ability to independently analyze complex system behavior under loadClear and effective communication skills, including the ability to articulate technical tradeoffs and reliability metrics to engineering stakeholdersWays to stand out from the crowd:Experience implementing reliability engineering practices within HPC scheduling environmentsDeep knowledge of job scheduling systems (LSF, Slurm, etc.) configuration tuning, scheduler internals, and advanced troubleshooting techniquesExperience building or enhancing observability systems, including metrics collection, monitoring pipelines, alerting strategies, and performance dashboardsBackground with container technologies such as Docker, Singularity, or Podman in HPC environmentsExperience influencing adoption of new infrastructure standards across multiple teams or sitesNVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most forward-thinking and hardworking people in the world on our team and our collaborative talent continues to drive NVIDIA's growth. We are seeking creative and independent engineers with real passion for technology!#LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 9, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, NC, DurhamType: Full time
$124k - $195.5k
As an HPC Operations Engineer at NVIDIA, you will play a pivotal role in ensuring the flawless operation of our high-performance computing (HPC)... ...operational tasksSolid understanding of workload schedulers such as LSF, Slurm, or similar systemsStrong grasp of network computing...SuggestedFull time$184k - $287.5k
...environment runs millions of cores across federated LSF cells, and every simulation, synthesis run... ...LSF platform, and we are looking for an engineer who knows LSF at the level of its... ...Engineering, or equivalent experience.8+ years in HPC or large-scale batch compute, with 5+...SeniorFull timeRemote work- AMD is seeking an Enterprise AI/HPC GPU architect to join the Datacenter System Architecture and Engineering team to develop world-class products around Instinct GPUs. In this role you will engage with customers and internal teams to define end-to-end architecture spanning...Senior
- CACI International seeks a SIGINT Operation Support Specialist to join our engineering and technical support team in the United States. The role supports military customers by resolving hardware and software issues and delivering training on CACI software such as APERTURE...Senior
- NextSilicon, a leader in HPC acceleration, is seeking an experienced HPC and AI application field support engineer to join our pre-sales engineering team. You will port, benchmark and optimize applications for CPUs, GPUs and FPGAs, and present value to customers across...Senior
$120k - $207k
NVIDIA is seeking a Senior HPC Support Engineer to provide customer support for cutting-edge networking solutions. The role involves both onsite and remote support, with duties including troubleshooting and resolving technical issues for customers. The ideal candidate will...SeniorRemote work$184k - $356.5k
A leading technology company in Austin, Texas, is seeking a Senior Software Architect to enhance communication performance in AI and HPC applications. Candidates should have at least 5 years of experience, a relevant Master's or Ph.D., and expertise in C/C++ programming...Senior- ...actively towards design and code reviews.Write and review designs and code in Swift/Obj-C with colleagues of different skills and seniority.Validate bug fixes and recommend product improvements to ensure seamless user experiences for iPhones, iPads and Apple Watch.Raise...Senior
$184k - $287.5k
...foundation for its EDA compute farm, and we need an automation engineer to own it end to end. You are joining at the point where this is... ...you'll be doing:Designing and owning the configuration schema for LSF cell deployment, so that a policy change is written once,...SeniorFull time$184k - $287.5k
...NVSHMEM, and UCX that are crucial for scaling Deep Learning and HPC. We're seeking a Senior Software Architect to help co‑design next‑gen data center... ..., NVSHMEM, OpenSHMEM, UCX, UCC). Deep understanding of operating systems, computer and system architecture. Solid...Senior$165k - $210k
...solve complex challenges, and stay at the forefront of data engineering and AI advancements. Remote first with casual, award-winning... ...background with most of your experience in infrastructure and operations (managing enterprise data platforms). Responsibilities Leading...SeniorCasual workRemote work- ...Description Job Description The Senior Director, Physical Design & Backend Engineering is accountable for end-to-end backend execution across HPC SoC and MCU programs, including implementation... ..., and area (PPA) targets. The role operates at the intersection of execution,...Senior
- ...involves delivering performance and feature enhancements for GPU products, leading a team, and designing software libraries for AI and HPC applications. Applicants should have over 10 years of professional experience in software development, a strong background in GPU...
- ...Senior Production Support Engineer Location: Austin, TX (Onsite) Exp. Level: 8+ yrs Role Purpose Owns AI-augmented incident triage,... ...uptime target during core PST business hours (8AM-5PM) Operate within the support portion of delivery (L1/L2/L3 tiers)...Senior
$120k - $207k
We are seeking a highly motivated Senior HPC Support Engineer - Ethernet / AI Infrastructure to be engaging with one of our prestige customers... ...solutions for sophisticated installations, maintenance, or operations for a broad scope of groundbreaking networking products....SeniorWork experience placementRemote work- ...Join our Austin lab team to operate and maintain sophisticated test fixtures for a world-renowned technology company. As a Senior Test Automation Engineer specializing in Hardware and Robotics , you'll deploy hardware test systems, execute daily testing protocols, and...Senior
- ...actively seeking passionate, collaborative, energetic, and forward-thinking individuals to join our team. We are seeking a Senior Software Engineer – Test Automation & Infrastructure to develop and scale automated test systems supporting the validation of complex RF...SeniorPermanent employmentFull timeContract workWork experience placementLocal area
$115k - $125k
...Role: Senior Production Support Engineer Location: Austin, TX (Onsite) Exp. Level: 8+ yrs Role Purpose Owns AI-augmented incident... ...target during core PST business hours (8AM-5PM) Operate within the support portion of delivery (L1/L2/L3 tiers)...SeniorContract work- ...maintain enterprise security monitoring systems, including SIEM, EDR/XDR, and SOAR, while driving detection engineering and automation initiatives. The role requires senior-level cybersecurity engineering expertise, collaboration with cross-functional teams, and a focus on...Senior
- ...Senior QA Automation EngineerAustin, Texas, United StatesSustainment is an AI-native software... ...(~10–20% of Your Time)Mentor two QA engineers and raise the team's automation... ...and QA's definition of done.Act as the QA operational lead and a trusted second to the Director...SeniorFull timeLocal areaWork visaShift work
- Expedia Group is seeking an experienced Salesforce Marketing Cloud specialist to own CRM production end-to-end, develop campaigns, and implement AI-assisted enhancements within a fast-paced environment. You will work on Journey Builder, Automation Studio, AMPscript/SSJS...Senior
- Renesas Electronics Corporation in Austin, TX, seeks a Sr. Director of Applications in the Performance Computing Power division to build and lead an organization that turns Renesas silicon into design wins among AI and compute customers. You own the technical relationship...Senior
- Esolvit Inc is seeking an experienced professional for a role focused on IBM Websphere Application Server Administration. The ideal candidate should have a strong background in Jython Scripting, UNIX, and Windows systems. Responsibilities include involvement in the full...Senior
$272k - $431.25k
NVIDIA in Austin, Texas is seeking a highly experienced CPU Architect to drive the development of CPU technology for innovative applications in AI, gaming, and autonomous vehicles. You will bridge architecture and physical design while analyzing power trade-offs and driving...Senior- Samsung Semiconductor in Austin, Texas is seeking a talented CPU Micro-architect to shape next-generation processors for AI and high-performance computing. The role involves optimizing microarchitecture and collaborating with teams to push boundaries in processor design...Senior
$120k - $145k
...Ranchers tag their cattle once and immediately start running their operation from their phone. We were founded by Callum Taylor (Harvard... ...in Austin, Texas. About the role As our device QA / Test Engineer, you'll be the backbone of reliability for Drover products....SeniorContract workImmediate startShift work- A technology firm is seeking an experienced Slack L3 Engineer to enhance collaboration across its organization. The ideal candidate will have over 5 years of expertise in enterprise Slack administration, focusing on security and compliance. This role involves managing...SeniorLocal area
- Siemens Healthineers is hiring a Senior Manufacturing Engineer in Austin, TX to lead major projects, drive manufacturing methods, and ensure product manufacturability. You will partner with design and manufacturing teams to optimize processes, reduce costs, and ensure regulatory...Senior
- Advanced Micro Devices (AMD) is seeking a CGC-Tools engineer to design and sustain advanced debug infrastructure for client, graphics, and custom product lines. You will enable end-to-end debug for on‑chip instrumentation, JTAG over USB‑C, and data pipelines across silicon...Senior
$53 - $63 per hour
...performs a variety of complicated tasks, a wide degree of creativity and latitude is expected. Job Description QA automation engineers design automated tests by creating scripts that run testing functions automatically. This includes determining the priority for...SeniorLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior HPC and LSF Operations Engineer. Be the first to apply!
- senior security operations engineer Austin, TX
- post production engineer Austin, TX
- network operations center engineer Austin, TX
- production network engineer Austin, TX
- remote operation drilling engineer Austin, TX
- production operations engineer Austin, TX
- data center operations engineer Austin, TX
- operations quality engineer Austin, TX
- application operations engineer Austin, TX
- senior production engineer Austin, TX




