Senior HPC Engineer - Fleet Engineering
Lambda Labs
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.If you'd like to build the world's best AI cloud, join us.What You’ll DoBuild and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and job-level signals — to detect and respond to issues proactivelyRemotely deploy and configure large-scale HPC clusters for AI workloads using automation wherever possibleAutomate cluster lifecycle: operating systems, firmware, drivers, and networking, managed as code (Ansible, Terraform) rather than by handCreate runbooks and automated remediations for common cluster failure modes, designed so Support and HPC Support can run them safelyTroubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power — working closely with on-site deployment teamsParticipate in on-call rotations and lead incident response for cluster-level problemsContribute to and maintain Standard Operating Procedures, and feed clear requirements back to other engineering teams on simplification, stability, and operational efficiencyYou7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar roleHave a strong understanding of modern AI infrastructure, from GPU architectures to hardware performance optimizationStrong understanding of Linux-based systems in a distributed environmentAre experienced configuring and troubleshooting InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, Ethernet/switching, GPU-direct, and NCCL environmentsSolid understanding of Python and Go, with experience working with SWE teams to improve internal tooling.Experience with monitoring and alerting tools (e.g., Prometheus, Grafana, Clickhouse)Proficiency in automation and configuration management tools (e.g., Ansible, Terraform)Have excellent problem-solving and troubleshooting skills and an innate attention to detailPassion for continuous improvement and innovationNice to HaveExperience with machine learning / deep learning frameworks (PyTorch, TensorFlow) and benchmarking tools (DeepSpeed, MLPerf)Knowledge of containerization and orchestration technologies (e.g., Docker, Kubernetes)Experience building and/or operating HPC resources.Depth in the NVIDIA hardware and firmware ecosystemExperience with data center power and thermal designBackground in chaos engineering or similar reliability testing methodologiesUnderstanding of compliance frameworks (SOC 2, ISO 27001, etc.)Salary Range InformationThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.About LambdaFounded in 2012, with 500+ employees, and growing fastOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent CoveWe have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOGOur values are publicly available: We offer generous cash & equity compensationHealth, dental, and vision coverage for you and your dependentsWellness and commuter stipends for select roles401k Plan with 2% company match (USA employees)Flexible paid time off plan that we all actually useEqual Opportunity EmployerLambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.Compensation Range: $227K - $356KLocationSan Francisco Office (Fremont St); Bellevue Office; Remote, USA; San Jose Office (First St)Employment TypeFull timeLocation TypeHybridDepartmentData Center BusinessCompensationSan Francisco / San JoseSan Francisco / San Jose $267K – $356KBellevueBellevue $240K – $320KRemote, USA$227K – $303K
$255k - $340k
...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and... ..., new product introduction (NPI), and fleet-scale maintenance — all engineered for gigawatt... ...lead for integrating OEM and white-label HPC AI/ML, general purpose compute, storage,...SeniorFleetWork at officeLocal areaWork from homeFlexible hours$152k - $241.5k
...intelligence.We’re looking for a Senior SRE to join our Compute Farm... ...they integrate cleanly with HPC schedulers, storage, and network... ...host lifecycle management, fleet reliability/auto-healing, E2E... ...Perl, or Ruby.Mentored other engineers and influenced technical direction...SeniorFleetFull time- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...recovery, resizing, and incident response using fleet management toolsParticipate in a well-... ...and authenticationWork closely with our HPC Ops and Datacenter Ops teams for low-...SeniorFleetWork at officeLocal areaWork from homeFlexible hours
$152k - $241.5k
...you to contribute to our team. We are looking for a Senior Software Engineer to join our DGX Cloud / Fleet Intelligence team and build agent-side systems that... .../log export pipelines and operating software in AI, HPC, cloud, or large-scale datacenter environments.Track...SeniorFleetFull timeLocal areaRemote work$267k - $356k
...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...capacity health of Lambda's production storage fleet across all data centers, operating behind... ...operating Linux systems in production or HPC environments, with hands-on storage...SeniorFleetWork experience placementWork at officeLocal areaWork from homeFlexible hours$184k - $287.5k
NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks....SeniorFull time$184k - $287.5k
...implement scalable, next-gen distributed storage services for HPC workloads, optimizing both performance and cost-effectiveness to... ...need to see:Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience.8+ years of experience...SeniorFull time$184k - $287.5k
NVIDIA Math Libraries team is looking for a senior engineer to join our development efforts in the area of kernel generation for AI and HPC, specifically targeting matrix operations, JITing and fusions. Around the world, leading commercial and academic organizations are...SeniorFull timeRemote work$152k - $241.5k
...environment remains resilient, measurable, and aligned with long-term engineering demands.What you'll be doing:Manage, scale, and optimize job... ...and tuning job scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency in Linux systems administration...SeniorFull time$224k - $356.5k
...-world production. We are looking for a Reinforcement Learning Engineer to join our mission in building intelligent, safe, and efficient... ...large-scale data flywheels, including mining scenarios from fleet telemetry logs, auto-labeling pipelines, and automated performance...SeniorFleetFull time$148.7k - $195k
...making the transition to electric easy for businesses, fleets and drivers. ChargePoint offers a once-in-a-lifetime... ...You Will Be DoingChargePoint is seeking a Senior Program Manager for Hardware Engineering to lead initiatives related to our EV charging hardware...SeniorFleetContract work$150k - $200k
Thanks for your interest in Oklo! We are searching for a senior maintenance program engineer to join our engineering team.Position DescriptionAt Oklo,... ...development, implementation, and continuous improvement of fleet maintenance standards across Oklo facilities. The...SeniorFleetRemote workFlexible hours$168k - $264.5k
...late in a program, it usually breaks here first.We are hiring a Senior Engineer to lead system validation, debug, and cross-functional... ...PVT stress, system-stress campaigns at scale, and multi-unit fleet testing.Lead debug of the hardest cross-stack issues — logic,...SeniorFleetFull timeFlexible hoursShift work$200k - $230k
...solely on making the transition to electric easy for businesses, fleets and drivers. ChargePoint offers a once-in-a-lifetime... ...looking for an experienced Power Electronics Controls and Firmware Engineer with more than seven years of hands-on expertise in developing...SeniorFleet$184k - $287.5k
...for AI research and production workloads. We are looking for Senior Software Engineers to help build the automation, tooling, and operational... ...infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation.Experience with SLOs, on-call, incident response...SeniorFleetFull time$152k - $241.5k
Join NVIDIA's Isaac Applications Engineering team and help build the platform for Physical AI robots — assembling the full robotics stack... ...managing GPU-backed CI infrastructure or self-hosted runner fleets.Your base salary will be determined based on your location, experience...SeniorFleetFull timeLive inNight shift$130k - $175k
...RoleTarana is scaling field deployments 5×, and we’re looking for a senior engineer to own the technical execution and continuous improvement of... ...cross-functional processes that allow HW QA to scale with the fleet.As a Senior Hardware QA & Diagnostics Engineer you will have...SeniorFleetWorldwideFlexible hours$148k - $235.75k
...make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw... ...into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer to operate the platform itself...SeniorFleetFull time$101k - $161k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...responsible for our global CloudVision service fleet, ensuring scalability, reliability, and... ...timeFunction: EngineeringExperience level: Mid-Senior LevelIndustry: Computer NetworkingSeniorFleet$143.3k
...the best customer experience in cloud computing.We aim to hire engineers who will thrive in a fast-paced, collaborative and open environment... ...responsible for solving operational challenges to our existing fleet with the goal of improving the current customer experience as...SeniorFleetLocal areaFlexible hours$183k - $247.6k
...You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers... ...servers to the data center. After launch you will oversee the fleet of servers you develop, monitoring their quality and how they are...SeniorFleetLocal areaFlexible hours- ...transition to electric easy for businesses, fleets and drivers. ChargePoint offers a once-in-... .... Reports To Director, Hardware Design Engineering What You Will Be Doing The Hardware Design Engineering Team is seeking a Senior Hardware Design Engineer - Interconnects to...SeniorFleet
$163.43k - $213.97k
...The Role: We are looking for a deeply hands-on Senior Distributed Systems Engineer to join the team building IonQ's Network and Security Platform... ...pipelines collecting telemetry from large device fleets, to time-series storage, real-time processing engines, and...SeniorFleetPermanent employmentContract workWork at office$173.9k - $235.2k
...workloads.We are looking for an experienced System development engineer to drive development for new EC2 machine learning platforms. In... ...recommending best practices for maintaining and improving code quality, fleet health, and security & reliability of our service.- Growing our...SeniorFleetInternshipLocal areaWork from homeWorldwideFlexible hours$183k - $247.6k
Do you like to apply deep hardware engineering expertise to deliver simple, sustainable, and scalable solutions for large networks? Would... ...global AWS network. Beyond product delivery we actively manage the fleet or routers in a network that grows by 70% annually. This means...SeniorFleetWork experience placementLocal areaFlexible hours$175k - $210k
...efficiencies. About the role We are looking for a talented Senior Robotics Software Engineer with expertise in robotics software development to join... ...the autonomous software powering the Bonsai Intelligence fleet across a multitude of applications and environments Write...SeniorFleetWork at office$154k - $200k
...industrial heat and power at scale. We are looking for a Senior / Staff Turbomachinery Systems Engineer to lead the definition, integration, and execution of... ...execution, commissioning, operations, and future fleet optimization. We are also looking for someone who is...SeniorFleetContract workFlexible hoursShift work$183k - $247.6k
...powers breakthrough innovation in AI/ML and HPC workloads. If you’re passionate about... ...team of software, hardware, and network engineers, supply chain specialists, security experts... ...solving operational challenges to our existing fleet with the goal of improving the current...SeniorFleetLocal areaFlexible hours$188k - $275k
...Learn more at What You'll Do: The Field Engineering organization at CoreWeave is dedicated to... ...InfiniBand/RoCE fabric validation and HPC performance benchmarking, defining how we operate customer bare-metal fleets at rack-level-and-up (IT service, break-fix,...SeniorFleetPermanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours$160.36k - $240.54k
...to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles.With years of real-world deployment experience... ...TeamOur hardware team is growing and we are looking for a Senior Electrical Engineer to help lead our programs in our mission to better everyday...SeniorFleetImmediate startFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior HPC Engineer - Fleet Engineering. Be the first to apply!
- senior operations technician San Jose, CA
- senior cloud service delivery manager San Jose, CA
- senior it service manager San Jose, CA
- senior project engineer San Jose, CA
- senior chief engineer San Jose, CA
- sr operations manager San Jose, CA
- senior physical design engineer San Jose, CA
- senior account director San Jose, CA
- senior director clinical development San Jose, CA
- sr accountant San Jose, CA


