Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior HPC Engineer - Fleet Engineering

Lambda Labs

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.If you'd like to build the world's best AI cloud, join us.What You’ll DoBuild and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and job-level signals — to detect and respond to issues proactivelyRemotely deploy and configure large-scale HPC clusters for AI workloads using automation wherever possibleAutomate cluster lifecycle: operating systems, firmware, drivers, and networking, managed as code (Ansible, Terraform) rather than by handCreate runbooks and automated remediations for common cluster failure modes, designed so Support and HPC Support can run them safelyTroubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power — working closely with on-site deployment teamsParticipate in on-call rotations and lead incident response for cluster-level problemsContribute to and maintain Standard Operating Procedures, and feed clear requirements back to other engineering teams on simplification, stability, and operational efficiencyYou7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar roleHave a strong understanding of modern AI infrastructure, from GPU architectures to hardware performance optimizationStrong understanding of Linux-based systems in a distributed environmentAre experienced configuring and troubleshooting InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, Ethernet/switching, GPU-direct, and NCCL environmentsSolid understanding of Python and Go, with experience working with SWE teams to improve internal tooling.Experience with monitoring and alerting tools (e.g., Prometheus, Grafana, Clickhouse)Proficiency in automation and configuration management tools (e.g., Ansible, Terraform)Have excellent problem-solving and troubleshooting skills and an innate attention to detailPassion for continuous improvement and innovationNice to HaveExperience with machine learning / deep learning frameworks (PyTorch, TensorFlow) and benchmarking tools (DeepSpeed, MLPerf)Knowledge of containerization and orchestration technologies (e.g., Docker, Kubernetes)Experience building and/or operating HPC resources.Depth in the NVIDIA hardware and firmware ecosystemExperience with data center power and thermal designBackground in chaos engineering or similar reliability testing methodologiesUnderstanding of compliance frameworks (SOC 2, ISO 27001, etc.)Salary Range InformationThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.About LambdaFounded in 2012, with 500+ employees, and growing fastOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent CoveWe have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOGOur values are publicly available: We offer generous cash & equity compensationHealth, dental, and vision coverage for you and your dependentsWellness and commuter stipends for select roles401k Plan with 2% company match (USA employees)Flexible paid time off plan that we all actually useEqual Opportunity EmployerLambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.Compensation Range: $227K - $356KLocationSan Francisco Office (Fremont St); Bellevue Office; Remote, USA; San Jose Office (First St)Employment TypeFull timeLocation TypeHybridDepartmentData Center BusinessCompensationSan Francisco / San JoseSan Francisco / San Jose $267K – $356KBellevueBellevue $240K – $320KRemote, USA$227K – $303K

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior HPC Engineer - Fleet Engineering in San Jose, CA vacancy
  • $255k - $340k

     ...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and...  ..., new product introduction (NPI), and fleet-scale maintenance — all engineered for gigawatt...  ...lead for integrating OEM and white-label HPC AI/ML, general purpose compute, storage,... 
    Senior
    Fleet
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...intelligence.We’re looking for a Senior SRE to join our Compute Farm...  ...they integrate cleanly with HPC schedulers, storage, and network...  ...host lifecycle management, fleet reliability/auto-healing, E2E...  ...Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Senior
    Fleet
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...recovery, resizing, and incident response using fleet management toolsParticipate in a well-...  ...and authenticationWork closely with our HPC Ops and Datacenter Ops teams for low-... 
    Senior
    Fleet
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $152k - $241.5k

     ...you to contribute to our team. We are looking for a Senior Software Engineer to join our DGX Cloud / Fleet Intelligence team and build agent-side systems that...  .../log export pipelines and operating software in AI, HPC, cloud, or large-scale datacenter environments.Track... 
    Senior
    Fleet
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    15 hours ago
  • $267k - $356k

     ...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-...  ...capacity health of Lambda's production storage fleet across all data centers, operating behind...  ...operating Linux systems in production or HPC environments, with hands-on storage... 
    Senior
    Fleet
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA has become the platform upon which every new AI-powered application is built. We are seeking a Sr. HPC Performance engineer to join our team of scientists and engineers passionate about building the next generation of scientific machine learning (ML) frameworks.... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...implement scalable, next-gen distributed storage services for HPC workloads, optimizing both performance and cost-effectiveness to...  ...need to see:Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience.8+ years of experience... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA Math Libraries team is looking for a senior engineer to join our development efforts in the area of kernel generation for AI and HPC, specifically targeting matrix operations, JITing and fusions. Around the world, leading commercial and academic organizations are... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...environment remains resilient, measurable, and aligned with long-term engineering demands.What you'll be doing:Manage, scale, and optimize job...  ...and tuning job scheduling systems (LSF, Slurm, etc.) in HPC or silicon design environmentsProficiency in Linux systems administration... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...-world production. We are looking for a Reinforcement Learning Engineer to join our mission in building intelligent, safe, and efficient...  ...large-scale data flywheels, including mining scenarios from fleet telemetry logs, auto-labeling pipelines, and automated performance... 
    Senior
    Fleet
    Full time

    Nvidia

    Santa Clara, CA
    15 hours ago
  • $148.7k - $195k

     ...making the transition to electric easy for businesses, fleets and drivers. ChargePoint offers a once-in-a-lifetime...  ...You Will Be DoingChargePoint is seeking a Senior Program Manager for Hardware Engineering to lead initiatives related to our EV charging hardware... 
    Senior
    Fleet
    Contract work

    ChargePoint

    Campbell, CA
    4 days ago
  • $150k - $200k

    Thanks for your interest in Oklo! We are searching for a senior maintenance program engineer to join our engineering team.Position DescriptionAt Oklo,...  ...development, implementation, and continuous improvement of fleet maintenance standards across Oklo facilities. The... 
    Senior
    Fleet
    Remote work
    Flexible hours

    Oklo

    Santa Clara, CA
    4 days ago
  • $168k - $264.5k

     ...late in a program, it usually breaks here first.We are hiring a Senior Engineer to lead system validation, debug, and cross-functional...  ...PVT stress, system-stress campaigns at scale, and multi-unit fleet testing.Lead debug of the hardest cross-stack issues — logic,... 
    Senior
    Fleet
    Full time
    Flexible hours
    Shift work

    Nvidia

    Santa Clara, CA
    15 hours ago
  • $200k - $230k

     ...solely on making the transition to electric easy for businesses, fleets and drivers. ChargePoint offers a once-in-a-lifetime...  ...looking for an experienced Power Electronics Controls and Firmware Engineer with more than seven years of hands-on expertise in developing... 
    Senior
    Fleet

    ChargePoint

    Campbell, CA
    4 days ago
  • $184k - $287.5k

     ...for AI research and production workloads. We are looking for Senior Software Engineers to help build the automation, tooling, and operational...  ...infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation.Experience with SLOs, on-call, incident response... 
    Senior
    Fleet
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

    Join NVIDIA's Isaac Applications Engineering team and help build the platform for Physical AI robots — assembling the full robotics stack...  ...managing GPU-backed CI infrastructure or self-hosted runner fleets.Your base salary will be determined based on your location, experience... 
    Senior
    Fleet
    Full time
    Live in
    Night shift

    Nvidia

    Santa Clara, CA
    4 days ago
  • $130k - $175k

     ...RoleTarana is scaling field deployments 5×, and we’re looking for a senior engineer to own the technical execution and continuous improvement of...  ...cross-functional processes that allow HW QA to scale with the fleet.As a Senior Hardware QA & Diagnostics Engineer you will have... 
    Senior
    Fleet
    Worldwide
    Flexible hours

    Tarana Wireless

    Milpitas, CA
    4 days ago
  • $148k - $235.75k

     ...make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw...  ...into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer to operate the platform itself... 
    Senior
    Fleet
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $101k - $161k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...responsible for our global CloudVision service fleet, ensuring scalability, reliability, and...  ...timeFunction: EngineeringExperience level: Mid-Senior LevelIndustry: Computer Networking
    Senior
    Fleet

    Arista Networks

    Santa Clara, CA
    2 days ago
  • $143.3k

     ...the best customer experience in cloud computing.We aim to hire engineers who will thrive in a fast-paced, collaborative and open environment...  ...responsible for solving operational challenges to our existing fleet with the goal of improving the current customer experience as... 
    Senior
    Fleet
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $183k - $247.6k

     ...You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers...  ...servers to the data center. After launch you will oversee the fleet of servers you develop, monitoring their quality and how they are... 
    Senior
    Fleet
    Local area
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    2 days ago
  •  ...transition to electric easy for businesses, fleets and drivers. ChargePoint offers a once-in-...  .... Reports To Director, Hardware Design Engineering What You Will Be Doing The Hardware Design Engineering Team is seeking a Senior Hardware Design Engineer - Interconnects to... 
    Senior
    Fleet

    ChargePoint

    Campbell, CA
    1 day ago
  • $163.43k - $213.97k

     ...The Role: We are looking for a deeply hands-on Senior Distributed Systems Engineer to join the team building IonQ's Network and Security Platform...  ...pipelines collecting telemetry from large device fleets, to time-series storage, real-time processing engines, and... 
    Senior
    Fleet
    Permanent employment
    Contract work
    Work at office

    IonQ Inc.

    Santa Clara, CA
    5 days ago
  • $173.9k - $235.2k

     ...workloads.We are looking for an experienced System development engineer to drive development for new EC2 machine learning platforms. In...  ...recommending best practices for maintaining and improving code quality, fleet health, and security & reliability of our service.- Growing our... 
    Senior
    Fleet
    Internship
    Local area
    Work from home
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  • $183k - $247.6k

    Do you like to apply deep hardware engineering expertise to deliver simple, sustainable, and scalable solutions for large networks? Would...  ...global AWS network. Beyond product delivery we actively manage the fleet or routers in a network that grows by 70% annually. This means... 
    Senior
    Fleet
    Work experience placement
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    15 hours ago
  • $175k - $210k

     ...efficiencies. About the role We are looking for a talented Senior Robotics Software Engineer with expertise in robotics software development to join...  ...the autonomous software powering the Bonsai Intelligence fleet across a multitude of applications and environments Write... 
    Senior
    Fleet
    Work at office

    Bonsai Robotics

    San Jose, CA
    2 days ago
  • $154k - $200k

     ...industrial heat and power at scale. We are looking for a  Senior / Staff Turbomachinery Systems Engineer to lead the definition, integration, and execution of...  ...execution, commissioning, operations, and future fleet optimization. We are also looking for someone who is... 
    Senior
    Fleet
    Contract work
    Flexible hours
    Shift work

    Antora Energy

    San Jose, CA
    9 days ago
  • $183k - $247.6k

     ...powers breakthrough innovation in AI/ML and HPC workloads. If you’re passionate about...  ...team of software, hardware, and network engineers, supply chain specialists, security experts...  ...solving operational challenges to our existing fleet with the goal of improving the current... 
    Senior
    Fleet
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    5 days ago
  • $188k - $275k

     ...Learn more at  What You'll Do: The Field Engineering organization at CoreWeave is dedicated to...  ...InfiniBand/RoCE fabric validation and HPC performance benchmarking, defining how we operate customer bare-metal fleets at rack-level-and-up (IT service, break-fix,... 
    Senior
    Fleet
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    13 days ago
  • $160.36k - $240.54k

     ...to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles.With years of real-world deployment experience...  ...TeamOur hardware team is growing and we are looking for a Senior Electrical Engineer to help lead our programs in our mission to better everyday... 
    Senior
    Fleet
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior HPC Engineer - Fleet Engineering. Be the first to apply!