Senior HPC Systems Engineer: Cloud & Scalable Infra
NVIDIA
NVIDIA is seeking a Senior Software Engineer in Westford, Massachusetts to improve their HPC infrastructure. The role includes designing scalable systems and supporting multi-cloud environments.
The ideal candidate will have 10+ years of experience, strong software development skills, and a B.S. degree in Computer Science or related field. NVIDIA values diversity and offers competitive salaries, equity, and benefits.
#J-18808-Ljbffr$152k - $241.5k
...NVIDIA Gruppe in Santa Clara is seeking a Senior Software Engineer to enhance their HPC infrastructure. The role involves applying distributed systems patterns, automation, and building scalable services in a hybrid multi-cloud environment. Candidates should have strong...Senior- Velaura seeks a Senior Systems Administrator / IT Engineer to administer, secure, and improve enterprise infrastructure... ...Linux servers, virtualization, cloud platforms, networking, identity,... ...and Operations to ensure reliable, scalable IT services #J-18808-Ljbffr VelauraSenior
- ...ASML Germany GmbH is seeking an experienced engineer to develop and maintain HPC infrastructure supporting scalable workloads across products. You will design compute... ...with strong observability. You will optimize system throughput and reliability, enable...Senior
- CoreWeave is seeking a Senior Software Engineer on the Systems Engineering team to own and evolve our testing framework. You will harden the Rust core, extend coverage into HPC and Slurm-on-Kubernetes, and ensure fast, hermetic CI at fleet scale. You will build test harnesses...Senior
- ...California, is seeking an experienced database Software Engineer to help develop the next generation of Apple's cloud services. You will work on a core component of the iCloud Platform, contributing to performance, scalability, and reliability across Apple’s services. The...Senior
$255k - $340k
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure... ...day is currently Tuesday.Hardware Engineering at Lambda is responsible for building... ...this is that team.What You’ll DoOwn system integration validation for new HPC AI/ML, general purpose compute,...SeniorWork at officeLocal areaWork from homeFlexible hours$255k - $340k
...The Superintelligence Cloud, is a leader in AI cloud... ...Tuesday.Hardware Engineering at Lambda is responsible... ...integrating OEM and white-label HPC AI/ML, general purpose... ...(NPI) for hardware systems, including system... ..., performance, and scalability of new systems.Serve as...SeniorWork at officeLocal areaWork from homeFlexible hours- ...Muon Space is seeking a Senior Software Engineer for its Software Platform team in San Jose. You will help build the CI/CD, Kubernetes, Terraform... ...to architectural decisions, and mentor engineers while delivering reliable, scalable systems for mission-critical #J-18808-LjbffrSenior3 days per week
- Crusoe Cloud is revolutionizing high-performance computing with sustainable GPU compute power. As a Cloud Support Engineer, you will be the primary contact for technical support, helping customers leverage Crusoe Cloud to achieve their goals and accelerate research and...SeniorRemote work
$152k - $253k
Join NVIDIA as a Senior System Mechanical Engineer and help build the systems powering the future of AI and accelerated computing!In this role, you... ....Translate product requirements into practical, scalable mechanical designs that meet performance, reliability, and...SeniorFull time$168k - $270.25k
...to our customers! Sr Site Reliability Engineer in this role will significantly impact and... ...orchestrators for performance, scalability and resilience, ensuring they meet real-... ...performance computing, Kubernetes and/or system administration would be an assetPrevious...SeniorFull timeWorldwide$203.45k - $344.3k
...and Digital World generation system, supporting large-scale model... ...build a high-fidelity, highly scalable virtual world generation and closed... ...a synthetic data production engine for perception, prediction,... ...robotics, simulation platforms, data Infra and training platforms to...SeniorFull time- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving... ...data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base...SeniorWork at officeLocal areaWork from homeFlexible hours
- ...Walmart Global Tech in Sunnyvale, CA is seeking a Senior Software Engineer to build scalable services for the Post Transactions team. You will develop... ...requires strong experience with Kotlin, Jetpack, MVVM, and distributed systems, plus cloud familiarity. #J-18808-LjbffrSenior
- Google is seeking a Senior Software Engineer in Sunnyvale, CA to drive performance improvements across large-scale systems. You will write and test critical code, participate in design... .... The team focuses on building scalable platforms, analyzing data for performance...Senior
$45 - $49 per hour
..., Information Technology, and Engineering Industries: Staffing and Recruiting... ...notified about new Network System Specialist jobs in Santa Clara... ...4 days ago Network Engineer, HPC Systems Network Strategy Menlo... ...Graduate (Physical Network Infra) - 2026 Start (BS/ MS) San Jose...SeniorTemporary work- ...Palo Alto Networks, Inc. is seeking a Senior Principal Software Engineer to lead development of tools, platforms, and infrastructure... ...experience building large-scale distributed systems, developer platforms, and cloud-native infrastructure while mentoring engineers...Senior
- ...ultra-scale GPU supercomputing systems to train next-generation... ...performance, fault tolerance, and scalability are co-designed across model... ...for a deeply technical engineer to co-design and optimize the... ...relevant distributed systems, HPC, or large-scale training projects...SeniorVisa sponsorship
- Apple Inc. in Cupertino, California, is seeking a senior ML/AI engineer to help turn research into practical features on Apple platforms... ...nimble team that collaborates across disciplines to build scalable AI systems, with hands‑on prototyping and observable, debuggable...Senior
- ...build and operate scalable, secure, and... ...empowers platform engineering teams to... ...premises, in the cloud, at the edge, or... ...customer's ML infra lead, and closed... ...2 / B200-class systems, Grace-Hopper superchips... ...GPU cloud, HPC, or specialized... ...OTE: senior people-leader band...SeniorFull timeRemote work
- KLA is seeking a Principal HPC Architect to design, build, optimize, and support large-scale... ...researchers, developers, and IT to deliver reliable, scalable HPC infrastructure and lead engineering initiatives. The role emphasizes systems engineering, performance tuning, and DevOps,...Senior
- General Motors is seeking a Staff ML Systems Engineer to help our Data Labeling Engineering team advance... ...data science, and product teams to deliver scalable labeling workflows and AI-assisted tooling. The role emphasizes cloud-scale, distributed systems, modern full-...SeniorRemote job
$200k - $322k
...NVIDIA’s DGX Cloud team helps some of the most advanced... ...requirements into scalable recommendations across... ...friction.gainsightWork across Engineering, Product, Operations,... ...with distributed systems, Kubernetes, schedulers... ...supporting AI, ML, or HPC workloads in cloud or hybrid...SeniorFull time$153k - $204k
CoreWeave is The Essential Cloud for AI™. Built for pioneers by... ...more at What You'll Do The Systems Engineering team owns the host software stack... ...that framework into HPC verification, Slurm-on-Kubernetes... ...the stack. About The Role As a Senior Software Engineer on the Systems...SeniorPermanent employmentFull timeTemporary workCasual workLive inWork at officeFlexible hours$166k - $244k
Overview Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have...SeniorFull time$170k - $205k
...center construction, and cloud services. If you... ...Production / Sustaining Engineer to strengthen Crusoe’s Hardware Systems Engineering team and close... ...performance, reliability, and scalability expectations.... ...to leverage them in AI/HPC environments. Expertise...Senior- ...Azure NetApp Files (ANF) team builds cloud-scale storage services and distributed systems spanning storage, networking, and... ...seek a highly motivated software engineer to own features end-to-end,... ...The role focuses on developing scalable backend services in Go, with Python...Senior
$187k - $270.7k
...Altera in San Jose is seeking a DevOps CI/CD Software Engineering Manager to lead the design and operation of advanced CI/CD platforms... ...strong technical expertise and leadership abilities to enhance scalability and efficiency in software development processes....Senior- A leading technology company in California is looking for a Senior Software Engineer to develop cutting-edge AI and ML solutions. Responsibilities include writing and testing code, collaborating through design and code reviews, and contributing to documentation. Candidates...SeniorFull time
- Unity Technologies is seeking a Senior Software Engineer to advance the reliability, scalability, and maintainability of AI authoring tools. You will shape infrastructure... ...The role focuses on designing scalable backend systems, cloud deployments on Azure/GC, and close...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior HPC Systems Engineer: Cloud & Scalable Infra. Be the first to apply!
- senior windows systems engineer Santa Clara, CA
- software system engineer Santa Clara, CA
- system test engineer Santa Clara, CA
- system validation engineer Santa Clara, CA
- mission system engineer Santa Clara, CA
- healthcare systems engineer Santa Clara, CA
- electronic systems engineer Santa Clara, CA
- system verification engineer Santa Clara, CA
- operating system engineer Santa Clara, CA
- system engineer remote Santa Clara, CA



