GPU Systems Engineer — HPC & AI Inference (On-site)
Vast.ai
Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background to push the bleeding edge of AI, working with CUDA/C++ and a modern tech stack. Ideal candidates have advanced C++, experience with parallel frameworks, and a strong track record in high-performance systems. This role is full-time and on-site, offering equity and a fast-paced startup environment. #J-18808-Ljbffr Vast.ai
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related... .... This role is based on-site in San Francisco or Los Angeles...Website
$120k - $170k
...vertically integrated AI cloud engineered for AI. We own and... ..., data centres, GPU superclusters,... ...hands-on with GPU, HPC, or large-scale data... ...or customer sites to provide onsite... ...on AI training and inference clusters. Confident... ...failures. ~ Linux systems engineering at...WebsiteFull timeRemote workFlexible hours- About Us Vast.ai’s cloud powers AI projects... .... We seek engineers with strong intrinsic... .... LOCATION: On-site at our office in... ...re looking for a systems engineer with HPC or parallel... ...to help scale AI inference. You’ll leverage... ...systems to optimize GPU performance at the...WebsiteFull timeWork at office
$250k
...a rapidly scaling AI cloud infrastructure... ...a next-generation GPU platform designed... ..., and inference at scale. The company... ...a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments... ...available infrastructure systems Improve CI/CD...WebsiteFull timeRemote work$170k - $250k
...company operating in the AI space, backed by a leading... ...is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges... ...continued growth without legacy system constraints. Don’t... ...early-stage equity On-site in San Francisco Salary...WebsiteFull timeVisa sponsorshipFlexible hours- ...—like the da Vinci surgical system and Ion—have transformed how... ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be... ...models to real-time onboard inference—while serving as a core...Local areaWorldwideFlexible hours
- Sail is building cutting-edge software to run AI inference and host agents at scale. You will own token processing at the... ...in production. You will work with state-of-the-art engines, profiling tools, and advanced GPU techniques while collaborating with a hands-on team in...Work at office
- Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-... .... You will work on high-throughput systems, latency-optimized serving, and... ...building efficient inference stacks, GPU-aware optimization, and deep learning...
$350k
...a rapidly growing AI infrastructure provider... ...AI training and inference across global cloud and GPU environments. The organization... ...is for a Staff Site Reliability Engineer to lead the... ...maintain observability systems, GPU telemetry... ...environments, Slurm, or HPC schedulers...WebsiteFull time- An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team...WebsiteWorldwide
- Together AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms, architectures, and... ...for low-latency, high-throughput inference. You will implement changes in production... ...systems, while profiling across GPU, networking, and memory to...
- Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize...
- ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion... ...help build the platform engineers turn to to ship AI... ...building the global operating system for distributed,... ...foundational engineers to lead our GPU Networking efforts,...Full timeFlexible hours
$150k - $200k
...Runpod is the AI Developer Cloud. More than... ...than 20 billion inference requests. We closed... ...infrastructure teams build the systems, the Reliability... ...standards across engineering Designing incident... ...systems. As a Site Reliability Engineer... ...visibility into GPU performance and distributed...WebsiteRemote workVisa sponsorshipWork visaFlexible hours- ...Cloud, is a leader in AI cloud... ...superintelligence. One person, one GPU. If you'd like to... ...large-scale HPC clusters for AI workloads... ...: operating systems, firmware, drivers,... ...working closely with on-site deployment teams... ...requirements back to other engineering teams on...WebsiteFull timeWork at officeLocal areaRemote workWork from homeFlexible hours
- A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning... .... You will design distributed training systems and optimize GPU utilization while collaborating with cross...
- MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability... ..., and have strong knowledge in GPU-accelerated inference. Excellent...
- A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving... ...candidates should have strong software engineering skills and experience with ML inference...
- ...looking for an experienced HPC infrastructure engineer to lead bringup,... ...probably the largest anime AI training cluster in... ...and the bare GPU machines, helping to make... ...anime and large-scale GPU systems. You’re familiar with... ...teams, and prefer on-site collaboration in...WebsiteWork at officeVisa sponsorship
$300k
...building out their AI and cloud platform,... ...model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer... ...automation of this GPU-powered infrastructure... ...workloads, automate systems at petascale, and be... ...performance computing (HPC) or AI/ML training...WebsitePermanent employment$180k - $230k
San Francisco, CAProduct Systems Engineering - Product Systems /Full time /On-siteWanna join the adventure... ..., product, and operations teams across sites, and your work will directly influence... ...observation, IoT connectivity, on-orbit AI, national security missions, and more....WebsiteFull timeTemporary work$264.8k - $331k
Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming vitally important in every function... ...and optimize our training and inference framework. Post-train state of... ...the architecture of the modern GPU cluster Experience with multi-node...WebsiteFull timeContract workFor contractorsFor subcontractorWork at office$145k - $195k
San Francisco, CAProduct Systems Engineering - Product Systems /Full time /On-siteWanna join the adventure... ..., product, and operations teams across sites to turn technical ambiguity into clear... ...observation, IoT connectivity, on-orbit AI, national security missions, and more....WebsiteFull timeTemporary work$194k - $266k
...to spot a "normalized" problem and the AI-native curiosity to create a solution using... ...or San Francisco, CA About The Role AI inference is becoming core infrastructure. Every... ...API change. We are looking for a Senior Systems Engineer to help build that layer. This is a...WebsiteTemporary workLocal areaFlexible hours$102k - $182.71k
...Position OverviewWe’re seeking a Senior Search Systems Engineer to build the intelligence and automation... ...that sits on top of our marketing and AI visibility data. This role focuses on... ...apply internally (not on this external site).SummaryLocation: San Francisco, CA, USA...WebsiteFull timeFor contractorsRemote workShift work- ...that is building the AI backbone for the... ...reinforcement learning, inference, and long-term... ...they make your AI system better. They are... ...a Backend Software Engineer (ML Infrastructure)... ...distributed systems, GPU workloads, and... ...Excited to work on-site in San Francisco with...WebsiteFull time
$156.86k - $191.72k
...Computing Center (NERSC) is seeking a System Infrastructure / Platform Engineer to help build and manage HPC systems and Linux-based... ...-edge technologies such as CPU/GPU clusters, parallel storage, high... ...position requires substantial on-site presence, but is eligible for a...WebsitePermanent employmentFull timeRemote workFlexible hours$140k - $210k
...destiny.Klaviyo is building an AI-first company—and that starts... ...for a Lead People Technology Engineer to be the first dedicated AI engineer... ...design and build intelligent systems that transform how we hire,... ...found on our official career site. Please be cautious of job...WebsiteShift work$160k - $220k
...Business Development Manager, AI Inference (Startup GTM) Position Type:... ...infrastructure, and direct engineering support for teams shipping to... ...Bay Area and able to work on-site in MountainView Preferred Prior... ...AI infrastructure, MaaS, GPU compute, or LLM API products...WebsiteWork experience placement- Modal is building an infrastructure layer for AI workloads, covering training, deployment, observation, and inference. You will perform hands-on inference research, selecting... ...work with customers alongside Forward Deployed Engineers to deploy and tune models, while expanding...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Systems Engineer — HPC & AI Inference (On-site). Be the first to apply!
- lead system engineer San Francisco, CA
- distributed systems engineer San Francisco, CA
- operations support system engineer San Francisco, CA
- computer systems engineer San Francisco, CA
- system performance engineer San Francisco, CA
- unix linux systems engineer San Francisco, CA
- microsoft systems engineer San Francisco, CA
- mission system engineer San Francisco, CA
- system engineer remote San Francisco, CA
- application system engineer San Francisco, CA



