Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU Systems Engineer — HPC & AI Inference (On-site)

Vast.ai

Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background to push the bleeding edge of AI, working with CUDA/C++ and a modern tech stack. Ideal candidates have advanced C++, experience with parallel frameworks, and a strong track record in high-performance systems. This role is full-time and on-site, offering equity and a fast-paced startup environment. #J-18808-Ljbffr Vast.ai

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the GPU Systems Engineer — HPC & AI Inference (On-site) in San Francisco, CA vacancy
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related...  .... This role is based on-site in San Francisco or Los Angeles... 
    Website

    Vast.ai Inc.

    San Francisco, CA
    2 days ago
  • $120k - $170k

     ...vertically integrated AI cloud engineered for AI. We own and...  ..., data centres, GPU superclusters,...  ...hands-on with GPU, HPC, or large-scale data...  ...or customer sites to provide onsite...  ...on AI training and inference clusters. Confident...  ...failures. ~ Linux systems engineering at... 
    Website
    Full time
    Remote work
    Flexible hours

    Nscale

    San Francisco, CA
    2 days ago
  • About Us Vast.ai’s cloud powers AI projects...  .... We seek engineers with strong intrinsic...  .... LOCATION: On-site at our office in...  ...re looking for a systems engineer with HPC or parallel...  ...to help scale AI inference. You’ll leverage...  ...systems to optimize GPU performance at the... 
    Website
    Full time
    Work at office

    Vast.ai Inc.

    San Francisco, CA
    1 day ago
  • $250k

     ...a rapidly scaling AI cloud infrastructure...  ...a next-generation GPU platform designed...  ..., and inference at scale. The company...  ...a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments...  ...available infrastructure systems Improve CI/CD... 
    Website
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $170k - $250k

     ...company operating in the AI space, backed by a leading...  ...is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges...  ...continued growth without legacy system constraints. Don’t...  ...early-stage equity On-site in San Francisco Salary... 
    Website
    Full time
    Visa sponsorship
    Flexible hours
    San Francisco, CA
    2 days ago
  •  ...—like the da Vinci surgical system and Ion—have transformed how...  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be...  ...models to real-time onboard inference—while serving as a core... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    2 days ago
  • Sail is building cutting-edge software to run AI inference and host agents at scale. You will own token processing at the...  ...in production. You will work with state-of-the-art engines, profiling tools, and advanced GPU techniques while collaborating with a hands-on team in... 
    Work at office

    Theory Ventures

    San Francisco, CA
    2 days ago
  • Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-...  .... You will work on high-throughput systems, latency-optimized serving, and...  ...building efficient inference stacks, GPU-aware optimization, and deep learning... 

    Kindredventures

    San Francisco, CA
    4 days ago
  • $350k

     ...a rapidly growing AI infrastructure provider...  ...AI training and inference across global cloud and GPU environments. The organization...  ...is for a Staff Site Reliability Engineer to lead the...  ...maintain observability systems, GPU telemetry...  ...environments, Slurm, or HPC schedulers... 
    Website
    Full time
    San Francisco, CA
    a month ago
  • An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team... 
    Website
    Worldwide

    Spellbrush

    San Francisco, CA
    2 days ago
  • Together AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms, architectures, and...  ...for low-latency, high-throughput inference. You will implement changes in production...  ...systems, while profiling across GPU, networking, and memory to... 

    Together

    San Francisco, CA
    3 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 

    Causal Labs

    San Francisco, CA
    1 day ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion...  ...help build the platform engineers turn to to ship AI...  ...building the global operating system for distributed,...  ...foundational engineers to lead our GPU Networking efforts,... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $150k - $200k

     ...Runpod is the AI Developer Cloud. More than...  ...than 20 billion inference requests. We closed...  ...infrastructure teams build the systems, the Reliability...  ...standards across engineering Designing incident...  ...systems. As a Site Reliability Engineer...  ...visibility into GPU performance and distributed... 
    Website
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    GrabJobs

    San Francisco, CA
    2 days ago
  •  ...Cloud, is a leader in AI cloud...  ...superintelligence. One person, one GPU. If you'd like to...  ...large-scale HPC clusters for AI workloads...  ...: operating systems, firmware, drivers,...  ...working closely with on-site deployment teams...  ...requirements back to other engineering teams on... 
    Website
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    1 day ago
  • A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning...  .... You will design distributed training systems and optimize GPU utilization while collaborating with cross... 

    Baseten

    San Francisco, CA
    4 days ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability...  ..., and have strong knowledge in GPU-accelerated inference. Excellent... 

    MakerMaker.AI

    San Francisco, CA
    21 hours ago
  • A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving...  ...candidates should have strong software engineering skills and experience with ML inference... 

    Gimlet Labs

    San Francisco, CA
    1 day ago
  •  ...looking for an experienced HPC infrastructure engineer to lead bringup,...  ...probably the largest anime AI training cluster in...  ...and the bare GPU machines, helping to make...  ...anime and large-scale GPU systems. You’re familiar with...  ...teams, and prefer on-site collaboration in... 
    Website
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    more than 2 months ago
  • $300k

     ...building out their AI and cloud platform,...  ...model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer...  ...automation of this GPU-powered infrastructure...  ...workloads, automate systems at petascale, and be...  ...performance computing (HPC) or AI/ML training... 
    Website
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $180k - $230k

    San Francisco, CAProduct Systems Engineering - Product Systems /Full time /On-siteWanna join the adventure...  ..., product, and operations teams across sites, and your work will directly influence...  ...observation, IoT connectivity, on-orbit AI, national security missions, and more.... 
    Website
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    1 day ago
  • $264.8k - $331k

    Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming vitally important in every function...  ...and optimize our training and inference framework. Post-train state of...  ...the architecture of the modern GPU cluster Experience with multi-node... 
    Website
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office

    Scale LLP

    San Francisco, CA
    3 days ago
  • $145k - $195k

    San Francisco, CAProduct Systems Engineering - Product Systems /Full time /On-siteWanna join the adventure...  ..., product, and operations teams across sites to turn technical ambiguity into clear...  ...observation, IoT connectivity, on-orbit AI, national security missions, and more.... 
    Website
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    1 day ago
  • $194k - $266k

     ...to spot a "normalized" problem and the AI-native curiosity to create a solution using...  ...or San Francisco, CA About The Role AI inference is becoming core infrastructure. Every...  ...API change. We are looking for a Senior Systems Engineer to help build that layer. This is a... 
    Website
    Temporary work
    Local area
    Flexible hours

    Cloudflare

    San Francisco, CA
    5 days ago
  • $102k - $182.71k

     ...Position OverviewWe’re seeking a Senior Search Systems Engineer to build the intelligence and automation...  ...that sits on top of our marketing and AI visibility data. This role focuses on...  ...apply internally (not on this external site).SummaryLocation: San Francisco, CA, USA... 
    Website
    Full time
    For contractors
    Remote work
    Shift work

    Autodesk

    San Francisco, CA
    1 day ago
  •  ...that is building the AI backbone for the...  ...reinforcement learning, inference, and long-term...  ...they make your AI system better. They are...  ...a Backend Software Engineer (ML Infrastructure)...  ...distributed systems, GPU workloads, and...  ...Excited to work on-site in San Francisco with... 
    Website
    Full time

    Rockstar

    San Francisco, CA
    1 day ago
  • $156.86k - $191.72k

     ...Computing Center (NERSC) is seeking a System Infrastructure / Platform Engineer to help build and manage HPC systems and Linux-based...  ...-edge technologies such as CPU/GPU clusters, parallel storage, high...  ...position requires substantial on-site presence, but is eligible for a... 
    Website
    Permanent employment
    Full time
    Remote work
    Flexible hours

    Berkeley Lab

    Berkeley, CA
    4 days ago
  • $140k - $210k

     ...destiny.Klaviyo is building an AI-first company—and that starts...  ...for a Lead People Technology Engineer to be the first dedicated AI engineer...  ...design and build intelligent systems that transform how we hire,...  ...found on our official career site. Please be cautious of job... 
    Website
    Shift work

    Klaviyo

    San Francisco, CA
    1 day ago
  • $160k - $220k

     ...Business Development Manager, AI Inference (Startup GTM) Position Type:...  ...infrastructure, and direct engineering support for teams shipping to...  ...Bay Area and able to work on-site in MountainView Preferred Prior...  ...AI infrastructure, MaaS, GPU compute, or LLM API products... 
    Website
    Work experience placement

    Intellipro, Inc.

    San Francisco, CA
    1 day ago
  • Modal is building an infrastructure layer for AI workloads, covering training, deployment, observation, and inference. You will perform hands-on inference research, selecting...  ...work with customers alongside Forward Deployed Engineers to deploy and tune models, while expanding... 

    Mixpeek

    San Francisco, CA
    21 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU Systems Engineer — HPC & AI Inference (On-site). Be the first to apply!