Staff Infra Engineer - Global GPU ML Inference
The Token Company
The Token Company in San Francisco is seeking a Member of Technical Staff for their infrastructure team. In this role, you will own the cloud systems that serve our compression API and build global low-latency, high-throughput GPU ML inference infrastructure. The ideal candidate will have solid experience in cloud infrastructure, including AWS and Docker, and a proven track record in building production environments. Additional benefits include equity, housing, food, and visa sponsorship. #J-18808-Ljbffr The Token Company
- B Capital is seeking a skilled engineer for GPU infrastructure in San Francisco. This role involves designing and operating high-performance systems for model inference, synthetic data generation, and reinforcement learning. The ideal candidate has strong GPU systems experience...Suggested
$173.5k - $331.05k
...are looking for a senior, hands-on engineer to own and evolve the cross-platform GPU rendering platform at the heart... ...collaboration skills, working in a global environment Ability to think... ...with GPU driver / hardware vendors ML inference integration (e.g., TensorRT, ONNX...SuggestedFull timeTemporary workLocal areaWorldwide- Jaide Health is seeking an engineer for their Model Efficiency team in... ...focuses on building reliable ML systems while enhancing core performance... ...techniques such as GPU/CUDA optimizations and collaborate... ...and insights into the LLM inference ecosystem. A commitment to diversity...SuggestedRemote job
$250k - $300k
...production. That means owning the inference stack end to end: profiling... ...work directly with customer engineering teams to tailor deployments... ...methods across many kinds of ML models, with an emphasis on large... ...~ Volunteer time off ~ Global travel insurance & emergency...SuggestedTemporary work- ...Senior Machine Learning Engineer for On-Device & Mobile... ...significant parts of the inference stack — from a trained... ...across NPU, mobile GPU, and desktop/laptop GPU... ...integration between the ML runtime and the game engine... ...eligible team members globally: Comprehensive health,...SuggestedFull timeWork at officeRemote workWorldwide
- Simplify in San Francisco is hiring Members of Technical Staff to build systems that accelerate LLM inference and own customer workloads end to end. You will work on high-performance kernels, inference engine internals, and production infrastructure for named accounts....
- Causal Labs is building a Large Physics foundation Model and GPU-driven compute environment to enable rapid research iteration at scale. You will design, deploy, and operate massive GPU clusters, extending Kubernetes and Slurm for efficient, multi-tenant workloads. You...
- Together AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms,... ...for low-latency, high-throughput inference. You will implement changes in production... ...style systems, while profiling across GPU, networking, and memory to improve latency...
$220k
We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures... ...management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our... ...for you. Good if you touched any of ML compilers and framework internals:...- United States Digital Space LLC is seeking an infrastructure leader to own a self-serve GPU compute platform for training and inference workloads. You will design and operate the system that lets researchers launch jobs across multi-cloud GPU fleets without manual provisioning...
$192k - $260k
A leading data and AI company is seeking a Staff Engineer to design and implement core systems for Foundation Model Serving. The ideal candidate... ...closely across teams to ensure operational excellence in GPU serving workloads. Competitive salary range of $192,000 to $26...$225k
Dormont Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and ensuring high reliability for RL and post-training workflows. The ideal candidate...$179k - $218k
...Silicon Reality" must be bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical... ...densities.Predictive Operations & Telemetry: Leverage AI/ML methodologies to analyze fleet-wide telemetry (power draws...Temporary work- ...technology company in San Francisco is seeking a Data Platform Engineer to drive architecture and implementation of core systems. The ideal... ...and demonstrates strong analytical skills in fields such as ML and statistics. Responsibilities include planning technical roadmaps...
$156k - $190k
....About the Role:As a Staff Cloud Support Engineer, you are a technical... ..., networking, and AI/ML infrastructure, and apply... ...AI infrastructure globally.What You’ll Be Working... ...NCCL, IB, GPU driver/firmware issues... ...workloads (training + inference) with performance tuning...Temporary work$207k - $290k
...portfolio serves 2,000+ global enterprise... ...an experienced AI Engineer with deep expertise... ...our team as a Senior Staff Architect. In this... ...experience in AI/ML engineering, including... ...techniques , including inference-time search, chain-... ...(Kubernetes, GPU/TPU clusters, and cloud...WorldwideFlexible hours$197.5k - $272k
...funding and proven by global deployment, we're solving... ...looking for a great Staff AI Engineer to join our seasoned... ...own the end-to-end ML pipeline—from data ingestion... ...for execution on CPU/GPU-bound targets or... ...modern C++ (C++14/17 for inference). ~ Deep proficiency...Work at officeWorldwideFlexible hoursShift work3 days per week$220k - $280k
...Staff MLOps Engineer — Machine Learning Platform Location: New York... ...will be our internal ML research scientists.... ...deployment, distributed inference pipelines, and... ...between execution latency, GPU/CPU throughput, and cloud... ..., and validate robust infra prototypes quickly...Remote work$160k - $300k
...product development. We empower global innovators in automotive,... ...mission is to revolutionize how engineering decisions are made, turning... ...About the Role As a Senior / Staff Infrastructure Engineer at... ...distributed systems) Exposure to ML infra Personality & Values:...Work at officeVisa sponsorshipFlexible hours- ...stake real consequences on. As a Staff Machine Learning Engineer, you’ll own AI-driven products end... ...low-latency, high-concurrency inference (Triton, vLLM, GPU-backed serving) that stays fast and... ...record of shipping and operating ML-driven functionality. ~ Mastery...Full timeContract workRemote workFlexible hours
$220k - $280k
...Together AI is building the best inference infrastructure for voice applications.... ...and reliability. We're looking for a Staff ML Engineer to drive the model serving layer for voice... ...throughput to the frontier. You'll profile GPU utilization, design batching strategies...Full time$215k - $322k
...Join GoFundMe as our next Staff Machine Learning Engineer (Pricing) . In this role, you... ...expertise in building production ML systems (data → training → online inference → measurement) with rigorous experimentation... ...@gofundme.com . Global Data Privacy Notice for Job...Full timeTemporary workWork at officeFlexible hours- ...We are looking for a Staff MLE to lead the technical... ...systems that power our global marketplace. What... ...state-of-the-art applied ML projects for ads... ...understand intention and infer interests from online activity... .... Coach and mentor engineers while collaborating...Full timeWork at officeRemote workRelocationRelocation package
$197.3k - $313.7k
...OPPORTUNITIES*Slack is looking for a Staff Machine Learning Engineer with deep expertise in... ...and finetuning to join our ML team. You'll design, train,... ...training pipelines on GPU infrastructure.Brainstorm with... ...model optimization for inference (quantization, pruning, speculative...Full time- ...over 2,000+ major global customers, approaching... ...an experienced AI Engineer with deep expertise... ...team as a Senior Staff Architect. In this... ...of experience in AI/ML engineering, including... ..., including inference‑time search, chain‑... ...infrastructure (Kubernetes, GPU/TPU clusters, and...Flexible hours
- A cutting-edge AI research firm in San Francisco is seeking talent to build and optimize GPU infrastructure for large-scale model inference and training workloads. The ideal candidate will have hands-on experience with GPU systems and optimization techniques, actively...
- ...Physics foundation Model and seeks an infrastructure engineer to design, deploy, and operate its GPU-driven compute environment. You will enable research... ...optimizing distributed clusters that power training and inference workloads. You will extend orchestration, implement...
- Claryo is seeking a Staff Software Engineer with a focus on Computer Vision Deployment based in San Francisco. The successful candidate will develop... ...include creating and managing distributed cloud GPU infrastructures and building comprehensive computer vision pipelines...Work at office3 days per week
- ...Description Job Description Staff Machine Learning Engineer, Artificial Intelligence (... ...-grade Machine Learning (ML) systems. This role sits... ..., evaluation systems, inference architecture, and deployment... ...production deployment, including GPU optimization, memory...Remote workWork from home
$190.9k - $232.8k
A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Infra Engineer - Global GPU ML Inference. Be the first to apply!
- software engineer staff San Francisco, CA
- assistant engineer San Francisco, CA
- engineering aide San Francisco, CA
- staff engineer San Francisco, CA
- staff security engineer San Francisco, CA
- assistant mechanical engineer San Francisco, CA
- assistant engineering manager San Francisco, CA
- senior staff systems engineer San Francisco, CA
- technology administrator San Francisco, CA
- project engineer assistant project manager San Francisco, CA




