GPU Performance Engineer: Scale AI Inference
$315kAnthropic
A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams, and troubleshooting performance issues. The position offers a competitive salary range of $315,000—$560,000, equity opportunities, and a supportive work environment that values communication and collaboration. #J-18808-Ljbffr Anthropic
- ...seeking a talented software engineer to join their dynamic Inference team. This role involves designing... ...infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image... ...to push the boundaries of AI technology, ensuring reliable...Performance
- A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal...Performance
- Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open...Performance
- ...powers mission-critical inference for the world's most dynamic AI companies, like... ...build the platform engineers turn to to ship AI... ...multi-modal workloads scale, the network is the... ...to lead our GPU Networking efforts,... ...validate networking performance on bleeding-edge clusters...PerformanceFull timeFlexible hours
$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure... ...a next-generation GPU platform designed for... ...experimentation, and inference at scale. The company... ...Site Reliability Engineer to support and scale large... ...reliability, scalability, and performance of HPC and cloud...PerformanceFull timeRemote work- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ..., our work helps care teams perform with greater precision and... ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be... ...models to real-time onboard inference—while serving as a core contributor...PerformanceLocal areaWorldwideFlexible hours
$220k - $320k
inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques...Performance- Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device... .... You will primarily work on uzu, our inference engine, and focus on supporting new... ...models work, and experience in writing high-performance GPU kernels or Rust systems programming. We...PerformanceLocal area
$325k
A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate... ...collaboration with researchers and focus on performance optimization. Compensation ranges from $325K to $4...Performance- Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize...Performance
- ...Team OpenAI’s Inference team ensures that... ...reliably, and at scale. We build and optimize... ...- to increase performance, flexibility, and... ...Role We’re hiring engineers to scale and optimize... ...across emerging GPU platforms. You’ll... ...OpenAI OpenAI is an AI research and...PerformanceFull time
- Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible for building high-performance schedulers and optimizing global routing while...Performance
$170k - $250k
...infrastructure company operating in the AI space, backed by a leading... ...within six months and is scaling rapidly with a small, high-performing team. This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The...PerformanceFull timeVisa sponsorshipFlexible hours$160k - $230k
...the RoleAt Together.ai, we are building... ...efficient and scalable inference for large language... ...the boundaries of performance, scalability, and... ...and Optimization Engineer to design, develop... ...models at scale. This role will focus... ...throughput inference, GPU/accelerator...PerformanceFull time- ...foundation Model and seeks infrastructure engineers for high-throughput, low-latency inference at scale in San Francisco. You will work on... ...collaborating with researchers to push model performance. We value deep learning expertise, GPU-aware optimization, and production-...Performance
- A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying... ...GPU systems for high-throughput inference and model performance optimization. The ideal...Performance
- ...NEAR AI was started by Illia Polosukhin, co... ...source AI at a global scale. We are specifically... ...an expert in high‑performance LLM serving systems and inference optimization. In this... ...major inference engines such as SGLang, vLLM... ...of state‑of‑the‑art GPU architectures, and effectively...Performance
$300k
...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling... ...team is building low-latency AI systems where milliseconds... ...runtime layers, profiling large-scale speech and multimodal models...PerformanceRelocationVisa sponsorshipFree visa$350k
Mirendil is looking for engineers to build infrastructure for frontier reasoning models at... ...Francisco location. This role focuses on large-scale reinforcement learning (RL) model... ...reliable training infrastructure, implement performance optimizations, and develop evaluation...Performance$100k - $120k
Coda Robotics is scaling the compute infrastructure that powers... ...models. As training and inference workloads grow, we need... ...team of kernel and system engineers focused on performance-critical code Design, implement... ...for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware...Performance- ...build and operate the inference systems that serve... .... This is an engineering role, not a research... ...Own the performance characteristics of... ...production‑grade, large‑scale serving infrastructure... ...Experience with GPU‑accelerated inference... ...#J-18808-Ljbffr MakerMaker.AIPerformance
- ...-a-Service (TaaS) Engineer to help build the... ...that convert large-scale infrastructure capacity... ...will work across performance benchmarking,... ...infrastructure stack, ensuring GPU capacity can be... ...GPU clusters, AI infrastructure,... ...model porting, inference/training workloads...PerformanceFull time
- About Us Vast.ai’s cloud powers AI projects and businesses... ...excellence. We seek engineers with strong intrinsic drive... ...programming experience to help scale AI inference. You’ll leverage your knowledge of high-performance systems to optimize GPU performance at the bleeding...PerformanceFull timeWork at office
- CoreWeave is seeking a Bare Metal Support Engineer in San Francisco, CA, to ensure high performance and reliability of our GPU infrastructure. You will engage directly with customers... ...support processes. Join us to be part of the AI infrastructure revolution! #J-18808-Ljbffr...Performance
$342k
...demands of advanced AI workloads. The... ...About the RoleAs an Engineer on our hardware optimization... ...and performance. You will work with... ...efficient training and inference on our models. If... ...decisions on scale up, scale out, front... ...understanding of GPU and/or other AI acceleratorsExperience...PerformanceWork at officeLocal areaRelocation packageFlexible hours$180k - $250k
...the next generation of AI products. We build the... ..., and do it at scale without compromise. For... ...unified platform where high-performance inference, orchestration, and... ...experienced software engineer who thrives on building... ...orchestration, scheduling, GPU autoscaling, large...PerformanceCurrently hiringRelocation package- A tech startup focused on AI workloads is seeking a Member of... ...Staff to design and optimize inference systems. The role involves managing... ...and improving execution performance across various components.... ...should have strong software engineering skills and experience with ML...Performance
$170k - $245k
...accelerate the progress of AI applications out into the real... ...or data scientist can scale an ML application from their... ...the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations... ...that push the boundaries of performance for inference at large scale...PerformanceWork at office$227.2k - $284k
...Scale's Physical AI business unit is dedicated to solving the... ...As an ML Systems Engineer on the Physical AI team... ...fault-tolerant, high-performance systems for serving robotics... ...tracking of model inference. Lead: Own... ...environments, including GPU-level algorithm optimizations...PerformanceFull time$140k - $200k
...pioneering the future of Physical AI. Our advanced vision... ...Data Infrastructure / Quality Engineering role will play a crucial... ...validation, and production-scale release Experience architecting... ...and optimizing high-performance GPU cloud inference services, with specific expertise...PerformanceFull timeWork experience placementLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Performance Engineer: Scale AI Inference. Be the first to apply!
- performance test architect San Francisco, CA
- performance food service San Francisco, CA
- performance improvement consultant San Francisco, CA
- senior performance engineer San Francisco, CA
- performance nutrition San Francisco, CA
- system performance engineer San Francisco, CA
- IT performance management San Francisco, CA
- acting performance San Francisco, CA
- application performance engineer San Francisco, CA
- performance testing San Francisco, CA



