GPU Performance Engineer: Scale AI Inference
$315kAnthropic
A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams, and troubleshooting performance issues. The position offers a competitive salary range of $315,000—$560,000, equity opportunities, and a supportive work environment that values communication and collaboration. #J-18808-Ljbffr Anthropic
- ...seeking a talented software engineer to join their dynamic Inference team. This role involves designing... ...infrastructure for large-scale multimodal models, focusing on high-performance delivery of audio and image... ...to push the boundaries of AI technology, ensuring reliable...Performance
- A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal...Performance
- Sciforium is seeking a GPU Kernel Engineer to push performance on modern accelerators. You will design and optimize custom GPU kernels, from... ...to integrating ops in ML frameworks used for large-scale training and inference. Ideal candidates have 5+ years in GPU kernel...Performance
- ...powers mission-critical inference for the world's most dynamic AI companies, like... ...build the platform engineers turn to to ship AI... ...multi-modal workloads scale, the network is the... ...to lead our GPU Networking efforts,... ...validate networking performance on bleeding-edge clusters...PerformanceFull timeFlexible hours
$250k
...opportunities? Join a rapidly scaling AI cloud infrastructure... ...a next-generation GPU platform designed for... ...experimentation, and inference at scale. The company... ...Site Reliability Engineer to support and scale large... ...reliability, scalability, and performance of HPC and cloud...PerformanceFull timeRemote work- ...patients worldwide.We’re a team of engineers, clinicians, and innovators... ..., our work helps care teams perform with greater precision and... ...PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be... ...models to real-time onboard inference—while serving as a core contributor...PerformanceLocal areaWorldwideFlexible hours
$325k
A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate... ...collaboration with researchers and focus on performance optimization. Compensation ranges from $325K to $4...Performance- ...About the Team Our Inference team brings OpenAI’s most... ...our state-of-the-art AI models, allowing them... ...before. We focus on performant and efficient model inference... ...Role We’re hiring engineers to scale and optimize OpenAI’s... ...across emerging GPU platforms. You’ll work...PerformanceFull time
$160k - $230k
...the RoleAt Together.ai, we are building... ...efficient and scalable inference for large language... ...the boundaries of performance, scalability, and... ...and Optimization Engineer to design, develop... ...models at scale. This role will focus... ...throughput inference, GPU/accelerator...PerformanceFull time$170k - $250k
...infrastructure company operating in the AI space, backed by a leading... ...within six months and is scaling rapidly with a small, high-performing team. This company is looking for an engineer to work directly with the CTO on complex GPU virtualisation challenges. The...PerformanceFull timeVisa sponsorshipFlexible hours$188k - $275k
...Essential Cloud for AI™. Built for... ...innovators to build and scale AI with confidence... ...infrastructure performance with deep technical... ...Do: The Field Engineering organization at CoreWeave... ...can train and inference on at scale,... ...lifecycle: leading new GPU cluster bring-up...PerformancePermanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours$140k - $200k
...pioneering the future of Physical AI. Our advanced vision... ...Data Infrastructure / Quality Engineering role will play a crucial... ...validation, and production-scale release Experience architecting... ...and optimizing high-performance GPU cloud inference services, with specific expertise...PerformanceWork experience placementLocal area- ...build and operate the inference systems that serve... .... This is an engineering role, not a research... ...Own the performance characteristics of... ...production‑grade, large‑scale serving infrastructure... ...Experience with GPU‑accelerated inference... ...#J-18808-Ljbffr MakerMaker.AIPerformance
$100k - $120k
Coda Robotics is scaling the compute infrastructure that powers... ...models. As training and inference workloads grow, we need... ...team of kernel and system engineers focused on performance-critical code Design, implement... ...for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware...Performance- Sciforium is an AI infrastructure company developing... ...-on support from AMD engineers the team is scaling rapidly to build the... ...a highly skilled GPU Kernel Engineer who... ...pushing the limits of performance on modern... ...large-scale training and inference. This role is ideal...PerformanceFlexible hours
- ...-a-Service (TaaS) Engineer to help build the... ...that convert large-scale infrastructure capacity... ...will work across performance benchmarking,... ...infrastructure stack, ensuring GPU capacity can be... ...GPU clusters, AI infrastructure,... ...model porting, inference/training workloads...PerformanceFull time
$342k
...demands of advanced AI workloads. The... ...About the RoleAs an Engineer on our hardware optimization... ...and performance. You will work with... ...efficient training and inference on our models. If... ...decisions on scale up, scale out, front... ...understanding of GPU and/or other AI acceleratorsExperience...PerformanceWork at officeLocal areaRelocation packageFlexible hours$170k - $245k
...accelerate the progress of AI applications out into the real... ...or data scientist can scale an ML application from their... ...the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations... ...that push the boundaries of performance for inference at large scale...PerformanceWork at office$264.8k - $331k
...Systems Research Engineer, Agent Post-training... ...our society. At Scale, our mission is to... ...the development of AI applications. For... ...algorithms to reach the performance necessary for... ...our training and inference framework.Post-train... ...of the modern GPU clusterExperience...PerformanceFull time$155k - $269k
...Description Waabi, founded by AI visionary Raquel... ...Scientists and Engineers building the content backbone... ...and curate assets at the scale of tens of thousands of... ..., and debugging GPU jobs in the cloud (AWS,... ...incentive awards and an annual performance bonus. Perks/...PerformanceFull timeWork at officeWork from homeFlexible hours$300 per month
...vertically integrated AI infrastructure company... ...urgency, who believe in the scale of our ambition and... ...and be part of a high-performing team that believes in... ..., our Production Engineering team ensures the reliability... ...AI pipelines and inference servicesDefine, measure...PerformanceTemporary work$350k
...Join a rapidly growing AI infrastructure... ...provider delivering large-scale compute solutions for AI training and inference across global cloud and GPU environments. The... ...build reliable, high-performance platforms supporting... ...Site Reliability Engineer to lead the reliability...PerformanceFull time$264.8k - $331k
...Systems Research Engineer, Agent Post-training... ...Enterprise GenAI AI is becoming vitally... ...our society. At Scale, our mission is to... ...algorithms to reach the performance necessary for... ...our training and inference framework. Post-train... ...of the modern GPU cluster Experience...PerformanceFull timeContract workFor contractorsFor subcontractorWork at office$315k
...interpretable, and steerable AI systems. We want AI to... ...researchers, engineers, policy experts, and business... ...innovations in GPU performance and systems engineering... ...performance at unprecedented scale, developing cutting-... ...dramatically improve inference efficiency. Working...PerformanceWork at officeVisa sponsorshipFlexible hours$350k
...knowledge and tools to make AI work for their unique... ...We are scientists, engineers, and builders who’ve created... ...for architecting and scaling the core infrastructure... ...clusters with GPU workloads, or building... ...controllers/operators, or performance profiling. Familiarity...PerformanceFull timeLocal areaImmediate startVisa sponsorshipWork visaRelocation packageFlexible hours$286.2k - $326.7k
...Senior Distinguished Engineer, AI Compute (Remote Eligible... ...and scalable, high-performance AI infrastructure. At... ...system delivering the high-scale developer and runtime... ...on top of CPU and GPU substrates. Your contributions... ...model training, model inference and feature generation...PerformanceFull timePart timeLocal areaRemote work- ...build, and operate the platform that schedules AI workloads across thousands of nodes. This... ...of architecture, reliability, and performance in a fast-growing environment. You will collaborate with founders and engineers from Nvidia, Google AI, Intel, and Pixie Labs...Performance
$189.6k - $237k
Scale’s ML platform (RLXF) team builds our internal... ...model training and inference. The platform has been... ...heart of the field of AI as an indispensable provider... ...software engineering skills, proficient in... ...qualifications, interview performance, and relevant education...PerformanceFull time- ...About the Team Our Inference team brings OpenAI... ...start-of-the-art AI models, allowing... ...before. We focus on performant and efficient... ...are looking for an engineer who wants to take... ...FLOP and every GB of GPU RAM of our hardware... ...rapidly increasing scale. Are self-...PerformanceFull time
$193k - $234k
...vertically integrated AI infrastructure company... ...urgency, who believe in the scale of our ambition and... ...and be part of a high-performing team that believes in each... ...Network Production Engineer to lead the physical and... ...performance compute (HPC) and GPU-based AI infrastructure...PerformanceTemporary workRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU Performance Engineer: Scale AI Inference. Be the first to apply!
- acting performance San Francisco, CA
- performance specialist San Francisco, CA
- high performance computing engineer San Francisco, CA
- performance windows San Francisco, CA
- system performance engineer San Francisco, CA
- performance improvement specialist San Francisco, CA
- performance testing San Francisco, CA
- human performance consultant San Francisco, CA
- performance coach San Francisco, CA
- application performance engineer San Francisco, CA




