AI Infrastructure Engineer: Scalable GPU Inference, On-Site
Spellbrush
An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team, you will work closely with top AI researchers, driving the evolution of generative AI in the gaming industry. If you are passionate about anime aesthetics and enjoy a fast-paced environment, this opportunity will allow you to contribute to a creative movement that impacts millions of users worldwide. #J-18808-Ljbffr Spellbrush
- ...Lead Software Engineer at JPMorgan Chase... ...Sector, Infrastructure Platforms team,... ...secure, stable, and scalable way. Drive significant... ...optimized for AI/ML workloads.... ...training, and inference.Experience with... ...of NVIDIA GPU infrastructure... ...care coverage, on-site health and wellness...WebsiteFor contractors
- ...building a self-serve compute platform that lets inference engineers run training jobs and inference services without worrying about GPU provisioning or cluster configuration. You... ...that hides complexity behind a unified platform and scalable #J-18808-Ljbffr Neura MarketSuggested
- ...’s leading generative AI studio behind niji・journey... ...is looking for an AI Infrastructure Engineer to join us in building... ...our next-generation inference architecture for... ...excellent understanding of GPU’s handling large... ...teams, and prefer on-site collaboration in either...WebsiteWork experience placementWork at officeVisa sponsorship
$269.1k - $307.2k
Distinguished AI Engineer (Agentic AI Platform)... ...investments in technology infrastructure and world-class talent... ...product experiences and scalable, high-performance AI... ...(e.g. LLM Inference, Similarity Search and... ...available through this site. Capital One...WebsiteFull timePart timeWork at officeLocal area- United States Digital Space LLC is seeking an infrastructure leader to own a self-serve GPU compute platform for training and inference workloads. You will design and operate the... ..., and a coherent platform roadmap for scalable AI workloads. #J-18808-Ljbffr United States...Suggested
$220k
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience...- ...Team OpenAI’s Inference team ensures that... ...resiliency across our infrastructure. We are forming a... ...We’re hiring engineers to scale and optimize... ...infrastructure across emerging GPU platforms. You’ll... ...performance, and scalability of model execution... ...OpenAI is an AI research and...Full time
- ...Artificial Intelligence Institute in San Francisco Bay Area seeks an AI Infrastructure Engineer to advance an inference stack that spans multiple chips and models. You will build optimization kernels, a scalable library generator, and a benchmark-driven pipeline that...Website
- ...Associates Limited is seeking a Staff-level engineer to architect and evolve the AI infrastructure software stack for large-scale GPU workloads. You’ll drive orchestration,... ...production environment. You’ll contribute to inference platforms, model serving, and high-...
$229.9k - $262.4k
Senior Lead AI Engineer (Gen AI Platform Services, Agentic... ...in technology infrastructure and world-class talent... ...product experiences and scalable, high-performance AI... ...large language model inference, similarity search, guardrails... ...through this site. Capital One...WebsiteFull timePart timeLocal area$151.8k - $265.35k
...organizations to create exceptional content effortlessly. The AI for Engineering team builds a scalable, production-grade AI platform that powers creativity... ...orchestration, tool integration, memory systems, inference services, data flows, evaluation loops, and real-time...WebsiteFull timeTemporary workLocal areaWorldwide- ...Job Description The AI Infrastructure team at Zensors builds the engine that powers our visual sensing... ...the training and inference of computer vision models... ...stream to enable massive scalability of our SaaS product.... ...Deep understanding of GPU hardware performance ,...
- OpenAI is seeking an experienced software engineer to join the GPT Infrastructure team and help build an automated inference optimization platform that scales research prototypes into production‑ready solutions. You will design durable APIs and control‑plane services,...
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience... ...to help scale AI inference. You will design and optimize GPU kernels and tensor... ...This role is based on-site in San Francisco or Los... ...HPC techniques and scalable workloads. You will collaborate...Website
- Hamilton Barnes Associates Limited is seeking an ambitious infrastructure engineer to help build and operate automated systems that bring GPU clusters from bare machines to customer-ready deployments. You will manage multi-tenant lifecycle, run Kubernetes and Postgres...Remote job
- ...vertically integrated AI infrastructure company built from... ...the Role:As an Engineering Manager on the Managed... ...of highly scalable, fault-tolerant infrastructure... ....This is an on-site role based in San Francisco... ...with CPU & GPU performance, inference frameworks, or LLM...WebsiteTemporary workWork at office
$314.8k - $359.3k
...Senior Distinguished AI Engineer At Capital One... ...in technology infrastructure and world-class talent... ...product experiences and scalable, high-performance AI... ...large language model inference, similarity search, guardrails... ...through this site. Capital One...WebsiteFull timePart timeLocal area$150k - $200k
...small, fast-moving AI infrastructure company building... ...including inference pipelines, distributed... ...orchestration, and GPU scheduling. Our... ...As a Platform Engineer, you'll work shoulder... ...high-performance, scalable cluster storage.... ...a full-time, on-site role based in San...WebsiteFull timeVisa sponsorship$286.2k - $326.7k
...Overview Senior Director, AI Engineering -Agentic AI Platform(... ...in technology infrastructure and world-class talent... ...experiences and scalable, high-performance AI... ...large language model inference, similarity search, guardrails... ...through this site. Capital One Financial...WebsiteFull timePart timeLocal areaRemote work- ...who are building AI systems to power magical... ...of researchers, engineers, designers, and... ...high-performance, scalable and reliable machine... ...production infrastructure at a large scale... ...with Kubernetes, and GPU workloads on those... ...and throughput of inference. ~ Strong understanding...Full timeWork experience placementWork at officeRemote workFlexible hours
- ...monitor, and scale autonomous AI agents with full visibility... ...for a Senior Applied AI Engineer to join our on-site Research & Intelligence... ...translating proven techniques into scalable, reliable components within... ..., PEFT, or cost-aware inference strategies. Experience working...WebsiteFull time
$250k
...a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for... ..., and inference at scale. The company... ...for a Senior / Staff Site Reliability Engineer to support and scale... ...infrastructure growth and scalability. Don’t miss out...WebsiteFull timeRemote work- Nscale seeks a Senior Infrastructure Support Engineer to own the health of GPU fleets and high‑performance fabrics. You will operate across GPU hardware, Linux... ...automation while mentoring mid‑level engineers. Travel to sites may be required, with a remote‑first team structure...WebsiteRemote work
$188k - $275k
...The Essential Cloud for AI™. Built for pioneers... ...CoreWeave combines superior infrastructure performance with deep... ...of the team: The Inference team is responsible... ...looking for an Applied AI Engineer to help us understand,... ...throughput, batching, GPU utilization,...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...client is a well-funded AI startup building production-grade ML infrastructure used by enterprise... ...for a Senior AI/ML Engineer to own model... ...evaluation systems, and inference serving at scale. Full-time, on-site in San Francisco.... ...distributed training, GPU optimization, or...WebsiteFull time
- ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor,... ...applied AI research, flexible infrastructure, and seamless developer... ...help build the platform engineers turn to to ship AI... ...foundational engineers to lead our GPU Networking efforts,...Full timeFlexible hours
$180k - $225k
As a Software Engineer on the ML Infrastructure team, you will design and build platforms for scalable, reliable, and efficient serving of LLMs.... ...TensorRT-LLM, or text-generation-inference.Compensation packages at... ...mission is to develop reliable AI systems for the world's most...Full time$310k - $400k
...picture and our vision at Postman.The OpportunityAs the Head of AI Platform Engineering at Postman, you will lead the alignment of AI development... ...for AI safety, ethical AI development, and building scalable, user-centric platforms.Preferred:Experience working in or...WebsiteWork at officeFlexible hours3 days per week$200k - $230k
...technology.Job DescriptionDirector, AI Platform... ...seeking a Director of AI Platform Engineering to lead the design, development... ...engineers, architect critical infrastructure, and drive the strategy for... ...technologies (new LLM providers, inference optimization, agentic...WebsiteOngoing contractFull timeCasual workWork at officeFlexible hours- .... Ltd. is seeking an Applied Research Engineer to design scalable pipelines for large-scale video understanding. You will work on multimodal AI applications, including CV, audio, and... ...production-ready systems. You will optimize inference performance, leverage foundation...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Infrastructure Engineer: Scalable GPU Inference, On-Site. Be the first to apply!
- ai engineer San Francisco, CA
- ai research engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- ai prompt engineer San Francisco, CA
- ai engineer remote San Francisco, CA
- ai developer San Francisco, CA
- principal infrastructure engineer San Francisco, CA
- remote infrastructure engineer San Francisco, CA



