Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Infrastructure Engineer: Scalable GPU Inference, On-Site

Spellbrush

An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team, you will work closely with top AI researchers, driving the evolution of generative AI in the gaming industry. If you are passionate about anime aesthetics and enjoy a fast-paced environment, this opportunity will allow you to contribute to a creative movement that impacts millions of users worldwide. #J-18808-Ljbffr Spellbrush

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Engineer: Scalable GPU Inference, On-Site in San Francisco, CA vacancy
  •  ...Lead Software Engineer at JPMorgan Chase...  ...Sector, Infrastructure Platforms team,...  ...secure, stable, and scalable way. Drive significant...  ...optimized for AI/ML workloads....  ...training, and inference.Experience with...  ...of NVIDIA GPU infrastructure...  ...care coverage, on-site health and wellness... 
    Website
    For contractors

    JP Morgan Chase

    San Francisco, CA
    2 days ago
  •  ...building a self-serve compute platform that lets inference engineers run training jobs and inference services without worrying about GPU provisioning or cluster configuration. You...  ...that hides complexity behind a unified platform and scalable #J-18808-Ljbffr Neura Market
    Suggested

    Neura Market

    San Francisco, CA
    4 days ago
  •  ...’s leading generative AI studio behind niji・journey...  ...is looking for an AI Infrastructure Engineer to join us in building...  ...our next-generation inference architecture for...  ...excellent understanding of GPU’s handling large...  ...teams, and prefer on-site collaboration in either... 
    Website
    Work experience placement
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    more than 2 months ago
  • $269.1k - $307.2k

    Distinguished AI Engineer (Agentic AI Platform)...  ...investments in technology infrastructure and world-class talent...  ...product experiences and scalable, high-performance AI...  ...(e.g. LLM Inference, Similarity Search and...  ...available through this site. Capital One... 
    Website
    Full time
    Part time
    Work at office
    Local area

    Capital One Financial Corporation

    San Francisco, CA
    19 hours ago
  • United States Digital Space LLC is seeking an infrastructure leader to own a self-serve GPU compute platform for training and inference workloads. You will design and operate the...  ..., and a coherent platform roadmap for scalable AI workloads. #J-18808-Ljbffr United States... 
    Suggested

    United States Digital Space LLC

    San Francisco, CA
    3 days ago
  • $220k

    Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 

    Perplexity

    San Francisco, CA
    2 days ago
  •  ...Team OpenAI’s Inference team ensures that...  ...resiliency across our infrastructure. We are forming a...  ...We’re hiring engineers to scale and optimize...  ...infrastructure across emerging GPU platforms. You’ll...  ...performance, and scalability of model execution...  ...OpenAI is an AI research and... 
    Full time

    OpenAI

    San Francisco, CA
    19 hours ago
  •  ...Artificial Intelligence Institute in San Francisco Bay Area seeks an AI Infrastructure Engineer to advance an inference stack that spans multiple chips and models. You will build optimization kernels, a scalable library generator, and a benchmark-driven pipeline that... 
    Website

    Touring Capital

    San Francisco, CA
    3 days ago
  •  ...Associates Limited is seeking a Staff-level engineer to architect and evolve the AI infrastructure software stack for large-scale GPU workloads. You’ll drive orchestration,...  ...production environment. You’ll contribute to inference platforms, model serving, and high-... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    1 day ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (Gen AI Platform Services, Agentic...  ...in technology infrastructure and world-class talent...  ...product experiences and scalable, high-performance AI...  ...large language model inference, similarity search, guardrails...  ...through this site. Capital One... 
    Website
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    4 days ago
  • $151.8k - $265.35k

     ...organizations to create exceptional content effortlessly. The AI for Engineering team builds a scalable, production-grade AI platform that powers creativity...  ...orchestration, tool integration, memory systems, inference services, data flows, evaluation loops, and real-time... 
    Website
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    1 day ago
  •  ...Job Description The AI Infrastructure team at Zensors builds the engine that powers our visual sensing...  ...the training and inference of computer vision models...  ...stream to enable massive scalability of our SaaS product....  ...Deep understanding of GPU hardware performance ,... 

    Zensors

    San Francisco, CA
    more than 2 months ago
  • OpenAI is seeking an experienced software engineer to join the GPT Infrastructure team and help build an automated inference optimization platform that scales research prototypes into production‑ready solutions. You will design durable APIs and control‑plane services,... 

    OpenAI

    San Francisco, CA
    3 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience...  ...to help scale AI inference. You will design and optimize GPU kernels and tensor...  ...This role is based on-site in San Francisco or Los...  ...HPC techniques and scalable workloads. You will collaborate... 
    Website

    Vast.ai Inc.

    San Francisco, CA
    1 day ago
  • Hamilton Barnes Associates Limited is seeking an ambitious infrastructure engineer to help build and operate automated systems that bring GPU clusters from bare machines to customer-ready deployments. You will manage multi-tenant lifecycle, run Kubernetes and Postgres... 
    Remote job

    Hamilton Barnes Associates Limited

    San Francisco, CA
    9 hours ago
  •  ...vertically integrated AI infrastructure company built from...  ...the Role:As an Engineering Manager on the Managed...  ...of highly scalable, fault-tolerant infrastructure...  ....This is an on-site role based in San Francisco...  ...with CPU & GPU performance, inference frameworks, or LLM... 
    Website
    Temporary work
    Work at office

    Crusoe

    San Francisco, CA
    14 hours ago
  • $314.8k - $359.3k

     ...Senior Distinguished AI Engineer At Capital One...  ...in technology infrastructure and world-class talent...  ...product experiences and scalable, high-performance AI...  ...large language model inference, similarity search, guardrails...  ...through this site. Capital One... 
    Website
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    San Francisco, CA
    19 hours ago
  • $150k - $200k

     ...small, fast-moving AI infrastructure company building...  ...including inference pipelines, distributed...  ...orchestration, and GPU scheduling. Our...  ...As a Platform Engineer, you'll work shoulder...  ...high-performance, scalable cluster storage....  ...a full-time, on-site role based in San... 
    Website
    Full time
    Visa sponsorship

    Clera

    San Francisco, CA
    19 hours ago
  • $286.2k - $326.7k

     ...Overview Senior Director, AI Engineering -Agentic AI Platform(...  ...in technology infrastructure and world-class talent...  ...experiences and scalable, high-performance AI...  ...large language model inference, similarity search, guardrails...  ...through this site. Capital One Financial... 
    Website
    Full time
    Part time
    Local area
    Remote work

    Capital One

    San Francisco, CA
    24 days ago
  •  ...who are building AI systems to power magical...  ...of researchers, engineers, designers, and...  ...high-performance, scalable and reliable machine...  ...production infrastructure at a large scale...  ...with Kubernetes, and GPU workloads on those...  ...and throughput of inference. ~ Strong understanding... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    19 hours ago
  •  ...monitor, and scale autonomous AI agents with full visibility...  ...for a Senior Applied AI Engineer to join our on-site Research & Intelligence...  ...translating proven techniques into scalable, reliable components within...  ..., PEFT, or cost-aware inference strategies. Experience working... 
    Website
    Full time

    Alterion, Inc.

    San Francisco, CA
    19 hours ago
  • $250k

     ...a rapidly scaling AI cloud infrastructure provider building a next-generation GPU platform designed for...  ..., and inference at scale. The company...  ...for a Senior / Staff Site Reliability Engineer to support and scale...  ...infrastructure growth and scalability. Don’t miss out... 
    Website
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • Nscale seeks a Senior Infrastructure Support Engineer to own the health of GPU fleets and high‑performance fabrics. You will operate across GPU hardware, Linux...  ...automation while mentoring mid‑level engineers. Travel to sites may be required, with a remote‑first team structure... 
    Website
    Remote work

    Nscale

    San Francisco, CA
    3 days ago
  • $188k - $275k

     ...The Essential Cloud for AI™. Built for pioneers...  ...CoreWeave combines superior infrastructure performance with deep...  ...of the team: The Inference team is responsible...  ...looking for an Applied AI Engineer to help us understand,...  ...throughput, batching, GPU utilization,... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    13 days ago
  •  ...client is a well-funded AI startup building production-grade ML infrastructure used by enterprise...  ...for a Senior AI/ML Engineer to own model...  ...evaluation systems, and inference serving at scale. Full-time, on-site in San Francisco....  ...distributed training, GPU optimization, or... 
    Website
    Full time

    Clera

    San Francisco, CA
    19 hours ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor,...  ...applied AI research, flexible infrastructure, and seamless developer...  ...help build the platform engineers turn to to ship AI...  ...foundational engineers to lead our GPU Networking efforts,... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    19 hours ago
  • $180k - $225k

    As a Software Engineer on the ML Infrastructure team, you will design and build platforms for scalable, reliable, and efficient serving of LLMs....  ...TensorRT-LLM, or text-generation-inference.Compensation packages at...  ...mission is to develop reliable AI systems for the world's most... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $310k - $400k

     ...picture and our vision at Postman.The OpportunityAs the Head of AI Platform Engineering at Postman, you will lead the alignment of AI development...  ...for AI safety, ethical AI development, and building scalable, user-centric platforms.Preferred:Experience working in or... 
    Website
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    1 day ago
  • $200k - $230k

     ...technology.Job DescriptionDirector, AI Platform...  ...seeking a Director of AI Platform Engineering to lead the design, development...  ...engineers, architect critical infrastructure, and drive the strategy for...  ...technologies (new LLM providers, inference optimization, agentic... 
    Website
    Ongoing contract
    Full time
    Casual work
    Work at office
    Flexible hours

    SS&C Technologies

    San Francisco, CA
    3 days ago
  •  .... Ltd. is seeking an Applied Research Engineer to design scalable pipelines for large-scale video understanding. You will work on multimodal AI applications, including CV, audio, and...  ...production-ready systems. You will optimize inference performance, leverage foundation... 

    Guidant Solutions Pvt. Ltd.

    San Francisco, CA
    9 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Infrastructure Engineer: Scalable GPU Inference, On-Site. Be the first to apply!