Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Inference Engineer

Acceler8 Talent

Member of Technical Staff, Inference Performance 200k base + equity I’m working with a AI infrastructure startup building a high-performance inference cloud for open models. The team optimizes the full path from model and serving engine through kernels, accelerators, and production infrastructure. Its platform is already processing trillions of tokens per month, and the company is expanding due to customer demand growing faster than its current capacity. This role will focus on making inference faster, more reliable, and more cost-efficient across different models and hardware architectures. You’ll work on problems such as: Profiling and optimizing inference latency, throughput, and memory usage Improving serving engines through batching, caching, quantization, and speculative decoding Developing and tuning CUDA, HIP, or Triton kernels Operating heterogeneous accelerator clusters Benchmarking models and hardware under realistic production workloads Building observability and reliability into the serving platform Debugging performance across models, runtimes, kernels, networking, and hardware. Looking for engineers who have strong evidence in one or more of: High-performance AI inference or model-serving systems GPU kernel development using CUDA, HIP, or Triton Quantization, speculative decoding, batching, or KV-cache optimization PyTorch, vLLM, SGLang, TensorRT-LLM, or similar frameworks GPU, accelerator, HPC, or distributed-compute infrastructure Low-level performance profiling and systems optimization Operating latency-sensitive systems in production. Strong candidates will be able to explain what they personally optimized, how they measured it, and the production impact it created. This is an intense, highly hands-on environment with direct founder access and broad ownership. It will suit engineers who want to move across models, kernels, hardware, and infrastructure rather than remain within a narrowly defined area. #J-18808-Ljbffr Acceler8 Talent

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Inference Engineer in San Francisco, CA vacancy
  • $175k - $250k

     ...compensation Excellent benefits (healthcare, vision, dental) Job Details We're looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across a variety of compute architectures. The role is... 
    Suggested
    Local area

    Jobot

    San Francisco, CA
    3 days ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-time local intelligence. You will primarily work on uzu, our inference engine, and focus on supporting new modalities and a variety of features. The ideal candidates... 
    Suggested
    Local area

    Mirai Labs

    San Francisco, CA
    9 hours ago
  •  ...platform as workloads and customer demand grow Most infrastructure engineers joining an AI cloud inherit a platform. Here, you’ll build it...  ...early-stage AI infrastructure company building a new kind of inference cloud, turning a heterogeneous fleet of AI compute into... 
    Suggested

    TecHire

    San Francisco, CA
    1 day ago
  •  ...Francisco, on‑site ABOUT THE ROLE You build and operate the inference systems that serve our models in production. The work spans serving...  ...that come with running real workloads. This is an engineering role, not a research role. You'll measure, profile, debug, and... 
    Suggested

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  • Distributed LLM Inference Engineer At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray, a popular open-source project that's creating an ecosystem of libraries... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    19 hours ago
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 

    inference.net

    San Francisco, CA
    1 day ago
  • About us General Compute is the neocloud for alternative chips. Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras...  ...line. Shape the roadmap, not just the backlog. As a founding engineer, you'll help decide what we build next in serving — multi-... 

    General Compute

    San Francisco, CA
    3 days ago
  • $160k - $252.5k

     ...peaker plants a thing of the past. Power Electronics Controls Engineer As a Power Electronics Controls Engineer, you will lead...  ...personal records, professional or employment information, and inferences drawn from your PI. We collect your PI for our purposes,... 
    Full time
    Shift work

    Redwood Materials

    San Francisco, CA
    2 days ago
  • $187.5k - $247.5k

     ...all from batteries we already have. Staff Mechanical Design Engineer, EPC Redwood Materials is hiring for a Staff...  ...personal records, professional or employment information, and inferences drawn from your PI. We collect your PI for our purposes, including... 
    Full time
    Work experience placement

    Redwood Materials

    San Francisco, CA
    1 day ago
  • $36.06 - $40.87 per hour

    Technical Support Field Engineer - San Francisco, CA Dentsply Sirona is the world’s largest manufacturer of professional dental products...  ..., certifications, transcripts and languages spoken); and inferences from personal information collected (e.g., a profile reflecting... 
    Hourly pay
    Work experience placement
    Work at office
    Remote work
    Worldwide
    Flexible hours
    Night shift

    Wellspect HealthCare

    San Francisco, CA
    1 day ago
  • $200.8k - $251k

     ...a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251... 
    Full time

    Scale AI

    San Francisco, CA
    19 hours ago
  • $147.5k - $235k

     ...months rather than years using proprietary energy architecture engineered entirely in-house. The company is profitable, the market it...  ...personal records, professional or employment information, and inferences drawn from your PI. We collect your PI for our purposes, including... 
    Full time

    Redwood Materials

    San Francisco, CA
    1 day ago
  •  ...Senior Systems Engineer San Francisco, California Onsite or Remote At Evidently, we are raising the quality of care for every...  ...work, security and compliance posture, database performance, AI inference infrastructure: you can cover these areas when they touch systems... 
    Remote work
    Work from home

    Evidently

    San Francisco, CA
    3 days ago
  • $75k - $85k

     ...yr - $85,000.00/yr We are searching for a Construction Project Engineer to support project teams on high-end residential projects in San...  ...your chances of interviewing at Level Recruiting by 2x Inferred from the description for this job Medical insurance 401(k) Vision... 
    Full time
    Work at office

    Level Recruiting

    San Francisco, CA
    19 hours ago
  • $200k - $300k

     ...job poster from Acceler8 Talent Senior Neuro-Symbolic Systems Engineer - San Francisco, CA A company building AI systems that can interact...  ...planners, or similar abstractions Build update rules, inference mechanisms, and dynamic graph operations that support multi‑step... 
    Full time
    Immediate start

    Acceler8 Talent

    San Francisco, CA
    19 hours ago
  • $160k - $195k

     ...Syntiant Corp. is looking for an experienced Field Applications Engineer to be a key technical partner to customers, sales, product,...  ...Experience in reviewing and debugging existing Neural Network inference models to improve performance, latency, and power. ~ Working... 
    Temporary work
    Flexible hours

    Syntiant

    San Francisco, CA
    2 days ago
  • $170k - $250k

     ...operations. About the role We are seeking a talented algorithm engineering expert for our vehicle systems development in the Drive-by-...  ...Referrals increase your chances of interviewing at Gatik by 2x Inferred from the description for this job Medical insurance Vision insurance... 
    Odd job
    Full time
    Internship
    Work at office

    Gatik

    San Francisco, CA
    4 days ago
  • $182k - $242k

     ...CX organization aligns closely with the internal and customer engineering teams, offering valuable insights from the field and having the...  ...Artificial Intelligence/Machine Learning (AI/ML) training and inference workloads on technologies such as Slurm and Kubernetes. You... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    4 days ago
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 

    Baseten

    San Francisco, CA
    1 day ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 

    Anthropic

    San Francisco, CA
    9 hours ago
  •  ...explorers. We are motivated by the challenge of solving tough engineering problems, committed to creating sustainable change, and driven...  ...discrimination using radiometric imagery, optimized for real-time embedded inference via quantization, pruning, or equivalent techniques, with full... 
    Flexible hours

    etc.

    San Francisco, CA
    4 days ago
  • $140k - $200k

     ...more efficient world.   The Sr. Data Infrastructure / Quality Engineering role will play a crucial part in architecting, building, and...  ...deploying and optimizing high-performance GPU cloud inference services, with specific expertise utilizing the NVIDIA architecture... 
    Full time
    Work experience placement
    Local area

    Ouster

    San Francisco, CA
    a month ago
  • $130k - $160k

    Embedded Autonomy Engineer Huntington Beach, California, United States; San Francisco, California, United States About Mach Industries...  ..., and make perception, planning, decision making, and on-edge inference run in real time under real SWaP, thermal, and vibration... 
    Permanent employment
    Work experience placement
    Work at office
    Local area
    Night shift

    Mach Industries

    San Francisco, CA
    2 days ago
  • $180k - $237.5k

     ...time, all from batteries we already have. Embedded Software Engineer - Power Electronics We are at the precipice of a global energy...  ...personal records, professional or employment information, and inferences drawn from your PI. We collect your PI for our purposes, including... 
    Full time

    Redwood Materials

    San Francisco, CA
    2 days ago
  •  ...Capital One in San Francisco is seeking a Staff AI Engineer to build responsible AI systems, partnering with engineers, researchers,...  ...with Capital One. The role covers foundation model training, inference, guardrails, evaluation, and observability, using Open Source... 

    Capital One

    San Francisco, CA
    19 hours ago
  • $150k - $250k

     ...platform we’re building will adapt to experts across fields: engineering, research, operations, and beyond. About the Role We’re hiring...  ...increase your chances of interviewing at Stealth AI Startup by 2x Inferred from the description for this job Medical insurance Vision... 
    Full time
    Summer work
    Internship
    Remote work
    Flexible hours

    Stealth AI Startup

    San Francisco, CA
    4 days ago
  • $176k - $209k

    Software Engineer (Agentic Systems) Bay Area, US Dialpad is the AI platform for customer experience, built to resolve customer problems...  ...: Research and implement emerging agent frameworks, LLM inference optimization, advanced retrieval systems, and cutting-edge safety... 
    Work at office

    Dialpad

    San Francisco, CA
    9 hours ago
  • $188k - $275k

     ...in March 2025. Learn more at  What You'll Do: The Field Engineering organization at CoreWeave is dedicated to ensuring every...  ..., firmware—into reliable compute that customers can train and inference on at scale, spanning infrastructure engineering, provisioning... 
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    18 days ago
  • $250k

     ...generation GPU platform designed for AI training, experimentation, and inference at scale. The company is developing a fully featured AI cloud...  ...The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $185.5k - $232k

     ...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and...  ...containerized services, internal tools, data systems, ML pipelines, inference, and agentic software. Research, develop, and maintain core... 
    Work experience placement
    Work at office
    Local area
    Relocation
    3 days per week

    Formation Bio (Formerly TrailSpark)

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Inference Engineer. Be the first to apply!