Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Engineer — Scalable AI Serving

Reducto, Inc.

A tech startup in AI model serving located in San Francisco is seeking a qualified candidate to architect scalable inference systems. The role focuses on optimizing model serving performance and integrating advanced techniques for AI deployment. Candidates should have strong expertise in Python and PyTorch, along with low-level systems knowledge. This in-person position offers a fast-paced work environment that is ideal for those eager to tackle complex challenges and shape the future of AI technology.#J-18808-Ljbffr

Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the LLM Inference Engineer — Scalable AI Serving in San Francisco, CA vacancy
  • $160k - $230k

    About the RoleAt Together.ai, we are building state...  ...enable efficient and scalable inference for large language...  ...Frameworks and Optimization Engineer to design, develop,...  ...shape the future of LLM inference infrastructure...  ...for high-performance serving.Apply CUDA graph... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  • Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open... 
    Suggested

    Anyscale

    San Francisco, CA
    4 days ago
  • $170k - $245k

     ...creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI...  ...tech stacks to accelerate the progress of AI applications out into the real world....  ...to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    6 hours ago
  •  ...Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in production....  ...that handles billions of inference requests, optimizing for latency, throughput...  ..., with responsibilities spanning scalable services, model serving, load... 
    Suggested

    Inception

    San Francisco, CA
    4 days ago
  • $300 per month

     ...vertically integrated AI infrastructure company...  ...Crusoe, our Production Engineering team ensures the reliability and scalability of Crusoe’s AI-...  ...services with a focus on serving and scaling LLM workloadsBuild automation...  ...distributed AI pipelines and inference servicesDefine,... 
    Suggested
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • $175k - $250k

    Global Inference Library Engineer $175000 - $250000 per year | San Francisco, CA |...  ...about us: We're a well-funded AI infrastructure startup...  ...who understands how modern LLM inference systems work under...  ...inference frameworks and model-serving infrastructure Hands-on experience... 
    Permanent employment
    Local area

    Australia-Employment

    San Francisco, CA
    4 days ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability...  ...in production-grade serving infrastructure, be fluent in Python... 

    MakerMaker.AI

    San Francisco, CA
    6 days ago
  • $300 per month

     ...vertically integrated AI infrastructure company...  ...Crusoe, our Production Engineering team ensures the reliability and scalability of Crusoe’s AI-...  ...services with a focus on serving and scaling LLM workloadsDefine, measure...  ...large-scale training and inference clustersAutomate... 
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • $91.1k - $179.5k

     ...success. We are hiring an AI Engineer to build and operate the data...  ...reliable, secure, and scalable AI solutions. This role is...  ...model training, real-time inference, and LLM applications using Claude-,...  ...datasets and feature engineering/serving for ML training and real-time... 
    Local area

    Deloitte

    San Francisco, CA
    1 day ago
  • Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more cost-efficient, and more reliable. You will extend orchestration frameworks (Kubernetes... 

    Inception

    San Francisco, CA
    4 days ago
  •  ...patients worldwide.We’re a team of engineers, clinicians, and innovators...  ...Systems GPU Engineer - AI & Robotics, you will be...  ...optimized, robust, validated and scalable medical device products and...  ...models to real-time onboard inference—while serving as a core contributor to team... 
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    2 days ago
  • Help build the inference stack behind the next generation...  ...architectures are served at scale. You’ll be building...  ...real-time multimodal AI capable of processing...  ...research and product engineering, designing the...  ...architectures. Design scalable, reliable distributed... 
    Work at office
    Relocation package

    techire.®

    San Francisco, CA
    3 days ago
  • $286.2k - $326.7k

     ...Senior Distinguished Engineer, AI Compute (Remote Eligible...  ...experiences and scalable, high-performance AI infrastructure...  ...to reimagine how we serve our customers and...  ...model training, model inference and feature generation...  ...workloads from LLM pre-training and reinforcement... 
    Full time
    Part time
    Local area
    Remote work

    Capital One

    San Francisco, CA
    10 days ago
  • $170k - $200k

     ...around the globe. Our AI-powered labor marketplace...  ...a Senior Software Engineer - Instawork Robotics to...  ...requirements into clear, scalable technical solutions, leveraging...  ...rolloutsExposure to LLM-based systems—includes...  ...AI-powered platform serves thousands of businesses... 
    Hourly pay
    Temporary work
    Local area
    Shift work

    Instawork

    San Francisco, CA
    4 days ago
  •  ...funded startup building AI-native solutions for...  ...looking for a Founding Engineer & CTO to be the technical...  ...our product vision into scalable, production-grade AI...  ...from data pipelines to LLM orchestration to client...  ...-tenant SaaS platforms serving enterprise clients. Familiarity... 
    Flexible hours

    NovusMinds AI

    San Francisco, CA
    6 days ago
  • Magic AI, Inc. is seeking a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows...  ...the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies... 

    Magic AI, Inc

    San Francisco, CA
    4 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging...  ...strong emphasis on HPC techniques and scalable workloads. You will collaborate with... 

    Vast.ai Inc.

    San Francisco, CA
    2 days ago
  •  ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion...  ...help build the platform engineers turn to to ship AI...  ...hardware. We believe that as LLM and multi-modal workloads...  ...compute for Disaggregated Serving, Wide Expert Parallelism... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $150k - $190k

     ...our 3D lidar technology will serve as the foundation of tomorrow...  ...Role SummaryAs a Staff Software Engineer in Test, you will be the...  ...Design, develop, and maintain a scalable and modular automated test...  ...and UDP.Experience leveraging AI and LLM tools to assist in analysis and... 
    Work experience placement
    Local area

    Ouster

    San Francisco, CA
    2 days ago
  •  ...Head of Internal Tools Engineering, Artificial Intelligence (AI) Required, Work From Home...  ...productivity. - Design scalable, secure architectures using...  ...unified technical vision. - Serve as the bridge between business...  ...Tools Engineering, LLM, SaaS, SDLC, Software as a... 
    Remote work
    Work from home

    Next Step Systems

    San Francisco, CA
    a month ago
  • $200.8k - $251k

     ...A leading AI technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large...  ...should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-... 
    Full time

    Scale AI

    San Francisco, CA
    11 hours ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere...  ...scale intelligence to serve humanity. We’re...  ...enterprises who are building AI systems to power magical...  ...fast, reliable, and scalable model training and build...  ...responsible for large-scale LLM training. Design... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    2 days ago
  • $206.3k - $388k

     ...looking for a Principal ML Engineer to architect and scale...  ...-ready data Scale up inference throughput across the...  ...compute scheduling ARCHITECT SCALABLE DATA INFRASTRUCTURE...  ...store, index, and serve billions of data points...  ...into impact, powered by AI and driven by human ingenuity... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    4 days ago
  • $221k - $247k

     ...reflects the people we serve.All full-time employees...  ...ourTotal Rewards philosophy.AI is a fundamental part...  ...ecosystem to improve scalability, efficiency, and...  ...EnablementPartner with EAIT engineers, administrators, AI architects...  ...knowledge of leading LLM platforms (e.g., OpenAI... 
    Full time
    Contract work
    Work at office
    Local area
    Shift work
    2 days per week
    3 days per week

    Gusto

    San Francisco, CA
    6 hours ago
  •  ...running the world’s best data and AI infrastructure platform so our...  ...their business. Founded by engineers — and customer obsessed — we...  ...metric views, and definitions that serve as the source of truth for...  ...applications or agentic workflows (LLM-powered apps and automations).... 
    Worldwide

    DataBricks

    San Francisco, CA
    4 days ago
  • $264.8k - $331k

     ...Learning Systems Research Engineer, Agent Post-training - Enterprise GenAI AI is becoming vitally...  ...research and resources that serve all of our enterprise...  ...optimize our training and inference framework. Post-train state...  ...: At least 1-3 years of LLM training in a production... 
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office

    SCALE INC

    San Francisco, CA
    3 days ago
  • $102k - $182.71k

     ...Senior Search Systems Engineer to build the intelligence...  ...of our marketing and AI visibility data. This role...  .... This role serves as the bridge between SEO...  ...teams to operationalize scalable datasets, transformation...  ...frameworks related to LLM analysis, AI agents, search... 
    Shift work

    Autodesk, Inc.

    San Francisco, CA
    4 days ago
  • $200k - $240k

     ...blockchain analytics and AI solutions to help law...  ...world for all. The AI Engineering Team is chartered with...  ...petabyte-scale pipelines, serve models with millisecond...  ...-edge tools in the LLM and agent space — including...  ...out a modular and scalable AI infrastructure stack... 
    Remote work
    Worldwide

    TRM Labs

    San Francisco, CA
    6 days ago
  • $105.4k - $124k

     ...the customers and businesses we serve to make better and smarter...  ...Intelligent Document Processing Engineer to join the Intelligent Document...  ...team delivers enterprise-wide AI, Machine Learning, Generative...  ...solutions, and deliver scalable applications that leverage Tungsten... 
    Full time
    Local area
    3 days per week

    US Bank

    San Francisco, CA
    2 days ago
  • $165k - $206k

     ...we are hiring the world’s best engineers, scientists, designers,...  ...Systems Engineer who leads with AI - not as a tool on the side, but...  ...to let AI do the heavy lifting.Serve as a cross-functional technical...  ...live supportLeverage Cursor with LLM-pair programming (Claude/... 

    Juul

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Engineer — Scalable AI Serving. Be the first to apply!