Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer- BIS (Baseten Inference Stack)

Full-time

Baseten

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $300M Series E , backed by investors including BOND, IVP, Spark Capital, Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products.

THE ROLE

Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use.

As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently.

This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users.

 

EXAMPLE INITIATIVES

Blog Posts

 

RESPONSIBILITIES

  • Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference

  • Work across the stack, from customer-facing features to low-level infrastructure components

  • Build platform capabilities related to routing, autoscaling, scheduling, observability, and runtime management

  • Improve the reliability, scalability, and usability of our inference stack

  • Collaborate closely with Model Performance engineers to make new inference optimizations broadly available to customers and easy to configure

  • Help define best practices around testing, release automation, benchmarking, and operational excellence

  • Debug complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads

  • Make thoughtful engineering tradeoffs balancing performance, reliability, operational simplicity, and developer experience

  • Own projects end-to-end: from architecture and implementation through deployment, monitoring, and iteration based on customer feedback

REQUIREMENTS

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field

  • Strong background in distributed systems, backend infrastructure, or platform engineering

  • Experience building and operating production systems where reliability, latency, and scale are first-class concerns

  • Strong sense of developer experience: you think about how systems are used, not just how they work

  • Motivated and willing to learn new languages, frameworks, and systems as needed

  • Ability to debug complex systems across multiple layers of the stack

  • Genuine interest in inference engineering. You don’t need to have hands on experience but are willing to learn

  • Excellent communication and collaboration skills

BONUS

  • Experience with Kubernetes, including concepts like operators and custom resources

  • Prior work on Dynamo, vLLM, SGLang, TensorRT-LLM, or similar inference frameworks

  • Experience with distributed scheduling, autoscaling, or service orchestration

  • Experience operating GPU workloads in production

  • Familiarity with observability tooling, CI/CD systems, or release automation

  • Experience contributing to open-source infrastructure or ML systems

BENEFITS

  • Competitive compensation, including meaningful equity.

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

  • Paid parental leave

  • Fertility and family-building stipend through Carrot

  • Company-facilitated 401(k)

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Software Engineer- BIS (Baseten Inference Stack) in San Francisco, CA vacancy
  •  ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion...  ...us and help build the platform engineers turn to to ship AI products....  ...Voice AI - our in-house inference stack to power Voice AI models - from product... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $120k - $180k

     ...yet, our team is tackling cutting-edge engineering challenges to bring revolutionary...  ...the role We are looking for a full-stack software enginee r to turn whiteboard ideas into...  ...features that showcase real-time sensing and inference in compelling, reliable ways.... 
    Suggested
    Full time
    Visa sponsorship

    Tacit Llc

    San Francisco, CA
    1 day ago
  • $125k - $160k

     ...started. Role Overview We are seeking a versatile Full Stack Software Engineer to join our engineering team. Reporting to the Software...  ...-Augmented Generation) architectures, or local model inference (Ollama). Experience in automated testing at multiple levels... 
    Suggested
    Full time
    Local area
    Visa sponsorship
    Work visa
    Shift work

    Cala Health

    San Francisco, CA
    1 day ago
  • $150k - $180k

     ...Reach Capital , and JFF Ventures , and are now hiring a Full Stack Engineer to help build the product that institutions use to interact...  ...layer up to our data stack (Postgres + DuckDB) and model inference, and keep query and inference latency low enough that the product... 
    Suggested
    Full time
    Work at office
    Immediate start

    Straia

    San Francisco, CA
    1 day ago
  •  ...frontier models at massive scale. As part of the inference team, you’ll be responsible for unlocking every...  ...model execution at the lowest levels of the stack. About the Role We are looking for a kernel-focused engineer to lead efforts in writing, porting, and optimizing... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for...  ...large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous,... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...focus on high-performance model inference and accelerating research...  ...systems. In this role, you’ll lead engineering efforts to ensure our largest...  ...the full inference stack - from model loading and memory...  ...performance issues across hardware and software layers. Have strong... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research and technology...  ...inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference...  ...GPU platforms. You’ll work across the stack - from low-level kernel performance to high... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $300k

     ...group of committed researchers, engineers, policy experts, and business...  ...About the role Our Inference team is responsible for building...  ...responsible for the entire stack from intelligent request...  ...you: Have significant software engineering experience, particularly... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  • $142.2k - $204.6k

    P-1284About This RoleAs a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks...  .... Your work will touch the full GenAI inference stack — from kernels and runtimes to orchestration and memory... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    5 days ago
  • $320k

     ...group of committed researchers, engineers, policy experts, and business...  ...Our mandate is to make inference deployment boring and unattended...  ...continuous and unattended. As a Software Engineer on the Launch...  ...~ Comfort working across the stack — from backend services and databases... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks...  ...benchmarking, analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems.... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...based I.T staffing and professional services company specializing in Web, Cloud & Mobility staffing solutions. Be it core Java, full-stack Java, Web/UI designers, Big Data or Cloud or Mobility developers/architects, we have them all.  Job Description 5+ years of... 
    Full time

    Jobsbridge

    San Francisco, CA
    1 day ago
  •  ...passionate about changing cancer care to join our growing team. Software engineer role We’re hiring a software engineer to build the...  ...bring our molecular insights to life. This role spans the full stack: on the front end, you'll craft AI-powered interfaces that... 
    Full time
    Work at office

    Valius Sciences

    San Francisco, CA
    1 day ago
  •  ...Series A , led by Felicis. About the role We are hiring Software Engineers to join our team. This is an opportunity to join us in-person...  ...solving business problems across all layers of the software stack. At Chalk, we build reliable data processing systems that can... 
    Full time
    Work at office
    Flexible hours

    Chalk

    San Francisco, CA
    1 day ago
  • $160k - $220k

     ...the world, Via is recognized as the leading transportation technology and service provider globally. As a  Senior   Full-Stack Software Engineer on the Remix engineering team, you’ll build software cities rely on to design and improve public transportation systems... 
    Full time
    2 days per week
    3 days per week

    Via

    San Francisco, CA
    1 day ago
  • $145k - $170k

     ...Francisco, CA Department: Product + Engineering Reports to: Director of Engineering...  ...and implement features across the full stack — from database schema to API layer to user...  ...You ~6+ years of experience as a software engineer, with strong full stack capabilities... 
    Full time
    Temporary work
    Work at office
    Remote work
    Work visa

    Upmetrics

    San Francisco, CA
    1 day ago
  •  ...re a team of passionate, mission-driven engineers from companies like Tesla, Amazon, SpaceX...  ...environments Work across the full stack, from React frontends to Node and Python...  ...to understand their workflows and build software that makes their work easier and faster... 
    Full time

    Medra

    San Francisco, CA
    1 day ago
  •  ...infrastructure. We work closely with model researchers, mobile engineers, frontend engineers, and platform teams to build intuitive experiences...  .... About the Role We are looking for an experienced Full Stack Engineer to join the Image Generation team and help shape the... 
    Full time
    Worldwide

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...of slowing down in 2025. We started 2024 with a small team of engineers and have doubled that in the past 12 months with plans to double...  ...What You'll Need: ~5+ years of work experience as a full stack engineer, ideally in a fast-growing technical SaaS company ~... 
    Remote job
    Full time
    Work experience placement
    Work at office
    Work from home
    Flexible hours
    3 days per week

    Miter

    San Francisco, CA
    1 day ago
  •  ...About Mariana Minerals Mariana Minerals is a software-first, vertically integrated minerals company on a mission to supply the...  ...Mariana Minerals is looking for an experienced Senior Full Stack Software Engineer to lead critical technical initiatives in building the... 
    Full time

    Mariana Minerals

    San Francisco, CA
    1 day ago
  • $196k - $220.5k

     ...mobile. Collaborating with the other engineers on your team to write, review, and ship...  ...continually raise the quality bar of the software we write. What you should have: ~...  ...with at least a couple parts of our tech stack: Python, Typescript/React, Elixir, Rust... 
    Full time

    Discord

    San Francisco, CA
    1 day ago
  • $163k - $246.5k

     ...organizations, including Vanta, Lyft, and Dropbox. Learn more at semgrep.dev . About the role As a Fullstack Engineer, you’ll work across the stack to design, build and maintain a fast and reliable user experience for our customers. You’ll collaborate closely... 
    Full time
    Currently hiring
    Local area
    Remote work
    Weekend work
    3 days per week

    Semgrep, Inc.

    San Francisco, CA
    1 day ago
  • $196k - $220k

     ...during, and after playing games. We're looking for a Senior Software Engineer to join the Growth team at Discord. Our team owns how new...  ...that turns them into engaged members. You'll work across the stack to build the systems that acquire and activate users at scale... 
    Full time

    Discord

    San Francisco, CA
    1 day ago
  • $196k - $220.5k

     ...for a highly technical, creative, hands-on, and impact-focused Software Engineer to join our growing Ads team. Our team is revolutionizing...  ...industry technologies What you should have ~ Full-stack experience with hands on experience with Typescript, React,... 
    Full time
    Relocation
    Relocation package

    Discord

    San Francisco, CA
    1 day ago
  •  ...performant and efficient model inference, as well as accelerating...  ...Role We are looking for an engineer who wants to take the world's...  ...least 3 years of professional software engineering experience. Have...  ...NVidia GPUs and the software stacks that optimize them (e.g. NCCL... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $170k - $216k

     ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve complex...  ...for a broad range of customers Software Engineers, Product, Data Science, System...  ...You will: Build and evolve ML inference infrastructure for simulations. Be responsible... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  •  ...in and out of the office. We are seeking a highly motivated Software Engineer (SWE) with experience in building robust and scalable web systems...  ...commerce functionalities.Debug and resolve issues across the stack, ensuring high availability and reliability for production... 
    Work at office
    Local area
    Remote work

    Oura

    San Francisco, CA
    4 days ago
  • $255k - $405k

     ...the TeamThe Coding team is reimagining how software is built in the AI era. We build tools and workflows that help software engineers work faster, tackle more ambitious projects...  ...the world.About the RoleWe’re hiring a Full Stack Software Engineer to help invent the next... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    5 days ago
  • $164k

    About the roleChime's Cards team is hiring a Senior Full-Stack Engineer to help build and scale our newest card products. You'll work across...  ..., you have5+ years of experience building production-grade software, with a track record of shipping reliable, user-facing... 
    Full time
    Work at office
    Local area
    Remote work

    Chime

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer- BIS (Baseten Inference Stack). Be the first to apply!