Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer- BIS (Baseten Inference Stack)

Full-time

Baseten

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $300M Series E , backed by investors including BOND, IVP, Spark Capital, Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products.

THE ROLE

Baseten’s Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We operate at the intersection of distributed systems, model performance, infrastructure, and developer experience. We enable customers to deploy and operate cutting-edge LLM models with industry-leading performance, scalability, reliability, and ease of use.

As a Software Engineer on the Inference Stack team, you’ll work across the stack - from the developer experience customers use to deploy models, the libraries used for features like tool calling and reasoning, all the way down to the systems we use to orchestrate deployments in Kubernetes and route traffic efficiently.

This is an ideal role for engineers who enjoy owning systems in production, solving hard integration problems, and making complex infrastructure simple and reliable for users.

 

EXAMPLE INITIATIVES

Blog Posts

 

RESPONSIBILITIES

  • Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference

  • Work across the stack, from customer-facing features to low-level infrastructure components

  • Build platform capabilities related to routing, autoscaling, scheduling, observability, and runtime management

  • Improve the reliability, scalability, and usability of our inference stack

  • Collaborate closely with Model Performance engineers to make new inference optimizations broadly available to customers and easy to configure

  • Help define best practices around testing, release automation, benchmarking, and operational excellence

  • Debug complex production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads

  • Make thoughtful engineering tradeoffs balancing performance, reliability, operational simplicity, and developer experience

  • Own projects end-to-end: from architecture and implementation through deployment, monitoring, and iteration based on customer feedback

REQUIREMENTS

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field

  • Strong background in distributed systems, backend infrastructure, or platform engineering

  • Experience building and operating production systems where reliability, latency, and scale are first-class concerns

  • Strong sense of developer experience: you think about how systems are used, not just how they work

  • Motivated and willing to learn new languages, frameworks, and systems as needed

  • Ability to debug complex systems across multiple layers of the stack

  • Genuine interest in inference engineering. You don’t need to have hands on experience but are willing to learn

  • Excellent communication and collaboration skills

BONUS

  • Experience with Kubernetes, including concepts like operators and custom resources

  • Prior work on Dynamo, vLLM, SGLang, TensorRT-LLM, or similar inference frameworks

  • Experience with distributed scheduling, autoscaling, or service orchestration

  • Experience operating GPU workloads in production

  • Familiarity with observability tooling, CI/CD systems, or release automation

  • Experience contributing to open-source infrastructure or ML systems

BENEFITS

  • Competitive compensation, including meaningful equity.

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

  • Paid parental leave

  • Fertility and family-building stipend through Carrot

  • Company-facilitated 401(k)

  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Software Engineer- BIS (Baseten Inference Stack) in San Francisco, CA vacancy
  •  ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion...  ...us and help build the platform engineers turn to to ship AI products....  ...Voice AI - our in-house inference stack to power Voice AI models - from product... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $125k - $160k

     ...started. Role Overview We are seeking a versatile Full Stack Software Engineer to join our engineering team. Reporting to the Software...  ...-Augmented Generation) architectures, or local model inference (Ollama). Experience in automated testing at multiple levels... 
    Suggested
    Full time
    Local area
    Visa sponsorship
    Work visa
    Shift work

    Cala Health

    San Francisco, CA
    1 day ago
  • $150k - $180k

     ...Reach Capital , and JFF Ventures , and are now hiring a Full Stack Engineer to help build the product that institutions use to interact...  ...layer up to our data stack (Postgres + DuckDB) and model inference, and keep query and inference latency low enough that the product... 
    Suggested
    Full time
    Work at office
    Immediate start

    Straia

    San Francisco, CA
    1 day ago
  •  ...About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for...  ...large-scale deployments across a rapidly evolving inference stack. About the Role We’re looking for an autonomous,... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research and technology...  ...inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference...  ...GPU platforms. You’ll work across the stack - from low-level kernel performance to high... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $320k

     ...group of committed researchers, engineers, policy experts, and business...  ...Our mandate is to make inference deployment boring and unattended...  ...continuous and unattended. As a Software Engineer on the Launch...  ...~ Comfort working across the stack — from backend services and databases... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks...  ...benchmarking, analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems.... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...FriendliAI FriendliAI is the fastest inference cloud for agents, built to run frontier...  .... With our world-class inference stack, we are building the platform teams can...  ...the Role We're seeking a Full-Stack Software Engineer to design, build, and scale our web platform... 
    Flexible hours

    FriendliAI Corp

    San Francisco, CA
    5 days ago
  •  ...Series A , led by Felicis. About the role We are hiring Software Engineers to join our team. This is an opportunity to join us in-person...  ...solving business problems across all layers of the software stack. At Chalk, we build reliable data processing systems that can... 
    Full time
    Work at office
    Flexible hours

    Chalk

    San Francisco, CA
    1 day ago
  • $160k - $220k

     ...the world, Via is recognized as the leading transportation technology and service provider globally. As a  Senior   Full-Stack Software Engineer on the Remix engineering team, you’ll build software cities rely on to design and improve public transportation systems... 
    Full time
    2 days per week
    3 days per week

    Via

    San Francisco, CA
    1 day ago
  • $145k - $170k

     ...Francisco, CA Department: Product + Engineering Reports to: Director of Engineering...  ...and implement features across the full stack — from database schema to API layer to user...  ...You ~6+ years of experience as a software engineer, with strong full stack capabilities... 
    Full time
    Temporary work
    Work at office
    Remote work
    Work visa

    Upmetrics

    San Francisco, CA
    1 day ago
  •  ...re a team of passionate, mission-driven engineers from companies like Tesla, Amazon, SpaceX...  ...environments Work across the full stack, from React frontends to Node and Python...  ...to understand their workflows and build software that makes their work easier and faster... 
    Full time

    Medra

    San Francisco, CA
    1 day ago
  •  ...based I.T staffing and professional services company specializing in Web, Cloud & Mobility staffing solutions. Be it core Java, full-stack Java, Web/UI designers, Big Data or Cloud or Mobility developers/architects, we have them all.  Job Description 5+ years of... 
    Full time

    Jobsbridge

    San Francisco, CA
    1 day ago
  •  ...infrastructure. We work closely with model researchers, mobile engineers, frontend engineers, and platform teams to build intuitive experiences...  .... About the Role We are looking for an experienced Full Stack Engineer to join the Image Generation team and help shape the... 
    Full time
    Worldwide

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...0 companies, and startups. We’ve raised $16M from top VCs and were YC W25. About the role We’re looking for a Full-Stack Software Engineer, Reinforcement Learning to build the product surfaces, backend systems, and internal tools that power HUD’s RL data engine.... 
    Full time
    Work at office
    Remote work
    Relocation
    Visa sponsorship

    Hud

    San Francisco, CA
    1 day ago
  • $196k - $220.5k

     ...mobile. Collaborating with the other engineers on your team to write, review, and ship...  ...continually raise the quality bar of the software we write. What you should have: ~...  ...with at least a couple parts of our tech stack: Python, Typescript/React, Elixir, Rust... 
    Full time

    Discord

    San Francisco, CA
    1 day ago
  • $163k - $246.5k

     ...organizations, including Vanta, Lyft, and Dropbox. Learn more at semgrep.dev. About the role As a Fullstack Engineer, you’ll work across the stack to design, build and maintain a fast and reliable user experience for our customers. You’ll collaborate closely... 
    Full time
    Currently hiring
    Local area
    Remote work
    Weekend work
    3 days per week

    Semgrep, Inc.

    San Francisco, CA
    1 day ago
  • $196k - $220k

     ...during, and after playing games. We're looking for a Senior Software Engineer to join the Growth team at Discord. Our team owns how new...  ...that turns them into engaged members. You'll work across the stack to build the systems that acquire and activate users at scale... 
    Full time

    Discord

    San Francisco, CA
    1 day ago
  • $196k - $220.5k

     ...for a highly technical, creative, hands-on, and impact-focused Software Engineer to join our growing Ads team. Our team is revolutionizing...  ...industry technologies What you should have ~ Full-stack experience with hands on experience with Typescript, React,... 
    Full time
    Relocation
    Relocation package

    Discord

    San Francisco, CA
    1 day ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research...  ...We are looking for an engineer who wants to take the world's...  ...efficiency of our model inference stack. Build tools to give us...  ...least 5 years of professional software engineering experience. Have... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $170k - $216k

     ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve complex...  ...for a broad range of customers Software Engineers, Product, Data Science, System...  ...You will: Build and evolve ML inference infrastructure for simulations. Be responsible... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $255k - $405k

     ...the TeamThe Coding team is reimagining how software is built in the AI era. We build tools and workflows that help software engineers work faster, tackle more ambitious projects...  ...the world.About the RoleWe’re hiring a Full Stack Software Engineer to help invent the next... 
    Work at office
    Local area
    Flexible hours

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...in and out of the office. We are seeking a highly motivated Software Engineer (SWE) with experience in building robust and scalable web systems...  ...commerce functionalities.Debug and resolve issues across the stack, ensuring high availability and reliability for production... 
    Work at office
    Local area
    Remote work

    Oura

    San Francisco, CA
    4 days ago
  • $180k - $214k

    San Francisco, CaliforniaEngineering - AI Engineering /US Full-time Salaried /HybridSamba TV is a media intelligence company. We use consented...  ...operates at global scale.We're looking for a Senior Full Stack Engineer to join our engineering team in San Francisco. This is... 
    Full time
    Contract work

    Samba TV

    San Francisco, CA
    2 days ago
  •  ...integrations because doing voice well requires controlling the entire stack. The way people work is changing. IC work is over; you...  ...), C# (Windows) - ML : Custom speech recognition models, inference pipelines - Infra : Terraform, Stripe You don't need experience... 
    Full time

    Aqua Voice

    San Francisco, CA
    1 day ago
  •  ...every exchange, and finding novel ways to show the value of our technology. About the Role We’re looking for Full Stack software engineers with a product mindset to join the GTM Innovation team. As a product engineer on this team, you’ll help OpenAI meet the... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...Team The ChatGPT team operates at the intersection of research, engineering, product, and design to bring OpenAI’s technology to a global...  ...ChatGPT. About the Role We are seeking an experienced Full Stack Engineer to join the ChatGPT Growth Partnerships team and help... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $140k - $260k

     ...Substack is building a new economic engine for culture, giving the brightest, most interesting...  ...an interest in working across our tech stack, willingness to prototype, and product-...  ...Requirements At least 3+ years of software engineering experience. Independent and... 
    Full time

    Substack

    San Francisco, CA
    1 day ago
  •  ...Netic is the AI revenue engine for essential services who are the backbone of the American economy. With $43M in funding from Founders...  ..., and the impact is immediate and tangible. Netic's Full-Stack Software Engineers working on the product team build the features that... 
    Full time
    Summer work
    Immediate start
    Sleeping nights

    Netic

    San Francisco, CA
    1 day ago
  • $196k - $220.5k

     ...and game-related launches at Discord. Our engineering culture values collaboration, and this...  ...have 5+ years of experience as a fullstack software engineer. You have experience with...  ...comfortable switching between different technical stacks and learning new ones. You enjoy... 
    Full time

    Discord

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer- BIS (Baseten Inference Stack). Be the first to apply!