Software Engineer - Voice AI (Inference Runtime)
Baseten
ABOUT BASETEN
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $300M Series E , backed by investors including BOND, IVP, Spark Capital, Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products.
THE ROLE:
Voice is becoming the internet’s next interface, but a production-grade Voice AI system is "hard to build" . You’ll join a small founding team of Baseten Voice AI, focused on bringing state-of-the-art open source models into production for Voice AI customers across productivity, customer service, clinical conversation, creator tools, education, and more. You’ll make a meaningful impact on people’s daily lives and help reshape these industries.
This is a high-impact, high-ownership role. You will be the primary owner of Baseten Voice AI - our in-house inference stack to power Voice AI models - from product roadmap through engineering implementation. You’ll partner closely with Forward Deployed Engineers, Model Performance Engineers, and sister engineering teams to push the boundaries of Voice AI.
EXAMPLE INITIATIVES:
Develop world-class model serving stack for state-of-the-art open-source voice models - reduce end-to-end and tail latency (p95/p99), increase throughput, and improve GPU efficiency via profiling, runtime tuning, and server-level optimizations.
Build large-scale, real-time infrastructure for multi-model voice agents - orchestrate STT, TTS, and agent components with streaming I/O to meet customer SLOs.
Design tight training and inference iteration loops for voice model customization - enable fast evaluation, safe rollout, and rapid experimentation for custom voice model development.
Past projects:
RESPONSIBILITIES:
Own and lead Voice AI product areas end-to-end - from architecture and system design through implementation, rollout, and long-term production operations.
Design, build, and operate real-time, large-scale, high-performance model serving systems for STT, TTS, and voice agent workloads for mission-critical customer deployments
Drive cross-team collaboration with sister engineering teams to solve full-stack technical problems, align on priorities, and coordinate end-to-end delivery across the product surface area
Mentor teammates through code reviews, design docs, and technical leadership.
REQUIREMENTS:
Bachelor's degree or higher in Computer Science or related field
Proven track record owning production-grade real-time, large-scale systems where tail latency (p99) matters.
Proficient coding abilities in one or more popular programming or scripting languages; Python proficiency is a plus.
Good taste in product, particularly developer-oriented tools
Interest in ML/AI infrastructure and willingness to learn
Strong collaboration and communication skills
Comfortable using AI coding assistants (e.g., Claude Code, Codex, Cursor) as a daily productivity multiplier — as an AI-native company, we see this as a must-have skill.
NICE TO HAVE:
Experience implementing pipeline-level model runtime optimizations such as dynamic batching, async scheduling, or decode-side throughput improvements.
Experience building developer platforms: SDKs, CLIs, APIs, and self-serve workflows for ML or infrastructure products.
Experience with containerization and orchestration technologies (Docker, Kubernetes), service meshes, or distributed scheduling.
Familiarity with speech/audio ML models (STT, TTS, speech-to-speech)
Familiarity with model-serving runtimes (vLLM, TensorRT, ONNX).
Familiarity with systems-level performance profiling across host-device boundaries (e.g. PyTorch Profiler), diagnosing GPU utilization issues
Exposure to customer-facing engineering: pre-sales prototyping, technical discovery, or working directly with customers to ship solutions.
BENEFITS
Competitive compensation, including meaningful equity.
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
Paid parental leave
Fertility and family-building stipend through Carrot
Company-facilitated 401(k)
Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).
- ...re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems... ...About OpenAI OpenAI is an AI research and deployment company dedicated... ...value the many different perspectives, voices, and experiences that form the full...SuggestedFull time
- ...scale. As part of the inference team, you’ll be responsible... ...for a kernel-focused engineer to lead efforts in... ...internal GPU libraries and runtime tools. Work closely... ...OpenAI OpenAI is an AI research and deployment... ...different perspectives, voices, and experiences that form...SuggestedFull time
- ...access state-of-the-art AI models - unlocking new... ...-performance model inference and accelerating research... ...role, you’ll lead engineering efforts to ensure our... ...issues across hardware and software layers. Have strong... ...perspectives, voices, and experiences that...SuggestedFull time
- ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion... ...help build the platform engineers turn to to ship AI... ...team builds the distributed runtime that powers large-scale LLM... ...and ease of use. As a Software Engineer on the Inference...SuggestedFull timeFlexible hours
- ...About the Team OpenAI’s Inference team powers the... ...small, fast-moving team of engineers focused on delivering a... ...the boundaries of what AI can do. We’re expanding... ...We’re looking for a software engineer to help us serve... ...different perspectives, voices, and experiences that...SuggestedFull time
- ...the Team With Codex we’re building an AI software engineer. One that you can pair with, delegate to... ...teammate. About the Role Codex Runtime is the execution and infrastructure... ...value the many different perspectives, voices, and experiences that form the full spectrum...Full time
- ...About the Team Our Inference team brings OpenAI’s most capable research... ...access our state-of-the-art AI models, allowing them to do... ...About the Role We’re hiring engineers to scale and optimize OpenAI’s... ...many different perspectives, voices, and experiences that form the...Full time
- ...the Team Our team analyzes inference stack performance across the... ...Enjoy collaborating with engineering and research teams to improve... ...About OpenAI OpenAI is an AI research and deployment company... ...many different perspectives, voices, and experiences that form the...Full time
- ...scale. About the Role As a software engineer on the Scaling team, you’ll... ...designing high-performance runtimes, building custom kernels, contributing... ...About OpenAI OpenAI is an AI research and deployment... ...many different perspectives, voices, and experiences that form...Full timeWork at officeLocal areaRelocation package3 days per week
- ...About the Team Our Inference team brings OpenAI’s most... ...access our start-of-the-art AI models, allowing them... ...We are looking for an engineer who wants to take the... ...years of professional software engineering experience.... ...different perspectives, voices, and experiences that form...Full time
- ...for building and orchestrating multi-agent AI systems, powering 300M+ agent executions... ...Role You'll work on the enterprise runtime layer that turns CrewAI's open-source Crews... ...Looking For Strong Python backend/platform engineering experience, especially building...Full timeRemote work
$320k
...interpretable, and steerable AI systems. We want AI to be safe... ...group of committed researchers, engineers, policy experts, and business... ...Claude—real-time, bidirectional voice conversations that feel... ...real-time media, low-latency inference, and distributed systems—building...Full timeWork at officeVisa sponsorshipFlexible hours$320k
...interpretable, and steerable AI systems. We want AI to be safe... ...of committed researchers, engineers, policy experts, and business... ...Our mandate is to make inference deployment boring and unattended... ...continuous and unattended. As a Software Engineer on the Launch...Full timeWork at officeVisa sponsorshipFlexible hoursShift work- ...OpenAI builds the low-level software that accelerates our most ambitious AI research. We work at... ...optimizations, and runtime improvements to make large-scale training and inference more efficient. Our work... ...different perspectives, voices, and experiences that form...Full time
$300k
...interpretable, and steerable AI systems. We want AI to be safe... ...of committed researchers, engineers, policy experts, and business... ...About the role Our Inference team is responsible for building... ...you: Have significant software engineering experience, particularly...Full timeWork at officeWorldwideVisa sponsorshipFlexible hours- ...underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for... ...APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and... ...Opportunity: We are seeking a Software Engineer to join Deepgram for Restaurants - a...Full timeHome officeFlexible hours
$160k - $225k
...building and running the world's best data and AI infrastructure platform so our customers... ...of Databricks' Mosaic AI mission. AI Runtime (AIR) is our managed platform for large‑... ...AI teams in the world. As a Senior Software Engineer for AI Runtime, you will play a critical...- Senior Software Engineer, Voice AI Nooks is an applied AI lab building the Agent Workspace for GTM. We design AI agents that operate across the full sales action set, from account strategy to prospect research and outreach. Our mission is to 10x every seller by automating...Work at office3 days per week
$220k
...Perplexity is looking for an engineer to join their team in San Francisco... ...on building and operating the inference engine, supporting new models,... ...a Rust-based serving runtime. The ideal candidate has 3+ years of experience in software engineering with a focus on ML...- ...Sciforium is an AI infrastructure company developing... ...-on support from AMD engineers the team is scaling... ...working across C++, Python, runtime execution, and... ...Monitoring, and distributed inference features. Collaborate... ...experience) ~3+ years of software engineering experience,...Full timeWork at officeFlexible hours
- ...also manages large-scale inference and platform... ...growth. Within Applied Engineering, the Ads Monetization... ...years of professional software engineering experience... ...OpenAI OpenAI is an AI research and deployment... ...different perspectives, voices, and experiences that...Full time
$300k
...interpretable, and steerable AI systems. We want AI to be safe... ...of committed researchers, engineers, policy experts, and business... ...About the Role The Cloud Inference team scales and optimizes Claude... ...You: Have significant software engineering experience, with...Full timeWork at officeVisa sponsorshipFlexible hours- ...About the Team We’re hiring software engineers to make OpenAI’s networking... ...support OpenAI’s training and inference infrastructure at frontier... ...About OpenAI OpenAI is an AI research and deployment... ...many different perspectives, voices, and experiences that form the...Full time
- ...tuning. We also operate inference infrastructure at... ...growth. The Fraud Engineering team works within our... ...We are looking for a software engineer with anti fraud... ...OpenAI OpenAI is an AI research and deployment... ...different perspectives, voices, and experiences that...Full timeImmediate start
- ...unblocked. You will work across engineering and infrastructure problems... ...and orchestration issues to inference bottlenecks, numerical problems... ...About OpenAI OpenAI is an AI research and deployment... ...many different perspectives, voices, and experiences that form the...Full time
- ...About the Team The Software Engineering team is responsible for designing and... .... We work across device runtimes, mobile/embedded apps, and backend... ...cloud integrations, applied AI tools, mobile or embedded... ...many different perspectives, voices, and experiences that form...Full timeWork at officeRelocation package
- ...as a service, and new stateful runtime environments for agentic... ...We’re looking for a backend engineer who can quickly understand OpenAI... ...built developer tools, especially AI-powered tools, communicate clearly... ...many different perspectives, voices, and experiences that form the...Full timeInternship
- ...also manages large-scale inference infrastructure. With... .... Within Applied Engineering, the Financial Engineering... ...years of professional software engineering experience... ...OpenAI is an AI research and deployment... ...different perspectives, voices, and experiences that...Full time
- ...About the Team The Applied AI team safely brings OpenAI's technology... ...fine-tuning. We also operate inference infrastructure at scale.... .... About the Role The Engineering Acceleration team designs,... ...many different perspectives, voices, and experiences that form the...Full timeImmediate startRelocation package
- ...frontier model training and inference workloads to run... ...and product teams. Engineers on this team own problems... ...industry experience in software or infrastructure engineering... ...OpenAI OpenAI is an AI research and deployment... ...perspectives, voices, and experiences that form...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer - Voice AI (Inference Runtime). Be the first to apply!
- software engineer full time San Francisco, CA
- software system engineer San Francisco, CA
- consulting software engineer San Francisco, CA
- software engineer travel San Francisco, CA
- real time software engineer San Francisco, CA
- network software engineer San Francisco, CA
- senior software engineer remote San Francisco, CA
- entry level software engineer remote San Francisco, CA
- software engineer intern San Francisco, CA
- new grad software engineer San Francisco, CA


