Software Engineer - Voice AI (Inference Runtime)
Baseten
ABOUT BASETEN
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $300M Series E , backed by investors including BOND, IVP, Spark Capital, Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products.
THE ROLE:
Voice is becoming the internet’s next interface, but a production-grade Voice AI system is "hard to build" . You’ll join a small founding team of Baseten Voice AI, focused on bringing state-of-the-art open source models into production for Voice AI customers across productivity, customer service, clinical conversation, creator tools, education, and more. You’ll make a meaningful impact on people’s daily lives and help reshape these industries.
This is a high-impact, high-ownership role. You will be the primary owner of Baseten Voice AI - our in-house inference stack to power Voice AI models - from product roadmap through engineering implementation. You’ll partner closely with Forward Deployed Engineers, Model Performance Engineers, and sister engineering teams to push the boundaries of Voice AI.
EXAMPLE INITIATIVES:
Develop world-class model serving stack for state-of-the-art open-source voice models - reduce end-to-end and tail latency (p95/p99), increase throughput, and improve GPU efficiency via profiling, runtime tuning, and server-level optimizations.
Build large-scale, real-time infrastructure for multi-model voice agents - orchestrate STT, TTS, and agent components with streaming I/O to meet customer SLOs.
Design tight training and inference iteration loops for voice model customization - enable fast evaluation, safe rollout, and rapid experimentation for custom voice model development.
Past projects:
RESPONSIBILITIES:
Own and lead Voice AI product areas end-to-end - from architecture and system design through implementation, rollout, and long-term production operations.
Design, build, and operate real-time, large-scale, high-performance model serving systems for STT, TTS, and voice agent workloads for mission-critical customer deployments
Drive cross-team collaboration with sister engineering teams to solve full-stack technical problems, align on priorities, and coordinate end-to-end delivery across the product surface area
Mentor teammates through code reviews, design docs, and technical leadership.
REQUIREMENTS:
Bachelor's degree or higher in Computer Science or related field
Proven track record owning production-grade real-time, large-scale systems where tail latency (p99) matters.
Proficient coding abilities in one or more popular programming or scripting languages; Python proficiency is a plus.
Good taste in product, particularly developer-oriented tools
Interest in ML/AI infrastructure and willingness to learn
Strong collaboration and communication skills
Comfortable using AI coding assistants (e.g., Claude Code, Codex, Cursor) as a daily productivity multiplier — as an AI-native company, we see this as a must-have skill.
NICE TO HAVE:
Experience implementing pipeline-level model runtime optimizations such as dynamic batching, async scheduling, or decode-side throughput improvements.
Experience building developer platforms: SDKs, CLIs, APIs, and self-serve workflows for ML or infrastructure products.
Experience with containerization and orchestration technologies (Docker, Kubernetes), service meshes, or distributed scheduling.
Familiarity with speech/audio ML models (STT, TTS, speech-to-speech)
Familiarity with model-serving runtimes (vLLM, TensorRT, ONNX).
Familiarity with systems-level performance profiling across host-device boundaries (e.g. PyTorch Profiler), diagnosing GPU utilization issues
Exposure to customer-facing engineering: pre-sales prototyping, technical discovery, or working directly with customers to ship solutions.
BENEFITS
Competitive compensation, including meaningful equity.
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
Paid parental leave
Fertility and family-building stipend through Carrot
Company-facilitated 401(k)
Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).
- ...re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems... ...About OpenAI OpenAI is an AI research and deployment company dedicated... ...value the many different perspectives, voices, and experiences that form the full...SuggestedFull time
- ...scale. As part of the inference team, you’ll be responsible... ...for a kernel-focused engineer to lead efforts in... ...internal GPU libraries and runtime tools. Work closely... ...OpenAI OpenAI is an AI research and deployment... ...different perspectives, voices, and experiences that form...SuggestedFull time
- ...the Team With Codex we’re building an AI software engineer. One that you can pair with, delegate to... ...teammate. About the Role Codex Runtime is the execution and infrastructure... ...value the many different perspectives, voices, and experiences that form the full spectrum...SuggestedFull time
- ...About the Team Our Inference team brings OpenAI’s most capable research... ...access our state-of-the-art AI models, allowing them to do... ...About the Role We’re hiring engineers to scale and optimize OpenAI’s... ...many different perspectives, voices, and experiences that form the...SuggestedFull time
- ...powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion... ...help build the platform engineers turn to to ship AI... ...team builds the distributed runtime that powers large-scale LLM... ...and ease of use. As a Software Engineer on the Inference...SuggestedFull timeFlexible hours
- ...About the Team OpenAI’s Inference team powers the... ...small, fast-moving team of engineers focused on delivering a... ...the boundaries of what AI can do. We’re expanding... ...We’re looking for a software engineer to help us serve... ...different perspectives, voices, and experiences that...Full time
- ...access state-of-the-art AI models - unlocking new... ...-performance model inference and accelerating research... ...role, you’ll lead engineering efforts to ensure our... ...issues across hardware and software layers. Have strong... ...perspectives, voices, and experiences that...Full time
- ...the Team Our team analyzes inference stack performance across the... ...Enjoy collaborating with engineering and research teams to improve... ...About OpenAI OpenAI is an AI research and deployment company... ...many different perspectives, voices, and experiences that form the...Full time
- ...scale. About the Role As a software engineer on the Scaling team, you’ll... ...designing high-performance runtimes, building custom kernels, contributing... ...About OpenAI OpenAI is an AI research and deployment... ...many different perspectives, voices, and experiences that form...Full timeWork at officeLocal areaRelocation package3 days per week
- ...About the Team Our Inference team brings OpenAI’s most... ...access our start-of-the-art AI models, allowing them... ...We are looking for an engineer who wants to take the... ...years of professional software engineering experience.... ...different perspectives, voices, and experiences that form...Full time
- ...for building and orchestrating multi-agent AI systems, powering 300M+ agent executions... ...Role You'll work on the enterprise runtime layer that turns CrewAI's open-source Crews... ...Looking For Strong Python backend/platform engineering experience, especially building...Full timeRemote work
$320k
...interpretable, and steerable AI systems. We want AI to be safe... ...group of committed researchers, engineers, policy experts, and business... ...Claude—real-time, bidirectional voice conversations that feel... ...real-time media, low-latency inference, and distributed systems—building...Full timeWork at officeVisa sponsorshipFlexible hours- ...OpenAI builds the low-level software that accelerates our most ambitious AI research. We work at... ...optimizations, and runtime improvements to make large-scale training and inference more efficient. Our work... ...different perspectives, voices, and experiences that form...Full time
$320k
...interpretable, and steerable AI systems. We want AI to be safe... ...of committed researchers, engineers, policy experts, and business... ...Our mandate is to make inference deployment boring and unattended... ...continuous and unattended. As a Software Engineer on the Launch...Full timeWork at officeVisa sponsorshipFlexible hoursShift work$300k
...interpretable, and steerable AI systems. We want AI to be safe... ...of committed researchers, engineers, policy experts, and business... ...About the role Our Inference team is responsible for building... ...you: Have significant software engineering experience, particularly...Full timeWork at officeWorldwideVisa sponsorshipFlexible hours$144k - $198k
...Golden, COSoftware - Onboard Software /Full time /On-siteWanna join... ...Earth orbit.The Deployment & Runtime Environments team is where software... ...someone scrappy — the kind of engineer who digs into an unfamiliar... ...observation, IoT connectivity, on-orbit AI, national security missions,...Full timeTemporary work$160k - $225k
...building and running the world's best data and AI infrastructure platform so our customers... ...of Databricks' Mosaic AI mission. AI Runtime (AIR) is our managed platform for large-... ...AI teams in the world.As a Senior Software Engineer for AI Runtime, you will play a critical...Local areaWorldwide$230k - $390k
...help businesses build better, more human customer experiences with AI. We are primarily an in-person company based in San Francisco,... ...is precise and matches each customer’s desired brand and style?Runtime: How do we adapt our agentic loop to accept real-time streams of...Full timeFlexible hours- ...underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for... ...APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and... ...Opportunity: We are seeking a Software Engineer to join Deepgram for Restaurants - a...Full timeHome officeFlexible hours
$190.9k - $232.8k
P-1285About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture, development... ...GenAI inference stack: kernels, runtimes, orchestration, memory, and integration... ...is the data and AI company. More than 10,000 organizations...Local areaWorldwide$293k - $385k
...turn advances in model inference and optimization... ...execution, compilers and runtimes, performance engineering, secure partner... ...RoleWe are seeking a software engineer to help build... ...with AI infrastructure, model... ...different perspectives, voices, and experiences that...Work at officeLocal areaRemote workFlexible hours$300k - $375k
A leading conversational AI platform in California seeks a Staff Software Engineer focused on Voice Agent. This role involves owning the architecture of the voice runtime and leading improvements in speech understanding and audio quality. The ideal candidate will have over...Work at office- ...Sciforium is an AI infrastructure company developing... ...-on support from AMD engineers the team is scaling... ...working across C++, Python, runtime execution, and... ...Monitoring, and distributed inference features. Collaborate... ...experience) ~3+ years of software engineering experience,...Full timeWork at officeFlexible hours
- Jetbridge is seeking a talented Software Engineer in San Francisco, California to drive the development of their real-time voice AI systems. The ideal candidate will spend the majority of their time coding and will be responsible for key product features ranging from backend...
- ...an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and runtime contexts. You will design the OpenAI-hosted control plane and partner-side software, focusing on reliable long-running workflows, reproducible...
$300k
...interpretable, and steerable AI systems. We want AI to be safe... ...of committed researchers, engineers, policy experts, and business... ...About the Role The Cloud Inference team scales and optimizes Claude... ...You: Have significant software engineering experience, with...Full timeWork at officeVisa sponsorshipFlexible hours- ...also manages large-scale inference and platform... ...growth. Within Applied Engineering, the Ads Monetization... ...years of professional software engineering experience... ...OpenAI OpenAI is an AI research and deployment... ...different perspectives, voices, and experiences that...Full time
- ...About the Team We’re hiring software engineers to make OpenAI’s networking... ...support OpenAI’s training and inference infrastructure at frontier... ...About OpenAI OpenAI is an AI research and deployment... ...many different perspectives, voices, and experiences that form the...Full time
- ...also manages large-scale inference infrastructure. With... .... Within Applied Engineering, the Financial Engineering... ...years of professional software engineering experience... ...OpenAI is an AI research and deployment... ...different perspectives, voices, and experiences that...Full time
- ...support requires human agents and AI in perfect balance, and... ...interfaces, data pipelines, and inference servers to predict support contact... ...with ML packages and software: Experience using Python libraries... ...advancing both statistical and runtime performance, ensuring reliable...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer - Voice AI (Inference Runtime). Be the first to apply!
- senior robotics software engineer San Francisco, CA
- software system engineer San Francisco, CA
- part time software developer San Francisco, CA
- fall software engineering internship San Francisco, CA
- security software engineer San Francisco, CA
- intel software engineer San Francisco, CA
- software developer fintech San Francisco, CA
- new graduate software engineer San Francisco, CA
- software development engineer aws San Francisco, CA
- information technology software engineer San Francisco, CA

