Software Engineer, Inference - Performance Optimization
OpenAI
About the Team
Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches.
In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. In this role, you will:
Build and refine performance models that translate microbenchmark results into cost-to-serve estimates.
Analyze inference workloads end to end across applications, models, and fleet infrastructure.
Enhance tooling to identify bottlenecks across layers for latency and throughput.
Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference.
You might thrive in this role if you:
Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency.
Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling.
Have deep expertise with performance profiling, benchmarking, analysis, and optimization.
Enjoy collaborating with engineering and research teams to improve real production systems.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement .
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form . No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link .
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
$229.9k - $262.4k
...Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we... ...product experiences and scalable, high-performance AI infrastructure. At Capital One,... ...develop, test, deploy, and support AI software components including foundation...PerformanceFull timePart timeLocal area- ...About the Team OpenAI’s Inference team ensures that our most advanced... ...and at scale. We build and optimize the systems that power our... ...AMD GPUs - to increase performance, flexibility, and resiliency... ...About the Role We’re hiring engineers to scale and optimize OpenAI...PerformanceFull time
- ...Team We’re building high-performance infrastructure to serve OpenAI... ...scale. As part of the inference team, you’ll be responsible... ...tuning memory layouts, and optimizing model execution at the lowest... ...looking for a kernel-focused engineer to lead efforts in writing,...PerformanceFull time
- ...and more. We focus on high-performance model inference and accelerating research... ...Lead to drive the design, optimization, and scaling of our inference... ...In this role, you’ll lead engineering efforts to ensure our... ...issues across hardware and software layers. Have strong familiarity...PerformanceFull time
- ...About the Team OpenAI’s Inference team powers the... ...models are available, performant, and scalable in production... ..., fast-moving team of engineers focused on delivering... ...We’re looking for a software engineer to help us serve... .... You'll build and optimize the systems that let users...PerformanceFull time
- ...small, fast-growing team of engineers in San Francisco powering Fortune... ...-latency, high-throughput inference for OCR and multimodal... ...smart batching and caching Optimize kernels, tokenization, and model... ...with clear SLOs Own performance dashboards and capacity planning...PerformanceFull timeWork at officeVisa sponsorshipRelocation package
- ...re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the... ...compromising reliability or performance. This role sits at the intersection... ...support model launches, inference optimizations, cloud provider integrations, and...PerformanceFull time
- ...powers mission-critical inference for the world's most... ...build the platform engineers turn to to ship AI products... ...systems, model performance, infrastructure, and... ...ease of use. As a Software Engineer on the Inference... ...make new inference optimizations broadly available to...PerformanceFull timeFlexible hours
- ...Baseten powers mission-critical inference for the world's most dynamic... ...and help build the platform engineers turn to to ship AI products.... ...Deployed Engineers, Model Performance Engineers, and sister... ...runtime tuning, and server-level optimizations. Build large-scale, real-...PerformanceFull timeFlexible hours
$300k
...committed researchers, engineers, policy experts, and... ...the role Our Inference team is responsible for... ...scientists the high-performance inference... ...Have significant software engineering experience... ...systems LLM inference optimization, batching, and caching...PerformanceFull timeWork at officeWorldwideVisa sponsorshipFlexible hours- ...able to before. We focus on performant and efficient model inference, as well as accelerating... ...We are looking for an engineer who wants to take the world... ...capable AI models and optimize them for use in a high-volume... ...3 years of professional software engineering experience....PerformanceFull time
- ...Cohere is a team of researchers, engineers, designers, and more, who... ...energized by building high-performance, scalable and reliable... ...closely with many teams to deploy optimized NLP models to production in... ...influence latency and throughput of inference. ~ Strong understanding or...PerformanceFull timeWork experience placementWork at officeRemote workFlexible hours
$300k
...of committed researchers, engineers, policy experts, and... ...the Role The Cloud Inference team scales and optimizes Claude to serve the massive... ...LLMs meet rigorous safety, performance, and security standards.... ...Have significant software engineering experience, with...PerformanceFull timeWork at officeVisa sponsorshipFlexible hours- ...Big More Better). You will own optimizations on both the training and on-robot inference stacks. We are still in a regime... ...Implementing ML, hardware, and software changes that lead to step-function... ...hundreds of millions of users, engineered the foundations of autonomous driving...Full time
$320k
...group of committed researchers, engineers, policy experts, and... ...Our mandate is to make inference deployment boring and unattended... ...continuous and unattended. As a Software Engineer on the Launch... ...This is a resource-constrained optimization problem at its core: validation...Full timeWork at officeVisa sponsorshipFlexible hoursShift work- ...Workspace. What you'll do As a Software Engineer on our Site Reliability team at... ...clear visibility into system health and performance. Partner with product and platform... ...with LLM infrastructure — optimizing inference performance, managing fine-tuned models...PerformanceFull timeFlexible hours
$170k - $216k
...evaluate the Waymo Driver's software stack at a massive... ...of customers Software Engineers, Product, Data Science... ...Build and evolve ML inference infrastructure for simulations... ...frameworks, TPUs and optimizing models for serving.... ..., if the role can be performed remote, the specific...Full timeRemote work- ...-on support from AMD engineers the team is scaling rapidly... ...-on experience with performance engineering, learn... ...large AI models are optimized and deployed at scale... ..., and distributed inference features. Collaborate... ...) ~3+ years of software engineering experience...PerformanceFull timeWork at officeFlexible hours
$195k - $225k
...and expanding the core engineering team in SF. The... ...realtime collaboration, GPU inference at scale, a modern... ...Role As the Senior Software Engineer – Backend (... ...GPU workloads to optimizing GraphQL resolvers and... ...ownership over APIs, performance, and data integrity —...PerformanceFull time- ...team at OpenAI builds the low-level software that accelerates our most... ...hardware and software, developing high-performance kernels, distributed system optimizations, and runtime improvements to make large-scale training and inference more efficient. Our work enables...PerformanceFull time
- ...powered workforce management that optimizes both human and AI capacity,... ..., data pipelines, and inference servers to predict support contact... ...with ML packages and software: Experience using Python libraries... ...team. Passion for performance: A strong commitment to advancing...PerformanceFull time
- ...and unblocked. You will work across engineering and infrastructure problems as they emerge... ...scaling and orchestration issues to inference bottlenecks, numerical problems, and... ...inference stacks. - Background in performance optimization, scaling, or production-critical...PerformanceFull time
- ...boundaries of data, scaling laws, optimization techniques, model... ...Role We’re looking for a Software Engineer focused on building and scaling... ...direct impact on system performance, reliability, and scale.... ...Collaborate across Pretraining, Inference, and Product teams to...PerformanceFull time
$300k - $320k
...of committed researchers, engineers, policy experts, and... ...Anthropic is looking for backend software engineers to work across... .... You'll own the performance, reliability, and efficiency... ...'ll partner closely with inference and safeguards to optimize the full stack. API Capabilities...PerformanceFull timeWork at officeVisa sponsorshipFlexible hours$125k - $160k
...seeking a versatile Full Stack Software Engineer to join our engineering... ...: Build delightful, performant, and accessible user experiences... ..., libraries, and tools to optimize performance and developer experience... ..., or local model inference (Ollama). Experience in...PerformanceFull timeLocal areaVisa sponsorshipWork visaShift work- ...for the architectural and engineering backbone of OpenAI’s infrastructure... .... Our work spans system software, networking, platform... ...-level monitoring, and performance optimization. About the Role We’re... ...benchmarks, porting existing inference and training workloads to...PerformanceFull time
$152.5k - $287.5k
...batteries we already have. Software Engineer, ML/Computer Vision (... ...battery sorting, spanning ML inference, image acquisition, sensor... ...services Monitor model performance in production to catch regressions... ...edge deployment or model optimization techniques for inference (e...PerformanceHourly payFull timeImmediate startShift work- ...Baseten powers mission-critical inference for the world's most dynamic... ...and help build the platform engineers turn to to ship AI products.... ...code directly impacts the performance of state-of-the-art machine... ...powers modern AI workloads, optimizing every microsecond of...PerformanceFull timeFlexible hours
- ...About the role As a founding software engineer at Bronco AI, you will be... ...coverage with our AI Develop and optimize systems for data processing and model inference at scale Collaborate with... ...Experience with high-performance computing and optimization...PerformanceFull timeNight shift
- ...to join our elite team of engineers. As part of our team, you'll... ...and promoting exceptional performance. HOW YOU’LL HAVE IMPACT... ...TypeScript) components of our software stack. AI First: Adapt... ..., prompt engineering, and optimizing inference for latency and cost. Prior...PerformanceFull timeImmediate startRemote workFlexible hours3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, Inference - Performance Optimization. Be the first to apply!
- software engineer full time San Francisco, CA
- software system engineer San Francisco, CA
- consulting software engineer San Francisco, CA
- software engineer travel San Francisco, CA
- real time software engineer San Francisco, CA
- network software engineer San Francisco, CA
- senior software engineer remote San Francisco, CA
- entry level software engineer remote San Francisco, CA
- software engineer intern San Francisco, CA
- new grad software engineer San Francisco, CA



