Software Engineer, Inference - Performance Optimization
OpenAI
About the Team
Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches.
In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs. In this role, you will:
Build and refine performance models that translate microbenchmark results into cost-to-serve estimates.
Analyze inference workloads end to end across applications, models, and fleet infrastructure.
Enhance tooling to identify bottlenecks across layers for latency and throughput.
Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference.
You might thrive in this role if you:
Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency.
Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling.
Have deep expertise with performance profiling, benchmarking, analysis, and optimization.
Enjoy collaborating with engineering and research teams to improve real production systems.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement .
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form . No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link .
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
- LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You will work across diverse... ...‑LLM. Join a technically focused startup building cutting-edge AI software. #J-18808-Ljbffr LeoforcePerformance
- ...About the Team Our Inference team brings OpenAI’s most capable research and technology... ...been able to before. We focus on performant and efficient model inference, as... .... About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across...PerformanceFull time
- ...About the Team OpenAI’s Inference team powers the... ...models are available, performant, and scalable in production... ..., fast-moving team of engineers focused on delivering... ...We’re looking for a software engineer to help us serve... .... You'll build and optimize the systems that let users...PerformanceFull time
- ...re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the... ...compromising reliability or performance. This role sits at the intersection... ...support model launches, inference optimizations, cloud provider integrations, and...PerformanceFull time
- ...powers mission-critical inference for the world's most... ...build the platform engineers turn to to ship AI products... ...systems, model performance, infrastructure, and... ...ease of use. As a Software Engineer on the Inference... ...make new inference optimizations broadly available to...PerformanceFull timeFlexible hours
- ...Baseten powers mission-critical inference for the world's most dynamic... ...and help build the platform engineers turn to to ship AI products.... ...Deployed Engineers, Model Performance Engineers, and sister... ...runtime tuning, and server-level optimizations. Build large-scale, real-...PerformanceFull timeFlexible hours
- ...About the Team Our Inference team brings OpenAI’s most... ...before. We focus on performant and efficient model inference... ...We are looking for an engineer who wants to take the... ...capable AI models and optimize them for use in a high... ...years of professional software engineering experience...PerformanceFull time
$190.9k - $232.8k
P-1285About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture, development, and optimization of the inference engine that powers Databricks... ...background (6+ years or equivalent) in performance-critical systemsProven track record of...PerformanceLocal areaWorldwide$215k - $260k
...and be part of a high-performing team that believes in... ...means owning the inference stack end to end: profiling... ...go, bringing modern optimization techniques into real... ...with customer engineering teams to tailor deployments... ...Build and support the software and product features...PerformanceTemporary work$190k - $265k
...improve their business. Founded by engineers — and customer-obsessed — we... ...of AI.The Foundation Model Inference team is the backbone of... ...customers to serve, scale, and optimize frontier models with enterprise-grade reliability and performance. Our Foundation Model APIs provide...PerformanceLocal areaWorldwide- ...develops the ocean intelligence platform for government maritime applications. We are seeking a Senior Software Engineer to advance maritime route optimization, vessel performance modeling, and voyage analysis. This hands-on role involves refining optimization algorithms,...Performance
- ...intelligence needed to ensure a sustainable future. The Role We are looking for an experienced Senior Software Engineer to work on maritime route optimization, vessel performance modeling, and voyage analysis capabilities for U.S. Government and defense customers. This is a...Performance
- ...Big More Better). You will own optimizations on both the training and on-robot inference stacks. We are still in a regime... ...Implementing ML, hardware, and software changes that lead to step-function... ...hundreds of millions of users, engineered the foundations of autonomous driving...Full time
$137.1k - $201.6k
...satisfaction and retention—and we’re pioneering new ways to optimize and personalize these decisions at scale using causal inference and optimization. About the Role We're seeking a Machine Learning Engineer to lead the development of state-of-the-art ML systems...Hourly payWork at officeLocal areaRemote workFlexible hours$170k - $216k
...evaluate the Waymo Driver's software stack at a massive... ...of customers Software Engineers, Product, Data Science... ...Build and evolve ML inference infrastructure for simulations... ...frameworks, TPUs and optimizing models for serving.... ..., if the role can be performed remote, the specific...Full timeRemote work$195k - $225k
...and expanding the core engineering team in SF. The... ...realtime collaboration, GPU inference at scale, a modern... ...Role As the Senior Software Engineer – Backend (... ...GPU workloads to optimizing GraphQL resolvers and... ...ownership over APIs, performance, and data integrity —...PerformanceFull time$166k - $244k
...5 years of experience with software development in one or more programming... ...following: Machine Learning Optimization (e.g., quantization,... ...Computer Science, Computer Engineering, or a related technical field... ...and low‑level hardware performance, architecting the transition...PerformanceFull timeTemporary work$140k - $210k
...the first robot in the world capable of performing all core warehouse functions. We... ...with high agency. The Role As a Software Engineer on the Multi-Agent Systems team , you... ...coordination capabilities, and build algorithm optimization that enables system robustness,...PerformanceLocal areaFlexible hours- ...powered workforce management that optimizes both human and AI capacity,... ..., data pipelines, and inference servers to predict support contact... ...with ML packages and software: Experience using Python libraries... ...team. Passion for performance: A strong commitment to advancing...PerformanceFull time
- ...team at OpenAI builds the low-level software that accelerates our most... ...hardware and software, developing high-performance kernels, distributed system optimizations, and runtime improvements to make large-scale training and inference more efficient. Our work enables...PerformanceFull time
- ...and unblocked. You will work across engineering and infrastructure problems as they emerge... ...scaling and orchestration issues to inference bottlenecks, numerical problems, and... ...inference stacks. - Background in performance optimization, scaling, or production-critical...PerformanceFull time
$125k - $160k
...seeking a versatile Full Stack Software Engineer to join our engineering... ...: Build delightful, performant, and accessible user experiences... ..., libraries, and tools to optimize performance and developer experience... ..., or local model inference (Ollama). Experience in...PerformanceFull timeLocal areaVisa sponsorshipWork visaShift work- ...boundaries of data, scaling laws, optimization techniques, model... ...Role We’re looking for a Software Engineer focused on building and scaling... ...direct impact on system performance, reliability, and scale.... ...Collaborate across Pretraining, Inference, and Product teams to...PerformanceFull time
- Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest layers of the stack, optimize kernel performance, develop new request scheduling and parallelism strategies, and help us use a heterogeneous...Performance
- ...Baseten powers mission-critical inference for the world's most dynamic... ...and help build the platform engineers turn to to ship AI products.... ...code directly impacts the performance of state-of-the-art machine... ...powers modern AI workloads, optimizing every microsecond of...PerformanceFull timeFlexible hours
- ...Specter is creating a software-defined "control plane... ...become the perception engine for a company's physical... ...devices and cloud inference services, then close the... ...those models actually perform in the field and feeding... ..., inference optimization, and deployment to edge...PerformanceFull timeShift work
- ...platform purpose-built for performance marketers. We leverage massive... ...science to automate and optimize TV advertising to drive business... .... We are seeking a Software Engineer to build out our simulation... ...modeling systems Causal inference — uplift modeling,...PerformanceFull timeWork at officeRemote workRelocationRelocation package
- ...for the architectural and engineering backbone of OpenAI’s infrastructure... .... Our work spans system software, networking, platform... ...-level monitoring, and performance optimization. About the Role We’re... ...benchmarks, porting existing inference and training workloads to...PerformanceFull time
$130k - $400k
...candidates. Title of Role: Software Engineer – Marketplace (Backend... ...,000 – $400,000 + Equity + Performance Bonus Company... ...modeling, and performance optimization skills ~ Experience with... ...-powered products or model inference infrastructure Experience...PerformanceFull timeWork at officeRelocation package$220k - $275k
...games. We are looking for a Senior Software Engineer specializing in Machine Learning to... ...models that enhance ad relevance, optimize performance, and drive revenue. This is a unique... ...Experience working with real-time ML inference, A/B testing, and optimization frameworks...PerformanceFull timeShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, Inference - Performance Optimization. Be the first to apply!
- software engineer part time San Francisco, CA
- software developer intern San Francisco, CA
- software developer fintech San Francisco, CA
- software support engineer San Francisco, CA
- ngo software engineer San Francisco, CA
- software engineer - web development San Francisco, CA
- intel software engineer San Francisco, CA
- IT software engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- senior software engineer remote San Francisco, CA


