Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Inference - Performance Optimization

Full-time

OpenAI

About the Team
Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and drive faster, cheaper inference. We combine systems profiling, benchmarking, and analysis to understand where time and cost are spent, then turn that understanding into performance optimizations and models that project performance and capacity needs for future launches.

About the Role
In this role, you will model inference performance across application, model, and fleet layers with higher fidelity. You will build cost-to-serve estimates from microbenchmarks and create tools that help cross-functional teams reason about latency, capacity, utilization, and cost tradeoffs.

In this role, you will:
  • Build and refine performance models that translate microbenchmark results into cost-to-serve estimates.

  • Analyze inference workloads end to end across applications, models, and fleet infrastructure.

  • Enhance tooling to identify bottlenecks across layers for latency and throughput.

  • Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference.

You might thrive in this role if you:

  • Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency.

  • Are comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling.

  • Have deep expertise with performance profiling, benchmarking, analysis, and optimization.

  • Enjoy collaborating with engineering and research teams to improve real production systems.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement .

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form . No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link .

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer, Inference - Performance Optimization in San Francisco, CA vacancy
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we...  ...product experiences and scalable, high-performance AI infrastructure. At Capital One,...  ...develop, test, deploy, and support AI software components including foundation... 
    Performance
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    2 days ago
  •  ...About the Team OpenAI’s Inference team ensures that our most advanced...  ...and at scale. We build and optimize the systems that power our...  ...AMD GPUs - to increase performance, flexibility, and resiliency...  ...About the Role We’re hiring engineers to scale and optimize OpenAI... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...Team We’re building high-performance infrastructure to serve OpenAI...  ...scale. As part of the inference team, you’ll be responsible...  ...tuning memory layouts, and optimizing model execution at the lowest...  ...looking for a kernel-focused engineer to lead efforts in writing,... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...and more. We focus on high-performance model inference and accelerating research...  ...Lead to drive the design, optimization, and scaling of our inference...  ...In this role, you’ll lead engineering efforts to ensure our...  ...issues across hardware and software layers. Have strong familiarity... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...About the Team OpenAI’s Inference team powers the...  ...models are available, performant, and scalable in production...  ..., fast-moving team of engineers focused on delivering...  ...We’re looking for a software engineer to help us serve...  .... You'll build and optimize the systems that let users... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...small, fast-growing team of engineers in San Francisco powering Fortune...  ...-latency, high-throughput inference for OCR and multimodal...  ...smart batching and caching Optimize kernels, tokenization, and model...  ...with clear SLOs Own performance dashboards and capacity planning... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    Pulse

    San Francisco, CA
    8 hours ago
  •  ...re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the...  ...compromising reliability or performance. This role sits at the intersection...  ...support model launches, inference optimizations, cloud provider integrations, and... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...powers mission-critical inference for the world's most...  ...build the platform engineers turn to to ship AI products...  ...systems, model performance, infrastructure, and...  ...ease of use. As a Software Engineer on the Inference...  ...make new inference optimizations broadly available to... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    8 hours ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic...  ...and help build the platform engineers turn to to ship AI products....  ...Deployed Engineers, Model Performance Engineers, and sister...  ...runtime tuning, and server-level optimizations. Build large-scale, real-... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    8 hours ago
  • $300k

     ...committed researchers, engineers, policy experts, and...  ...the role Our Inference team is responsible for...  ...scientists the high-performance inference...  ...Have significant software engineering experience...  ...systems LLM inference optimization, batching, and caching... 
    Performance
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    8 hours ago
  •  ...able to before. We focus on performant and efficient model inference, as well as accelerating...  ...We are looking for an engineer who wants to take the world...  ...capable AI models and optimize them for use in a high-volume...  ...3 years of professional software engineering experience.... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...Cohere is a team of researchers, engineers, designers, and more, who...  ...energized by building high-performance, scalable and reliable...  ...closely with many teams to deploy optimized NLP models to production in...  ...influence latency and throughput of inference. ~ Strong understanding or... 
    Performance
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    8 hours ago
  • $300k

     ...of committed researchers, engineers, policy experts, and...  ...the Role The Cloud Inference team scales and optimizes Claude to serve the massive...  ...LLMs meet rigorous safety, performance, and security standards....  ...Have significant software engineering experience, with... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    8 hours ago
  •  ...Big More Better). You will own optimizations on both the training and on-robot inference stacks. We are still in a regime...  ...Implementing ML, hardware, and software changes that lead to step-function...  ...hundreds of millions of users, engineered the foundations of autonomous driving... 
    Full time

    The Generalist

    San Francisco, CA
    8 hours ago
  • $320k

     ...group of committed researchers, engineers, policy experts, and...  ...Our mandate is to make inference deployment boring and unattended...  ...continuous and unattended. As a Software Engineer on the Launch...  ...This is a resource-constrained optimization problem at its core: validation... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    8 hours ago
  •  ...Workspace.  What you'll do As a Software Engineer on our Site Reliability team at...  ...clear visibility into system health and performance. Partner with product and platform...  ...with LLM infrastructure — optimizing inference performance, managing fine-tuned models... 
    Performance
    Full time
    Flexible hours

    Sierra

    San Francisco, CA
    24 days ago
  • $170k - $216k

     ...evaluate the Waymo Driver's software stack at a massive...  ...of customers Software Engineers, Product, Data Science...  ...Build and evolve ML inference infrastructure for simulations...  ...frameworks, TPUs and optimizing models for serving....  ..., if the role can be performed remote, the specific... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    8 hours ago
  •  ...-on support from AMD engineers the team is scaling rapidly...  ...-on experience with performance engineering, learn...  ...large AI models are optimized and deployed at scale...  ..., and distributed inference features. Collaborate...  ...) ~3+ years of software engineering experience... 
    Performance
    Full time
    Work at office
    Flexible hours

    Sciforium

    San Francisco, CA
    8 hours ago
  • $195k - $225k

     ...and expanding the core engineering team in SF. The...  ...realtime collaboration, GPU inference at scale, a modern...  ...Role As the Senior Software Engineer – Backend (...  ...GPU workloads to optimizing GraphQL resolvers and...  ...ownership over APIs, performance, and data integrity —... 
    Performance
    Full time

    Vizcom

    San Francisco, CA
    8 hours ago
  •  ...team at OpenAI builds the low-level software that accelerates our most...  ...hardware and software, developing high-performance kernels, distributed system optimizations, and runtime improvements to make large-scale training and inference more efficient. Our work enables... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...powered workforce management that optimizes both human and AI capacity,...  ..., data pipelines, and inference servers to predict support contact...  ...with ML packages and software: Experience using Python libraries...  ...team. Passion for performance: A strong commitment to advancing... 
    Performance
    Full time

    Assembled

    San Francisco, CA
    8 hours ago
  •  ...and unblocked. You will work across engineering and infrastructure problems as they emerge...  ...scaling and orchestration issues to inference bottlenecks, numerical problems, and...  ...inference stacks. - Background in performance optimization, scaling, or production-critical... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  •  ...boundaries of data, scaling laws, optimization techniques, model...  ...Role We’re looking for a Software Engineer focused on building and scaling...  ...direct impact on system performance, reliability, and scale....  ...Collaborate across Pretraining, Inference, and Product teams to... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  • $300k - $320k

     ...of committed researchers, engineers, policy experts, and...  ...Anthropic is looking for backend software engineers to work across...  .... You'll own the performance, reliability, and efficiency...  ...'ll partner closely with inference and safeguards to optimize the full stack. API Capabilities... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    8 hours ago
  • $125k - $160k

     ...seeking a versatile Full Stack Software Engineer to join our engineering...  ...: Build delightful, performant, and accessible user experiences...  ..., libraries, and tools to optimize performance and developer experience...  ..., or local model inference (Ollama). Experience in... 
    Performance
    Full time
    Local area
    Visa sponsorship
    Work visa
    Shift work

    Cala Health

    San Francisco, CA
    8 hours ago
  •  ...for the architectural and engineering backbone of OpenAI’s infrastructure...  .... Our work spans system software, networking, platform...  ...-level monitoring, and performance optimization. About the Role We’re...  ...benchmarks, porting existing inference and training workloads to... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    8 hours ago
  • $152.5k - $287.5k

     ...batteries we already have.   Software Engineer, ML/Computer Vision (...  ...battery sorting, spanning ML inference, image acquisition, sensor...  ...services Monitor model performance in production to catch regressions...  ...edge deployment or model optimization techniques for inference (e... 
    Performance
    Hourly pay
    Full time
    Immediate start
    Shift work

    Redwood Materials

    San Francisco, CA
    8 hours ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic...  ...and help build the platform engineers turn to to ship AI products....  ...code directly impacts the performance of state-of-the-art machine...  ...powers modern AI workloads, optimizing every microsecond of... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    8 hours ago
  •  ...About the role As a founding software engineer at Bronco AI, you will be...  ...coverage with our AI Develop and optimize systems for data processing and model inference at scale Collaborate with...  ...Experience with high-performance computing and optimization... 
    Performance
    Full time
    Night shift

    Bronco Ai

    San Francisco, CA
    8 hours ago
  •  ...to join our elite team of engineers. As part of our team, you'll...  ...and promoting exceptional performance. HOW YOU’LL HAVE IMPACT...  ...TypeScript) components of our software stack. AI First: Adapt...  ..., prompt engineering, and optimizing inference for latency and cost. Prior... 
    Performance
    Full time
    Immediate start
    Remote work
    Flexible hours
    3 days per week

    Freed

    San Francisco, CA
    8 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Inference - Performance Optimization. Be the first to apply!