Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Inference - Multi Modal

Full-time

OpenAI

About the Team

OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms. Our work ensures these models are available, performant, and scalable in production, and we partner closely with Research to bring the next generation of models into the world. We're a small, fast-moving team of engineers focused on delivering a world-class developer experience while pushing the boundaries of what AI can do.

We’re expanding into multimodal inference, building the infrastructure needed to serve models that handle image, audio, and other non-text modalities. These workloads are inherently more heterogeneous and experimental, involving diverse model sizes and interactions, more complex input/output formats, and tighter coordination with product and research.

About the Role

We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale. You’ll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production.

This work is inherently cross-functional: you’ll collaborate directly with researchers training these models and with product teams defining new modalities of interaction. You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text.

In this role, you will:

  • Design and implement inference infrastructure for large-scale multimodal models.

  • Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs.

  • Enable experimental research workflows to transition into reliable production services.

  • Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities.

  • Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers.

You might thrive in this role if you:

  • Have experience building and scaling inference systems for LLMs or multimodal models.

  • Have worked with GPU-based ML workloads and understand the performance dynamics of large models, especially with complex data like images or audio.

  • Enjoy experimental, fast-evolving work and collaborating closely with research.

  • Are comfortable dealing with systems that span networking, distributed compute, and high-throughput data handling.

  • Have familiarity with inference tooling like vLLM, TensorRT-LLM, or custom model parallel systems.

  • Own problems end-to-end and are excited to operate in ambiguous, fast-moving spaces.

Nice to Have:

  • Experience working with image generation or audio synthesis models in production.

  • Exposure to distributed ML training or system-efficient model design.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement .

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form . No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link .

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Vacancy posted more than 2 months ago
Similar jobs that could be interesting for youBased on the Software Engineer, Inference - Multi Modal in San Francisco, CA vacancy
  •  ...About the Team Our Inference team brings OpenAI’s most capable research and technology...  ...inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference...  ...-impact opportunity to shape OpenAI’s multi-platform inference capabilities from the... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE...  ...-scale, real-time infrastructure for multi-model voice agents - orchestrate STT, TTS... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE...  ..., reliability, and ease of use. As a Software Engineer on the Inference Stack team,... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  • $160k - $210k

     ...hardcore and obsessed team of the world’s best engineers and operators. If you are obsessed with...  ...Robotics organization is looking for a software engineer that will design, develop, and...  ...with robotics engineers, ML engineers, multi‑robot coordination teams, and product... 
    Suggested
    Full time
    Local area
    Flexible hours

    Nimble Nimble

    San Francisco, CA
    more than 2 months ago
  •  ...About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic...  ...and help build the platform engineers turn to to ship AI products....  ...We believe that as LLM and multi-modal workloads scale, the network...  ...to architect the software fabric that unifies thousands... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  • $224.5k - $251.5k

     ...help pioneer an advanced multi-agent orchestration...  ...fostering an AI-native engineering culture. This...  ...deploy scalable, multi-modal AI agents capable of autonomous...  ...agent frameworks, LLM inference optimization, advanced...  ...10+ years of relevant software engineering experience... 
    Full time
    Work at office

    Dialpad

    San Francisco, CA
    a month ago
  • $320k

     ...group of committed researchers, engineers, policy experts, and business...  ...the role The Cloud Inference team scales and optimizes Claude...  ...Have significant software engineering experience, with...  ...environments Solid understanding of multi-region deployments, geographic... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    more than 2 months ago
  • $176k - $220k

     ...Work together with engineers, scientists, operators...  ...hosted or self-hosted inference. You’ll also contribute...  ...Strong production software engineering experience...  ...gateways, proxies, or multi-tenant platform services...  ...AI gateway. vLLM, Modal, Ray, Triton, PyTorch,... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Handshake

    San Francisco, CA
    8 days ago
  •  ...About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks...  ..., analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About... 
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  •  ...Background: Specter is creating a software-defined "control plane" for...  ...software ecosystem on top of multi-modal wireless mesh sensing...  ...ultimately become the perception engine for a company's physical footprint...  ...of edge devices and cloud inference services, then close the loop... 
    Full time
    Shift work

    S.e. Specter

    San Francisco, CA
    27 days ago
  •  ...model innovation and systems engineering paired with a design-minded...  ..., and we are looking for a Software Engineer, Data Infrastructure...  ...closely with research and inference teams. Your work will directly...  ...Define Cartesia's multi-modal data strategy across pre-training... 
    Work at office
    Visa sponsorship
    Flexible hours

    Cartesia, Inc.

    San Francisco, CA
    3 days ago
  •  ...Background Specter is creating a software-defined "control plane" for...  ...software ecosystem on top of multi-modal wireless mesh sensing...  ...ultimately become the perception engine for a company's physical...  ...power real-time perception and inference across our edge-cloud platform... 
    Full time

    S.e. Specter

    San Francisco, CA
    more than 2 months ago
  • $170k - $216k

     ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve...  ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering...  ...You will: Build and evolve ML inference infrastructure for simulations. Be responsible... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    more than 2 months ago
  •  ...the Job We’re seeking an Agent Engineer to design and build agentic...  ...Qualifications ~3+ years of experience in software engineering, preferably in...  ...building the world’s best AI inference platform that makes large language and multi-modal models fast, efficient, and... 
    Full time
    Worldwide
    Flexible hours

    FriendliAI

    San Francisco, CA
    more than 2 months ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research and technology to...  ...About the Role We are looking for an engineer who wants to take the world's largest and...  ...Have at least 5 years of professional software engineering experience. Have or can quickly... 
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE...  ...generation), tool/function calling and multi-modal serving Profile and optimize... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  • $137.1k - $201.6k

     ...high-throughput batch inference, and fine-tuning on autoscaling...  ...serving and inference engines, fine-tuning and...  ...industry experience in software engineering ~ Deep backend...  ...with distributed/multi-node fine-tuning and training...  ...GPU platforms (e.g., Modal), or high-throughput... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    more than 2 months ago
  • $320k

     ...group of committed researchers, engineers, policy experts, and business...  ...Our mandate is to make inference deployment boring and unattended...  ...continuous and unattended. As a Software Engineer on the Launch...  ...manage complex state machines and multi-stage pipelines ~... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    more than 2 months ago
  • Senior Backend Engineer We believe using large language and multimodal...  ...authentication, billing, multi-tenant isolation, and zero tolerance...  ...layer that sits between our inference engine and every customer who...  ...large language and multi-modal models fast, efficient, and... 
    Worldwide
    Flexible hours

    FriendliAI Corp

    San Francisco, CA
    5 hours ago
  •  ...California. The Role: As a Full-Stack Software Engineer , you will be a core contributor to...  ...scaffolding frameworks for autonomous, multi-step decision-making What matters...  ...and reasoning systems that span multiple modalities (text, vision, actions) Experience using... 
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    a month ago
  • $213k - $263k

     ...Waymo builds technology that powers the Waymo Driver. Our software allows the Waymo Driver to perceive the world around...  ...data from a diverse set of sensors, enabling software engineers like you to develop multi-modal models and techniques at scale. Our mission is to... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    more than 2 months ago
  •  ...their business.   We are seeking a Software Engineer to build out our simulation and AI capabilities...  ...or modeling systems Causal inference — uplift modeling, synthetic controls,...  ...Experience building agentic AI systems or multi-agent simulations Big data... 
    Full time
    Work at office
    Remote work
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    more than 2 months ago
  • $320k - $405k

     ...group of committed researchers, engineers, policy experts, and business...  ...optimize how we use it. As a Software Engineer for Compute...  ...attribution frameworks for our multi-tenant infrastructure, enabling...  ...utilization across AI training and inference workloads—including large-... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    more than 2 months ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Capital. Join us and help build the platform engineers turn to to ship AI products. THE ROLE...  ...~ Experience developing and operating multi-tenant systems at scale, where authorization... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    22 days ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE...  ...in GEMM tuning and distributed/multi-GPU compute Contributions to open-source... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic...  ...and help build the platform engineers turn to to ship AI products....  ...seeking talented and experienced Software Engineers to join our...  ...and traces across Baseten’s multi-cloud infrastructure Own and... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    a month ago
  • $170k - $216k

     ...Waymo builds technology that powers the Waymo Driver. Our software allows the Waymo Driver to perceive the world around...  ...data from a diverse set of sensors, enabling software engineers like you to develop multi-modal models and techniques at scale. Our mission is to... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    more than 2 months ago
  •  ...is looking for a Cloud Infrastructure Engineer to own the architecture and evolution...  ...behind our GPU-accelerated AI inference cloud. As a Software Engineer, Cloud Infrastructure, you will...  ...inelastic, tenants must stay isolated, and multi-node serving depends on the network... 
    Permanent employment
    Full time
    Flexible hours

    FriendliAI

    San Francisco, CA
    a month ago
  •  ...Company Background: Specter is creating a software-defined “control plane” for the physical...  ...hardware-software ecosystem on top of multi-modal wireless mesh sensing technology. This...  ...platform will ultimately become the perception engine for a company’s physical footprint,... 
    Full time

    S.e. Specter

    San Francisco, CA
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Inference - Multi Modal. Be the first to apply!