Software Engineer, Inference - TL
OpenAI
About the Team
Our team brings OpenAI’s most capable research and technology to the world through our products. We empower consumers, enterprises, and developers alike to access state-of-the-art AI models - unlocking new capabilities across productivity, creativity, and more. We focus on high-performance model inference and accelerating research through efficient and reliable infrastructure.
About the Role
We’re looking for a hands-on Tech Lead to drive the design, optimization, and scaling of our inference systems. In this role, you’ll lead engineering efforts to ensure our largest models run with exceptional efficiency in high-throughput, low-latency environments. You’ll be responsible for shaping our CUDA strategy, driving performance at the kernel level, and collaborating across teams to deliver end-to-end production readiness.
In this role, you will:
Lead the design and implementation of core inference infrastructure for serving frontier AI models in production.
Own and optimize CUDA-based systems and kernels to maximize performance across our fleet.
Partner with researchers to integrate novel model architectures into performant, scalable inference pipelines.
Build tooling and observability to detect bottlenecks, guide system tuning, and ensure stable deployment at scale.
Collaborate cross-functionally to align technical direction across research, infra, and product teams.
Mentor engineers on GPU performance, CUDA development, and distributed inference best practices.
You may thrive in this role if you:
Have deep expertise in CUDA, including writing and optimizing high-performance kernels for inference or training workloads.
Have experience leading complex engineering efforts, particularly at the systems and performance layer of large-scale ML infrastructure.
Understand the full inference stack - from model loading and memory management to communication libraries and deployment orchestration.Are comfortable working in large, distributed GPU environments and debugging performance issues across hardware and software layers.
Have strong familiarity with PyTorch and NVIDIA’s GPU software stack (NCCL, NVLink, MIG, etc.).
Take a systems-level view, but aren’t afraid to dive into low-level code when performance is on the line.
Bonus:
Experience with inference frameworks like TensorRT, vLLM, SGLang, or custom model parallelism infrastructure.
Familiarity with TPU, AMD GPUs, ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron, MPI, or Horovod.
Familiarity with profiling tools (Nsight, nvprof, or custom observability stacks).
Background in HPC or large-scale distributed systems engineering.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or any other legally protected status.
For US Based Candidates: Pursuant to the San Francisco Fair Chance Ordinance, we will consider qualified applicants with arrest and conviction records.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link .
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
- ...About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re...SuggestedFull time
- ...serve OpenAI’s frontier models at massive scale. As part of the inference team, you’ll be responsible for unlocking every last FLOP from... ...stack. About the Role We are looking for a kernel-focused engineer to lead efforts in writing, porting, and optimizing GPU...SuggestedFull time
- ...consistently fail. We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises, YC startups... ...About the Role Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own profiling, batching, and autoscaling...SuggestedFull timeWork at officeVisa sponsorshipRelocation package
- ...About the Team OpenAI’s Inference team powers the deployment of our most advanced models... ...world. We're a small, fast-moving team of engineers focused on delivering a world-class... ...About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal...SuggestedFull time
- ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI... ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE... ..., reliability, and ease of use. As a Software Engineer on the Inference Stack team,...SuggestedFull timeFlexible hours
- ...About the Team OpenAI’s Inference team ensures that our most advanced models run efficiently, reliably, and at scale. We build and... ...hardware architectures like AMD. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across...Full time
$300k
...growing group of committed researchers, engineers, policy experts, and business leaders working... .... About the role Our Inference team is responsible for building and maintaining... ...fit if you: Have significant software engineering experience, particularly...Full timeWork at officeWorldwideVisa sponsorshipFlexible hours$142.2k - $204.6k
P-1284About This RoleAs a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks’ Foundation Model API. You’ll work at the intersection of research and production, ensuring our large language model...Local areaWorldwide$320k
...growing group of committed researchers, engineers, policy experts, and business leaders working... ...the Role Our mandate is to make inference deployment boring and unattended.... ...deployment continuous and unattended. As a Software Engineer on the Launch Engineering team,...Full timeWork at officeVisa sponsorshipFlexible hoursShift work- ...About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks... ..., analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About...Full time
- ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Voice is becoming...Full timeFlexible hours
$170k - $216k
...products that evaluate the Waymo Driver's software stack at a massive scale. We solve... ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering... ...You will: Build and evolve ML inference infrastructure for simulations. Be responsible...Full timeRemote work- ...About the Team Our Inference team brings OpenAI’s most capable research and technology to... ...About the Role We are looking for an engineer who wants to take the world's largest and... ...Have at least 5 years of professional software engineering experience. Have or can quickly...Full time
$300k
...growing group of committed researchers, engineers, policy experts, and business leaders working... .... About the Role The Cloud Inference team scales and optimizes Claude to serve... ...Fit If You: Have significant software engineering experience, with a strong background...Full timeWork at officeVisa sponsorshipFlexible hours$150k - $180k
...Capital , and JFF Ventures , and are now hiring a Full Stack Engineer to help build the product that institutions use to interact... ...application layer up to our data stack (Postgres + DuckDB) and model inference, and keep query and inference latency low enough that the...Full timeWork at officeImmediate start$125k - $160k
...Role Overview We are seeking a versatile Full Stack Software Engineer to join our engineering team. Reporting to the Software Engineering... ...-Augmented Generation) architectures, or local model inference (Ollama). Experience in automated testing at multiple levels...Full timeLocal areaVisa sponsorshipWork visaShift work$120k - $180k
...yet, our team is tackling cutting-edge engineering challenges to bring revolutionary products... ...We are looking for a full-stack software enginee r to turn whiteboard ideas into... ...features that showcase real-time sensing and inference in compelling, reliable ways....Full timeVisa sponsorship$250k - $300k
...reliably in production. That means owning the inference stack end to end: profiling where time... ...will also work directly with customer engineering teams to tailor deployments to their... ...service.Build and support the software and product features around the inference...Temporary work- ...video. Our team also manages large-scale inference and platform infrastructure that... ...over unchecked growth. Within Applied Engineering, the Ads Monetization team in Financial... ...Possess a minimum of 5 years of professional software engineering experience. Bring...Full time
- ...s best for our customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft. Each... ...), especially how they influence latency and throughput of inference. ~ Strong understanding or working experience with distributed...Full timeWork experience placementWork at officeRemote workFlexible hours
- ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Senior Enterprise...Full timeFlexible hours
$190k - $265k
...use deep data insights to improve their business. Founded by engineers — and customer-obsessed — we leap at every opportunity to solve... ...infrastructure that power the next generation of AI.The Foundation Model Inference team is the backbone of Databricks’ generative AI capabilities...Local areaWorldwide$230k - $265k
...tools they need. About The Position We’re looking for a software engineer to join Parafin’s Infrastructure team and lead the evolution... ...systems for model experimentation, training, evaluation, inference, and retraining that power underwriting and other ML-driven...Full timeWork from homeFlexible hours- ...spend worldwide. Using frontier causal inference-based econometric models to run experiments... ...product managers, economists, and engineers from Google, Netflix, Meta, and Amazon,... ...experience building and shipping production software systems ~ Must have strong Python proficiency...Full timeWork at officeWork from homeWorldwideFlexible hours
- ...tools they need.About the Position:We’re looking for a seasoned software engineer to join Parafin’s Infrastructure team and lead the... ...lifecycle (training, deployment, monitoring, retraining), real-time inference.Contributions to internal tooling or open-source projects in...Flexible hours
$238k - $290k
...started.Role OverviewAs a Backend Platform Engineer at Harvey, you will help build and... ...such as supporting high-throughput model inference, managing streaming and long-running API... ...with confidenceWhat You Have5+ years of software engineering experience (post-BS/MS), including...Flexible hoursShift work- ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a...Full timeFlexible hours
$225k
...approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal. About the role: As a Software Engineer on the product team, you’ll be responsible for building and maintaining our product...Full timeLocal areaRelocationVisa sponsorship- ...About the Role We’re hiring three exceptional Founding Software Engineers to help us scale the computational biology platform that... ...Computational Biology. We deal with problems ranging from scaling ML inference on AWS for hundreds of GPUs to dissecting pdb files with...Full timeRelocation
- ...latency voice interactions. We partner closely with Research and Inference to bring frontier model capabilities to developers and use... ...feedback to improve our models. About the Role As a software engineer on API Multimodal, you will build and operate the products...Full timeInternship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, Inference - TL. Be the first to apply!
- agile software developer San Francisco, CA
- software developer internship no experience San Francisco, CA
- intermediate software engineer San Francisco, CA
- software engineer staff San Francisco, CA
- experienced software developer San Francisco, CA
- work from home software developer San Francisco, CA
- software developer no experience San Francisco, CA
- software developer fintech San Francisco, CA
- software data engineer San Francisco, CA
- financial software developer San Francisco, CA

