Software Engineer, Inference
Luma AI
About Luma AI Luma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change. Role & Responsibilities Ship new model architectures by integrating them into our inference engine Collaborate closely across research, engineering and infrastructure to streamline and optimize model efficiency and deployments Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows Automate, test and maintain our inference services to ensure maximum uptime and reliability Optimize deployment workflows to scale across thousands of machines Manage and optimize our inference workloads across different clusters & hardware providers Build sophisticated scheduling systems to optimally leverage our expensive GPU resources while meeting internal SLOs Build and maintain CI/CD pipelines for processing/optimizing model checkpoints, platform components, and SDKs for internal teams to integrate into our products/internal tooling Background Strong Python and system architecture skills Experience with model deployment using PyTorch, Huggingface, vLLM, SGLang, tensorRT-LLM, or similar Experience with queues, scheduling, traffic-control, fleet management at scale Experience with Linux, Docker, and Kubernetes Bonus points: Experience with modern networking stacks, including RDMA (RoCE, Infiniband, NVLink) Experience with high performance large scale ML systems (>100 GPUs) Experience with FFmpeg and multimedia processing Example Projects Create a resilient artifact store that manages all checkpoints across multiple versions of multiple models Enable hotswapping of models for our GPU workers based on live traffic patterns Build a robust queueing system for our jobs that take into account cluster availability and user priority Architect a e2e model serving deployment pipeline for a custom vendor Integrate our inference stack into an online reinforcement learning pipeline Regression & precision testing across different hardware platforms Building a full tracing system to trace the end-to-end lifetime of any inference workload Tech stack Must have Python Redis S3-compatible Storage Model serving (one of: PyTorch, vLLM, SGLang, Huggingface) Understanding of large-scale orchestration, deployment, scheduling (via Kubernetes or similar) Nice to have CUDA FFmpeg #J-18808-Ljbffr Luma AI
$185k - $250k
About Baseten Baseten powers mission‑critical inference for the world’s most dynamic AI companies, like Cursor... ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer on the Core Product team at Baseten, you...SuggestedFlexible hours$310k
About the Team OpenAI's Inference team powers the deployment of our most advanced models -... ...world. We're a small, fast-moving team of engineers focused on delivering a world-class... ...research. About the Role We're looking for a software engineer to help us serve OpenAI's...Suggested- ...tools consistently fail. We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises, YC startups... ...plus About the Role Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own profiling, batching, and...SuggestedWork at officeVisa sponsorshipRelocation package
- Senior Software Engineer - AI Inference Systems AI needs far more compute than exists today, but simply adding more hardware isn't the answer. We're partnering with a well-funded Series A startup, founded by entrepreneurs with a proven track record of building successful...Suggested
$295k
About The Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and... ..., analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. Background...Suggested$150k - $230k
...Model Performance team. The role involves designing and operating Model APIs to enhance AI model performance focusing on advanced inference capabilities. The ideal candidate should have over 3 years of experience in distributed systems or APIs and strong communication skills...- About The Role We are hiring Software Engineers focused on AI Infrastructure to build the systems that enable frontier multimodal AI to operate... ...engineering — including GPU orchestration, large-scale inference systems, performance optimization, and developer platforms that...InternshipImmediate start
$325k
About the Team Our Inference team brings OpenAI's most capable research and technology to... ...inference. About the Role We are looking for an engineer who wants to take the world's largest... .... Have at least 5 years of professional software engineering experience. Have or can...- Baseten is hiring a Product Engineer on the Dedicated Inference team to shape the developer experience for deploying and operating AI workloads in production. You will build CLI, SDKs, APIs, observability tools, and debugging workflows that customers rely on to manage mission...
$160k - $250k
Together AI is building the Inference Platform that brings the most advanced generative AI... ...balancing across data centers and model engine pods. Develop auto‑scaling systems to dynamically... ...is a strong plus. Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC...Full timeLocal area$160k - $250k
Together AI is searching for a skilled engineer to join their team, focusing on building and optimizing a large-scale inference platform. This role involves developing low-latency systems and collaborating with research teams to deploy advanced AI models. Candidates should...$200k
Platform Engineer - Inference Optimization We build and operate large-scale LLM inference and training infrastructure serving millions of users. This role focuses on deep optimization of SOTA serving frameworks and building a scalable, low-latency, cost-efficient AI platform...- Together AI in San Francisco is seeking a Research Engineer to help build a platform that lets users customize open-source models, bridging post-training and production inference. You will contribute across Fine-Tuning, RL, and Evaluation services and collaborate with product...Full timeWorldwide
- Perplexity is seeking an experienced platform engineer to own a unified, self-serve compute platform for training and inference workloads. You will design systems that launch training jobs and operate inference services without GPU provisioning burdens. You will manage...
- Perplexity seeks an experienced platform engineer to own and evolve a self-serve GPU compute platform. You will design and operate GPU... ...-cloud orchestration to support both training and real-time inference workloads. You’ll implement fault-tolerant scheduling, multi-cluster...
- Lightning AI in San Francisco or Seattle is seeking a Senior Application Security Engineer to secure our AI/ML platforms and inference services. You will work with platform, ML, and infrastructure teams to identify risks and implement secure architectures. The role emphasizes...
- Gravity Engineering Services Pvt Ltd. is seeking a Staff Engineer to lead technical efforts for their Inference Runtime. This role requires a deep understanding of systems engineering... ...have extensive experience in managing software engineering for inference runtimes, striving...
- ...San Francisco is seeking a Member of Technical Staff in Product Engineering to build the platform for PI's models. You will enable... ...to access models, perform data ingestion, and deploy validated inference end to end. You will be embedded with partner engagements to diagnose...
- ...Francisco is seeking a Member of Technical Staff, ML Product Engineer to develop APIs and systems that support genome-scale workloads... ...involves building reliable batch systems and ensuring efficient inference for scientific applications. The ideal candidate will have...
- ...About the Role We are seeking an experienced Engineering Manager to lead the Cloud Inference team for AWS. You will lead your team to scale and optimize... ...years of experience in high‑scale, high‑reliability software development , particularly infrastructure or capacity...
- MeshyAI is seeking a Platform Engineer to enhance our AI inference platform in San Francisco. You will design and develop core capabilities, focusing on resource management and service orchestration. The ideal candidate holds a relevant degree and has experience in backend...Remote jobFlexible hours
$300k
...growing group of committed researchers, engineers, policy experts, and business leaders working... ...AI systems. About The Role The Cloud Inference team scales and optimizes Claude to... ...May Be a Good Fit If You Have significant software engineering experience, with a strong background...Visa sponsorship$300k
A forward-thinking AI company in San Francisco is seeking an experienced software engineer to join their Inference team. The role focuses on building and optimizing systems for AI model deployment, with a strong emphasis on technical excellence and societal impact. Candidates...- ...At Inductive Bio, our goal is to build software that can dramatically improve how molecules... .... We are seeking a full-stack software engineer to join our talented, ambitious, and... ...infrastructure for model management and low-latency inference, including security features,...
- ...queries a month, and every one of them fans out into multiple AI inference requests running in real time. Behind that sits a large GPU... ...fleet spread across several cloud providers. Today, our inference engineers and researchers build models while also managing networking,...Shift work
- Tech Lead, Data & Inference Engineer Location: San Francisco Work type: Full Time Compensation: above market base + bonus + equity Roles & Responsibilities Lead the design, development and scaling of an end‑to‑end data platform from ingestion to insights, ensuring...Full time
$167.2k - $209k
A leading cloud service provider is seeking a Senior Engineer 2 for their AI Inference Data Plane team. This remote role focuses on designing and developing high-scale, resilient data plane services that enhance AI-driven applications. The ideal candidate will have strong...Remote job- Gravity Engineering Services Pvt Ltd. is looking for an experienced Engineering Manager to lead the Cloud Inference team for AWS. In this role, you will manage the end-to-end product of... ...10 years of experience in high-scale software development and at least 5 years in engineering...
- Gravity Engineering Services Pvt Ltd. in San Francisco is looking for a specialized engineer to advance the efficiency of ML inference systems. The role encompasses algorithm design, system optimization, and the integration of RL-driven training techniques. Ideal candidates...
- Acceler8 Talent is seeking a Senior Software Engineer to design and optimize AI inference systems for production workloads. You’ll work across runtime behavior, scheduling, memory management and system performance to deliver faster, more scalable AI inference. You’ll collaborate...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer, Inference. Be the first to apply!
- ngo software engineer San Francisco, CA
- software developer San Francisco, CA
- software developer internship no experience San Francisco, CA
- junior software developer San Francisco, CA
- part time software developer remote San Francisco, CA
- financial software developer San Francisco, CA
- senior software engineer ruby on rails San Francisco, CA
- software engineer amazon San Francisco, CA
- senior software design engineer San Francisco, CA
- software engineer part time San Francisco, CA

