Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Inference

Luma AI

About Luma AI Luma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change. Role & Responsibilities Ship new model architectures by integrating them into our inference engine Collaborate closely across research, engineering and infrastructure to streamline and optimize model efficiency and deployments Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows Automate, test and maintain our inference services to ensure maximum uptime and reliability Optimize deployment workflows to scale across thousands of machines Manage and optimize our inference workloads across different clusters & hardware providers Build sophisticated scheduling systems to optimally leverage our expensive GPU resources while meeting internal SLOs Build and maintain CI/CD pipelines for processing/optimizing model checkpoints, platform components, and SDKs for internal teams to integrate into our products/internal tooling Background Strong Python and system architecture skills Experience with model deployment using PyTorch, Huggingface, vLLM, SGLang, tensorRT-LLM, or similar Experience with queues, scheduling, traffic-control, fleet management at scale Experience with Linux, Docker, and Kubernetes Bonus points: Experience with modern networking stacks, including RDMA (RoCE, Infiniband, NVLink) Experience with high performance large scale ML systems (>100 GPUs) Experience with FFmpeg and multimedia processing Example Projects Create a resilient artifact store that manages all checkpoints across multiple versions of multiple models Enable hotswapping of models for our GPU workers based on live traffic patterns Build a robust queueing system for our jobs that take into account cluster availability and user priority Architect a e2e model serving deployment pipeline for a custom vendor Integrate our inference stack into an online reinforcement learning pipeline Regression & precision testing across different hardware platforms Building a full tracing system to trace the end-to-end lifetime of any inference workload Tech stack Must have Python Redis S3-compatible Storage Model serving (one of: PyTorch, vLLM, SGLang, Huggingface) Understanding of large-scale orchestration, deployment, scheduling (via Kubernetes or similar) Nice to have CUDA FFmpeg #J-18808-Ljbffr Luma AI

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Software Engineer, Inference in San Francisco, CA vacancy
  • $185k - $250k

    About Baseten Baseten powers mission‑critical inference for the world’s most dynamic AI companies, like Cursor...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Software Engineer on the Core Product team at Baseten, you... 
    Suggested
    Flexible hours

    Baseten

    San Francisco, CA
    13 hours ago
  • $310k

    About the Team OpenAI's Inference team powers the deployment of our most advanced models -...  ...world. We're a small, fast-moving team of engineers focused on delivering a world-class...  ...research. About the Role We're looking for a software engineer to help us serve OpenAI's... 
    Suggested

    OpenAI

    San Francisco, CA
    13 hours ago
  •  ...tools consistently fail. We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises, YC startups...  ...plus About the Role Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own profiling, batching, and... 
    Suggested
    Work at office
    Visa sponsorship
    Relocation package

    Trypulse

    San Francisco, CA
    13 hours ago
  • Senior Software Engineer - AI Inference Systems AI needs far more compute than exists today, but simply adding more hardware isn't the answer. We're partnering with a well-funded Series A startup, founded by entrepreneurs with a proven track record of building successful... 
    Suggested

    Acceler8 Talent

    San Francisco, CA
    13 hours ago
  • $295k

    About The Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks and...  ..., analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. Background... 
    Suggested

    OpenAI

    San Francisco, CA
    13 hours ago
  • $150k - $230k

     ...Model Performance team. The role involves designing and operating Model APIs to enhance AI model performance focusing on advanced inference capabilities. The ideal candidate should have over 3 years of experience in distributed systems or APIs and strong communication skills... 

    Dormont Manufacturing Company

    San Francisco, CA
    2 days ago
  • About The Role We are hiring Software Engineers focused on AI Infrastructure to build the systems that enable frontier multimodal AI to operate...  ...engineering — including GPU orchestration, large-scale inference systems, performance optimization, and developer platforms that... 
    Internship
    Immediate start

    SPREEAI

    San Francisco, CA
    13 hours ago
  • $325k

    About the Team Our Inference team brings OpenAI's most capable research and technology to...  ...inference. About the Role We are looking for an engineer who wants to take the world's largest...  .... Have at least 5 years of professional software engineering experience. Have or can... 

    OpenAI

    San Francisco, CA
    13 hours ago
  • Baseten is hiring a Product Engineer on the Dedicated Inference team to shape the developer experience for deploying and operating AI workloads in production. You will build CLI, SDKs, APIs, observability tools, and debugging workflows that customers rely on to manage mission... 

    Baseten

    San Francisco, CA
    1 day ago
  • $160k - $250k

    Together AI is building the Inference Platform that brings the most advanced generative AI...  ...balancing across data centers and model engine pods. Develop auto‑scaling systems to dynamically...  ...is a strong plus. Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC... 
    Full time
    Local area

    Together AI

    San Francisco, CA
    13 hours ago
  • $160k - $250k

    Together AI is searching for a skilled engineer to join their team, focusing on building and optimizing a large-scale inference platform. This role involves developing low-latency systems and collaborating with research teams to deploy advanced AI models. Candidates should... 

    Together AI

    San Francisco, CA
    13 hours ago
  • $200k

    Platform Engineer - Inference Optimization We build and operate large-scale LLM inference and training infrastructure serving millions of users. This role focuses on deep optimization of SOTA serving frameworks and building a scalable, low-latency, cost-efficient AI platform... 

    Kaon (prev. FlowGPT)

    San Francisco, CA
    13 hours ago
  • Together AI in San Francisco is seeking a Research Engineer to help build a platform that lets users customize open-source models, bridging post-training and production inference. You will contribute across Fine-Tuning, RL, and Evaluation services and collaborate with product... 
    Full time
    Worldwide

    Together AI

    San Francisco, CA
    13 hours ago
  • Perplexity is seeking an experienced platform engineer to own a unified, self-serve compute platform for training and inference workloads. You will design systems that launch training jobs and operate inference services without GPU provisioning burdens. You will manage... 

    Apply

    San Francisco, CA
    2 days ago
  • Perplexity seeks an experienced platform engineer to own and evolve a self-serve GPU compute platform. You will design and operate GPU...  ...-cloud orchestration to support both training and real-time inference workloads. You’ll implement fault-tolerant scheduling, multi-cluster... 

    Perplexity

    San Francisco, CA
    1 day ago
  • Lightning AI in San Francisco or Seattle is seeking a Senior Application Security Engineer to secure our AI/ML platforms and inference services. You will work with platform, ML, and infrastructure teams to identify risks and implement secure architectures. The role emphasizes... 

    Lightning-Ai

    San Francisco, CA
    5 days ago
  • Gravity Engineering Services Pvt Ltd. is seeking a Staff Engineer to lead technical efforts for their Inference Runtime. This role requires a deep understanding of systems engineering...  ...have extensive experience in managing software engineering for inference runtimes, striving... 

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    13 hours ago
  •  ...San Francisco is seeking a Member of Technical Staff in Product Engineering to build the platform for PI's models. You will enable...  ...to access models, perform data ingestion, and deploy validated inference end to end. You will be embedded with partner engagements to diagnose... 

    Physical Intelligence

    San Francisco, CA
    1 day ago
  •  ...Francisco is seeking a Member of Technical Staff, ML Product Engineer to develop APIs and systems that support genome-scale workloads...  ...involves building reliable batch systems and ensuring efficient inference for scientific applications. The ideal candidate will have... 

    Jobr

    San Francisco, CA
    2 days ago
  •  ...About the Role We are seeking an experienced Engineering Manager to lead the Cloud Inference team for AWS. You will lead your team to scale and optimize...  ...years of experience in high‑scale, high‑reliability software development , particularly infrastructure or capacity... 

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    4 days ago
  • MeshyAI is seeking a Platform Engineer to enhance our AI inference platform in San Francisco. You will design and develop core capabilities, focusing on resource management and service orchestration. The ideal candidate holds a relevant degree and has experience in backend... 
    Remote job
    Flexible hours

    MeshyAI

    San Francisco, CA
    13 hours ago
  • $300k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  ...AI systems. About The Role The Cloud Inference team scales and optimizes Claude to...  ...May Be a Good Fit If You Have significant software engineering experience, with a strong background... 
    Visa sponsorship

    Anthropic

    San Francisco, CA
    13 hours ago
  • $300k

    A forward-thinking AI company in San Francisco is seeking an experienced software engineer to join their Inference team. The role focuses on building and optimizing systems for AI model deployment, with a strong emphasis on technical excellence and societal impact. Candidates... 

    Anthropic

    San Francisco, CA
    13 hours ago
  •  ...At Inductive Bio, our goal is to build software that can dramatically improve how molecules...  .... We are seeking a full-stack software engineer to join our talented, ambitious, and...  ...infrastructure for model management and low-latency inference, including security features,... 

    Inductive Bio, Inc.

    San Francisco, CA
    1 day ago
  •  ...queries a month, and every one of them fans out into multiple AI inference requests running in real time. Behind that sits a large GPU...  ...fleet spread across several cloud providers. Today, our inference engineers and researchers build models while also managing networking,... 
    Shift work

    Apply

    San Francisco, CA
    2 days ago
  • Tech Lead, Data & Inference Engineer Location: San Francisco Work type: Full Time Compensation: above market base + bonus + equity Roles & Responsibilities Lead the design, development and scaling of an end‑to‑end data platform from ingestion to insights, ensuring... 
    Full time

    Catalyst Labs

    San Francisco, CA
    13 hours ago
  • $167.2k - $209k

    A leading cloud service provider is seeking a Senior Engineer 2 for their AI Inference Data Plane team. This remote role focuses on designing and developing high-scale, resilient data plane services that enhance AI-driven applications. The ideal candidate will have strong... 
    Remote job

    DigitalOcean

    San Francisco, CA
    1 day ago
  • Gravity Engineering Services Pvt Ltd. is looking for an experienced Engineering Manager to lead the Cloud Inference team for AWS. In this role, you will manage the end-to-end product of...  ...10 years of experience in high-scale software development and at least 5 years in engineering... 

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    13 hours ago
  • Gravity Engineering Services Pvt Ltd. in San Francisco is looking for a specialized engineer to advance the efficiency of ML inference systems. The role encompasses algorithm design, system optimization, and the integration of RL-driven training techniques. Ideal candidates... 

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    1 day ago
  • Acceler8 Talent is seeking a Senior Software Engineer to design and optimize AI inference systems for production workloads. You’ll work across runtime behavior, scheduling, memory management and system performance to deliver faster, more scalable AI inference. You’ll collaborate... 

    Acceler8 Talent

    San Francisco, CA
    13 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Inference. Be the first to apply!