Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Inference

Luma AI

Luma Model Serving EngineerYou'll own how Luma's models get served — integrating new architectures into the inference engine, scaling deployments across thousands of machines, and keeping expensive GPU fleets busy while meeting internal SLOs.This is large-scale inference systems work: scheduling, fleet management, deployment pipelines, and reliability across clusters and hardware providers. It fits a strong systems engineer comfortable with model serving and Kubernetes at scale. If you want pure modeling rather than the systems that run models, this is firmly the systems side.What You'll OwnShip new model architectures by integrating them into the inference engine.Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.Automate, test, and maintain inference services for maximum uptime and reliability.Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines.Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.First 90 DaysOne way the first 90 could unfold.Days 1–30 — Immerse & Diagnose: Learn the inference stack, the fleets, and where reliability or utilization break.Days 30–60 — Ship & Validate: Integrate a model or ship tooling/scheduling that improves uptime or GPU utilization.Days 60–90 — Scale & Systemize: Harden deployment pipelines and scheduling across clusters and providers.What You BringStrong Python and system-architecture skills.Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar.Experience with queues, scheduling, traffic control, and fleet management at scale.Experience with Linux, Docker, and Kubernetes, and with orchestration, deployment, and scheduling.Familiarity with Redis and S3-compatible storage.Nice to HaveModern networking stacks including RDMA (RoCE, InfiniBand, NVLink).High-performance large-scale ML systems (100+ GPUs).CUDA, and FFmpeg or multimedia processing.About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer, Inference in Redwood City, CA vacancy
  • $160.36k - $240.54k

     ...components. Develop observability to track ML model lifecycles from data generation to on-road validation.Maintain an in-house ML inference platform to serve large language models efficiently.Maintain an in-house ML compiler platform to compile, deploy, and validate Nuro... 
    Suggested
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    2 days ago
  •  ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application software team is the central nervous system of SpaceX - we create mission critical... 
    Suggested
    Permanent employment
    Temporary work
    Remote work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    3 days ago
  •  ...: AI), is the Enterprise AI application software company. C3 AI delivers a family of fully...  .... We are looking for a senior software engineer to help build the C3 Agentic AI Platform...  ...optimize latency, throughput, scalability, and inference cost across the platform.Strengthen... 
    Suggested
    Work experience placement

    C3 IoT

    Redwood City, CA
    3 days ago
  • $180k - $300k

     ...to outperform larger models despite using far less compute at inference time, substantially reducing the cost of deployment. For more...  ...area and has the deep expertise on both data research and data engineering necessary to solve this incredibly challenging problem and make... 
    Suggested
    Work at office
    Work from home
    Relocation package

    DatologyAI

    Redwood City, CA
    16 hours ago
  • $180k - $300k

     ...to outperform larger models despite using far less compute at inference time, substantially reducing the cost of deployment. For more...  ...area and has the deep expertise on both data research and data engineering necessary to solve this incredibly challenging problem and make... 
    Suggested
    Work at office
    Work from home
    Visa sponsorship
    Relocation package

    DatologyAI

    Redwood City, CA
    16 hours ago
  • $200k - $287.5k

     ...Senior Software Engineer On Billing PlatformAt Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we...  ...with Cortex AI, Intelligence, and Apps — including token-based inference, agent workflows, and AI-native subscriptions.Use AI as a... 
    Flexible hours

    Streamlit

    Menlo Park, CA
    4 days ago
  • $180k - $300k

     ...models despite using far less compute at inference time, substantially reducing the cost...  ...expertise on both data research and data engineering necessary to solve this incredibly...  ...environment. In this role, you will build software from the ground up to solve critical bottlenecks... 
    Work at office
    Work from home
    Relocation package

    DatologyAI

    Redwood City, CA
    3 days ago
  • $110k - $270k

     ...unit (GPNPU) architecture. Quadric's co-optimized software and hardware is targeted to run neural network (NN) inference workloads in a wide variety of edge and...  ...conventional C++ DSP and control code.RoleThe Full-Stack Engineer is key to making the Quadric product and... 
    Full time
    Work at office
    Local area
    Immediate start

    Quadric

    Burlingame, CA
    16 hours ago
  • $193.93k - $352.29k

     ...components. Develop observability to track ML model lifecycles from data generation to on-road validation. Maintain an in-house ML inference platform to serve large language models efficiently. Maintain an in-house ML compiler platform to compile, deploy, and validate... 
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    3 days ago
  • $160.36k - $240.54k

     ...components.  Develop observability to track ML model lifecycles from data generation to on-road validation. Maintain an in-house ML inference platform to serve large language models efficiently. Maintain an in-house ML compiler platform to compile, deploy, and validate... 
    Full time
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    15 days ago
  • $207k - $300k

    Analyze and optimize AI inference workloads across the application, model, and distributed...  ...s degree in Computer Science, Computer Engineering, Electrical Engineering, Applied...  ...practical experience.8 years of experience in software development.Experience in Python and... 

    Google

    Mountain View, CA
    2 days ago
  • $180k - $300k

     ...to outperform larger models despite using far less compute at inference time, substantially reducing the cost of deployment. For more...  ...area and has the deep expertise on both data research and data engineering necessary to solve this incredibly challenging problem and make... 
    Work at office
    Work from home
    Relocation package

    DatologyAI

    Redwood City, CA
    16 hours ago
  • $195k - $261k

     ...defense startup founded by two former Navy electrical engineers with a proven track record in robotics and software. We are developing an autonomous gun turret using...  ...for a Senior Embedded Software Engineer - Inference AI/ML to own the end-to-end process of taking... 
    Local area

    Allen Control Systems

    Mountain View, CA
    3 days ago
  •  ...Runtimes: Build and scale the orchestration engines that execute complex agentic workflows,...  ...theory in algorithms, systems, and software design. Experience: You bring 7+ years of...  ...that move efficiently from ingestion to inference. Customer‑Facing Experience: You have a... 
    Full time

    CVIN LLC

    Menlo Park, CA
    1 day ago
  • $215k - $290k

     ...talent firm that focuses on placing the best product managers, software, and hardware talent at innovative companies. Our team is 100%...  ...across the United States to help them hire. Full-Stack Software Engineer Location - Redwood City, CA (On-site) On-site role... 
    H1b
    Work at office
    Remote work
    Visa sponsorship

    Recruiting from Scratch

    Redwood City, CA
    3 days ago
  • $110k - $270k

     ...unit (GPNPU) architecture. Quadric's co-optimized software and hardware is targeted to run neural network (NN) inference workloads in a wide variety of edge and...  ...C++ DSP and control code.RoleThe AI Inference Engineer in Quadric is the key bridge between the world... 
    Full time
    Work at office
    Local area
    Immediate start
    2 days per week

    Quadric

    Burlingame, CA
    16 hours ago
  • $130k - $152k

     ...epicenter of this historic cultural and financial shift, keep reading.About the team + role Robinhood’s Credit Card & Banking Product Engineering team conceptualizes, designs and builds our entire customer-facing product: from mobile application to underlying backend systems... 
    Work at office
    Immediate start
    Flexible hours
    Shift work

    Robinhood Financial

    Menlo Park, CA
    2 days ago
  • C3 AI (NYSE: AI), is the Enterprise AI application software company. C3 AI delivers a family of fully integrated products including the...  .... Learn more at: C3 AIC3 AI is looking for Senior Software Engineers to join the rapidly growing Data org within the Platform Engineering... 
    Work experience placement

    C3 IoT

    Redwood City, CA
    16 hours ago
  • $240k - $360k

     ...direction, technical governance, and the evolution of software and network platforms within Network Platform Engineering (NPE).NPE designs, builds, and operates Equinix’...  ...of model lifecycle management and real-time inference systems (nice to have)Technical ExpertiseDeep... 
    Full time
    Work at office
    Shift work

    Equinix

    Redwood City, CA
    2 days ago
  • $160k - $190k

     ...Syntiant Corp., a leader in the high-growth AI software and semiconductor solutions space, is...  ...and talented Senior Software Engineer to take on a critical role with expansive...  ...architectures that enable world-leading inference speed and minimized memory footprint across... 
    Temporary work
    Flexible hours

    Syntiant

    Redwood City, CA
    3 days ago
  • $124.7k - $208.85k

     ...unique and trusted finds, from everyday pieces to one-of-a-kind vintage and luxury. About the RolePoshmark is looking for a Senior Software Engineer to support the rapid growth of our Cloud Platform and marketplace. In this role, you’ll help build and operate a high-traffic,... 

    Poshmark

    Redwood City, CA
    3 days ago
  • C3 AI (NYSE: AI), is the Enterprise AI application software company. C3 AI delivers a family of fully integrated products including the...  ...: C3 AIC3 AI is looking for a highly motivated Senior Software Engineer - Platform to join our Platform Engineering team. You will play... 

    C3 IoT

    Redwood City, CA
    3 days ago
  •  ...Software Engineer – SaaS PlatformHow often do you get the chance to make a global impact developing the latest AI inside of the "built world"? Reconstruct's Visual Command Center (VCC) uses AI and Machine Learning inside of computer vision to track the lifecycle of large... 
    Work at office

    Reconstruct

    Menlo Park, CA
    1 day ago
  • $120k - $160k

     ...learning capabilities. Our systems are software-driven, hardware-agnostic, and have already...  ...the Role:As an Application Support Engineer you are the front line between Dexterity...  ...Specific StrengthsExperience supporting AI/ML inference platforms (model drift, retraining... 
    Permanent employment
    Work at office

    Dexterity

    Redwood City, CA
    16 hours ago
  • $85k - $150k

    Full Stack Software Engineer at Blink.new - an AI App Creation Platform. Join us to design and develop end‑to‑end features, provide technical customer support, and build innovative AI‑powered applications. Base pay range $85,000.00/yr - $150,000.00/yr Seniority level... 
    Full time

    Blink

    Redwood City, CA
    16 hours ago
  • $165.2k - $223.6k

     ...behavior and continuously improve at massive scale.As a Software Development Engineer, you will solve challenging data and distributed-systems...  ...fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniquesPreferred qualification... 
    Internship
    Local area
    Flexible hours

    Amazon

    Palo Alto, CA
    16 hours ago
  • $153k - $179k

     ...team works closely with security, infrastructure, and product engineering partners to ensure data is both usable and protected. You will...  ...party integrations, and emerging AI-driven use cases.As a Senior Software Engineer, you will design and build backend systems that... 
    Work at office
    Flexible hours
    Shift work
    3 days per week

    Robinhood Financial

    Menlo Park, CA
    1 day ago
  • $165k - $206.5k

     ..., and reliability of our services. You will also partner with engineering teams across the organization to understand their needs and drive...  ..., Mathematics or a related field3+ years of professional software development experienceStrong understanding of common algorithms... 
    Live in
    Work at office
    Shift work
    3 days per week

    Box

    Redwood City, CA
    1 day ago
  • $130k - $152k

     ...mission is to close critical attack paths before they can be exploited and to mitigate active threats with speed and precision.As a Software Engineer on the Proactive Capabilities team , you will design, build, and scale engineering tools that empower Security Engineers,... 
    Work at office
    Flexible hours
    Shift work
    3 days per week

    Robinhood Financial

    Menlo Park, CA
    11 hours ago
  • $186 per hour

     ...overlooked industry.About the role You’ll be based out of our Redwood City office and work in person with the rest of our engineering team. As a Software Engineer, you’ll ship features end-to-end across our stack: a Go backend, a TypeScript/React web app, and React Native... 
    Internship
    Work at office
    Local area
    Remote work
    Relocation

    Rundoo

    Redwood City, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Inference. Be the first to apply!