Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer - GenAI inference

Full-time

Databricks Inc.

P-1284

About This Role


As a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks’ Foundation Model API. You’ll work at the intersection of research and production, ensuring our large language model (LLM) serving systems are fast, scalable, and efficient. Your work will touch the full GenAI inference stack — from kernels and runtimes to orchestration and memory management.

What You Will Do



  • Contribute to the design and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference

  • Collaborate with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine

  • Optimize for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators

  • Build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations

  • Develop and enhance scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads

  • Support reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioning

  • Integrate with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overhead

  • Collaborate cross-functionally: with platform engineers, cloud infrastructure, and security/compliance teams

  • Document and share learnings, contributing to internal best practices and open-source efforts when possible

What We Look For



  • BS/MS/PhD in Computer Science, or a related field

  • Strong software engineering background (3+ years or equivalent) in performance-critical systems

  • Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc.

  • Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)

  • Comfortable designing and operating distributed systems, including RPC frameworks, queuing, RPC batching, sharding, memory partitioning

  • Demonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)

  • Experience building instrumentation, tracing, and profiling tools for ML models

  • Ability to work closely with ML researchers, translate novel model ideas into production systems

  • Ownership mindset and eagerness to dive deep into complex system challenges

  • Bonus: published research or open-source contributions in ML systems, inference optimization, or model serving

 

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer - GenAI inference in Remote vacancy
  • $2,000 per month

     ...architecture and design of the Sohu host software stack Implement high-performance,...  ...handling continuous batching and real time inference Implement inference-time...  ...person team in Cupertino, and greatly value engineering skills. We do not have boundaries between... 
    Suggested
    Full time
    Work at office
    Relocation package

    Etched

    Remote
    17 hours ago
  • $170k - $216k

     ...solutions to speed up developer velocity. We’re looking for a software engineer to join the team to build and maintain the critical data and...  ...Software Engineer.   You will: Develop Waymo's inference platform to make it scalable, high throughput, and low... 
    Suggested
    Full time
    Remote work

    Waymo

    Remote
    17 hours ago
  • $150k - $205k

     ...offices around the world, Aeris is the preeminent IoT software company globally powering critical projects across energy...  ...to expand, we are seeking a experienced Software Engineer to join our Generative AI (GenAI) team. In this critical role, you will be responsible... 
    Suggested
    Full time
    Shift work

    Aeris Communications

    Remote
    17 hours ago
  • $92k - $135k

     ...intelligence that drives innovation.  What You’ll Do: Join the Inference team to ship production features that improve latency,...  ...practices, and grow quickly with mentorship from experienced engineers. About the role: Implement well-scoped features and... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Remote work
    Flexible hours

    Coreweave

    Washington DC
    17 hours ago
  • $190.8k - $267.1k

     ...Learning teams. What You’ll Do: As a Senior Software Engineer, you will lead the development of a large-scale GenAI Platform at Reddit. Contribute to the...  ...lifecycle. ~ Strong knowledge of model serving, inference pipelines, monitoring, and observability for... 
    Suggested
    Full time
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    Remote
    21 days ago
  • $160k - $190k

     ...Sponsorship available*** About e360’s App Engineering e360 is a 30+ year privately-...  ....    What You’ll Do: Deliver GenAI solutions to customers as Professional...  ...and evaluate the applicability of new software technologies to platform development efforts... 
    Full time
    Remote work

    E360

    Remote
    17 hours ago
  • $170k - $216k

     ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve...  ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering...  ...You will: Build and evolve ML inference infrastructure for simulations. Be responsible... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    17 hours ago
  •  ...frameworks. Strong track record of working with machine learning systems and/or platforms. Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between them. Experience serving fine-tuned LLMs (PEFT, DPO, RL... 
    Full time

    Snowflake

    Remote
    17 hours ago
  • $152.2k - $243.7k

     ...influence how modern cloud and GenAI-powered platforms are...  ...cloud technologies, platform engineering, and Generative AI enablement...  ...safer, and smarter delivery of software at scale. This is a...  ...model training, deployment, inference, and experimentation. Drive... 
    Full time
    Work experience placement
    Work at office
    Local area
    Worldwide
    Relocation package

    Visa

    Remote
    17 hours ago
  • $200k - $250k

     ...experience, and we’re looking for a Senior MLOps Engineer to help us run them reliably and...  ..., and scaling — across a custom-built inference platform powering a live conversational...  ...For ~5–8+ years of experience in software, ML, platform, or infrastructure engineering... 
    Full time
    Remote work
    Flexible hours

    Wizard

    United States
    17 hours ago
  • $320k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  .... About the role The Cloud Inference team scales and optimizes Claude to serve...  ...qualifications Have significant software engineering experience, with a strong background... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Remote
    17 hours ago
  • $88k - $136.9k

     ...starts with you. Job Description The SW Engineer is responsible for conducting, planning, or overseeing one or...  ...computer applications. ~ Experience in debugging and modifying software programs. ~ Experience in writing and maintaining technical... 
    Full time
    Work experience placement
    Work at office
    Local area

    Visa

    Remote
    17 hours ago
  • Role Description In the Senior Engineer role, you will own meaningful subsystems of Stack AV's inference platform and drive them from design through production. You will...  .... This position may also involve working with software and technologies subject to U.S. export... 
    Full time

    Stack AV

    Remote
    4 days ago
  • $110k - $270k

     ...unit (GPNPU) architecture. Quadric's co-optimized software and hardware is targeted to run neural network (NN) inference workloads in a wide variety of edge and...  ...and control code. Role The Full-Stack Engineer is key to making the Quadric product and toolchain... 
    Full time
    Work at office
    Local area
    Immediate start
    Worldwide
    Flexible hours

    Quadric.Io, Inc

    Remote
    17 hours ago
  • $166k - $220k

     ...for Mission Autonomy, Anduril’s premier software platform that enables masses of Fury, Barracuda...  ..., Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems...  ...are also pioneering the integration of GenAI agents into our systems, allowing... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Remote
    17 hours ago
  • $8k

     ...Visionist has an exciting new opportunity for a Full Stack Software Engineer. You will be joining a critical mission supporting our customers...  ...implement, and optimize infrastructure to support AI model inference at scale - Support the development and ongoing... 
    Permanent employment
    Full time
    Contract work
    Temporary work
    Flexible hours

    Visionist

    Laurel, MD
    17 hours ago
  • $320k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  ...the Role Our mandate is to make inference deployment boring and unattended....  ...deployment continuous and unattended. As a Software Engineer on the Launch Engineering team,... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    Remote
    17 hours ago
  • $139.2k - $174k

     ...generation of AI-driven applications. We are seeking a Senior Engineer 2 to join our AI Inference Data Plane team. In this role, you will be a key...  ...with gRPC. ~Proven experience shipping customer-facing software products and running critical services in a high-scale environment... 
    Full time
    Remote work

    DigitalOcean

    Remote
    1 day ago
  •  ...s best for our customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft. Each...  ...), especially how they influence latency and throughput of inference. ~ Strong understanding or working experience with distributed... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    17 hours ago
  •  ...We Are Synopsys is the leader in engineering solutions from silicon to systems, enabling customers...  ...of what you are actually solving for. GenAI coding tools do not intimidate you. You...  ..., and maintainable code following software engineering best practices, with unit test... 

    Synopsys

    Mountain View, CA
    3 days ago
  •  ...facing platforms. We are looking for an AI Engineer to design, build, and operate production-...  ...product managers, data engineers, and software engineers in an agile environment, owning...  ...development, including at least 1-2 years building GenAI-driven applications. Multi-Agent... 
    Full time
    Work at office

    Dynasty Financial Partners

    Remote
    17 hours ago
  • $202.5k - $247.5k

     ...universal gateway for API delivery, AI inference, device fleets, and site-to-site connectivity...  ...are vital to our success! We like software that’s serious and culture that’s not...  ...Our Data Platform team is part of the Engineering organization and doesn’t live in a silo... 
    Permanent employment
    Full time
    Live in
    Work at office
    Local area
    Remote work
    Home office
    Flexible hours

    Ngrok

    United States
    17 hours ago
  • $2,000 - $4,500 per month

     ...expertise comes into play: We are looking for a talented software engineer to help us build and scale our image capture and analysis platform...  ..., manage, and optimize data processing and machine learning inference pipelines on our servers. Troubleshoot problems across... 
    Remote job
    Full time
    Work experience placement

    The Glacier

    United States
    17 hours ago
  • $202.5k - $247.5k

     ...universal gateway for API delivery, AI inference, device fleets, and site-to-site connectivity...  ...are vital to our success! We like software that’s serious and culture that’s not...  ...Platform team builds the systems ngrok engineers rely on to build, deploy, and operate ngrok... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work
    Home office
    Flexible hours

    Ngrok

    Remote
    17 hours ago
  • $187k

     ...company. We enable data scientists and ML engineers to develop, train, deploy, and monitor...  ...training, feature computation, real-time inference, and experimentation. You’ll work at the...  ...foundation in computer science and software engineering principles ~ Deeply interested... 
    Full time
    Work at office
    Local area
    Remote work
    Night shift

    Chime

    Remote
    17 hours ago
  • $100k

     ...innovation. The Opportunity We are looking for a driven Software Engineer to join the Training Platform team under our Machine...  ...engineering on production systems dealing with training or inference of deep learning models. Proven track record of building... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    Remote
    17 hours ago
  • $175k - $215k

     ...across cities. The Platform team is part of the Marketplace engineering: we provide specialized business platforms supporting...  ...life cycle, such as feature engineering, training workflows and inference services; we are working on next-generation Marketplace Simulation... 
    Full time
    Remote work

    Waymo

    Remote
    17 hours ago
  •  ...spend worldwide. Using frontier causal inference-based econometric models to run experiments...  ...product managers, economists, and engineers from Google, Netflix, Meta, and Amazon,...  ...experience building and shipping production software systems ~ Must have strong Python proficiency... 
    Full time
    Work at office
    Work from home
    Worldwide
    Flexible hours

    Haus Analytics

    Remote
    17 hours ago
  • $104k - $130k

     ...paying for.  About the Role The New York Times is hiring a Software Engineer to join the New A.I. Products & Platforms mission. We are a...  ...within established backend infrastructure ~ Proficiency in GenAI‑assisted developer tooling (e.g., Cursor, Copilot, or Claude... 
    Full time
    Work at office
    Local area
    Flexible hours
    2 days per week

    The New York Times Company

    Remote
    17 hours ago
  •  ...practicing MDs, AI scientists, PhDs, creatives, technologists, and engineers working together to empower people and make care make more...  ...and LLM-driven workflows. What You’ll Do Design and build GenAI systems that turn LLMs into composable, dependable tools—leveraging... 
    Hourly pay
    Full time
    Flexible hours

    Abridge

    Remote
    17 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer - GenAI inference. Be the first to apply!