Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer - GenAI inference

Full-time

Databricks Inc.

P-1284

About This Role


As a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks’ Foundation Model API. You’ll work at the intersection of research and production, ensuring our large language model (LLM) serving systems are fast, scalable, and efficient. Your work will touch the full GenAI inference stack — from kernels and runtimes to orchestration and memory management.

What You Will Do



  • Contribute to the design and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference

  • Collaborate with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine

  • Optimize for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators

  • Build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations

  • Develop and enhance scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads

  • Support reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioning

  • Integrate with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overhead

  • Collaborate cross-functionally: with platform engineers, cloud infrastructure, and security/compliance teams

  • Document and share learnings, contributing to internal best practices and open-source efforts when possible

What We Look For



  • BS/MS/PhD in Computer Science, or a related field

  • Strong software engineering background (3+ years or equivalent) in performance-critical systems

  • Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc.

  • Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)

  • Comfortable designing and operating distributed systems, including RPC frameworks, queuing, RPC batching, sharding, memory partitioning

  • Demonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)

  • Experience building instrumentation, tracing, and profiling tools for ML models

  • Ability to work closely with ML researchers, translate novel model ideas into production systems

  • Ownership mindset and eagerness to dive deep into complex system challenges

  • Bonus: published research or open-source contributions in ML systems, inference optimization, or model serving

 

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Software Engineer - GenAI inference in Remote vacancy
  • $190.8k - $267.1k

     ...Learning teams. What You’ll Do: As a Senior Software Engineer, you will lead the development of a large-scale GenAI Platform at Reddit. Contribute to the...  ...lifecycle. ~ Strong knowledge of model serving, inference pipelines, monitoring, and observability for... 
    Suggested
    Full time
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    Remote
    a month ago
  • $165.2k - $223.6k

     ...(AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on Amazon’s custom machine...  ...JAX enabling unparalleled ML inference and training performance.The...  ...-software boundary, our engineers build systematic infrastructure... 
    Suggested
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $8k

     ...Secret (TS/SCI) clearance with polygraph is required.  Visionist has an exciting new, fully FUNDED opportunity for a Software Engineer - Inference on our largest PRIME contract. Our team of Analysts and Engineers is motivated by the direct impact on the mission, crafting... 
    Suggested
    Permanent employment
    Full time
    Contract work
    Temporary work
    Immediate start
    Flexible hours

    Visionist, Inc.

    Remote
    1 day ago
  • $2,000 per month

     ...architecture and design of the Sohu host software stack Implement high-performance,...  ...handling continuous batching and real time inference Implement inference-time...  ...person team in Cupertino, and greatly value engineering skills. We do not have boundaries between... 
    Suggested
    Full time
    Work at office
    Relocation package

    Etched

    Remote
    1 day ago
  • $120.4k - $198.7k

     ...What Is the Opportunity? Travelers is seeking a Software Engineer II to join our organization as we grow and transform our Technology...  ...Experience with GitHub Experience with AI, GenAI Delivery - Intermediate delivery skills including the ability... 
    Suggested
    Full time
    Work experience placement
    Local area
    Immediate start

    The Travelers Indemnity Company

    Remote
    1 day ago
  • $140k - $150k

     ...Fitch Solutions is currently seeking a Senior Software Engineer, AI based out of our Chicago office. Fitch Solutions is a leading provider...  ...Lead enterprise-wide AI standardization initiatives – Build GenAI solutions across multiple business units while creating unified... 
    Full time
    Temporary work
    Apprenticeship
    Work at office
    Local area
    Immediate start
    2 days per week

    Fitch Group

    Remote
    1 day ago
  •  ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application software team is the central nervous system of SpaceX - we create mission critical... 
    Permanent employment
    Temporary work
    Remote work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    3 days ago
  • $160k - $240k

    Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779 Description & Requirements Our team: Join the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI Inference... 
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    5 hours ago
  • $152k - $241.5k

    NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and optimize the GPU-accelerated software that powers today’s most sophisticated AI applications. Our team is responsible... 
    Full time
    Remote work

    Nvidia

    Texas
    1 day ago
  • $106.9k - $176.5k

     ...Sector - Technology Consulting - AI & Data - GenAI Developer - Senior ConsultantFrom...  ...generative AI models, working with AI/ML engineers, developing web applications and ensuring...  ...disciplinesExperience with vector databases and AI inference optimizationsExperience with open-source... 
    For contractors
    Summer holiday
    Work at office
    Local area
    Flexible hours

    EY (Ernst & Young)

    McLean, VA
    2 days ago
  •  ...help build and scale infrastructure for inference, evaluation, and continuous model...  ...with a strong blend of infrastructure engineering, production ML systems, LLM/LVM inference...  ...background, especially in Python, with solid software engineering fundamentals. ~... 
    Full time
    Work at office
    Local area
    Flexible hours
    3 days per week

    Ambient.ai

    Remote
    1 day ago
  • $170k - $216k

     ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve...  ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering...  ...You will: Build and evolve ML inference infrastructure for simulations. Be responsible... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $165k - $242k

     ...degradation, rollback/traffic-shift strategies. Mentor IC1/IC2 engineers; review cross-team designs and elevate coding/testing...  ...(Prometheus, Grafana, OpenTelemetry). Practical knowledge of inference internals: batching, caching, mixed precision (BF16/FP8), streaming... 
    Permanent employment
    Temporary work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Bellevue, WA
    5 days ago
  • $150k - $180k

     ...operating posture.   We embed engineers and leaders inside client...  ...do: ~8+ years building software , a substantial share of it...  ...it for a living. ~ Shipped GenAI/LLM systems to production — not...  ...-tuning, distillation, or inference/serving optimization. Graph... 
    Full time
    Work at office
    Remote work
    Worldwide

    Provectus, Inc.

    Remote
    1 day ago
  • $139k - $204k

    What You’ll Do Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements to...  ...orchestration, and hardware teams to evolve our Kubernetes‑native inference platform and meet strict P99 SLAs at scale. About The Role... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Bellevue, WA
    5 days ago
  • $165k - $242k

    Apply for the Senior Software Engineer II, Inference role at CoreWeave. CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    5 days ago
  • $175k - $220k

     ...Software Engineer, AI/ML GenAI Title of Role: Software Engineer, AI/ML GenAI Location: San Francisco, on-site or remote Company Stage of Funding: Venture Round Office Type: On-site or remote Salary: $175K-$220K Company Description We're representing... 
    Work at office
    Remote work

    Recruiting from Scratch

    San Francisco, CA
    2 days ago
  • $125.5k - $230.2k

     ...Sector - Technology Consulting - AI & Data - GenAI Developer - ManagerFrom strategy to...  ...generative AI models, working with AI/ML engineers, developing web applications and...  ...vector databases (Pinecone, Weaviate) and AI inference optimizations· Experience with open-source... 
    For contractors
    Summer holiday
    Work at office
    Local area
    Flexible hours

    EY (Ernst & Young)

    McLean, VA
    2 days ago
  •  ...deliver top-notch technology products. As a Senior Lead Software Engineer at JPMorgan Chase within the Enterprise Technology - Public...  ..., deployment, and ongoing optimization. Apply modern GenAI workflows, including prompt engineering techniques, tracing,... 
    Full time

    JPMorgan Chase & Co.

    Remote
    1 day ago
  •  ...frameworks. Strong track record of working with machine learning systems and/or platforms. Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between them. Experience serving fine-tuned LLMs (PEFT, DPO, RL... 
    Full time

    Snowflake

    Remote
    1 day ago
  • $120k - $140k

     ...intuitive low-code design. Join the Agentic Integration movement at snaplogic.com . The Role:  We are looking for a Software Engineer to join our Agent Creator team, focusing on building and maintaining LLM integrations within the SnapLogic integration... 
    Full time
    Work experience placement
    Immediate start

    Snaplogic

    Remote
    1 day ago
  • $152k - $241.5k

     ...company”.We are looking for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning & AI Compiler (DLC) team...  ...of AI, our DLC has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices,... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 hours ago
  • Reddit, Inc. is seeking a Senior Software Engineer for the Ads Creative Effectiveness team to build GenAI and predictive products for ad creative at scale. You will lead architecture for high-performance backends, design editing pipelines for images/videos, and orchestrate... 
    Remote job

    Reddit, Inc.

    New York, NY
    2 days ago
  • $119.8k - $234.7k

     ...ContributorTravel: Less than 25%Profession: Software EngineeringDiscipline:...  ....As a Senior Software Engineer, you will design and deliver...  ...generative artificial intelligence [GenAI], approaches to source code...  ...of machine learning inference, graph compilation, memory management... 
    Ongoing contract
    Local area
    Remote work
    Flexible hours
    3 days per week

    Microsoft

    Redmond, WA
    1 day ago
  •  ...organization, apply now.We are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina (US-NC),...  ...-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an... 
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Charlotte, NC
    2 days ago
  • $200k - $250k

     ...experience, and we’re looking for a Senior MLOps Engineer to help us run them reliably and...  ..., and scaling — across a custom-built inference platform powering a live conversational...  ...For ~5–8+ years of experience in software, ML, platform, or infrastructure engineering... 
    Full time
    Remote work
    Flexible hours

    Wizard

    United States
    1 day ago
  • $152.2k - $243.7k

     ...influence how modern cloud and GenAI-powered platforms are...  ...cloud technologies, platform engineering, and Generative AI enablement...  ...safer, and smarter delivery of software at scale. This is a...  ...model training, deployment, inference, and experimentation. Drive... 
    Full time
    Work experience placement
    Work at office
    Local area
    Worldwide
    Relocation package

    Visa

    Remote
    1 day ago
  • $165.2k - $223.6k

     ...consistency of product identity and to infer relationships between products in Amazon...  ...for an innovative and customer-focused software engineer to help us make the world's best product...  ...traceability. You will pioneer advanced GenAI / Agentic solutions that power next-generation... 
    Hourly pay
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon

    Sunnyvale, CA
    5 hours ago
  • $320k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  .... About the role The Cloud Inference team scales and optimizes Claude to serve...  ...qualifications Have significant software engineering experience, with a strong background... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Remote
    1 day ago
  • $125.5k - $230.2k

     ...Sector – Technology Consulting - AI & Data - GenAI Developer – Manager From strategy to...  ...AI models, working with AI/ML engineers, developing web applications and ensuring...  ...vector databases (Pinecone, Weaviate) and AI inference optimizations · Experience with open-source... 
    For contractors
    Summer holiday
    Work at office
    Local area
    Flexible hours

    Ernst & Young

    McLean, VA
    27 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer - GenAI inference. Be the first to apply!