Software Engineer - GenAI inference
Databricks Inc.
P-1284
About This Role
As a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks’ Foundation Model API. You’ll work at the intersection of research and production, ensuring our large language model (LLM) serving systems are fast, scalable, and efficient. Your work will touch the full GenAI inference stack — from kernels and runtimes to orchestration and memory management.
What You Will Do
- Contribute to the design and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference
- Collaborate with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine
- Optimize for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators
- Build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations
- Develop and enhance scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads
- Support reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioning
- Integrate with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overhead
- Collaborate cross-functionally: with platform engineers, cloud infrastructure, and security/compliance teams
- Document and share learnings, contributing to internal best practices and open-source efforts when possible
What We Look For
- BS/MS/PhD in Computer Science, or a related field
- Strong software engineering background (3+ years or equivalent) in performance-critical systems
- Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc.
- Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)
- Comfortable designing and operating distributed systems, including RPC frameworks, queuing, RPC batching, sharding, memory partitioning
- Demonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)
- Experience building instrumentation, tracing, and profiling tools for ML models
- Ability to work closely with ML researchers, translate novel model ideas into production systems
- Ownership mindset and eagerness to dive deep into complex system challenges
- Bonus: published research or open-source contributions in ML systems, inference optimization, or model serving
$193.3k - $261.5k
...(AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on Amazon’s custom machine... ...JAX enabling unparalleled ML inference and training performance.The... ...-software boundary, our engineers build systematic infrastructure...SuggestedWork experience placementInternshipLocal areaFlexible hours$150k - $205k
...offices around the world, Aeris is the preeminent IoT software company globally powering critical projects across energy... ...to expand, we are seeking a experienced Software Engineer to join our Generative AI (GenAI) team. In this critical role, you will be responsible...SuggestedFull timeShift work$170k - $216k
...solutions to speed up developer velocity. We’re looking for a software engineer to join the team to build and maintain the critical data and... ...Software Engineer. You will: Develop Waymo's inference platform to make it scalable, high throughput, and low...SuggestedFull timeRemote work$2,000 per month
...architecture and design of the Sohu host software stack Implement high-performance,... ...handling continuous batching and real time inference Implement inference-time... ...person team in Cupertino, and greatly value engineering skills. We do not have boundaries between...SuggestedFull timeWork at officeRelocation package$190.9k - $232.8k
...P-1285 About This Role As a staff software engineer for GenAI inference, you will lead the architecture, development, and optimization of the inference engine that powers Databricks Foundation Model API.. You’ll bridge research advances and production demands, ensuring...SuggestedFull timeLocal areaWorldwide$92k - $135k
...intelligence that drives innovation. What You’ll Do: Join the Inference team to ship production features that improve latency,... ...practices, and grow quickly with mentorship from experienced engineers. About the role: Implement well-scoped features and...Permanent employmentFull timeTemporary workCasual workInternshipWork at officeRemote workFlexible hours$152k - $241.5k
NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and optimize the GPU-accelerated software that powers today’s most sophisticated AI applications. Our team is responsible...Full timeRemote work$152k - $241.5k
NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by contributing directly to upstream inference engines like vLLM and SGLang-ensuring they run best‑in...Full timeRemote work$160k - $240k
Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779 Description & Requirements Our team: Join the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI Inference...Temporary workFor contractorsWork experience placement$106.9k - $176.5k
...Sector - Technology Consulting - AI & Data - GenAI Developer - Senior ConsultantFrom... ...generative AI models, working with AI/ML engineers, developing web applications and ensuring... ...disciplinesExperience with vector databases and AI inference optimizationsExperience with open-source...For contractorsSummer holidayWork at officeLocal areaFlexible hours$170k - $216k
...products that evaluate the Waymo Driver's software stack at a massive scale. We solve... ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering... ...You will: Build and evolve ML inference infrastructure for simulations. Be responsible...Full timeRemote work$139.2k - $174k
...generation of AI-driven applications. We are seeking a Senior Engineer 2 to join our AI Inference Data Plane team. In this role, you will be a key... ...with gRPC. ~Proven experience shipping customer-facing software products and running critical services in a high-scale environment...Full timeRemote work$160k - $190k
...Sponsorship available*** About e360’s App Engineering e360 is a 30+ year privately-... .... What You’ll Do: Deliver GenAI solutions to customers as Professional... ...and evaluate the applicability of new software technologies to platform development efforts...Full timeRemote work$190.8k - $267.1k
...Learning teams. What You’ll Do: As a Senior Software Engineer, you will lead the development of a large-scale GenAI Platform at Reddit. Contribute to the... ...lifecycle. ~ Strong knowledge of model serving, inference pipelines, monitoring, and observability for...Full timeFor contractorsWork experience placementFlexible hours$150k - $180k
...operating posture. We embed engineers and leaders inside client... ...do: ~8+ years building software , a substantial share of it... ...it for a living. ~ Shipped GenAI/LLM systems to production — not... ...-tuning, distillation, or inference/serving optimization. Graph...Full timeWork at officeRemote workWorldwide$125.5k - $230.2k
...Sector - Technology Consulting - AI & Data - GenAI Developer - ManagerFrom strategy to... ...generative AI models, working with AI/ML engineers, developing web applications and... ...vector databases (Pinecone, Weaviate) and AI inference optimizations· Experience with open-source...For contractorsSummer holidayWork at officeLocal areaFlexible hours$175k - $220k
...Software Engineer, AI/ML GenAI Title of Role: Software Engineer, AI/ML GenAI Location: San Francisco, on-site or remote Company Stage of Funding: Venture Round Office Type: On-site or remote Salary: $175K–$220K Company Description We're representing...Work at officeRemote work- ...deliver top-notch technology products. As a Senior Lead Software Engineer at JPMorgan Chase within the Enterprise Technology - Public... ..., deployment, and ongoing optimization. Apply modern GenAI workflows, including prompt engineering techniques, tracing,...Full time
- ...frameworks. Strong track record of working with machine learning systems and/or platforms. Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between them. Experience serving fine-tuned LLMs (PEFT, DPO, RL...Full time
$120k - $140k
...intuitive low-code design. Join the Agentic Integration movement at snaplogic.com . The Role: We are looking for a Software Engineer to join our Agent Creator team, focusing on building and maintaining LLM integrations within the SnapLogic integration...Full timeWork experience placementImmediate start$152k - $241.5k
...company”.We are looking for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning & AI Compiler (DLC) team... ...of AI, our DLC has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices,...Full timeRemote work$200k - $250k
...experience, and we’re looking for a Senior MLOps Engineer to help us run them reliably and... ..., and scaling — across a custom-built inference platform powering a live conversational... ...For ~5–8+ years of experience in software, ML, platform, or infrastructure engineering...Full timeRemote workFlexible hours$152.2k - $243.7k
...influence how modern cloud and GenAI-powered platforms are... ...cloud technologies, platform engineering, and Generative AI enablement... ...safer, and smarter delivery of software at scale. This is a... ...model training, deployment, inference, and experimentation. Drive...Full timeWork experience placementWork at officeLocal areaWorldwideRelocation package- ...building cloud‑native AI data platforms that power analytics, machine learning, and GenAI solutions across global, multi‑industry programs. We are looking for a Senior AI/Data Platform Engineer with a strong focus on Google Cloud Platform to design, build, and operate...Full timeContract workWork at officeWork from home
$165.2k - $223.6k
...consistency of product identity and to infer relationships between products in Amazon... ...for an innovative and customer-focused software engineer to help us make the world's best product... ...traceability. You will pioneer advanced GenAI / Agentic solutions that power next-generation...Hourly payInternshipLocal areaWorldwideFlexible hours$320k
...growing group of committed researchers, engineers, policy experts, and business leaders working... .... About the role The Cloud Inference team scales and optimizes Claude to serve... ...qualifications Have significant software engineering experience, with a strong background...Full timeWork at officeVisa sponsorshipFlexible hours$170.6k - $261.3k
Job DescriptionAs a Senior Software Engineer on the SimCore team, you will build and deploy applied... ...with state-of-the-art multimodal GenAI and/or 3D reconstruction models, and excel... ...excel at building robust, high-performance inference pipelines. This role is not focused on...Full timeLocal areaWork from homeFlexible hours$188k - $275k
...at . What You'll Do Description of the team: The Inference team is responsible for delivering high-performance model serving... ...stack. About the role: We are looking for an Applied AI Engineer to help us understand, measure, and improve the real-world...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- Job Title: Sr. GenAI Cloud Engineer Location: North Chicago, IL 60064 Hybrid Assignment Duration: End of 2026 We’re hiring a Contractor to help Integrate AI/MI with Go/AI platform. You'll design cloud-native, event-driven systems that integrate Generative AI (LLMs),working...For contractorsRemote work
- ...can do with AI. We're looking for a Software Engineer who thrives at the intersection of systems... ...value. Enable next-generation GenAI workloads — Create infrastructure for multimodal... ..., real-time, near-real-time and batch inference, and asynchronous GPU pipelines....Hourly payFull timeImmediate startRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer - GenAI inference. Be the first to apply!
- agile software developer Remote
- software developer internship no experience Remote
- intermediate software engineer Remote
- software engineer staff Remote
- experienced software developer Remote
- work from home software developer Remote
- software developer no experience Remote
- software developer fintech Remote
- software data engineer Remote
- financial software developer Remote



