Staff Software Engineer - GenAI inference
$190.9k - $232.8kDatabricks Inc.
P-1285About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture, development, and optimization of the inference engine that powers Databricks Foundation Model API.. You’ll bridge research advances and production demands, ensuring high throughput, low latency, and robust scaling. Your work will encompass the full GenAI inference stack: kernels, runtimes, orchestration, memory, and integration with frameworks and orchestration systems.What You Will DoOwn and drive the architecture, design, and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inferencePartner closely with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engineLead the end-to-end optimization for latency, throughput, memory efficiency, and hardware utilization across GPUs, and acceleratorsDefine and guide standards to build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizationsArchitect scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloadsEnsure reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioningCollaborate cross-functionally on Integrating with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overheadDrive cross-team collaboration: with platform engineers, cloud infrastructure, and security/compliance teamsRepresent the team externally through benchmarks, whitepapers, and open-source contributionsWhat We Look ForBS/MS/PhD in Computer Science, or a related fieldStrong software engineering background (6+ years or equivalent) in performance-critical systemsProven track record of owning complex system components and driving architectural decisions end-to-endDeep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc.Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)Strong background in distributed systems design, including RPC frameworks, queuing, RPC batching, sharding, memory partitioningDemonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)Experience building instrumentation, tracing, and profiling tools for ML modelsAbility to lead through influence - work closely with ML researchers, translate novel model ideas into production systemsExcellent communication and leadership skills, with a proactive and ownership-driven mindsetBonus: published research or open-source contributions in ML systems, inference optimization, or model servingPay Range TransparencyDatabricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here.Local Pay Range$190,900—$232,800 USDAbout DatabricksDatabricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog. To learn more, follow Databricks on LinkedIn, X, YouTube, and Instagram.BenefitsAt Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here.Our Commitment to Diversity and InclusionAt Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics.ComplianceIf access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.
$190.9k - $232.8k
...P-1285About This RoleAs a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness... ...of the high-performance GPU kernels powering our GenAI inference stack. You will lead development of highly-tuned, low-...SuggestedLocal areaWorldwide$252k - $315k
...platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale... ...candidate will have a strong understanding of software engineering principles and practices, as well...SuggestedFull time- ...About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re...SuggestedFull time
- ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI... ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE... ..., reliability, and ease of use. As a Software Engineer on the Inference Stack team,...SuggestedFull timeFlexible hours
- ...About the Team OpenAI’s Inference team powers the deployment of our most advanced models... ...world. We're a small, fast-moving team of engineers focused on delivering a world-class... ...About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal...SuggestedFull time
- ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Voice is becoming...Full timeFlexible hours
- ...About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower... ...via model inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across...Full time
- ...About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks... ..., analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About...Full time
$190k - $265k
...use deep data insights to improve their business. Founded by engineers — and customer-obsessed — we leap at every opportunity to solve... ...infrastructure that power the next generation of AI.The Foundation Model Inference team is the backbone of Databricks’ generative AI capabilities...Local areaWorldwide$188k - $282k
...beyond, this team often provides the first line of defense for Auth0 customers. The Staff Software Engineer At Okta, we’re building the next generation of authentication for the GenAI era. We’re looking for a Staff Software Engineer to join the AI DevEx team at...Full timeLocal areaWorldwideFlexible hours- ...About the Team Our Inference team brings OpenAI’s most capable research and technology to... ...About the Role We are looking for an engineer who wants to take the world's largest and... ...Have at least 5 years of professional software engineering experience. Have or can quickly...Full time
$170k - $216k
...products that evaluate the Waymo Driver's software stack at a massive scale. We solve... ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering... ...You will: Build and evolve ML inference infrastructure for simulations. Be responsible...Full timeRemote work$229.9k - $262.4k
Senior Lead AI Engineer (GenAI Platform Services) Overview: At Capital One, we are creating responsible... ..., develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model...Full timePart timeLocal area$240k - $280k
...About the Role We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software... ...its full lifecycle — turning racks of GPUs into running inference clusters without a human touching a runbook. The Research...Full time$200k - $300k
...Staff Software Engineer F2 is redefining how the financial sector operates by bridging the gap between legacy workflows in institutional finance... ...(Temporal), streaming pipelines, vector search, and LLM inference paths; balancing latency, quality, and cost. Build for...Full time- ...Staff Engineer Lambda, the superintelligence cloud, is a leader in AI cloud infrastructure... ...the next generation of AI training and inference at scale. As a Staff Engineer on our... ...Qualifications ~10+ years of experience in software engineering, platform engineering, or...Work at officeImmediate startWork from home
- ...Staff Software Engineer Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley... ...Points Production experience with AI/LLM systems — inference pipelines, evaluation workflows, model integration, or AI-...Permanent employmentWork at office
- ...our growing team. About the Role Plenful is hiring a Staff Software Engineer to lead the design and development of systems that power... ...with ML/AI teams on data pipelines for model training and inference. Technical Leadership Lead technical projects and...Full timeWork at officeRemote workFlexible hours2 days per week
$176.5k
...content that powers our products and our inference-driven features. The core of the... ...join an existing team of Senior, Staff, and Principal engineers building the Content Library and... ...a systems engineer who loves owning software end-to-end: design, implementation,...Full timeFor contractorsLocal areaHome officeFlexible hours- ...future. DataRobot’s Fleet team is the engine behind how our platform runs across... ...velocity. That’s where you come in. As a Staff Software Engineer, you’ll be responsible for... ...with GPU infrastructure for training and inference. Why Join the Fleet Management team?...Full timeLocal areaRemote workWorldwideFlexible hours
$231k - $314.35k
...- and we're just getting started. Role Overview As a Software Engineer on the Product Engineering team at Harvey, you will own and lead... ...excited about building the future of application layer and genAI products. This role is based in San Francisco, CA. We use...Relocation package$180k - $225k
....About Data EngineOur Generative AI Data Engine powers the world’s most advanced LLMs and... ...opportunities across several teams within the GenAI Engineering organization, based on your... ...improvementsRequirements:5+ years of software engineering experience, ideally in high-growth...Full time$265k - $295k
...educational institutions Work together with engineers, scientists, operators, and more from... ...scale. About the Role As a Staff Software Engineer on the Consumer Experience team... ...entity resolution and real-time inference Experience building AI-powered systems...Full timeWork at officeRemote workFlexible hours$262k - $329k
...the phone in someone's hand, the native engine underneath, on-device ML, and the backend pipelines behind all of it. As a Staff Software Engineer on the Video Performance team,... ...device ML performance work, such as tuning inference latency with CoreML, TFLite or TensorRT,...Full timeLive inWork at officeLocal areaFlexible hours$207k - $300k
...fresh in real-time as YouTube’s data schemas evolve.Mentor Senior Engineers, drive technical roadmap planning, and collaborate with... ...or equivalent practical experience. 8 years of experience in software development.5 years of experience testing, and launching software...$231k - $340k
...'re just getting started. Role Overview As a Backend Software Engineer on the Product Engineering team at Harvey, you will own and lead... ...excited about building the future of application layer and genAI products. We use an in-person work model and offer...Relocation package$210k - $300k
...hardcore and obsessed team of the world's best engineers and operators. If you are obsessed with... ...About the Role We're looking for a Software Engineer to join our ML Infrastructure... ..., you'll help build the training and inference systems that power our general-purpose warehouse...Local areaFlexible hours$237.6k - $318.24k
...build with us at Crusoe. About This Role: The Senior Staff Software Engineer for the AI Model Lifecycle team will play a crucial role in... .... Performance optimizations on GPU systems and inference frameworks. Benefits: ~ Competitive compensation ~...Temporary work$192k - $260k
...for hosting and serving frontier AI model inference for open source models like Llama, Qwen,... ...is necessary. We’re looking for engineers who have owned high scale operational sensitive... ...LLM APIs and runtimes at scale.As a Staff Engineer, you’ll play a critical role in...Local areaWorldwide- ...data centers. About the Role At Watney, ML Infrastructure engineers turn data collected from a live fleet of robots into better... ...optimal GPU utilization. What You’ll Do Own training and inference infrastructure Build the data pipelines that these training...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Software Engineer - GenAI inference. Be the first to apply!
- software sales representative San Francisco, CA
- embedded software San Francisco, CA
- software applications developer San Francisco, CA
- entry level software sales San Francisco, CA
- software technology San Francisco, CA
- software implementation project manager San Francisco, CA
- software support San Francisco, CA
- government software San Francisco, CA
- software technical writer San Francisco, CA
- software engineer - cloud services San Francisco, CA




