Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer - GenAI inference

$190.9k - $232.8k

Databricks Inc.

P-1285About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture, development, and optimization of the inference engine that powers Databricks Foundation Model API.. You’ll bridge research advances and production demands, ensuring high throughput, low latency, and robust scaling. Your work will encompass the full GenAI inference stack: kernels, runtimes, orchestration, memory, and integration with frameworks and orchestration systems.What You Will DoOwn and drive the architecture, design, and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inferencePartner closely with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engineLead the end-to-end optimization for latency, throughput, memory efficiency, and hardware utilization across GPUs, and acceleratorsDefine and guide standards to build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizationsArchitect scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloadsEnsure reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioningCollaborate cross-functionally on Integrating with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overheadDrive cross-team collaboration: with platform engineers, cloud infrastructure, and security/compliance teamsRepresent the team externally through benchmarks, whitepapers, and open-source contributionsWhat We Look ForBS/MS/PhD in Computer Science, or a related fieldStrong software engineering background (6+ years or equivalent) in performance-critical systemsProven track record of owning complex system components and driving architectural decisions end-to-endDeep understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc.Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)Strong background in distributed systems design, including RPC frameworks, queuing, RPC batching, sharding, memory partitioningDemonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)Experience building instrumentation, tracing, and profiling tools for ML modelsAbility to lead through influence - work closely with ML researchers, translate novel model ideas into production systemsExcellent communication and leadership skills, with a proactive and ownership-driven mindsetBonus: published research or open-source contributions in ML systems, inference optimization, or model servingPay Range TransparencyDatabricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here.Local Pay Range$190,900—$232,800 USDAbout DatabricksDatabricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and 70% of the Fortune 500 — rely on the Databricks Data + AI Platform to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified platform that includes Genie, Lakebase, Agent Bricks, Lakeflow, Lakehouse, and Unity Catalog. To learn more, follow Databricks on LinkedIn, X, YouTube, and Instagram.BenefitsAt Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here.Our Commitment to Diversity and InclusionAt Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics.ComplianceIf access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer - GenAI inference in San Francisco, CA vacancy
  • $190.9k - $232.8k

     ...P-1285About This RoleAs a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness...  ...of the high-performance GPU kernels powering our GenAI inference stack. You will lead development of highly-tuned, low-... 
    Suggested
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    a month ago
  • $252k - $315k

     ...platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale...  ...candidate will have a strong understanding of software engineering principles and practices, as well... 
    Suggested
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $190k - $265k

     ...use deep data insights to improve their business. Founded by engineers — and customer-obsessed — we leap at every opportunity to solve...  ...infrastructure that power the next generation of AI.The Foundation Model Inference team is the backbone of Databricks’ generative AI capabilities... 
    Suggested
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (GenAI Platform Services) Overview: At Capital One, we are creating responsible...  ..., develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    4 days ago
  •  ...future. DataRobot’s Fleet team is the engine behind how our platform runs across...  ...velocity. That’s where you come in. As a Staff Software Engineer, you’ll be responsible for...  ...with GPU infrastructure for training and inference. Why Join the Fleet Management team?... 
    Suggested
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    DataRobot

    San Francisco, CA
    3 days ago
  • $231k - $314.35k

     ...- and we're just getting started. Role Overview As a Software Engineer on the Product Engineering team at Harvey, you will own and lead...  ...excited about building the future of application layer and genAI products. This role is based in San Francisco, CA. We use... 
    Relocation package

    Harvey

    San Francisco, CA
    4 days ago
  • $200k - $300k

     ...Staff Software Engineer F2 is redefining how the financial sector operates by bridging the gap between legacy workflows in institutional finance...  ...(Temporal), streaming pipelines, vector search, and LLM inference paths; balancing latency, quality, and cost. Build for... 
    Full time

    F2

    San Francisco, CA
    21 hours ago
  •  ...Staff Software Engineer Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley...  ...Points Production experience with AI/LLM systems — inference pipelines, evaluation workflows, model integration, or AI-... 
    Permanent employment
    Work at office

    Arena AI

    San Francisco, CA
    1 day ago
  •  ...Staff Engineer Lambda, the superintelligence cloud, is a leader in AI cloud infrastructure...  ...the next generation of AI training and inference at scale. As a Staff Engineer on our...  ...Qualifications ~10+ years of experience in software engineering, platform engineering, or... 
    Work at office
    Immediate start
    Work from home

    Lambda

    San Francisco, CA
    4 days ago
  •  ...our growing team. About the Role Plenful is hiring a Staff Software Engineer to lead the design and development of systems that power...  ...with ML/AI teams on data pipelines for model training and inference. Technical Leadership Lead technical projects and... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    2 days ago
  • $180k - $225k

     ....About Data EngineOur Generative AI Data Engine powers the world’s most advanced LLMs and...  ...opportunities across several teams within the GenAI Engineering organization, based on your...  ...improvementsRequirements:5+ years of software engineering experience, ideally in high-growth... 
    Full time

    Scale AI

    San Francisco, CA
    a month ago
  • $179.4k - $224.25k

     ...platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale...  ...candidate will have a strong understanding of software engineering principles and practices, as well... 
    Full time

    Scale AI

    San Francisco, CA
    a month ago
  • $264.66k - $369.8k

     ...high-quality frontend applications that bring those systems to life.Work with other AI engineers, software engineers and machine learning engineers to architect, design and implement GenAI-powered products and featuresCollaborate across functions to understand user needs,... 
    Full time
    Work experience placement
    Work at office
    Local area
    Shift work

    Plaid Financial

    San Francisco, CA
    a month ago
  • $207k - $300k

     ...fresh in real-time as YouTube’s data schemas evolve.Mentor Senior Engineers, drive technical roadmap planning, and collaborate with...  ...or equivalent practical experience. 8 years of experience in software development.5 years of experience testing, and launching software... 

    Google

    San Bruno, CA
    29 days ago
  • $231k - $340k

     ...'re just getting started. Role Overview As a Backend Software Engineer on the Product Engineering team at Harvey, you will own and lead...  ...excited about building the future of application layer and genAI products. We use an in-person work model and offer... 
    Relocation package

    Harvey

    San Francisco, CA
    2 days ago
  •  ...reliable on-demand, logistics engine for last-mile retail delivery...  ...to join our team. As a Staff Machine Learning Engineer, you...  ..., Codex, Cursor) in the full software development lifecycle, including...  ...fieldExpertise in applied ML for Causal Inference and Recommendation Systems -... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    a month ago
  • $262k - $329k

     ...the phone in someone's hand, the native engine underneath, on-device ML, and the backend pipelines behind all of it. As a Staff Software Engineer on the Video Performance team,...  ...device ML performance work, such as tuning inference latency with CoreML, TFLite or TensorRT,... 
    Full time
    Live in
    Work at office
    Local area
    Flexible hours

    Canva

    San Francisco, CA
    4 days ago
  • $210k - $300k

     ...hardcore and obsessed team of the world's best engineers and operators. If you are obsessed with...  ...About the Role We're looking for a Software Engineer to join our ML Infrastructure...  ..., you'll help build the training and inference systems that power our general-purpose warehouse... 
    Local area
    Flexible hours

    Nimble Robotics

    San Francisco, CA
    3 days ago
  • $192k - $260k

     ...proprietary large language models. It offers real-time, low-latency inference, governance, monitoring, and lineage. As AI adoption...  ...operationalize models at scale with strong SLAs and cost efficiency.As a Staff Engineer, you’ll play a critical role in shaping both the product... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    a month ago
  •  ...data centers. About the Role At Watney, ML Infrastructure engineers turn data collected from a live fleet of robots into better...  ...optimal GPU utilization. What You’ll Do Own training and inference infrastructure Build the data pipelines that these training... 

    Watney

    San Francisco, CA
    25 days ago
  • $192k - $260k

     ...for hosting and serving frontier AI model inference for open source models like Llama, Qwen,...  ...is necessary. We’re looking for engineers who have owned high scale operational sensitive...  ...LLM APIs and runtimes at scale.As a Staff Engineer, you’ll play a critical role in... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    a month ago
  • $237.6k - $318.24k

     ...other, come build with us at Crusoe.About This Role:The Senior Staff Software Engineer for the AI Model Lifecycle team will play a crucial role in...  ...frameworks.Performance optimizations on GPU systems and inference frameworks.Benefits:Competitive compensationRestricted... 
    Temporary work

    Crusoe

    San Francisco, CA
    a month ago
  •  ...founded in 2024 by a team of former Scale AI engineers and operators. In less than a year, we'...  ...models. About This Role As a Staff Software Engineer, Platform at David AI, you'll...  ...of audio or video data. Scaled up inference and train compute for large scale... 
    Work at office

    David AI

    San Francisco, CA
    2 days ago
  • $215k - $260k

     ...reliably in production. That means owning the inference stack end to end: profiling where time...  ...will also work directly with customer engineering teams to tailor deployments to their...  ...service. Build and support the software and product features around the inference... 
    Temporary work

    Crusoe

    San Francisco, CA
    25 days ago
  •  ...processing solutions for production training and inference workloads. Solve problems involving...  ...of experience building production-grade software, infrastructure, or developer-facing systems, with strong Python engineering experience. At least 6 years of personally... 
    Full time
    Flexible hours

    Anyscale

    San Francisco, CA
    10 days ago
  • $193.93k - $352.29k

     ...investors. About the Role We are looking for a Senior/Staff Software Engineer to serve as a technical leader for Nuro's ML Data engine. You...  ...methods. E.g. build systems that compute embeddings or run inference at scale, manage vector databases, and automatically sample... 
    Immediate start
    Flexible hours
    Shift work

    Nuro

    San Francisco, CA
    17 days ago
  • $250k - $300k

     ...us at Crusoe. About the Role: We are seeking Sr. Staff Software Engineers to serve as the Managed Platform Services (MAPS) organization...  ...: Deep dive AI-native customer profiles (training and inference) to identify what features and intelligent insights — both... 
    Permanent employment
    Temporary work
    Shift work

    Crusoe

    San Francisco, CA
    25 days ago
  • $190.9k - $334.1k

     ...It all started when engineer Fred Luddy wrote code that automated a tedious task for...  ...Qualifications ~8+ years of software engineering with strong fundamentals in...  ...retrieval. Exposure to LLM fine-tuning or inference optimization in production.   Why join... 
    Full time
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    San Francisco, CA
    5 days ago
  •  ...Calling All Full-Stack Software Engineers We're looking for an exceptional full-stack engineer to join us as an early team member. Truewind is changing the way accountants and financial analysts work by building AI-native workflows. We're developing cutting-edge... 

    Truewind

    San Francisco, CA
    3 days ago
  • $200k - $250k

     ...About Us At 3Y Health, we are building AI-driven software to empower healthcare providers and solve the overwhelming administrative complexity...  ..., and 8VC. About the Role We are seeking a Frontend Engineer to help us craft intuitive, high-performance user interfaces... 
    Work experience placement
    Private practice
    Work at office

    3Y Health

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer - GenAI inference. Be the first to apply!