Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer - GenAI inference

$142.2k - $204.6k

Databricks Inc.

P-1284About This RoleAs a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks’ Foundation Model API. You’ll work at the intersection of research and production, ensuring our large language model (LLM) serving systems are fast, scalable, and efficient. Your work will touch the full GenAI inference stack — from kernels and runtimes to orchestration and memory management.What You Will DoContribute to the design and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inferenceCollaborate with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engineOptimize for latency, throughput, memory efficiency, and hardware utilization across GPUs, and acceleratorsBuild and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizationsDevelop and enhance scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloadsSupport reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioningIntegrate with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overheadCollaborate cross-functionally: with platform engineers, cloud infrastructure, and security/compliance teamsDocument and share learnings, contributing to internal best practices and open-source efforts when possibleWhat We Look ForBS/MS/PhD in Computer Science, or a related fieldStrong software engineering background (3+ years or equivalent) in performance-critical systemsSolid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc.Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)Comfortable designing and operating distributed systems, including RPC frameworks, queuing, RPC batching, sharding, memory partitioningDemonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)Experience building instrumentation, tracing, and profiling tools for ML modelsAbility to work closely with ML researchers, translate novel model ideas into production systemsOwnership mindset and eagerness to dive deep into complex system challengesBonus: published research or open-source contributions in ML systems, inference optimization, or model servingPay Range TransparencyDatabricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here.Local Pay Range$142,200—$204,600 USDAbout DatabricksDatabricks is the data and AI company. More than 10,000 organizations worldwide — including Comcast, Condé Nast, Grammarly, and over 50% of the Fortune 500 — rely on the Databricks Data Intelligence Platform to unify and democratize data, analytics and AI. Databricks is headquartered in San Francisco, with offices around the globe and was founded by the original creators of Lakehouse, Apache Spark, Delta Lake and MLflow. To learn more, follow Databricks on Twitter, LinkedIn and Facebook.BenefitsAt Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here.Our Commitment to Diversity and InclusionAt Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics.ComplianceIf access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Software Engineer - GenAI inference in San Francisco, CA vacancy
  • $300k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  .... About the role Our Inference team is responsible for building and maintaining...  ...fit if you: Have significant software engineering experience, particularly... 
    Suggested
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower...  ...via model inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...more. We focus on high-performance model inference and accelerating research through...  ...inference systems. In this role, you’ll lead engineering efforts to ensure our largest models...  ...performance issues across hardware and software layers. Have strong familiarity with... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $216.2k - $270.25k

     ....About Data EngineOur Generative AI Data Engine powers the world’s most advanced LLMs and...  ...opportunities across several teams within the GenAI Engineering organization, based on your...  ...improvementsRequirements:5+ years of software engineering experience, ideally in high-growth... 
    Suggested
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE...  ..., reliability, and ease of use. As a Software Engineer on the Inference Stack team,... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...About the Team OpenAI’s Inference team powers the deployment of our most advanced models...  ...world. We're a small, fast-moving team of engineers focused on delivering a world-class...  ...About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...serve OpenAI’s frontier models at massive scale. As part of the inference team, you’ll be responsible for unlocking every last FLOP from...  ...stack. About the Role We are looking for a kernel-focused engineer to lead efforts in writing, porting, and optimizing GPU... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...consistently fail. We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises, YC startups...  ...About the Role Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own profiling, batching, and autoscaling... 
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    Pulse

    San Francisco, CA
    1 day ago
  •  ...practicing MDs, AI scientists, PhDs, creatives, technologists, and engineers working together to empower people and make care make more...  ...Liberty in Pittsburgh. The Role We are looking for passionate GenAI Engineers of all levels who are passionate about making a... 
    Hourly pay
    Full time
    Work at office
    Relocation package
    Flexible hours

    Abridge

    San Francisco, CA
    1 day ago
  • $188k - $235k

     ...platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale...  ...candidate will have a strong understanding of software engineering principles and practices, as well... 
    Full time
    Shift work

    Scale Ai

    San Francisco, CA
    1 day ago
  •  ...About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks...  ..., analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Voice is becoming... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $320k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  ...the Role Our mandate is to make inference deployment boring and unattended....  ...deployment continuous and unattended. As a Software Engineer on the Launch Engineering team,... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    1 day ago
  • $190.9k - $232.8k

    P-1285About This RoleAs a staff software engineer for GenAI Performance and Kernel, you will own the design, implementation, optimization, and correctness...  ...of the high-performance GPU kernels powering our GenAI inference stack. You will lead development of highly-tuned, low-... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • $179.4k - $224.25k

     ...platform that provides APIs for knowledge retrieval, inference, evaluation, and more. We are looking for a strong engineer to join our team and help us build and scale...  ...candidate will have a strong understanding of software engineering principles and practices, as well... 
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  •  ...focus on performant and efficient model inference, as well as accelerating research progression...  ...About the Role We are looking for an engineer who wants to take the world's largest...  ...Have at least 3 years of professional software engineering experience. Have or can quickly... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $170k - $216k

     ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve...  ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering...  ...You will: Build and evolve ML inference infrastructure for simulations. Be responsible... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $175k - $220k

     ...Software Engineer, AI/ML GenAI Title of Role: Software Engineer, AI/ML GenAI Location: San Francisco, on-site or remote Company Stage of Funding: Venture Round Office Type: On-site or remote Salary: $175K-$220K Company Description We're representing... 
    Work at office
    Remote work

    Recruiting from Scratch

    San Francisco, CA
    5 days ago
  • $238k - $290k

     ...OverviewAs a Backend Platform Engineer at Harvey, you will help...  ...tackle challenges unique to GenAI-native applications — such as...  ...supporting high-throughput model inference, managing streaming and long-...  ...confidenceWhat You Have5+ years of software engineering experience (post-... 
    Flexible hours
    Shift work

    Harvey

    San Francisco, CA
    3 days ago
  • $189.2k - $372.9k

     ...Summary At Deloitte, Forward Deployed Engineers (FDE) don’t just build AI solutions,...  ...2026. Work you’ll do As a Lead Frontier GenAI FDE, you will serve as the senior...  ...integrated/verticalized sector solutions in software, data, AI, network, and hybrid cloud infrastructure... 
    Local area
    Visa sponsorship

    Deloitte

    San Francisco, CA
    2 days ago
  • $250k - $300k

     ...reliably in production. That means owning the inference stack end to end: profiling where time...  ...will also work directly with customer engineering teams to tailor deployments to their...  ...service.Build and support the software and product features around the inference... 
    Temporary work

    Crusoe

    San Francisco, CA
    20 hours ago
  • $300k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  .... About the Role The Cloud Inference team scales and optimizes Claude to serve...  ...Fit If You: Have significant software engineering experience, with a strong background... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  • $264.8k - $331k

     ...enterprise clients. As an ML Sys Research Engineer, you'll work on building out the...  ..., profile and optimize our training and inference framework. Post-train state of the art models...  ...-node LLM training and inference Strong software engineering skills, proficient in frameworks... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $125k - $160k

     ...Role Overview We are seeking a versatile Full Stack Software Engineer to join our engineering team. Reporting to the Software Engineering...  ...-Augmented Generation) architectures, or local model inference (Ollama). Experience in automated testing at multiple levels... 
    Full time
    Local area
    Visa sponsorship
    Work visa
    Shift work

    Cala Health

    San Francisco, CA
    1 day ago
  • About the TeamDoorDash’s GenAI Platform team sits within Machine...  ...weights model serving and batch inference, guardrails, and cost...  ...observability. This role is ideal for an engineer who enjoys building reliable...  ...of industry experience in software engineeringStrong backend... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    4 days ago
  • $150k - $180k

     ...Capital , and JFF Ventures , and are now hiring a Full Stack Engineer to help build the product that institutions use to interact...  ...application layer up to our data stack (Postgres + DuckDB) and model inference, and keep query and inference latency low enough that the... 
    Full time
    Work at office
    Immediate start

    Straia

    San Francisco, CA
    1 day ago
  • THE GLOBAL LEADER IN DATA & ANALYTICS RECRUITMENTHarnham Search and Selection Company Number: 05723485Harnham Search and Selection is a registered company in England and Wales. Reg no. 05723485Harnham Europe Limited Company Number: 09956940Harnham GmbH HRB: 196954Harnham...
    Work at office

    Harnham

    San Francisco, CA
    2 days ago
  • $151k - $204.3k

     ...leveraging the state-of-the-art AI/ML/GenAI tools on Amazon Web Service (...  ...Amazon.com’s recommendations engine is driven by machine learning...  ...the AI stack: 1) API-driven Inference Services like Amazon Bedrock...  ...closely with ISV(Independent Software Vendors) to enable large-... 
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    San Francisco, CA
    2 days ago
  • $350k

     .... About the Role We’re looking for an engineer to design, build, and operate the GPU supercomputing...  ...that powers large‑scale training and inference. You will deliver high‑performant,...  ...imaging, and capacity planning. Write software that abstracts cluster management and... 
    Full time
    Immediate start
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    17 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer - GenAI inference. Be the first to apply!