Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Engineer, Inference Optimizations

$191.2k - $239k
Full-time

DigitalOcean

Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you’ll find your place here. We value winning together—while learning, having fun, and making a profound difference for the dreamers and builders in the world.

DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be the Inference Cloud of choice for digitally native companies and you will help ensure we can offer the industry-leading performance for our inference services. You will be responsible for the architectural decisions that maximize throughput and minimize latency for the world’s most advanced large models. As an IC leader, you will act as a force multiplier for the engineering organization, solving the most complex bottlenecks in memory bandwidth and compute utilization while guiding the technical roadmap for our high-performance inference fleet.

What You’ll Do:

  • Performance Architecture: Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers, ensuring our infrastructure extracts maximum value from every TFLOP.
  • Deep-Dive Optimization: Engineer solutions for complex performance issues, including attention layer optimizations, memory and precision management, and advanced parallelization across multi-node GPU clusters.
  • Technological Innovation: Proactively implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of the Gen AI landscape. Some examples of projects you may work on:
    • Improving batch size performance using AMD's AITER library for AMD MI355X - identify and tune AITER's CK (composable kernel) or ASK (assembly) to optimize FP8 / BF16
    • Identify kernel fusion opportunities for GLM-5 kernels for different layers of the Transformer block (FlashAttention, RMS Norm)
    • Tune expert gateway router kernels for MoE models like Qwen3-235B, DeepSeek V3, GLM-5 etc
  • Hardware & Ecosystem Mastery: Act as the subject matter expert on modern GPU families (NVIDIA/AMD) and their software stacks (CUDA, ROCm, TensorRT, OpenAI Triton), advising on hardware procurement and software integration.
  • Precision Optimization: Develop and deploy state-of-the-art quantization techniques (FP8, INT8, and experimental FP4) to double throughput without losing accuracy.
  • Technical Mentorship: Lead by example through high-quality code and design reviews, elevating the technical bar for the team without the administrative overhead of direct management.
  • Strategic Collaboration: Partner with Product Management and TPMs to translate "theoretical hardware limits" into "shippable product features," ensuring our platform is both powerful and developer-friendly.
  • Community Leadership: Maintain a strong presence in the GPU infrastructure and model performance optimization communities, contributing to and integrating the best of open-source AI.

What You’ll Bring to DigitalOcean:

  • Technical Depth: 5+ years of experience in high-performance computing or AI infrastructure, with a proven track record of solving compute utilization and memory bandwidth bottlenecks.
  • Gen AI Literacy: Deep familiarity with the Gen AI (LLM, VLM, LMM) landscape, including the specific quirks and architectural requirements of major model families.
  • Optimization Expert: Hands-on experience with attention-layer optimizations and parallelization strategies across distributed GPU environments.
  • Hardware Fluency: Comprehensive understanding of NVIDIA and AMD GPU architectures and their respective software ecosystems (CUDA, ROCm, etc.).
  • Open Source Mastery: Extensive experience integrating, building with, and contributing to open-source software projects.
  • Systems Design: Excellent system design skills, particularly related to low-level GPU programming - optimization, memory access patterns, and parallel execution.
  • Leadership through Influence: Experience acting as a technical lead, driving design and delivery through cross-functional alignment and expert-level delegation.
  • Low-Level Mastery: Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores).
  • The Toolkit: Expert-level Triton or CUDA. If you’ve contributed to the Triton compiler or wrote custom CUDA kernels for a major LLM, we want you.

Compensation Range:

  • $191,200 - $239,000

*This is a hybrid role

JR: 2026-7625

#LI-Hybrid

Why You’ll Like Working for DigitalOcean

  • We innovate with purpose. You’ll be a part of a cutting-edge technology company with an upward trajectory, who are proud to simplify cloud and AI so builders can spend more time creating software that changes the world. As a member of the team, you will be a Shark who thinks big, bold, and scrappy, like an owner with a bias for action and a powerful sense of responsibility for customers, products, employees, and decisions.
  • We prioritize career development. At DO, you’ll do the best work of your career. You will work with some of the smartest and most interesting people in the industry. We are a high-performance organization that will always challenge you to think big. Our organizational development team will provide you with resources to ensure you keep growing. We provide employees with reimbursement for relevant conferences, training, and education. All employees have access to LinkedIn Learning's 10,000+ courses to support their continued growth and development.
  • We care about your well-being. Regardless of your location, we will provide you with a competitive array of benefits to support you from our Employee Assistance Program to Local Employee Meetups to flexible time off policy, to name a few. While the philosophy around our benefits is the same worldwide, specific benefits may vary based on local regulations and preferences.
  • We reward our employees. The salary range for this position is based on market data, relevant years of experience, and skills. You may qualify for a bonus in addition to base salary; bonus amounts are determined based on company and individual performance. We also provide equity compensation to eligible employees, including equity grants upon hire and the option to participate in our Employee Stock Purchase Program.
  • DigitalOcean is an equal-opportunity employer. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service.

Application Limit: You may apply to a maximum of 3 positions within any 180-day period. This policy promotes better role-candidate matching and encourages thoughtful applications where your qualifications align most strongly.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff Engineer, Inference Optimizations in Seattle, WA vacancy
  • $203.5k - $299.3k

     ...RoleWe are hiring a Causal Machine Learning Engineer to help build the causal ML foundation...  ...policy evaluation, promotion optimization, or marketplace decisioning systems.You...  ...…Deep practical experience with causal inference, econometrics, experimentation, or causal... 
    Suggested
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    Seattle, WA
    6 hours ago
  • $229k - $343k

     ...infrastructure, and on-device and server-side inference. Our team creates intuitive tools,...  ....We’re looking for a Machine Learning Engineer to join Snap Inc!What you’ll do:Develop...  ...of GPU, CPU, or NPU optimization techniquesExperience building and optimizing... 
    Suggested
    Full time
    Live in
    Work at office
    Local area
    Worldwide

    Snap

    Seattle, WA
    4 days ago
  • Nuance Labs in Seattle is seeking a Member of Technical Staff focused on model optimization and inference. This role demands expertise in refining AI models for real-time interactions, requiring a strong foundation in ML systems and familiarity with frameworks like vLLM... 
    Suggested

    Nuance Labs

    Seattle, WA
    3 days ago
  • $203.5k - $299.3k

     ...are hiring a Causal Machine Learning Engineer to help build the causal ML foundation...  ...counterfactual policy evaluation, promotion optimization, or marketplace decisioning systems....  ...Deep practical experience with causal inference, econometrics, experimentation, or... 
    Suggested
    Hourly pay
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours

    DoorDash USA

    Seattle, WA
    3 days ago
  •  ...generation of logistics technology. We seek a Machine Learning Engineer to own ML models end-to-end and work with backend teams...  .... Expect strong focus on predictive modeling, causal inference, and marketplace optimization in a high-growth setting. The role emphasizes hands-on... 
    Suggested

    Uber

    Seattle, WA
    1 day ago
  • $173.5k - $331.05k

     .... We are looking for a senior, hands-on engineer to own and evolve the cross-platform GPU...  ...leadership. What You'll Do Own, evolve, and optimize the core GPU rendering platform that...  ...working with GPU driver / hardware vendors ML inference integration (e.g., TensorRT, ONNX... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    Seattle, WA
    1 day ago
  • $190.2k - $345.65k

     ...AI Services team is seeking a SeniorStaffMachine Learning Engineer for our GenAI Services area. In this high-impact role,...  ...Stock, and Premiere. You will design and develop efficient inference pipelines, optimize models for latency and through at inference, and build APIs... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    Seattle, WA
    4 days ago
  • $197.3k - $313.7k

     ...OPPORTUNITIES*Slack is looking for a Staff Machine Learning Engineer with deep expertise in model...  ...low level with training frameworks, optimize model architectures, build finetuning...  ...Familiarity with model optimization for inference (quantization, pruning, speculative... 
    Full time

    Salesforce

    Seattle, WA
    4 days ago
  • $208k - $260k

     ...truly global scale.About the Role:As a Staff Machine Learning Engineer in Remitly's Core AI/ML team, you'll...  ..., customer-facing ML systems.Optimize models and pipelines using MLOps best...  ...algorithms, Large Language Models, Causal Inference, Personalization, Knowledge Graphs,... 
    Full time
    Work at office
    Flexible hours
    3 days per week

    Remitly

    Seattle, WA
    9 hours ago
  •  ...role in shaping how AI models are built, optimized, and scaled. We develop a platform for...  ...multimodal support), AI Ops, efficient inference, and a modern feature platform designed...  ...and drive innovation. We’re looking for engineers and researchers passionate about generative... 
    Flexible hours

    Apple

    Seattle, WA
    1 day ago
  • $195k - $230k

     ...Staff ML Systems EngineerFieldAI's Irvine team is where embodied...  ...architectures, combining rigorous engineering with learning systems proven...  ...pipeline development.Optimize performance across distributed...  ...machine learning training and inference systems.Familiarity with modern... 
    Local area

    Field AI

    Seattle, WA
    1 day ago
  • $209k

     ...pricing, matching, recommendation, and optimization systems that will disrupt the freight industry...  ...a highly motivated Machine Learning Engineer to join Uber’s Marketplace team to...  ...challenges in predictive modeling, causal inference, constrained optimization,... 
    Full time
    Work at office
    Remote work

    Uber

    Seattle, WA
    1 day ago
  •  ...marketing platform top businesses trust to optimize billions in ad spend worldwide. With...  ...optimization, machine learning, and causal inference. We are looking for individuals who not...  ...scientists, data scientists, data engineers and other MLEs to deliver trustworthy results... 
    Work at office
    Work from home
    Worldwide
    Flexible hours

    Haus Analytics

    Seattle, WA
    3 days ago
  • $207k - $275k

     ...capture a large fraction of the total inference market, which is worth tens of billions...  ...qualifications, we’re looking for strong engineers with great taste. The most important...  ...distillation, reward modeling, and policy optimization. Strong research judgment, including... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    4 days ago
  • $131.6k - $210.3k

     ...distributed applications across our global ecosystem.We’re seeking a Staff Software Engineer who will design and build scalable backend systems that...  ..., Mistral, and Gemini into backend systems. Build and optimize RAG pipelines using vector databases (Pinecone, Weaviate,... 
    Full time
    Work experience placement
    Work at office
    Local area

    Visa

    Bellevue, WA
    6 hours ago
  • $220k - $292k

     ...Planning, Hardware, and Test Engineering to solve some of the hardest...  ...are looking for a founding Staff AI Infrastructure Engineer to...  ...our AI Research Scientists, optimizing their experimentation...  ...custom ASICs) for training and inference workloads. Model Observability... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Seattle, WA
    3 days ago
  • $144k - $200k

     ...Staff Propulsion Design EngineerAgile Space Industries is a rapidly growing supplier of propulsion hardware and engineering services for the space industry. Headquartered in Durango, Colorado...  ...the design, development, and optimization of propulsion components and systems... 
    Full time
    Temporary work
    Part time
    Work at office
    Flexible hours

    Agile Space Industries

    Seattle, WA
    4 days ago
  •  ...career, and the financial world.The role: We’re seeking a Staff Security Detection Engineer to build and mature SoFi’s machine learning-driven...  ...Glue, Athena, Lambda) for training, feature pipelines, and inference.MLOps practices - feature stores, model registries, experiment... 
    Remote work

    SoFi

    Seattle, WA
    4 days ago
  •  ...complex, contested environments. We are seeking a Structures Design Engineer to lead the design, development, and testing of a novel launch...  ...and mission operations.Perform trade studies and system optimization to enhance portability, ruggedness, and maintainability of support... 
    Full time
    Temporary work
    Part time

    ClearanceJobs

    Seattle, WA
    3 days ago
  •  ...The 3D Simulation group at Zoox is looking for machine learning engineers to bring the latest research in 3D vision to improve diversity...  .... \n In this role, you will: Research, implement, and optimize state of the art machine learning approaches to improve simulation... 
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    1 day ago
  • $242.8k - $357k

    About the TeamThe Consumer Engineering Team is responsible for helping consumers discover and...  ...our customers.About the RoleAs a Senior Staff Machine Learning Engineer on Core Cx,...  ...learning, and fine-tuning. Hands-on experience optimizing LLM systems (e.g., fine-tuning, advanced... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    Seattle, WA
    3 days ago
  • $229k - $343k

     ...Saturn, and other digital services.Snap Engineering teams build fun and technically...  ...at the forefront.​We’re looking for a Staff Machine Learning Engineer to join Snap...  ...ranking models, embeddings, deep learning, optimization, evaluation, and experimentationStrong... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    Seattle, WA
    4 days ago
  • $208k - $298.5k

     ...people across the globe who think that’s work worth doing.Staff Machine Learning Engineer, Foundation - Seattle Why We Have This RoleWe are...  ...deploying, ML systems in production. Develop a strategy for optimizing models and systems for performance, scalability, efficiency... 
    Full time
    Work at office
    Relocation package
    3 days per week

    Qualtrics

    Seattle, WA
    9 hours ago
  •  ...Job Description: We are seeking a highly motivated FPGA Design Engineer with experience developing high-performance digital designs....  ...implementation, timing closure, bitstream generation. Debug, test, and optimize for reliability, latency, and throughput. Collaborate with... 
    Full time
    Temporary work
    Part time
    Worldwide

    Shield AI

    Seattle, WA
    7 days ago
  • $160k - $210k

     ...function, but to help redefine the future of how work gets done.Staff Network Engineer Location: MPK/Bellevue/DublinSnowflake's Enterprise...  ...Senior Network Engineer to lead the design, operation, and optimization of our Zero Trust and secure-access platform. This role is... 
    Work at office

    Snowflake

    Bellevue, WA
    2 days ago
  • $131.6k - $233.7k

     ...Progress starts with you.Job DescriptionWe are seeking a Staff Systems Engineer, Microsoft Teams, to join Visa’s enterprise-wide subject matter...  ...Systems Engineer, you will own the end‑to‑end deployment, optimization, and evolution of the Teams environment, including... 
    Full time
    Contract work
    Work experience placement
    Work at office
    Local area

    Visa

    Bellevue, WA
    1 day ago
  • $136k - $204k

     ...items to the Internet. Team Overview:The Silicon Characterization Engineering team supports and helps coordinate lab-based probe station...  ...standard operating procedures (SOPs) and provide feedback to optimize and improve them.What You Will Bring:Bachelor's degree in electrical... 
    Work experience placement
    Remote work

    Impinj

    Seattle, WA
    1 day ago
  • $174k - $299k

     ...Role Overview We are looking for experienced and innovative ML engineers with an entrepreneurial mindset to be part of the technical...  ...for e-commerce scenarios at Coupang. We pioneer in innovative optimization techniques for AI reasoning and token completion models.... 
    Temporary work
    Flexible hours

    Coupang

    Seattle, WA
    2 days ago
  • $190.2k - $345.65k

     ...expanding rapidly into adjacent verticals.We are hiring a Senior Staff Machine Learning Engineer to architect and lead the data processing, indexing, and...  ...plus; hands-on familiarity with embedding models and the inference paths that produce them (PyTorch).A track record building... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    Seattle, WA
    9 hours ago
  • $164k - $282k

     ...services and platforms that take care of delivering customer orders optimally right from the moment the order is placed up until it is...  ...this challenge, we are looking for an experienced and passionate engineer who can design and build large-scale, multi-tiered,... 
    Temporary work
    Flexible hours

    Coupang

    Seattle, WA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Engineer, Inference Optimizations. Be the first to apply!