Staff Engineer, Inference Optimizations
$191.2k - $239kDigitalOcean
Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you’ll find your place here. We value winning together—while learning, having fun, and making a profound difference for the dreamers and builders in the world.
DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be the Inference Cloud of choice for digitally native companies and you will help ensure we can offer the industry-leading performance for our inference services. You will be responsible for the architectural decisions that maximize throughput and minimize latency for the world’s most advanced large models. As an IC leader, you will act as a force multiplier for the engineering organization, solving the most complex bottlenecks in memory bandwidth and compute utilization while guiding the technical roadmap for our high-performance inference fleet.
What You’ll Do:
- Performance Architecture: Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers, ensuring our infrastructure extracts maximum value from every TFLOP.
- Deep-Dive Optimization: Engineer solutions for complex performance issues, including attention layer optimizations, memory and precision management, and advanced parallelization across multi-node GPU clusters.
- Technological Innovation: Proactively implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of the Gen AI landscape. Some examples of projects you may work on:
- Improving batch size performance using AMD's AITER library for AMD MI355X - identify and tune AITER's CK (composable kernel) or ASK (assembly) to optimize FP8 / BF16
- Identify kernel fusion opportunities for GLM-5 kernels for different layers of the Transformer block (FlashAttention, RMS Norm)
- Tune expert gateway router kernels for MoE models like Qwen3-235B, DeepSeek V3, GLM-5 etc
- Hardware & Ecosystem Mastery: Act as the subject matter expert on modern GPU families (NVIDIA/AMD) and their software stacks (CUDA, ROCm, TensorRT, OpenAI Triton), advising on hardware procurement and software integration.
- Precision Optimization: Develop and deploy state-of-the-art quantization techniques (FP8, INT8, and experimental FP4) to double throughput without losing accuracy.
- Technical Mentorship: Lead by example through high-quality code and design reviews, elevating the technical bar for the team without the administrative overhead of direct management.
- Strategic Collaboration: Partner with Product Management and TPMs to translate "theoretical hardware limits" into "shippable product features," ensuring our platform is both powerful and developer-friendly.
- Community Leadership: Maintain a strong presence in the GPU infrastructure and model performance optimization communities, contributing to and integrating the best of open-source AI.
What You’ll Bring to DigitalOcean:
- Technical Depth: 5+ years of experience in high-performance computing or AI infrastructure, with a proven track record of solving compute utilization and memory bandwidth bottlenecks.
- Gen AI Literacy: Deep familiarity with the Gen AI (LLM, VLM, LMM) landscape, including the specific quirks and architectural requirements of major model families.
- Optimization Expert: Hands-on experience with attention-layer optimizations and parallelization strategies across distributed GPU environments.
- Hardware Fluency: Comprehensive understanding of NVIDIA and AMD GPU architectures and their respective software ecosystems (CUDA, ROCm, etc.).
- Open Source Mastery: Extensive experience integrating, building with, and contributing to open-source software projects.
- Systems Design: Excellent system design skills, particularly related to low-level GPU programming - optimization, memory access patterns, and parallel execution.
- Leadership through Influence: Experience acting as a technical lead, driving design and delivery through cross-functional alignment and expert-level delegation.
- Low-Level Mastery: Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores).
- The Toolkit: Expert-level Triton or CUDA. If you’ve contributed to the Triton compiler or wrote custom CUDA kernels for a major LLM, we want you.
Compensation Range:
- $191,200 - $239,000
*This is a hybrid role
JR: 2026-7625
#LI-Hybrid
Why You’ll Like Working for DigitalOcean
- We innovate with purpose. You’ll be a part of a cutting-edge technology company with an upward trajectory, who are proud to simplify cloud and AI so builders can spend more time creating software that changes the world. As a member of the team, you will be a Shark who thinks big, bold, and scrappy, like an owner with a bias for action and a powerful sense of responsibility for customers, products, employees, and decisions.
- We prioritize career development. At DO, you’ll do the best work of your career. You will work with some of the smartest and most interesting people in the industry. We are a high-performance organization that will always challenge you to think big. Our organizational development team will provide you with resources to ensure you keep growing. We provide employees with reimbursement for relevant conferences, training, and education. All employees have access to LinkedIn Learning's 10,000+ courses to support their continued growth and development.
- We care about your well-being. Regardless of your location, we will provide you with a competitive array of benefits to support you from our Employee Assistance Program to Local Employee Meetups to flexible time off policy, to name a few. While the philosophy around our benefits is the same worldwide, specific benefits may vary based on local regulations and preferences.
- We reward our employees. The salary range for this position is based on market data, relevant years of experience, and skills. You may qualify for a bonus in addition to base salary; bonus amounts are determined based on company and individual performance. We also provide equity compensation to eligible employees, including equity grants upon hire and the option to participate in our Employee Stock Purchase Program.
- DigitalOcean is an equal-opportunity employer. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service.
Application Limit: You may apply to a maximum of 3 positions within any 180-day period. This policy promotes better role-candidate matching and encourages thoughtful applications where your qualifications align most strongly.
$203.5k - $299.3k
...RoleWe are hiring a Causal Machine Learning Engineer to help build the causal ML foundation... ...policy evaluation, promotion optimization, or marketplace decisioning systems.You... ...…Deep practical experience with causal inference, econometrics, experimentation, or causal...SuggestedHourly payWork at officeLocal areaRemote workFlexible hours$229k - $343k
...infrastructure, and on-device and server-side inference. Our team creates intuitive tools,... ....We’re looking for a Machine Learning Engineer to join Snap Inc!What you’ll do:Develop... ...of GPU, CPU, or NPU optimization techniquesExperience building and optimizing...SuggestedFull timeLive inWork at officeLocal areaWorldwide- Nuance Labs in Seattle is seeking a Member of Technical Staff focused on model optimization and inference. This role demands expertise in refining AI models for real-time interactions, requiring a strong foundation in ML systems and familiarity with frameworks like vLLM...Suggested
$203.5k - $299.3k
...are hiring a Causal Machine Learning Engineer to help build the causal ML foundation... ...counterfactual policy evaluation, promotion optimization, or marketplace decisioning systems.... ...Deep practical experience with causal inference, econometrics, experimentation, or...SuggestedHourly payFull timeWork at officeLocal areaRemote workFlexible hours- ...generation of logistics technology. We seek a Machine Learning Engineer to own ML models end-to-end and work with backend teams... .... Expect strong focus on predictive modeling, causal inference, and marketplace optimization in a high-growth setting. The role emphasizes hands-on...Suggested
$173.5k - $331.05k
.... We are looking for a senior, hands-on engineer to own and evolve the cross-platform GPU... ...leadership. What You'll Do Own, evolve, and optimize the core GPU rendering platform that... ...working with GPU driver / hardware vendors ML inference integration (e.g., TensorRT, ONNX...Full timeTemporary workLocal areaWorldwide$190.2k - $345.65k
...AI Services team is seeking a SeniorStaffMachine Learning Engineer for our GenAI Services area. In this high-impact role,... ...Stock, and Premiere. You will design and develop efficient inference pipelines, optimize models for latency and through at inference, and build APIs...Full timeTemporary workLocal areaWorldwide$197.3k - $313.7k
...OPPORTUNITIES*Slack is looking for a Staff Machine Learning Engineer with deep expertise in model... ...low level with training frameworks, optimize model architectures, build finetuning... ...Familiarity with model optimization for inference (quantization, pruning, speculative...Full time$208k - $260k
...truly global scale.About the Role:As a Staff Machine Learning Engineer in Remitly's Core AI/ML team, you'll... ..., customer-facing ML systems.Optimize models and pipelines using MLOps best... ...algorithms, Large Language Models, Causal Inference, Personalization, Knowledge Graphs,...Full timeWork at officeFlexible hours3 days per week- ...role in shaping how AI models are built, optimized, and scaled. We develop a platform for... ...multimodal support), AI Ops, efficient inference, and a modern feature platform designed... ...and drive innovation. We’re looking for engineers and researchers passionate about generative...Flexible hours
$195k - $230k
...Staff ML Systems EngineerFieldAI's Irvine team is where embodied... ...architectures, combining rigorous engineering with learning systems proven... ...pipeline development.Optimize performance across distributed... ...machine learning training and inference systems.Familiarity with modern...Local area$209k
...pricing, matching, recommendation, and optimization systems that will disrupt the freight industry... ...a highly motivated Machine Learning Engineer to join Uber’s Marketplace team to... ...challenges in predictive modeling, causal inference, constrained optimization,...Full timeWork at officeRemote work- ...marketing platform top businesses trust to optimize billions in ad spend worldwide. With... ...optimization, machine learning, and causal inference. We are looking for individuals who not... ...scientists, data scientists, data engineers and other MLEs to deliver trustworthy results...Work at officeWork from homeWorldwideFlexible hours
$207k - $275k
...capture a large fraction of the total inference market, which is worth tens of billions... ...qualifications, we’re looking for strong engineers with great taste. The most important... ...distillation, reward modeling, and policy optimization. Strong research judgment, including...Permanent employmentTemporary workCasual workWork at officeFlexible hours$131.6k - $210.3k
...distributed applications across our global ecosystem.We’re seeking a Staff Software Engineer who will design and build scalable backend systems that... ..., Mistral, and Gemini into backend systems. Build and optimize RAG pipelines using vector databases (Pinecone, Weaviate,...Full timeWork experience placementWork at officeLocal area$220k - $292k
...Planning, Hardware, and Test Engineering to solve some of the hardest... ...are looking for a founding Staff AI Infrastructure Engineer to... ...our AI Research Scientists, optimizing their experimentation... ...custom ASICs) for training and inference workloads. Model Observability...Full timeWork experience placementImmediate start$144k - $200k
...Staff Propulsion Design EngineerAgile Space Industries is a rapidly growing supplier of propulsion hardware and engineering services for the space industry. Headquartered in Durango, Colorado... ...the design, development, and optimization of propulsion components and systems...Full timeTemporary workPart timeWork at officeFlexible hours- ...career, and the financial world.The role: We’re seeking a Staff Security Detection Engineer to build and mature SoFi’s machine learning-driven... ...Glue, Athena, Lambda) for training, feature pipelines, and inference.MLOps practices - feature stores, model registries, experiment...Remote work
- ...complex, contested environments. We are seeking a Structures Design Engineer to lead the design, development, and testing of a novel launch... ...and mission operations.Perform trade studies and system optimization to enhance portability, ruggedness, and maintainability of support...Full timeTemporary workPart time
- ...The 3D Simulation group at Zoox is looking for machine learning engineers to bring the latest research in 3D vision to improve diversity... .... \n In this role, you will: Research, implement, and optimize state of the art machine learning approaches to improve simulation...Full timeTemporary workRelocation package
$242.8k - $357k
About the TeamThe Consumer Engineering Team is responsible for helping consumers discover and... ...our customers.About the RoleAs a Senior Staff Machine Learning Engineer on Core Cx,... ...learning, and fine-tuning. Hands-on experience optimizing LLM systems (e.g., fine-tuning, advanced...Hourly payWork at officeLocal areaRemote workFlexible hours$229k - $343k
...Saturn, and other digital services.Snap Engineering teams build fun and technically... ...at the forefront.We’re looking for a Staff Machine Learning Engineer to join Snap... ...ranking models, embeddings, deep learning, optimization, evaluation, and experimentationStrong...Full timeLive inWork at officeLocal area$208k - $298.5k
...people across the globe who think that’s work worth doing.Staff Machine Learning Engineer, Foundation - Seattle Why We Have This RoleWe are... ...deploying, ML systems in production. Develop a strategy for optimizing models and systems for performance, scalability, efficiency...Full timeWork at officeRelocation package3 days per week- ...Job Description: We are seeking a highly motivated FPGA Design Engineer with experience developing high-performance digital designs.... ...implementation, timing closure, bitstream generation. Debug, test, and optimize for reliability, latency, and throughput. Collaborate with...Full timeTemporary workPart timeWorldwide
$160k - $210k
...function, but to help redefine the future of how work gets done.Staff Network Engineer Location: MPK/Bellevue/DublinSnowflake's Enterprise... ...Senior Network Engineer to lead the design, operation, and optimization of our Zero Trust and secure-access platform. This role is...Work at office$131.6k - $233.7k
...Progress starts with you.Job DescriptionWe are seeking a Staff Systems Engineer, Microsoft Teams, to join Visa’s enterprise-wide subject matter... ...Systems Engineer, you will own the end‑to‑end deployment, optimization, and evolution of the Teams environment, including...Full timeContract workWork experience placementWork at officeLocal area$136k - $204k
...items to the Internet. Team Overview:The Silicon Characterization Engineering team supports and helps coordinate lab-based probe station... ...standard operating procedures (SOPs) and provide feedback to optimize and improve them.What You Will Bring:Bachelor's degree in electrical...Work experience placementRemote work$174k - $299k
...Role Overview We are looking for experienced and innovative ML engineers with an entrepreneurial mindset to be part of the technical... ...for e-commerce scenarios at Coupang. We pioneer in innovative optimization techniques for AI reasoning and token completion models....Temporary workFlexible hours$190.2k - $345.65k
...expanding rapidly into adjacent verticals.We are hiring a Senior Staff Machine Learning Engineer to architect and lead the data processing, indexing, and... ...plus; hands-on familiarity with embedding models and the inference paths that produce them (PyTorch).A track record building...Full timeTemporary workLocal areaWorldwide$164k - $282k
...services and platforms that take care of delivering customer orders optimally right from the moment the order is placed up until it is... ...this challenge, we are looking for an experienced and passionate engineer who can design and build large-scale, multi-tiered,...Temporary workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Engineer, Inference Optimizations. Be the first to apply!
- project engineer assistant project manager Seattle, WA
- senior staff systems engineer Seattle, WA
- staff data engineer Seattle, WA
- assistant engineer Seattle, WA
- engineering aide Seattle, WA
- software engineer staff Seattle, WA
- senior staff engineer Seattle, WA
- staff security engineer Seattle, WA
- technology administrator Seattle, WA
- staff engineer Seattle, WA



