Staff Engineer, Inference Optimizations
$191.2k - $239kDigitalOcean
Dive in and do the best work of your career at DigitalOcean. Journey alongside a strong community of top talent who are relentless in their drive to build the simplest scalable cloud. If you have a growth mindset, naturally like to think big and bold, and are energized by the fast-paced environment of a true industry disruptor, you’ll find your place here. We value winning together—while learning, having fun, and making a profound difference for the dreamers and builders in the world.
DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be the Inference Cloud of choice for digitally native companies and you will help ensure we can offer the industry-leading performance for our inference services. You will be responsible for the architectural decisions that maximize throughput and minimize latency for the world’s most advanced large models. As an IC leader, you will act as a force multiplier for the engineering organization, solving the most complex bottlenecks in memory bandwidth and compute utilization while guiding the technical roadmap for our high-performance inference fleet.
What You’ll Do:
- Performance Architecture: Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers, ensuring our infrastructure extracts maximum value from every TFLOP.
- Deep-Dive Optimization: Engineer solutions for complex performance issues, including attention layer optimizations, memory and precision management, and advanced parallelization across multi-node GPU clusters.
- Technological Innovation: Proactively implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of the Gen AI landscape. Some examples of projects you may work on:
- Improving batch size performance using AMD's AITER library for AMD MI355X - identify and tune AITER's CK (composable kernel) or ASK (assembly) to optimize FP8 / BF16
- Identify kernel fusion opportunities for GLM-5 kernels for different layers of the Transformer block (FlashAttention, RMS Norm)
- Tune expert gateway router kernels for MoE models like Qwen3-235B, DeepSeek V3, GLM-5 etc
- Hardware & Ecosystem Mastery: Act as the subject matter expert on modern GPU families (NVIDIA/AMD) and their software stacks (CUDA, ROCm, TensorRT, OpenAI Triton), advising on hardware procurement and software integration.
- Precision Optimization: Develop and deploy state-of-the-art quantization techniques (FP8, INT8, and experimental FP4) to double throughput without losing accuracy.
- Technical Mentorship: Lead by example through high-quality code and design reviews, elevating the technical bar for the team without the administrative overhead of direct management.
- Strategic Collaboration: Partner with Product Management and TPMs to translate "theoretical hardware limits" into "shippable product features," ensuring our platform is both powerful and developer-friendly.
- Community Leadership: Maintain a strong presence in the GPU infrastructure and model performance optimization communities, contributing to and integrating the best of open-source AI.
What You’ll Bring to DigitalOcean:
- Technical Depth: 5+ years of experience in high-performance computing or AI infrastructure, with a proven track record of solving compute utilization and memory bandwidth bottlenecks.
- Gen AI Literacy: Deep familiarity with the Gen AI (LLM, VLM, LMM) landscape, including the specific quirks and architectural requirements of major model families.
- Optimization Expert: Hands-on experience with attention-layer optimizations and parallelization strategies across distributed GPU environments.
- Hardware Fluency: Comprehensive understanding of NVIDIA and AMD GPU architectures and their respective software ecosystems (CUDA, ROCm, etc.).
- Open Source Mastery: Extensive experience integrating, building with, and contributing to open-source software projects.
- Systems Design: Excellent system design skills, particularly related to low-level GPU programming - optimization, memory access patterns, and parallel execution.
- Leadership through Influence: Experience acting as a technical lead, driving design and delivery through cross-functional alignment and expert-level delegation.
- Low-Level Mastery: Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores).
- The Toolkit: Expert-level Triton or CUDA. If you’ve contributed to the Triton compiler or wrote custom CUDA kernels for a major LLM, we want you.
Compensation Range:
- $191,200 - $239,000
*This is a hybrid role
JR: 2026-7625
#LI-Hybrid
Why You’ll Like Working for DigitalOcean
- We innovate with purpose. You’ll be a part of a cutting-edge technology company with an upward trajectory, who are proud to simplify cloud and AI so builders can spend more time creating software that changes the world. As a member of the team, you will be a Shark who thinks big, bold, and scrappy, like an owner with a bias for action and a powerful sense of responsibility for customers, products, employees, and decisions.
- We prioritize career development. At DO, you’ll do the best work of your career. You will work with some of the smartest and most interesting people in the industry. We are a high-performance organization that will always challenge you to think big. Our organizational development team will provide you with resources to ensure you keep growing. We provide employees with reimbursement for relevant conferences, training, and education. All employees have access to LinkedIn Learning's 10,000+ courses to support their continued growth and development.
- We care about your well-being. Regardless of your location, we will provide you with a competitive array of benefits to support you from our Employee Assistance Program to Local Employee Meetups to flexible time off policy, to name a few. While the philosophy around our benefits is the same worldwide, specific benefits may vary based on local regulations and preferences.
- We reward our employees. The salary range for this position is based on market data, relevant years of experience, and skills. You may qualify for a bonus in addition to base salary; bonus amounts are determined based on company and individual performance. We also provide equity compensation to eligible employees, including equity grants upon hire and the option to participate in our Employee Stock Purchase Program.
- DigitalOcean is an equal-opportunity employer. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service.
Application Limit: You may apply to a maximum of 3 positions within any 180-day period. This policy promotes better role-candidate matching and encourages thoughtful applications where your qualifications align most strongly.
$173.5k - $331.05k
.... We are looking for a senior, hands-on engineer to own and evolve the cross-platform GPU... .... What You’ll Do Own, evolve, and optimize the core GPU rendering platform that underpins... ...GPU driver / hardware vendors ML inference integration (e.g., TensorRT, ONNX...SuggestedTemporary workLocal areaWorldwide$195k - $239k
...the world. We are looking for a Staff Forward Deployed Engineer who is passionate about solving... ...mission is to accelerate “time-to-inference” in production at scale and serve as... ...systems, benchmarking frameworks, model optimization agents, and deployment automation...SuggestedLocal areaWorldwideFlexible hours$190.2k - $345.65k
...Firefly’s Generative AI Services team is seeking a Senior Staff Machine Learning Engineer for our GenAI Services area. In this high-impact... ...and Premiere. You will design and develop efficient inference pipelines, optimize models for latency and throughput, and build APIs...SuggestedTemporary workLocal areaWorldwide$195k - $239k
...in the world. We are looking for a Staff Forward Deployed Engineer (FDE) who is passionate about... ...internal engineering teams to deploy, optimize, and scale production AI systems on... ...infrastructure deployment. You will work across Inference Engine, runtime systems,...SuggestedFull timeLocal areaWorldwideFlexible hours$209.1k - $282.9k
...We seek a Staff Robotics Engineer to help create the next generation of Physical AI systems on Arm... ...performance, reliability, and safety Optimize end-to-end latency, determinism,... ...Familiarity with AI and perception inference pipelines deployed on embedded or edge...SuggestedWork at officeLocal areaVisa sponsorshipRelocation package$136k - $204k
...items to the Internet. Team Overview:The Silicon Characterization Engineering team supports and helps coordinate lab-based probe station... ...standard operating procedures (SOPs) and provide feedback to optimize and improve them.What You Will Bring:Bachelor's degree in electrical...Work experience placementRemote work$144k - $200k
...growing supplier of propulsion hardware and engineering services for the space industry,... ...reliability, and purpose. Position Overview Staff Propulsion Design Engineer – Seattle‑... ...studies, modeling, and simulation to optimize performance, efficiency, and reliability...Full timeWork at office- ...group of committed researchers, engineers, policy experts, and business... ...sandboxed code execution, or inference and RL infrastructure... ...another field that models and optimizes complex systems Strong candidates... ...: Currently, we expect all staff to be in one of our offices...Full timeWork at officeVisa sponsorshipFlexible hours
- ...architectures, combining rigorous engineering with learning systems proven... ...field. We are seeking a Staff ML Systems Engineer to... ...distributed pipeline development. Optimize performance across... ...machine learning training and inference systems. Familiarity with...Local area
$210k - $234k
...to deliver exceptional service to seller and buyer clients. Engineering @ Compass Compass has built the first modern end-to-end... ...retention, and training-use terms of third-party foundation models, inference providers, and vector stores used in production. Instrument...Minimum wageFlexible hours- ...a new battery pack, from concept through production Drive optimal performance while managing size, weight, manufacturing, and reliability... ...REQUIRED QUALIFICATIONS: ~ Bachelor's Degree in Mechanical Engineering, or equivalent ~5+ years of experience ~ Strong technical...Full timeTemporary workPart timeWorldwide
$188k - $275k
...in March 2025. Learn more at What You'll Do: The Storage Engine Team at CoreWeave is responsible for the product capabilities... ..., and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency. Lead efforts to improve...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$131.6k - $210.3k
...applications across our global ecosystem. We’re seeking a Staff Software Engineer who will design and build scalable backend systems that... ...Claude, Mistral, and Gemini into backend systems. Build and optimize RAG pipelines using vector databases (Pinecone, Weaviate, FAISS...Full timeWork experience placementWork at officeLocal area- ...contested environments. We are seeking a Structures Design Engineer to lead the design, development, and testing of a novel launch... ...and mission operations. Perform trade studies and system optimization to enhance portability, ruggedness, and maintainability of support...Full timeTemporary workPart timeWorldwide
- ...we are looking for a Lead Survivability Enginee r to shape and optimize the low observable (LO) performance of our next-generation autonomous aircraft. This role is ideal for an experienced engineer with deep expertise in radar cross-section (RCS) shaping, LO materials...Full timeTemporary workPart timeWorldwide
- ...Job Description: We are seeking a highly motivated FPGA Design Engineer with experience developing high-performance digital designs.... ...implementation, timing closure, bitstream generation. Debug, test, and optimize for reliability, latency, and throughput. Collaborate with...Full timeTemporary workPart timeWorldwide
$115k - $230k
...Culture, Great Rewards, and Great Careers. GEICO is seeking a Staff Engineer, Applied AI to help shape how Generative AI enhances... ...competency in distributed systems, service design, performance optimization, and reliability engineering. Nice to Have Experience...Hourly payFull timeWork experience placementLocal area$164k - $282k
...hyper-connected world. Role Overview Coupang’s Rocket Growth Engineering Team is responsible for building the next generation fulfillment... ...to Settlements. We provide intelligence for Sellers to optimize their Selection, prices, and fulfillment experience to Coupang...Temporary workFlexible hours$170k - $260k
...YouTube. Job Description: We are seeking a highly motivated engineer to support the design, development, integration, and fielding... ..., and analysis-to-test correlation. Drive structural optimization across performance, weight, cost, and manufacturability targets...Full timeTemporary workPart timeWorldwide- ...Staff Network EngineerLocation: MPK/Bellevue/DublinSnowflake's Enterprise Technology Network Services team is looking for a Senior Network Engineer to lead the design, operation, and optimization of our Zero Trust and secure-access platform. This role is Zscaler-centric...Work at office
- ...and integrate state-of-the-art high-power systems tailored for optimal performance and reliability in harsh environments Guide teammates... ..., manufacturing, etc. Help build the team and lead multiple engineers dedicated to the power system design Drive a cross-...Full timeTemporary workPart timeWorldwide
$140k - $192.5k
...We are looking for a highly curious, self‑directed full‑stack engineer who thrives on solving complex technical challenges for our Digital... ...with new technologies, including AI‑driven tools, to optimize our systems and improve developer workflows. You will: Participate...Local areaWorldwideFlexible hours- ...complex, contested environments. We are seeking a Staff Mechanical Design Engineer to lead the design, development, and testing of Ground... ...mission operations. Perform trade studies and system optimization to enhance portability, ruggedness, and maintainability...Full timeTemporary workPart timeWorldwide
- ...Shield AI on LinkedIn, X, Instagram, and YouTube. Job Description: We are seeking a highly motivated Staff Advance Industrial Engineer to design, optimize, and scale world-class manufacturing systems for advanced aerospace products. In this role, you will lead factory...Full timeTemporary workPart timeWorldwideRelocation
$210k - $320k
...will push the boundaries of innovation and performance. Shield AI is seeking a Communications Systems Engineer to lead the design, integration, and optimization of advanced communication systems for our next-generation autonomous aircraft. This is a high-impact role...Full timeTemporary workPart timeWorldwide- ...LinkedIn, X, Instagram, and YouTube. Position Overview: We are seeking a highly experienced Senior Staff Network Engineer to lead the design, implementation, and optimization of complex network infrastructures. This role requires deep expertise across network engineering,...Full timeTemporary workPart timeWorldwide
$134.5k - $214.7k
...Join Visa and do work that matters - to you, to your community, and to the world. Progress starts with you.Job DescriptionThe Staff Software Engineer is a senior technical leader responsible for designing, building, and evolving secure, scalable solutions for Visa's...Full timeContract workWork experience placementWork at officeLocal area- Senior Staff Network Engineer Lambda, the superintelligence cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers... ...engineers building one of the largest AI training and inference networks in the world. Lambda has 10x'd over the last three...Local areaFlexible hours
- ...meaningful business workflows.You will work with the AI Platform engineering team, product leader, and subject-matter experts across the... ...experience building and operating sophisticated software systems at Sr Staff Engineer equivalent scope.A track record of leading technically...Remote work
$170k - $250k
...intelligent autonomous aircraft designed to operate and survive in highly contested environments. We are seeking a Survivability Materials Engineer to lead the selection, qualification, specification, and production implementation of low-observable and signature-control...Full timeTemporary workPart time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Engineer, Inference Optimizations. Be the first to apply!
- software engineer staff Seattle, WA
- technology administrator Seattle, WA
- assistant engineer Seattle, WA
- assistant chief engineer Seattle, WA
- staff engineer Seattle, WA
- assistant engineering manager Seattle, WA
- senior staff systems engineer Seattle, WA
- senior staff engineer Seattle, WA
- assistant mechanical engineer Seattle, WA
- project engineer assistant project manager Seattle, WA





