Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Engineer, Inference Optimizations

$191.2k - $239k
Full-time

DigitalOcean

Role Description

DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. You will help ensure we can offer industry-leading performance for our inference services.

  • Performance Architecture: Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers.
  • Deep-Dive Optimization: Engineer solutions for complex performance issues, including attention layer optimizations, memory and precision management, and advanced parallelization across multi-node GPU clusters.
  • Technological Innovation: Proactively implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of the Gen AI landscape.
  • Hardware & Ecosystem Mastery: Act as the subject matter expert on modern GPU families (NVIDIA/AMD) and their software stacks (CUDA, ROCm, TensorRT, OpenAI Triton).
  • Precision Optimization: Develop and deploy state-of-the-art quantization techniques (FP8, INT8, and experimental FP4) to double throughput without losing accuracy.
  • Technical Mentorship: Lead by example through high-quality code and design reviews.
  • Strategic Collaboration: Partner with Product Management and TPMs to translate "theoretical hardware limits" into "shippable product features."
  • Community Leadership: Maintain a strong presence in the GPU infrastructure and model performance optimization communities.

Qualifications

  • 5+ years of experience in high-performance computing or AI infrastructure.
  • Deep familiarity with the Gen AI (LLM, VLM, LMM) landscape.
  • Hands-on experience with attention-layer optimizations and parallelization strategies across distributed GPU environments.
  • Comprehensive understanding of NVIDIA and AMD GPU architectures and their respective software ecosystems (CUDA, ROCm, etc.).
  • Extensive experience integrating, building with, and contributing to open-source software projects.
  • Excellent system design skills, particularly related to low-level GPU programming.
  • Experience acting as a technical lead, driving design and delivery through cross-functional alignment.
  • Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores).
  • Expert-level Triton or CUDA experience.

Requirements

  • Proven track record of solving compute utilization and memory bandwidth bottlenecks.
  • Experience with attention-layer optimizations and parallelization strategies.
  • Expertise in low-level GPU programming and optimization.

Benefits

  • Competitive array of benefits including Employee Assistance Program and flexible time off policy.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning's 10,000+ courses for continued growth and development.
  • Salary range based on market data, relevant years of experience, and skills.
  • Potential for bonus based on company and individual performance.
  • Equity compensation to eligible employees, including equity grants and Employee Stock Purchase Program.

Company Description

DigitalOcean is an equal-opportunity employer. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff Engineer, Inference Optimizations in Remote vacancy
  • $298k - $368k

     ...work jointly with downstream teams on the optimization and integration into the Waymo Driver....  ...a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently...  ...expertise in low-latency on-device inference techniques and a deep understanding of... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    11 hours ago
  • $141k - $249k

     ...Collaborate closely with autonomy and algorithm engineers to scale safe self-driving systems using...  ...such as TensorRT and modelopt to optimize the models running on the truck. Create and benchmark new CUDA kernels for inference. Comprehensively profile model... 
    Suggested
    Work at office
    Work from home
    Flexible hours

    Waabi

    Pittsburgh, PA
    2 days ago
  • $3,000 per month

     ...excellence. THE WORK As the Systems Engineer you'll be preparing plans, implements,...  ...needs, functions that may be logically inferred and implied as crucial to system operations...  ...posting date in order to receive optimal consideration. At Lockheed Martin, we... 
    Suggested
    Full time
    Temporary work
    Part time
    Work experience placement
    Work at office
    Remote work
    Relocation
    Relocation package
    Flexible hours
    Shift work

    Lockheed Martin - US

    Maryland
    24 days ago
  •  ...surface, and sustain next. In this Sr. Staff ML Engineer role, you’ll be the technical lead...  ...that make content succeed, and designing optimization approaches that balance relevance, quality...  ...learning, mechanism design, or causal inference applied to ecosystems/marketplaces.... 
    Suggested
    Full time
    Work at office
    Remote work
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    11 hours ago
  • $260k - $310k

     ...Join the Affirm team as a Senior Staff Machine Learning Engineer and become a pivotal part of our innovative...  ...priorities, and exercising judgment optimized for the broader engineering...  ...training pipelines, model serving and inference infrastructure, monitoring, and automated... 
    Suggested
    Full time
    Work at office
    Remote work
    Flexible hours

    Affirm

    United States
    11 hours ago
  • $137.1k - $201.6k

     ...About the Role We’re looking for a Staff Machine Learning Engineer to drive the design and development of large-scale ML/optimization systems to target personalization efforts...  ..., you will: Contribute to Causal inference modeling to measure the incremental impact... 
    Hourly pay
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours

    From Restaurants Near You

    San Francisco, CA
    11 hours ago
  •  ...ForeFlight is seeking a Senior Machine Learning Engineer to help build and scale domain-...  ...focuses on developing vertical ASR models optimized for high-accuracy transcription in noisy...  ...transcription systems. Optimize inference pipelines for edge, cloud, and low-latency... 
    Full time
    Immediate start
    Worldwide

    Jeppesen Foreflight Careers

    Remote
    11 hours ago
  • $244k - $293k

     ...About the Role: We are hiring a Staff Machine Learning Engineer to drive the design, development,...  ...the right user at the right time to optimizing the efficacy of our paid offerings....  ...facing products. Experience with causal inference, uplift modeling, and interventional... 
    Full time
    Work experience placement
    Remote work

    Match Group

    New York, NY
    11 hours ago
  • $215k - $322k

     ...010. Join GoFundMe as our next Staff Machine Learning Engineer (Pricing) . In this role, you will...  ...checkout experiences, donation yield optimization (one-time and recurring), recurring...  ...systems (data → training → online inference → measurement) with rigorous experimentation... 
    Full time
    Temporary work
    Work at office
    Flexible hours

    Gofundme

    San Francisco, CA
    11 hours ago
  •  ...Technologies Inc. \\ \ We’re Hiring: Staff Machine Learning Engineer \\\\ \ At SA Technologies Inc., we...  ...down to the bare -metal model optimization. \ \ This unique role blends...  ...just operational metrics but also inference quality and concept drift. \\\... 
    Full time

    Dba Technologies, Llc

    Remote
    11 hours ago
  •  ...workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we...  ...This role will be critical in building, optimizing and scaling end-to-end machine learning...  ...including distributed training and inference pipeline for large language models(LLM),... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Servicenow

    Remote
    11 hours ago
  • $170k - $230k

     ...deeply impactful for members.  As a Staff Machine Learning Engineer on our Clinical Health team, you...  ...prototypes into production ML systems optimized for scale, latency, and cost efficiency...  ..., deploying, and operating ML inference systems at scale (real-time streaming... 
    Full time
    Work at office
    Relocation

    Whoop

    Remote
    11 hours ago
  • $180.9k - $265.32k

     ...autonomous driving.   The Role   We are seeking a  Staff Software Engineer to lead the integration and deployment of advanced...  ...scheduling and diagnostics.   ~ Performance Optimization    ~ Optimize inference pipelines using CUDA, TensorRT, and mixed-precision... 
    Full time
    Immediate start
    Night shift

    Lucid

    Remote
    11 hours ago
  • $126k - $193k

     ...talented Machine Learning & AI Engineer to support our programs...  ...Intelligence. We are seeking a Staff/Senior Staff Machine Learning...  ...to design, train, and optimize high‑performance ML models and...  ...fusion Deploy high‑performance inference pipelines for large‑volume, real... 
    Full time
    Temporary work
    For contractors
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Scitec Corp

    Remote
    11 hours ago
  • $251k - $310k

     ...group of machine learning (ML) engineers, software engineers, and ML...  ...you will report to a Senior Staff Engineering Manager   You...  ...bottlenecks in training and inference performance (e.g., memory bandwidth...  ...attention mechanisms. Optimize model code for specific hardware... 
    Full time
    Remote work

    Waymo

    Remote
    11 hours ago
  •  ...everyone. Join us. The Role As a Staff Applied Machine Learning Engineer focused on Fraud & Abuse, you will...  ...activity across Block. The team optimizes for reliable decisions, safe...  ...to end: data contracts, low-latency inference, batch scoring, feature quality, online... 
    Full time
    Local area

    Block

    United States
    11 hours ago
  •  ...A unified model gateway for inference Seamless collection of metrics and feedback Optimization tools for prompts, models, and...  ...Optimization: From prompt engineering to fine-tuning and reinforcement...  ...Founding Member of Technical Staff with expertise in front-end... 
    Full time

    Right Hire Consulting Llc

    Remote
    11 hours ago
  • $151k - $177.5k

     ...battery from the inside out today. We engineer and manufacture ground-breaking battery...  ...engineering logic, statistical methods, optimization, machine learning, and foundation models...  ...workflows for feature generation, inference, event detection, and feedback into operational... 
    Full time

    Sila

    Remote
    11 hours ago
  • $220k - $280k

     ...reimagine the DFS industry together? As a Staff Machine Learning Engineer, you will lead the technical charge...  ...production services. Real-Time Inference at Scale: Steer the design and...  ...pipelines. You will lead the creation and optimization of a centralized feature store... 
    Full time
    Remote work
    Work visa
    Flexible hours

    PrizePicks

    United States
    1 day ago
  •  ...AI Fabrik builds an edge inference delivery network for high-performance...  ...are builders, architects, engineers, and researchers with hands‑...  ...Role We're looking for a Staff Edge Network Engineer to own...  ...Apply provider middle‑mile optimization (e.g., Cloudflare Argo, Akamai... 
    Work at office

    AI Fabrik

    San Francisco, CA
    5 days ago
  • $175k - $250k

     ...from real-time market outcomes. The Staff AI Engineer will be responsible for moving beyond...  ...Key Responsibilities: Learning & Optimization Feedback Loop Implementation: Design...  ...of proven strategies. Model & Inference Infrastructure Ownership: Transition... 
    Remote job
    Full time
    Immediate start
    Shift work

    Mlabs

    Florida, FL
    11 hours ago
  •  ...are looking for an exceptional Search/AI Engineer with experience in Search Relevance to...  ...will lead the design, development, and optimization of intelligent search systems that leverage...  ...language, resolve ambiguity, and infer user intent Design and deploy learning... 
    Full time
    Remote work
    Flexible hours

    Workato

    Remote
    11 hours ago
  • $227k - $300k

     ...Edge. We are looking for a great Senior Staff AI Engineer to join our seasoned AI team and lead...  ...-constrained edge devices and model optimization. You will work in a fast-paced startup...  ...knowledge of modern C++ (C++14/17 for inference). ~ Deep proficiency with PyTorch or... 
    Full time
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus, Inc

    Remote
    11 hours ago
  •  ...are seeking an experienced AI Engineer with deep expertise in...  ...to join our team as a Senior Staff Architect. In this role, you...  ...the design, development, and optimization of cutting-edge RL solutions,...  ...optimization techniques , including inference-time search, chain-of-thought... 
    Full time
    Worldwide

    Jazzx Ai

    Remote
    11 hours ago
  • $178.4k - $267.6k

     ...Qualcomm Technologies, Inc. Job Area: Engineering Group, Engineering Group Machine...  ...acceleration, model quantization, edge inference and related fields. Come join us on...  ...Architect, design, develop and test model optimization techniques that include - but are not limited... 
    Work experience placement
    Work from home

    Qualcomm

    San Diego, CA
    16 hours ago
  • $140k - $165k

     ...the Role: We are seeking a hands-on AI Engineer to design, deploy, and maintain on-prem...  ...Kubernetes, Helm, Docker). Build and optimize RAG pipelines — including document...  ...specific tasks; optimize for edge/on-device inference. Implement Model Control Protocols (MCP... 
    Full time

    SK Hynix Memory Solutions America Inc.

    Remote
    11 hours ago
  •  ...Job Overview As a Staff AI Engineer, you define and drive the architecture of AI and agentic...  ...design of large‑scale AI/LLM systems: inference platforms, APIs, and distributed architectures...  ...to capacity planning and cost optimization strategies for AI/LLM infrastructure GenAI... 
    Work at office
    Remote work
    Flexible hours

    Modernizing Medicine

    Boca Raton, FL
    2 days ago
  • $190k - $250k

     ...patient-facing agentic platform to optimize patient outcomes through a...  .... We are looking for expert engineers and leads to join our team...  ...Responsibilities As a Staff Backend AI Engineer, you will...  ...infrastructure, from low-latency inference pipelines to workflow automation... 
    Full time
    Work experience placement
    Remote work

    Arbiter Ai

    New York, NY
    11 hours ago
  • $157k - $200k

     ...performance, security, and scalability. Establish and enforce engineering best practices, including rigorous code reviews, automated testing...  ...teams to improve system stability, enhance observability, and optimize incident response procedures. Drive AI adoption by identifying... 
    Full time
    Temporary work
    Work at office
    Immediate start
    Remote work

    Vizient

    Irving, TX
    5 days ago
  • $206.32k - $221.4k

     ...Description CyberSource Corporation, a Visa Inc. company, needs a Staff SW Engineer (multiple openings) in Foster City, CA to Responsible for...  .... Formulate methods to enable consistent data loading and optimize data operations. Monitor health of platforms, generate... 
    Work at office
    Local area
    Remote work
    2 days per week
    3 days per week

    Visa

    Foster, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Engineer, Inference Optimizations. Be the first to apply!