Staff Engineer, Inference Optimizations
$191.2k - $239kDigitalOcean
Role Description
DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. You will help ensure we can offer industry-leading performance for our inference services.
- Performance Architecture: Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers.
- Deep-Dive Optimization: Engineer solutions for complex performance issues, including attention layer optimizations, memory and precision management, and advanced parallelization across multi-node GPU clusters.
- Technological Innovation: Proactively implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of the Gen AI landscape.
- Hardware & Ecosystem Mastery: Act as the subject matter expert on modern GPU families (NVIDIA/AMD) and their software stacks (CUDA, ROCm, TensorRT, OpenAI Triton).
- Precision Optimization: Develop and deploy state-of-the-art quantization techniques (FP8, INT8, and experimental FP4) to double throughput without losing accuracy.
- Technical Mentorship: Lead by example through high-quality code and design reviews.
- Strategic Collaboration: Partner with Product Management and TPMs to translate "theoretical hardware limits" into "shippable product features."
- Community Leadership: Maintain a strong presence in the GPU infrastructure and model performance optimization communities.
Qualifications
- 5+ years of experience in high-performance computing or AI infrastructure.
- Deep familiarity with the Gen AI (LLM, VLM, LMM) landscape.
- Hands-on experience with attention-layer optimizations and parallelization strategies across distributed GPU environments.
- Comprehensive understanding of NVIDIA and AMD GPU architectures and their respective software ecosystems (CUDA, ROCm, etc.).
- Extensive experience integrating, building with, and contributing to open-source software projects.
- Excellent system design skills, particularly related to low-level GPU programming.
- Experience acting as a technical lead, driving design and delivery through cross-functional alignment.
- Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores).
- Expert-level Triton or CUDA experience.
Requirements
- Proven track record of solving compute utilization and memory bandwidth bottlenecks.
- Experience with attention-layer optimizations and parallelization strategies.
- Expertise in low-level GPU programming and optimization.
Benefits
- Competitive array of benefits including Employee Assistance Program and flexible time off policy.
- Reimbursement for relevant conferences, training, and education.
- Access to LinkedIn Learning's 10,000+ courses for continued growth and development.
- Salary range based on market data, relevant years of experience, and skills.
- Potential for bonus based on company and individual performance.
- Equity compensation to eligible employees, including equity grants and Employee Stock Purchase Program.
Company Description
DigitalOcean is an equal-opportunity employer. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service.
$298k - $368k
...work jointly with downstream teams on the optimization and integration into the Waymo Driver.... ...a diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently... ...expertise in low-latency on-device inference techniques and a deep understanding of...SuggestedFull timeRemote work$141k - $249k
...Collaborate closely with autonomy and algorithm engineers to scale safe self-driving systems using... ...such as TensorRT and modelopt to optimize the models running on the truck. Create and benchmark new CUDA kernels for inference. Comprehensively profile model...SuggestedWork at officeWork from homeFlexible hours$3,000 per month
...excellence. THE WORK As the Systems Engineer you'll be preparing plans, implements,... ...needs, functions that may be logically inferred and implied as crucial to system operations... ...posting date in order to receive optimal consideration. At Lockheed Martin, we...SuggestedFull timeTemporary workPart timeWork experience placementWork at officeRemote workRelocationRelocation packageFlexible hoursShift work- ...surface, and sustain next. In this Sr. Staff ML Engineer role, you’ll be the technical lead... ...that make content succeed, and designing optimization approaches that balance relevance, quality... ...learning, mechanism design, or causal inference applied to ecosystems/marketplaces....SuggestedFull timeWork at officeRemote workRelocationRelocation package
$260k - $310k
...Join the Affirm team as a Senior Staff Machine Learning Engineer and become a pivotal part of our innovative... ...priorities, and exercising judgment optimized for the broader engineering... ...training pipelines, model serving and inference infrastructure, monitoring, and automated...SuggestedFull timeWork at officeRemote workFlexible hours$137.1k - $201.6k
...About the Role We’re looking for a Staff Machine Learning Engineer to drive the design and development of large-scale ML/optimization systems to target personalization efforts... ..., you will: Contribute to Causal inference modeling to measure the incremental impact...Hourly payFull timeWork at officeLocal areaRemote workFlexible hours- ...ForeFlight is seeking a Senior Machine Learning Engineer to help build and scale domain-... ...focuses on developing vertical ASR models optimized for high-accuracy transcription in noisy... ...transcription systems. Optimize inference pipelines for edge, cloud, and low-latency...Full timeImmediate startWorldwide
$244k - $293k
...About the Role: We are hiring a Staff Machine Learning Engineer to drive the design, development,... ...the right user at the right time to optimizing the efficacy of our paid offerings.... ...facing products. Experience with causal inference, uplift modeling, and interventional...Full timeWork experience placementRemote work$215k - $322k
...010. Join GoFundMe as our next Staff Machine Learning Engineer (Pricing) . In this role, you will... ...checkout experiences, donation yield optimization (one-time and recurring), recurring... ...systems (data → training → online inference → measurement) with rigorous experimentation...Full timeTemporary workWork at officeFlexible hours- ...Technologies Inc. \\ \ We’re Hiring: Staff Machine Learning Engineer \\\\ \ At SA Technologies Inc., we... ...down to the bare -metal model optimization. \ \ This unique role blends... ...just operational metrics but also inference quality and concept drift. \\\...Full time
- ...workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we... ...This role will be critical in building, optimizing and scaling end-to-end machine learning... ...including distributed training and inference pipeline for large language models(LLM),...Full timeWork at officeRemote workFlexible hours
$170k - $230k
...deeply impactful for members. As a Staff Machine Learning Engineer on our Clinical Health team, you... ...prototypes into production ML systems optimized for scale, latency, and cost efficiency... ..., deploying, and operating ML inference systems at scale (real-time streaming...Full timeWork at officeRelocation$180.9k - $265.32k
...autonomous driving. The Role We are seeking a Staff Software Engineer to lead the integration and deployment of advanced... ...scheduling and diagnostics. ~ Performance Optimization ~ Optimize inference pipelines using CUDA, TensorRT, and mixed-precision...Full timeImmediate startNight shift$126k - $193k
...talented Machine Learning & AI Engineer to support our programs... ...Intelligence. We are seeking a Staff/Senior Staff Machine Learning... ...to design, train, and optimize high‑performance ML models and... ...fusion Deploy high‑performance inference pipelines for large‑volume, real...Full timeTemporary workFor contractorsWork experience placementWork at officeImmediate startRemote workFlexible hours$251k - $310k
...group of machine learning (ML) engineers, software engineers, and ML... ...you will report to a Senior Staff Engineering Manager You... ...bottlenecks in training and inference performance (e.g., memory bandwidth... ...attention mechanisms. Optimize model code for specific hardware...Full timeRemote work- ...everyone. Join us. The Role As a Staff Applied Machine Learning Engineer focused on Fraud & Abuse, you will... ...activity across Block. The team optimizes for reliable decisions, safe... ...to end: data contracts, low-latency inference, batch scoring, feature quality, online...Full timeLocal area
- ...A unified model gateway for inference Seamless collection of metrics and feedback Optimization tools for prompts, models, and... ...Optimization: From prompt engineering to fine-tuning and reinforcement... ...Founding Member of Technical Staff with expertise in front-end...Full time
$151k - $177.5k
...battery from the inside out today. We engineer and manufacture ground-breaking battery... ...engineering logic, statistical methods, optimization, machine learning, and foundation models... ...workflows for feature generation, inference, event detection, and feedback into operational...Full time$220k - $280k
...reimagine the DFS industry together? As a Staff Machine Learning Engineer, you will lead the technical charge... ...production services. Real-Time Inference at Scale: Steer the design and... ...pipelines. You will lead the creation and optimization of a centralized feature store...Full timeRemote workWork visaFlexible hours- ...AI Fabrik builds an edge inference delivery network for high-performance... ...are builders, architects, engineers, and researchers with hands‑... ...Role We're looking for a Staff Edge Network Engineer to own... ...Apply provider middle‑mile optimization (e.g., Cloudflare Argo, Akamai...Work at office
$175k - $250k
...from real-time market outcomes. The Staff AI Engineer will be responsible for moving beyond... ...Key Responsibilities: Learning & Optimization Feedback Loop Implementation: Design... ...of proven strategies. Model & Inference Infrastructure Ownership: Transition...Remote jobFull timeImmediate startShift work- ...are looking for an exceptional Search/AI Engineer with experience in Search Relevance to... ...will lead the design, development, and optimization of intelligent search systems that leverage... ...language, resolve ambiguity, and infer user intent Design and deploy learning...Full timeRemote workFlexible hours
$227k - $300k
...Edge. We are looking for a great Senior Staff AI Engineer to join our seasoned AI team and lead... ...-constrained edge devices and model optimization. You will work in a fast-paced startup... ...knowledge of modern C++ (C++14/17 for inference). ~ Deep proficiency with PyTorch or...Full timeWork at officeWorldwideFlexible hoursShift work3 days per week- ...are seeking an experienced AI Engineer with deep expertise in... ...to join our team as a Senior Staff Architect. In this role, you... ...the design, development, and optimization of cutting-edge RL solutions,... ...optimization techniques , including inference-time search, chain-of-thought...Full timeWorldwide
$178.4k - $267.6k
...Qualcomm Technologies, Inc. Job Area: Engineering Group, Engineering Group Machine... ...acceleration, model quantization, edge inference and related fields. Come join us on... ...Architect, design, develop and test model optimization techniques that include - but are not limited...Work experience placementWork from home$140k - $165k
...the Role: We are seeking a hands-on AI Engineer to design, deploy, and maintain on-prem... ...Kubernetes, Helm, Docker). Build and optimize RAG pipelines — including document... ...specific tasks; optimize for edge/on-device inference. Implement Model Control Protocols (MCP...Full time- ...Job Overview As a Staff AI Engineer, you define and drive the architecture of AI and agentic... ...design of large‑scale AI/LLM systems: inference platforms, APIs, and distributed architectures... ...to capacity planning and cost optimization strategies for AI/LLM infrastructure GenAI...Work at officeRemote workFlexible hours
$190k - $250k
...patient-facing agentic platform to optimize patient outcomes through a... .... We are looking for expert engineers and leads to join our team... ...Responsibilities As a Staff Backend AI Engineer, you will... ...infrastructure, from low-latency inference pipelines to workflow automation...Full timeWork experience placementRemote work$157k - $200k
...performance, security, and scalability. Establish and enforce engineering best practices, including rigorous code reviews, automated testing... ...teams to improve system stability, enhance observability, and optimize incident response procedures. Drive AI adoption by identifying...Full timeTemporary workWork at officeImmediate startRemote work$206.32k - $221.4k
...Description CyberSource Corporation, a Visa Inc. company, needs a Staff SW Engineer (multiple openings) in Foster City, CA to Responsible for... .... Formulate methods to enable consistent data loading and optimize data operations. Monitor health of platforms, generate...Work at officeLocal areaRemote work2 days per week3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Engineer, Inference Optimizations. Be the first to apply!
- assistant engineer Remote
- staff design engineer Remote
- assistant engineering manager Remote
- senior staff systems engineer Remote
- assistant chief engineer Remote
- staff data engineer Remote
- project engineer assistant project manager Remote
- engineering aide Remote
- senior staff engineer Remote
- staff security engineer Remote




