Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff, TPU & AMD GPU Performance Engineering

$200k - $400k
Full-time

Inferact

Role Description

We're looking for a TPU and AMD GPU performance engineer to make vLLM a first-class inference engine across non-NVIDIA accelerators. Frontier inference cannot be locked to one hardware stack. As AMD GPUs, TPUs, and other accelerators become increasingly important, vLLM needs backend paths that are fast, correct, benchmarked, and maintainable across heterogeneous hardware platforms.

  • Build and optimize AMD GPU and TPU backends, kernels, compiler integrations, runtime paths, and benchmarking infrastructure.
  • Work at the boundary of inference systems, kernels, compilers, and hardware architecture.
  • Improve paths such as attention, GEMM, sampling, KV-cache, communication-heavy operations, and model serving on non-NVIDIA hardware.
  • Your work will directly impact how broadly and efficiently the world can run AI inference with vLLM.

Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, machine learning systems, hardware systems, compilers, or similar.
  • Hands-on experience optimizing workloads on AMD GPUs, TPUs, or another non-NVIDIA accelerator stack.
  • Experience with AMD ecosystem tools such as ROCm, HIP, Triton, CK, AITER, or equivalent GPU performance libraries and tooling.
  • Experience with TPU, XLA, JAX, Pallas, or related compiler and runtime tooling for accelerator workloads.
  • Ability to optimize ML inference paths such as attention, GEMM, sampling, KV-cache, fused kernels, backend runtimes, or communication-heavy operations.
  • Strong performance profiling and benchmarking discipline, including tokens/second, latency, throughput, correctness parity, hardware counters, and reproducible measurement methodology.
  • Ability to navigate immature tooling, incomplete documentation, backend-specific rough edges, and cross-platform performance differences without getting stuck.

Requirements

  • Experience with vLLM, SGLang, TensorRT-LLM, ATOM, JAX-based serving framework, or other LLM inference systems.
  • Deep understanding of inference architecture and serving tradeoffs, including batching, KV-cache, decoding, prefill/decode scheduling, and backend performance constraints.
  • Experience with compiler technologies such as XLA, MLIR, LLVM, Triton, Pallas, or other compiler/kernel DSLs, including lowering, fusion, and backend code generation.
  • Knowledge of quantization techniques such as MXFP8, MXFP4, mixed precision, or hardware-specific numeric formats, and the ability to reason about accuracy/performance tradeoffs.
  • Experience with distributed inference performance, including communication, memory movement, hardware topology, and scale-out bottlenecks across multi-accelerator workloads.
  • Open-source contributions to vLLM, JAX/XLA, ROCm, Triton, PyTorch, compiler projects, or related ML systems infrastructure.

Benefits

  • Generous health, dental, and vision benefits.
  • 401(k) company match.

Logistics

  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
  • Visa sponsorship: We sponsor visas on a case-by-case basis.
Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff, TPU & AMD GPU Performance Engineering in Remote vacancy
  • Role Description We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton... 
    Performance
    Full time

    Inferact

    Remote
    a month ago
  • $150k - $300k

     ...On-site Department Engineering Building Open Superintelligence...  ...system with performance engineering at its...  ...reliable at scale. Core Technical Responsibilities...  ...heterogeneous hardware (CPU, GPU, TPU) Platform...  ...development and encourage team members to contribute to the... 
    Performance
    Full time
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Menlo Ventures

    San Francisco, CA
    2 days ago
  • Member of Technical Staff - Engineer, RL and Control Position Summary You will help build and refine the learned...  ...tools for evaluating controller performance in simulation and on hardware. Collaboration...  ...-of-the-art machine learning and GPU programming frameworks, such as... 
    Performance
    Work from home
    Flexible hours

    Walden Robotics, Inc.

    Cambridge, MA
    2 days ago
  •  ...Consulting Member Of Technical Staff As a Consulting Member of Technical Staff, you will be a...  ...technologies. Your expertise in data platform engineering and service development will drive...  ...decisions where analysis of data, performance, privacy, security, and healthcare... 
    Performance
    Immediate start
    Remote work
    Worldwide

    Hackajob

    United States
    3 days ago
  •  ...inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: ~...  ...(e.g. Kueue, KAI, KServe) ~You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled... 
    Performance
    Full time

    SkyPilot

    Remote
    a month ago
  • $250k - $300k

     ...built a platform that deploys GPU clusters into third-party...  ...to GB300. As part of the engineering team, you'll help shape a platform...  ...designing and building high-performance systems spanning storage, networking...  ...A track record of impressive technical work you can speak to in... 
    Performance
    Full time
    Remote work
    San Francisco, CA
    22 days ago
  • Job Role: Member of Technical Staff (Austin, TX Remote) The Member of Technical Staff (MTS) - Systems...  ...and expertise for complex systems engineering projects. MTS engineers serve as...  ...patches. Ensure kernel stability and performance. Feature Development Design and implement... 
    Performance
    Remote work

    ALTEN

    Austin, TX
    4 days ago
  • $110.4k - $165.5k

     ...problems and providing unmatched technical expertise. As the operator...  ...for an RF/Microwave Engineer who wants their work to...  ...complex systems – Generate performance budgets and run circuit simulations...  ...to be Successful – Senior Member of the Technical Staff Minimum Requirements:... 
    Performance
    Full time
    For contractors
    Work at office
    Immediate start
    Remote work
    Relocation package
    Flexible hours

    AERO

    El Segundo, CA
    4 days ago
  • $110.4k - $165.5k

     ...RF/Microwave Engineer The Aerospace Corporation is the trusted...  ...and providing unmatched technical expertise. As the operator...  ...systems – Generate performance budgets and run circuit simulations...  ...to be Successful – Senior Member of the Technical Staff Minimum Requirements:... 
    Performance
    Full time
    For contractors
    Work at office
    Immediate start
    Remote work
    Relocation package
    Flexible hours

    The Aerospace Corporation

    El Segundo, CA
    2 days ago
  •  ...enterprise finance workflows. As a Member of Technical Staff in Finance Research, you will develop...  ...agentic systems, working across research, engineering, and product teams in a remote, full...  ...turning findings into measurable AI performance improvements. Develop and curate... 
    Performance
    Full time
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • $180k - $238.1k

     ...Member of Technical Staff (Performance Engineering) New York, NY Category-defining tech. Career-defining work. Lots of tech companies disrupt. But, many fail when they try to scale. We're different. CockroachDB makes it easier for companies to build and scale... 
    Performance
    Local area
    Remote work
    Worldwide
    Flexible hours

    Cockroach Labs

    New York, NY
    4 days ago
  • Member of Technical Staff: Mechanical Engineering Position Summary: We are hiring a Mechanical Engineer to design, build, and ship high-performance actuators for our robot. In this role, you will own the design of defined actuator mechanical components and subassemblies... 
    Performance
    Work from home
    Flexible hours

    Walden Robotics, Inc.

    Cambridge, MA
    2 days ago
  • $95.2k - $165.5k

     ...problems and providing unmatched technical expertise. As the operator of a...  ...MPD, the Communication Systems Engineering Department (CSED) performs analysis, modeling and simulation...  ...Communications Systems Engineer (Member of Technical Staff/Senior Member of Technical Staff... 
    Performance
    Full time
    Immediate start
    Remote work
    Relocation package
    Flexible hours

    aero

    El Segundo, CA
    3 days ago
  •  ...About the job TL;DR: A founding-style, full-stack engineer who ships end to end across a high-performance Go backend and a cross-platform iOS and Android app (React Native and Expo), with AI agents as your force multiplier. If mobile UI/UX is your strength, you can... 
    Performance
    Remote work
    Flexible hours
    Shift work

    Andreessen Horowitz

    New York, NY
    4 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for...  ...As a founding member of the engineering team, you will impact the design...  ...compute resources (CPU and GPU) efficiently? What data...  ...to building a high-performing, inclusive team with a professional... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    1 day ago
  •  ...Role Overview Lead technical ownership at the intersection of research, data, and deployed...  ...AI systems to improve model and agent performance through rigorous evaluation, failure...  ...domains such as finance, healthcare, and engineering. Key Responsibilities Own... 
    Performance
    Full time
    Remote work

    SaidGig

    United States
    26 days ago
  •  ...systems. We are seeking a Senior Member of Technical Staff to build core systems, solve challenging engineering problems, and contribute...  ...internal tooling. ~Solve complex performance, reliability, and scalability...  ...or scale-up experience. ~GPU or data-intensive systems.... 
    Performance
    Full time

    Pragmatike

    Remote
    14 days ago
  • Member of the Technical Staff, Systems Location: North America Remote / San Francisco...  ...and billions of GPU-hours supported, on everything...  ...We are looking for strong engineers with experience and interest...  ...designing and building high performance systems across, but not... 
    Performance
    Full time
    Remote work

    Andromeda

    San Francisco, CA
    1 day ago
  •  ...step-function improvements in performance and efficiency. Customers deploy...  ...an Infrastructure platform Engineer to design, build, and operate...  ...and operate large‑scale CPU, GPU, and accelerator clusters powering...  ...environments across NVIDIA, AMD, Intel, ARM, or emerging accelerators... 
    Performance

    Gimlet Labs

    San Francisco, CA
    21 hours ago
  •  ...new physics at scale. We are seeking engineers to build platform infrastructure at...  ...physics, at scale. We are seeking a Member of Technical Staff, Performance & Capacity to answer two questions honestly...  ...Looking For Five or more years with GPU and large-scale compute workloads,... 
    Performance
    Remote work

    Physical Superintelligence

    Boston, MA
    1 day ago
  • Mount Thor is hiring a software engineer with deep networking expertise to build and operate...  ...(macOS and Apple Silicon) available and performant at datacenter scale for AI workloads. We...  ...role you will Own the architecture and technical roadmap for Mount Thor’s production... 
    Performance

    Mount Thor

    San Francisco, CA
    4 days ago
  •  ...building AI systems to discover new physics at scale. We are seeking engineers to build platform infrastructure at the intersection of...  ...dynamics, FEA, quantum, or comparable), with comfort in high‑performance computing environments. You can read and write production code... 
    Performance
    Remote work

    Physical Superintelligence

    Boston, MA
    4 days ago
  • $230k

     ...over 10 times faster than GPU-based hyperscale cloud inference...  ...multiple openings for Sr. Member of Technical Staff. Title: Sr. Member of...  ...latency and scalable system performance. Develop Python-based scripts...  ...Security Analyst, Software Engineer, Sr. Member of Technical... 
    Performance
    Remote work

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  •  ...is a team of researchers, engineers, designers, and more, who are...  ...focused team, breakthrough performance doesn’t require...  ..., and join the team. As a member of technical staff with a focus on multimodal...  ...Experience in writing efficient GPU kernels using CUDA, optimising... 
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    2 days ago
  • Member of Technical Staff - AI Cloud Infrastructure Location Bay Area; Boston; Washington...  ...a senior infrastructure engineer to architect and stand up...  ...AI platform, whether at a GPU cloud, an internal machine...  ...clouds. Lead high-performance storage strategy. Deploy and... 
    Performance
    Full time
    Work from home
    Flexible hours
    2 days per week

    Emerald AI, Inc.

    Washington DC
    1 day ago
  • $150k - $300k

     ...training stack. Core Technical Responsibilities LLM Serving...  ...across our cloud GPU fleets. GPU‑Aware...  ...Inference Optimization & Performance Framework Development...  ...PyTorch: LLM Inference engine development and...  ...development and encourage team members to contribute to the... 
    Performance
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    1 day ago
  •  ...at About the Role As a Member of Technical Staff, you will help invent and build...  ...with hands‑on software engineering experience and are excited...  ...computing Systems for AI or HPC Performance optimization Excellent...  .... Experience with GPU systems, accelerators, or performance... 
    Performance
    Work from home
    Flexible hours
    2 days per week

    Emerald AI

    Boston, MA
    2 days ago
  • $150k - $300k

     ...tuning runs on managed GPU clusters with a single...  ...runs the jobs. Core Technical Responsibilities Hosted...  ...We're looking for engineers who are fluent across...  ...networking, namespaces, performance tuning Programming &...  ...development and encourage team members to contribute to the... 
    Performance
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Kubelt

    San Francisco, CA
    2 days ago
  •  ...systems. We are seeking a Staff level Member of Technical Staff to lead high-impact...  ...initiatives and solve complex engineering problems spanning multiple...  ..., reliability, and performance bottlenecks. ~Establish...  ...experience (Nice to Have). ~GPU or data-intensive systems... 
    Performance
    Full time

    Pragmatike

    Remote
    14 days ago
  •  ...Description Job Description Member of Technical Staff, Machine Learning,...  ...Technical Staff role is for engineers who want to develop strong...  ...data. - Debug model issues, performance problems, and production...  ...improvement. - Tech Stack:  GPU, JAX, ML, Machine Learning,... 
    Performance
    Remote work
    Work from home

    Ginas Tech Jobs

    San Francisco, CA
    11 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff, TPU & AMD GPU Performance Engineering. Be the first to apply!