Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff - Kernels & GPU Performance

$150k - $350k

Gimlet Labs, Inc.

About Us Gimlet Labs is building the first heterogeneous neocloud for AI workloads. As AI systems scale, the industry is hitting fundamental limits in power, capacity, and cost with today’s homogeneous, vertically integrated infrastructure. Gimlet addresses this by decoupling AI workloads from the underlying hardware. Our platform intelligently partitions workloads into components and orchestrates each component to hardware that best fits its performance and efficiency needs. This approach enables heterogeneous systems across multi-vendor and multi-generation hardware, including the latest emerging accelerators. These systems unlock step‑function improvements in performance and cost efficiency at scale. On top of this foundation, Gimlet is building a production‑grade neocloud for agentic workloads. Customers use Gimlet to deploy and manage their workloads through stable, production‑ready APIs, without having to reason about hardware selection, placement, or low‑level performance optimization. Gimlet works with foundation labs, hyperscalers, and AI native companies to power real production workloads built to scale to gigawatt‑class AI datacenters. Mission Gimlet Labs is seeking a Member of Technical Staff focused on kernels and GPU performance. In this role, you will work close to accelerators and execution hardware to extract maximum performance from AI workloads across diverse and rapidly evolving platforms. You will analyze low‑level execution behavior, design and optimize kernels, and ensure performance is reliable across both established and emerging hardware. This role is ideal for engineers who enjoy deep performance work, reasoning about hardware tradeoffs, and turning theoretical peak performance into real‑world results. Responsibilities Design, implement, and optimize GPU and accelerator kernels for AI workloads Analyze and tune performance across the GPU execution stack, including memory access patterns, synchronization, and instruction scheduling Work with compilers and runtimes to ensure kernels integrate cleanly and perform well in end‑to‑end systems Bring up and optimize execution on new or emerging accelerators Profile, benchmark, and debug performance issues across kernels, runtimes, and hardware Ensure performance optimizations are robust, correct, and production‑ready at scale Qualifications Strong software engineering fundamentals Experience working on performance‑critical systems close to hardware Comfort reasoning about low‑level execution behavior, memory hierarchies, and performance tradeoffs Preferred Qualifications Experience with CUDA, Triton, CUTLASS, or other accelerator programming models Deep understanding of GPU execution models (warps/wavefronts, blocks, grids) Experience optimizing memory access patterns (coalescing, shared memory, cache behavior) Familiarity with occupancy, latency hiding, and instruction‑level parallelism Experience using profiling and performance analysis tools Familiarity with multi‑GPU or distributed execution is a plus Compensation Range: $150K - $350K #J-18808-Ljbffr Gimlet Labs, Inc.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff - Kernels & GPU Performance in San Francisco, CA vacancy
  • $285k - $315k

     ...building the future of high-performance compute We firmly believe...  ...this, we are building our Kernel Optimizer, which takes code...  ...Role We build the fastest GPU compiler in the world. Most...  ...production kernels. We're hiring a Member of Technical Staff for GPU Kernel Engineering... 
    Performance
    Full time
    Work at office
    Immediate start
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    1 day ago
  • $225k

     ...compute to achieve this goal. About the role As a Kernel Engineer, you will design, implement, and maintain high-performance kernels to optimize throughput and latency...  ...Blackwell or Google TPUs Develop and optimize GPU kernels in frameworks such as NCCL, MSCCLPP, CUTLASS... 
    Performance
    Relocation
    Visa sponsorship

    Magic

    San Francisco, CA
    1 day ago
  • $150k - $300k

     .... As our Solutions Architect for GPU Infrastructure, you'll be the technical expert who transforms customer requirements...  ...workloads Implement high‑performance networking with InfiniBand, RoCE,...  ...performance Tune system performance from kernel parameters to CUDA configurations... 
    Performance

    Prime Intellect

    San Francisco, CA
    21 hours ago
  •  ...language model training and inference. You will develop high‑performance ML kernels, enable efficient low‑precision arithmetic, and improve the...  ...multiplication, gating, and normalization, optimized for modern GPU architectures. Design compute primitives to reduce memory... 
    Performance

    Inception

    San Francisco, CA
    1 day ago
  •  ...datacentre demand and supply, GPU total cost of...  ..., distils our deep technical research and knowledge...  ...for a highly motivated member of technical staff to join our...  ...architecture & model performance for the next generation...  ...as an ML engineer or kernel programmer Growth Areas... 
    Performance
    Full time
    Work at office
    Remote work
    Worldwide

    S27a

    San Francisco, CA
    3 days ago
  •  ...to state-of-the-art AI Models, CUDA kernels, and GPU cloud infrastructure. We are...  ...benchmark that continuous benchmarks performance of popular frontier models MI300X vs...  ...seeking a highly motivated & skilled Member of Technical Staff to join our growing engineering team... 
    Performance
    Full time
    Work at office
    Remote work
    Worldwide

    S27a

    San Francisco, CA
    3 days ago
  • Member of the Technical Staff, Systems Location: North America Remote / San Francisco...  ...management and billions of GPU-hours supported, on...  ...designing and building high performance systems across, but not limited...  ...low-level Linux foundations: kernel, drivers, filesystems, containers... 
    Performance
    Full time
    Remote work

    Andromeda

    San Francisco, CA
    1 day ago
  • $200k - $260k

    Member of Technical Staff, Inference Engine You will build the core of our inference engine: the runtime...  ...ours from first principles: custom GPU kernels, a purpose-built runtime, and a...  ...every release honest. 5+ years building performance-critical systems in C++, Rust, CUDA,... 
    Performance

    ATBF Labs Inc.

    San Francisco, CA
    3 days ago
  • $275k - $315k

     ...building the future of high-performance compute We firmly believe...  ...this, we are building our Kernel Optimizer, which takes code...  ...Role We build the fastest GPU compiler in the world. Most...  ...which is why we're hiring a Member of Technical Staff for Sandbox Infrastructure... 
    Performance
    Full time
    Work at office
    Relocation
    Relocation package

    SF Tensor

    San Francisco, CA
    3 days ago
  • $275k - $315k

     ...building the future of high-performance compute We firmly believe...  ...this, we are building our Kernel Optimizer, which takes code...  ...Role We build the fastest GPU compiler in the world. Most...  ...deployable models. We're hiring a Member of Technical Staff to own the modeling side of... 
    Performance
    Full time
    Work at office
    Relocation package

    SF Tensor

    San Francisco, CA
    3 days ago
  • $200k - $350k

    Member of Technical Staff — LLM Research & Training About the Role We are looking for an exceptional...  ...infrastructure, post‑training, and GPU/kernel optimization. We are particularly...  ...distributed infrastructure, or low‑level performance optimization. You will work in a... 
    Performance
    H1b
    Visa sponsorship

    Pragmatike

    San Francisco, CA
    2 days ago
  • $150k - $300k

     ...training stack. Core Technical Responsibilities LLM Serving...  ...across our cloud GPU fleets. GPU‑Aware...  ...Inference Optimization & Performance Framework Development...  ...Performance: Profile kernels, memory bandwidth and...  ...development and encourage team members to contribute to the... 
    Performance
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    1 day ago
  •  ...With a focused team, breakthrough performance doesn’t require breakthrough compute...  ...that matter, and join the team. As a member of technical staff with a focus on multimodal AI, you will...  ...: Experience in writing efficient GPU kernels using CUDA, optimising performance... 
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    San Francisco, CA
    2 days ago
  • $150k - $300k

     ...distributed system with performance engineering at its...  ...skills, from deep Linux kernel topics to high-level distributed...  ...at scale. Core Technical Responsibilities...  ...heterogeneous hardware (CPU, GPU, TPU) Platform...  ...development and encourage team members to contribute to the... 
    Performance
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Prime Intellect, Inc.

    San Francisco, CA
    2 days ago
  • $200k - $400k

     ...run. About the Role As a Member of Technical Staff in Research Infrastructure,...  ...to find where the FLOPs and GPU memory are going, and others...  ...complex codebase for both performance and architectural integrity...  ...there quickly. Writing custom kernels is a bonus here, not the... 
    Performance
    Live in
    Flexible hours

    Simile

    San Francisco, CA
    1 day ago
  • Member of Technical Staff, Model Efficiency Who are we? Our mission is to scale intelligence...  ...inference stack to improve core performance metrics by diving deep into model...  ...performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model... 
    Performance
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    2 days ago
  •  ...component to hardware that best fits its performance and efficiency needs. This approach...  .... Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference....  ...boundaries Work closely with compilers, kernels, networking, and distributed systems... 
    Performance

    Gimlet Labs

    San Francisco, CA
    4 days ago
  •  ...Job Title Member of Technical Staff: Infrastructure Salary Not Disclosed Company Description Observable Intuition is an early‑stage...  ...platform, optimizing model serving, orchestration, and GPU workloads for performance, latency, and cost. Architect a portable, multi‑... 
    Performance

    Jack & Jill

    San Francisco, CA
    3 days ago
  •  ..., enabling step-function improvements in performance and efficiency. Customers deploy through...  ...Design, deploy, and operate large‑scale CPU, GPU, and accelerator clusters powering...  ...performance, networking, storage, processes, and kernel‑level issues. Experience operating... 
    Performance

    Gimlet Labs

    San Francisco, CA
    3 hours ago
  •  ...responsibility to defend. About the role As a Member of Technical Staff, ML Product Engineer, this role owns...  ...layer: job queues, autoscaling GPU inference, retries, fault tolerance, and...  ...and latency per inference, and make performance predictable enough to price and guarantee... 
    Performance
    Local area

    Radical Numerics

    San Francisco, CA
    3 hours ago
  • $180k - $290k

    Member of Technical Staff - Research Engineer About Black Forest Labs We're the team...  ..., and correctly across large GPU fleets. In this role, you will...  ...the hardest systems and performance problems arise: attention performance, custom kernels, low-precision training, profiling... 
    Performance
    Remote work
    Worldwide
    2 days per week

    BlackForestLabs

    San Francisco, CA
    21 hours ago
  • $225k - $275k

     ...building the future of high-performance compute We firmly believe...  ...this, we are building our Kernel Optimizer, which takes code...  ...Role We build the fastest GPU compiler in the world. Most...  ...though, so we're hiring a Member of Technical Staff for Product Engineering to... 
    Performance
    16 hours
    Full time
    Live in
    Work at office
    Relocation package
    Night shift

    San Francisco Tensor Company

    San Francisco, CA
    1 day ago
  • $250k - $300k

     ...built a platform that deploys GPU clusters into third-party...  ...designing and building high-performance systems spanning storage, networking...  ...of the Linux stack, across kernel, drivers, filesystems,...  ...A track record of impressive technical work you can speak to in depth... 
    Performance
    Full time
    Remote work
    San Francisco, CA
    a month ago
  •  ...Job Description Job Description Member of Technical Staff, Machine Learning, Artificial Intelligence...  ...data. - Debug model issues, performance problems, and production incidents....  ...continuous improvement. - Tech Stack:  GPU, JAX, ML, Machine Learning, Python, and... 
    Performance
    Remote work
    Work from home

    Ginas Tech Jobs

    San Francisco, CA
    3 days ago
  • $250k

     ...generation of agentic infrastructure for GPU-intensive workloads. Operating across...  ...stack, the organisation combines high-performance hardware with intelligent software to...  ...offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's growth... 
    Performance
    Full time
    San Francisco, CA
    a month ago
  •  ...processing down to the lowest layers of the stack. You'll optimize kernel performance, develop new scheduling and parallelism strategies, and help...  ...like vLLM and SGLang. Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every... 
    Performance

    Sail Research

    San Francisco, CA
    1 day ago
  • Member of Technical Staff - ML Systems & Inference Bay Area, CA | Onsite Join a well-funded AI infrastructure...  ...inference, distributed systems, and performance engineering. You'll build production...  ..., and work closely with compiler, kernel, and distributed systems engineers to... 
    Performance

    Acceler8 Talent

    San Francisco, CA
    21 hours ago
  • Job Description - Member of Technical Staff (Hardware) Location: San Francisco (on-site at our offices...  ...understand them. You’ll design and run performance benchmarks across accelerators and...  ...Etched, MatX, Tenstorrent or similar), GPU and accelerator teams at larger players... 
    Performance

    Artificial Analysis, Inc.

    San Francisco, CA
    21 hours ago
  •  ...training fast, efficient, and reliable, so that every GPU cycle accelerates research progress....  ...communication overlap. Ability to profile and debug performance in complex codebases, from framework internals down to kernels and collectives Deep understanding of deep... 
    Performance

    Uncover

    San Francisco, CA
    1 day ago
  •  ...component to hardware that best fits its performance and efficiency needs. This approach...  .... Gimlet Labs is seeking a Member of Technical Staff focused on compilers. In this role, you...  ...spanning graph-level, tensor-level, and kernel-level representations Implement partitioning... 
    Performance

    Gimlet Labs

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff - Kernels & GPU Performance. Be the first to apply!