Member of Technical Staff - Kernels & GPU Performance
Gimlet Labs
Gimlet Labs is building the first heterogeneous neocloud for AI workloads. As AI systems scale, the industry is hitting fundamental limits in power, capacity, and cost with today’s homogeneous, vertically integrated infrastructure. Gimlet addresses this by decoupling AI workloads from the underlying hardware. Our platform intelligently partitions workloads into components and orchestrates each component to hardware that best fits its performance and efficiency needs. This approach enables heterogeneous systems across multi-vendor and multi-generation hardware, including the latest emerging accelerators. These systems unlock step-function improvements in performance and cost efficiency at scale. On top of this foundation, Gimlet is building a production-grade neocloud for agentic workloads. Customers use Gimlet to deploy and manage their workloads through stable, production-ready APIs, without having to reason about hardware selection, placement, or low-level performance optimization. Gimlet works with foundation labs, hyperscalers, and AI native companies to power real production workloads built to scale to gigawatt-class AI datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on kernels and GPU performance. In this role, you will work close to accelerators and execution hardware to extract maximum performance from AI workloads across diverse and rapidly evolving platforms. You will analyze low-level execution behavior, design and optimize kernels, and ensure performance is reliable across both established and emerging hardware. This role is ideal for engineers who enjoy deep performance work, reasoning about hardware tradeoffs, and turning theoretical peak performance into real-world results. Responsibilities Design, implement, and optimize GPU and accelerator kernels for AI workloads Analyze and tune performance across the GPU execution stack, including memory access patterns, synchronization, and instruction scheduling Work with compilers and runtimes to ensure kernels integrate cleanly and perform well in end-to-end systems Bring up and optimize execution on new or emerging accelerators Profile, benchmark, and debug performance issues across kernels, runtimes, and hardware Ensure performance optimizations are robust, correct, and production-ready at scale Qualifications Strong software engineering fundamentals Experience working on performance-critical systems close to hardware Comfort reasoning about low-level execution behavior, memory hierarchies, and performance tradeoffs Preferred Qualifications Experience with CUDA, Triton, CUTLASS, or other accelerator programming models Deep understanding of GPU execution models (warps/wavefronts, blocks, grids) Experience optimizing memory access patterns (coalescing, shared memory, cache behavior) Familiarity with occupancy, latency hiding, and instruction-level parallelism Experience using profiling and performance analysis tools Familiarity with multi-GPU or distributed execution is a plus #J-18808-Ljbffr Gimlet Labs
$225k
...compute to achieve this goal. About the role As a Kernel Engineer, you will design, implement, and maintain high-performance kernels to optimize throughput and latency... ...Blackwell or Google TPUs Develop and optimize GPU kernels in frameworks such as NCCL, MSCCLPP, CUTLASS...PerformanceRelocationVisa sponsorship$150k - $300k
.... As our Solutions Architect for GPU Infrastructure, you'll be the technical expert who transforms customer requirements... ...workloads Implement high‑performance networking with InfiniBand, RoCE,... ...performance Tune system performance from kernel parameters to CUDA configurations...Performance$350k
...for an engineer to design, implement, and optimize custom ML kernels that bolster our model development stack. Your work will be deep... ...system, combining hardware and software insights to optimize performance. Some example areas you might work on (not limited to): Design...Performance$150k - $300k
...training stack. Core Technical Responsibilities LLM Serving... ...across our cloud GPU fleets. GPU‑Aware... ...Inference Optimization & Performance Framework Development... ...Performance: Profile kernels, memory bandwidth and... ...development and encourage team members to contribute to the...PerformanceWork at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work- ...With a focused team, breakthrough performance doesn’t require breakthrough compute... ...that matter, and join the team. As a member of technical staff with a focus on multimodal AI, you will... ...: Experience in writing efficient GPU kernels using CUDA, optimising performance...PerformanceFull timeWork at officeLocal areaRemote workHome office
$200k - $400k
...run. About the Role As a Member of Technical Staff in Research Infrastructure,... ...to find where the FLOPs and GPU memory are going, and others... ...complex codebase for both performance and architectural integrity... ...there quickly. Writing custom kernels is a bonus here, not the...PerformanceLive inFlexible hours$150k - $300k
...distributed system with performance engineering at its... ...skills, from deep Linux kernel topics to high-level distributed... ...at scale. Core Technical Responsibilities... ...heterogeneous hardware (CPU, GPU, TPU) Platform... ...development and encourage team members to contribute to the...PerformanceFull timeWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours- ...Performance Engineer Morph builds the inference infrastructure behind the fastest open models. Our stack spans kernels, model serving, routing, autoscaling, and capacity. We are hiring a performance... ...production systems Understand GPU performance, memory bandwidth,...PerformanceImmediate start
$250k - $300k
...built a platform that deploys GPU clusters into third-party... ...designing and building high-performance systems spanning storage, networking... ...of the Linux stack, across kernel, drivers, filesystems,... ...A track record of impressive technical work you can speak to in depth...PerformanceFull timeRemote work- ...Job Description Job Description Member of Technical Staff, Machine Learning, Artificial Intelligence... ...data. - Debug model issues, performance problems, and production incidents.... ...continuous improvement. - Tech Stack: GPU, JAX, ML, Machine Learning, Python, and...PerformanceRemote workWork from home
$250k
...generation of agentic infrastructure for GPU-intensive workloads. Operating across... ...stack, the organisation combines high-performance hardware with intelligent software to... ...offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's growth...PerformanceFull time- ...processing down to the lowest layers of the stack. You'll optimize kernel performance, develop new scheduling and parallelism strategies, and help... ...like vLLM and SGLang. Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every...Performance
- ...efficiency Dataloaders, fusion, activation remat, gradient checkpointing. FSDP/ZeRO/tensor+pipeline parallel; NCCL tuning. GPU + kernel performance Nsight profiling, Triton/CUDA kernels, fused ops. Flash-attention-style speedups, sequence packing, KV-cache tricks....Performance
- ...component to hardware that best fits its performance and efficiency needs. This approach... .... Gimlet Labs is seeking a Member of Technical Staff focused on compilers. In this role, you... ...spanning graph-level, tensor-level, and kernel-level representations Implement partitioning...Performance
- ...lowest layers of the stack. You'll optimize kernel performance, develop new request scheduling and... .... Understand every microsecond of GPU time spent during a forward pass. You'll... ...about your experience, and share as much technical detail about Sail as you want to hear....PerformanceWork at officeImmediate start
- Job Description - Member of Technical Staff (Hardware) Location: San Francisco (on-site at our offices... ...understand them. You’ll design and run performance benchmarks across accelerators and... ...Etched, MatX, Tenstorrent or similar), GPU and accelerator teams at larger players...Performance
- ...component to hardware that best fits its performance and efficiency needs. This approach... .... Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference.... ...boundaries Work closely with compilers, kernels, networking, and distributed systems...Performance
- ...training fast, efficient, and reliable, so that every GPU cycle accelerates research progress.... ...communication overlap Ability to profile and debug performance in complex codebases, from framework internals down to kernels and collectives Deep understanding of deep...Performance
$225k
...large-scale model training across massive GPU clusters. You will work at the boundary... ...systems, ensuring that training runs are performant, reliable, and reproducible under extreme... ...and training throughput Collaborate with Kernels and Research to align model architecture...PerformanceRelocationVisa sponsorship$225k
...reliable. What you’ll work on Design and scale high-performance inference serving systems Optimize KV-cache... ...Profile and eliminate performance bottlenecks across GPU, networking, and storage layers Collaborate with Kernels and Research to align execution systems with...PerformanceRelocationVisa sponsorship- ...optimizes AI itself. Our journey starts with GPU kernels, but will expand into every corner of... ...systems that help the agent diagnose performance bottlenecks Ship features that... ...You're a strong fit if you: Have deep technical intuition and can learn new domains quickly...PerformanceRemote work
- ...with a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own... ...global low-latency, high-throughput GPU ML inference infra that sits in the... ...Terraform, Docker and CI/CD, and building for performance and reliability at scale Have owned...PerformanceVisa sponsorship
$200k - $300k
...growing AI inference company in San Francisco to hire Members of Technical Staff — engineers who build the systems that make LLM... ...outright. What you'll work on Develop and optimize high-performance computing kernels, and work across inference engine internals and serving...PerformanceH1bWork at office- ...enabling step-function improvements in performance and efficiency. Customers deploy through... ...runtime systems, scheduling, memory movement, kernel orchestration, and serving optimization... ...build and run this company. As an early member of the team, you will have significant...Performance
- ..., enabling step-function improvements in performance and efficiency. Customers deploy through... ...Design, deploy, and operate large‑scale CPU, GPU, and accelerator clusters powering... ...performance, networking, storage, processes, and kernel‑level issues. Experience operating...Performance
- ..., DoorDash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage... ...work with SGLang on specdec and multimodal inference performance our work on Flash Attention 4 kernels Work with engineering to turn frontier serving...PerformanceWork at office
- ...Design, build, and operate large-scale GPU infrastructure for high-throughput model... ...learning pipelines at scale. Build high-performance inference platforms capable of serving and... ...Improve performance of model execution through kernel-level optimization, model parallelism...PerformanceRelocation package
- ...evaluations, secure sandboxes, high-performance training, and deployment into... ...they own. Core Technical Responsibilities This hybrid... ...heterogeneous hardware (CPU, GPU, TPU) Backend & Feature Development... ...development and encourage team members to contribute to the broader...Performance
- ...and backers at . About the Role As a Member of Technical Staff, you will help invent and build the next... ...Cloud computing Systems for AI or HPC Performance optimization Excellent software... ...scale AI infrastructure. Experience with GPU systems, accelerators, or performance...PerformanceWork from homeFlexible hours2 days per week
$150k - $300k
...fine-tuning runs on managed GPU clusters with a single API call... ...that runs the jobs. Core Technical Responsibilities Hosted Training... ...: networking, namespaces, performance tuning Programming & Platform... ...development and encourage team members to contribute to the broader...PerformanceWork at officeLocal areaRemote workVisa sponsorshipRelocation packageFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff - Kernels & GPU Performance. Be the first to apply!
- work from home technical support specialist San Francisco, CA
- product support technician San Francisco, CA
- helpdesk support technician San Francisco, CA
- help desk assistant San Francisco, CA
- senior technical associate San Francisco, CA
- IT help desk technician San Francisco, CA
- technical solutions specialist San Francisco, CA
- desktop support analyst San Francisco, CA
- trade support analyst San Francisco, CA
- senior IT support technician San Francisco, CA


