Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Engineer - AI Performance & Kernel Optimization

Zyphra

Job Description

Job Description

Zyphra is an artificial intelligence company based in San Francisco, California.

The Role:

As a Research Engineer - AI Performance & Kernel Optimization , you will improve and optimize the performance of our large-scale language model training and inference stacks. You will work closely with our pretraining and inference teams to identify bottlenecks, design and implement highly optimized kernels, and push the limits of throughput, latency, and hardware utilization across a range of accelerator platforms. This role is suited for someone who enjoys deep systems work, cares about performance at every level of the stack, and is excited to translate low-level optimizations into meaningful gains for frontier-scale AI systems.

You’ll Work Across:
  • Kernel development and optimization for large-scale ML workloads, using any level of the stack from PTX/assembly to CUDA, HIP, Triton, or other GPU DSLs

  • Performance tuning for training and inference stacks across GPUs and other accelerators

  • Profiling and eliminating bottlenecks in memory movement, communication, scheduling, and compute utilization

  • Optimizing distributed training and inference systems for large MoE models, including large-scale model parallelism

  • Portability and optimization across non-NVIDIA hardware, with special interest in AMD hardware such as the MI300x and MI355x

  • Collaboration with research and infrastructure teams to turn systems improvements into real-world model training and inference gains

What We're Looking For / Requirements:
  • Strong engineering aptitude for building reliable, high-performance systems

  • Excellent low-level performance intuition and the ability to reason about hardware-software interactions

  • Are excited to rapidly learn new systems, tools, and hardware environments

  • Excellent communication and collaboration skills, with the ability to work effectively across research and engineering teams

  • Enjoy diving deep into the weeds and hunting down the last 10–20% of performance

Qualifications / Additional Skills:
  • Experience writing highly performant GPU kernels at any level of abstraction–PTX, CUDA, HIP, Triton, or other kernel DSLs

  • Experience optimizing ML workloads for large-scale training, ideally in language model pretraining or inference environments

  • Experience with non-NVIDIA accelerator hardware, such as AMD, AWS Trainium, Google TPU, Qualcomm, ARM, Intel, and custom ASICs

  • Strong understanding of distributed training systems and parallelism schemes, including data parallelism, tensor/model parallelism, pipeline parallelism, sharding, and communication/computation overlap

  • Experience with performance engineering in other demanding parallel computing environments such as HPC, quantitative finance, scientific computing, graphics, compilers, or numerical simulation

  • Strong systems intuition around memory hierarchy, bandwidth constraints, kernel fusion, launch overhead, communication overhead, and hardware utilization

  • Experience using profiling and debugging tools to drive performance improvements

  • Familiarity with infrastructure underlying large-scale training and inference, including collective communication libraries, and runtime performance analysis

  • Background in a highly technical field such as physics, mathematics, theoretical computer science, computer science, or electrical engineering

  • Any HPC experience is a strong plus

Why Work at Zyphra:
  • Our research methodology is grounded in methodical, step-by-step approaches to ambitious goals. Both deep research and engineering excellence are equally valued

  • We strongly value new and crazy ideas and are very willing to bet big on new ideas

  • We move as quickly as we can; we aim to minimize the bar to impact as low as possible

  • We all enjoy what we do and love discussing AI

Benefits and Perks:
  • Comprehensive medical, dental, vision, and FSA plans

  • Competitive compensation and 401(k) plan

  • Relocation and immigration support on a case-by-case basis

  • In-office snacks and meals provided

  • Unlimited PTO and company holidays

  • In-person team in San Francisco with a collaborative, high-energy environment

Vacancy posted 18 days ago
Similar jobs that could be interesting for youBased on the Research Engineer - AI Performance & Kernel Optimization in San Francisco, CA vacancy
  • $120k - $200k

     ...We are actively seeking a Research Engineer specializing in Machine Learning and AI to play a pivotal role in pioneering...  ...improvement processes to optimize the performance and accuracy of machine learning...  ...level optimizations, such as CUDA kernel programming, is highly... 
    Performance
    Casual work
    Work at office

    Erth.AI Inc.

    San Francisco, CA
    2 days ago
  •  ...methods cannot reach. AI is reinventing life...  ...software engineering, and Chai is at the...  ...are seeking an AI Research Engineer to help design...  ...train, evaluate, and optimize Chai's core models...  ...backgrounds include high-performance computing, custom CUDA kernels and GPU programming... 
    Performance
    Shift work

    Chai Discovery, Inc

    San Francisco, CA
    5 days ago
  • $120k - $250k

    WHO WE ARE Lightning AI is the company behind PyTorch Lightning...  ...to take ideas from research to production with less friction...  ...a highly skilled Research Engineer to help optimize training and inference...  ...systems, AI infrastructure, performance engineering, and practical... 
    Performance
    Full time
    Work at office
    Remote work
    Work from home
    Flexible hours
    2 days per week

    Lightning AI

    San Francisco, CA
    2 days ago
  • $175k - $250k

     ...team to push the boundaries of AI research and development. Their mission...  ...The Role: As a Research Engineer in Pre-Training, you'll develop...  ...What You’ll Do: Design and optimize novel pre-training methods to improve model performance. Conduct large-scale training... 
    Performance
    Full time
    Relocation package

    HartleyCo

    San Francisco, CA
    6 hours ago
  •  ...Join to apply for the Research Engineer role at Jobright.ai 2 days ago Be among the first 25 applicants Join...  ...Engineer, you will design, implement, and optimize large-scale ML systems,...  ...translate algorithmic ideas into robust, performant systems. • Own projects end-to-end—... 
    Performance
    Full time
    Internship
    H1b

    jobright.com

    San Francisco, CA
    3 days ago
  • $225k - $400k

     ...Research Engineer Title of Role: Research Engineer Location: San Francisco, onsite Company...  ...-Backed — Software Development, AI, Devtools, Data, Enterprise, B2B Office...  ...techniques to extract insights and optimize performance metrics. Conduct experiments and A/... 
    Performance
    Work at office

    Recruiting from Scratch

    San Francisco, CA
    1 day ago
  •  ...Research Engineer On Physical Ai Team Hedra is a pioneering generative modeling company — first models...  ...action sequences Evaluate model performance using both benchmark datasets and...  ...fundamentals in machine learning, optimization, and large-scale data processing... 
    Performance
    Work at office

    HEDRA INC

    San Francisco, CA
    1 day ago
  • Get AI-powered advice on this job and more exclusive...  ...empowers software engineers by automating...  ...end‑to‑end, balancing research and engineering to create...  ...AI models Build and optimize data pipelines to process...  ...evaluate AI models, improve performance, and reduce compute... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Resolve AI

    San Francisco, CA
    3 days ago
  • $159k - $296k

     ...Description Waabi, founded by AI visionary Raquel...  ...driving trucks. As a research engineer for Learnable Planner...  ...learning, optimization-based approaches, search...  ...like TensorRT, CUDA kernels. The US yearly salary...  ...awards and an annual performance bonus. Perks/Benefits... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    17 days ago
  • $250k - $300k

     ...opening. ZeroRFI is the AI company building the...  ...a Principal Software Engineer to architect the systems...  ...vision systems, and optimization models that fundamentally...  ...apply cutting-edge AI research to one of humanity's...  ...networks for building performance simulation and predictive... 
    Performance
    Full time
    Work at office
    Home office

    Zero Rfi

    San Francisco, CA
    1 day ago
  • ML/AI Research Engineer — Agentic AI Lab (Founding Team) Location: San Francisco Bay Area Type:...  ...the design, training, evaluation, and optimization of agent‑native AI models. You'll...  ...evaluation harnesses for LLM and agent performance, including synthetic evals, trace capture... 
    Performance
    Full time

    Fabrion

    San Francisco, CA
    6 hours ago
  •  ...Senior Experimental Research Engineer, Electromagnetics Senior Experimental Research Engineer, Electromagnetics...  ...code to analyze experimental results, perform exploratory data analysis, and model...  ...directly into each article, started with the help of AI. #J-18808-Ljbffr... 
    Performance
    Full time

    Gridware

    San Francisco, CA
    6 hours ago
  •  ...Beta, Felicis, Figma Ventures, AI Grant, and more, and we're...  ...accomplished a lot as just one engineer and one designer!) to a clan...  ...breakthrough AI service with leading performance on relevant benchmarks,...  ...". Specifically, in an AI research engineer role, we are looking... 
    Performance

    Poly

    San Francisco, CA
    2 days ago
  •  ...Ando is a messaging platform where AI agents take on work alongside their human teammates...  ...about as well. If you want to have your research come into contact with reality, Ando is...  ..., retrieval, and context impact agent performance, vs where do we genuinely need RL and continual... 
    Performance
    Work from home

    Ando

    San Francisco, CA
    5 days ago
  • $200k - $400k

     ...is the leading conversational AI platform empowering every brand...  ...the Team Read more about the research team's work here: The...  ...About the Role As a Research Engineer, you’ll be responsible for building...  ...end models and pipelines that optimize for quality, efficiency, and user... 
    Full time
    Work at office
    Local area

    Decagon

    San Francisco, CA
    4 days ago
  •  ...Description We are Genmo, a research lab dedicated to building...  ...us in shaping the future of AI and pushing the boundaries...  ...seeking an exceptional Software Engineer to join our research team...  ...GANs, Transformers) Help optimize model performance and scaling capabilities... 
    Performance
    Work at office

    Genmo

    San Francisco, CA
    5 days ago
  •  ...Francisco, California. The Role: As a Research Engineer - Brain Computer Interface Models ,...  ...for EEG and other BCI modalities Performance optimization of the training stack Integration...  ...enjoy what we do and love discussing AI Benefits and Perks:... 
    Performance
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    18 days ago
  •  ...Francisco, California. The Role: As a Research Engineer - Audio & Speech Models , you will...  ...Large-scale audio training runs Performance optimization of our training stack Audio...  ...enjoy what we do and love discussing AI Benefits and Perks: Comprehensive... 
    Performance
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    7 days ago
  • $155k - $269k

     ...Description Waabi, founded by AI visionary Raquel Urtasun, is...  ...efficient simulation. As a Research Engineer in the World Models team,...  ...and inference pipelines. - Optimize model training and inference...  ...incentive awards and an annual performance bonus.   Perks/Benefits:... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    1 day ago
  •  ...A leading conversational AI platform in San Francisco seeks an AI/ML Engineer to build advanced systems for unprecedented performance. The ideal candidate will have over 8 years of experience and a strong track record in AI/ML projects. You'll design state-of-the-art... 
    Performance
    Full time

    Decagon

    San Francisco, CA
    6 hours ago
  • $180.6k - $315k

     ...AI is becoming vitally important in every function of our society. At Scale, our mission...  ...post-training algorithms to reach the performance necessary for complex agents in enterprises...  ...around the world. The Enterprise ML Research Lab works on the front lines of this AI revolution... 
    Performance
    Full time

    Scale AI, Inc.

    San Francisco, CA
    6 hours ago
  •  ...Francisco, California. The Role: As a Research Engineer - Language Model Pre-Training , you'...  ...runs and model parallelization Performance optimization of our pretraining stack Dataset...  ...enjoy what we do and love discussing AI Benefits and Perks:... 
    Performance
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    a month ago
  • $100k - $300k

     ...Company Overview At Skild AI, we are building the world's first general purpose...  .... Position Overview We are hiring Research Engineers to develop scalable robotic systems aimed...  ...new algorithms for training and optimizing general-purpose robot foundation models... 
    Full time

    Skild AI

    San Francisco, CA
    6 hours ago
  • $160k - $250k

     ...Join to apply for the Founding Research Engineer role at Adam Join to apply...  ...a frontier problem: training AI models to intelligently interpret...  ...accuracy in 3D space Optimize inference for faster CAD edits...  ...teach long-horizon agents to perform design tasks end-to-end We Are... 
    Full time

    Adam

    San Francisco, CA
    6 hours ago
  • $176k - $255k

     ...the industry’s leading AI labs to provide high...  ...accelerate progress in GenAI research. We are looking for...  ...and Research Engineers with expertise in LLM...  ...This role will focus on optimizing data curation and algorithmic...  ...experience, interview performance, and relevant... 
    Performance
    Full time
    Shift work

    Scale AI

    San Francisco, CA
    more than 2 months ago
  • $155k - $269k

     ...Description Waabi, founded by AI visionary Raquel Urtasun,...  ...multidisciplinary team of Research Scientists and Engineers building the content...  ...modeling, and throughput optimization for batch and pipeline workloads...  ...awards and an annual performance bonus. Perks/Benefits:... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    2 days ago
  • $200k - $350k

     ...Roam Roam is an Applied AI lab building World Models for...  ...-time technical founders, engineers that made 100+ games for Voodoo...  ...methods, and inference-time optimization for real-time generation. Your...  ...3D environments. Our current research spans: Distributed multi-... 
    Visa sponsorship
    Relocation package

    ROAM

    San Francisco, CA
    1 day ago
  •  ...Adam Founding Research Team Opportunity We're building the founding...  ...a frontier problem: training AI models to intelligently interpret...  ...accuracy in 3D space Optimize inference for faster CAD edits...  ...teach long-horizon agents to perform design tasks end-to-end We... 

    adam.ai

    San Francisco, CA
    16 hours ago
  • $225k

     ...to safe AGI lies in automating research and code generation to improve...  ...the role As a Research Engineer, you'll work on training, evaluating, and serving large AI models and new inference-time...  ...What you'll work on Optimize inference throughput for novel... 
    Relocation
    Visa sponsorship

    Magic AI Corp.

    San Francisco, CA
    5 days ago
  •  ...’re building high-performance infrastructure to...  ...GPUs by designing kernels, tuning memory layouts, and optimizing model execution at...  ...a kernel-focused engineer to lead efforts in...  ...Collaborate with researchers to port or re-architect...  ...OpenAI is an AI research and deployment... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Engineer - AI Performance & Kernel Optimization. Be the first to apply!