Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Research Engineer - AI Performance & Kernel Optimization

Zyphra

Job Description

Job Description

Zyphra is an artificial intelligence company based in San Francisco, California.

The Role:

As a Research Engineer - AI Performance & Kernel Optimization , you will improve and optimize the performance of our large-scale language model training and inference stacks. You will work closely with our pretraining and inference teams to identify bottlenecks, design and implement highly optimized kernels, and push the limits of throughput, latency, and hardware utilization across a range of accelerator platforms. This role is suited for someone who enjoys deep systems work, cares about performance at every level of the stack, and is excited to translate low-level optimizations into meaningful gains for frontier-scale AI systems.

You’ll Work Across:
  • Kernel development and optimization for large-scale ML workloads, using any level of the stack from PTX/assembly to CUDA, HIP, Triton, or other GPU DSLs

  • Performance tuning for training and inference stacks across GPUs and other accelerators

  • Profiling and eliminating bottlenecks in memory movement, communication, scheduling, and compute utilization

  • Optimizing distributed training and inference systems for large MoE models, including large-scale model parallelism

  • Portability and optimization across non-NVIDIA hardware, with special interest in AMD hardware such as the MI300x and MI355x

  • Collaboration with research and infrastructure teams to turn systems improvements into real-world model training and inference gains

What We're Looking For / Requirements:
  • Strong engineering aptitude for building reliable, high-performance systems

  • Excellent low-level performance intuition and the ability to reason about hardware-software interactions

  • Are excited to rapidly learn new systems, tools, and hardware environments

  • Excellent communication and collaboration skills, with the ability to work effectively across research and engineering teams

  • Enjoy diving deep into the weeds and hunting down the last 10–20% of performance

Qualifications / Additional Skills:
  • Experience writing highly performant GPU kernels at any level of abstraction–PTX, CUDA, HIP, Triton, or other kernel DSLs

  • Experience optimizing ML workloads for large-scale training, ideally in language model pretraining or inference environments

  • Experience with non-NVIDIA accelerator hardware, such as AMD, AWS Trainium, Google TPU, Qualcomm, ARM, Intel, and custom ASICs

  • Strong understanding of distributed training systems and parallelism schemes, including data parallelism, tensor/model parallelism, pipeline parallelism, sharding, and communication/computation overlap

  • Experience with performance engineering in other demanding parallel computing environments such as HPC, quantitative finance, scientific computing, graphics, compilers, or numerical simulation

  • Strong systems intuition around memory hierarchy, bandwidth constraints, kernel fusion, launch overhead, communication overhead, and hardware utilization

  • Experience using profiling and debugging tools to drive performance improvements

  • Familiarity with infrastructure underlying large-scale training and inference, including collective communication libraries, and runtime performance analysis

  • Background in a highly technical field such as physics, mathematics, theoretical computer science, computer science, or electrical engineering

  • Any HPC experience is a strong plus

Why Work at Zyphra:
  • Our research methodology is grounded in methodical, step-by-step approaches to ambitious goals. Both deep research and engineering excellence are equally valued

  • We strongly value new and crazy ideas and are very willing to bet big on new ideas

  • We move as quickly as we can; we aim to minimize the bar to impact as low as possible

  • We all enjoy what we do and love discussing AI

Benefits and Perks:
  • Comprehensive medical, dental, vision, and FSA plans

  • Competitive compensation and 401(k) plan

  • Relocation and immigration support on a case-by-case basis

  • In-office snacks and meals provided

  • Unlimited PTO and company holidays

  • In-person team in San Francisco with a collaborative, high-energy environment

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Research Engineer - AI Performance & Kernel Optimization in San Francisco, CA vacancy
  •  ...the human layer of AI. Our mission is to...  ...through pioneering research in multimodal AI for...  ...Research Scientist/Engineer with a focus on model optimization to join our core AI...  ...understanding of inference performance and GPU/accelerator...  ...custom Triton/CUDA kernels or low-level... 
    Performance
    Full time
    Remote work
    Relocation package
    Flexible hours

    Tavus

    San Francisco, CA
    3 days ago
  •  ...methods cannot reach. AI is reinventing life...  ...software engineering, and Chai is at the...  ...are seeking an AI Research Engineer to help design...  ...train, evaluate, and optimize Chai's core models,...  ...include high-performance computing, custom CUDA kernels and GPU programming... 
    Performance
    Shift work

    Chai Discovery, Inc

    San Francisco, CA
    3 days ago
  • $120k - $200k

     ...We are actively seeking a Research Engineer specializing in Machine Learning and AI to play a pivotal role in pioneering...  ...improvement processes to optimize the performance and accuracy of machine learning...  ...level optimizations, such as CUDA kernel programming, is highly... 
    Performance
    Casual work
    Work at office

    Erth.AI Inc.

    San Francisco, CA
    5 days ago
  •  ...well known in the AI community for seminal research accomplishments at...  ...Discovery is seeking an AI Engineer to play a crucial...  ...team to build, optimize, and maintain the software...  ...and optimize high-performance computing (HPC)...  ...CUDA and Triton kernels, and so forth)... 
    Performance
    Full time

    Menlo Ventures

    San Francisco, CA
    3 days ago
  •  ...interpretable, and steerable AI systems. We want AI...  ...group of committed researchers, engineers, policy experts, and...  ...Own end-to-end performance of our largest RL runs...  ...behavior of large-scale optimization Hands-on...  ...before anyone writes the kernel The annual compensation... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    4 days ago
  • $159k - $296k

     ...Description Waabi, founded by AI visionary Raquel...  ...driving trucks. As a research engineer for Learnable Planner...  ...learning, optimization-based approaches, search...  ...like TensorRT, CUDA kernels. The US yearly salary...  ...awards and an annual performance bonus. Perks/Benefits... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    11 days ago
  •  ...conversation for every patient. An AI agent that knows who you are...  ...About the Role As a Research Engineer, you’ll be responsible for...  ...models and pipelines that optimize for quality, efficiency, and...  ...for different tasks based on performance, latency, and cost... 
    Performance
    Work at office
    Flexible hours

    Assort Health

    San Francisco, CA
    22 hours ago
  •  ...Get AI-powered advice on this job and more exclusive...  ...empowers software engineers by automating...  ...end‑to‑end, balancing research and engineering to create...  ...ready AI models Build and optimize data pipelines to process...  ...AI models, improve performance, and reduce compute costs... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Resolve AI

    San Francisco, CA
    1 day ago
  •  ...market — now building a Physical AI team to bring these models...  ...and economy use cases. As a Research Engineer on our Physical AI team, you...  ...sequences Evaluate model performance using both benchmark...  ...fundamentals in machine learning, optimization, and large-scale data... 
    Performance
    Work at office

    HEDRA INC

    San Francisco, CA
    1 day ago
  •  ...Francisco, California. The Role: As a Research Engineer - Audio & Speech Models , you will...  ...Large-scale audio training runs Performance optimization of our training stack Audio...  ...enjoy what we do and love discussing AI Benefits and Perks: Comprehensive... 
    Performance
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    a month ago
  •  ...Francisco, California. The Role: As a Research Engineer - Brain Computer Interface Models ,...  ...for EEG and other BCI modalities Performance optimization of the training stack Integration...  ...enjoy what we do and love discussing AI Benefits and Perks:... 
    Performance
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    a month ago
  •  ...Description We are Genmo, a research lab developing the world’s...  ...an exceptional Software Engineer to join our research team in...  ...frontiers of visual generative AI. As a Research Engineer,...  ...GANs, Transformers) Help optimize model performance and scaling capabilities... 
    Performance
    Work at office

    Genmo

    San Francisco, CA
    17 days ago
  • $155k - $269k

     ...Description Waabi, founded by AI visionary Raquel Urtasun, is...  ...efficient simulation. As a Research Engineer in the World Models team,...  ...and inference pipelines. - Optimize model training and inference...  ...incentive awards and an annual performance bonus.   Perks/Benefits:... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    a month ago
  • $155k - $269k

     ...Waabi, founded by AI visionary Raquel Urtasun, is the leader...  ...learn more visit: As a Research Engineer in Sensor Signal Processing,...  ...fusion, and filtering. - Optimize signal processing algorithms...  ...accelerators. - Solid knowledge in performance profiling and optimization.... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    1 day ago
  • $50 per hour

     ...‑agent solutions that optimize Corporate Procurement...  ...roadmaps, and mentor junior engineers to elevate the overall...  ...best practices for performance optimization, cost...  ...robust, collaborative AI engineering culture. Documentation...  ...to cutting‑edge research, mentorship from... 
    Performance
    Full time
    Contract work
    Temporary work
    Work experience placement
    Casual work
    Flexible hours

    Lockheed Martin

    San Francisco, CA
    3 days ago
  • $180k - $250k

     ...Open role Research Engineer San Francisco (On-site), Full-time About...  ...is creating the AI-human translation layer that...  ...faster. You will orchestrate and optimize training runs on long-horizon...  ...and profiling to find which of kernel, dataloader, communication or... 
    Full time
    Work at office
    Relocation package

    Breakout Ventures

    San Francisco, CA
    1 day ago
  • $155k - $269k

     ...Description Waabi, founded by AI visionary Raquel Urtasun,...  ...multidisciplinary team of Research Scientists and Engineers building the content...  ...modeling, and throughput optimization for batch and pipeline workloads...  ...awards and an annual performance bonus. Perks/Benefits:... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    5 days ago
  •  ...methods cannot reach. AI is reinventing life...  ...reinvented software engineering, and Chai is at the...  ...Make our models performant, resource efficient...  ...partnerships with fellow researchers and engineers....  ...across model, layer, and kernel levels; optimize workloads through... 
    Shift work

    Chai Discovery Inc.

    San Francisco, CA
    22 hours ago
  • $180k - $240k

    Our client, a venture-backed AI Startup, is hiring a talented ML/AI Research Engineer to join their team in San Francisco...  ..., training, evaluation and optimization of agent-native AI systems, working...  ...systems with efficient, high-performance retrieval strategies. ~Implement... 
    Performance
    Full time

    Alldus International Consulting Ltd

    San Francisco, CA
    more than 2 months ago
  • $180k - $280k

     ...SuperAnnotate helps the world’s leading AI teams build responsible, next...  ...Impact You’ll Make Our research team is expanding to keep...  ...of the field. As a Research Engineer, you’ll take a research...  ...that move the needle on model performance, with your work feeding directly... 
    Performance
    Full time

    SuperAnnotate AI

    San Francisco, CA
    1 day ago
  •  ...By applying to this role, you will be considered for Research Engineer roles across all teams at OpenAI. About the Role As...  ...Research Engineer here, you will be responsible for building AI systems that can perform previously impossible tasks or achieve unprecedented... 
    Performance

    openai

    San Francisco, CA
    3 days ago
  •  ...Ando Research Team Member Ando is a messaging platform where AI agents take on work alongside their human teammates. We're rebuilding Slack from the ground...  ...basic memory, retrieval, and context impact agent performance, vs where do we genuinely need RL and continual learning... 
    Performance
    Work from home

    ASARI S.A de CV

    San Francisco, CA
    1 day ago
  • $250k - $290k

     ...the United States to help them hire. Research Engineer Location San Francisco, CA /...  ...Industry Artificial Intelligence, AI Infrastructure, Developer Tools, Machine...  ...capable of measuring model and product performance at scale Generate and curate high-... 
    Performance
    H1b
    Remote work

    Recruiting from Scratch

    San Francisco, CA
    4 days ago
  • $134k - $235k

     ...Job Description Waabi, founded by AI visionary Raquel Urtasun, is the leader...  ...positive way. To learn more visit: As a Research Engineer in Neural Rendering, you will create the...  ...awards and discretionary annual performance bonus. Perks/Benefits: - Competitive... 
    Performance
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    a month ago
  •  ...Job Description Job Description Research Engineer – Code Generation & Model Evaluation...  ...generation, software engineering, and AI model evaluation. You'll apply...  ...generation workflows. Refactor and optimize code for performance, scalability, and maintainability.... 
    Performance
    Remote job
    For contractors

    YO AI Labs

    San Francisco, CA
    16 days ago
  •  ...that — with models that need to perform reliably on messy, real-world...  ...'re looking for strong CV/ML engineers with good product judgment who...  ...segmentation, detection, document AI, or vision-language)...  ...Practical mindset — you read the research, but your focus is on shipping... 
    Performance
    Full time
    For contractors

    Bobyard, Inc.

    San Francisco, CA
    3 days ago
  •  ...Francisco, California. The Role: As a Research Engineer - Language Model Pre-Training , you'...  ...runs and model parallelization Performance optimization of our pretraining stack Dataset...  ...enjoy what we do and love discussing AI Benefits and Perks:... 
    Performance
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    3 days ago
  •  ...reliable, interpretable, and steerable AI systems. We want AI to be safe and...  ...a quickly growing group of committed researchers, engineers, policy experts, and business leaders...  ...safely Work with researchers and performance engineers to make sure systems changes... 
    Performance
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    4 days ago
  •  ...Beta, Felicis, Figma Ventures, AI Grant, and more, and we're...  ...accomplished a lot as just one engineer and one designer!) to a clan...  ...breakthrough AI service with leading performance on relevant benchmarks,...  ...". Specifically, in an AI research engineer role, we are looking... 
    Performance

    Poly

    San Francisco, CA
    1 day ago
  • $200k - $400k

     ...Do you have experience in AI research across LLMs, agents, evaluations, benchmarking, AI safety, alignment, agent understanding, RL environments...  ...and environments that improve model reasoning and agent performance. Fresh PhD graduates can be considered. Salary on offer is $2... 
    Performance
    Internship

    TTN Talent

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Research Engineer - AI Performance & Kernel Optimization. Be the first to apply!