Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Research Engineer (Kernel & Inference Optimization) [Remote]

Full-time

jobgether

United States
  • Remote job

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in United States.

You will work at the intersection of AI research, systems engineering, and high-performance model inference.
Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments.
You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices.
The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels.
You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers.
Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements.
You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems.

Accountabilities

  • Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.
  • Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.
  • Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.
  • Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.
  • Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.
  • Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).
  • Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
  • Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads.
  • Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications.
  • Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies.
  • Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability.

Requirements:

  • Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences.
  • Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch.
  • Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices.
  • Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications.
  • Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment.
  • Strong experience writing GPU kernels for mobile devices such as smartphones.
  • Practical experience developing and deploying end-to-end inference pipelines, from model optimization through production integration on constrained hardware.
  • Strong ability to apply empirical research and systematic experimentation to solve latency, computational, and memory challenges.
  • Experience designing robust evaluation and benchmarking frameworks for inference systems.
  • Knowledge of distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU clusters.
  • Deep understanding of the mathematical foundations and architecture of diffusion models and Vision Transformers.
  • Familiarity with modern inference optimization techniques including pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding such as EAGLE.
  • Strong analytical and problem-solving abilities, with an ability to investigate complex system bottlenecks and turn research findings into practical engineering solutions.
  • Excellent English communication skills and the ability to collaborate effectively with distributed, cross-functional technical teams.

Benefits:

  • Opportunity to work on advanced AI systems spanning model serving, inference optimization, mobile computing, edge deployment, and large-scale distributed inference.
  • Remote-first working environment with an international team.
  • Exposure to cutting-edge AI research and practical systems engineering challenges.
  • Opportunity to contribute to performance-critical infrastructure where improvements can have a measurable impact on real-world AI applications.
  • Collaborative environment combining research-driven experimentation with hands-on engineering.
  • Opportunity to work with advanced model architectures including diffusion models, Vision Transformers, and multimodal systems.
  • Access to challenging technical problems involving GPU kernels, inference engines, memory optimization, and distributed computing.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the AI Research Engineer (Kernel & Inference Optimization) [Remote] in United States vacancy
  •  ...on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in United States. You will work at the intersection of AI research, systems engineering, and high-... 
    Suggested
    Full time
    Remote work

    Jobgether

    Remote
    7 days ago
  •  ...AfterQuery seeks an ML Research Engineer focused on inference and GPU kernels. The role is fully remote and contract-based, offering competitive hourly compensation and the chance to work on frontier AI research problems. You will define problems, develop reference... 
    Suggested
    Remote job
    Hourly pay
    Contract work

    Jobleads-US

    Bridgewater, MA
    3 days ago
  •  ...comPhone: (***) ***-****Job Title: AI Research Engineer - Deep Learning...  ...acceleration, building low-latency inference systems that process...  ...ResponsibilitiesDrive performance optimizations across all aspects of large...  ..., including custom kernel creation, data streaming pipelines... 
    Suggested

    Objective Paradigm

    Chicago, IL
    5 days ago
  •  ...based in San Francisco, California. The Role: As a Research Engineer - AI Performance & Kernel Optimization , you will improve and optimize the performance of our large-scale language model training and inference stacks. You will work closely with our pretraining... 
    Suggested
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    6 days ago
  •  ...the human layer of AI. Our mission is to...  ...through pioneering research in multimodal AI for...  ...Research Scientist/Engineer with a focus on model optimization to join our core AI...  ...Strong understanding of inference performance and GPU...  ...custom Triton/CUDA kernels or low-level... 
    Suggested
    Full time
    Remote work
    Relocation package
    Flexible hours

    Tavus

    San Francisco, CA
    7 days ago
  • $236k - $330k

     ...this new era, we seek AI-native thinkers...  ...systems developers and researchers to join the Snowflake...  ...of the art in LLM inference systems and optimization.Our mission is to build...  ...systems to GPU kernels and model-system co-...  ...We embrace AI-native engineering, using AI not only as... 

    Snowflake

    Bellevue, WA
    1 day ago
  •  ...AfterQuery is building a research lab focused on ML inference and GPU kernel engineering, covering inference serving systems, kernel optimization, and deployment infrastructure for large models. This is research-and-evaluation work, not production engineering, defining... 
    Remote job
    Hourly pay

    Jobleads-US

    Laurel, MD
    3 days ago
  •  ...AfterQuery is assembling a research cohort focused on ML inference and GPU kernel engineering, spanning inference serving systems, kernel optimization, and deployment infrastructure. This is research-and-evaluation work, not production engineering, centered on defining... 
    Remote job
    Contract work
    Work at office

    Jobleads-US

    Boise, ID
    3 days ago
  • $250k - $300k

    Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible...  ...-scale model inference, including but not limited to GPU kernel development, novel inference devices like ASICs and FPGAs... 
    Work experience placement
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    3 days ago
  •  ...AfterQuery is a research lab exploring AI boundaries through novel datasets and evaluation, focusing on ML inference and GPU kernel engineering. This remote, contract role invites senior experts to define correct and excellent solutions for hard, real-world problems.... 
    Remote job
    Contract work

    Jobleads-US

    Lansing, MI
    3 days ago
  •  ...AfterQuery is a research lab exploring the boundaries of artificial intelligence...  .... We recruit for an ML Research Engineer focusing on inference and GPU kernel engineering, working remotely on research...  ...and rubrics, and evaluating AI outputs for correctness and depth.... 
    Remote job

    Jobleads-US

    New York, NY
    2 days ago
  •  ...intelligence. As a Model Optimization & Deployment Engineer, you will focus on bringing highly...  ...ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time,...  ...maximize memory bandwidth on AI accelerators. Write production... 
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    a month ago
  •  ...AfterQuery is a research lab exploring the boundaries of artificial intelligence through novel datasets and experimentation...  ...performance on hard, real-world problems, focusing on inference and GPU kernel engineering. This is research-and-evaluation work, not production... 
    Remote job
    Flexible hours

    Jobleads-US

    Ann Arbor, MI
    3 days ago
  •  ...AfterQuery seeks an ML Research Engineer focusing on inference and GPU kernels for remote, contract work. You will architect challenging evaluation problems, craft reference solutions, and judge AI outputs with rigor and domain insight. The role emphasizes research... 
    Contract work
    Remote work

    Jobleads-US

    Tempe, AZ
    3 days ago
  • $135 per hour

     ...AfterQuery is a research lab exploring the boundaries of artificial intelligence...  ...and experimentation. We believe great AI comes from exceptional, human-...  ...remote, contract role focuses on ML inference and GPU kernel engineering. Competitive hourly pay ($135/hr) based... 
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Jobleads-US

    Rockville, MD
    3 days ago
  •  ...AfterQuery is seeking an ML Research Engineer focusing on inference and GPU kernels. This remote contract role involves designing research scenarios, creating reference solutions, and grading AI outputs. You’ll contribute to frontier AI evaluation rather than production... 
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Jobleads-US

    Irving, TX
    3 days ago
  •  ...AfterQuery is a research lab exploring AI frontiers through novel datasets and experimentation. We are building a cohort of senior ML researchers focused on inference and GPU kernel engineering to define criteria for correctness and excellence on real-world problems.... 
    Contract work
    Remote work
    Flexible hours

    Jobleads-US

    Irvine, CA
    3 days ago
  • FMX is seeking a Senior AI Research Engineer to help design, build, and scale Skynapse, FMX's...  ...fine-tuning, governed deployment, and inference optimization.Bachelor's degree in computer...  ...LangGraph, LangChain, LlamaIndex, Semantic Kernel, AutoGen, CrewAI, or similar... 

    Cantor Fitzgerald

    New York, NY
    5 days ago
  • $110k - $270k

     ...processing unit (GPNPU) architecture. Quadric's co-optimized software and hardware is targeted to run neural network (NN) inference workloads in a wide variety of edge and...  ...C++ DSP and control code.RoleThe AI Kernel Engineer in Quadric plays the key role to enable a... 
    Full time
    Work at office
    Local area
    Immediate start
    2 days per week

    Quadric

    Burlingame, CA
    4 days ago
  • $215k - $260k

     ...vertically integrated AI infrastructure company...  ...That means owning the inference stack end to end: profiling...  ...go, bringing modern optimization techniques into real...  ...with customer engineering teams to tailor deployments...  ...and SGLang to the CUDA kernels underneath, profiling... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...Embedded Software Engineer to develop and optimize high-performance compute...  ...RISC-V-based vision AI accelerator. You...  ...performance compute kernels, runtime components,...  ...deployment and inference workflows Develop...  ...industry, graduate research, doctoral research,... 
    Full time
    Flexible hours

    Mentium Technologies Inc.

    Santa Barbara, CA
    4 days ago
  • $207k - $300k

    Analyze and optimize AI inference workloads across the application, model, and...  ...Computer Science, Computer Engineering, Electrical Engineering,...  ...collective communication, kernel), and can reason their implications...  ...Learning and Neuroscience research scientists. Your... 

    Google

    Mountain View, CA
    3 days ago
  •  ...enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application software...  ...in Palo Alto, you will design and optimize large-scale model serving systems end-...  ...production workloads, including low-level GPU kernel work, quantization, speculative... 
    Permanent employment
    Temporary work
    Remote work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    3 days ago
  •  ...AfterQuery is building a research cohort focused on ML inference and GPU kernel engineering, spanning inference serving systems to deployment infrastructure. This is...  ...authoring reference solutions and rubrics, and evaluating AI outputs for depth and domain judgment. #J-18808-... 
    Remote job

    Jobleads-US

    Draper, UT
    2 days ago
  •  ...Driving innovation in model serving and inference architectures, the full-time AI Research Engineer will optimize deployment strategies for advanced AI systems in a fully...  ...discipline Proven experience in low-level kernel optimizations and inference optimization on mobile... 
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  •  ...semiconductors where AI can design and...  ...Stanford professors, SAIL researchers, Olympiad medalists...  ..., integrate, and optimize state‑of‑the‑art CUDA kernels to power AI models...  ...model training, inference, and reinforcement...  ...with researchers and engineers, you’ll help make Voltai... 

    Voltai

    Palo Alto, CA
    5 days ago
  • $182k - $242k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers,...  ..., rendering, and real-time inference. Our stack is engineered for speed, scale, and cost-efficiency...  ...& Performance team, focused on kernel authoring and optimization. You will write, profile, and tune... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Washington DC
    a month ago
  • $2,500 per month

     ...heavily focused on inference . Backed by...  ...staffed by leading engineers, Etched is redefining...  ...push the frontier on kernel engineering....  ...breakthroughs will come from AI systems that can...  ...implementations optimized for Etched...  ...a uniquely tight research loop: proprietary... 
    Work at office
    Relocation package

    Etched

    San Jose, CA
    6 days ago
  •  ...computing, cloud, and AI. Whether you’re designing...  ...Software Development Engineer to own the end-to-end model...  ...and high-performance inference serving. THE PERSON:...  ...hardware, written GPU kernels that moved production metrics...  ...Enablement Enable and optimize large-scale model... 

    AMD

    Markham, IL
    5 days ago
  • $200k - $250k

     ...consensus. About the Role We're looking for an AI Inference Platform Engineer to build, operate, and optimize the systems that serve large language, vision,...  ...to identify bottlenecks from individual GPU kernels through multi-node inference systems. Design... 
    Temporary work
    Flexible hours

    DRW

    Chicago, IL
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Research Engineer (Kernel & Inference Optimization) [Remote]. Be the first to apply!