AI Research Engineer (Kernel & Inference Optimization) [Remote]
jobgether
- Remote job
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in United States.
You will work at the intersection of AI research, systems engineering, and high-performance model inference.
Your focus will be on developing and optimizing model-serving architectures for advanced AI systems across a range of hardware environments.
You will tackle challenges involving latency, throughput, memory efficiency, and scalability, including deployment on resource-constrained mobile and edge devices.
The role combines hands-on research with low-level engineering, giving you the opportunity to develop novel inference strategies and GPU kernels.
You will work with complex architectures spanning text, image, audio, diffusion models, and vision transformers.
Your work will involve rigorous benchmarking, production testing, and iterative optimization to translate research into measurable performance improvements.
You will collaborate with cross-functional teams in a highly technical, remote environment focused on pushing the boundaries of efficient AI systems.
Accountabilities
- Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.
- Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.
- Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.
- Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.
- Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.
- Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.
- Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).
- Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
- Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads.
- Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications.
- Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies.
- Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability.
Requirements:
- Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences.
- Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch.
- Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices.
- Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications.
- Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment.
- Strong experience writing GPU kernels for mobile devices such as smartphones.
- Practical experience developing and deploying end-to-end inference pipelines, from model optimization through production integration on constrained hardware.
- Strong ability to apply empirical research and systematic experimentation to solve latency, computational, and memory challenges.
- Experience designing robust evaluation and benchmarking frameworks for inference systems.
- Knowledge of distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU clusters.
- Deep understanding of the mathematical foundations and architecture of diffusion models and Vision Transformers.
- Familiarity with modern inference optimization techniques including pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding such as EAGLE.
- Strong analytical and problem-solving abilities, with an ability to investigate complex system bottlenecks and turn research findings into practical engineering solutions.
- Excellent English communication skills and the ability to collaborate effectively with distributed, cross-functional technical teams.
Benefits:
- Opportunity to work on advanced AI systems spanning model serving, inference optimization, mobile computing, edge deployment, and large-scale distributed inference.
- Remote-first working environment with an international team.
- Exposure to cutting-edge AI research and practical systems engineering challenges.
- Opportunity to contribute to performance-critical infrastructure where improvements can have a measurable impact on real-world AI applications.
- Collaborative environment combining research-driven experimentation with hands-on engineering.
- Opportunity to work with advanced model architectures including diffusion models, Vision Transformers, and multimodal systems.
- Access to challenging technical problems involving GPU kernels, inference engines, memory optimization, and distributed computing.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
- ...on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Research Engineer (Kernel & Inference Optimization) based in United States. You will work at the intersection of AI research, systems engineering, and high-...SuggestedFull timeRemote work
- ...AfterQuery seeks an ML Research Engineer focused on inference and GPU kernels. The role is fully remote and contract-based, offering competitive hourly compensation and the chance to work on frontier AI research problems. You will define problems, develop reference...SuggestedRemote jobHourly payContract work
- ...comPhone: (***) ***-****Job Title: AI Research Engineer - Deep Learning... ...acceleration, building low-latency inference systems that process... ...ResponsibilitiesDrive performance optimizations across all aspects of large... ..., including custom kernel creation, data streaming pipelines...Suggested
- ...based in San Francisco, California. The Role: As a Research Engineer - AI Performance & Kernel Optimization , you will improve and optimize the performance of our large-scale language model training and inference stacks. You will work closely with our pretraining...SuggestedWork at officeRelocation package
- ...the human layer of AI. Our mission is to... ...through pioneering research in multimodal AI for... ...Research Scientist/Engineer with a focus on model optimization to join our core AI... ...Strong understanding of inference performance and GPU... ...custom Triton/CUDA kernels or low-level...SuggestedFull timeRemote workRelocation packageFlexible hours
$236k - $330k
...this new era, we seek AI-native thinkers... ...systems developers and researchers to join the Snowflake... ...of the art in LLM inference systems and optimization.Our mission is to build... ...systems to GPU kernels and model-system co-... ...We embrace AI-native engineering, using AI not only as...- ...AfterQuery is building a research lab focused on ML inference and GPU kernel engineering, covering inference serving systems, kernel optimization, and deployment infrastructure for large models. This is research-and-evaluation work, not production engineering, defining...Remote jobHourly pay
- ...AfterQuery is assembling a research cohort focused on ML inference and GPU kernel engineering, spanning inference serving systems, kernel optimization, and deployment infrastructure. This is research-and-evaluation work, not production engineering, centered on defining...Remote jobContract workWork at office
$250k - $300k
Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible... ...-scale model inference, including but not limited to GPU kernel development, novel inference devices like ASICs and FPGAs...Work experience placementWork at officeLocal areaImmediate start- ...AfterQuery is a research lab exploring AI boundaries through novel datasets and evaluation, focusing on ML inference and GPU kernel engineering. This remote, contract role invites senior experts to define correct and excellent solutions for hard, real-world problems....Remote jobContract work
- ...AfterQuery is a research lab exploring the boundaries of artificial intelligence... .... We recruit for an ML Research Engineer focusing on inference and GPU kernel engineering, working remotely on research... ...and rubrics, and evaluating AI outputs for correctness and depth....Remote job
- ...intelligence. As a Model Optimization & Deployment Engineer, you will focus on bringing highly... ...ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time,... ...maximize memory bandwidth on AI accelerators. Write production...Temporary workRelocation package
- ...AfterQuery is a research lab exploring the boundaries of artificial intelligence through novel datasets and experimentation... ...performance on hard, real-world problems, focusing on inference and GPU kernel engineering. This is research-and-evaluation work, not production...Remote jobFlexible hours
- ...AfterQuery seeks an ML Research Engineer focusing on inference and GPU kernels for remote, contract work. You will architect challenging evaluation problems, craft reference solutions, and judge AI outputs with rigor and domain insight. The role emphasizes research...Contract workRemote work
$135 per hour
...AfterQuery is a research lab exploring the boundaries of artificial intelligence... ...and experimentation. We believe great AI comes from exceptional, human-... ...remote, contract role focuses on ML inference and GPU kernel engineering. Competitive hourly pay ($135/hr) based...Remote jobHourly payContract workFlexible hours- ...AfterQuery is seeking an ML Research Engineer focusing on inference and GPU kernels. This remote contract role involves designing research scenarios, creating reference solutions, and grading AI outputs. You’ll contribute to frontier AI evaluation rather than production...Remote jobHourly payContract workFlexible hours
- ...AfterQuery is a research lab exploring AI frontiers through novel datasets and experimentation. We are building a cohort of senior ML researchers focused on inference and GPU kernel engineering to define criteria for correctness and excellence on real-world problems....Contract workRemote workFlexible hours
- FMX is seeking a Senior AI Research Engineer to help design, build, and scale Skynapse, FMX's... ...fine-tuning, governed deployment, and inference optimization.Bachelor's degree in computer... ...LangGraph, LangChain, LlamaIndex, Semantic Kernel, AutoGen, CrewAI, or similar...
$110k - $270k
...processing unit (GPNPU) architecture. Quadric's co-optimized software and hardware is targeted to run neural network (NN) inference workloads in a wide variety of edge and... ...C++ DSP and control code.RoleThe AI Kernel Engineer in Quadric plays the key role to enable a...Full timeWork at officeLocal areaImmediate start2 days per week$215k - $260k
...vertically integrated AI infrastructure company... ...That means owning the inference stack end to end: profiling... ...go, bringing modern optimization techniques into real... ...with customer engineering teams to tailor deployments... ...and SGLang to the CUDA kernels underneath, profiling...Temporary work- ...Embedded Software Engineer to develop and optimize high-performance compute... ...RISC-V-based vision AI accelerator. You... ...performance compute kernels, runtime components,... ...deployment and inference workflows Develop... ...industry, graduate research, doctoral research,...Full timeFlexible hours
$207k - $300k
Analyze and optimize AI inference workloads across the application, model, and... ...Computer Science, Computer Engineering, Electrical Engineering,... ...collective communication, kernel), and can reason their implications... ...Learning and Neuroscience research scientists. Your...- ...enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application software... ...in Palo Alto, you will design and optimize large-scale model serving systems end-... ...production workloads, including low-level GPU kernel work, quantization, speculative...Permanent employmentTemporary workRemote workWorldwideWeekend work
- ...AfterQuery is building a research cohort focused on ML inference and GPU kernel engineering, spanning inference serving systems to deployment infrastructure. This is... ...authoring reference solutions and rubrics, and evaluating AI outputs for depth and domain judgment. #J-18808-...Remote job
- ...Driving innovation in model serving and inference architectures, the full-time AI Research Engineer will optimize deployment strategies for advanced AI systems in a fully... ...discipline Proven experience in low-level kernel optimizations and inference optimization on mobile...Full timeRemote work
- ...semiconductors where AI can design and... ...Stanford professors, SAIL researchers, Olympiad medalists... ..., integrate, and optimize state‑of‑the‑art CUDA kernels to power AI models... ...model training, inference, and reinforcement... ...with researchers and engineers, you’ll help make Voltai...
$182k - $242k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers,... ..., rendering, and real-time inference. Our stack is engineered for speed, scale, and cost-efficiency... ...& Performance team, focused on kernel authoring and optimization. You will write, profile, and tune...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$2,500 per month
...heavily focused on inference . Backed by... ...staffed by leading engineers, Etched is redefining... ...push the frontier on kernel engineering.... ...breakthroughs will come from AI systems that can... ...implementations optimized for Etched... ...a uniquely tight research loop: proprietary...Work at officeRelocation package- ...computing, cloud, and AI. Whether you’re designing... ...Software Development Engineer to own the end-to-end model... ...and high-performance inference serving. THE PERSON:... ...hardware, written GPU kernels that moved production metrics... ...Enablement Enable and optimize large-scale model...
$200k - $250k
...consensus. About the Role We're looking for an AI Inference Platform Engineer to build, operate, and optimize the systems that serve large language, vision,... ...to identify bottlenecks from individual GPU kernels through multi-node inference systems. Design...Temporary workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Research Engineer (Kernel & Inference Optimization) [Remote]. Be the first to apply!
- senior ai engineer United States
- ai developer United States
- ai engineer United States
- ai ml engineer United States
- ai engineer remote United States
- machine learning ai engineer United States
- ai prompt engineer United States
- ai research engineer United States
- deep learning research engineer United States
- junior machine learning research engineer United States




