ML Systems Engineer: Trainium Inference & Kernels
Slope
OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software stack to run frontier models efficiently on the platform. This deeply technical, cross‑stack role covers kernels, compilers, and model execution. You will develop high‑performance kernels, improve compiler support, and ensure scalable execution of the model forward pass on Trainium. #J-18808-Ljbffr Slope
- ...Inc. is looking for a Member of Technical Staff focused on ML systems and inference in San Francisco. You will design and build inference... ...allocation. Candidates should have strong foundations in software engineering, experience with ML inference systems, and performance...Suggested
- ...a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and... ...components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly in Python and C++....Suggested
- MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal candidate will have 3+ years of experience in production...Suggested
$200.8k - $251k
...member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary...SuggestedFull time- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...Suggested
- TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and...
- ...seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference...
- Reactor in San Francisco is seeking an ML Inference Engineer to maximize performance of generative media models and push ultra-low-latency, high-throughput inference. You will craft an in-house runtime, implement optimizations with PyTorch tools, and collaborate with partner...
- Oscar is hiring a Senior Machine Learning Inference Engineer for a full-time role in the San Francisco Bay Area. You will focus on improving... ..., with significant ownership over production inference systems. The ideal candidate has 3+ years of professional experience...Full time
- A media technology company in San Francisco is seeking a Founding Engineer specializing in ML Inference. This highly technical role requires expertise in the ML infrastructure stack and aims to optimize generative media performance. The ideal candidate will drive innovations...Relocation package
- Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying... ...JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration...Full timeWork at officeRemote workFlexible hours
$110 per hour
...Dorsey . Position: MLOps Engineer (JAX, PyTorch, Pallas/Triton)... ...training infrastructure, and ML framework-level topics . Design... ...solutions to MLOps and ML systems problems . Evaluate MLOps... ...systems reasoning, and kernel-level optimization across tasks...Remote jobContract workSummer workWeekday work- Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more cost-efficient, and more reliable. You will extend orchestration frameworks (Kubernetes...
- ...Associates Limited (US) in San Francisco offers a hybrid Senior ML Inference Engineer role. You will enhance AI-native infrastructure powering... ...generative and multimodal models, with ownership across inference systems and production performance. The ideal candidate has 3+ years...
$160k - $230k
...About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models models and...Full time$180k - $270k
...deploying high-throughput, ultra-low-latency inference engines for large language models or... ...critical intersection between the core ML training team and the backend infrastructure... ...moving environments and genuinely enjoy the systems-engineering challenge of squeezing every...Full timeWork at officeWorldwide- Together AI is building the best inference infrastructure for voice applications. We seek a Staff ML Engineer to own the model serving stack and optimize latency and throughput for real-time voice workloads. You'll work with state-of-the-art accelerators and collaborate...
- Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model... ...or Python and insights into the LLM inference ecosystem. A commitment to diversity...Remote job
- ...up. We're seeking talented MLOps Engineers with deep, hands-on expertise in PyTorch and kernel-level programming (Triton/Pallas... ...training data for frontier AI systems. This is a W-2 employment position... ..., training infrastructure, and ML framework-level topics. Design challenging...Full timeWeekday work
$170.1k - $258.3k
...-capable fully self-driving systems, to move us toward safer, more... ...mobility. For the AI Kernels & Compilers team, that mission... ...development, and performance engineering so that every cycle on our accelerators... ...the heart of our on‑vehicle ML inference for ADAS and autonomous...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$128.7k - $261.3k
...capable fully self-driving systems, to move us toward... ...mobility. For the AI Kernels & Compilers team, that... ...development, and performance engineering so that every cycle on... ...into fast, reliable inference across GPUs powering... ..., and effortless for ML engineers across the...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$209k - $313k
...Saturn, and other digital services.Snap Engineering teams build fun and technically sophisticated... ...:Strong understanding of causal inference and modern approaches to estimating treatment... ...(A/B tests) and leveraging causal ML in production systemsPreferred Qualifications...Full timeLive inWork at officeLocal area- United States Digital Space LLC in New York seeks an experienced ML Research Engineer to join the Enterprise ML Research Lab. You will build, profile and optimize our training and inference framework and post-train state-of-the-art models for enterprise engagements. You...
$203.5k - $299.3k
...next generation of causal decisioning systems for New Verticals: grocery,... ...are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash grows... ...Deep practical experience with causal inference, econometrics, experimentation, or...Hourly payWork at officeLocal areaRemote workFlexible hours$155k - $180k
...roles (not only product and engineering), so Roboflow employs developers... ...center of all of this is inference — one of our most important... ...encode that judgment into the system itself. Streamline how new... ...the latest computer vision and ML models to our users. Teach...Full timeSecond jobRemote workWork from homeRelocation packageFlexible hoursNight shift$189.6k - $237k
Scale’s ML platform (RLXF) team builds our internal distributed... ...language model training and inference. The platform has been... ...have:Strong excitement about system optimizationExperience with multi... ...distributed ML systemsStrong software engineering skills, proficient in...Full time- ...AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms,... ...scheduling for low-latency, high-throughput inference. You will implement changes in... ...production-grade inference engines, including kernel backends and ATLAS-style systems,...
- Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize...
- ...a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training... ...training loops. You will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies,...
$401k
An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role. You will be responsible for improving efficiency... ...autonomous role with significant ownership across inference systems and model performance in production. This role is hybrid in...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Systems Engineer: Trainium Inference & Kernels. Be the first to apply!
- data scientist machine learning engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- machine learning engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- graduate machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- senior ml engineer San Francisco, CA
- system engineer remote San Francisco, CA



