Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer: Trainium Inference & Kernels

Slope

OpenAI in San Francisco seeks an experienced Software Engineer to help bring inference workloads to AWS Trainium and build the software stack to run frontier models efficiently on the platform. This deeply technical, cross‑stack role covers kernels, compilers, and model execution. You will develop high‑performance kernels, improve compiler support, and ensure scalable execution of the model forward pass on Trainium. #J-18808-Ljbffr Slope

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer: Trainium Inference & Kernels in San Francisco, CA vacancy
  •  ...Inc. is looking for a Member of Technical Staff focused on ML systems and inference in San Francisco. You will design and build inference...  ...allocation. Candidates should have strong foundations in software engineering, experience with ML inference systems, and performance... 
    Suggested

    Gimlet Labs, Inc.

    San Francisco, CA
    1 day ago
  •  ...a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and...  ...components. Ideal candidates should have strong software engineering skills and experience with ML inference systems, particularly in Python and C++.... 
    Suggested

    Gimlet Labs

    San Francisco, CA
    4 days ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal candidate will have 3+ years of experience in production... 
    Suggested

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  • $200.8k - $251k

     ...member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary... 
    Suggested
    Full time

    Scale AI

    San Francisco, CA
    5 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 
    Suggested

    Vast.ai Inc.

    San Francisco, CA
    4 days ago
  • TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and... 

    TensorScale AI

    San Francisco, CA
    1 day ago
  •  ...seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference... 

    Reflection AI

    San Francisco, CA
    1 day ago
  • Reactor in San Francisco is seeking an ML Inference Engineer to maximize performance of generative media models and push ultra-low-latency, high-throughput inference. You will craft an in-house runtime, implement optimizations with PyTorch tools, and collaborate with partner... 

    Reactor

    San Francisco, CA
    3 days ago
  • Oscar is hiring a Senior Machine Learning Inference Engineer for a full-time role in the San Francisco Bay Area. You will focus on improving...  ..., with significant ownership over production inference systems. The ideal candidate has 3+ years of professional experience... 
    Full time

    Oscar

    San Francisco, CA
    5 days ago
  • A media technology company in San Francisco is seeking a Founding Engineer specializing in ML Inference. This highly technical role requires expertise in the ML infrastructure stack and aims to optimize generative media performance. The ideal candidate will drive innovations... 
    Relocation package

    Reactor.am

    San Francisco, CA
    1 day ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying...  ...JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    4 days ago
  • $110 per hour

     ...Dorsey . Position: MLOps Engineer (JAX, PyTorch, Pallas/Triton)...  ...training infrastructure, and ML framework-level topics . Design...  ...solutions to MLOps and ML systems problems . Evaluate MLOps...  ...systems reasoning, and kernel-level optimization across tasks... 
    Remote job
    Contract work
    Summer work
    Weekday work

    Mercor

    San Francisco, CA
    10 days ago
  • Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more cost-efficient, and more reliable. You will extend orchestration frameworks (Kubernetes... 

    Inception

    San Francisco, CA
    1 day ago
  •  ...Associates Limited (US) in San Francisco offers a hybrid Senior ML Inference Engineer role. You will enhance AI-native infrastructure powering...  ...generative and multimodal models, with ownership across inference systems and production performance. The ideal candidate has 3+ years... 

    Oscar Technology

    San Francisco, CA
    2 days ago
  • $160k - $230k

     ...About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models models and... 
    Full time

    Together Ai

    San Francisco, CA
    1 day ago
  • $180k - $270k

     ...deploying high-throughput, ultra-low-latency inference engines for large language models or...  ...critical intersection between the core ML training team and the backend infrastructure...  ...moving environments and genuinely enjoy the systems-engineering challenge of squeezing every... 
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    1 day ago
  • Together AI is building the best inference infrastructure for voice applications. We seek a Staff ML Engineer to own the model serving stack and optimize latency and throughput for real-time voice workloads. You'll work with state-of-the-art accelerators and collaborate... 

    Together

    San Francisco, CA
    1 day ago
  • Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model...  ...or Python and insights into the LLM inference ecosystem. A commitment to diversity... 
    Remote job

    Jaide Health

    San Francisco, CA
    4 days ago
  •  ...up. We're seeking talented MLOps Engineers with deep, hands-on expertise in PyTorch and kernel-level programming (Triton/Pallas...  ...training data for frontier AI systems. This is a W-2 employment position...  ..., training infrastructure, and ML framework-level topics. Design challenging... 
    Full time
    Weekday work

    Obsidian

    San Francisco, CA
    2 days ago
  • $170.1k - $258.3k

     ...-capable fully self-driving systems, to move us toward safer, more...  ...mobility. For the AI Kernels & Compilers team, that mission...  ...development, and performance engineering so that every cycle on our accelerators...  ...the heart of our on‑vehicle ML inference for ADAS and autonomous... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    2 hours ago
  • $128.7k - $261.3k

     ...capable fully self-driving systems, to move us toward...  ...mobility. For the AI Kernels & Compilers team, that...  ...development, and performance engineering so that every cycle on...  ...into fast, reliable inference across GPUs powering...  ..., and effortless for ML engineers across the... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    2 days ago
  • $209k - $313k

     ...Saturn, and other digital services.Snap Engineering teams build fun and technically sophisticated...  ...:Strong understanding of causal inference and modern approaches to estimating treatment...  ...(A/B tests) and leveraging causal ML in production systemsPreferred Qualifications... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    San Francisco, CA
    2 days ago
  • United States Digital Space LLC in New York seeks an experienced ML Research Engineer to join the Enterprise ML Research Lab. You will build, profile and optimize our training and inference framework and post-train state-of-the-art models for enterprise engagements. You... 

    United States Digital Space LLC

    San Francisco, CA
    2 days ago
  • $203.5k - $299.3k

     ...next generation of causal decisioning systems for New Verticals: grocery,...  ...are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash grows...  ...Deep practical experience with causal inference, econometrics, experimentation, or... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    3 days ago
  • $155k - $180k

     ...roles (not only product and engineering), so Roboflow employs developers...  ...center of all of this is inference — one of our most important...  ...encode that judgment into the system itself. Streamline how new...  ...the latest computer vision and ML models to our users. Teach... 
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    1 day ago
  • $189.6k - $237k

    Scale’s ML platform (RLXF) team builds our internal distributed...  ...language model training and inference. The platform has been...  ...have:Strong excitement about system optimizationExperience with multi...  ...distributed ML systemsStrong software engineering skills, proficient in... 
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  •  ...AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms,...  ...scheduling for low-latency, high-throughput inference. You will implement changes in...  ...production-grade inference engines, including kernel backends and ATLAS-style systems,... 

    Together

    San Francisco, CA
    5 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 

    Causal Labs

    San Francisco, CA
    4 days ago
  •  ...a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training...  ...training loops. You will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies,... 

    Magic AI, Inc

    San Francisco, CA
    1 day ago
  • $401k

    An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role. You will be responsible for improving efficiency...  ...autonomous role with significant ownership across inference systems and model performance in production. This role is hybrid in... 
    Full time

    Oscar

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer: Trainium Inference & Kernels. Be the first to apply!