Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Systems Architect - Inference & RL at Scale

Togetherai

Togetherai is seeking a qualified engineer to join their Turbo team in San Francisco. The role focuses on advancing inference efficiency and operating Reinforcement Learning pipelines for ML systems. A candidate should have a strong background in systems and algorithms, with experience in Python and distributed systems. With a commitment to innovation, Togetherai offers competitive compensation, equity, health benefits, and opportunities for growth in a dynamic work environment. #J-18808-Ljbffr Togetherai

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Systems Architect - Inference & RL at Scale in San Francisco, CA vacancy
  • $180k - $270k

     ...building the world's most trusted AI work companion for...  ...throughput, ultra-low-latency inference engines for large language models...  ...environments and genuinely enjoy the systems-engineering challenge of...  ...naturalness or ASR accuracy. Large-Scale Distributed Systems: Deploying... 
    Suggested
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    18 hours ago
  • $160k - $230k

     ...About the Role Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing...  ...of our AI inference systems. This role involves working with...  ...efficiently and effectively at scale. If you are passionate about AI... 
    Suggested
    Full time

    Together Ai

    San Francisco, CA
    18 hours ago
  •  ...Internals Accellor is an AI-native services firm...  ...operationalize AI at scale and unlock sustained enterprise...  .... Technical Architect — AI Systems & Platform Internals Experience...  ...— AI Systems, Inference & Platform Internals to...  ...distributed training, RL infrastructure,... 
    Suggested

    Accellor

    San Francisco, CA
    5 days ago
  • $155k - $180k

     ...the center of all of this is inference — one of our most important...  ...increasingly authored with the help of AI agents. That's a great problem...  ...as contribution volume scales. Build an agentic-driven...  ...encode that judgment into the system itself. Streamline how new... 
    Suggested
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    18 hours ago
  •  ...seeking an experienced Revenue Operations Lead to architect our revenue systems and CRM as the single source of truth. You’ll own the...  ...to enable data-driven GTM decisions as the company scales its ARR in a fast-moving, AI‑first environment. #J-18808-Ljbffr RevOps Report
    Suggested

    RevOps Report

    San Francisco, CA
    5 days ago
  • A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate... 

    Reflection AI

    San Francisco, CA
    3 days ago
  •  ...large distributed ML training and inference clusters Develop efficient,...  ...pipelines to manage petabyte-scale datasets and model training...  ..., AWS, or Azure) and their ML/AI service offerings Familiarity...  ...on distributed task management systems and scalable model serving & deployment... 

    Kindredventures

    San Francisco, CA
    5 days ago
  •  ...first-of-its-kind voice interface that scales to billions of users. As a ML engineer,...  ...will prototype features, design scalable systems, and work on a fast-moving team of researchers...  ...infrastructure for ultra-low latency inference and personalize speech models with fine-... 

    DevExplore

    San Francisco, CA
    2 days ago
  •  ...collaborate with researchers, product managers, and systems engineers to optimize, deploy, and scale models like GPT-4 and custom architectures....  ...years in ML systems, plus deep knowledge of distributed training and modern inference engines. #J-18808-Ljbffr AI Breaking Wire

    AI Breaking Wire

    San Francisco, CA
    3 days ago
  •  ...the intersection of LLM inference, browser understanding, and low-latency systems, shipping models that need...  ...barriers, or consumer-focused "AI browsers," we run AI...  ...improve model quality at scale Experiment with...  ...understanding Background in RL or online learning from user... 
    Full time
    Sleeping nights

    Composite

    San Francisco, CA
    18 hours ago
  •  ...Sciforium is an AI infrastructure company developing...  ...engineers the team is scaling rapidly to build the...  ...learning , and deployment + inference optimization . You’ll...  ...Post-training & RL Develop post-training...  ...writing robust, performant systems) Experience with... 
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    18 hours ago
  • $220k - $320k

     ...Help us build the systems that train specialized AI models for the fastest-growing companies...  ...to meet you. About Inference.net Inference.net trains...  ..., evaluation, and planet-scale hosting. We are a well-funded...  ...latest techniques in SFT, RL, and model optimization to... 
    Full time
    Work at office

    Inference

    San Francisco, CA
    18 hours ago
  • $203.5k - $299.3k

     ...the next generation of causal decisioning systems for New Verticals: grocery, convenience,...  ...is to build the causal spine for a large-scale consumer marketplace.You're excited about...  ...have…Deep practical experience with causal inference, econometrics, experimentation, or causal... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    9 hours ago
  •  ...opportunities posted on AI Chopping Block Design, build...  ...machine learning systems including data ingestion...  ...state of the art in LLMs, RL, and code generation. Develop...  ...for tuning training and inference end-to-end for high...  ...training pipelines that scale reliably across domains.... 
    Flexible hours

    AI Chopping Block, Inc.

    San Francisco, CA
    3 days ago
  • $227.2k - $284k

     ...mission is to develop reliable AI systems for the world’s most important...  ...Systems (AIS) is part of the Scale Generative AI Platform (SGP),...  ...needed — training/fine-tuning, inference, memory and retrieval, evaluation...  ...traces, online or offline RL — and validate them with rigorous... 
    Full time

    Scale AI

    San Francisco, CA
    18 hours ago
  • $148.5k - $237.6k

     ...trailblazing Dynamics 365 Finance Functional Architect to help us scale. If you're skilled in architecting...  ...You will help shape the future of Finance systems by leveraging modern ERP capabilities, automation, analytics, and AI-driven solutions to improve efficiency, decision... 
    Work experience placement

    Axon

    San Francisco, CA
    18 hours ago
  • $154.2k - $192.8k

    We’re looking for a Customer Support Systems & Analytics Architect to own the data strategy and reporting...  ...expanding BPO model, and sophisticated AI automation—we are moving away from fragmented...  ...reporting infrastructure is built to scale. Your work will be foundational to how... 

    Mercury

    San Francisco, CA
    3 days ago
  • $342k

     ...Hardware organization develops system and infrastructure solutions...  ...tailored to the demands of advanced AI workloads. We work across the...  ...the RoleWe are seeking a 3P Architect to define and drive rack- and...  ...system architecture for large-scale infrastructure or data center... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    18 hours ago
  • A media technology company in San Francisco is seeking a Founding Engineer specializing in ML Inference. This highly technical role requires expertise in the ML infrastructure stack and aims to optimize generative media performance. The ideal candidate will drive innovations... 
    Relocation package

    Reactor.am

    San Francisco, CA
    4 days ago
  • $195k - $365k

     ...building the real-world AI interface for professionals...  ...and training large-scale audio or speech models from...  ...with building AI systems that natively understand...  ...Reinforcement Learning (RL) techniques (like RLHF or...  ...Optimization: End‑to‑end inference and performance optimization... 
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    3 days ago
  • $176.6k - $239k

     ...motivated Senior Solutions Architect to join our North...  ...- Point of Sale (POS) systems processing thousands of...  ...IoT-based tracking, and AI-assisted anomaly detection...  ...enabling AI inference, computer vision, and operational...  ...technology at massive global scale.Retail enterprises... 
    Local area
    Worldwide
    Flexible hours

    AmazonWebServices

    San Francisco, CA
    2 days ago
  • $112k - $168k

     ...their own destiny.Klaviyo is seeking an AI Solutions Architect to accelerate how our Customer Success...  ...and prototypes into production-grade systems, working closely with Klaviyo's...  ...team and leveraging their support to scale larger, cross-functional initiatives.This... 
    Work at office

    Klaviyo

    San Francisco, CA
    1 day ago
  • Modal is building an infrastructure layer for AI and is seeking strong engineers to optimize ML systems for performance at scale. You will contribute to Modal’s container runtime and open-source projects, pushing language and diffusion models toward higher throughput and... 

    Modal

    San Francisco, CA
    3 days ago
  •  ...its Safeguards team to own the infrastructure behind cutting-edge AI research. You will build the tooling, pipelines, and production...  ...run experiments, train detection methods, and deploy results at scale. The role emphasizes reliability and correctness as models evolve... 

    Neura Market

    San Francisco, CA
    1 day ago
  •  ...enterprises to build and orchestrate AI workforces. Our AI workers don...  ...voice, email, and enterprise systems. Born in Y Combinator (S23)...  ...deployment, monitoring, and scaling in production environments....  ..., generative AI, or real-time inference systems. Hands-on experience... 
    Full time
    Worldwide
    Shift work

    HappyRobot

    San Francisco, CA
    18 hours ago
  •  ...tutoring) is hard to access at scale and hasn’t been meaningfully...  ...Speak is building a human-level, AI-powered tutor in your pocket:...  ...best-in-class assessment systems across multiple products (Speak...  ...extraction → model training → inference → feedback generation) Own monitoring... 
    Full time
    Live in
    Immediate start

    Speak

    San Francisco, CA
    18 hours ago
  • $200k - $400k

     ...Behind our name: Like fire, AI holds the potential for both immense...  ...and designing AI systems that people can trust. Our mission...  ...design of models at frontier scale. Training infrastructure –...  ...interpretability, training, and inference. Integrate new machine learning... 
    Full time

    Goodfire

    San Francisco, CA
    18 hours ago
  • $150k - $200k

     ...safety-critical deployment ● Architect clean interfaces and...  ...software stack optimized for inference performance ● Advance SOTA...  ...developing and deploying ML systems from research through production...  ...with imitation learning or RL-based methods ● Background... 
    Full time

    Deft Ai, Inc.

    San Francisco, CA
    18 hours ago
  •  ...Tilde Research is a moonshot AI lab advancing mechanistic interpretability...  ...the ability to rapidly test, scale, and iterate on them—and that...  .... You’ll work on the systems that support training and...  ...you might work on: Optimize inference and training throughput for novel... 
    Full time
    Internship

    Tilde Research

    San Francisco, CA
    18 hours ago
  •  ...Knowtex is building the future of voice AI operating systems for clinicians, transforming how...  ...with our ambient documentation platform scaling to thousands of clinicians across hundreds...  ...models for low-latency, real-time inference at production scale Large Language... 
    Full time

    Knowtex

    San Francisco, CA
    6 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Systems Architect - Inference & RL at Scale. Be the first to apply!