Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff, Inference & RL Systems

$225k

Magic

About the role As a Software Engineer on the Inference & RL Systems team, you will design and operate the distributed systems that serve our models in production and power large-scale post-training workflows. This role sits at the boundary between model execution and distributed infrastructure. You will work on systems that determine inference latency, throughput, stability, and the reliability of RL and post-training training loops. Magic’s long-context models introduce demanding execution constraints: KV-cache scaling, memory pressure under long sequences, batching trade-offs, long-horizon trajectory rollouts, and sustained throughput under real-world workloads. You will own the infrastructure that makes both production inference and large-scale RL iteration fast and reliable. What you’ll work on Design and scale high-performance inference serving systems Optimize KV-cache management, batching strategies, and scheduling Improve throughput and latency for long-context workloads Build and maintain distributed RL and post-training infrastructure Improve reliability of rollout, evaluation, and reward pipelines Automate fault detection and recovery for serving and RL systems Profile and eliminate performance bottlenecks across GPU, networking, and storage layers Collaborate with Kernels and Research to align execution systems with model architecture What we’re looking for Strong software engineering and distributed systems fundamentals Experience building or operating large-scale inference or training systems Deep understanding of GPU execution constraints and memory trade-offs Experience debugging performance issues in production ML systems Ability to reason about system-level trade-offs between latency, throughput, and cost Track record of owning critical production infrastructure Compensation, benefits, and perks (US) Annual salary range: $225K - $550K Equity is a significant part of total compensation, in addition to salary 401(k) plan with 6% salary matching Generous health, dental and vision insurance for you and your dependents Unlimited paid time off Visa sponsorship and relocation stipend to bring you to SF, if possible A small, fast-paced, highly focused team Magic strives to be the place where high-potential individuals can do their best work. We value quick learning and grit just as much as skill and experience. #J-18808-Ljbffr Magic

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff, Inference & RL Systems in San Francisco, CA vacancy
  •  ...first heterogeneous neocloud for AI workloads. As AI systems scale, the industry is hitting fundamental limits...  ...-class AI datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference. In this role, you will design and build the inference... 
    Suggested

    Gimlet Labs

    San Francisco, CA
    1 day ago
  • $150k - $300k

     ...and pair it with the full RL post‑training stack: environments...  ...cloud LLM serving, LLM inference optimization and RL systems. You will be working on...  ...RL training stack. Core Technical Responsibilities LLM Serving...  ...and encourage team members to contribute to the broader... 
    Suggested
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    3 days ago
  • $200k - $400k

     ...The Role We're looking for an inference runtime engineer to push the...  ...Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang,...  ...model serving. Familiarity with RL frameworks and algorithms for...  ...etc). Written widely-shared technical blogs or side projects on... 
    Suggested
    Remote work
    Visa sponsorship
    Shift work

    Inferact

    San Francisco, CA
    1 day ago
  • $225k

     ...Our approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal. About the role As a Software Engineer on the Pre-training Systems team, you will design and operate the distributed infrastructure... 
    Suggested
    Relocation
    Visa sponsorship

    Magic

    San Francisco, CA
    3 days ago
  • $350k

     ...areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments. Our team...  ..., and MIT. The Role We are looking for an engineer to own the inference systems that power our models in production and research. You'll... 
    Suggested

    Mirendil

    San Francisco, CA
    3 days ago
  •  ...will do in our immensely competitive market. Build the systems that make AI inference fast, reliable, and cost-efficient at global scale. You’ll...  ...CTO, who will ask about your experience, and share as much technical detail about Sail as you want to hear. Come in to Sail's... 
    Work at office
    Immediate start

    Sail

    San Francisco, CA
    3 days ago
  • $150k - $350k

    Mission Gimlet Labs is seeking a Member of Technical Staff focused on distributed systems. In this role, you will build the core platform that schedules, routes, and operates AI workloads reliably at production scale. You will work on systems that coordinate execution... 

    Gimlet Labs, Inc.

    San Francisco, CA
    1 day ago
  •  ...and other industry leaders. The Opportunity Language model inference is the fastest-moving market in AI: dozens of providers,...  ...benchmarks are the industry’s reference point, and we’re hiring a Member of Technical Staff to drive them. You’ll own coverage of the serverless... 
    Shift work

    Artificial Analysis

    San Francisco, CA
    2 days ago
  •  ...building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and...  ...across model scales. Responsibilities Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against... 

    Causal Labs

    San Francisco, CA
    3 days ago
  •  ...building the first heterogeneous neocloud for AI workloads. As AI systems scale, the industry is hitting fundamental limits in power,...  ...to gigawatt-class AI datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on distributed systems. In this role, you will... 

    Gimlet Labs

    San Francisco, CA
    1 day ago
  • $350k

     ...areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large‑scale experiments. Our team includes...  ...automate data collection pipelines for complex, long‑horizon RL tasks. Build robust systems to identify and prevent reward hacking... 

    Mirendil

    San Francisco, CA
    3 days ago
  • Member of Technical Staff - Post‑Training Join to apply for the Member of Technical...  .... About The Role Build systems that transform powerful pre...  ...learning algorithms, and inference‑time scaling techniques. Collaborate...  ...data, reward modeling, or RL techniques. Evidence of... 
    Full time
    Relocation package

    Reflection AI

    San Francisco, CA
    3 days ago
  • $200k

    Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is to build...  ...pre-training, domain-specific RL, ultra-long context, and inference-time compute to achieve this goal...  ...compute, networking, and storage systems. Long-running distributed jobs, high... 
    Relocation
    Visa sponsorship

    Magic AI, Inc

    San Francisco, CA
    1 day ago
  • Eragon is seeking a Member of Technical Staff to own end-to-end AI system development in production, reporting directly to the founder/CEO. The on-site SF office hosts a small, intense team working at the frontier of AI tools for enterprise use. You will build and deploy... 
    Work at office

    davidjoseph-co

    San Francisco, CA
    3 days ago
  •  ...starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at...  ...will help guide our future. Requirements: A research-leaning or systems background in LLM inference, with work you can point to. Fluency... 
    Work at office

    Mixpeek

    San Francisco, CA
    1 day ago
  • $350k

     ...areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments. Our team includes...  ...and infrastructure. You will work to push the scale of our RL stack, whether it is novel recipe ideas, reliability, or performance... 

    Mirendil

    San Francisco, CA
    3 days ago
  •  ...candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making...  ...superoptimization, autotuning, or a genuine appetite to build one of these systems from scratch if you haven\'t yet. A solid grasp of inference... 
    Full time

    Touring Capital

    San Francisco, CA
    1 day ago
  • $350k

     ...areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments. Our team includes...  ...on include: Post-training recipes : Develop and iterate on RL, SFT, and distillation recipes. Understand how choices in objectives... 

    Mirendil

    San Francisco, CA
    3 days ago
  •  ...interactive world models: systems that generate, simulate, and...  ...AI The Role We are hiring a Member of Technical Staff to lead reinforcement...  ...the company for larger-scale RL across both digital and physical...  ...or high-throughput inference systems Familiarity with supervised... 

    Moonlake AI

    San Francisco, CA
    3 days ago
  •  ...present bottleneck is the lack of high-quality RL training environments. Our first step is...  ...frontier of ML research, training, and inference infrastructure. Collaborate with...  ...RLVR problems. Proficiency in Python and systems programming and at least one of PyTorch or... 
    Full time
    Visa sponsorship
    Relocation package

    Preference Model

    San Francisco, CA
    23 hours ago
  • $120k - $300k

     ...inside browsers, APIs, and internal systems just like human analysts — automating...  ...FinTech The Role A backend-leaning Member of Technical Staff building the end-to-end systems that...  ...Tech stack: Python, AWS (distributed inference, caching, queue orchestration, self-... 
    Full time
    H1b
    Visa sponsorship

    David Joseph & Company

    San Francisco, CA
    20 days ago
  • $240k - $280k

     ...poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $2...  ...their founding engineers, working on: AI systems that make healthcare administration...  ...of interviewing at Cabana by 2x Inferred from the description for this job 401... 
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    2 days ago
  • $200k

     ...move beyond SFT and human supervision to RL and Learning from Experience. Our...  ...and F500 enterprises. THE ROLE As the Member of the Technical Staff, you are the “PM of the model,” architecting...  ...rapid deployment of your work in live systems. Model Alignment & Data Strategy :... 
    Flexible hours

    Abundant Energy Inc.

    San Francisco, CA
    1 day ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical...  ...transformation, training/fine-tuning, and inference? You will also: Find opportunities to...  ..., or a related field 5+ years of systems engineering experience in an industry... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    3 days ago
  •  ...human attention, and an agentic operating system can lift that ceiling by an order of...  ...to copy from. About the Role Members of Technical Staff (MTS) are the senior engineers who build...  ...observability. Background in offline RL, contextual bandits, or sequential decision... 

    Beacon Software

    San Francisco, CA
    4 days ago
  • $150k - $300k

     ...plane and pair it with the full RL post-training stack:...  ...Infrastructure, you'll be the technical expert who transforms customer...  ...requirements into production‑ready systems capable of training the world...  ...strategies for LLM training, inference, and HPC workloads Present... 

    Prime Intellect

    San Francisco, CA
    2 days ago
  •  ...parts of inputs that are unimportant, reducing inference costs for scale-ups and enterprises that...  ...people with a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own the cloud systems that serve our compression API end-to-end.... 
    Visa sponsorship

    The Token Company

    San Francisco, CA
    2 days ago
  •  ...This Role We're looking for an engineer with deep Rust expertise and strong algorithmic fundamentals to work on performance‑critical systems. You'll build the low‑level infrastructure that powers real‑time robotic perception, planning, and control. Core Responsibilities... 

    Dimensional Inc.

    San Francisco, CA
    2 days ago
  •  ...Transformer or next AlphaFold breakthrough. As a Member of Technical Staff, you will build this autonomous AI research system to usher in the golden era of discovery....  ...Expertise in One or More Areas Evolutionary Search & RL: Evolutionary algorithms, meta-learning,... 

    Thesis (YC F25)

    San Francisco, CA
    23 hours ago
  •  ...and parallelism strategies, and help us squeeze every FLOP out of our hardware. What you’ll do Modify and extend state-of-the-art inference engines like vLLM and SGLang. Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain... 

    Sail Research

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff, Inference & RL Systems. Be the first to apply!