Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

RL Systems Engineer - Scale, Reliability & Observability

Jobleads-US

United States Digital Space LLC in San Francisco is seeking an experienced software engineer to design, build, and operate RL-scale distributed systems. You will work across training, sampling, and environment execution on a large fleet, with a focus on reliability and performance.

You will collaborate with researchers and performance engineers to preserve training correctness, reduce tail latency, and implement observability, fault tolerance, and automation across the stack.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the RL Systems Engineer - Scale, Reliability & Observability in San Francisco, CA vacancy
  •  ...United States Digital Space LLC is seeking a senior RL researcher to scale reinforcement learning systems from small results to frontier-scale runs. You will...  ...throughput and cost implications, and collaborating across research and engineering #J-18808-Ljbffr Jobleads-US
    Suggested

    Jobleads-US

    San Francisco, CA
    2 days ago
  •  ...mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe...  ...committed researchers, engineers, policy experts, and business...  ...horizons. At frontier scale, an RL run is an unusually...  ...needs shift Build observability that makes it possible... 
    Suggested
    Visa sponsorship
    Shift work

    Jobleads-US

    San Francisco, CA
    2 days ago
  •  ...Cloudflare is seeking a Systems Engineer to scale PostgreSQL infrastructure across multiple regions. You will deploy...  ...infrastructure and application teams. Emphasis on reliability and performance with strong focus on observability. The role offers a comprehensive benefits... 
    Suggested

    Jobleads-US

    San Francisco, CA
    2 days ago
  •  ...Site Reliability Engineer - AI Infrastructure Location: Global Remote...  ...startups access to the kind of scaled AI infrastructure once...  ...been quietly building the systems, network, and orchestration...  ...monitoring, alerting, and observability for critical systems. Collaborate... 
    Suggested
    Full time
    Remote work

    Andromeda Cluster

    San Francisco, CA
    1 day ago
  • $152.5k - $205k

     ...to power trusted, internet-scale financial innovation. Learn...  ...responsible for:As a Senior Site Reliability Engineer on Circle’s platform team,...  ...solving hard distributed-systems problems, taking ownership...  ....Define and evolve observability practices across metrics, logs... 
    Suggested
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  •  ...native financial operating system for a real-time,...  ...full.About the teamThe Engineering team at Airwallex is a...  ...together to build scalable, reliable, and secure products...  ...incident response, observability, and automation across...  ...strategy for large-scale, cross-functional projects... 
    Temporary work
    Local area

    Airwallex

    San Francisco, CA
    3 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco,...  ...deploying, and maintaining large-scale distributed systems that power AI workloads....  ..., performance, and reliability across environments. What...  ...systems for orchestration, observability, distributed storage, and... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    4 days ago
  • $148.5k - $223.9k

     ...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San...  ...that blend automation, observability, and AI-powered...  ...but proactively design systems that prevent them, applying...  ...improve reliability at scale. By leveraging cutting-... 
    Worldwide
    Weekend work

    Salesforce.Com Inc

    San Francisco, CA
    1 day ago
  • $210.38k - $243.21k

    Manager, Site Reliability Engineer (Hybrid in South San Francisco)About the...  ...across cloud infrastructure, observability, automation, networking,...  ...improve deployment processes, system resilience, and...  ...Platform provides precision at a scale that is otherwise unavailable... 

    Twist Bioscience

    San Francisco, CA
    1 day ago
  •  ...in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that...  ...scalable and highly reliable systems. Our SREs ensure our production...  ...systems are resilient (observable, fault-tolerant, recoverable...  ...and services at scale ~ History of end-to-end project... 
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    1 day ago
  •  ...make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build...  ...deep experience operating production systems at scale, an automation-first mindset, and...  ...production readiness, incident management, observability, resilience testing, and toil... 

    Autodesk

    San Francisco, CA
    4 days ago
  • $153k - $191.3k

     ...processing, and software engineering, our office is a truly...  ...and support a robust system for reproducible...  ...environments, to guarantee the reliability, scalability, and...  ...gapped deployments at scale Clarify and surface...  ...environment ~ Ability to observe and troubleshoot... 
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Remote work
    Home office
    3 days per week

    Planet Labs PBC

    San Francisco, CA
    22 hours ago
  •  ...anywhere in their environments. Our systems can automatically detect and...  ...Role We're hiring a Site Reliability Engineer to own the operational health...  ...recurring, and owning the observability that keeps us ahead of problems as we scale. You set your own priorities... 
    Remote work

    Specter Services LLC

    San Francisco, CA
    1 day ago
  •  ...daily users while enabling our engineering teams to ship fast. You'll...  ...and tooling that improves reliability and partnering with engineering to design systems that are observable, resilient, and easy to operate...  ..., and help shape how Gamma scales to serve its next 100... 
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    10 hours ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies...  ...Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at...  ..., and real-time analytics systems. This is a hands‑on, high‑... 
    Full time

    Alembic Technologies

    San Francisco, CA
    4 days ago
  •  ...own controls, with the reliability and operational...  ...expect from any critical system. Retool's Core Infrastructure...  ...operate at enterprise scale. The strongest...  ...migration steps. Improve observability for Retool Cloud, self...  ...Partner with product engineers on infrastructure... 

    re-tool®

    San Francisco, CA
    2 days ago
  • $150k

     ...seeking an experienced Site Reliability Engineer (SRE) with a strong focus...  ...repositories to ensure all systems remain compliant, secure, and...  ...and dashboards using observability tooling (e.g., CloudWatch,...  ...vulnerability remediation at scale, including OS‑level patching... 

    VantageScore®

    San Francisco, CA
    3 days ago
  •  ...Open role Site Reliability Engineer (SRE) San Francisco, CA (On-site...  ...to ensure the platform can scale ahead of demand without performance...  ...one component fails, the system remains operational)....  ...of systems monitoring and observability — experience with tools like... 

    Methodic

    San Francisco, CA
    2 days ago
  •  ...RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You...  ...pipelines, ML workloads, and real-time analytics systems.This is a hands-on, high-impact role with... 

    Alembic Limited

    San Francisco, CA
    2 days ago
  •  ...autonomously across entire enterprise systems. Born in Y Combinator (S23) and...  ...Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You’ll own the stability, observability, and debugging workflows that... 
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    2 days ago
  •  ...senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and observability of our platform, and help...  ...You'll work on petabyte-scale data processing, SaaS...  ...and billing, and the query systems that power our APIs.... 
    Full time
    Local area
    Remote work

    Databento

    San Francisco, CA
    1 day ago
  •  ...All roles San Francisco, CA Site Reliability Engineer San Francisco, CAFull-timeMid to...  ...pipelines, hard isolation for untrusted code, observability, and the cost controls that keep large...  ...operated production infrastructure at scale and treats security, reliability, and... 
    Full time

    Zof AI, Inc.

    San Francisco, CA
    2 days ago
  •  ...We are looking for a Site Reliability Engineer with experience in incident...  ...focus on the intersection of systems engineering and data...  ...production environments at scale. - Data Proficiency: Strong...  ...Java, Python, or C++. - Observability Expertise: Deep understanding... 

    BayOne Solutions

    San Francisco, CA
    4 days ago
  • $164k - $205k

     ...troubleshoot, and maintain production systems Build and operate cloud...  ...environment Manage and scale Kubernetes clusters that...  ...Design intelligent alerting and observability systems Collaborate with engineering teams to embed reliability into the development... 
    Work experience placement
    Summer holiday
    Live out
    Work at office
    Local area
    Flexible hours
    Shift work
    2 days per week

    SupportFinity

    San Francisco, CA
    1 day ago
  •  ...We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for...  ...resilient cloud-native systems that power critical business...  ...will drive initiatives across observability, incident management,...  ...combines deep expertise in large-scale distributed systems with a... 

    Engg

    San Francisco, CA
    3 days ago
  • $181k - $263k

    ## Senior Staff Site Reliability EngineerApplylocations:...  ...Staff Site Reliability Engineer who will set the...  ...Drive engineering-wide system design, automation, and...  ...as Code (Terraform) at scale across multi-environment...  ...across teams* Expertise in observability engineering—SLOs, SLI... 
    Work from home
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    1 day ago
  • $194k - $267k

    Okta# Staff Site Reliability Engineer - KubernetesBellevue, Washington; Chicago...  ...passion for solving large-scale automation, testing, and...  ...communication, security, and observability within the Kubernetes...  ...troubleshoot, and resolve system issues related to performance... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    2 days ago
  • $200k - $260k

     ...standard data model and system-level visualization...  ...technical leader driving reliability, automation, and...  ...which include IaC, CI/CD, observability, incident response and...  ...teams, mentor senior engineers, and be a primary escalation...  ...run them reliably at scale Architect and... 
    Casual work
    Work at office
    Remote work
    Flexible hours

    Sight Machine

    San Francisco, CA
    4 days ago
  •  ...Role We’re hiring a Senior Database Reliability Engineer to own the reliability, performance,...  ...Make the Django ORM a strength at scale: catch N+1 patterns in review, extend...  ...before they ship Build and improve observability across pganalyze, CloudWatch, and Honeycomb... 
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours
    3 days per week

    scribehow.com

    San Francisco, CA
    1 day ago
  •  ...Senior Database Reliability EngineerSan Francisco, CA...  ...Database Operations Engineering team provides a seamless...  ..., identifying SLAs, system vulnerabilities, and...  ...Operations: Manage large-scale data infrastructures,...  ...in monitoring and observability tools (e.g., Datadog,... 
    Flexible hours

    Crunchyroll

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to RL Systems Engineer - Scale, Reliability & Observability. Be the first to apply!