RL Systems Engineer - Scale, Reliability & Observability
Jobleads-US
United States Digital Space LLC in San Francisco is seeking an experienced software engineer to design, build, and operate RL-scale distributed systems. You will work across training, sampling, and environment execution on a large fleet, with a focus on reliability and performance.
You will collaborate with researchers and performance engineers to preserve training correctness, reduce tail latency, and implement observability, fault tolerance, and automation across the stack.
#J-18808-Ljbffr Jobleads-US- ...United States Digital Space LLC is seeking a senior RL researcher to scale reinforcement learning systems from small results to frontier-scale runs. You will... ...throughput and cost implications, and collaborating across research and engineering #J-18808-Ljbffr Jobleads-USSuggested
- ...mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe... ...committed researchers, engineers, policy experts, and business... ...horizons. At frontier scale, an RL run is an unusually... ...needs shift Build observability that makes it possible...SuggestedVisa sponsorshipShift work
- ...Cloudflare is seeking a Systems Engineer to scale PostgreSQL infrastructure across multiple regions. You will deploy... ...infrastructure and application teams. Emphasis on reliability and performance with strong focus on observability. The role offers a comprehensive benefits...Suggested
- ...Site Reliability Engineer - AI Infrastructure Location: Global Remote... ...startups access to the kind of scaled AI infrastructure once... ...been quietly building the systems, network, and orchestration... ...monitoring, alerting, and observability for critical systems. Collaborate...SuggestedFull timeRemote work
$152.5k - $205k
...to power trusted, internet-scale financial innovation. Learn... ...responsible for:As a Senior Site Reliability Engineer on Circle’s platform team,... ...solving hard distributed-systems problems, taking ownership... ....Define and evolve observability practices across metrics, logs...SuggestedFlexible hours- ...native financial operating system for a real-time,... ...full.About the teamThe Engineering team at Airwallex is a... ...together to build scalable, reliable, and secure products... ...incident response, observability, and automation across... ...strategy for large-scale, cross-functional projects...Temporary workLocal area
$175k - $250k
...Senior Cloud Infrastructure Engineer Location: San Francisco,... ...deploying, and maintaining large-scale distributed systems that power AI workloads.... ..., performance, and reliability across environments. What... ...systems for orchestration, observability, distributed storage, and...Full timeRemote workRelocationRelocation package$148.5k - $223.9k
...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San... ...that blend automation, observability, and AI-powered... ...but proactively design systems that prevent them, applying... ...improve reliability at scale. By leveraging cutting-...WorldwideWeekend work$210.38k - $243.21k
Manager, Site Reliability Engineer (Hybrid in South San Francisco)About the... ...across cloud infrastructure, observability, automation, networking,... ...improve deployment processes, system resilience, and... ...Platform provides precision at a scale that is otherwise unavailable...- ...in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that... ...scalable and highly reliable systems. Our SREs ensure our production... ...systems are resilient (observable, fault-tolerant, recoverable... ...and services at scale ~ History of end-to-end project...Immediate startRemote workWorldwide
- ...make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build... ...deep experience operating production systems at scale, an automation-first mindset, and... ...production readiness, incident management, observability, resilience testing, and toil...
$153k - $191.3k
...processing, and software engineering, our office is a truly... ...and support a robust system for reproducible... ...environments, to guarantee the reliability, scalability, and... ...gapped deployments at scale Clarify and surface... ...environment ~ Ability to observe and troubleshoot...Full timeTemporary workFor contractorsWork at officeLocal areaRemote workHome office3 days per week- ...anywhere in their environments. Our systems can automatically detect and... ...Role We're hiring a Site Reliability Engineer to own the operational health... ...recurring, and owning the observability that keeps us ahead of problems as we scale. You set your own priorities...Remote work
- ...daily users while enabling our engineering teams to ship fast. You'll... ...and tooling that improves reliability and partnering with engineering to design systems that are observable, resilient, and easy to operate... ..., and help shape how Gamma scales to serve its next 100...Work at officeWork from home
$210k - $240k
...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies... ...Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at... ..., and real-time analytics systems. This is a hands‑on, high‑...Full time- ...own controls, with the reliability and operational... ...expect from any critical system. Retool's Core Infrastructure... ...operate at enterprise scale. The strongest... ...migration steps. Improve observability for Retool Cloud, self... ...Partner with product engineers on infrastructure...
$150k
...seeking an experienced Site Reliability Engineer (SRE) with a strong focus... ...repositories to ensure all systems remain compliant, secure, and... ...and dashboards using observability tooling (e.g., CloudWatch,... ...vulnerability remediation at scale, including OS‑level patching...- ...Open role Site Reliability Engineer (SRE) San Francisco, CA (On-site... ...to ensure the platform can scale ahead of demand without performance... ...one component fails, the system remains operational).... ...of systems monitoring and observability — experience with tools like...
- ...RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You... ...pipelines, ML workloads, and real-time analytics systems.This is a hands-on, high-impact role with...
- ...autonomously across entire enterprise systems. Born in Y Combinator (S23) and... ...Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You’ll own the stability, observability, and debugging workflows that...WorldwideShift work
- ...senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and observability of our platform, and help... ...You'll work on petabyte-scale data processing, SaaS... ...and billing, and the query systems that power our APIs....Full timeLocal areaRemote work
- ...All roles San Francisco, CA Site Reliability Engineer San Francisco, CAFull-timeMid to... ...pipelines, hard isolation for untrusted code, observability, and the cost controls that keep large... ...operated production infrastructure at scale and treats security, reliability, and...Full time
- ...We are looking for a Site Reliability Engineer with experience in incident... ...focus on the intersection of systems engineering and data... ...production environments at scale. - Data Proficiency: Strong... ...Java, Python, or C++. - Observability Expertise: Deep understanding...
$164k - $205k
...troubleshoot, and maintain production systems Build and operate cloud... ...environment Manage and scale Kubernetes clusters that... ...Design intelligent alerting and observability systems Collaborate with engineering teams to embed reliability into the development...Work experience placementSummer holidayLive outWork at officeLocal areaFlexible hoursShift work2 days per week- ...We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for... ...resilient cloud-native systems that power critical business... ...will drive initiatives across observability, incident management,... ...combines deep expertise in large-scale distributed systems with a...
$181k - $263k
## Senior Staff Site Reliability EngineerApplylocations:... ...Staff Site Reliability Engineer who will set the... ...Drive engineering-wide system design, automation, and... ...as Code (Terraform) at scale across multi-environment... ...across teams* Expertise in observability engineering—SLOs, SLI...Work from homeFlexible hoursNight shift$194k - $267k
Okta# Staff Site Reliability Engineer - KubernetesBellevue, Washington; Chicago... ...passion for solving large-scale automation, testing, and... ...communication, security, and observability within the Kubernetes... ...troubleshoot, and resolve system issues related to performance...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$200k - $260k
...standard data model and system-level visualization... ...technical leader driving reliability, automation, and... ...which include IaC, CI/CD, observability, incident response and... ...teams, mentor senior engineers, and be a primary escalation... ...run them reliably at scale Architect and...Casual workWork at officeRemote workFlexible hours- ...Role We’re hiring a Senior Database Reliability Engineer to own the reliability, performance,... ...Make the Django ORM a strength at scale: catch N+1 patterns in review, extend... ...before they ship Build and improve observability across pganalyze, CloudWatch, and Honeycomb...Full timeWork at officeRemote workHome officeFlexible hours3 days per week
- ...Senior Database Reliability EngineerSan Francisco, CA... ...Database Operations Engineering team provides a seamless... ..., identifying SLAs, system vulnerabilities, and... ...Operations: Manage large-scale data infrastructures,... ...in monitoring and observability tools (e.g., Datadog,...Flexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to RL Systems Engineer - Scale, Reliability & Observability. Be the first to apply!
- distributed systems engineer San Francisco, CA
- unix linux systems engineer San Francisco, CA
- digital communications systems engineer San Francisco, CA
- space systems engineer San Francisco, CA
- sr systems engineer San Francisco, CA
- system engineer contract San Francisco, CA
- senior linux systems engineer San Francisco, CA
- mission system engineer San Francisco, CA
- operations support system engineer San Francisco, CA
- computer systems engineer San Francisco, CA


