Customer Reliability Engineer
Fluidstack
The Data Center Operations Team
Examples of key problems the team is working on:
- Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 100 GW.
- Fly the plane while it’s being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
- Write the playbook, don’t inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.
Role Scope
- Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
- Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
- Run customer‑facing incident communication with technical depth and no spin.
- Turn recurring customer pain into engineering fixes with the production teams.
What We’re Looking For
- You’ve supported large‑scale compute customers (HPC, cloud, or AI labs) at a technical level.
- You debug distributed systems methodically across layers you don’t own.
- You have written incident updates customers trusted more after reading.
- You push internal teams to fix causes, not symptoms, and follow up until they do.
Bonus: GPU training workloads. InfiniBand or RoCE. Slurm or Kubernetes. NCCL debugging.
We are committed to pay equity and transparency.
Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
#J-18808-Ljbffr- ...Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco · Full-Time About Andromeda Andromeda Cluster... ..., configure, and operate Kubernetes-based clusters for customers across multiple providers. Build automation and tooling...CustomerFull timeRemote work
$150k - $180k
...cutting-edge autonomous technologies, we are seeking a Senior Reliability Engineer (REL) to lead efforts in ensuring the long-term performance,... ...to uncover failure modes, validate off-the-shelf and custom hardware, and model expected field performance. Translate...CustomerFull timeImmediate startWorldwideFlexible hoursNight shift- ...gets done — 94% of the Fortune 500 use it, and 45% are paying customers. We hit $100M ARR in May 2026 and have grown to over 5... ...customers. About the Role We’re hiring a Senior Database Reliability Engineer to own the reliability, performance, and scalability of...CustomerFull timeWork at officeRemote workHome officeFlexible hours3 days per week
$152.5k - $205k
...stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and... ...performance, security, and cost-effectiveness of the systems our customers depend on.What you'll work on: Design, build, and operate Kubernetes...CustomerFlexible hours$150k - $190k
...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to... ...100M+ monthly open source downloads, 6,000+ active LangSmith customers, and 5 of the Fortune 10 use LangSmith in production (+ 35% of...CustomerPermanent employmentWork at officeFlexible hours$210.38k - $243.21k
Manager, Site Reliability Engineer (Hybrid in South San Francisco)About the RoleWe are seeking an experienced and hands-on Site Reliability... ...precision at a scale that is otherwise unavailable to our customers.Twist Bioscience Corporation is an Equal Opportunity Employer...Customer- ...Zof AI is seeking a Site Reliability Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply. This role owns the execution layer of our control plane: Kubernetes and container orchestration, CI/CD pipelines, hard...CustomerFull time
- ...Open role Site Reliability Engineer (SRE) San Francisco, CA (On-site) Responsibilities Develop and maintain advanced monitoring... ...that detect and address issues before they impact customers. Perform regular capacity planning and load testing to ensure...Customer
- ...YOU: Good software has to run where customers need it. For many of Retool's largest customers... ..., behind their own controls, with the reliability and operational clarity they would... ...TAMs to trust. Partner with product engineers on infrastructure requirements for new...Customer
- ...looking for a midlevel or senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and... ...on petabyte-scale data processing, SaaS problems like customer management and billing, and the query systems that power...CustomerFull timeLocal areaRemote work
$98.58k - $138.02k
...office locations: Austin, TX; Irvine, CA; or Akron, OH. Site Reliability Engineer II will be responsible for supporting, enhancing, and... ...Familiarity with restaurant industry SaaS platforms and customer‑facing applications. R365 Team Member Benefits & Compensation...CustomerWork at office- ...be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future... ...move fast and learn faster; obsess about creating customer value; value impact over activity; and embrace healthy...CustomerWorldwideHome officeFlexible hours
- ...also run Akka Automated Operations, our managed platform for customer workloads on dedicated, BYOC, and BYOK8s deployments, backed by... ...and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions...CustomerRemote workFlexible hours
- ...solutions for data centers. The organization helps operators enhance customer experience, streamline day-to-day operations, and stay ahead... ...Skills/Qualifications: BS/MS degree in Computer Science, Engineering, or a related subject. Equivalent experience accepted....CustomerFull timeWork experience placementRemote workFlexible hours
- ...our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we... ...-code Experience improving system observability (e.g., custom metrics, traces, log pipelines) Why join us? Join a world...CustomerWorldwideShift work
$148.5k - $223.9k
...Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech... ...Job Details Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in SanFrancisco. Working closely with...CustomerWorldwideWeekend work- ...and the U.S. Special Forces. The Role We're hiring a Site Reliability Engineer to own the operational health of our connected sensor platform — spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it. This is a...CustomerRemote work
$300 per month
...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you will define and... ...Define and instrument SLOs and SLIs across customer workloads and internal services. Navigate ambiguity...CustomerFlexible hours- ...most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure... ...an essential part of our company, ensuring that we’re setting our businesses, clients, customers and employees up for success.Customer
- ...culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software... ...(SLOs) and thus delivering a smooth and frictionless Customer Experience. Site Reliability Engineer Role As an SRE...CustomerImmediate startRemote workWorldwide
$189k - $283.6k
...money. Afterpay is transforming the way customers manage their spending over time. TIDAL... ...proactively and reactively improve the reliability of Block's platform and critical infrastructure... ...desire to perform and grow as an engineer ~5+ years of software development...CustomerFull timeRelocation packageFlexible hoursShift work$117k - $209.33k
...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure... ..., and engineering standards needed to support critical customer-facing services.You will combine software engineering and production...CustomerFull timeFor contractors$153k - $191.3k
...and data company all rolled into one. Customers and users across the globe use Planet's... ...manufacturing, data processing, and software engineering, our office is a truly inspiring mix of... ...environments, to guarantee the reliability, scalability, and availability of our services...CustomerFull timeTemporary workFor contractorsWork at officeLocal areaRemote workHome office3 days per week- ...future of our cloud platform and champion engineering excellence across Ironclad. In this... ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud... ..., and provide support with internal or customer-facing incidents Translate the near...CustomerFull timeContract workWork at office
- ...this space is still largely uncharted. Engineers here are building agentic solutions to... ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud... ..., and provide support with internal or customer-facing incidents Translate the near,...CustomerFull timeContract workWork at office
$233.4k - $291.8k
...future of anime! About The Role We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US... ...cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability...CustomerFlexible hours$140k - $200k
...world. The Industrial Autonomy Team is seeking a senior reliability and verification engineer to design the test frameworks and validation strategies proving our products for autonomy and physical AI customers are safe and rugged. Within our small hardware team, you...CustomerFull timeWork experience placementLocal area$181k - $263k
## Senior Staff Site Reliability EngineerApplylocations: San Franciscotime type: Full timeposted... ...new standard for building a connected customer view with unmatched clarity and context... ...for a Senior Staff Site Reliability Engineer who will set the technical direction for...CustomerWork from homeFlexible hoursNight shift$200k - $260k
...enables the creation and operation of customer instances in our ecosystem in a standardized... ...Team as a technical leader driving reliability, automation, and scalability across the... ...practices across teams, mentor senior engineers, and be a primary escalation point for...CustomerCasual workWork at officeRemote workFlexible hours$173k - $230k
...employment Visa sponsorship. Role Summary The Principal Site Reliability Engineer applies software engineering and systems engineering... ...the ability to communicate with internal and/or external customers. Employee must be able to perform essential functions and...CustomerHourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Customer Reliability Engineer. Be the first to apply!
- senior reliability engineer San Francisco, CA
- sr reliability engineer San Francisco, CA
- reliability engineer San Francisco, CA
- reliability maintenance engineering technician San Francisco, CA
- customer San Francisco, CA
- customer engineer San Francisco, CA
- customer retention San Francisco, CA
- customer satisfaction San Francisco, CA
- amazon customer San Francisco, CA
- work from home customer San Francisco, CA



