Customer Reliability Engineer
Fluidstack
About Fluidstack We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.
We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI. We hire people who care deeply about this problem space. If that is you, please apply! How We Operate
We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI. We hire people who care deeply about this problem space. If that is you, please apply! How We Operate
- Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.
- Insane urgency. We drive everything forward as fast as possible.
- Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.
- Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.
- Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.
- Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.
- Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
- Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.
- Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
- Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
- Run customer-facing incident communication with technical depth and no spin.
- Turn recurring customer pain into engineering fixes with the production teams.
- The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.
- You've supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
- You debug distributed systems methodically across layers you don't own.
- You've written incident updates customers trusted more after reading.
- You push internal teams to fix causes, not symptoms, and follow up until they do.
- Bonus: GPU training workloads. InfiniBand or RoCE. Slurm or Kubernetes. NCCL debugging.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Customer Reliability Engineer in San Francisco, CA vacancy
$150k - $180k
...cutting-edge autonomous technologies, we are seeking a Senior Reliability Engineer (REL) to lead efforts in ensuring the long-term performance,... ...to uncover failure modes, validate off-the-shelf and custom hardware, and model expected field performance. Translate...CustomerFull timeImmediate startWorldwideFlexible hoursNight shift$140k - $200k
...efficient world.The Industrial Autonomy Team is seeking a senior reliability and verification engineer to design the test frameworks and validation strategies proving our products for autonomy and physical AI customers are safe and rugged. Within our small hardware team, you...CustomerWork experience placementLocal area$113.4k - $162k
...for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,... ...valuesCustomer ObsessedWe strive to have a deep understanding of our customers. Do Right By Our PeopleWe treat each other with fairness,...CustomerTemporary work$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB... ..., and edge load balancing, ensuring customer data remains safe in transit. Ultimately... ...plays a pivotal role in engineering the reliable, globally connected, multi-cloud network...CustomerLocal areaRemote workWorldwideFlexible hours$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range... ...build and ship products to delight our customers. We manage the end-to-end lifecycle of... ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager...CustomerWork at officeLocal areaRemote workWorldwideFlexible hours$165k - $227k
...this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products... ...scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset...CustomerLocal areaWorldwideFlexible hours$117k - $209.33k
...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure... ..., and engineering standards needed to support critical customer-facing services.You will combine software engineering and production...CustomerFull timeFor contractors$148.5k - $223.9k
...SalesforceSalesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech... ...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with...CustomerFull timeWorldwideWeekend work- ...world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology,... ...company, ensuring that we’re setting our businesses, clients, customers and employees up for success.Full timePosting Date: 2026-09...Customer
$150k - $190k
...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to... ...100M+ monthly open source downloads, 6,000+ active LangSmith customers, and 5 of the Fortune 10 use LangSmith in production (+ 35% of...CustomerPermanent employmentWork at officeFlexible hours$200k - $240k
...Senior Site Reliability Engineer (SRE)Location: San Francisco, CAWork Model: OnsiteIndustry: Renewable EnergyComp: $200,000 - $240,000 We... ...operation of production services. Occasionally travel to customer sites to support infrastructure deployments and troubleshoot...Customer- ...manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we... ...as-code Experience improving system observability (e.g., custom metrics, traces, log pipelines) Why join us? Join...CustomerWorldwideShift work
$207k - $300k
...systems by pushing for changes that improve reliability and velocity.Define the technical... ...Master's degree in Computer Science or Engineering, or a related field.Experience architecting... ...reliability, uptime appropriate to customer's needs and a fast rate of improvement....CustomerWorldwide$220k - $235k
...future of our cloud platform and champion engineering excellence across Ironclad. In this... ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud... ..., and provide support with internal or customer-facing incidentsTranslate the near, mid...CustomerFull timeContract workWork at office- ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI... ...platform — spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it. This is a...CustomerRemote work
$167.7k - $245.2k
...across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance... ...Observability portfolios. Your Impact As a Senior Site Reliability Engineer (SRE), you will lead the design and management of large-scale...CustomerFull timeTemporary workWork experience placementWork at officeLocal areaFlexible hours$150k - $220k
...incredible interest from investors, demand from customers, and a need to grow our team to meet... ...in this way. The Role: As an engineering organization, we pride ourselves on... ..., and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team...CustomerLocal area$194k - $267k
...strategic priorities—like reducing costs, and doing more for your customers.If you like to be challenged and have a passion for solving... ...on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing...CustomerPermanent employmentWork at officeLocal areaWorldwideFlexible hours- ...be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future... ...move fast and learn faster; obsess about creating customer value; value impact over activity; and embrace healthy...CustomerWorldwideHome officeFlexible hours
$170k - $220k
...Senior Site Reliability Engineer Supio is a trusted AI platform purpose-built for law firms, reshaping how data drives impactful outcomes... .... We go beyond surface-level AI to deeply understand our customers' daily needs, empowering law firms with unparalleled data insights...CustomerWork at officeRemote workFlexible hours$194k - $237k
...employment Visa sponsorship. Role Summary The Principal Site Reliability Engineer applies software engineering and systems engineering... ...the ability to communicate with internal and/or external customers. Employee must be able to perform essential functions and...CustomerHourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations... ...understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively...CustomerTemporary workImmediate startFlexible hoursShift work- ...infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on.... ...adapt as we grow. We have real paying customers and a playbook, and we still move at startup speed...Customer
$157k - $239k
...Infrastructure /Full time /On-siteWanna join the adventure?As a Site Reliability Engineer with strong networking skills in our Cloud Infrastructure (... ..., Loft’s flight heritage and proven technologies enable customers to focus on their mission objectives.With a growing fleet...CustomerFull timeTemporary work$167.7k - $245.2k
...across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance... ...portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in...CustomerFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$200k - $260k
...enables the creation and operation of customer instances in our ecosystem in a standardized... ...Team as a technical leader driving reliability, automation, and scalability across the... ...practices across teams, mentor senior engineers, and be a primary escalation point for...CustomerCasual workWork at officeRemote workFlexible hours$181k - $263k
...LiveRamp is setting the new standard for building a connected customer view with unmatched clarity and context while protecting... ...operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering...CustomerFull timeWork from homeWorldwideFlexible hoursNight shift$221.2k - $300k
Lead a team of engineers to maintain service uptime while managing global on-call rotations... ...improve operational practices to drive reliability, maintainability, and stakeholder... ...help developers build more sustainably. Customers in more than 200 countries and territories...CustomerFull timeWork at office$209.5k - $286.9k
Director, Technical Program Manager (Resiliency and Reliability Engineering) What You’ll Do: As a Director of Technical Program Management (TPM... ...scale products & platforms that will help Capital One customers have incredible experiences. In addition to the technical...CustomerFull timePart timeH1bLocal area- ...also run Akka Automated Operations, our managed platform for customer workloads on dedicated, BYOC, and BYOK8s deployments, backed by... ...and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions...CustomerRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Customer Reliability Engineer. Be the first to apply!
Related searches
- senior reliability engineer San Francisco, CA
- reliability engineer San Francisco, CA
- sr reliability engineer San Francisco, CA
- customer liaison San Francisco, CA
- customer retention San Francisco, CA
- work from home customer San Francisco, CA
- customer satisfaction San Francisco, CA
- customer marketing San Francisco, CA
- customer project program manager San Francisco, CA
- customer engineer San Francisco, CA


