Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Customer Reliability Engineer

FluidStack

About Fluidstack

We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.


We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.

We hire people who care deeply about this problem space. If that is you, please apply!

How We Operate
  • Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.
  • Insane urgency. We drive everything forward as fast as possible.
  • Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.
  • Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.
  • Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.
The Data Center Operations Team

Examples of key problems the team is working on
  • Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.
  • Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
  • Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.
Role Scope
  • Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
  • Run customer-facing incident communication with technical depth and no spin.
  • Turn recurring customer pain into engineering fixes with the production teams.
What We're Looking For
  • The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.
  • You've supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • You debug distributed systems methodically across layers you don't own.
  • You've written incident updates customers trusted more after reading.
  • You push internal teams to fix causes, not symptoms, and follow up until they do.
  • Bonus: GPU training workloads. InfiniBand or RoCE. Slurm or Kubernetes. NCCL debugging.

We are committed to pay equity and transparency.

Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email View email address on click.appcast.io with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Customer Reliability Engineer in San Francisco, CA vacancy
  • $150k - $180k

     ...cutting-edge autonomous technologies, we are seeking a Senior Reliability Engineer (REL) to lead efforts in ensuring the long-term performance,...  ...to uncover failure modes, validate off-the-shelf and custom hardware, and model expected field performance. Translate... 
    Customer
    Full time
    Immediate start
    Worldwide
    Flexible hours
    Night shift

    Eight Sleep

    San Francisco, CA
    4 days ago
  • $140k - $200k

     ...efficient world.The Industrial Autonomy Team is seeking a senior reliability and verification engineer to design the test frameworks and validation strategies proving our products for autonomy and physical AI customers are safe and rugged. Within our small hardware team, you... 
    Customer
    Work experience placement
    Local area

    Ouster

    San Francisco, CA
    2 days ago
  •  ...be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future...  ...move fast and learn faster; obsess about creating customer value; value impact over activity; and embrace healthy... 
    Customer
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    3 days ago
  • $113.4k - $162k

     ...for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,...  ...valuesCustomer ObsessedWe strive to have a deep understanding of our customers. Do Right By Our PeopleWe treat each other with fairness,... 
    Customer
    Temporary work

    TextNow

    San Francisco, CA
    1 day ago
  • $152.5k - $205k

     ...stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and...  ...performance, security, and cost-effectiveness of the systems our customers depend on.What you'll work on: Design, build, and operate Kubernetes... 
    Customer
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range...  ...build and ship products to delight our customers. We manage the end-to-end lifecycle of...  ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... 
    Customer
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    5 days ago
  • $165k - $227k

     ...this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products...  ...scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset... 
    Customer
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $117k - $209.33k

     ...OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure...  ..., and engineering standards needed to support critical customer-facing services.You will combine software engineering and production... 
    Customer
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    4 days ago
  •  ...world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology,...  ...company, ensuring that we’re setting our businesses, clients, customers and employees up for success.Full timePosting Date: 2026-09... 
    Customer

    JP Morgan Chase

    San Francisco, CA
    4 days ago
  • $165k - $225.6k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability...  ...best practices, deliver excellent internal customer service, and actively contribute to Agile... 
    Customer
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $167.7k - $245.2k

     ...technology portfolio and beyond, helping customers deploy at scale while also delivering...  ...very effective.We’re looking for talented engineers with a software or operations...  ...application development teams to ensure the reliability, performance and security of our infrastructure... 
    Customer
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech...  ...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with... 
    Customer
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    2 days ago
  • $150k - $190k

     ...intelligent agents ubiquitous. We build the foundation for agent engineering in the real world, helping developers move from prototypes to...  ...100M+ monthly open source downloads, 6,000+ active LangSmith customers, and 5 of the Fortune 10 use LangSmith in production (+ 35% of... 
    Customer
    Permanent employment
    Work at office
    Flexible hours

    LangChain

    San Francisco, CA
    4 days ago
  • $140k - $180k

     ...delivering critical supplies quickly and reliably. Today, Zipline operates on four...  ...supplies, food, and retail products.  Our customers include the world's largest and most prominent...  ...Senior Mechanical Hardware Reliability Engineer, you will own the reliability of... 
    Customer
    Local area

    Zipline

    San Francisco, CA
    8 days ago
  • $207k - $300k

     ...systems by pushing for changes that improve reliability and velocity.Define the technical...  ...Master's degree in Computer Science or Engineering, or a related field.Experience architecting...  ...reliability, uptime appropriate to customer's needs and a fast rate of improvement.... 
    Customer
    Worldwide

    Google

    San Francisco, CA
    3 days ago
  • $150k - $220k

     ...incredible interest from investors, demand from customers, and a need to grow our team to meet...  ...in this way. The Role: As an engineering organization, we pride ourselves on...  ..., and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team... 
    Customer
    Local area

    Forge Global

    San Francisco, CA
    2 days ago
  • $220k - $235k

     ...future of our cloud platform and champion engineering excellence across Ironclad. In this...  ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud...  ..., and provide support with internal or customer-facing incidentsTranslate the near, mid... 
    Customer
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    2 days ago
  • $194k - $267k

     ...strategic priorities—like reducing costs, and doing more for your customers.If you like to be challenged and have a passion for solving...  ...on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing... 
    Customer
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $194k - $267k

     ...talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is...  ...highly reliable, scalable, and secure cloud services that our customers can trust. We embrace an automation-first mindset and... 
    Customer
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $174k - $239k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for...  ...by product teamsDeliver excellent internal customer service and advocate for SRE and DevOps practices... 
    Customer
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    5 days ago
  •  ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI...  ...platform — spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it. This is a... 
    Customer
    Remote work

    Specter Services LLC

    San Francisco, CA
    4 days ago
  •  ...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the...  ...infrastructure for our users that scales, is reliable, and makes the complexities of operating...  .... Build the systems enterprise customers expect: rate limiting, authentication, usage... 
    Customer
    Permanent employment
    Shift work

    Arena AI

    San Francisco, CA
    7 hours ago
  •  ...manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we...  ...as-code Experience improving system observability (e.g., custom metrics, traces, log pipelines) Why join us? Join... 
    Customer
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    4 days ago
  •  ...infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on....  ...adapt as we grow. We have real paying customers and a playbook, and we still move at startup speed... 
    Customer

    Alembic

    San Francisco, CA
    5 days ago
  • $170k - $220k

     ...Senior Site Reliability Engineer Supio is a trusted AI platform purpose-built for law firms, reshaping how data drives impactful outcomes...  .... We go beyond surface-level AI to deeply understand our customers' daily needs, empowering law firms with unparalleled data insights... 
    Customer
    Work at office
    Remote work
    Flexible hours

    Supio

    San Francisco, CA
    3 days ago
  • $100k - $160k

     ...delivering critical supplies quickly and reliably. Today, Zipline operates on four...  ...supplies, food, and retail products.  Our customers include the world's largest and most prominent...  ...As a Mechanical Hardware Reliability Engineer, you will be responsible for defining and... 
    Customer
    Local area

    Zipline

    San Francisco, CA
    8 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure...  ...MongoDBMongoDB is built for change, empowering our customers and our people to innovate at the speed of the market. We have... 
    Customer
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    3 days ago
  • $167.7k - $245.2k

    Meet the TeamThe Cisco Meraki cloud supports millions of customer devices from 8 physical data centers and several cloud regions...  ...that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available... 
    Customer
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Francisco, CA
    2 days ago
  • $157k - $239k

     ...Infrastructure /Full time /On-siteWanna join the adventure?As a Site Reliability Engineer with strong networking skills in our Cloud Infrastructure (...  ..., Loft’s flight heritage and proven technologies enable customers to focus on their mission objectives.With a growing fleet... 
    Customer
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    1 day ago
  • $167.7k - $245.2k

     ...across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance...  ...portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in... 
    Customer
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Customer Reliability Engineer. Be the first to apply!