Customer Reliability Engineer
Andromeda Cluster
Site Reliability Engineer - AI Infrastructure
Location: Global Remote / San Francisco · Full-Time
About Andromeda
Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.
We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.
Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.
Our long-term vision is to build the liquidity layer for global AI compute — a marketplace that moves the infrastructure and workloads powering AGI not dissimilar to the flows of capital in the world's financial markets.
We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.
What You’ll Do
-
Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers.
-
Build automation and tooling to streamline cluster deployments and integrations.
-
Debug customer issues across networking, storage, scheduling, and system layers.
-
Improve reliability and scalability of both training and inference infrastructure.
-
Design and implement monitoring, alerting, and observability for critical systems.
-
Collaborate with engineering and product teams to plan and deliver infrastructure for new services.
-
Participate in on-call and incident response, leading postmortems and reliability improvements.
What We’re Looking For
-
5+ years experience in SRE, DevOps, or infrastructure engineering roles.
-
Strong Linux systems and networking fundamentals.
-
Deep experience with Kubernetes and container orchestration at scale.
-
Proficiency with Infrastructure-as-Code (Terraform, Helm, Ansible, etc.).
-
Strong automation and scripting skills (Python, Go, or Bash).
-
Experience with observability stacks (Prometheus, Grafana, Loki, Datadog, etc.).
-
Track record of operating production systems and leading incident response.
Nice to Have
-
Exposure to ML/AI infrastructure or GPU-based systems (CUDA, Slurm, Triton, etc.).
-
Familiarity with high-performance networking (InfiniBand, NVLink) or distributed storage (VAST, Weka, Ceph).
-
Customer-facing support or consulting experience.
Why You’ll Love It Here
This is a builder’s role. You’ll have ownership and autonomy to shape how our systems run, working directly with customers and providers while building the foundation for reliable, scalable AI infrastructure.
#J-18808-Ljbffr- ...something that matters to the world. The Production Engineering Team Examples of key problems the team is working on... ...you set become the standard. Role Scope Own reliability for named customer workloads: their clusters, their SLAs, their escalations...Customer
$150k - $180k
...cutting-edge autonomous technologies, we are seeking a Senior Reliability Engineer (REL) to lead efforts in ensuring the long-term performance,... ...to uncover failure modes, validate off-the-shelf and custom hardware, and model expected field performance. Translate...CustomerFull timeImmediate startWorldwideFlexible hoursNight shift$115k - $160k
...Hardware Reliability Engineer Reno, Nevada, United States; San Francisco, California, United States Amperesand is reinventing how the... ...meeting timeline and quality expectation to serve infrastructure customers around the world. Amperesand is led by breakthrough...CustomerTemporary workLocal areaShift work- ...gets done — 94% of the Fortune 500 use it, and 45% are paying customers. We hit $100M ARR in May 2026 and have grown to over 5... ...customers. About the Role We’re hiring a Senior Database Reliability Engineer to own the reliability, performance, and scalability of...CustomerFull timeWork at officeRemote workHome officeFlexible hours3 days per week
- ...Zof AI is seeking a Site Reliability Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply. This role owns the execution layer of our control plane: Kubernetes and container orchestration, CI/CD pipelines, hard...CustomerFull time
- ...manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we... ...as-code Experience improving system observability (e.g., custom metrics, traces, log pipelines) Why join us? Join...CustomerWorldwideShift work
- ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI... ...platform — spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it. This is a...CustomerRemote work
- ...believe there's a fit worth exploring Position Title: Site Reliability Engineer Location: Remote Duration: 12 Months Overview:... ...this role: Deploy software for Cloud Prem and SAAS customers. Respond to and diagnose system incidents in a timely...CustomerRemote work
- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...management. Define and instrument SLOs and SLIs across customer workloads and internal services. Navigate ambiguity, make...CustomerFlexible hours
- ...be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future... ...move fast and learn faster; obsess about creating customer value; value impact over activity; and embrace healthy...CustomerWorldwideHome officeFlexible hours
$150k - $250k
...Site Reliability Engineer role USC or GC only are considered at this time. San Francisco - Local to Bay area only but role... ...role is critical for us right now - we have enterprise customers with urgent reliability issues that need immediate attention...CustomerWork experience placementCasual workLocal areaImmediate startRemote work- ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And... ...engineering candidate to join the Site Reliability organization in San Francisco. Working...CustomerWorldwideWeekend work
- ...OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software... ...Objectives (SLOs) and thus delivering a smooth and frictionless Customer Experience. Site Reliability Engineer Role As an...CustomerImmediate startRemote workWorldwide
$200k - $240k
...Senior Site Reliability Engineer (SRE) Location: San Francisco, CA Work Model: Onsite Industry: Renewable Energy Comp: $200... ...operation of production services. • Occasionally travel to customer sites to support infrastructure deployments and troubleshoot...Customer$117k - $209.33k
...Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...automation, and engineering standards needed to support critical customer-facing services. You will combine software engineering...CustomerFor contractors$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations... ...understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively...CustomerTemporary workImmediate startFlexible hoursShift work$98.58k - $138.02k
...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique... ...Familiarity with restaurant industry SaaS platforms and customer-facing applications. R365 Team Member Benefits & Compensation...CustomerWork at office$163.71k - $306k
...YOU: Good software has to run where customers need it. For many of Retool's largest... ...infrastructure, behind their own controls, with the reliability and operational clarity they would... ...TAMs to trust. Partner with product engineers on infrastructure requirements for new...Customer$160k - $250k
...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions... ...-trained AI models, serving billions of customer API requests every month. Hive also offers... ...we also need to grow our DevOps and Site Reliability team to maintain the reliability of our...Customer- ...platform for industrial and frontline teams. More than 13,000 customers , including Duracell, McDonald's, Shell, DHL and Volvo,... ...happened, and it is ours. We’re looking for a Site Reliability Engineer to help advance MaintainX’s reliability, observability, and...CustomerRemote work
$189k - $283.6k
...money. Afterpay is transforming the way customers manage their spending over time. TIDAL... ...proactively and reactively improve the reliability of Block's platform and critical infrastructure... ...desire to perform and grow as an engineer ~5+ years of software development...CustomerFull timeRelocation packageFlexible hoursShift work- ...Site Reliability Engineer San Francisco, CA We believe communication belongs to everyone. We exist to democratize phone service. TextNow... ...sharing feedback, embracing change and more! Our values Customer Obsessed We strive to have a deep understanding of our customers...CustomerTemporary work
- ...Job Title At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition...CustomerTemporary workWork experience placement
$153k - $191.3k
...and data company all rolled into one. Customers and users across the globe use Planet's... ...manufacturing, data processing, and software engineering, our office is a truly inspiring mix of... ...environments, to guarantee the reliability, scalability, and availability of our services...CustomerFull timeTemporary workFor contractorsWork at officeLocal areaRemote workHome office3 days per week- ...future of our cloud platform and champion engineering excellence across Ironclad. In this... ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud... ..., and provide support with internal or customer-facing incidents Translate the near...CustomerFull timeContract workWork at office
$150k - $220k
...Manager, Site Reliability Engineer San Francisco, California, United States At Forge, we know our team is our greatest asset. As technology... ...generated incredible interest from investors, demand from customers, and a need to grow our team to meet the needs of more...CustomerLocal area$194k - $267k
...strategic priorities-like reducing costs, and doing more for your customers. If you like to be challenged and have a passion for... ...new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing...CustomerPermanent employmentWork at officeLocal areaWorldwideFlexible hours$181k - $263k
...LiveRamp is setting the new standard for building a connected customer view with unmatched clarity and context while protecting... ...operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering...CustomerFull timeWork from homeWorldwideFlexible hoursNight shift$200k - $260k
...enables the creation and operation of customer instances in our ecosystem in a standardized... ...Team as a technical leader driving reliability, automation, and scalability across the... ...practices across teams, mentor senior engineers, and be a primary escalation point for...CustomerCasual workWork at officeRemote workFlexible hours- ...future of anime! About the role We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US... ...cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability...Customer
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Customer Reliability Engineer. Be the first to apply!
- senior reliability engineer San Francisco, CA
- sr reliability engineer San Francisco, CA
- reliability engineer San Francisco, CA
- reliability maintenance engineering technician San Francisco, CA
- customer San Francisco, CA
- customer engineer San Francisco, CA
- customer retention San Francisco, CA
- customer satisfaction San Francisco, CA
- amazon customer San Francisco, CA
- work from home customer San Francisco, CA

