Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Customer Reliability Engineer

Andromeda Cluster

Site Reliability Engineer - AI Infrastructure

Location: Global Remote / San Francisco · Full-Time

About Andromeda

Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.

We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.

Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.

Our long-term vision is to build the liquidity layer for global AI compute — a marketplace that moves the infrastructure and workloads powering AGI not dissimilar to the flows of capital in the world's financial markets.

We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.

What You’ll Do
  • Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers.

  • Build automation and tooling to streamline cluster deployments and integrations.

  • Debug customer issues across networking, storage, scheduling, and system layers.

  • Improve reliability and scalability of both training and inference infrastructure.

  • Design and implement monitoring, alerting, and observability for critical systems.

  • Collaborate with engineering and product teams to plan and deliver infrastructure for new services.

  • Participate in on-call and incident response, leading postmortems and reliability improvements.

What We’re Looking For
  • 5+ years experience in SRE, DevOps, or infrastructure engineering roles.

  • Strong Linux systems and networking fundamentals.

  • Deep experience with Kubernetes and container orchestration at scale.

  • Proficiency with Infrastructure-as-Code (Terraform, Helm, Ansible, etc.).

  • Strong automation and scripting skills (Python, Go, or Bash).

  • Experience with observability stacks (Prometheus, Grafana, Loki, Datadog, etc.).

  • Track record of operating production systems and leading incident response.

Nice to Have
  • Exposure to ML/AI infrastructure or GPU-based systems (CUDA, Slurm, Triton, etc.).

  • Familiarity with high-performance networking (InfiniBand, NVLink) or distributed storage (VAST, Weka, Ceph).

  • Customer-facing support or consulting experience.

Why You’ll Love It Here

This is a builder’s role. You’ll have ownership and autonomy to shape how our systems run, working directly with customers and providers while building the foundation for reliable, scalable AI infrastructure.

#J-18808-Ljbffr
Vacancy posted 10 hours ago
Similar jobs that could be interesting for youBased on the Customer Reliability Engineer in San Francisco, CA vacancy
  •  ...something that matters to the world. The Production Engineering Team Examples of key problems the team is working on...  ...you set become the standard. Role Scope Own reliability for named customer workloads: their clusters, their SLAs, their escalations... 
    Customer

    FluidStack

    San Francisco, CA
    4 days ago
  • $150k - $180k

     ...cutting-edge autonomous technologies, we are seeking a Senior Reliability Engineer (REL) to lead efforts in ensuring the long-term performance,...  ...to uncover failure modes, validate off-the-shelf and custom hardware, and model expected field performance. Translate... 
    Customer
    Full time
    Immediate start
    Worldwide
    Flexible hours
    Night shift

    Eight Sleep

    San Francisco, CA
    3 days ago
  • $115k - $160k

     ...Hardware Reliability Engineer Reno, Nevada, United States; San Francisco, California, United States Amperesand is reinventing how the...  ...meeting timeline and quality expectation to serve infrastructure customers around the world. Amperesand is led by breakthrough... 
    Customer
    Temporary work
    Local area
    Shift work

    Ampersand

    San Francisco, CA
    17 hours ago
  •  ...gets done — 94% of the Fortune 500 use it, and 45% are paying customers. We hit $100M ARR in May 2026 and have grown to over 5...  ...customers. About the Role We’re hiring a Senior Database Reliability Engineer to own the reliability, performance, and scalability of... 
    Customer
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours
    3 days per week

    scribehow.com

    San Francisco, CA
    10 hours ago
  •  ...Zof AI is seeking a Site Reliability Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply. This role owns the execution layer of our control plane: Kubernetes and container orchestration, CI/CD pipelines, hard... 
    Customer
    Full time

    Zof AI

    San Francisco, CA
    10 hours ago
  •  ...manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we...  ...as-code Experience improving system observability (e.g., custom metrics, traces, log pipelines) Why join us? Join... 
    Customer
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    3 days ago
  •  ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI...  ...platform — spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it. This is a... 
    Customer
    Remote work

    Specter Services LLC

    San Francisco, CA
    3 days ago
  •  ...believe there's a fit worth exploring Position Title: Site Reliability Engineer Location: Remote Duration: 12 Months Overview:...  ...this role: Deploy software for Cloud Prem and SAAS customers. Respond to and diagnose system incidents in a timely... 
    Customer
    Remote work

    HonorVet Technologies

    San Francisco, CA
    3 days ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...management. Define and instrument SLOs and SLIs across customer workloads and internal services. Navigate ambiguity, make... 
    Customer
    Flexible hours

    Baseten

    San Francisco, CA
    17 hours ago
  •  ...be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future...  ...move fast and learn faster; obsess about creating customer value; value impact over activity; and embrace healthy... 
    Customer
    Worldwide
    Home office
    Flexible hours

    Zoomcar

    San Francisco, CA
    10 hours ago
  • $150k - $250k

     ...Site Reliability Engineer role USC or GC only are considered at this time. San Francisco - Local to Bay area only but role...  ...role is critical for us right now - we have enterprise customers with urgent reliability issues that need immediate attention... 
    Customer
    Work experience placement
    Casual work
    Local area
    Immediate start
    Remote work

    3B Staffing LLC

    San Francisco, CA
    1 day ago
  •  ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And...  ...engineering candidate to join the Site Reliability organization in San Francisco. Working... 
    Customer
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    1 day ago
  •  ...OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software...  ...Objectives (SLOs) and thus delivering a smooth and frictionless Customer Experience. Site Reliability Engineer Role As an... 
    Customer
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    3 days ago
  • $200k - $240k

     ...Senior Site Reliability Engineer (SRE) Location: San Francisco, CA Work Model: Onsite Industry: Renewable Energy Comp: $200...  ...operation of production services. • Occasionally travel to customer sites to support infrastructure deployments and troubleshoot... 
    Customer

    Lawrence Harvey

    San Francisco, CA
    17 hours ago
  • $117k - $209.33k

     ...Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...automation, and engineering standards needed to support critical customer-facing services. You will combine software engineering... 
    Customer
    For contractors

    Autodesk

    San Francisco, CA
    2 days ago
  • $81.1k - $187k

     ...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations...  ...understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively... 
    Customer
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    San Francisco, CA
    2 days ago
  • $98.58k - $138.02k

     ...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique...  ...Familiarity with restaurant industry SaaS platforms and customer-facing applications. R365 Team Member Benefits & Compensation... 
    Customer
    Work at office

    Restaurant365

    San Francisco, CA
    1 day ago
  • $163.71k - $306k

     ...YOU: Good software has to run where customers need it. For many of Retool's largest...  ...infrastructure, behind their own controls, with the reliability and operational clarity they would...  ...TAMs to trust. Partner with product engineers on infrastructure requirements for new... 
    Customer

    Retool

    San Francisco, CA
    2 days ago
  • $160k - $250k

     ...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions...  ...-trained AI models, serving billions of customer API requests every month. Hive also offers...  ...we also need to grow our DevOps and Site Reliability team to maintain the reliability of our... 
    Customer

    Hive

    San Francisco, CA
    1 day ago
  •  ...platform for industrial and frontline teams. More than 13,000 customers , including Duracell, McDonald's, Shell, DHL and Volvo,...  ...happened, and it is ours. We’re looking for a Site Reliability Engineer to help advance MaintainX’s reliability, observability, and... 
    Customer
    Remote work

    MaintainX

    San Francisco, CA
    17 hours ago
  • $189k - $283.6k

     ...money. Afterpay is transforming the way customers manage their spending over time. TIDAL...  ...proactively and reactively improve the reliability of Block's platform and critical infrastructure...  ...desire to perform and grow as an engineer ~5+ years of software development... 
    Customer
    Full time
    Relocation package
    Flexible hours
    Shift work

    Block Inc

    San Francisco, CA
    10 hours ago
  •  ...Site Reliability Engineer San Francisco, CA We believe communication belongs to everyone. We exist to democratize phone service. TextNow...  ...sharing feedback, embracing change and more! Our values Customer Obsessed We strive to have a deep understanding of our customers... 
    Customer
    Temporary work

    TextNow

    San Francisco, CA
    2 days ago
  •  ...Job Title At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition... 
    Customer
    Temporary work
    Work experience placement

    Phenom People

    San Francisco, CA
    1 day ago
  • $153k - $191.3k

     ...and data company all rolled into one. Customers and users across the globe use Planet's...  ...manufacturing, data processing, and software engineering, our office is a truly inspiring mix of...  ...environments, to guarantee the reliability, scalability, and availability of our services... 
    Customer
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Remote work
    Home office
    3 days per week

    Planet Labs PBC

    San Francisco, CA
    17 hours ago
  •  ...future of our cloud platform and champion engineering excellence across Ironclad. In this...  ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud...  ..., and provide support with internal or customer-facing incidents Translate the near... 
    Customer
    Full time
    Contract work
    Work at office

    Ironclad Inc

    San Francisco, CA
    3 days ago
  • $150k - $220k

     ...Manager, Site Reliability Engineer San Francisco, California, United States At Forge, we know our team is our greatest asset. As technology...  ...generated incredible interest from investors, demand from customers, and a need to grow our team to meet the needs of more... 
    Customer
    Local area

    FORGE

    San Francisco, CA
    2 days ago
  • $194k - $267k

     ...strategic priorities-like reducing costs, and doing more for your customers. If you like to be challenged and have a passion for...  ...new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing... 
    Customer
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    3 days ago
  • $181k - $263k

     ...LiveRamp is setting the new standard for building a connected customer view with unmatched clarity and context while protecting...  ...operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering... 
    Customer
    Full time
    Work from home
    Worldwide
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    3 days ago
  • $200k - $260k

     ...enables the creation and operation of customer instances in our ecosystem in a standardized...  ...Team as a technical leader driving reliability, automation, and scalability across the...  ...practices across teams, mentor senior engineers, and be a primary escalation point for... 
    Customer
    Casual work
    Work at office
    Remote work
    Flexible hours

    Sight Machine

    San Francisco, CA
    17 hours ago
  •  ...future of anime! About the role We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US...  ...cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability... 
    Customer

    Engg

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Customer Reliability Engineer. Be the first to apply!