Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Hyperbolic Labs

Who We Are

Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By aggregating computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.

As we prepare for growth after our Series A, our team - led by co-founders with PhDs in AI, Math, and Computer Science - is poised to redefine computing.

About the Role

We're seeking a Site Reliability Engineer to ensure Hyperbolic's GPU marketplace and AI infrastructure operate with exceptional reliability, performance, and security. As an aggregator of compute resources from hundreds of global suppliers, our SLOs, trust, and economic efficiency are product-critical. You'll be responsible for defining and maintaining service level objectives for job success rates, building robust incident response systems, managing capacity across our distributed GPU network, and implementing secure rollout and rollback mechanisms that keep our platform running smoothly 24/7.

In this role, you'll establish the reliability standards that define customer trust in our platform, design monitoring and alerting systems that provide deep visibility into our infrastructure, build automation for capacity management and resource allocation, lead incident response and post-mortem processes, and work closely with engineering teams to improve system resilience. You'll also focus on security and infrastructure hardening, ensuring strong isolation between tenants and suppliers, implementing key management systems, and building compliance frameworks. This is a high-impact position where your work directly influences our ability to deliver on our promise of affordable, accessible AI compute at scale.

Who You Are
  • Architected, deployed, and managed large-scale Kubernetes environments, including cluster administration, container orchestration, autoscaling, service discovery, and high-availability infrastructure to ensure reliability and scalability of mission-critical systems.
  • Led troubleshooting and performance optimization efforts across Kubernetes-based production environments, proactively identifying system bottlenecks, automating remediation workflows, and improving overall platform stability and uptime.
  • Strong automation mindset with experience using infrastructure-as-code, configuration management, and CI/CD pipelines
  • Strong background in capacity planning and management, including forecasting, resource allocation, and cost optimization for distributed systems
  • Experienced in incident response, on-call rotations, and post-mortem processes with a track record of reducing MTTR and improving system resilience
  • Deep knowledge of deployment systems including progressive rollouts, canary deployments, feature flags, and automated rollback mechanisms
  • Proficient in observability tools and practices including metrics, logging, tracing, and alerting systems (Prometheus, Grafana, ELK stack, or similar)
  • Strong understanding of infrastructure security including tenant isolation, workload isolation, network segmentation, and security hardening
  • Experience with secrets management, key management systems (KMS), certificate management, and secure credential rotation
  • Expert in site reliability engineering with proven experience defining, monitoring, and maintaining SLOs and SLAs for production systems
  • Knowledge of compliance frameworks and security best practices for cloud platforms (SOC 2, ISO 27001, or similar)
  • Excellent problem-solving skills with ability to debug complex distributed systems issues under pressure
Preferred Qualifications
  • Experience operating GPU infrastructure, AI/ML platforms, or compute marketplaces at scale
  • Background in distributed systems, peer-to-peer networks, or decentralized infrastructure
  • Knowledge of multi-tenancy security patterns, container security, and runtime security tools
  • Experience with chaos engineering, fault injection, and resilience testing
  • Familiarity with cost optimization strategies for cloud infrastructure and GPU resources
  • Experience building and operating systems with demanding uptime requirements (99.9%+ SLAs)
  • Background at companies like AWS, Google Cloud, Azure, or fast-growing infrastructure startups
  • Contributions to open-source reliability, observability, or security tools

Hyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in San Francisco, CA vacancy
  •  ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient... 
    Senior

    TechChain Talent

    San Francisco, CA
    4 days ago
  •  ...Udaip Cloud-Based Data And Ai Platform Engineer At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it... 
    Senior
    Temporary work
    Work experience placement

    Phenom People

    San Francisco, CA
    3 days ago
  • $210k - $240k

    Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $... 
    Senior
    Full time

    Alembic Technologies

    San Francisco, CA
    1 day ago
  • $210.8k - $272.8k

    About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user... 
    Senior
    Local area

    Thumbtack

    San Francisco, CA
    14 hours ago
  • $175k - $250k

     ...00.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance...  ...scalability, performance, and reliability across environments. What You’ll Do... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    1 day ago
  •  ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and...  ...platform capabilities in partnership with architects and product engineering Build a world-class observability platform and monitoring... 
    Senior

    Gravity Engineering Services Pvt Ltd.

    San Francisco, CA
    4 days ago
  • $174.92k - $209.91k

     ...: to make access to data as simple and reliable as electricity. With Fivetran, customer...  ...canonical and ready to query, with no engineering or maintenance required. We’re proud that...  ...integrate our teams, systems, and career sites. About the Role Fivetran is building... 
    Senior
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    2 days ago
  •  ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will... 
    Senior

    Kody

    San Francisco, CA
    2 days ago
  • $300k

     ...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring... 
    Senior
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $250k

     ...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves... 
    Senior
    Permanent employment
    Remote work
    San Francisco, CA
    a month ago
  •  ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical role in ensuring these systems...  ...hands-on engineering and technical leadership role. You will own reliability for major platform domains, design scalable solutions on Kubernetes... 
    Senior

    Saviynt

    San Francisco, CA
    15 days ago
  •  ...human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to...  ...The Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible... 
    Remote work
    1 day per week

    Runloop AI, Inc

    San Francisco, CA
    3 days ago
  • $100k - $170k

     ...Site Reliability Engineer Houston; San Francisco; Seattle About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services... 
    Flexible hours
    Shift work

    Nscale

    San Francisco, CA
    4 days ago
  • $98.58k - $138.02k

     ...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering... 
    Work at office

    Restaurant365

    San Francisco, CA
    4 days ago
  • $150k

     ...Site Reliability Engineer San Francisco, CA About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security... 

    VantageScore®

    San Francisco, CA
    14 hours ago
  •  ...mission-critical industries, helping partners move more quickly and reliably from algorithm to silicon. Our platform accelerates deployment...  .... The Roles We are looking for an experienced software engineer to help us build a new generation of transpilation tools... 
    Senior
    Full time
    Remote work
    Relocation package
    Flexible hours

    Code Metal

    San Francisco, CA
    14 hours ago
  •  ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built... 
    Relocation package

    1872 Consulting

    San Francisco, CA
    3 days ago
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description... 
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    San Francisco, CA
    1 day ago
  • $86k - $105k

     ...generation of application infrastructure and to be responsible for reliability, automation and scalability using and the latest best...  ...certifications. Minimum of 2 years prior DevOps, software engineering or related experience. Must be able to work different schedules... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    San Francisco, CA
    1 day ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,... 
    Work at office
    Local area
    1 day per week

    Mithril

    San Francisco, CA
    3 days ago
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge... 
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    3 days ago
  • $200k - $300k

     ...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round - Healthcare, AI Office Type: Onsite Salary: $200K-$300K Company Description We're representing a dynamic... 
    Work at office

    Recruiting from Scratch

    San Francisco, CA
    3 days ago
  • The company The future of data lies in decentralization, and the concept of a data mesh is the proven approach for implementing this at Enterprise scale. We’re here to make it a reality. Nextdata OS is a data-mesh-native platform built to meet the challenge of decentralizing...
    Senior
    Full time

    Nextdata

    San Francisco, CA
    14 hours ago
  • $205k - $305k

     ...Director Of Site Reliability Engineering Interested in working on cutting-edge blockchain technology and creating equitable access to the global...  ..., operate, and improve production services. This is a senior engineering leadership role reporting to the CTO. You will... 
    Temporary work
    Work at office
    Local area
    Worldwide
    Flexible hours

    Stellar

    San Francisco, CA
    9 days ago
  •  ...Staff Site Reliability Engineer (SRE) Location: San Francisco, CA Job Responsibilities As our Staff SRE, you'll be the primary expert...  ...entire SRE and engineering organization. Mentor mid-level and senior engineers on design patterns, operational rigor, and... 

    United IT

    San Francisco, CA
    4 days ago
  • $204k - $281k

     .... This is an opportunity to do career-defining work. We’re all in on this mission. If you are too, let’s talk. Manager, Site Reliability Engineering San Francisco, California Okta authenticates, authorizes and provisions millions of users a day. The service is hosted on... 
    Permanent employment
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    4 days ago
  •  ...General Intelligence (AGI). Safety is more important to us than unfettered growth. About the Role We are looking for a senior software engineer to build the foundational platform for identity across all OpenAI products. This involves building authentication,... 
    Senior
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  •  ...focused on building native actions — an agentic framework that understands you, and works reliably. We’re a team of AI researchers, designers, growth experts, and engineers rethinking human-computer interaction from the ground up. We value high-agency teammates who... 
    Senior
    Full time

    Wispr Flow

    San Francisco, CA
    14 hours ago
  •  ...company valued at $10 billion. We work in‑person five days a week in our new San Francisco headquarters. About the Role As a Site Reliability Engineer (SRE) at Mercor, you’ll own production reliability across our most critical systems, partnering directly with... 

    Mercor

    San Francisco, CA
    1 day ago
  •  ...What we do Idler builds reinforcement learning environments that teach AI models to code like 0.01% engineers. Our training environments are based on real-world coding scenarios that frontier models will actually encounter. We've closed a multimillion-dollar contract... 
    Senior
    Full time
    Contract work
    Relocation package

    Idler

    San Francisco, CA
    14 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!