Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$130k - $200k

Nscale

Site Reliability Engineer

Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each other the truth, and everyone here stays close to the infrastructure that makes AI work.

The Role

This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll be expected to make the systems you touch quieter over time.

What You'll Do

• Build and own the automation and tooling that keeps the platform running; treat operational toil as a bug to be fixed, not a fact of life. • Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance. • Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-incident reviews that actually change the system. • Investigate performance and reliability problems across Linux, networking, and distributed services, then fix them at the source. • Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the stack. • Improve availability, scalability, and efficiency through code, not manual effort.

What You'll Bring • 3-6 years in SRE, systems engineering, or software engineering, including time running production in a data center or cloud environment. • Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work away. • Solid command of Linux, networking fundamentals, and distributed systems. • A track record of troubleshooting live production issues and owning the fix through to the retro. • Fluency with monitoring and observability; metrics, logs, dashboards, and alerting. • Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be asked.

Nice to Have

• Experience with AI or GPU workloads, or high-performance computing (HPC). • Familiarity with high-performance networking (InfiniBand, RDMA). • Kubernetes, plus virtualized or bare-metal environments.

On-Call and Pace A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation, and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth acting on rather than just an interruption. The goal is to make the systems quieter over time, so each rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in better shape than you found it, you'll do well here.

What We Offer • Competitive base plus equity, reviewed every 12 months. • Real scope early, and a progression plan built around the skills you want to sharpen. • Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape your day.

Salary Range $130,000 - $200,000 USD. Actual compensation varies with skill set, experience, and location, and the role may be eligible for bonus and equity.

Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there's anything we can do to accommodate your specific situation, please let us know.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in San Francisco, CA vacancy
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Suggested

    Alembic

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,... 
    Suggested
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    1 day ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  •  ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production... 
    Suggested
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    2 days ago
  • $139.76k - $287.75k

     ...their business.We are seeking a Senior Site ReliabilityEngineer to help operate, scale...  ...will be instrumental in advancing the reliability, scalability, automation, observability,...  ...The ideal candidate is a highly hands-on engineer with strong production experience and a... 
    Suggested
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    1 day ago
  • $167.7k - $245.2k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    1 day ago
  • $117k - $209.33k

    Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    3 days ago
  • $113.4k - $162k

     ...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at... 
    Temporary work

    TextNow

    San Francisco, CA
    5 days ago
  • $152.5k - $205k

     ...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical... 
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  •  ...let’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of...  ..., working together to build scalable, reliable, and secure products that empower businesses...  ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work closely... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    2 days ago
  • $190.8k - $267.1k

     ...while helping Reddit grow its business. The reliability of our Ads systems directly impacts...  ...Reliability team partners closely with Ads Engineering teams to improve reliability,...  ...advertising ecosystem.We're looking for a Staff Site Reliability Engineer who will define and... 
    For contractors
    Work experience placement
    Remote work
    Flexible hours

    Reddit

    San Francisco, CA
    1 day ago
  • $165k - $227k

     ...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure... 
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • $150k

     ...About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our... 

    VantageScore

    San Francisco, CA
    1 day ago
  •  ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring through a focused Engineering Hiring Sprint...  ...: Platform Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers AI Agents Engineers... 
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    2 days ago
  •  ...human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to...  ...The Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible... 
    Remote work
    1 day per week

    Runloop AI, Inc

    San Francisco, CA
    15 hours ago
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description... 
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    San Francisco, CA
    3 days ago
  • $200k - $300k

     ...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round - Healthcare, AI Office Type: Onsite Salary: $200K-$300K Company Description We're representing a dynamic... 
    Work at office

    Recruiting from Scratch

    San Francisco, CA
    15 hours ago
  • $98.58k - $138.02k

     ...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering... 
    Work at office

    Restaurant365

    San Francisco, CA
    16 hours ago
  •  ...access to life-saving treatment. What We Look for in a Great Engineer You have the intensity and technical mastery to own mission...  ...high-velocity feature release while maintaining the highest reliability. DevX Support: Support Developer Experience (DevX) work to... 
    Work at office

    Latent

    San Francisco, CA
    3 days ago
  •  ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built... 
    Relocation package

    1872 Consulting

    San Francisco, CA
    16 hours ago
  • $150k - $250k

     ...Site Reliability Engineer role USC or GC only are considered at this time. San Francisco - Local to Bay area only but role is remote and occasion meeting required Latest update, 03/31/2026: The Site Reliability Engineer role is critical for... 
    Work experience placement
    Casual work
    Local area
    Immediate start
    Remote work

    3B Staffing LLC

    San Francisco, CA
    15 hours ago
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge... 
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    15 hours ago
  • $160k - $250k

     ...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content...  ...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering... 

    Hive

    San Francisco, CA
    14 hours ago
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    1 day ago
  • $163.71k - $306k

     ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system...  ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially... 

    Retool

    San Francisco, CA
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad business... 

    JP Morgan Chase

    San Francisco, CA
    3 days ago
  • $217k - $303.9k

     ...information, visit .As Reddit continues to scale globally, reliability and performance are more critical than ever. The Site Experience SRE team sits at the intersection of infrastructure, product engineering, and user experience - ensuring that every interaction across... 
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    2 days ago
  • $174k - $239k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!