Site Reliability Engineer
$130k - $200kNscale
Site Reliability Engineer
Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each other the truth, and everyone here stays close to the infrastructure that makes AI work.
The Role
This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll be expected to make the systems you touch quieter over time.
What You'll Do
• Build and own the automation and tooling that keeps the platform running; treat operational toil as a bug to be fixed, not a fact of life. • Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance. • Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-incident reviews that actually change the system. • Investigate performance and reliability problems across Linux, networking, and distributed services, then fix them at the source. • Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the stack. • Improve availability, scalability, and efficiency through code, not manual effort.
What You'll Bring • 3-6 years in SRE, systems engineering, or software engineering, including time running production in a data center or cloud environment. • Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work away. • Solid command of Linux, networking fundamentals, and distributed systems. • A track record of troubleshooting live production issues and owning the fix through to the retro. • Fluency with monitoring and observability; metrics, logs, dashboards, and alerting. • Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be asked.
Nice to Have
• Experience with AI or GPU workloads, or high-performance computing (HPC). • Familiarity with high-performance networking (InfiniBand, RDMA). • Kubernetes, plus virtualized or bare-metal environments.
On-Call and Pace A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation, and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth acting on rather than just an interruption. The goal is to make the systems quieter over time, so each rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in better shape than you found it, you'll do well here.
What We Offer • Competitive base plus equity, reviewed every 12 months. • Real scope early, and a progression plan built around the skills you want to sharpen. • Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape your day.
Salary Range $130,000 - $200,000 USD. Actual compensation varies with skill set, experience, and location, and the role may be eligible for bonus and equity.
Equal Opportunities Statement
At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there's anything we can do to accommodate your specific situation, please let us know.
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice.
- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Suggested
$148.5k - $223.9k
...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,...SuggestedFull timeWorldwideWeekend work$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,...SuggestedPermanent employmentLocal areaWorldwideFlexible hours- ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production...SuggestedWorldwideHome officeFlexible hours
$139.76k - $287.75k
...their business.We are seeking a Senior Site ReliabilityEngineer to help operate, scale... ...will be instrumental in advancing the reliability, scalability, automation, observability,... ...The ideal candidate is a highly hands-on engineer with strong production experience and a...SuggestedWork at officeLocal areaRelocationRelocation package$167.7k - $245.2k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$117k - $209.33k
Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...Full timeFor contractors$113.4k - $162k
...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at...Temporary work$152.5k - $205k
...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical...Flexible hours- ...let’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ..., working together to build scalable, reliable, and secure products that empower businesses... ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work closely...Temporary workLocal areaWorldwide
$190.8k - $267.1k
...while helping Reddit grow its business. The reliability of our Ads systems directly impacts... ...Reliability team partners closely with Ads Engineering teams to improve reliability,... ...advertising ecosystem.We're looking for a Staff Site Reliability Engineer who will define and...For contractorsWork experience placementRemote workFlexible hours$165k - $227k
...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable...Local areaWorldwideFlexible hours$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure...Flexible hours$150k
...About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our...- ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring through a focused Engineering Hiring Sprint... ...: Platform Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers AI Agents Engineers...Work at officeLocal areaFlexible hours
- ...human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to... ...The Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible...Remote work1 day per week
$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description...Work at officeVisa sponsorshipFlexible hours$200k - $300k
...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round - Healthcare, AI Office Type: Onsite Salary: $200K-$300K Company Description We're representing a dynamic...Work at office$98.58k - $138.02k
...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering...Work at office- ...access to life-saving treatment. What We Look for in a Great Engineer You have the intensity and technical mastery to own mission... ...high-velocity feature release while maintaining the highest reliability. DevX Support: Support Developer Experience (DevX) work to...Work at office
- ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built...Relocation package
$150k - $250k
...Site Reliability Engineer role USC or GC only are considered at this time. San Francisco - Local to Bay area only but role is remote and occasion meeting required Latest update, 03/31/2026: The Site Reliability Engineer role is critical for...Work experience placementCasual workLocal areaImmediate startRemote work- ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge...Work at officeWeekend work
$160k - $250k
...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content... ...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering...- ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely... ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale...Permanent employmentWork experience placementWork at officeLocal area
$163.71k - $306k
...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system... ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially...$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad business...
$217k - $303.9k
...information, visit .As Reddit continues to scale globally, reliability and performance are more critical than ever. The Site Experience SRE team sits at the intersection of infrastructure, product engineering, and user experience - ensuring that every interaction across...For contractorsWork experience placement$174k - $239k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability...Work experience placementLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer San Francisco, CA
- site reliability engineer remote San Francisco, CA
- site reliability engineer sre San Francisco, CA
- site recruiter San Francisco, CA
- junior website developer San Francisco, CA
- on site coordinator San Francisco, CA
- construction site safety San Francisco, CA
- site services specialist San Francisco, CA
- website content developer San Francisco, CA
- website coordinator San Francisco, CA

