Site Reliability Engineer
$100k - $170kNscale
About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native
startups and global enterprises, from bare metal up through the platform services teams actually build
on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each
other the truth, and everyone here stays close to the infrastructure that makes AI work. The Role
This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take
real surface area: the automation and tooling other engineers depend on, and the reliability of
production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll
be expected to make the systems you touch quieter over time. What You'll Do
• Build and own the automation and tooling that keeps the platform running; treat operational toil as
a bug to be fixed, not a fact of life.
• Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance.
• Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-
incident reviews that actually change the system.
• Investigate performance and reliability problems across Linux, networking, and distributed services,
then fix them at the source.
• Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the
stack.
• Improve availability, scalability, and efficiency through code, not manual effort.
What You'll Bring
• 3-6 years in SRE, systems engineering, or software engineering, including time running production in
a data center or cloud environment.
• Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work
away.
• Solid command of Linux, networking fundamentals, and distributed systems.
• A track record of troubleshooting live production issues and owning the fix through to the retro.
• Fluency with monitoring and observability; metrics, logs, dashboards, and alerting.
• Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be
asked. Nice to Have
• Experience with AI or GPU workloads, or high-performance computing (HPC).
• Familiarity with high-performance networking (InfiniBand, RDMA).
• Kubernetes, plus virtualized or bare-metal environments.
On-Call and Pace
A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation,
and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth
acting on rather than just an interruption. The goal is to make the systems quieter over time, so each
rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in
better shape than you found it, you'll do well here.
What We Offer
• Competitive base plus equity, reviewed every 12 months.
• Real scope early, and a progression plan built around the skills you want to sharpen.
• Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape
your day. Salary Range
$100,000 - $170,000 USD. Actual compensation varies with skill set, experience, and location, and the
role may be eligible for bonus and equity. Equal Opportunities Statement At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know. The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $100,000-$170,000 USD For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
startups and global enterprises, from bare metal up through the platform services teams actually build
on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each
other the truth, and everyone here stays close to the infrastructure that makes AI work. The Role
This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take
real surface area: the automation and tooling other engineers depend on, and the reliability of
production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll
be expected to make the systems you touch quieter over time. What You'll Do
• Build and own the automation and tooling that keeps the platform running; treat operational toil as
a bug to be fixed, not a fact of life.
• Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance.
• Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-
incident reviews that actually change the system.
• Investigate performance and reliability problems across Linux, networking, and distributed services,
then fix them at the source.
• Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the
stack.
• Improve availability, scalability, and efficiency through code, not manual effort.
What You'll Bring
• 3-6 years in SRE, systems engineering, or software engineering, including time running production in
a data center or cloud environment.
• Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work
away.
• Solid command of Linux, networking fundamentals, and distributed systems.
• A track record of troubleshooting live production issues and owning the fix through to the retro.
• Fluency with monitoring and observability; metrics, logs, dashboards, and alerting.
• Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be
asked. Nice to Have
• Experience with AI or GPU workloads, or high-performance computing (HPC).
• Familiarity with high-performance networking (InfiniBand, RDMA).
• Kubernetes, plus virtualized or bare-metal environments.
On-Call and Pace
A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation,
and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth
acting on rather than just an interruption. The goal is to make the systems quieter over time, so each
rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in
better shape than you found it, you'll do well here.
What We Offer
• Competitive base plus equity, reviewed every 12 months.
• Real scope early, and a progression plan built around the skills you want to sharpen.
• Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape
your day. Salary Range
$100,000 - $170,000 USD. Actual compensation varies with skill set, experience, and location, and the
role may be eligible for bonus and equity. Equal Opportunities Statement At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know. The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $100,000-$170,000 USD For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Seattle, WA vacancy
- ...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III - DevOps Engineer at JPMorgan Chase within the Commercial and Investment Bank, you will solve complex and broad business...Suggested
- ...Job Title: Site Reliability Engineer Location: Seattle, WA FTE Only Job Description Must Have Technical/Functional Skills • 15+ years of IT experience with at least 5+ years in API management and integration architecture. •...Suggested
- ...Lululemon Site Reliability Engineering Engineer We are a yoga-inspired technical apparel company up to big things. The practice and philosophy of yoga informs our overall purpose to elevate the world through the power of practice. We are proud to be a growing global...Suggested
- ...Position Overview SingleStore is seeking a Site Reliability Engineer to help optimize and scale our managed service offering across all three major cloud providers. In this role, you will be at the intersection of leading technology trends - A highly performant distributed...SuggestedWorldwide
- ...Site Reliability Engineer Join the innovators connecting just about anything—from families to cars to now things—on T-Mobile's biggest and best network yet. The SyncUP Things platform team has an immediate need for a Site Reliability Engineer. Responsibilities:...SuggestedContract workImmediate startRemote work
- ...Site Reliability Engineer (SRE) Location: Seattle, WA (Onsite – 4 days/week) Industry: Quick Service Restaurant (QSR) Employment Type: Contract Rate: DOE Key Responsibilities Manage and enhance the enterprise vulnerability management program using tools such...Contract work
- ...Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA. We offer IT solutions across the disciplines of program/project management, applications development, infrastructure, Cyber security, and enterprise content...
- Technical/Functional Skills Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper-V High Availability Clusters ...
$91.4k - $187k
...infrastructure and/or service according to terms for reliability and functionality. ~Assists team... .... ~Gains basic knowledge of site reliability trends and shares relevant information... ...are seeking a skilled Site Reliability Engineer to design, build, operate, and automate...Temporary workImmediate startFlexible hoursShift work$159.2k - $301.6k
...running Graphs on the cloud. In this reliability-focused role, you will own the availability... .... You'll partner with the backend engineers building these APIs to make sure the system... ...Science. ~5-10 years of experience in site reliability engineering, infrastructure,...Temporary workLocal areaWorldwide- Job Title Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB ...
$142.3k - $263.3k
...Senior Site Reliability Engineer Apple Services Engineering Cloud Service Infrastructure team is one of the most exciting examples of Apple's long-held passion for combining art and technology. Join Apple Services Engineering Cloud Service Infrastructure team, as a...Relocation$160k - $250k
...public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is able...- ...Sr. Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA. We offer IT solutions across the disciplines of program/project management, applications development, infrastructure, Cyber security, and enterprise...Local area
- ...Site Reliability Engineer (SRE) – Backend/Cloud Client – Airbnb/Altimetrik Client remote / Hybrid from Altimetrik office Pay Rate -65 C2C Position: Site Reliability Engineer (SRE) – Backend/Cloud Must-Have Experience: Familiarity with cloud-hosted...Work at officeRemote work
- Job Title: Site Reliability Engineer (SRE) - Middleware API Key Skills: SRE, Middleware API, Scripting and Automation tools Location: Seattle, Washington & Atlanta, Georgia We at Coforge are hiring experienced professionals with strong knowledge of Site Reliability Engineering...
$160k - $210k
...to advertising and deliver unparalleled precision, relevance, and impact at scale. The Role We are looking for a senior site reliability engineer to expand our global datacenter footprint and improve service management across Cognitiv. The primary focus is to rapidly...Work at officeRemote workWork from home- ...certification), ISO 27001:2005 Information Security Management System (ISMS), and CMMI-DEV Level 3. Job Description Sr. Site Reliability Engineer Location - Seattle, WA Duration - 12 months Interview - in-person if local or Phone + Skype Minimum Requirements 30% C#...Local areaWorldwide
- SingleStore is seeking a Site Reliability Engineer to help optimize and scale our managed service offering across all three major cloud providers. In this role, you will be at the intersection of leading technology trends - a highly performant distributed database, managed...
- ...airplane, or remote military base, Ditto’s peer‑to‑peer sync engine ensures devices stay connected and data stays consistent, even... ...the demands of our enterprise customers, we need experienced Site Reliability Engineers to ensure our infrastructure delivers enterprise‑...Remote workFlexible hours
- Optomi, in partnership with an industry-leading technology organization, is seeking a Senior Site Reliability Engineer (SRE) to join their team. This individual will play a key role in driving automation, reducing operational toil, modernizing legacy platforms, and improving...
$120k - $150k
This is an engineering-first Senior SRE role. We’re looking for Senior Engineers Who Have the following: Built and shipped significant... ...services end-to-end in production (design → launch → on-call → reliability improvements) Led incident response and drove durable follow-...$127k - $249k
THE TEAM Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational... ..., alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours$194k - $267k
...something more than once, automate it” and who can rapidly self-educate on new concepts and tools. POSITION OVERVIEW: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$163.62k - $212.71k
...Principal Site Reliability Engineer Bellevue, WA Immigration / Work Authorization Notice: Applicants must be currently authorized to work in the United States. iSpot is not able to sponsor or take over sponsorship of an employment visa for this position at this time...Full timePart timeWork at officeLocal areaImmediate startRemote workWork from homeFlexible hoursShift work3 days per week1 day per week- ...that the residential, construction & building product industries operate across the globe. We are looking for a Manager, Site Reliability Engineering to be part of revolutionizing these industries. What You Will Do Lead and grow a team of site reliability engineers....
$121.5k - $306.4k
...infrastructure and service and provides input on best practices for reliability and functionality. Establishes direction to ensure accurate... ...with new technology, executing improvements, building site reliability knowledge, and providing clear data. #LI-ES2 Responsibilities...Temporary workFlexible hours- ...SRE / DevOps Engineer Seattle based client. Seattle-WA (3 days onsite). U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. We are unable to sponsor currently. Must Have Skills F5 load balancer Chef and Terraform - must have...
$194k - $267k
...defining work. We're all in on this mission. If you are too, let's talk. We are seeking a highly technical Observability Site Reliability Engineer with a specialty in Google Cloud, to own and expand our Observability ecosystem into GCP. In this role, you will move...Permanent employmentLocal areaWorldwideFlexible hours$194k - $267k
...in on this mission. If you are too, let's talk. The Team The Site Reliability team is dedicated to architecting and owning the... ...durable, automated systems that maximize platform reliability and engineering velocity. The ideal candidate is someone who enjoys analyzing...WorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
Related searches
- site reliability engineer sre Seattle, WA
- site reliability engineer Seattle, WA
- site safety Seattle, WA
- historic site Seattle, WA
- IT site lead Seattle, WA
- site leader Seattle, WA
- site recruiter Seattle, WA
- website coordinator Seattle, WA
- website content developer Seattle, WA
- site services specialist Seattle, WA


