Site Reliability Engineer
$100k - $170kNscale
About Nscale Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native
startups and global enterprises, from bare metal up through the platform services teams actually build
on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each
other the truth, and everyone here stays close to the infrastructure that makes AI work. The Role
This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take
real surface area: the automation and tooling other engineers depend on, and the reliability of
production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll
be expected to make the systems you touch quieter over time. What You'll Do
• Build and own the automation and tooling that keeps the platform running; treat operational toil as
a bug to be fixed, not a fact of life.
• Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance.
• Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-
incident reviews that actually change the system.
• Investigate performance and reliability problems across Linux, networking, and distributed services,
then fix them at the source.
• Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the
stack.
• Improve availability, scalability, and efficiency through code, not manual effort.
What You'll Bring
• 3-6 years in SRE, systems engineering, or software engineering, including time running production in
a data center or cloud environment.
• Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work
away.
• Solid command of Linux, networking fundamentals, and distributed systems.
• A track record of troubleshooting live production issues and owning the fix through to the retro.
• Fluency with monitoring and observability; metrics, logs, dashboards, and alerting.
• Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be
asked. Nice to Have
• Experience with AI or GPU workloads, or high-performance computing (HPC).
• Familiarity with high-performance networking (InfiniBand, RDMA).
• Kubernetes, plus virtualized or bare-metal environments.
On-Call and Pace
A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation,
and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth
acting on rather than just an interruption. The goal is to make the systems quieter over time, so each
rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in
better shape than you found it, you'll do well here.
What We Offer
• Competitive base plus equity, reviewed every 12 months.
• Real scope early, and a progression plan built around the skills you want to sharpen.
• Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape
your day. Salary Range
$100,000 - $170,000 USD. Actual compensation varies with skill set, experience, and location, and the
role may be eligible for bonus and equity. Equal Opportunities Statement At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know. The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $100,000-$170,000 USD For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
startups and global enterprises, from bare metal up through the platform services teams actually build
on. Our culture runs on ownership, accountability, and speed. We move with urgency, we tell each
other the truth, and everyone here stays close to the infrastructure that makes AI work. The Role
This is a career-level SRE role for someone who wants to own systems, not just watch them. You'll take
real surface area: the automation and tooling other engineers depend on, and the reliability of
production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll
be expected to make the systems you touch quieter over time. What You'll Do
• Build and own the automation and tooling that keeps the platform running; treat operational toil as
a bug to be fixed, not a fact of life.
• Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance.
• Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-
incident reviews that actually change the system.
• Investigate performance and reliability problems across Linux, networking, and distributed services,
then fix them at the source.
• Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the
stack.
• Improve availability, scalability, and efficiency through code, not manual effort.
What You'll Bring
• 3-6 years in SRE, systems engineering, or software engineering, including time running production in
a data center or cloud environment.
• Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work
away.
• Solid command of Linux, networking fundamentals, and distributed systems.
• A track record of troubleshooting live production issues and owning the fix through to the retro.
• Fluency with monitoring and observability; metrics, logs, dashboards, and alerting.
• Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be
asked. Nice to Have
• Experience with AI or GPU workloads, or high-performance computing (HPC).
• Familiarity with high-performance networking (InfiniBand, RDMA).
• Kubernetes, plus virtualized or bare-metal environments.
On-Call and Pace
A quick note on the shape of the job. This role sits close to production, so there is an on-call rotation,
and some weeks are busier than others. We share it fairly, and we treat every page as a signal worth
acting on rather than just an interruption. The goal is to make the systems quieter over time, so each
rotation asks less of the person carrying it. If you take ownership of what you run and like leaving it in
better shape than you found it, you'll do well here.
What We Offer
• Competitive base plus equity, reviewed every 12 months.
• Real scope early, and a progression plan built around the skills you want to sharpen.
• Flexibility that treats you as an adult; we care that the work gets done, and we trust you to shape
your day. Salary Range
$100,000 - $170,000 USD. Actual compensation varies with skill set, experience, and location, and the
role may be eligible for bonus and equity. Equal Opportunities Statement At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enrich our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know. The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $100,000-$170,000 USD For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
Vacancy posted 15 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Seattle, WA vacancy
- Company DescriptionComtech LLC is a woman-owned small business focused on delivering end-to-end solutions and products. Since 1998, we have successfully serviced enterprises across the public and private sectors, and the Department of Defense. Our services span all aspects...Suggested
$170k - $220k
Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating...Suggested$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance...SuggestedWork experience placementWork at officeRemote workFlexible hours$143k - $191k
...mission critical capabilities to our customers. System Deployment Engineers work in complex environments with shared environmental... ...customersParticipate in customer demonstrations and exercisesWork with site reliability engineers to provide and refine requirements for tooling and...SuggestedFull timeTemporary workWork experience placementImmediate start$134.25k - $214.8k
...upload. Every piece of digital evidence. Every chain of custody log that holds up in court. That's us.Axon's Platform team is the engine behind what hundreds of thousands of officers rely on every day. We're one of the world's largest blob storage customers, ingesting...SuggestedWork experience placementWork at officeRemote work$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,...Permanent employmentLocal areaWorldwideFlexible hours$166k - $258k
...Seattle office a minimum of 4 days/week in order to be considered for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could be you.You'll help build the scalable, reliable, and...Full timeWork at office$160k - $250k
...public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is able...- ...provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their multi-... ...role involves leading programs in partnership with senior engineering leaders, requiring over 10 years of experience in cloud infrastructure...
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours- ...A leading tech company is seeking a Site Reliability Engineer in Seattle to ensure seamless operation of physical infrastructure. Responsibilities include automation solutions, system monitoring, and collaboration with engineering teams. The ideal candidate has a degree...
$198.36k - $416.1k
...provides excellent experiences for billions of users around the world. Provide technical leadership and mentorship to a team of Site Reliability Engineers focused on building observable, fault-tolerant video services with high availability Drive architectural decisions for...Temporary workShift work- ...Lululemon Site Reliability Engineering Engineer We are a yoga-inspired technical apparel company up to big things. The practice and philosophy of yoga informs our overall purpose to elevate the world through the power of practice. We are proud to be a growing global...
- Job Title Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB ...
$210.6k - $305.1k
...Minimum Qualifications: You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...Full timeTemporary workLocal areaFlexible hours$194k - $267k
...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$163.62k - $212.71k
...maintaining the tools, platforms, and processes that improve our engineering teams' productivity and streamline the software development... ...:We are seeking a seasoned and strategic Lead/Principal Site Reliability Engineer to drive the reliability, scalability, and performance...Full timePart timeWork experience placementWork at officeLocal areaImmediate startRemote workWork from homeFlexible hoursShift work3 days per week1 day per week- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III - DevOps Engineer at JPMorgan Chase within the Commercial and Investment Bank, you will solve complex and broad business...
$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$194k - $267k
...in on this mission. If you are too, let's talk.The TeamThe Site Reliability team is dedicated to architecting and owning the foundational... ...durable, automated systems that maximize platform reliability and engineering velocity.The ideal candidate is someone who enjoys analyzing...Local areaWorldwideFlexible hours- ...future together. We are responsible for the reliability of all the company's major data warehouse products, services, and query engines. We serve business needs across domains... ...practices, and emerging technologies related to site reliability and infrastructure engineering....
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...Local areaRemote workWorldwideFlexible hours$113.3k - $205.52k
...important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering...Work at officeRemote workWorldwideFlexible hours$160k - $210k
...departmental collaboration and a unified sense of purpose, making teamwork a cornerstone of our success. We are looking for a Senior Site Reliability engineer to work on expanding our global footprint of datacenters and improve service management across Cognitiv. Our immediate...Work at officeLocal areaImmediate startRemote work$166k - $244k
# Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them...Temporary work- Overview Site Reliability Engineer, Compute - USDS TikTok is the leading destination for short-form mobile video. U.S. Data Security (USDS) is a subsidiary of TikTok in the U.S. This security-first division was created to bring heightened focus and governance to data protection...Work experience placement
$129.96k - $246.24k
Site Reliability Engineer, Product - USDS 3 weeks ago Be among the first 25 applicants Responsibilities About the team: The Product Engineering team monitors and maintains the availability of TikTok, including services such as video playback, content discovery/recommendations...Full timeTemporary workWork at officeLocal area3 days per week$129.96k - $246.24k
Overview Site Reliability Engineer, Edge Services - USDS Base pay range: $129,960.00/yr - $246,240.00/yr Responsibilities Architect and implement solutions that enable internal and external customers to harness the power of TikTok’s content delivery network. Contribute...$198.36k - $416.1k
...provides excellent experiences for billions of users around the world. Provide technical leadership and mentorship to a team of Site Reliability Engineers focused on building observable, fault-tolerant video services with high availability Drive architectural decisions for...Temporary workShift work- ...Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core... ...not about building features, but about engineering the resilience and performance of the... ...to maintain system stability.As a Site Reliability Engineer, you will be on the...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
Related searches
- site reliability engineer Seattle, WA
- site reliability engineer remote Seattle, WA
- site reliability engineer sre Seattle, WA
- website content developer Seattle, WA
- site leader Seattle, WA
- on-site clinical research associate (traveling/remote) Seattle, WA
- on site coordinator Seattle, WA
- official site Seattle, WA
- site recruiter Seattle, WA
- historic site Seattle, WA

