Senior Staff Site Reliability Operations
$184k - $264.5kNVIDIA
For over 25 years, NVIDIA has been at the forefront of transforming computer graphics, PC gaming, and accelerated computing, driven by a legacy of continuous innovation and exceptional talent! We are now bringing to bear the immense potential of AI to usher in the next era of computing, where our GPUs power the "brains" of computers, robots, and autonomous vehicles that can comprehend the world. This pioneering work demands vision, innovation, and the world's best talent. Join our diverse and supportive environment, whereNVIDIANs are inspired to excel and make a profound global impact. We are seeking a Site Reliability Operations Technical Lead to serve as the senior technical individual contributor for reliability and support at our Seattle, WA site. This role owns technical service delivery locally, acts as the support point of last resort before regional and global platform teams, and provides technical leadership to site support engineers. Be responsible for the hardest issues across Active Directory, Exchange, database platforms, and compute infrastructure, lead the site through major incidents, and drive out the recurring problems that consume the team’s capacity. You will also lead site-level projects, represent local requirements in global initiatives, and set the technical standard the site support team works to. The successful candidate is equally comfortable running a root cause analysis, supporting an executive before an all-hands, and briefing IT leadership on site risk.What you'll be doing:Own day-to-day site operations — incidents, requests, critical issues, and support coverage — with accountability for queue health, SLA attainment, backlog, and service quality, plus site asset and inventory management across lifecycle, refresh, procurement, and compliance.Serve as Tier 3 escalation owner for the site and AMER across identity (AD, hybrid Entra ID, GPO, Kerberos/LDAP, SSO, MFA), messaging (Exchange hybrid mail flow, mailbox, SMTP relay), compute (Windows, Linux, macOS, virtualization, storage, and hands-on datacenter and lab hardware), and endpoint (M365, Teams, Intune, Autopilot, imaging through migrations) driving root cause and permanent fixes rather than repeat break-fix.Own endpoint compliance, vulnerability remediation, patch management, and hardening; audit readiness and evidence; and partnership with InfoSec on incident response and privileged access.Drive critical issues into global platform teams and vendors with reproduction cases and diagnostic evidence through to a committed fix.Act as technical lead for site SRO engineers setting standards, reviewing work, directing blocking issues, building diagnostic rigor through mentorship, and owning the site knowledge base and runbook library.Serve as the primary technical contact for site IT, partnering with employees, site and executive leadership, Facilities, Security, HR, and Procurement on incidents, planned changes, onboarding and moves, and office and lab expansions.Build automation in PowerShell, Python, or Bash for diagnostics, remediation, health checks, and reporting; analyze ticket and reliability trends to eliminate top recurring drivers; and champion AI-driven and agentic solutions that advance SRO strategy.Represent site and AMER priorities in regional and global IT initiatives, standards, and architecture forums, and lead operational decisions in the manager's absence.What we need to see:12+ years in enterprise support engineering, infrastructure, or end user services, including 5+ years in a senior, lead, or escalation-tier role in a multi-site environment.Deep hands-on solving across Active Directory and hybrid Entra ID, Exchange hybrid, Windows and Linux server, virtualization, enterprise storage, and datacenter hardware.Enterprise endpoint management (Intune, Autopilot, MECM/SCCM, Jamf, or equivalent), Windows 11, and the Microsoft 365 ecosystem, plus endpoint security and vulnerability remediation.Database operations support and networking fundamentals — DNS, DHCP, VLAN, wireless, firewall policy, and switch-level troubleshooting.Scripting and automation in Python, PowerShell, or Bash applied to real support problems, and ServiceNow or similar ITSM.Demonstrated technical leadership without formal authority, excellent executive-level communication during incidents, and the rigor to pursue root cause over symptom clearing.Willingness to work on-site and hands-on (including lifting and moving equipment), join an on-call rotation, and support after-hours maintenance windows and cutovers.Bachelor's degree in Computer Science, Information Systems, or related field, or equivalent experience.Ways to Stand Out from the crowd:Experience supporting engineering, lab, R&D, or manufacturing environments with specialized equipment and non-standard availability requirements.Local technical lead through a site buildout, relocation, or major migration; or experience influencing global standards and tooling roadmaps for site and regional needs.Executive support programs, AV and hybrid conference room technologies, or build automation with measurable efficiency and experience benefits.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 264,500 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 15, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, WA, SeattleType: Full time
- ...depend on every day. We're hiring a senior, hands-on engineer to own the reliability, availability, security, and... ...the rest of engineering builds and operates reliable services through your work... ...of hands-on Cloud Operations and Site Reliability Engineering, operating...OperationsSeniorWebsiteFull time
- ...team is responsible for the reliability, scalability, and efficiency... ...maintain system stability.As a Site Reliability Engineer, you will... .... You will focus on hands-on operational work, from responding to... ...practices while working alongside senior engineers to solve...OperationsSeniorWebsite
- Salesforce seeks a Site Reliability Engineer for Missionforce and related teams to design, implement, and operate highly available, scalable backend systems in public-cloud environments. You will partner with product engineers and forward-deployed teams to deliver reliable...OperationsSeniorWebsite
$232k - $319k
...real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.This is an opportunity... ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and...OperationsSeniorWebsitePermanent employmentLocal areaWorldwideFlexible hours$55k - $151.47k
...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the... ...robust, secure IT systems that support business operations. You will work to optimize the performance of networks,...OperationsSeniorWebsiteFull timeH1b$160k - $250k
...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS... ...same task twice. Responsibilities Automate manual operational processes Improve workflows of developer, data, and machine...OperationsSeniorWebsite- ...resident relations. The role collaborates with vendors and on-site staff to meet residential needs and maintain a welcoming, well-run community... ..., manage vendor relationships, and support events and daily operations with a focus on service excellence and solutions for residents...OperationsSeniorWebsiteWorldwide
$230k - $280k
...from the ground up, we own and operate each layer of the stack — from... ...toward. This is the most senior individual contributor role in... ...technical authority the exec staff, the board, and Crusoe's capital... ...versus partner, regional siting, technology transition timing,...OperationsSeniorWebsiteTemporary work- ...This is an engineering-first Senior SRE role. We’re looking for senior engineers who... ...production (design → launch → on-call → reliability improvements) Led incident response and... ...the tooling and platforms that make operating services safer and easier for every engineer...OperationsSeniorWebsite
$13 per hour
...opportunities in modern enterprise. Agentforce Operations is reimagining the supply chain with an AI-... ...Operations & Missionforce Team As a Site Reliability Engineer for Missionforce Operations, you will be a senior technical contributor who helps shape the infrastructure...OperationsSeniorWebsite$173.9k - $235.2k
...organizations to design scalable, reliable systems for our accelerated... ...hardware, software, and operations teams.Key job responsibilitiesFleet... ...and Manufacturing Partner sites.A day in the lifeYou start... ...employees, supervisors, and staff; adhere to standards of excellence...OperationsSeniorWebsitePermanent employmentWork experience placementInternshipLocal areaWorldwideFlexible hoursNight shiftDay shift$134.25k - $214.8k
...that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on... ...tools that power reliable, scalable, and secure engineering operations across the company. You will:Build robust, easy-to-use...OperationsSeniorWebsiteWork experience placementWork at officeRemote workFlexible hours$134.25k - $214.8k
...that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on... ...tools that power reliable, scalable, and secure engineering operations across the company. You will:Build robust, easy-to-use...OperationsSeniorWebsiteWork experience placementWork at officeRemote workFlexible hours- ...Library (ITIL) v.3 Framework across enterprise infrastructure operations. These methodologies and processes are reinforced through... ...System (ISMS), and CMMI-DEV Level 3. Job Description Sr. Site Reliability Engineer Location – Seattle, WA Duration – 12 months...OperationsSeniorWebsiteLocal areaWorldwide
- ...Sr. Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA. We offer IT solutions... ...(ITIL) v.3 Framework across enterprise infrastructure operations. These methodologies and processes are reinforced through our...OperationsSeniorWebsiteLocal area
- ...billion users. The core goals of the team are to ensure high system reliability, uninterrupted service, and smooth data processing. We are... ...problems that occur in the Paimon-Flink architecture during operation, design and implement necessary mechanisms and tools, such as...OperationsSeniorFlexible hours
- ...Senior Cloud Engineer (Azure) Location: Seattle, WA Visa: GC or Citizen Or H1B Note: Its a Senior position... ...Azure or AWS) ~8+ years of experience working in Operations, Engineering, DevOps, or Site Reliability in a medium to large company. ~ Strong experience...OperationsSeniorWebsiteH1b
$160k - $200k
...industry. Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve... ...expertise is the priority, working familiarity with datacenter operations is important so you can help provide multi-DC coverage...OperationsSeniorWebsiteWork at officeImmediate startRemote workWork from home- ...Senior Automation Engineer Location: Seattle, WA Visa: GC or Citizen Or H1B Note: Its a Senior position need... ...Qualifications: ~10+ years experience working in Operations, Engineering, DevOps, or Site Reliability in a medium to large company. ~ Can describe...OperationsSeniorWebsiteH1b
$200k - $287.5k
...future of how work gets done.Senior Software Engineer, Capacity EngineeringAt... ...essential for Snowflake's operations and ongoing growth. Capacity... ..., prefer to write scalable, reliable, and testable software, are... ...on the Snowflake Careers Site for salary and benefits...OperationsSeniorWebsite$168k - $200k
..., providing secure, simple, and reliable ways to manage their money, ensuring... ...Role:Remitly is looking for a Senior Enterprise Risk Management... ...sanctions and adverse-media screening operations across partner network.Participate in on-site engagements with key...OperationsSeniorWebsiteFull timeWork at officeWorldwideFlexible hours$183k - $247.6k
AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and... ...to Design and Manufacturing Partner sites.A day in the lifeYou start the day reviewing... ...cooperatively with other employees, supervisors, and staff; adhere to standards of excellence...OperationsSeniorWebsiteWork at officeLocal areaWorldwideFlexible hoursDay shift- ...We are seeking a hands-on Senior Embedded Firmware Engineer to... ...implementation, testing, integration and operational support. This is an... .... Implement efficient, reliable movement of data between... ...infrastructure. Willingness to work on‑site in the Seattle/Redmond area....OperationsSeniorWebsite
$152k - $241.5k
...will design and develop high-throughput, reliable telemetry pipelines and modern data... ...production-grade coding, and a passion for operational excellence.What You Will Be Doing:Design... ...platform engineering, infrastructure, and site reliability teams to deliver production-...OperationsSeniorWebsiteFull time$140k - $165k
...Our platform is engineered for reliability and scale and harnesses the... ...Zenoti visit: Customer Success Senior Manager (SaaS Customer... ...technical account managers, and operations teams to engage customers in... ...to travel as needed to be on-site with key customers or attend...OperationsSeniorWebsiteLocal area- ...and implementation, maintains application reliability by working to identify systemic issues... ...implementation and deployment documentation for operations and internal customers• Mange project... ...• Work as part of a team• Work with on-site equipmentAdditional Responsibilities•...OperationsSeniorWebsiteFull timeWork experience placementWork at office
$191k - $297k
..., and we're looking for a senior leader to own the operational engine that keeps it running... ...end to end — the health, reliability, security posture, and... ...offshore, and contingent staff, and you're accountable for... ...of the Nordstrom Careers site. Applicants with disabilities...OperationsSeniorWebsiteFull timeContract workWork at office- ...data and AI infrastructure provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their... ...strategies and drive program execution to ensure operational excellence and support mission-critical workloads....Senior
$151.3k
We are seeking a Senior Software Development Engineer to join our... ...builder who designs, develops, and operates the platform infrastructure... ...operational excellence and reliability - Build automated deployment... ...should apply via our internal or external career site.OperationsSeniorWebsiteCasual workInternshipLocal area- ...across electrical, mechanical, and HVAC disciplines to support reliable facility operations. Join a team dedicated to patient-centered care and safety,... ...health system and a comprehensive benefits package. On-site work at Swedish First Hill with competitive wage range and...OperationsSeniorWebsiteDay shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Staff Site Reliability Operations. Be the first to apply!
- senior service associate Seattle, WA
- senior safety specialist Seattle, WA
- senior vice president of business development Seattle, WA
- senior service designer Seattle, WA
- senior sales recruiter Seattle, WA
- senior mulesoft developer Seattle, WA
- senior media manager Seattle, WA
- senior business manager Seattle, WA
- senior linux systems engineer Seattle, WA
- senior mainframe developer Seattle, WA




