Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Staff Site Reliability Operations

$184k - $264.5k

NVIDIA

For over 25 years, NVIDIA has been at the forefront of transforming computer graphics, PC gaming, and accelerated computing, driven by a legacy of continuous innovation and exceptional talent! We are now bringing to bear the immense potential of AI to usher in the next era of computing, where our GPUs power the "brains" of computers, robots, and autonomous vehicles that can comprehend the world. This pioneering work demands vision, innovation, and the world's best talent. Join our diverse and supportive environment, where NVIDIANs are inspired to excel and make a profound global impact. We are seeking a Site Reliability Operations Technical Lead to serve as the senior technical individual contributor for reliability and support at our Seattle, WA site. This role owns technical service delivery locally, acts as the support point of last resort before regional and global platform teams, and provides technical leadership to site support engineers. Be responsible for the hardest issues across Active Directory, Exchange, database platforms, and compute infrastructure, lead the site through major incidents, and drive out the recurring problems that consume the team’s capacity. You will also lead site-level projects, represent local requirements in global initiatives, and set the technical standard the site support team works to. The successful candidate is equally comfortable running a root cause analysis, supporting an executive before an all-hands, and briefing IT leadership on site risk. What you'll be doing: Own day-to-day site operations — incidents, requests, critical issues, and support coverage — with accountability for queue health, SLA attainment, backlog, and service quality, plus site asset and inventory management across lifecycle, refresh, procurement, and compliance. Serve as Tier 3 escalation owner for the site and AMER across identity (AD, hybrid Entra ID, GPO, Kerberos/LDAP, SSO, MFA), messaging (Exchange hybrid mail flow, mailbox, SMTP relay), compute (Windows, Linux, macOS, virtualization, storage, and hands-on datacenter and lab hardware), and endpoint (M365, Teams, Intune, Autopilot, imaging through migrations) driving root cause and permanent fixes rather than repeat break-fix. Own endpoint compliance, vulnerability remediation, patch management, and hardening; audit readiness and evidence; and partnership with InfoSec on incident response and privileged access. Drive critical issues into global platform teams and vendors with reproduction cases and diagnostic evidence through to a committed fix. Act as technical lead for site SRO engineers setting standards, reviewing work, directing blocking issues, building diagnostic rigor through mentorship, and owning the site knowledge base and runbook library. Serve as the primary technical contact for site IT, partnering with employees, site and executive leadership, Facilities, Security, HR, and Procurement on incidents, planned changes, onboarding and moves, and office and lab expansions. Build automation in PowerShell, Python, or Bash for diagnostics, remediation, health checks and reporting; analyze ticket and reliability trends to eliminate top recurring drivers; and champion AI-driven and agentic solutions that advance SRO strategy. Represent site and AMER priorities in regional and global IT initiatives, standards, and architecture forums, and lead operational decisions in the manager's absence. What we need to see: 12+ years in enterprise support engineering, infrastructure, or end user services, including 5+ years in a senior, lead, or escalation-tier role in a multi-site environment. Deep hands-on solving across Active Directory and hybrid Entra ID, Exchange hybrid, Windows and Linux server, virtualization, enterprise storage, and datacenter hardware. Enterprise endpoint management (Intune, Autopilot, MECM/SCCM, Jamf, or equivalent), Windows 11, and the Microsoft 365 ecosystem, plus endpoint security and vulnerability remediation. Database operations support and networking fundamentals — DNS, DHCP, VLAN, wireless, firewall policy, and switch-level troubleshooting. Scripting and automation in Python, PowerShell, or Bash applied to real support problems, and ServiceNow or similar ITSM. Demonstrated technical leadership without formal authority, excellent executive-level communication during incidents, and the rigor to pursue root cause over symptom clearing. Willingness to work on-site and hands-on (including lifting and moving equipment), join an on-call rotation, and support after-hours maintenance windows and cutovers. Bachelor's degree in Computer Science, Information Systems, or related field, or equivalent experience. Ways to Stand Out from the crowd: Experience supporting engineering, lab, R&D, or manufacturing environments with specialized equipment and non-standard availability requirements. Local technical lead through a site buildout, relocation, or major migration; or experience influencing global standards and tooling roadmaps for site and regional needs. Executive support programs, AV and hybrid conference room technologies, or build automation with measurable efficiency and experience benefits. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 264,500 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 15, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr NVIDIA

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Staff Site Reliability Operations in Seattle, WA vacancy
  • NVIDIA is seeking a Site Reliability Operations Technical Lead for our Seattle, WA site. You will own day-to-day site operations, drive incidents to resolution, and provide technical leadership to the site support engineers. The role requires deep hands-on expertise across... 
    Operations
    Senior
    Website

    NVIDIA

    Seattle, WA
    3 days ago
  • NVIDIA is seeking a Senior Staff Site Reliability Operations Technical Lead in Seattle to own local reliability, incident management, and service delivery. You will coordinate across AD/Entra ID, Exchange, Windows, Linux, and data-center hardware, while guiding site engineers... 
    Operations
    Senior
    Website
    Local area

    Nvidia Corporation in

    Seattle, WA
    17 hours ago
  • NVIDIA is seeking a Site Reliability Operations Technical Lead for our Seattle, WA site. You will own and optimize day-to-day site operations, act as Tier 3 escalation for AD, Exchange Hybrid, and compute infrastructure, and drive permanent fixes with global teams. You... 
    Operations
    Senior
    Website
    Permanent employment
    Local area

    NVIDIA Gruppe

    Seattle, WA
    2 days ago
  • $170k - $220k

     ...We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our...  ...coordination. Build safe, repeatable, and observable workflows.GitHub Operations: Manage GitHub branching strategies, pull request flows,... 
    Operations
    Senior
    Website

    Supio

    Seattle, WA
    3 days ago
  • $139k - $242k

     ...-up, hardware RMA, data center operations, and platform services into a cohesive, high-reliability engine of fleet management. This...  .... About the Role:  As a Senior Software Engineer on the Fleet...  ...to problems of scale for multi-site deployment and management of CoreWeave... 
    Operations
    Senior
    Website
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    24 days ago
  • $160k - $210k

     ...industry. Now, we're growing! We are looking for a Senior Site Reliability Engineer to strengthen our AWS infrastructure and improve...  ...expertise is the priority, working familiarity with datacenter operations is important so you can help provide multi-DC coverage... 
    Operations
    Senior
    Website
    Work at office
    Immediate start
    Remote work
    Work from home

    Cognitiv

    Bellevue, WA
    24 days ago
  •  ...team is responsible for the reliability, scalability, and efficiency...  ...maintain system stability.As a Site Reliability Engineer, you will...  .... You will focus on hands-on operational work, from responding to...  ...practices while working alongside senior engineers to solve... 
    Operations
    Senior
    Website

    TikTok

    Seattle, WA
    1 day ago
  • $139k - $242k

     ...-up, hardware RMA, data center operations, and platform services into a cohesive, high-reliability engine of fleet management. This...  .... About the Role:  As a Senior Software Engineer on the Fleet...  ...to problems of scale for multi-site deployment and management of CoreWeave... 
    Operations
    Senior
    Website
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    8 days ago
  • $134.25k - $214.8k

     ...that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on...  ...tools that power reliable, scalable, and secure engineering operations across the company. You will:Build robust, easy-to-use... 
    Operations
    Senior
    Website
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Seattle, WA
    2 days ago
  •  ...Sr. Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA. We offer IT solutions...  ...(ITIL) v.3 Framework across enterprise infrastructure operations. These methodologies and processes are reinforced through our... 
    Operations
    Senior
    Website
    Local area

    Comtech LLC

    Seattle, WA
    3 days ago
  • $232k - $319k

     ...real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.This is an opportunity...  ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... 
    Operations
    Senior
    Website
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    2 days ago
  •  ...Seattle is seeking a Fleet Manager to lead maintenance operations, ensuring safety, reliability, and regulatory compliance for the regional fleet. You will...  ...strong leadership, and proficiency with CMMS tools. On-site work in Seattle is expected, with opportunities to... 
    Operations
    Senior
    Website

    LSG Lufthansa Service Holding AG

    Seattle, WA
    2 days ago
  •  ...Dynamics Information Technology is seeking an Operations and Maintenance Lead to oversee enterprise IT operations...  ...exceptional stakeholder engagement to ensure reliable, secure, and modern infrastructure across distributed sites. The ideal candidate offers 10+ years in... 
    Operations
    Senior
    Website

    General Dynamics Information Technology

    Seattle, WA
    4 days ago
  • $55k - $151.47k

     ...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the...  ...robust, secure IT systems that support business operations. You will work to optimize the performance of networks,... 
    Operations
    Senior
    Website
    Full time
    H1b

    PwC

    Seattle, WA
    3 days ago
  •  ...Senior Automation EngineerLocation: Seattle, WA Duration: Contract Visa: GC or Citizen only or H1 where there...  ...Basic Qualifications:12+ years experience working in Operations, Engineering, DevOps, or Site Reliability in a medium to large company.Can describe specific tools... 
    Operations
    Senior
    Website
    Contract work

    Georgia IT Inc

    Seattle, WA
    5 days ago
  • Tiktok in Seattle is seeking a Senior Systems Administrator to design, implement, and support complex server and network...  ...SharePoint), collaborating with clients and vendors to deliver reliable IT services. This on-site role includes on-call rotation and a strong emphasis on... 
    Senior
    Website

    Tiktok

    Seattle, WA
    1 day ago
  • Nscale is seeking a Senior Infrastructure Support Engineer to own...  ...hands-on L2/L3 role. You will operate across GPU hardware, Linux,...  ...Ops, and Engineering to ensure reliability. Responsibilities include...  ...engineers. On-call travel to sites may be required. #J-18808-Ljbffr... 
    Operations
    Senior
    Website
    Remote job

    Nscale

    Seattle, WA
    1 day ago
  •  ...data and AI infrastructure provider is seeking a Senior Staff Technical Program Manager for Reliability to enhance the reliability and performance of their...  ...strategies and drive program execution to ensure operational excellence and support mission-critical workloads.... 
    Senior

    Menlo Ventures

    Bellevue, WA
    4 days ago
  •  ...across electrical, mechanical, and HVAC disciplines to support reliable facility operations. Join a team dedicated to patient-centered care and safety,...  ...health system and a comprehensive benefits package. On-site work at Swedish First Hill with competitive wage range and... 
    Operations
    Senior
    Website
    Day shift

    Providence Health & Services

    Seattle, WA
    17 hours ago
  • $180.2k - $243.34k

     ...data and AI infrastructure platform is seeking a Senior Staff Technical Program Manager for Reliability in Seattle. This high-visibility role will lead critical...  ..., focusing on enhancing the reliability and operational excellence of their multi-cloud infrastructure. Applicants... 
    Senior

    Databricks

    Seattle, WA
    17 hours ago
  • $250k - $350k

    Artera is hiring a Senior Staff AI Builder to help define what comes...  ...to planning, evaluation, and reliability across multiple squads, graduating...  ...’s roadmap. You’ll own the operational plans that shape our short‑...  ...be moving to full‑time on‑site (5 days/week) by September 1... 
    Senior
    Website
    Full time
    Temporary work
    Summer work
    Summer holiday
    Work at office
    Remote work
    Relocation
    Relocation package
    Flexible hours
    Shift work
    3 days per week

    Artera

    Seattle, WA
    3 days ago
  • Senior Cloud Engineer (Azure) Location: Seattle, WA Visa: GC or Citizen Or H1B Note: Its a Senior position need...  ...as Azure or AWS) 8+ years of experience working in Operations, Engineering, DevOps, or Site Reliability in a medium to large company. Strong experience with... 
    Operations
    Senior
    Website
    H1b

    Staffing the Universe

    Seattle, WA
    2 days ago
  •  ...support line cooks, and ensure sanitation and safety standards are met in daily tasks. The role emphasizes reliability, teamwork, and adherence to food safety rules, with on-site duties and opportunities to grow within a busy dining operation. #J-18808-Ljbffr Workstream
    Operations
    Website

    Workstream

    Bellevue, WA
    3 days ago
  • Kong Inc. is seeking a Senior Site Reliability Engineer for Managed Gateways in Washington. You will own production reliability for a cloud-native platform spanning AWS, GCP, and Azure, leading a high-performing SRE team and shaping the enterprise deployment experience.... 
    Senior
    Website

    Cacheflow

    Seattle, WA
    1 day ago
  • $128k - $135k

     ...Job Description Sevan Multi-Site Solutions is a veteran-owned business...  ...Medallion Award. The Senior Construction Manager acts as...  ...’s design and construction staff, overseeing the work of general...  ...Support the Project Executive or Operations Director in developing project... 
    Operations
    Senior
    Website
    Full time
    Contract work
    For contractors
    Work experience placement
    For subcontractor
    Live in
    Work at office
    Remote work
    Flexible hours
    Shift work

    Sevan Multi-Site Solutions, Inc.

    Seattle, WA
    6 days ago
  •  ...Administrator to manage provisioning, installation, configuration, operation, and maintenance of IT infrastructure and related software and...  ...across servers, networks, and applications with emphasis on reliability and incident response in a fast-paced setting. #J-18808-... 
    Operations
    Senior

    Cisco Systems

    Seattle, WA
    1 day ago
  •  ...Senior Systems Engineer Hybrid About the Role This is a hands...  ...plane that connect customer sites to Cloudflare. As a Senior...  ...partial failures, and safe to operate across a large fleet carrying...  ...appliance dataplane. Build secure, reliable networking capabilities using... 
    Operations
    Senior
    Website
    Local area
    Remote work

    Cloudflare Inc

    Seattle, WA
    3 days ago
  • $86.4k - $199.5k

     ...seeking a skilled professional in Seattle, WA, to manage Oracle database environments. The candidate will collaborate with the Site Reliability Engineering team to enhance service architecture, ensuring performance and security. Requirements include U.S. citizenship... 
    Senior
    Website
    Shift work
    Weekend work

    Oracle

    Seattle, WA
    3 days ago
  • $21.65 - $23 per hour

     ...great benefits, but we also offer reliable work opportunities and assist...  ...Paid sick leave Kitchen Staff Duties: Support the kitchen manager in daily operations Operate standard kitchen...  ...lift items over 30 lbs Work Site: Address: 120 Andover Park... 
    Operations
    Website
    Hourly pay
    Full time
    Local area
    Shift work

    Dough Zone

    Seattle, WA
    3 days ago
  • $171k - $231.4k

    AWS operates the world's largest fleet of GPU-accelerated...  ....We are seeking a Senior Technical Program Manager...  ...decisions reflected in fleet reliability metrics within weeks of...  ...Manufacturing Partner sites.A day in the lifeYou...  ..., supervisors, and staff; adhere to standards of... 
    Operations
    Senior
    Website
    Interim role
    Local area
    Worldwide
    Flexible hours
    Day shift

    Amazon

    Seattle, WA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Staff Site Reliability Operations. Be the first to apply!