Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer - Hardware Infrastructure

$184k - $287.5k

NVIDIA

At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation with consistent reliability and uptime. As an SRE here, you will be part of a welcoming team that values collaboration and creativity, empowering developers to make significant updates while sustaining efficient system function.What you'll be doing:Develop and support guidelines for incident management, planned maintenance, and blameless postmortems.Assist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions.Define reliability and supportability metrics, Service Level Objectives, and error budgets.Develop and drive the adoption of actionable, customer-centric monitoring and alerting.Apply automation and Generative AI/Agentic solutions to minimize manual and tedious activities and boost customer support.Guide teams on establishing sustainable on-call and operational standards.What we need to see:Degree in Computer Science or a related technical field involving coding, or equivalent experience.8+ years of experience in SRE, DevOps, or Production Engineering.Strong understanding of SRE principles, including incident management, error budgets, SLOs, and SLAs.Experience crafting and deploying systems that are fault-tolerant, performant, and supportable.Background with infrastructure automation.Experience running critical services in production.Experience in one or more of the following: Python, Go, Perl, or Ruby.Hands-on experience with observability platforms (e.g., Prometheus, Grafana).Strong communication skills with the ability to convey technical concepts effectively to diverse audiences.Flexibility and adaptability working in a fast-paced environment with evolving requirements.Ways to stand out from the crowd:Expertise in establishing incident management and postmortem processes.Experience driving adoption of common tools and processes across diverse groups.Experience working with LLM/Generative AI/Agentic solutions to shorten mitigation time, lessen toil, and ensure Service Level Objectives are met.Hands-on expertise operating and scaling distributed systems with tight SLAs, ensuring high availability and performance.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer - Hardware Infrastructure in Santa Clara, CA vacancy
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global...  ...like Windows, Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra Processors... 
    Suggested
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $132k - $190k

     ...requirements, define architecture, execute hardware design, and product validation.Lead the...  ...:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a related...  ...deployed in the data center.Our Platforms Infrastructure Engineering team designs and builds... 
    Suggested
    Worldwide

    Google

    Sunnyvale, CA
    4 days ago
  • $184k - $287.5k

     ...push the boundaries of innovation and engineering? At NVIDIA, we lead the world in accelerated...  ...high‑performance systems.As a Senior Hardware Systems Engineer, you will help build...  ...Familiarity with hyperscale data center infrastructure, including cooling methods, facility... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $132k - $190k

    Execute functional validation planning and participate in hardware design reviews for key sub-modules and interfaces to ensure specification...  ...cycle.Minimum qualifications:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a related field, or equivalent... 
    Suggested
    Worldwide

    Google

    San Jose, CA
    5 days ago
  • $255k - $340k

     ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of...  ...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building...  ...with the quality team and fleet reliability team during hardware NPI and after production... 
    Suggested
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $172k - $246k

     ...tomorrow’s standard —from breakthrough hardware and battery systems to intuitive design,...  ...services. We are seeking a Staff Hardware Engineer to define and drive the architecture,...  ...tradeoffs across performance, power, cost, reliability, and scalability.Establish design... 
    Full time
    Local area
    Work from home
    Relocation
    Relocation package
    Flexible hours
    3 days per week

    General Motors

    Sunnyvale, CA
    4 days ago
  •  ...for an Open Source Software Engineer to build and optimize machine...  ...that integrate with AMD hardware. You will work at the intersection...  ...testing and validation infrastructure.Participate in code reviews...  ...advanced AI workloads fast, reliable, and broadly accessible on AMD... 

    AMD

    San Jose, CA
    4 days ago
  • $188k - $274k

    Lead a system and hardware design on data center hardware products.Bring up systems and execute engineering validation in the lab.Gather requirements, define architecture,...  ...world's largest and most effective computing infrastructure. You will see those systems from concepts... 

    Google

    Sunnyvale, CA
    2 days ago
  • $159k - $230k

     ...facilitate system performance, design reliability and failure reproduction.Design network, power and cooling infrastructure for end-to-end system testing.Drive...  ....Contribute within a team of hardware/software designers, test engineers, on project planning within hardware... 

    Google

    Sunnyvale, CA
    4 days ago
  • Dawar Consulting is hiring a System Engineer specializing in Hardware Support in Santa Clara, CA. This long-term contract role involves troubleshooting and supporting next-generation sequencing systems and sample prep platforms. The ideal candidate should have a B.S. in... 
    Long term contract

    Dawar Consulting

    Santa Clara, CA
    2 days ago
  • $159k - $230k

     ...issues.Support design engineers with debug, component...  ...thoroughly down to a hardware interface or...  ...sustaining efforts.The AI and Infrastructure team is redefining...  ...unparalleled scale, efficiency, reliability and velocity. Our...  ...must be performed on site.Bachelor’s degree in... 
    Worldwide

    Google

    Sunnyvale, CA
    1 hour ago
  • $166k - $244k

    Overview Site Reliability Engineering (SRE) combines software and systems engineering to build and run...  ...existing systems, building infrastructure and eliminating work through automation...  ...IP, routing, network topologies and hardware, SDN). 2 years of experience leading... 
    Full time

    Google

    Sunnyvale, CA
    5 days ago
  • $147k - $210k

     ...development code.Review code developed by other engineers and provide feedback to ensure best...  ...sources of issues and the impact on hardware, network, or service operations and...  ...troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when... 

    Google

    Sunnyvale, CA
    5 days ago
  • $267k - $356k

     ...a leader in AI cloud infrastructure serving tens of thousands...  ....Lambda's Storage Engineering team is the backbone...  ...industry, which means reliability and performance aren'...  ..., capacity, and hardware failures.Investigate...  ...across new and existing sites using tools such as Ansible... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    7 days ago
  • $100k

     ...team and looking for contributors of all seniorities. We are seeking a junior-to-mid level SOC Emulation Engineer to support our hardware emulation infrastructure and internal chip design teams. This role focuses on integrating vendor and custom hardware transactors, developing... 
    Permanent employment
    Full time

    Tenstorrent

    Santa Clara, CA
    3 days ago
  • $157.3k - $212.8k

    Amazon Web Services (AWS) Hardware Engineering Services (HWEngS) owns the new product development...  ...NPI) and operation of all AWS global infrastructure. In other words, we’re the people who...  ...as total cost of ownership, quality, reliability, performance, and serviceability. You... 
    Local area
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    2 days ago
  • $122.44k - $232.19k

     ...will be joining the Intel Government Technologies Customer Engineering team as a Hardware Platform Applications Engineer (PAE). This is an exciting...  ...support for customer developed systems including providing on-site power on support Participating in the defining and... 
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    5 hours ago
  • $147k - $211k

     ...platforms.Inform direction for research where engineering gaps are identified that merit improved...  ...Science.2 years of experience with hardware design, and data structures or...  ...the architecture built by the Technical Infrastructure team to keep it running. From developing... 

    Google

    Sunnyvale, CA
    3 days ago
  • $136k - $218.5k

     ...dedicated and motivated Software developer with particular interest in algorithms and RTL Design. Understanding both Software and Hardware principles will be a key requirement for this role.What you'll be doing:Architect, design, develop and support tools for RTL generation... 
    Full time

    Nvidia

    Santa Clara, CA
    5 hours ago
  • $138k - $197k

     ...Bachelor's degree in Electrical Engineering, Computer Engineering,...  ...shape the future of AI/ML hardware acceleration. You will have...  ...methodologies and flows.The AI and Infrastructure team is redefining what’s...  ...scale, efficiency, reliability and velocity. Our customers... 
    Worldwide

    Google

    Sunnyvale, CA
    5 hours ago
  • $131k - $175k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...and cross-functional teams, including hardware, software, thermal, and manufacturing engineers...  ...design and deployment of AI and cloud infrastructure—translating cluster architectures into... 
    Remote work
    Flexible hours

    Arista Networks

    Santa Clara, CA
    2 days ago
  • $138k - $197k

     ...and upgrade our emulation infrastructure and act as a primary interface...  ...team members in debug of hardware, tooling, and project specific...  ...'s degree in Electrical Engineering, Computer Engineering, Computer...  ...scale, efficiency, reliability and velocity. Our customers... 
    Worldwide

    Google

    Sunnyvale, CA
    4 days ago
  • $146.7k - $339.3k

     ...available for this positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid...  ...teams on architecture roadmaps. Influence vendor and hardware strategy for on-prem and cloud workloads. Design self-healing... 
    Full time
    Work at office
    Remote work
    Worldwide
    Shift work
    Weekend work

    Zoom

    San Jose, CA
    5 days ago
  • $236k - $329k

     ...manufacturing teams to understand and analyze hardware quality issues.Partner with internal...  ...:Bachelor's degree in Electrical Engineering, Computer Engineering, Physics, a related...  ...and the Data Center team.Our Platforms Infrastructure Engineering team designs and builds the... 

    Google

    Sunnyvale, CA
    1 day ago
  • $65 - $85 per hour

     ...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming...  ...Engineer (Contract) to work in IPP (Infrastructure, Planning and Process). IPP is a...  .../Linux/Android), a multitude of hardware platforms both NVIDIA GPUs and Tegra... 
    Full time
    Contract work
    Worldwide

    Sustainable Talent

    Santa Clara, CA
    1 day ago
  •  ...challenges with our customers. Our global team of more than 3,000 engineers works across electrical, mechanical, software, design...  ...chooseSummaryWe are seeking a detail-oriented Lead Engineer, Hardware Test to join our hardware validation team. In this role, you will... 
    Local area

    Celestica

    San Jose, CA
    1 day ago
  • $145k - $165k

     ...Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate...  ...highly available, fault-tolerant infrastructure and services. Install, maintain, and...  ...server, storage, and networking hardware in office and colocation facilities.... 
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    1 day ago
  •  ...Senior Site Reliability Engineer Location: Remote Duration: 12 month contract to start...  ...optimizing existing systems, building infrastructure and eliminating work through automation...  ...sources of issues and the impact on hardware, network, or service operations and... 
    Contract work
    Local area
    Remote work

    My3Tech Inc

    Sunnyvale, CA
    3 days ago
  • $134.9k - $185k

     ...and development company that designs and engineers high-profile electronic devices. Amazon...  ..., Fire TV, and Amazon Echo. Amazon reliability team aims to develop reliable and robust...  ...delight our customers. In this role, as a Hardware Reliability Engineer, you will be... 
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    2 days ago
  • $148.32k - $203.94k

     ...higher performance, smaller size, lower power, and better reliability. With more than 4 billion devices shipped, SiTime is...  ...visit: .Job SummaryWe are seeking a hands-on Principal Infrastructure Hardware Engineer to architect, design, and deliver system platforms supporting... 

    SiTime

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer - Hardware Infrastructure. Be the first to apply!