Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer - Hardware Infrastructure

$184k - $287.5k

NVIDIA

At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation with consistent reliability and uptime. As an SRE here, you will be part of a welcoming team that values collaboration and creativity, empowering developers to make significant updates while sustaining efficient system function.What you'll be doing:Develop and support guidelines for incident management, planned maintenance, and blameless postmortems.Assist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions.Define reliability and supportability metrics, Service Level Objectives, and error budgets.Develop and drive the adoption of actionable, customer-centric monitoring and alerting.Apply automation and Generative AI/Agentic solutions to minimize manual and tedious activities and boost customer support.Guide teams on establishing sustainable on-call and operational standards.What we need to see:Degree in Computer Science or a related technical field involving coding, or equivalent experience.8+ years of experience in SRE, DevOps, or Production Engineering.Strong understanding of SRE principles, including incident management, error budgets, SLOs, and SLAs.Experience crafting and deploying systems that are fault-tolerant, performant, and supportable.Background with infrastructure automation.Experience running critical services in production.Experience in one or more of the following: Python, Go, Perl, or Ruby.Hands-on experience with observability platforms (e.g., Prometheus, Grafana).Strong communication skills with the ability to convey technical concepts effectively to diverse audiences.Flexibility and adaptability working in a fast-paced environment with evolving requirements.Ways to stand out from the crowd:Expertise in establishing incident management and postmortem processes.Experience driving adoption of common tools and processes across diverse groups.Experience working with LLM/Generative AI/Agentic solutions to shorten mitigation time, lessen toil, and ensure Service Level Objectives are met.Hands-on expertise operating and scaling distributed systems with tight SLAs, ensuring high availability and performance.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer - Hardware Infrastructure in Santa Clara, CA vacancy
  • $159k - $230k

     ...requirements, define architecture, execute hardware design, and product validation.Lead the...  ...:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a related...  ...deployed in the data center.Our Platforms Infrastructure Engineering team designs and builds the... 
    Suggested
    Worldwide

    Google

    Sunnyvale, CA
    3 days ago
  • $184k - $287.5k

     ...push the boundaries of innovation and engineering? At NVIDIA, we lead the world in accelerated...  ...high‑performance systems.As a Senior Hardware Systems Engineer, you will help build...  ...Familiarity with hyperscale data center infrastructure, including cooling methods, facility... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $132k - $190k

    Execute functional validation planning and participate in hardware design reviews for key sub-modules and interfaces to ensure specification...  ...cycle.Minimum qualifications:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a related field, or equivalent... 
    Suggested
    Worldwide

    Google

    San Jose, CA
    1 day ago
  • $255k - $340k

     ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of...  ...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building...  ...with the quality team and fleet reliability team during hardware NPI and after production... 
    Suggested
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    17 hours ago
  • $132k - $190k

    Support various hardware prototype bringup and qualification testing activities by developing...  ...reproduction.Develop software and infrastructure required for managing fleet of test...  ...hardware designers, qualification and test engineers, on project planning within hardware... 
    Suggested

    Google

    Sunnyvale, CA
    1 day ago
  • $147k - $211k

     ...development code.Review code developed by other engineers and provide feedback to ensure best...  ...sources of issues and the impact on hardware, network, or service operations and...  ...troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when... 

    Google

    Sunnyvale, CA
    17 hours ago
  • $157.3k - $212.8k

    Amazon Web Services (AWS) Hardware Engineering Services (HWEngS) owns the new product development...  ...NPI) and operation of all AWS global infrastructure. In other words, we’re the people who...  ...as total cost of ownership, quality, reliability, performance, and serviceability. You... 
    Local area
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    4 days ago
  • $183k - $247.6k

    Amazon Web Services (AWS) Hardware Engineering Services (HWEngS) owns the new product development...  ...NPI) and operation of all AWS global infrastructure. In other words, we’re the people who...  ...as total cost of ownership, quality, reliability, performance, and serviceability. You... 
    Local area
    Flexible hours

    AmazonWebServices

    Cupertino, CA
    4 days ago
  • $122.44k - $232.19k

     ...will be joining the Intel Government Technologies Customer Engineering team as a Hardware Platform Applications Engineer (PAE). This is an exciting...  ...support for customer developed systems including providing on-site power on support Participating in the defining and... 
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    2 days ago
  • $147k - $211k

     ...platforms.Inform direction for research where engineering gaps are identified that merit improved...  ...Science.2 years of experience with hardware design, and data structures or...  ...the architecture built by the Technical Infrastructure team to keep it running. From developing... 

    Google

    Sunnyvale, CA
    17 hours ago
  • $136k - $218.5k

     ...dedicated and motivated Software developer with particular interest in algorithms and RTL Design. Understanding both Software and Hardware principles will be a key requirement for this role.What you'll be doing:Architect, design, develop and support tools for RTL generation... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $100k

     ...our team and looking for contributors of all seniorities.We are seeking a junior-to-mid level SOC Emulation Engineer to support our hardware emulation infrastructure and internal chip design teams. This role focuses on integrating vendor and custom hardware transactors,... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    2 days ago
  • $138k - $197k

     ...Bachelor's degree in Electrical Engineering, Computer Engineering,...  ...shape the future of AI/ML hardware acceleration. You will have...  ...methodologies and flows.The AI and Infrastructure team is redefining what’s...  ...scale, efficiency, reliability and velocity. Our customers... 
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago
  • $138k - $198k

     ...and upgrade our emulation infrastructure and act as a primary interface...  ...team members in debug of hardware, tooling, and project specific...  ...'s degree in Electrical Engineering, Computer Engineering, Computer...  ...scale, efficiency, reliability and velocity. Our customers... 
    Worldwide

    Google

    Sunnyvale, CA
    17 hours ago
  • $131k - $175k

     ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,...  ...and cross-functional teams, including hardware, software, thermal, and manufacturing engineers...  ...design and deployment of AI and cloud infrastructure—translating cluster architectures into... 
    Remote work
    Flexible hours

    Arista Networks

    Santa Clara, CA
    4 days ago
  • $120k - $250k

     ...human level intelligence. Its robots are engineered to perform a variety of tasks in the...  ...collaboration. We are looking for a Reliability Test Engineer to design and execute test...  ...are looking for someone with a strong hardware test background and is able to design and... 
    Full time
    Contract work
    Work at office

    Figure

    San Jose, CA
    21 hours ago
  •  ...challenges with our customers. Our global team of more than 3,000 engineers works across electrical, mechanical, software, design...  ...chooseSummaryWe are seeking a detail-oriented Lead Engineer, Hardware Test to join our hardware validation team. In this role, you will... 
    Local area

    Celestica

    San Jose, CA
    2 days ago
  • $145k - $165k

     ...Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate...  ...highly available, fault-tolerant infrastructure and services. Install, maintain, and...  ...server, storage, and networking hardware in office and colocation facilities.... 
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    3 days ago
  • $164.8k - $226.6k

     ...higher performance, smaller size, lower power, and better reliability. With more than 4 billion devices shipped, SiTime is...  ...visit: .Job SummaryWe are seeking a hands-on Principal Infrastructure Hardware Engineer to architect, design, and deliver system platforms supporting... 

    SiTime

    Santa Clara, CA
    1 day ago
  • $136k - $218.5k

    NVIDIA is seeking capable customer-facing hardware engineers to work directly with Cloud Scale Providers (CSP’s) deploying next generation...  ...as AI Factories, are vital to scale compute and networking infrastructure needed for agentic AI processing. The CSP HW Systems... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $168k - $264.5k

     ...? NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA...  ...across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.Author, test, and activate shared and non-shared Cloudlet... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers...  ...across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...THE ROLEWe are hiring Applied AI Engineers to work directly with hardware and software engineering teams on...  ...feedback.Work with research and infrastructure teams to generalize repeated patterns...  ...systems while preserving reliability and auditability.Collaboration with... 

    AMD

    Santa Clara, CA
    17 hours ago
  • $132k - $190k

     ...management, etc.) to design, implement, and deliver prototype hardware to enable development activities.Plan and execute bring-up and...  ...needed).Minimum qualifications:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a related field, or equivalent... 
    Worldwide

    Google

    Mountain View, CA
    1 day ago
  • $140k - $224.25k

     ...Technical Sourcing Lead to help us build the strength of our Hardware Sourcing team. NVIDIA has been transforming computer graphics,...  ...business to develop an inclusive candidate pool pipeline.Work with engineering leaders and recruiters in defining and maintaining hiring... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $236k - $330k

     ...system and subsystem hardware validation for data center...  ...and roadmap for test infrastructure and processes to meet...  ...closely with engineering, manufacturing, product...  ...unparalleled scale, efficiency, reliability and velocity. Our...  ...must be performed on site.Bachelor's degree in... 
    Work at office
    Worldwide

    Google

    Sunnyvale, CA
    4 days ago
  • $116k - $189.75k

     ...architecture, validation, system integration, and external partners to deliver robust, high-performance hardware solutions.What we need to see:BS/MS in Electrical Engineering, Computer Engineering, or related field (or equivalent experience).Strong background in system‑... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $195k - $205k

     ...Title: Staff Software Test Engineer, Hardware Controls This position is based in our Campbell, California offices. This position is on-site & full-time Why Imperative Care? At Imperative Care, we are developing novel robotic-assisted technologies and interventional... 
    Full time
    Work experience placement
    Worldwide

    Imperative Care

    Campbell, CA
    21 hours ago
  • $165.2k - $223.6k

    AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we...  ...we’re looking for talented people who want to help.The AWS Hardware Engineering team creates server designs for Amazon’s innovative web... 
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    17 hours ago
  • $175k - $215k

     ..., and Italy. We’re looking for top-notch Engineers to join our global team. If you’re interested...  ...Role OverviewWe are seeking a Principal Hardware Applications Engineer to act as the...  ...hardware and firmwareSupport customer on-site visits, technical reviews, and failure analysis... 
    Work experience placement
    Remote work

    TDK InvenSense

    San Jose, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer - Hardware Infrastructure. Be the first to apply!