Site Reliability Engineer - Hardware Infrastructure
$184k - $287.5kNVIDIA
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation with consistent reliability and uptime. As an SRE here, you will be part of a welcoming team that values collaboration and creativity, empowering developers to make significant updates while sustaining efficient system function.What you'll be doing:Develop and support guidelines for incident management, planned maintenance, and blameless postmortems.Assist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions.Define reliability and supportability metrics, Service Level Objectives, and error budgets.Develop and drive the adoption of actionable, customer-centric monitoring and alerting.Apply automation and Generative AI/Agentic solutions to minimize manual and tedious activities and boost customer support.Guide teams on establishing sustainable on-call and operational standards.What we need to see:Degree in Computer Science or a related technical field involving coding, or equivalent experience.8+ years of experience in SRE, DevOps, or Production Engineering.Strong understanding of SRE principles, including incident management, error budgets, SLOs, and SLAs.Experience crafting and deploying systems that are fault-tolerant, performant, and supportable.Background with infrastructure automation.Experience running critical services in production.Experience in one or more of the following: Python, Go, Perl, or Ruby.Hands-on experience with observability platforms (e.g., Prometheus, Grafana).Strong communication skills with the ability to convey technical concepts effectively to diverse audiences.Flexibility and adaptability working in a fast-paced environment with evolving requirements.Ways to stand out from the crowd:Expertise in establishing incident management and postmortem processes.Experience driving adoption of common tools and processes across diverse groups.Experience working with LLM/Generative AI/Agentic solutions to shorten mitigation time, lessen toil, and ensure Service Level Objectives are met.Hands-on expertise operating and scaling distributed systems with tight SLAs, ensuring high availability and performance.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global... ...like Windows, Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra Processors...SuggestedFull timeWork experience placementWorldwide$132k - $190k
...requirements, define architecture, execute hardware design, and product validation.Lead the... ...:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a related... ...deployed in the data center.Our Platforms Infrastructure Engineering team designs and builds...SuggestedWorldwide$184k - $287.5k
...push the boundaries of innovation and engineering? At NVIDIA, we lead the world in accelerated... ...high‑performance systems.As a Senior Hardware Systems Engineer, you will help build... ...Familiarity with hyperscale data center infrastructure, including cooling methods, facility...SuggestedFull time$132k - $190k
Execute functional validation planning and participate in hardware design reviews for key sub-modules and interfaces to ensure specification... ...cycle.Minimum qualifications:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a related field, or equivalent...SuggestedWorldwide$255k - $340k
...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of... ...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building... ...with the quality team and fleet reliability team during hardware NPI and after production...SuggestedWork at officeLocal areaWork from homeFlexible hours$172k - $246k
...tomorrow’s standard —from breakthrough hardware and battery systems to intuitive design,... ...services. We are seeking a Staff Hardware Engineer to define and drive the architecture,... ...tradeoffs across performance, power, cost, reliability, and scalability.Establish design...Full timeLocal areaWork from homeRelocationRelocation packageFlexible hours3 days per week- ...for an Open Source Software Engineer to build and optimize machine... ...that integrate with AMD hardware. You will work at the intersection... ...testing and validation infrastructure.Participate in code reviews... ...advanced AI workloads fast, reliable, and broadly accessible on AMD...
$188k - $274k
Lead a system and hardware design on data center hardware products.Bring up systems and execute engineering validation in the lab.Gather requirements, define architecture,... ...world's largest and most effective computing infrastructure. You will see those systems from concepts...$159k - $230k
...facilitate system performance, design reliability and failure reproduction.Design network, power and cooling infrastructure for end-to-end system testing.Drive... ....Contribute within a team of hardware/software designers, test engineers, on project planning within hardware...- Dawar Consulting is hiring a System Engineer specializing in Hardware Support in Santa Clara, CA. This long-term contract role involves troubleshooting and supporting next-generation sequencing systems and sample prep platforms. The ideal candidate should have a B.S. in...Long term contract
$159k - $230k
...issues.Support design engineers with debug, component... ...thoroughly down to a hardware interface or... ...sustaining efforts.The AI and Infrastructure team is redefining... ...unparalleled scale, efficiency, reliability and velocity. Our... ...must be performed on site.Bachelor’s degree in...Worldwide$166k - $244k
Overview Site Reliability Engineering (SRE) combines software and systems engineering to build and run... ...existing systems, building infrastructure and eliminating work through automation... ...IP, routing, network topologies and hardware, SDN). 2 years of experience leading...Full time$147k - $210k
...development code.Review code developed by other engineers and provide feedback to ensure best... ...sources of issues and the impact on hardware, network, or service operations and... ...troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when...$267k - $356k
...a leader in AI cloud infrastructure serving tens of thousands... ....Lambda's Storage Engineering team is the backbone... ...industry, which means reliability and performance aren'... ..., capacity, and hardware failures.Investigate... ...across new and existing sites using tools such as Ansible...Work experience placementWork at officeLocal areaWork from homeFlexible hours$100k
...team and looking for contributors of all seniorities. We are seeking a junior-to-mid level SOC Emulation Engineer to support our hardware emulation infrastructure and internal chip design teams. This role focuses on integrating vendor and custom hardware transactors, developing...Permanent employmentFull time$157.3k - $212.8k
Amazon Web Services (AWS) Hardware Engineering Services (HWEngS) owns the new product development... ...NPI) and operation of all AWS global infrastructure. In other words, we’re the people who... ...as total cost of ownership, quality, reliability, performance, and serviceability. You...Local areaFlexible hours$122.44k - $232.19k
...will be joining the Intel Government Technologies Customer Engineering team as a Hardware Platform Applications Engineer (PAE). This is an exciting... ...support for customer developed systems including providing on-site power on support Participating in the defining and...Full timeInternshipLocal areaImmediate startShift work$147k - $211k
...platforms.Inform direction for research where engineering gaps are identified that merit improved... ...Science.2 years of experience with hardware design, and data structures or... ...the architecture built by the Technical Infrastructure team to keep it running. From developing...$136k - $218.5k
...dedicated and motivated Software developer with particular interest in algorithms and RTL Design. Understanding both Software and Hardware principles will be a key requirement for this role.What you'll be doing:Architect, design, develop and support tools for RTL generation...Full time$138k - $197k
...Bachelor's degree in Electrical Engineering, Computer Engineering,... ...shape the future of AI/ML hardware acceleration. You will have... ...methodologies and flows.The AI and Infrastructure team is redefining what’s... ...scale, efficiency, reliability and velocity. Our customers...Worldwide$131k - $175k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...and cross-functional teams, including hardware, software, thermal, and manufacturing engineers... ...design and deployment of AI and cloud infrastructure—translating cluster architectures into...Remote workFlexible hours$138k - $197k
...and upgrade our emulation infrastructure and act as a primary interface... ...team members in debug of hardware, tooling, and project specific... ...'s degree in Electrical Engineering, Computer Engineering, Computer... ...scale, efficiency, reliability and velocity. Our customers...Worldwide$146.7k - $339.3k
...available for this positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid... ...teams on architecture roadmaps. Influence vendor and hardware strategy for on-prem and cloud workloads. Design self-healing...Full timeWork at officeRemote workWorldwideShift workWeekend work$236k - $329k
...manufacturing teams to understand and analyze hardware quality issues.Partner with internal... ...:Bachelor's degree in Electrical Engineering, Computer Engineering, Physics, a related... ...and the Data Center team.Our Platforms Infrastructure Engineering team designs and builds the...$65 - $85 per hour
...Site Reliability Engineer Sustainable Talent is partnering with a global leader who's been transforming... ...Engineer (Contract) to work in IPP (Infrastructure, Planning and Process). IPP is a... .../Linux/Android), a multitude of hardware platforms both NVIDIA GPUs and Tegra...Full timeContract workWorldwide- ...challenges with our customers. Our global team of more than 3,000 engineers works across electrical, mechanical, software, design... ...chooseSummaryWe are seeking a detail-oriented Lead Engineer, Hardware Test to join our hardware validation team. In this role, you will...Local area
$145k - $165k
...Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate... ...highly available, fault-tolerant infrastructure and services. Install, maintain, and... ...server, storage, and networking hardware in office and colocation facilities....Work at officeImmediate start- ...Senior Site Reliability Engineer Location: Remote Duration: 12 month contract to start... ...optimizing existing systems, building infrastructure and eliminating work through automation... ...sources of issues and the impact on hardware, network, or service operations and...Contract workLocal areaRemote work
$134.9k - $185k
...and development company that designs and engineers high-profile electronic devices. Amazon... ..., Fire TV, and Amazon Echo. Amazon reliability team aims to develop reliable and robust... ...delight our customers. In this role, as a Hardware Reliability Engineer, you will be...Local areaFlexible hours$148.32k - $203.94k
...higher performance, smaller size, lower power, and better reliability. With more than 4 billion devices shipped, SiTime is... ...visit: .Job SummaryWe are seeking a hands-on Principal Infrastructure Hardware Engineer to architect, design, and deliver system platforms supporting...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer - Hardware Infrastructure. Be the first to apply!
- site reliability engineer remote Santa Clara, CA
- site reliability engineer Santa Clara, CA
- site reliability engineer sre Santa Clara, CA
- data infrastructure engineer Santa Clara, CA
- infrastructure engineering manager Santa Clara, CA
- senior infrastructure engineer Santa Clara, CA
- infrastructure automation engineer Santa Clara, CA
- principal infrastructure engineer Santa Clara, CA
- remote infrastructure engineer Santa Clara, CA
- lead infrastructure engineer Santa Clara, CA

