Site Reliability Engineer - Hardware Infrastructure
$184k - $287.5kNVIDIA
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation with consistent reliability and uptime. As an SRE here, you will be part of a welcoming team that values collaboration and creativity, empowering developers to make significant updates while sustaining efficient system function.What you'll be doing:Develop and support guidelines for incident management, planned maintenance, and blameless postmortems.Assist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions.Define reliability and supportability metrics, Service Level Objectives, and error budgets.Develop and drive the adoption of actionable, customer-centric monitoring and alerting.Apply automation and Generative AI/Agentic solutions to minimize manual and tedious activities and boost customer support.Guide teams on establishing sustainable on-call and operational standards.What we need to see:Degree in Computer Science or a related technical field involving coding, or equivalent experience.8+ years of experience in SRE, DevOps, or Production Engineering.Strong understanding of SRE principles, including incident management, error budgets, SLOs, and SLAs.Experience crafting and deploying systems that are fault-tolerant, performant, and supportable.Background with infrastructure automation.Experience running critical services in production.Experience in one or more of the following: Python, Go, Perl, or Ruby.Hands-on experience with observability platforms (e.g., Prometheus, Grafana).Strong communication skills with the ability to convey technical concepts effectively to diverse audiences.Flexibility and adaptability working in a fast-paced environment with evolving requirements.Ways to stand out from the crowd:Expertise in establishing incident management and postmortem processes.Experience driving adoption of common tools and processes across diverse groups.Experience working with LLM/Generative AI/Agentic solutions to shorten mitigation time, lessen toil, and ensure Service Level Objectives are met.Hands-on expertise operating and scaling distributed systems with tight SLAs, ensuring high availability and performance.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$159k - $230k
...requirements, define architecture, execute hardware design, and product validation.Lead the... ...:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a related... ...deployed in the data center.Our Platforms Infrastructure Engineering team designs and builds the...SuggestedWorldwide$184k - $287.5k
...push the boundaries of innovation and engineering? At NVIDIA, we lead the world in accelerated... ...high‑performance systems.As a Senior Hardware Systems Engineer, you will help build... ...Familiarity with hyperscale data center infrastructure, including cooling methods, facility...SuggestedFull time$132k - $190k
Execute functional validation planning and participate in hardware design reviews for key sub-modules and interfaces to ensure specification... ...cycle.Minimum qualifications:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a related field, or equivalent...SuggestedWorldwide$255k - $340k
...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of... ...from home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building... ...with the quality team and fleet reliability team during hardware NPI and after production...SuggestedWork at officeLocal areaWork from homeFlexible hours$132k - $190k
Support various hardware prototype bringup and qualification testing activities by developing... ...reproduction.Develop software and infrastructure required for managing fleet of test... ...hardware designers, qualification and test engineers, on project planning within hardware...Suggested$147k - $211k
...development code.Review code developed by other engineers and provide feedback to ensure best... ...sources of issues and the impact on hardware, network, or service operations and... ...troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when...$157.3k - $212.8k
Amazon Web Services (AWS) Hardware Engineering Services (HWEngS) owns the new product development... ...NPI) and operation of all AWS global infrastructure. In other words, we’re the people who... ...as total cost of ownership, quality, reliability, performance, and serviceability. You...Local areaFlexible hours$183k - $247.6k
Amazon Web Services (AWS) Hardware Engineering Services (HWEngS) owns the new product development... ...NPI) and operation of all AWS global infrastructure. In other words, we’re the people who... ...as total cost of ownership, quality, reliability, performance, and serviceability. You...Local areaFlexible hours$122.44k - $232.19k
...will be joining the Intel Government Technologies Customer Engineering team as a Hardware Platform Applications Engineer (PAE). This is an exciting... ...support for customer developed systems including providing on-site power on support Participating in the defining and...Full timeInternshipLocal areaImmediate startShift work$147k - $211k
...platforms.Inform direction for research where engineering gaps are identified that merit improved... ...Science.2 years of experience with hardware design, and data structures or... ...the architecture built by the Technical Infrastructure team to keep it running. From developing...$136k - $218.5k
...dedicated and motivated Software developer with particular interest in algorithms and RTL Design. Understanding both Software and Hardware principles will be a key requirement for this role.What you'll be doing:Architect, design, develop and support tools for RTL generation...Full time$100k
...our team and looking for contributors of all seniorities.We are seeking a junior-to-mid level SOC Emulation Engineer to support our hardware emulation infrastructure and internal chip design teams. This role focuses on integrating vendor and custom hardware transactors,...Permanent employment$138k - $197k
...Bachelor's degree in Electrical Engineering, Computer Engineering,... ...shape the future of AI/ML hardware acceleration. You will have... ...methodologies and flows.The AI and Infrastructure team is redefining what’s... ...scale, efficiency, reliability and velocity. Our customers...Worldwide$138k - $198k
...and upgrade our emulation infrastructure and act as a primary interface... ...team members in debug of hardware, tooling, and project specific... ...'s degree in Electrical Engineering, Computer Engineering, Computer... ...scale, efficiency, reliability and velocity. Our customers...Worldwide$131k - $175k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...and cross-functional teams, including hardware, software, thermal, and manufacturing engineers... ...design and deployment of AI and cloud infrastructure—translating cluster architectures into...Remote workFlexible hours$120k - $250k
...human level intelligence. Its robots are engineered to perform a variety of tasks in the... ...collaboration. We are looking for a Reliability Test Engineer to design and execute test... ...are looking for someone with a strong hardware test background and is able to design and...Full timeContract workWork at office- ...challenges with our customers. Our global team of more than 3,000 engineers works across electrical, mechanical, software, design... ...chooseSummaryWe are seeking a detail-oriented Lead Engineer, Hardware Test to join our hardware validation team. In this role, you will...Local area
$145k - $165k
...Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate... ...highly available, fault-tolerant infrastructure and services. Install, maintain, and... ...server, storage, and networking hardware in office and colocation facilities....Work at officeImmediate start$164.8k - $226.6k
...higher performance, smaller size, lower power, and better reliability. With more than 4 billion devices shipped, SiTime is... ...visit: .Job SummaryWe are seeking a hands-on Principal Infrastructure Hardware Engineer to architect, design, and deliver system platforms supporting...$136k - $218.5k
NVIDIA is seeking capable customer-facing hardware engineers to work directly with Cloud Scale Providers (CSP’s) deploying next generation... ...as AI Factories, are vital to scale compute and networking infrastructure needed for agentic AI processing. The CSP HW Systems...Full time$168k - $264.5k
...? NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA... ...across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.Author, test, and activate shared and non-shared Cloudlet...Full time- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers... ...across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational...Work at officeLocal areaWork from homeFlexible hours
- ...THE ROLEWe are hiring Applied AI Engineers to work directly with hardware and software engineering teams on... ...feedback.Work with research and infrastructure teams to generalize repeated patterns... ...systems while preserving reliability and auditability.Collaboration with...
$132k - $190k
...management, etc.) to design, implement, and deliver prototype hardware to enable development activities.Plan and execute bring-up and... ...needed).Minimum qualifications:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a related field, or equivalent...Worldwide$140k - $224.25k
...Technical Sourcing Lead to help us build the strength of our Hardware Sourcing team. NVIDIA has been transforming computer graphics,... ...business to develop an inclusive candidate pool pipeline.Work with engineering leaders and recruiters in defining and maintaining hiring...Full time$236k - $330k
...system and subsystem hardware validation for data center... ...and roadmap for test infrastructure and processes to meet... ...closely with engineering, manufacturing, product... ...unparalleled scale, efficiency, reliability and velocity. Our... ...must be performed on site.Bachelor's degree in...Work at officeWorldwide$116k - $189.75k
...architecture, validation, system integration, and external partners to deliver robust, high-performance hardware solutions.What we need to see:BS/MS in Electrical Engineering, Computer Engineering, or related field (or equivalent experience).Strong background in system‑...Full time$195k - $205k
...Title: Staff Software Test Engineer, Hardware Controls This position is based in our Campbell, California offices. This position is on-site & full-time Why Imperative Care? At Imperative Care, we are developing novel robotic-assisted technologies and interventional...Full timeWork experience placementWorldwide$165.2k - $223.6k
AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we... ...we’re looking for talented people who want to help.The AWS Hardware Engineering team creates server designs for Amazon’s innovative web...InternshipLocal areaFlexible hours$175k - $215k
..., and Italy. We’re looking for top-notch Engineers to join our global team. If you’re interested... ...Role OverviewWe are seeking a Principal Hardware Applications Engineer to act as the... ...hardware and firmwareSupport customer on-site visits, technical reviews, and failure analysis...Work experience placementRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer - Hardware Infrastructure. Be the first to apply!
- site reliability engineer Santa Clara, CA
- senior infrastructure engineer Santa Clara, CA
- infrastructure engineering manager Santa Clara, CA
- infrastructure engineer Santa Clara, CA
- infrastructure developer Santa Clara, CA
- principal infrastructure engineer Santa Clara, CA
- remote infrastructure engineer Santa Clara, CA
- data infrastructure engineer Santa Clara, CA
- site services specialist Santa Clara, CA
- construction site safety Santa Clara, CA


