Senior Reliability Engineer
$116k - $184kJobleads-US
NVIDIA is the world leader in accelerated computing, developing breakthroughs that tackle challenges no one else can solve. Our work in AI and digital twins is transforming the world's largest industries and profoundly impacting society. Come join the team and help build the next era of computing! We're seeking an outstanding Senior HTOL Reliability Engineer to join our Santa Clara lab. This role requires deep device-circuitry knowledge and hands-on hardware development. You will build next-generation HTOL boards and run HTOL processes on advanced ovens. This ensures world-class reliability of the silicon powering the AI era.
What you'll be doing:
- Implement and optimize HTOL test programs aligned with JEDEC standards.
- Operate and maintain HTOL ovens, ensuring efficient test conditions and high data accuracy.
- Design, debug, and bring up HTOL burn in boards.
- Debug and bring up HTOL patterns.
- Apply sophisticated thermal management techniques to deliver detailed temperature control and mitigate thermal stress in HTOL environments.
- Work alongside lab technicians and reliability engineers to solve technical challenges and continuously improving test processes.
- Contribute to cross-functional teams to debug and resolve hardware and software product issues in the HTOL environment.
- Maintain and improve our reliability database, finding opportunities for improvement.
- Collaborate with vendors to develop and implement improvements to burn-in boards, HTOL systems, and thermal interface materials.
What we need to see:
- Master's or Bachelor's degree in Electrical Engineering or a related field (or equivalent experience).
- 5+ years of experience in HTOL test system operation and data analysis for semiconductor devices.
- Proven expertise in HTOL stress testing, JEDEC standards, and environmental stress tests, including Temperature Cycling (TC), Reflow, Thermal Shock, and HAST.
- Strong ATE or TE skills, including pattern bring up and pcb board debugging and design.
- Hands‑on experience with High power HTOL chambers including operation, repair, and preventative maintenance of HTOL chamber.
- Proficiency with oscilloscopes, current probes, and other test equipment for data acquisition and analysis.
- Experience with pattern vector debugging, test script development/modification, and data analysis tools.
- Programming experience with Python or MATLAB for data analysis and automation.
- Excellent communication, teamwork, and problem solving skills, with strong attention to detail.
Ways to stand out from the crowd:
- Experience with multi-die HTOL testing and the associated testing challenges.
- Background in HTOL board design for high power GPU or SoC devices.
- Pattern translations, debugging, bringup, experience working with DFT.
- Familiarity with reliability analytics platforms (e.g., JMP) and statistical lifetime modeling (e.g., Weibull, Arrhenius).
- Track record of driving vendor qualification and component selection for reliability test hardware.
With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers.
We have some of the most forward-thinking and hardworking people on the planet working for us.
If you're creative and autonomous, with a genuine passion for technology, we want to hear from you.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 116,000 USD - 184,000 USD. You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until August 9, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.
#J-18808-Ljbffr Jobleads-US- ...Zealogics.com is seeking a Hardware FMEA & Reliability Engineer to drive DFMEA, PFMEA, and reliability activities across the product lifecycle. You will collaborate with design, manufacturing, quality, and suppliers to ensure robust, high-quality hardware solutions....Senior
- ...Everpure seeks a Senior Software Engineer to drive technical resolution and stability of our hyperscale cloud storage deployments for top-tier... ...multi‑million‑dollar account success, collaborating with hyperscale teams to architect reliable #J-18808-Ljbffr Jobleads-USSeniorWork at office
- ...NVIDIA Corporation seeks a Senior HTOL Reliability Engineer for our Santa Clara lab. You will lead HTOL test programs, operate HTOL ovens, and bring up burn-in boards and patterns, applying advanced thermal management to ensure device reliability. Collaboration with...Senior
$168k - $264.5k
...industry. As the computational power increases with every GPU generation, developing efficient and reliable systems is an imperative. We are looking for a System Reliability Engineer to join NVIDIA's existing Reliability Engineering team, involved in NVIDIA's diverse system...Senior$90k - $130k
...Senior Database Reliability EngineerMust Have Technical/Functional Skills • 5+ years of experience designing, operating, and troubleshooting... ...network layers. • 3+ years of experience in Linux systems engineering (performance tuning, memory management, I/O tuning, configuration...SeniorRemote work- ...Meta Reality Labs seeks a Sustaining Reliability and Quality Engineer to own long-term hardware reliability across the device portfolio. You will set the reliability strategy, drive field issue resolution, and influence design and supplier processes to sustain performance...Senior
$174k - $252k
...automation, and evolve systems by pushing for changes that improve reliability and velocity.Practice sustainable incident response and... ...Minimum qualifications:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience.5 years...Senior- ...About the Team: SK HMS Systems engineering / Quality assurance group is looking for a motivated individual to join our reliability testing team. Job Description:... ...track record working as a contractor or senior consultant, capable of rapidly onboarding...SeniorContract workFor contractors
$160k - $240k
...another millions of times a day - quickly, reliably, and securely. Any time you swipe your... ...does a successful Site Reliability Engineer do at Fiserv?You will join our global team... ...reliability, operations or DevOps at a mid-to-senior level.Strong shell scripting skills and...SeniorFull time- ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation...Senior
$104.9k - $174.7k
About the Role:The SRE role is responsible for improving the reliability, availability, performance, and operational quality of production... ...and preventive actions through completion.Follow up with engineering, development, security, support, and business stakeholders to...SeniorFull timeLocal area- ...Apple Inc. is seeking a Hardware Reliability Engineer for Watch System Reliability in Cupertino, CA. You will define durability from architecture to mass production, develop accelerated tests, and lead cross-functional teams to mitigate risk and improve product robustness...Senior
$262k - $364k
...infrastructure from SRE side, ensuring it is reliable, scalable, cost effective and performant, while working closely with senior technical leads in the development teams.... ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines...Senior$332k
...by great technology—and amazing people. NVIDIA's hardware reliability is foundational to some of the world's most demanding computing... ...of Board and System Level Reliability, you will set the engineering standard for how NVIDIA's products perform and thrive across...SeniorWork at office$132.6k - $214.5k
...you will collaborate closely with our engineering teams to develop innovative solutions that... ...’ performance and health. As a Senior Staff SRE with the Cortex Observability... ...operability of the product and ensure the reliability and availability of our services....SeniorFull timeWork at officeVisa sponsorshipWork visa- ...Adobe’s Database Reliability Engineering (DBRE) team is seeking a Senior Database Reliability Engineer (Individual Contributor) to own and scale database platforms across global services. You will drive reliability, performance, and resilience, partnering with engineering...Senior
$150.4k - $277.6k
...Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple’s long... ...field with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role Advanced...SeniorRelocationDay shift- ...NVIDIA is seeking a System Reliability Engineer to join the Reliability Engineering team, focusing on GPUs, data center servers, and related products. You will define reliability tests, lead failure analyses, and drive corrective actions to improve design and manufacturing...Senior
$152k - $241.5k
...infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience who have worked... ...initial provisioning through repair.Experience managing production reliability through on-call duties, incident response, observability, and...SeniorPermanent employmentFull time- ...Own the architecture and design of reliable, scalable, cost-effective, and performant AI inference and training infrastructure, partnering... ...is required, while a master's degree in computer science or engineering is preferred. Key Skills Software Development, Technical...Senior
$148k - $235.75k
...where everyone is inspired to do their best work. Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job‑centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...Senior$262k - $364k
...Senior Staff Software Engineer, Site Reliability Engineering, Workspace AI Google Sunnyvale, CA, USA X In most instances, this position requires in-person interviews as part of the hiring process. ~ Bachelor’s degree in Computer Science, a related field, or equivalent...Senior- ...research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality. As the Senior Reliability Engineer, you own reliability as an engineering discipline, not just a test outcome. Every mission profile we define, every...Senior
$164k - $248k
DescriptionJob Description: Staff reliability engineer will serve as the lead engineer to plan, coordinate, and execute on accelerated life testing, environmental testing, characterization, and other tasks required for timely qualification of Power Integrations products...SeniorWork experience placement- ...SUV miles with ones on vehicles that are more affordable, more enjoyable and 10-50x more efficient. ALSO is looking for a Reliability Engineer to play a key role in developing and leading the reliability of multiple electric mobility vehicles over the full...SeniorWork at officeLocal areaRemote workFlexible hours1 day per week
- ...KLA in Milpitas, California, seeks a Sr. Reliability Engineer for SEM systems to drive reliability and uptime of wafer inspection tools. You will investigate failure mechanisms, perform root-cause analysis, and lead DOE studies, collaborating with hardware, firmware,...Senior
$55 - $60 per hour
...Description Job Description Pay Range: $55.00hr - $60.00hr Job Overview: Our client is looking for an experienced Site Reliability Engineer (SRE) to join the Infrastructure Platform Engineering team. In this role, the successful candidate will help design, scale,...SeniorTemporary workLocal area$190k - $222.5k
...miles with ones on vehicles that are more affordable, more enjoyable and 10-50x more efficient. ALSO is looking for a Senior Reliability Engineer to play a key role in developing and leading the reliability of multiple electric mobility vehicles over the full...SeniorWork at officeLocal areaRemote workFlexible hours1 day per week- ...requirements Drive continuous improvement initiatives to enhance reliability, scalability, and efficiency of infrastructure and services,... ...Qualifications ~ Bachelor’s degree in computer science, Engineering, or related field; or equivalent work experience ~5+ years...SeniorWork experience placement
$131.1k - $222.76k
Job Summary:We are now looking for a Principal Reliability Engineer. The position is an individual contributor role, working on NPD Reliability Engineering & qualifications, and MP change qualifications as needed.Job Description & Responsibilities:For NPD projects, the...Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Reliability Engineer. Be the first to apply!
- senior reliability engineer Santa Clara, CA
- sr reliability engineer Santa Clara, CA
- reliability engineer Santa Clara, CA
- senior computer engineer Santa Clara, CA
- senior manager customer operations Santa Clara, CA
- senior development engineer Santa Clara, CA
- senior software engineer ruby on rails Santa Clara, CA
- sr marketing manager Santa Clara, CA
- senior customer service Santa Clara, CA
- senior business manager Santa Clara, CA



