Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior System Reliability Engineer

$168k - $264.5k

NVIDIA

NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing — with the GPU acting as the brains of computers, robots, and self-driving cars that can perceive and understand the world. Today, we are increasingly known as “the AI computing company.” We're looking to grow our company and build our teams with the most thoughtful people in the world. Join us at the forefront of technological advancement. GPU Servers are one of the fastest-growing segments for NVIDIA and the Artificial Intelligence industry. As the computational power increases with every GPU generation, developing efficient and reliable systems is an imperative. We are looking for a System Reliability Engineer to join NVIDIA's existing Reliability Engineering team, involved in NVIDIA's diverse system product range specifically Graphics and High-Performance Computing printed circuit boards and Data Center Servers.What you'll be doing:Represent Product Reliability Engineering in development teams.Develop and execute reliability test plans for product qualification. Collaborate with peer reliability groups for successful implementation.Define product reliability tests for products ranging from embedded, automotive, graphics cards to server, rack, and cluster products.Lead reliability testing, failure analysis, and root cause investigations; drive corrective actions to improve design and manufacturing quality.Establish and continuously improve product reliability standards, metrics, and methodologies.Participate in product and engineering design reviews, assess the reliability budget and influence changes that enhance product reliability.Collaborate cross-functionally with engineering teams, suppliers, and partners to achieve reliability targets using Design for Reliability (DfR) methods including FMEA and DoE approaches.Provide reliability predictions to access and drive product reliability to meet product requirements.What we need to see:Bachelor’s or Master’s degree in Electrical Engineering, Mechanical Engineering, or equivalent experience.8+ years of experience in hardware reliability or hardware engineering from datacenter, systems, or computer industries.Hands-on experience in theoretical and practical Reliability concepts as it relates to high-tech electronic enterprise and consumer products.Have a strong understanding of statistical concepts and how they relate to product reliability & life analysis.Good verbal and writing skills as well as the ability to communicate at a high level.Good project management skills and ability to balance multiple simultaneous projects during development.Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 264,500 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 7, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the Senior System Reliability Engineer in Santa Clara, CA vacancy
  • $116k - $184k

     ...build the next era of computing!We're seeking an outstanding Senior HTOL Reliability Engineer to join our Santa Clara lab. This role requires deep...  ...develop and implement improvements to burn-in boards, HTOL systems, and thermal interface materials.What we need to see:... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $167.3k - $284.4k

     ...hands without us. KLA invents systems and solutions for the...  ...expert teams of physicists, engineers, data scientists and problem-...  ...application development engineers, and senior product technology process...  .../Preferred QualificationsSr. Reliability Engineer - SEM SystemsJoin a... 
    Senior
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    3 days ago
  • $210.16k - $271.98k

    Senior Principal Systems Development EngineerHelp architect and deliver Dell's L11 rack-scale AI solutions...  .... As a Principal Systems Development Engineer, you'll own the system-level...  ...requirements.Ensure the manufacturability, reliability, and serviceability of rack designs,... 
    Senior

    Dell Technologies

    Santa Clara, CA
    1 day ago
  •  ...Senior Principal Systems Development EngineerHelp architect and deliver Dell's L11 rack-scale AI solutions: fully integrated racks that bring...  ...our customers' AI growth. As a Principal Systems Development Engineer, you'll own the system-level architecture and engineering of... 
    Senior

    Dell Careers

    Santa Clara, CA
    1 day ago
  • $332k

     ...great technology—and amazing people.NVIDIA's hardware reliability is foundational to some of the world's most...  ...critical conditions. As Sr. Director of Board and System Level Reliability, you will set the engineering standard for how NVIDIA's products perform and thrive... 
    Senior
    Full time
    Work at office

    Nvidia

    Santa Clara, CA
    4 days ago
  • $136k - $218.5k

    We are seeking Systems Quality and Reliability Engineer to join our LPU team!NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $115k - $140k

     ...Job Description Job Description Sr. NPI Systems Electrical Engineer The Company Halo Industries has developed breakthrough technology...  ...automation and data analysis. · Experience with high-reliability and safety-critical systems. · Any relevant... 
    Senior
    Full time
    Contract work
    Temporary work
    Work at office

    Halo Industries, Inc.

    Santa Clara, CA
    5 days ago
  • $106.9k

     ...jobs that impact everyone's life. Build the future nobody’s dreamed of yet... are you ready to rethink the impossible? As a Senior Engineer Systems in our Research & Development team, you'll have the opportunity to merge creativity with your technical expertise by... 
    Senior
    Local area

    Infineon Technologies

    San Jose, CA
    2 days ago
  • $155.5k - $248.9k

     ...deliver exceptional results.Opportunity OverviewTeradyne's Memory Test Division is seeking a Senior Reliability Design Engineer to help develop next-generation semiconductor test systems by driving reliability into products from concept through qualification and release. In... 
    Senior
    Local area
    Relocation
    Flexible hours

    Universal Robots

    San Jose, CA
    2 days ago
  • $136k - $218.5k

     ...Automotive, and Embedded markets. As a Silicon Speed Features Engineer, you will co-design system-level speed features, build the validation and...  ...system architects, hardware, firmware/software, process/reliability, and operations teams to co-design system-level speed features... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $168k - $264.5k

     ...products. With our invention of the GPU - the engine of innovative visual computing - the...  ..., as the car is becoming a connected system, and graphics and computing power is...  ...automotive industry, we are now looking for a Senior Reliability Engineer. The position is an individual... 
    Senior

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $133.1k - $306.4k

    OCI Network Availability is seeking a Senior Manager to lead a Networking Reliability Engineering team responsible for driving operational excellence across the...  ...services.Solve complex problems across distributed systems, network infrastructure, and highly available... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Santa Clara, CA
    3 days ago
  •  ...technologies—like the da Vinci surgical system and Ion—have transformed how care is delivered...  ...of patients worldwide.We’re a team of engineers, clinicians, and innovators united by one...  ...DescriptionPrimary Function of Position Senior Systems Analysts (Robotic Control Engineers... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    3 days ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually...  ...Sumo’s developers to deliver features more rapidly.Scale systems sustainably through mechanisms like automation, and... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    2 days ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building...  ...includes the Lambda website, cloud APIs and systems as well as internal tooling for system...  ...networking teams to improve service reliability and deployment workflowsDeploy and maintain... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...functional/non-functional requirements to system requirements.• Experience and...  ...of in Production support and performance engineering.• Technical Skills.• Ability to work in...  ...Information TechnologyExperience level: Mid-Senior LevelIndustry: Information Technology And... 
    Senior
    Permanent employment
    Full time

    Sonsoft

    Sunnyvale, CA
    1 day ago
  •  ...automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud...  ...our environment. Your mission is to take a high-velocity system and implement the best practices, guardrails, and automated... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    4 days ago
  • $101k - $161k

     ...prestigious awards, such as Best Engineering Team, Best Company for...  ...WithWe’re looking for Site Reliability Engineers to join our growing...  ...software engineering background, systems architecture knowledge, with...  ...level: Mid-Senior LevelIndustry: Computer Networking
    Senior

    Arista Networks

    Santa Clara, CA
    2 days ago
  • $174k - $253k

     ...they go live through activities such as system design consulting, developing software...  ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... 
    Senior

    Google

    Sunnyvale, CA
    2 days ago
  • $148k - $235.75k

     ...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps...  ...turns raw, high-volume telemetry into reliable, job-centric insights and automation for...  ...You’ll partner Software Engineering and Systems Engineering team to translate platform signals... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $90k - $180k

     ...more than 160 countries.JOB DESCRIPTION:About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale...  ...building and maintaining the resilient backbone for systems where failure is not an option, and where our success directly... 
    Senior
    Remote work
    Shift work

    Abbott

    Sunnyvale, CA
    11 hours ago
  • $262k - $365k

     ...they go live through activities such as system design consulting, developing software...  ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines... 
    Senior

    Google

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...intelligence.We’re looking for a Senior SRE to join our Compute Farm...  ...ll keep critically important systems running while working on the...  ...lifecycle management, fleet reliability/auto-healing, E2E...  ...Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Apple Inc. in Sunnyvale, CA, is seeking a senior product design engineer to lead mechanical design of mechanisms and mechatronic systems from concept through validation. You will partner with electrical and firmware teams, craft 3D CAD models and 2D drawings, and drive... 
    Senior

    Apple

    Sunnyvale, CA
    4 days ago
  • $135.2k - $306.4k

    We are seeking a highly skilled Linux Systems Engineer with deep expertise across the Linux operating system stack. The ideal candidate is...  ....• Collaborate with engineering teams to improve platform reliability, performance, and maintainability.• Participate in root-cause... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    NVIDIA is now looking for a Senior Memory System Engineer to join our ASIC Memory Subsystem team! As a Senior Systems Engineer at NVIDIA, you'll join a group of hardworking engineers to develop and architect innovative Memory Solution for Tegra SoCs. In this position, you... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $168k - $258.75k

    Join NVIDIA's datacenter product engineering team in our Operations organization and be at the forefront of technological advancement! As a Senior System Debug Engineer, you will drive failure analysis and debug efforts during our New Product Introduction (NPI) phase.... 
    Senior
    Full time
    Work experience placement
    Overseas

    Nvidia

    Santa Clara, CA
    3 days ago
  • $136k - $212.75k

     ...Make the choice to join us today. We are now looking for a Senior Validation Engineer in the DGX Server Product Engineering Team. In this role you...  ...GPU accelerated computing products.What you will be doing:System architecture, design, performance modelling, estimation across... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    4 days ago
  • $184k - $260k

     ...Senior Principal System Validation Engineer San Jose, CA Astera Labs (NASDAQ: ALAB)provides rack-scale AI infrastructure through purpose-built connectivity solutions grounded in open standards. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables... 
    Senior
    Flexible hours

    Astera Labs

    San Jose, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior System Reliability Engineer. Be the first to apply!