Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Staff Site Reliability Engineer - Compute Core Engineering

$200k - $322k

NVIDIA

NVIDIA has been reinventing computer graphics, PC gaming, and accelerated computing for 30 years. It is a unique legacy of innovation that’s fueled by great technology and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, generative AI, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work.We are seeking a highly skilled Senior Staff SRE to join our dynamic team. Our company is at the forefront of technological innovation, and we are dedicated to driving efficiency and optimizing the performance of our infrastructure both on-prem and cloud. Join us in this exciting endeavor!What you will be doing:Lead initiatives to transform IT Compute Core Team, architecture to build new service offerings across On-Prem and CloudYou will design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP. This includes building for performance and reliability at global scale, covering automation, monitoring, high availability, capacity planning, and lifecycle management.Define and implement metrics to measure the efficiency of services and drive efficiency with software and hardware optimizations (SR-IOV/ DPU)Experience with Technologies like eBPF and XDP for Observability & DDoS mitigationCollect and review system data for capacity and planning purposes, analyze capacity data and develop plans for appropriate level enterprise-wide systems, and coordinate with management personnel in implementing changes.Develop and maintain tools for collecting, analyzing, and visualizing data for reporting, alerting, monitoring.Collaborate with NVIDIA leadership, senior engineers, program managers, and product managers to develop compelling IT products and services that meet customer needs.What we need to see:Bachelor’s degree in Engineering, Computer Science, Mathematics, or related field, or equivalent experience12+ years of proven experience in compute platform engineering with a focus on automation.Experience in designing and deploying Containerization architectures and Distributed Systems InfrastructureProven experience evaluating existing application architectures and identify opportunities for containerization to improve scalability, reliability, and efficiency.Strong analytical skills with the ability to define and track key performance metrics.Experience in developing tools for data analysis and performance profiling, Development with Terraform, Config Management tools.Proficiency in programming languages such as Go and/or Python.Linux OS Proficiency with Kernel InternalsExperience with running large environments consisting of BareMetal Build InfrastructureUnderstanding of Network Protocols and Architectures (VLAN/VxLAN/SDN/BGP/Anycast)Ways to stand out from the crowd:Deep understanding of other infrastructure components like, DNS, LDAP, Security Tools etc..Hands-on experience with containers and its implementationDeploying and Managing Services like DNS , LDAP at scaleSolid understanding of microservices architecture, infrastructure as code (IaC) and configuration management tools.NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and passionate people on the planet working for us. If you're creative and autonomous, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 200,000 USD - 322,000 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until September 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, Remote; US, CA, Santa ClaraType: Full time

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Staff Site Reliability Engineer - Compute Core Engineering in Santa Clara, CA vacancy
  •  ...hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give...  ...Tuesday.About the RoleLambda’s Core Cloud Platform powers compute provisioning...  ...data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...Lambda's mission is to make compute as ubiquitous as electricity...  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...networking teams to improve service reliability and deployment...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  •  ...hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give...  ...from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    19 hours ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually...  ...technology stackDeep understanding of AWS Networking, Compute, Storage, and managed services.Competency with modern CI/... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    4 days ago
  •  ...automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud...  ...capacity planning and autoscaling. You will optimize our compute resources (EC2, Lambda) to handle bursty traffic patterns... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    1 day ago
  • $101k - $161k

     ...latest advancements in cloud computing, artificial intelligence,...  ...prestigious awards, such as Best Engineering Team, Best Company for...  ...Work WithWe’re looking for Site Reliability Engineers to join our growing...  ...EngineeringExperience level: Mid-Senior LevelIndustry: Computer... 
    Senior

    Arista Networks

    Santa Clara, CA
    4 days ago
  • $168k - $270.25k

     ...developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as...  ...of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $148k - $235.75k

    NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for...  ...on the world.Join our team of innovative engineers who are building an AI Data Center AIOps...  ...that turns raw, high-volume telemetry into reliable, job-centric insights and automation for... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $267k - $356k

     ...'s mission is to make compute as ubiquitous as electricity...  ....Lambda's Storage Engineering team is the backbone...  ...industry, which means reliability and performance aren't...  ...across new and existing sites using tools such as...  ...Solid understanding of core storage protocols across... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $152k - $241.5k

    NVIDIA has been transforming computer graphics, PC gaming, and accelerated...  ....We’re looking for a Senior SRE to join our Compute Farm...  ...lifecycle management, fleet reliability/auto-healing, E2E observability...  ...Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $168k - $270.25k

    The NVIDIA Experience (NVEX) Solutions Engineering team is looking for an experienced solution engineer focused on customer support of NVIDIA...  ...to support customersWhat we need to see:Minimum of a BS in Computer Engineering, Electrical Engineering, or equivalent experience.... 
    Senior
    Full time
    Weekend work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA is growing a senior engineering team focused on making our compute software stack first-class on NVIDIA CPU platforms. The team turns modern toolchains, build and code-health practices, performance-analysis workflows, and optimization techniques into repeatable improvements... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $262k - $364k

     ...a team of Software/Systems Engineers on projects for users and be...  ...qualifications:Bachelor’s degree in Computer Science, a related field, or...  ..., or a related field.Site Reliability Engineering (SRE) combines software...  ...chose to join SRE.As the Senior Engineering Manager for... 
    Senior

    Google

    Sunnyvale, CA
    3 days ago
  • $187.04k - $359.72k

     ...for changes that improve reliability and velocity. Qualifications...  ...: BS or MS degree in Computer Science, Electrical Engineering, Computer Engineering or...  ...Corporate Functions and more. On-site presence across teams...  ...creativity is at the core of TikTok's mission. Our... 
    Senior
    Temporary work
    Local area
    Overseas
    Shift work

    Tik Tok

    San Jose, CA
    1 day ago
  • $184k - $287.5k

     ...hand does not scale, and we are not going to try. NVIDIA is building a config-as-code foundation for its EDA compute farm, and we need an automation engineer to own it end to end. You are joining at the point where this is still partially manual and inconsistently applied... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

    Quantum computing is a strategic priority for NVIDIA, and our goal is to help accelerate the entire ecosystem. In this role, you’ll join...  ...to enable fault-tolerant quantum computing.Do you love engineering complex systems for controlling a quantum computer and working... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $160k - $240k

     ...millions of times a day - quickly, reliably, and securely. Any time you...  ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our...  ...operations or DevOps at a mid-to-senior level.Strong shell scripting... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

    We are looking for an experienced Senior System Software Engineer to help build NVIDIA Halos. Halos is our full-stack safety platform for physical...  ...systems, safety, security, virtualization, and accelerated computing! If you have a good understanding of System Software... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    19 hours ago
  • $203.45k - $344.3k

     ...a strategic perspective, making it the core infrastructure that supports weekly model...  ...costs, data architecture and closed-loop engineering system.Job ResponsbilitiesResponsible...  ...RequirementsMaster's degree or above in Computer Science, Artificial Intelligence, Automation... 
    Senior
    Full time
    Temporary work
    Work experience placement

    XPENG Motors

    Santa Clara, CA
    2 days ago
  • $262k - $364k

     ...within the AViD ecosystem have reliability and uptime appropriate to...  ...and performance.Build creative engineering solutions to operations and infrastructure...  ...:Bachelor's degree in Computer Science, a related technical...  ...in a strategic way.Site Reliability Engineering (SRE)... 
    Senior

    Google

    Mountain View, CA
    1 day ago
  • $222k - $300.5k

     ...TeamIntuit's Infrastructure and Site Reliability organization owns the...  ...The Fintech Platform Systems Engineering team builds and operates the...  ...The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering...  ...AWS cloud infrastructure (compute, networking, storage,... 
    Senior
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    4 days ago
  • $207.4k - $259.2k

     ...We’re seeking exceptional engineers, operators and builders to...  ...experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our...  ..., and security of our core systems and services. You will...  ...team.Bachelor's degree in Computer Science, Engineering, or a... 
    Senior
    Permanent employment
    Local area
    Visa sponsorship
    Night shift

    Archer Aviation

    San Jose, CA
    4 days ago
  •  ...hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give...  ...from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...provisioning, upgrades, and scaling.Own the reliability, performance, and security of... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $145k - $165k

     ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
    Senior

    Bolt Graphics, Inc.

    Sunnyvale, CA
    1 day ago
  • $192.4k - $275.8k

     ...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines...  ...this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    2 days ago
  •  ...Oracle Cloud Infrastructure (OCI) seeks a Senior Principal Engineer to lead the design and implementation of reliability validation for OCI control plane services, focusing on a high-performance, low-level systems approach. You will mentor engineers, define validation... 
    Senior

    Oracle

    Santa Clara, CA
    1 day ago
  • $200k - $230k

     ...looking for an experienced Power Electronics Controls and Firmware Engineer with more than seven years of hands-on expertise in developing...  ...and controllers.Validate design concepts through both computer simulations and hands-on laboratory testing.Create automated test... 
    Senior

    ChargePoint

    Campbell, CA
    2 days ago
  • $165k - $280k

     ...with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARLINK) At SpaceX we’re leveraging our experience in...  ...-region environment Manage petabyte scale bare metal compute clusters Closely collaborate with engineers across all programs... 
    Senior
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    InvestedintheMission

    Palo Alto, CA
    2 days ago
  • $203.45k - $344.3k

     ...learning, and smart connectivity.As a core member of our AI Infrastructure team...  .... We look forward to building a reliable, observable, and cost-effective data...  ...QualificationsBachelor's degree or higher in Computer Science, Software Engineering, Artificial Intelligence, or related... 
    Senior
    Full time
    Overseas

    XPENG Motors

    Santa Clara, CA
    2 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building...  ...anomaly detection.Partner with senior leaders and engineers across Cloud, Platform...  ...engineering roles.BS or MS degree in Computer Science or a related technical field... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Staff Site Reliability Engineer - Compute Core Engineering. Be the first to apply!