Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Engineer, Cloud Site Reliability Engineering

$272k - $431.25k

NVIDIA

NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Autonomous Vehicles to cater to their infrastructure needs. These cloud services provide almost half a million automated jobs per day on thousands of servers helping with the efficiency of thousands of NVIDIA's software engineers worldwide. The cloud hosts various machines and devices with operating systems like Windows, Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra Processors. It delivers unified CI/CD solutions and cloud-based software development. Are you passionate about distributed infrastructure and looking for sophisticated, critical issues, ready to build the next generation of cloud services, design creative solutions, mine through data to uncover real problems and fix them?What you'll be doing:Serve as an SRE Architect part of GPU Private Cloud team used by thousands of NVIDIANs globally for interactive development, centralized CI/CD, and QA testing.Evaluating, identifying and developing software solutions to optimize critical software development workflows across various organizations within NVIDIA.Architecting, implementing, and supporting end-to-end CI/CD system using open-source and NVIDIA proprietary software.Customer (NVIDIA Internal development teams) onboarding to Private cloud infrastructure with a good discovery of the use case and available solutions within the cloud.Identify performance bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.Leading software development projects and technically direct a team of brilliant engineers and guide them to provide efficient and impactful solutions.Looking for problems within software systems and resolving the issuesCraft and implement critical metrics using various analytics methods and dashboards.What we need to see:BS or MS in Electrical Engineering, Computer Science, or relevant field (or equivalent experience).15+ years of systems software development including at least 1 year dedicated to developing/exploring AI.Experience of maintaining cloud infrastructure and highly available production environment.Strong programming and software development skills in JAVA, Python, Shell-script along with good understanding of distributed systems and REST APIs.Experience in working with SQL/NoSQL database systems such as MySQL, Cassandra, MongoDB or Elasticsearch.Excellent knowledge and working experience with Docker containers and Virtual Machines.Good background of Cloud technologies like: OpenStack, Docker, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, Kafka.Ability to work across organizational boundaries effectively to improve alignment and productivity between teams in a multi-national, multi-time-zone corporate environment.Ways to stand out from the crowd:Depth in AI, Machine Learning and Deep Learning algorithms and techniques.Strong collaborative and interpersonal skills, with a consistent record of guiding and influencing others in dynamic environments.Experience developing large-scale software systems using modular architecture under real-time performance requirements.Background in designing high-performance, scalable software systems with a strong focus on hardware cost optimization.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 9, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Principal Engineer, Cloud Site Reliability Engineering in Santa Clara, CA vacancy
  • $248k - $396.75k

     ...environment, where NVIDIANs are inspired to excel and make a profound global impact.NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  • $210.6k - $305.1k

     ...own. Powered by AI and an unmatched set of cloud, internet and enterprise network...  ...:  You have led a distributed team of 5+ engineers, can demonstrate strong technical vision...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    12 hours ago
  • $168k - $264.5k

     ...outstanding? NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA...  ...across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.Author, test, and activate shared and non-shared... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $168k - $270.25k

     ...deploy and run an AI data center. We take great pride in providing excellent, comprehensive support to our customers! ​Sr Site Reliability Engineer in this role will significantly impact and contribute to the overall success of both external customers running their clusters... 
    Suggested
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $124k - $271.2k

    What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security... 
    Suggested
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    2 days ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers...  ...our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 

    Google

    Sunnyvale, CA
    3 days ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    12 hours ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens...  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building...  ...Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE,... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is designed for a builder -... 
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    3 days ago
  • $230k - $250k

     ...foundation for autonomous networking, giving engineers and AI agents the ability to know the...  ...and secure networks across every major cloud and vendor environment.Global leaders...  ...been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep... 
    Night shift

    Forward Networks

    Santa Clara, CA
    3 days ago
  • $168k - $270.25k

     ...artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing...  ...(HPC) storage solutions while harnessing the power of cloud computing. You will be responsible for crafting and deploying... 
    Full time

    Nvidia

    Santa Clara, CA
    12 hours ago
  • $128.6k - $184.9k

     ...infrastructure that powers our global cloud platform. As a team of six engineers distributed across the US, Canada,...  ...with a strong focus on automation, reliability, and operational excellence. We are...  ...7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    CISCO Systems

    Santa Clara, CA
    2 days ago
  • $267k - $356k

    Lambda, The Superintelligence Cloud, is a leader in AI cloud...  ...currently Tuesday.Lambda's Storage Engineering team is the backbone behind...  ...in the industry, which means reliability and performance aren't just goals...  ...across new and existing sites using tools such as Ansible,... 
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    12 hours ago
  • $168k - $270.25k

    NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $101k - $161k

     ...industry leader in data-driven, client-to-cloud networking for large data center,...  ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation...  ...You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s... 

    Arista Networks

    Santa Clara, CA
    1 day ago
  • $147k - $210k

     ...product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e.g., style...  ..., and troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when you treat operations... 

    Google

    Sunnyvale, CA
    3 days ago
  • $90k - $180k

     ...people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale,...  ...evolving business demands, including distributed systems and cloud deployments in Azure. Work closely with software... 
    Remote work

    Abbott

    Sunnyvale, CA
    12 hours ago
  • Lambda, The Superintelligence Cloud, is a leader in AI cloud...  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...networking teams to improve service reliability and deployment...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering,... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    12 hours ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually...  ...to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture... 
    Flexible hours

    Sumo Logic

    San Jose, CA
    1 day ago
  • $160k - $240k

     ...millions of times a day - quickly, reliably, and securely. Any time you...  ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our...  ...continuous improvement across our cloud-native environments.What you... 
    Full time

    Fiserv

    Sunnyvale, CA
    2 days ago
  • $166k - $244k

    Overview Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have... 
    Full time

    Google

    Sunnyvale, CA
    3 days ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $146.7k - $339.3k

     ...available for this positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid...  ..., you will patch and maintain thousands of physical and cloud systems worldwide. To streamline operations, you will develop... 
    Full time
    Work at office
    Remote work
    Worldwide
    Shift work
    Weekend work

    Zoom

    San Jose, CA
    3 days ago
  • $122.5k - $175k

     ...leveraging the world’s largest security data lake to power our cloud-native Zero Trust Exchange platform. This innovation...  ...the future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    4 days ago
  • $207k - $300k

     ...of solutions to enhance the reliability of systems that support F1.Scale...  ...teams.Engage in software engineering on services written in Java,...  ...technical field.Experience in a Site Reliability Engineering role....  ...systems. SRE ensures that Google Cloud's services—both our... 

    Google

    San Jose, CA
    4 days ago
  • $222k - $300.5k

     ...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational...  .... The Fintech Platform Systems Engineering team builds and operates the AWS-based...  ...blameless postmortems.Build and scale AWS cloud infrastructure (compute, networking, storage... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    8 hours ago
  • $150k - $195k

     ...do and are proud of our work to secure clouds and container environments for thousands...  ...team is growing, and we are looking for engineers with passion for automation. You will help...  ...teams to improve the scalability and reliability of internal processes. Participate in an... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    2 days ago
  • $141k - $208k

     ...About ClickHouse Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most...  ...committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building... 
    Local area
    Remote work
    Home office
    Flexible hours

    GrabJobs

    San Jose, CA
    6 hours ago
  •  ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior...  ...using GitHub Actions, Flux, Kustomize Design and implement cloud solutions, build MLOps on cloud AWS Data science model... 
    Full time

    Saransh

    Sunnyvale, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Engineer, Cloud Site Reliability Engineering. Be the first to apply!