Principal Engineer, Cloud Site Reliability Engineering
$272k - $431.25kNVIDIA
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Autonomous Vehicles to cater to their infrastructure needs. These cloud services provide almost half a million automated jobs per day on thousands of servers helping with the efficiency of thousands of NVIDIA's software engineers worldwide. The cloud hosts various machines and devices with operating systems like Windows, Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra Processors. It delivers unified CI/CD solutions and cloud-based software development. Are you passionate about distributed infrastructure and looking for sophisticated, critical issues, ready to build the next generation of cloud services, design creative solutions, mine through data to uncover real problems and fix them?What you'll be doing:Serve as an SRE Architect part of GPU Private Cloud team used by thousands of NVIDIANs globally for interactive development, centralized CI/CD, and QA testing.Evaluating, identifying and developing software solutions to optimize critical software development workflows across various organizations within NVIDIA.Architecting, implementing, and supporting end-to-end CI/CD system using open-source and NVIDIA proprietary software.Customer (NVIDIA Internal development teams) onboarding to Private cloud infrastructure with a good discovery of the use case and available solutions within the cloud.Identify performance bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.Leading software development projects and technically direct a team of brilliant engineers and guide them to provide efficient and impactful solutions.Looking for problems within software systems and resolving the issuesCraft and implement critical metrics using various analytics methods and dashboards.What we need to see:BS or MS in Electrical Engineering, Computer Science, or relevant field (or equivalent experience).15+ years of systems software development including at least 1 year dedicated to developing/exploring AI.Experience of maintaining cloud infrastructure and highly available production environment.Strong programming and software development skills in JAVA, Python, Shell-script along with good understanding of distributed systems and REST APIs.Experience in working with SQL/NoSQL database systems such as MySQL, Cassandra, MongoDB or Elasticsearch.Excellent knowledge and working experience with Docker containers and Virtual Machines.Good background of Cloud technologies like: OpenStack, Docker, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, Kafka.Ability to work across organizational boundaries effectively to improve alignment and productivity between teams in a multi-national, multi-time-zone corporate environment.Ways to stand out from the crowd:Depth in AI, Machine Learning and Deep Learning algorithms and techniques.Strong collaborative and interpersonal skills, with a consistent record of guiding and influencing others in dynamic environments.Experience developing large-scale software systems using modular architecture under real-time performance requirements.Background in designing high-performance, scalable software systems with a strong focus on hardware cost optimization.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 9, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$248k - $396.75k
...environment, where NVIDIANs are inspired to excel and make a profound global impact.NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build...SuggestedFull time$210.6k - $305.1k
...own. Powered by AI and an unmatched set of cloud, internet and enterprise network... ...: You have led a distributed team of 5+ engineers, can demonstrate strong technical vision... ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees...SuggestedFull timeTemporary workLocal areaFlexible hours$168k - $264.5k
...outstanding? NVIDIA's Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA team. As an SRE at NVIDIA... ...across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.Author, test, and activate shared and non-shared...SuggestedFull time$168k - $270.25k
...deploy and run an AI data center. We take great pride in providing excellent, comprehensive support to our customers! Sr Site Reliability Engineer in this role will significantly impact and contribute to the overall success of both external customers running their clusters...SuggestedFull timeWorldwide$124k - $271.2k
What You Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security...SuggestedFull timeWork at officeRemote work- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers... ...our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational...Work at officeLocal areaWork from homeFlexible hours
$174k - $252k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance...Full timeWorldwide- Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens... ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building... ...Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE,...Work at officeLocal areaWork from homeFlexible hours
- ...s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is designed for a builder -...Full timeWork at office2 days per week
$230k - $250k
...foundation for autonomous networking, giving engineers and AI agents the ability to know the... ...and secure networks across every major cloud and vendor environment.Global leaders... ...been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep...Night shift$168k - $270.25k
...artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing... ...(HPC) storage solutions while harnessing the power of cloud computing. You will be responsible for crafting and deploying...Full time$128.6k - $184.9k
...infrastructure that powers our global cloud platform. As a team of six engineers distributed across the US, Canada,... ...with a strong focus on automation, reliability, and operational excellence. We are... ...7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours$267k - $356k
Lambda, The Superintelligence Cloud, is a leader in AI cloud... ...currently Tuesday.Lambda's Storage Engineering team is the backbone behind... ...in the industry, which means reliability and performance aren't just goals... ...across new and existing sites using tools such as Ansible,...Work experience placementWork at officeLocal areaWork from homeFlexible hours$168k - $270.25k
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers...Full time$101k - $161k
...industry leader in data-driven, client-to-cloud networking for large data center,... ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation... ...You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s...$147k - $210k
...product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e.g., style... ..., and troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when you treat operations...$90k - $180k
...people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale,... ...evolving business demands, including distributed systems and cloud deployments in Azure. Work closely with software...Remote work- Lambda, The Superintelligence Cloud, is a leader in AI cloud... ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...networking teams to improve service reliability and deployment... ...rotationYouHave 5+ years of experience in Site Reliability Engineering,...Work at officeLocal areaWork from homeFlexible hours
- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually... ...to optimize operations, increase efficiency in our use of cloud resources and our developer’s time, harden security posture...Flexible hours
$160k - $240k
...millions of times a day - quickly, reliably, and securely. Any time you... ...at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our... ...continuous improvement across our cloud-native environments.What you...Full time$166k - $244k
Overview Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have...Full time$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation...Full time$146.7k - $339.3k
...available for this positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to work on our hybrid... ..., you will patch and maintain thousands of physical and cloud systems worldwide. To streamline operations, you will develop...Full timeWork at officeRemote workWorldwideShift workWeekend work$122.5k - $175k
...leveraging the world’s largest security data lake to power our cloud-native Zero Trust Exchange platform. This innovation... ...the future of cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role going into the San...Full timeWork at officeLocal area3 days per week$207k - $300k
...of solutions to enhance the reliability of systems that support F1.Scale... ...teams.Engage in software engineering on services written in Java,... ...technical field.Experience in a Site Reliability Engineering role.... ...systems. SRE ensures that Google Cloud's services—both our...$222k - $300.5k
...OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational... .... The Fintech Platform Systems Engineering team builds and operates the AWS-based... ...blameless postmortems.Build and scale AWS cloud infrastructure (compute, networking, storage...WorldwideShift work$150k - $195k
...do and are proud of our work to secure clouds and container environments for thousands... ...team is growing, and we are looking for engineers with passion for automation. You will help... ...teams to improve the scalability and reliability of internal processes. Participate in an...Full timeWorldwide$141k - $208k
...About ClickHouse Recognized on the 2025 Forbes Cloud 100 list, ClickHouse is one of the most... ...committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building...Local areaRemote workHome officeFlexible hours- ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior... ...using GitHub Actions, Flux, Kustomize Design and implement cloud solutions, build MLOps on cloud AWS Data science model...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Engineer, Cloud Site Reliability Engineering. Be the first to apply!
- senior principal engineer Santa Clara, CA
- director data engineering Santa Clara, CA
- data center chief engineer Santa Clara, CA
- director quality engineering Santa Clara, CA
- director of product engineering Santa Clara, CA
- chief design engineer Santa Clara, CA
- chief engineer Santa Clara, CA
- senior chief engineer Santa Clara, CA
- principal engineer Santa Clara, CA
- senior director engineering Santa Clara, CA

