Site Reliability Engineer
$104.4k - $171kAlibaba Cloud
The mission of the Cloud Intelligence Group SRE (Site Reliability Engineering) Team is to ensure the stability of production environments, enterprise-grade cloud data reliability, and service continuity for the Cloud Intelligence Group. Our greatest challenge lies in guaranteeing uninterrupted business operations for cloud-based customers and achieving availability that exceeds 99.99%.
Objectives of the Cloud Intelligence Group SRE Team
Our goal is to establish a systematic stability assurance framework that integrates technology and management, including but not limited to:
- 1. Developing stability standards and metrics
- Covering robust architecture, R&D quality, release management, production environment operations, and more.
- Embedding stability into Alibaba Cloud's technical R&D system.
- 2. Driving major stability governance campaigns
- Initiatives such as full-stack disaster recovery, phased change rollout, the 1-5-10 emergency response mechanism (1-minute alerting, 5-minute triage, 10-minute recovery), and financial-loss prevention.
- Rapidly and continuously mitigating stability risks.
- 3. Building a stability-focused technical platform
- Platform capabilities for unattended change management, red/blue team drills, emergency collaboration, risk and vulnerability inspection, and monitoring/alerting.
- Simplifying stability engineering through automation and tooling.
- 4. Executing production incident management
- Emergency response, cross-team coordination, root cause analysis, rapid recovery, and post-incident reviews to drive systemic improvements.
- 5. Ensuring stability for large-scale customer events
- Technical and operational support for critical activities such as Olympics and customer business peak periods.
- 6. On-call responsibilities
- Responding to customer issues within Service Level Agreement (SLA) timeframes, resolving problems proactively, and enhancing customer experience.
Responsibilities
The objective of the Cloud Intelligence Group's SRE team is to establish a systematic stability assurance framework that integrates technology and management, including but not limited to:
- 1. Daily operations and maintenance of applications, databases, and middleware, as well as troubleshooting and answering customer inquiries;
- 2. Collaborating with R&D to develop critical support plans based on customer business requirements during peak periods, including preparation during the standby period, on-duty support during critical periods, and post-standby review;
- A degree in Computer Science or related field
- 3 years of experience as a Site Reliability Engineer (SRE) or above.
- Proficiency with Linux environments or cloud infrastructure
- Exceptional system diagnostic and problem-solving skills
- Strong teamwork spirit and ability to work well under pressure
- In-depth understanding of Kubernetes or monitoring systems
- Expertise in programming languages such as Golang, Python, or Java
- Experience in designing and implementing distributed systems
The pay range for this position at commencement of employment is expected to be between $104,400 and $171,000/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.
If hired, employee will be in an "at-will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.
#J-18808-Ljbffr$160k - $240k
...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit... ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in...SuggestedFull time$170k - $200k
...We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance...SuggestedFull time- ...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity... ...looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background in...SuggestedWork experience placement
- ...Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You’ll Do (Day-to-Day) Own and manage our cloud infrastructure (GCP or AWS, on-prem). Build, maintain, and optimize Kubernetes clusters (including GPU-backed clusters...Suggested
$230k - $250k
...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change... ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"...SuggestedNight shift$132.6k - $214.5k
...As part of this role, you will collaborate closely with our engineering teams to develop innovative solutions that provide clear and... ...team to influence the operability of the product and ensure the reliability and availability of our services. Qualifications...Full timeWork at officeVisa sponsorshipWork visa$230k - $250k
...minds are shaping the future of network reliability, security, and AI‑ready operations. About... ...you will be building the reliability engineering function at Forward — defining how we... ...Looking For ~6+ years of experience in site reliability engineering, DevOps, or...Night shift- ...Investigate and resolve performance and reliability issues across application, infrastructure, database, Kubernetes, and Linux layers... ..., plan capacity, improve observability, and collaborate with engineering teams and business stakeholders. Requirements: Requires hands...
$145k - $165k
...: Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining...Work at officeImmediate start$65 - $85 per hour
...Talent is partnering with Nvidia, a global leader in computer graphics, PC gaming, and accelerated computing, to bring a Site Reliability Engineer (Contract) to the team based in Santa Clara, CA. This is a full‑time (W‑2) contract role. Pay ranges from $65/hr to $85...Full timeContract workWorldwide$150.4k - $277.6k
...Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple’s long... ...field with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role Advanced...RelocationDay shift- ...of Huobi globe spanning infrastructure. • Work with engineering teams to make sure new features and changes are deployed quickly... .... • Constantly improve our system performance and reliability through better tools, process and monitoring system. •...Worldwide
- ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the...Work at office
$150k - $195k
...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the... ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation....Full timeWorldwide$120k - $150k
...interested in working with the World's leading AI-first Quality Engineering Company? Ready to advance your career, team up with global... ...every day? Join us at QualityAI! We are looking for a Site Reliability Engineer to join our growing team in Riverwoods, IL United States...Casual workLocal areaFlexible hours$174k - $252k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...$262k - $364k
...and training AI infrastructure from SRE side, ensuring it is reliable, scalable, cost effective and performant, while working closely... ...qualifications:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and systems engineering...$195k - $285k
...purpose-built AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we build. This role builds and leads d-Matrix's Site Reliability Engineering function from the ground up, owning the...Remote work$175k - $265k
...Overviewd-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities, on-... .... This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab,...- ...Up to 25% Job ID: 1874 The Role The Platform Engineering team builds, secures, and operates scalable infrastructure... ...products with on‑premises components deployed at customer sites. The Site Reliability Engineering discipline keeps the platform stable and reliable...Work at officeRemote work
$152k - $287.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...such as Python, Go, Perl, or Ruby. ~ Mentored other engineers and influenced technical direction through design reviews, architecture...Full time$248k - $396.75k
...Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems with exceptional efficiency, resilience, and availability. It combines software and systems engineering practices with...$150k - $180k
..., environmental, and innovation outcomes. Role Verrus is looking for candidates to serve as software-focused Senior Site Reliability Engineer at Verrus. This is a full‑time position based out of the Mountain View, CA office. Verrus takes a very technology‑forward...Full timeWork at officeLocal areaFlexible hours$209.7k - $238.25k
...this particular role covers. ️ About our Vehicle Software Engineering Teams Our Vehicle Software Engineering team builds and operates... ...world. Within it, the Tools & Infrastructure team owns the reliability stack for vehicle software: observability, incident...Full timeWork at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work3 days per week$262k - $364k
...Senior Staff Software Engineer, Site Reliability Engineering, Workspace AI Google Sunnyvale, CA, USA X In most instances, this position requires in-person interviews as part of the hiring process. ~ Bachelor’s degree in Computer Science, a related field, or equivalent...- ...Own the architecture and design of reliable, scalable, cost-effective, and performant AI... ...master's degree in computer science or engineering is preferred. Key Skills Software... ...Machine Learning, Artificial Intelligence, Site Reliability Engineering, AI Infrastructure...
- ...Role :- Site Reliability Engineer (SRE) Infrastructure & Agentic Automation Location :- Santa Clara, CA (Hybrid) Work Authorization: USC/GC only Position Summary & Job Description:- Client is looking for an experienced Site Reliability Engineer (SRE) to...
$104.9k - $174.7k
...SRE role is responsible for improving the reliability, availability, performance, and... ...actions through completion.Follow up with engineering, development, security, support, and business... ...Qualifications5+ years of experience in Site Reliability Engineering, Systems Engineering...Full timeLocal area$255.7k - $300k
...Manager, Software Engineer, Site Reliability Engineering Share Manager, Software Engineer, Site Reliability Engineering ~ link Copy link corporate_fare Google place Sunnyvale, CA, USA Advanced Experience owning outcomes and decision making, solving ambiguous...Full timeWork at office$100k - $200k
...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Sunnyvale, CA
- site reliability engineer sre Sunnyvale, CA
- site safety Sunnyvale, CA
- on-site clinical research associate (traveling/remote) Sunnyvale, CA
- site services specialist Sunnyvale, CA
- construction site safety Sunnyvale, CA
- junior website developer Sunnyvale, CA
- site recruiter Sunnyvale, CA
- historic site Sunnyvale, CA
- IT site lead Sunnyvale, CA



