Senior Site Reliability Engineer
$81.1k - $187kOracle
Site Reliability Engineer 3
We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection and resolution of issues.
The engineer will work closely with development, infrastructure, security, and operations teams to monitor service health, troubleshoot production issues, participate in incident response, improve observability, and implement reliability best practices. This role also includes analyzing recurring failures, building automation, supporting deployments, and contributing to capacity planning, disaster recovery, and operational readiness.
Also works on number of different region/realm rollouts, deployments. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Performs data collection to maintain and optimize operations and reliability. Leverages knowledge to perform incident response and/or maintenance tasks. Provides health and performance reporting. Identifies opportunities for automation. Communicates about services and identifies and explains the potential impact of changes. Provides support for technology and document incidents. Experiments with new tools and assesses potential impact and develops knowledge of site reliability trends.
Responsibilities
Key Responsibilities Capacity Ingestion and Management:
- Takes proactive steps to design and architect infrastructure and/or service according to terms for reliability and functionality.
- Forecasts demands for infrastructure and responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads.
- Collaborates with the software development team to develop infrastructures and features that are reliable and scalable according to deployment requirements.
- Independently identifies opportunities for and drives prototyping (e.g., testing new applications or infrastructures, assisting in onboarding).
Incident and Service Lifecycle Management:
- Performs data collection, triage, technical analysis, and redirection to maintain and optimize operations and infrastructure reliability.
- Independently monitors services, maintains up-to-date knowledge of their performance, and documents their condition.
- Leverages comprehensive knowledge to perform incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).
- Provides health and performance reporting and takes appropriate actions based on trends in data.
- May independently perform provisioning to support infrastructure, applications, and services.
- May perform standard and non-standard decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.
Automation:
- Identifies opportunities for automation and assesses potential benefits.
- Develops automation tools or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures.
- Independently conducts testing to ensure automation performs the task correctly and produces expected results.
Technical Communication and Guidance:
- Communicates the scale, capacity, security, performance attributes, and requirements of services and technology within and sometimes beyond immediate team.
- Identifies and explains the potential impact of infrastructure, feature, and tool changes, considering their impact on team operations.
Troubleshooting and Resolution:
- Provides operational support for technology, escalating incidents and other standard and non-standard issues arising within Oracle services.
- Participates in on-call shifts to address issues.
- Resolves technical issues spanning various services, investigating and debugging products in order to reach SLOs (service level objectives).
- Documents incidents and performs root cause analyses according to standard reporting methods.
- Independently performs post-mortem procedures to prevent incident reoccurrence.
Innovation and Improvement:
- Experiments with new tools and technologies to assess their potential impact on and improve infrastructure performance and reliability, ensuring adherence to security standards.
- Independently identifies and executes improvements for performance bottlenecks and deployments to ensure efficient resource usage, speed, and scalability.
- Develops knowledge of site reliability trends and shares new information with team members, management, and beyond to help others build, test, deploy and run services.
- Performs standard and non-standard analyses and provides clear data on production to contribute to business development decisions (e.g., design changes).
Core Responsibilities Planning & Execution:
- Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements.
- Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.
Collaboration & Partnership:
- Collaborates across teams to align on expectations and achieve shared objectives.
- Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships.
- Actively listens to diverse perspectives and asks questions to ensure understanding of others.
Problem Solving:
- Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate.
- Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors.
- Contributes to knowledge sharing and best practices.
Continuous Learning:
- Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices.
- Seeks out and leverages feedback and training to improve skills.
- Contributes to a culture of continuous learning and knowledge sharing with team members.
Continuous Improvement:
- Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team.
- Seeks input from team members on alternative approaches and methods for improving work.
Qualifications
Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US: Hiring Range in USD from: $81,100 to $187,000 per annum. May be eligible for bonus and equity. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following:
- Medical, dental, and vision insurance, including expert medical opinion
- Short term disability and long term disability
- Life insurance and AD&D
- Supplemental life insurance (Employee/Spouse/Child)
- Health care and dependent care Flexible Spending Accounts
- Pre-tax commuter and parking benefits
- 401(k) Savings and Investment Plan with company match
- Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
- 11 paid holidays
- Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
- Paid parental leave
- Adoption assistance
- Employee Stock Purchase Plan
- Financial planning and group legal
- Voluntary benefits including auto, homeowner and pet insurance
Required Skills
- Bash Scripting
- CI/CD Tools
- Cloud Implementation
- DevOps
- Grafana Visualization
- OCI - DevOps
- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
- ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services... ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role...Senior
$190.8k - $267.1k
...helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser... ...team partners closely with Ads Engineering to improve reliability, scalability, operational... ...advertiser trust. We’re looking for a Senior Site Reliability Engineer to build, operate,...SeniorFor contractorsWork experience placement- ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves...Senior
$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB... ...a pivotal role in engineering the reliable, globally connected, multi-cloud... ...Role OverviewWe are seeking a talented Senior Site Reliability Engineer (SRE) with a strong...SeniorLocal areaRemote workWorldwideFlexible hours$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure...SeniorFlexible hours$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$117k - $209.33k
Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...SeniorFull timeFor contractors$165k - $227k
...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable...SeniorLocal areaWorldwideFlexible hours- ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ...ownership, working together to build scalable, reliable, and secure products that empower... ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work...SeniorTemporary workLocal areaWorldwide
$148.5k - $223.9k
...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations...SeniorFull timeWorldwideWeekend work$200.7k - $250.9k
...washed away in a flood in 1942, the Royal Engineers rebuilt it. Then it washed away again in... ...opportunities for improvements in reliability/observability/performance/preparedness and... ...candidate for the role: Has past Site Reliability Engineering or DevOps experience...Senior$200k - $240k
...systems across all product teams. You will collaborate closely with engineering leadership, product managers, and cross-functional teams to... ...and Helm ~ Understand the importance of performant and reliable systems ~ Education - Ideally looking for a B.A. / B.S. degree...SeniorWork at officeImmediate start3 days per week- ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'...Senior
$170k - $220k
...Senior Site Reliability Engineer Supio is a trusted AI platform purpose-built for law firms, reshaping how data drives impactful outcomes. Our innovative approach blends technology with deep legal expertise, making us a leader in our field. We go beyond surface-level...SeniorWork at officeRemote workFlexible hours$167.7k - $245.2k
...assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Your Impact As a Senior Site Reliability Engineer (SRE), you will lead the design and management of large-scale, highly available distributed systems, collaborating...SeniorFull timeTemporary workWork experience placementWork at officeLocal areaFlexible hours$167.7k - $245.2k
...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale...SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$232k - $319k
...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... ...enabled with self-serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and...SeniorPermanent employmentLocal areaWorldwideFlexible hours$250k
...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves...SeniorFull timeRemote work$55k - $187k
...Not Applicable Specialism IFS - Internal Firm Services - Other Management Level Senior Associate Job Description & Summary The Opportunity As a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability,...SeniorFull timeH1b$15k
...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage...SeniorWork at officeLocal areaRemote work$139.76k - $287.75k
...to grow their business.We are seeking a Senior Site ReliabilityEngineer to help operate,... ...will be instrumental in advancing the reliability, scalability, automation, observability... ...The ideal candidate is a highly hands-on engineer with strong production experience and a...SeniorWork at officeLocal areaRelocationRelocation package$106k - $130k
..., for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Role Summary The Senior Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and...SeniorHourly payFull timeImmediate startVisa sponsorshipWork visaFlexible hours$120k - $175k
...level of sports fandom. Ready to reimagine the DFS industry together? We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting-edge solutions and pushing the boundaries of what's possible...SeniorFull timeRemote workWork visaFlexible hours$262k - $364k
Lead a team of software/systems engineers on projects for users and be directly responsible for uptime.Own end-to-end availability... ...Engineering.Experience with machine learning infrastructure.Site Reliability Engineering (SRE) combines software and systems engineering...Senior$232.34k - $290.42k
...same: to make access to data as simple and reliable as electricity. With Fivetran, customer... ..., canonical and ready to query, with no engineering or maintenance required. We’re proud... ...integrate our teams, systems, and career sites. About the Role Fivetran and dbt...SeniorFull timeWork at officeRemote work$300k
...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring...SeniorPermanent employment- ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely... ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale...Permanent employmentWork experience placementWork at officeLocal area
$113.4k - $162k
...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at...Temporary work- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad business...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre San Francisco, CA
- site reliability engineer San Francisco, CA
- site reliability engineer remote San Francisco, CA
- senior operations technician San Francisco, CA
- senior operations associate San Francisco, CA
- senior cloud service delivery manager San Francisco, CA
- senior it service manager San Francisco, CA
- senior project engineer San Francisco, CA
- senior chief engineer San Francisco, CA
- sr operations manager San Francisco, CA


