Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$81.1k - $187k

Oracle

Site Reliability Engineer 3

We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection and resolution of issues.

The engineer will work closely with development, infrastructure, security, and operations teams to monitor service health, troubleshoot production issues, participate in incident response, improve observability, and implement reliability best practices. This role also includes analyzing recurring failures, building automation, supporting deployments, and contributing to capacity planning, disaster recovery, and operational readiness.

Also works on number of different region/realm rollouts, deployments. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Performs data collection to maintain and optimize operations and reliability. Leverages knowledge to perform incident response and/or maintenance tasks. Provides health and performance reporting. Identifies opportunities for automation. Communicates about services and identifies and explains the potential impact of changes. Provides support for technology and document incidents. Experiments with new tools and assesses potential impact and develops knowledge of site reliability trends.

Responsibilities

Key Responsibilities Capacity Ingestion and Management:

  • Takes proactive steps to design and architect infrastructure and/or service according to terms for reliability and functionality.
  • Forecasts demands for infrastructure and responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads.
  • Collaborates with the software development team to develop infrastructures and features that are reliable and scalable according to deployment requirements.
  • Independently identifies opportunities for and drives prototyping (e.g., testing new applications or infrastructures, assisting in onboarding).

Incident and Service Lifecycle Management:

  • Performs data collection, triage, technical analysis, and redirection to maintain and optimize operations and infrastructure reliability.
  • Independently monitors services, maintains up-to-date knowledge of their performance, and documents their condition.
  • Leverages comprehensive knowledge to perform incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).
  • Provides health and performance reporting and takes appropriate actions based on trends in data.
  • May independently perform provisioning to support infrastructure, applications, and services.
  • May perform standard and non-standard decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.

Automation:

  • Identifies opportunities for automation and assesses potential benefits.
  • Develops automation tools or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures.
  • Independently conducts testing to ensure automation performs the task correctly and produces expected results.

Technical Communication and Guidance:

  • Communicates the scale, capacity, security, performance attributes, and requirements of services and technology within and sometimes beyond immediate team.
  • Identifies and explains the potential impact of infrastructure, feature, and tool changes, considering their impact on team operations.

Troubleshooting and Resolution:

  • Provides operational support for technology, escalating incidents and other standard and non-standard issues arising within Oracle services.
  • Participates in on-call shifts to address issues.
  • Resolves technical issues spanning various services, investigating and debugging products in order to reach SLOs (service level objectives).
  • Documents incidents and performs root cause analyses according to standard reporting methods.
  • Independently performs post-mortem procedures to prevent incident reoccurrence.

Innovation and Improvement:

  • Experiments with new tools and technologies to assess their potential impact on and improve infrastructure performance and reliability, ensuring adherence to security standards.
  • Independently identifies and executes improvements for performance bottlenecks and deployments to ensure efficient resource usage, speed, and scalability.
  • Develops knowledge of site reliability trends and shares new information with team members, management, and beyond to help others build, test, deploy and run services.
  • Performs standard and non-standard analyses and provides clear data on production to contribute to business development decisions (e.g., design changes).

Core Responsibilities Planning & Execution:

  • Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements.
  • Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.

Collaboration & Partnership:

  • Collaborates across teams to align on expectations and achieve shared objectives.
  • Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships.
  • Actively listens to diverse perspectives and asks questions to ensure understanding of others.

Problem Solving:

  • Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate.
  • Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors.
  • Contributes to knowledge sharing and best practices.

Continuous Learning:

  • Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices.
  • Seeks out and leverages feedback and training to improve skills.
  • Contributes to a culture of continuous learning and knowledge sharing with team members.

Continuous Improvement:

  • Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team.
  • Seeks input from team members on alternative approaches and methods for improving work.
Qualifications

Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US: Hiring Range in USD from: $81,100 to $187,000 per annum. May be eligible for bonus and equity. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following: 1. Medical, dental, and vision insurance, including expert medical opinion 2. Short term disability and long term disability 3. Life insurance and AD&D 4. Supplemental life insurance (Employee/Spouse/Child) 5. Health care and dependent care Flexible Spending Accounts 6. Pre-tax commuter and parking benefits 7. 401(k) Savings and Investment Plan with company match 8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation. 9. 11 paid holidays 10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours. 11. Paid parental leave 12. Adoption assistance 13. Employee Stock Purchase Plan 14. Financial planning and group legal 15. Voluntary benefits including auto, homeowner and pet insurance The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Required Skills
  • Bash Scripting
  • CI/CD Tools
  • Cloud Implementation
  • DevOps
  • Grafana Visualization
  • OCI - DevOps
  • Oracle Cloud Guard
  • Oracle Observability and Management
Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in San Francisco, CA vacancy
  •  ...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate... 
    Senior

    Alembic Technologies

    San Francisco, CA
    4 days ago
  •  ...Senior SRE Unify is building the first AI-native outbound platform where agents and...  ...Senior SRE, you'll tackle the scaling and reliability challenges that come with adding...  ...tracing, metrics, and alerting that give engineers clear visibility into system behavior and... 
    Senior

    Unify

    San Francisco, CA
    1 day ago
  • $181.69k - $213.75k

     ...Senior Site Reliability Engineer San Francisco, California; Santa Clara, California; Seattle, WA The Company You'll Join Carta connects founders, investors, and limited partners through world-class software, purpose-built for everyone in venture capital, private... 
    Senior
    Full time
    Work at office

    Carta

    San Francisco, CA
    4 days ago
  •  ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and... 
    Senior
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    1 day ago
  • $195k - $240k

     ...Senior Site Reliability Engineer San Francisco (Hybrid) At You.com, we are building the AI Search Infrastructure that powers modern AI systems. Our goal is to create the trusted knowledge layer that agents, applications, and enterprises rely on to retrieve real-... 
    Senior
    Full time
    Immediate start
    Remote work
    Work from home
    Flexible hours

    Y.O.U.

    San Francisco, CA
    4 days ago
  • $127k - $249k

    THE TEAM Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational...  ..., alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  • $117k - $209.33k

     ...Job Requisition ID # 26WD99273 Position Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products. As part of a... 
    Senior
    For contractors

    Autodesk

    San Francisco, CA
    5 days ago
  • $159.2k - $301.6k

     ...running Graphs on the cloud. In this reliability-focused role, you will own the availability...  .... You'll partner with the backend engineers building these APIs to make sure the system...  ...Science. ~5-10 years of experience in site reliability engineering, infrastructure,... 
    Senior
    Temporary work
    Local area
    Worldwide

    Adobe

    San Francisco, CA
    4 days ago
  • $148.5k - $223.9k

     ...duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce...  ...future of Salesforce. Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely... 
    Senior
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  • $166.9k - $225.9k

     ...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team...  ...What you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building... 
    Senior
    Work at office
    Immediate start
    Worldwide
    Monday to Friday
    Flexible hours

    Drata Inc

    San Francisco, CA
    1 day ago
  • $287k

     ...Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was...  ...expect us to hit our SLAs. What? We're looking for a Senior or Staff Site level Reliability Engineer as part of Infrastructure team to: Own... 
    Senior
    Contract work
    Work at office
    Remote work

    IVO Inc

    San Francisco, CA
    4 days ago
  • $170k - $190k

    Medrio is seeking a Senior Site Reliability Engineer in San Francisco, California. The role involves maintaining and supporting the Medrio Platform, troubleshooting issues, and implementing solutions. Candidates should have experience in Kubernetes, cloud services, and... 
    Senior
    Flexible hours

    Medrio

    San Francisco, CA
    3 days ago
  • A tech company specializing in AI is seeking a Site Reliability Engineer to ensure the reliability and observability of their production services. You will instrument services, develop SRE standards, and manage incident response. The ideal candidate has experience in AWS... 
    Senior
    Remote work
    Flexible hours

    You.com

    San Francisco, CA
    11 hours ago
  • $220k - $235k

     ...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements into assets. Contracts move faster, insights surface instantly, and agents push work forward, all with you in control. Whether you're buying or selling... 
    Senior
    Full time
    Contract work
    Work at office

    Ironclad Inc

    San Francisco, CA
    1 day ago
  • $160k - $300k

    Hebbia is looking for a Site Reliability Engineer to manage and improve critical production systems. This role requires writing production-quality code and collaborating with product engineering teams to enhance system reliability. Applicants should have 5+ years in software... 
    Senior

    Hebbia

    San Francisco, CA
    11 hours ago
  • $180k - $200k

    Parabola is seeking a Senior Site Reliability Engineer to join their team in San Francisco. In this role, you will monitor and improve software performance, maintain infrastructure, and collaborate with engineering teams. Candidates should have over 5 years of experience... 
    Senior

    Parabola

    San Francisco, CA
    11 hours ago
  • $210.8k - $272.8k

    Thumbtack is hiring for a Site Reliability Engineer to enhance the reliability and scalability of our services in San Francisco, CA. You'll design and support resilient systems, ensuring a smooth user experience. The ideal candidate will have extensive experience with AWS... 
    Senior

    Thumbtack

    San Francisco, CA
    11 hours ago
  • Early Warning is seeking a Staff Site Reliability Engineer to enhance application performance and resiliency while guiding development teams. This role involves designing automation and monitoring systems, improving scalability and availability, and participating in a 2... 
    Senior

    Early Warning

    San Francisco, CA
    11 hours ago
  • $181k - $263k

     ...and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering across LiveRamp's global infrastructure. This is a... 
    Senior
    Work from home
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    1 day ago
  • Location San Francisco, CA Employment Type Full time Department Engineering Who We Are Hyperbolic Labs is on a mission to democratize AI...  ...to redefine computing. About the Role We\'re seeking a Site Reliability Engineer to ensure Hyperbolic\'s GPU marketplace and AI infrastructure... 
    Senior
    Full time

    Hyperbolic

    San Francisco, CA
    11 hours ago
  • $170k - $220k

    Who We're Looking For We’re looking for a hands‑on, high‑agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end‑to‑end—managing daily releases, weekly deploys, and hotfixes—while also automating infrastructure... 
    Senior

    Supio

    San Francisco, CA
    1 day ago
  • $200k - $260k

    Senior Software Engineer, Site Reliability Engineer (SRE) Why Harvey At Harvey, we’re transforming how legal and professional services operate — not incrementally, but end‑to‑end. By combining frontier agentic AI, an enterprise‑grade platform, and deep domain expertise... 
    Senior
    Relocation package

    Harvey

    San Francisco, CA
    11 hours ago
  • $175k - $250k

     ...00.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance...  ...scalability, performance, and reliability across environments. What You’ll Do... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    1 day ago
  •  ...systems. It's designed so Stellar's ecosystem can make a real-world, lasting impact. About the Role SDF is looking for a Senior Site Reliability Engineer to help build and operate the foundation that powers our engineering teams. You'll ensure the reliability and... 
    Senior

    TechChain Talent

    San Francisco, CA
    11 hours ago
  • $210.8k - $272.8k

    About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user... 
    Senior
    Local area

    Thumbtack

    San Francisco, CA
    11 hours ago
  •  ...multimodality is critical for intelligence. This requires a massive, reliable, and performant GPU infrastructure that pushes the boundaries...  ...You Come In We are looking for a hands‑on, first‑principles engineer who is fluent in Linux, comfortable operating close to the... 
    Senior
    Work experience placement

    Luma AI

    San Francisco, CA
    2 days ago
  • CloudDevs works with fast-moving, venture-backed startups across the US. We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities. You will either be placed directly into one of our partner startups or added to our... 
    Senior
    Local area

    Breakout Tools

    San Francisco, CA
    1 day ago
  • $170k - $190k

    Position Medrio Senior Site Reliability Engineer Responsibilities Build, maintain, and support all environments which host the Medrio Platform Monitor environments for issues, configuring and building alerting/self‑healing of issues Update and maintain documentation... 
    Senior
    Temporary work
    Flexible hours

    Medrio

    San Francisco, CA
    3 days ago
  • $213k - $263k

     ...driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. Waymo's Software Reliability Engineers (SREs) are responsible for the stable operation of Waymo's fully autonomous systems and supporting infrastructure. As an SRE... 
    Senior
    Full time
    Remote work

    Waymo

    San Francisco, CA
    11 hours ago
  •  ...respond when things go wrong, helping every organization be more reliable. We do this by building an industry‑leading incident...  ...Build tools and automation to eliminate manual toil, improve engineering velocity and developer experience, and improve system reliability... 
    Senior
    Home office

    Rootly

    San Francisco, CA
    11 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!