Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$91.4k - $187k

Oracle

Job Description

OCI Incident Response is the first line of defense in maintaining the high availability of Oracle's cloud. We minimize customer-impacting events by making them shorter, less frequent, and less impactful through large-scale incident management. We are at the forefront of reducing event duration by leveraging our operational experience, knowledge of best practices, and ability to develop tools that automate incident management.

Description

We are looking for a Senior Site Reliability Engineer to join our OCI team. This role is part of a globally distributed team responsible for detecting, triaging, and mitigating OCI service-impacting events as quickly as possible. You will be part of one of these regional teams and will be responsible for minimizing the downtime of OCI services. You will achieve this by delivering excellent major incident management and operating systems with high scalability, performance, and security that help prevent incidents from occurring.

Oracle's Cloud is state-of-the-art and constantly evolving. When issues arise, your team will respond within minutes to ensure customer impact is minimized. This role will expose you to the inner workings of OCI's systems and organization. You will interact with and influence leaders across Oracle and drive broad, cross-organization programs aimed at iteratively improving OCI-wide service availability. We are an agile team with significant impact. If you want to be part of a fast-moving team breaking new ground, we would love to speak with you!

We are looking for candidates who are flexible to work AMER shift hours (9:30 AM to 5:30 PM PST) on a rotating roster, including occasional weekends and public holidays.

Career Level - IC3

Responsibilities

Responsibilities:

  • Solve complex problems related to infrastructure cloud services and automate common tasks to ensure continuous availability with minimal human intervention.

  • Command and coordinate SMEs and service leaders to restore services as quickly as possible during major incidents, while keeping accurate and timely data on the progress of such incidents.

  • Utilize a deep understanding of cloud computing design patterns and their dependencies to mitigate complex major incidents.

  • Embed a methodical approach to troubleshoot large, complex, interconnected systems used in incident detection and orchestration.

  • Document pertinent information related to incidents that aids process improvement, identifies deviations, and enables the creation of an incident knowledge base.

  • Monitor and evaluate high-level service and infrastructure dashboards, taking action to address identified anomalies.

  • Identify opportunities and take ownership of automation and/or continuous improvement of incident management process steps and best practices.

  • Define and document the technical architecture of large-scale distributed systems.

  • Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services.

  • Be responsible for the design and delivery of the mission-critical stack, with a focus on security, resiliency, scalability, and performance.

  • Partner with development teams to define operational requirements for product roadmaps.

  • Articulate the technical characteristics of services and technology areas, and guide development teams to engineer and add premier capabilities to the Oracle Cloud service portfolio.

  • Act as the ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs).

Minimum Qualifications:

Bachelor's degree or higher in Computer Science or relevant work experience..

  • 3+ years' experience in Site Reliability Engineering, DevOps, or System Engineering.

  • Must have public cloud operations experience (e.g., AWS, Azure, GCP, OCI).

  • Extensive experience with Major Incident Management in a cloud-based environment.

  • Demonstrate clear understanding of automation and orchestration principles.

  • Experience having worked in at least one modern object-oriented programming language.

  • Experience with professional software engineering standard methodologies such as Agile project management, coding standards, code reviews, source control management, build processes, testing, and operations.

  • Familiarity with infrastructure automation tools such as Chef, Ansible, Jenkins, Terraform

  • Excellent expertise with several of following technologies: Infrastructure-as-a-Service, CI/CD systems, Docker, RESTful APIs, log analysis tools, debugging tools

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $91,400 to $187,000 per annum. May be eligible for bonus and equity.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

  1. Medical, dental, and vision insurance, including expert medical opinion

  2. Short term disability and long term disability

  3. Life insurance and AD&D

  4. Supplemental life insurance (Employee/Spouse/Child)

  5. Health care and dependent care Flexible Spending Accounts

  6. Pre-tax commuter and parking benefits

  7. 401(k) Savings and Investment Plan with company match

  8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

  9. 11 paid holidays

  10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

  11. Paid parental leave

  12. Adoption assistance

  13. Employee Stock Purchase Plan

  14. Financial planning and group legal

  15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC3

About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on click.appcast.io or by calling View phone number on click.appcast.io in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Boston, MA vacancy
  • $160k - $200k

     ...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    4 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    18 hours ago
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Senior
    Work experience placement
    Work at office
    Remote work

    Axon

    Boston, MA
    1 day ago
  • $166k - $220k

     ...failure. As such, it is critical that Anduril services are reliable and maintainable. This means that all services &...  ...ground systems & Kubernetes infrastructure.ABOUT THE JOBAs a Site Reliability Engineer on the Observability team, you will build & operate Anduril... 
    Senior
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Boston, MA
    18 hours ago
  • $127k - $249k

     ...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the...  ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background.... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    1 day ago
  • $168k - $200k

     ...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and... 
    Senior

    Datavant

    Boston, MA
    2 days ago
  • $160k - $200k

     ...Senior Site Reliability Engineer This role is located in Somerville, MA - We are a hybrid work environment and are in the office 3+ days/per week. Tulip, the leader in AI-native frontline operations, is helping companies around the world equip their workforce with... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interfaces

    Somerville, MA
    3 days ago
  • $140k - $210.9k

     ...position will be primarily on-site with residency commutable to...  ...DevOps backgrounds or software engineering backgrounds (e.g., Java...  ...interest in operating and improving reliability of distributed production...  ...Responsibilities As a Senior Engineer of the SRE / Production... 
    Senior
    Full time
    Temporary work
    Part time
    Work at office
    Shift work

    Federal Reserve Bank

    Boston, MA
    4 days ago
  • $121.4k - $218.6k

     ...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner...  ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling... 
    Senior
    Work experience placement
    Work at office

    Akamai

    Boston, MA
    1 day ago
  •  ...commuting distance of one of our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team, you will operate the...  ...candidate is someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and automating cloud-based... 
    Senior
    Full time

    Federal Reserve Bank of Boston

    Boston, MA
    1 day ago
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection... 
    Senior
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Boston, MA
    4 days ago
  •  ...Information Technology group delivers secure, reliable technology solutions that enable...  ...You Will Have in This RoleAs a Senior Application Support Engineer, you will help power DTCC's global...  ...processing and settlement.Leveraging Site Reliability Engineering (SRE) principles... 
    Senior
    Remote work
    Flexible hours

    DTCC- The Depository Trust & Clearing Corporation

    Boston, MA
    18 hours ago
  • $60 - $70 per hour

     ...Job Description- 100% REMOTE! This DevOps Automation Engineer role sits within a platform operations team and focuses on supporting...  ...a global, regulated MedTech context. The position emphasizes site reliability engineering and platform operations over CI/CD-heavy... 
    Senior
    Contract work
    Temporary work
    Remote work

    Actalent

    Boston, MA
    3 days ago
  • $139k - $257.55k

     ...Individual Contributor The Challenge The Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe... 
    Senior
    Temporary work
    Local area
    Remote work
    Relocation

    Adobe

    Waltham, MA
    3 days ago
  • The Depository Trust & Clearing Corporation (DTCC) seeks a Senior Application Support Engineer to ensure reliability and performance of its critical trade processing platforms. You will apply SRE principles, drive automation, and partner with global teams to support AWS... 
    Senior

    The Depository Trust & Clearing Corporation (DTCC)

    Boston, MA
    5 days ago
  • $130k - $150k

     ...technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are...  ...career mentoring and performance coaching from an assigned senior colleague. Additional leadership and collaboration opportunities... 
    Work at office
    Work from home
    3 days per week

    CRA International

    Boston, MA
    2 days ago
  • $160k - $200k

    Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware... 
    Local area
    Remote work

    QuEra Computing

    Boston, MA
    4 days ago
  •  ...mission-critical industries, helping partners move more quickly and reliably from algorithm to silicon. Our platform accelerates deployment...  .... The Roles We are looking for an experienced software engineer to help us build a new generation of transpilation tools... 
    Senior
    Full time
    Remote work
    Relocation package
    Flexible hours

    Code Metal

    Boston, MA
    1 day ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Cambridge, MA
    1 day ago
  •  ...Site Reliability Engineer Cambridge, MA About Watershed Our vision is to become the leading biocomputing platform. The future of biology is in big data analysis, and we are on a mission to accelerate digital drug discovery with the Watershed platform. Watershed... 

    Watershed Informatics

    Cambridge, MA
    18 hours ago
  •  ...efforts and improve strategic decision-making. As a member of our engineering team, you'll be working closely with other team members and our...  ...engineering teams can work and the tools they useLocation: on-site in BostonWe believe that it takes a diverse team to build the... 
    Senior

    Roberts Recruiting

    Boston, MA
    4 days ago
  •  ...ISEE is seeking an experienced Senior Software Engineer to join our team. The ideal candidate has several years of work experience, and have worked on complex, performance-critical code-bases. Role responsibilities include: - Support full software development life... 
    Senior
    Full time
    Work experience placement

    Isee

    Cambridge, MA
    1 day ago
  • $108k - $209k

     ...seeking an experienced, creative, and talented Principal / Senior Software Engineer. The ideal candidate will have a strong background in software...  .... Leverage AWS cloud infrastructure to build scalable, reliable, and efficient applications and AI-powered services. Uphold... 
    Senior

    Seres Therapeutics

    Cambridge, MA
    2 days ago
  • $148k - $185k

     ...the future together. The Crown Is Yours As a Lead Site Reliability Engineer, you'll set the reliability standard across our Infrastructure...  ...into clear, actionable insights that help teams and senior leaders make better decisions about reliability, risk, and... 
    Full time
    Immediate start

    DraftKings

    Boston, MA
    20 hours ago
  • $160k - $225k

     ...Staff Site Reliability Engineer Manifold is the AI platform for life sciences, accelerating life-changing medicines to patients. Our products speed up workflows in areas from target identification and clinical development to market access and precision medicine in the... 

    Manifold

    Cambridge, MA
    2 days ago
  • $150k - $215k

     ...individuals optimize their health, fitness, and recovery. As a Senior Software Engineer on the AI team, you will play a key role in building and...  ...is ideal for an engineer who is passionate about building reliable, scalable applications and thrives in a fast-paced,... 
    Senior
    Full time
    Work at office
    Relocation

    Whoop

    Boston, MA
    1 day ago
  • $130k - $140k

     ...mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocation(s): Waltham, MA | HybridAbout the RoleSr Site Reliability Engineer- Guardian of the products to ensuring systems are reliable, scalable, and efficient... 
    Senior
    Ongoing contract
    Full time
    Temporary work
    Work experience placement

    SS&C Technologies

    Waltham, MA
    3 days ago
  • $150k - $195k

     ...empowers members to perform at a higher level through a deeper understanding of their bodies and daily lives.WHOOP is seeking a Senior Reliability Engineer to lead the charge in ensuring our hardware products deliver a consistent, high-reliability experience for members. In... 
    Senior
    Full time
    Work at office
    Relocation

    WHOOP

    Boston, MA
    1 day ago
  • $138k - $252k

     ...scaled autonomy Solicit and incorporate feedback from end users of the APIs and implementations, and collaborate with adjacent engineering teams to help make the product vision a reality Develop and improve our APIs for commanding and controlling teams of... 
    Senior
    Full time
    Work experience placement
    Local area
    Relocation package
    Flexible hours

    Anduril Industries

    Boston, MA
    1 day ago
  • $140k - $160k

     ...About this role: Pickle is seeking a dynamic, driven Senior Software Engineer, Navigation, to enhance the speed and safety of our autonomous...  ...algorithms and capable of optimizing for performance and reliability. Detail-oriented, but with a system-level mindset.... 
    Senior
    Full time
    Work at office
    3 days per week

    Pickle Robot Company

    Charlestown, MA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!