Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$81.1k - $187k

Oracle

Job Description

OCI Incident Response is the first line of defense in maintaining the high availability of Oracle’s cloud. We minimize customer-impacting events by making them shorter, less frequent, and less impactful through large-scale incident management. We are at the forefront of reducing event duration by leveraging our operational experience, knowledge of best practices, and ability to develop tools that automate incident management.

Description

We are looking for a Senior Site Reliability Engineer to join our OCI team. This role is part of a globally distributed team responsible for detecting, triaging, and mitigating OCI service-impacting events as quickly as possible. You will be part of one of these regional teams and will be responsible for minimizing the downtime of OCI services. You will achieve this by delivering excellent major incident management and operating systems with high scalability, performance, and security that help prevent incidents from occurring.

Oracle’s Cloud is state-of-the-art and constantly evolving. When issues arise, your team will respond within minutes to ensure customer impact is minimized. This role will expose you to the inner workings of OCI’s systems and organization. You will interact with and influence leaders across Oracle and drive broad, cross-organization programs aimed at iteratively improving OCI-wide service availability. We are an agile team with significant impact. If you want to be part of a fast-moving team breaking new ground, we would love to speak with you!

We are looking for candidates who are flexible to work AMER shift hours (9:30 AM to 5:30 PM PST) on a rotating roster, including occasional weekends and public holidays.

Career Level - IC3

Responsibilities

Responsibilities:

  • Solve complex problems related to infrastructure cloud services and automate common tasks to ensure continuous availability with minimal human intervention.

  • Command and coordinate SMEs and service leaders to restore services as quickly as possible during major incidents, while keeping accurate and timely data on the progress of such incidents.

  • Utilize a deep understanding of cloud computing design patterns and their dependencies to mitigate complex major incidents.

  • Embed a methodical approach to troubleshoot large, complex, interconnected systems used in incident detection and orchestration.

  • Document pertinent information related to incidents that aids process improvement, identifies deviations, and enables the creation of an incident knowledge base.

  • Monitor and evaluate high-level service and infrastructure dashboards, taking action to address identified anomalies.

  • Identify opportunities and take ownership of automation and/or continuous improvement of incident management process steps and best practices.

  • Define and document the technical architecture of large-scale distributed systems.

  • Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services.

  • Be responsible for the design and delivery of the mission-critical stack, with a focus on security, resiliency, scalability, and performance.

  • Partner with development teams to define operational requirements for product roadmaps.

  • Articulate the technical characteristics of services and technology areas, and guide development teams to engineer and add premier capabilities to the Oracle Cloud service portfolio.

  • Act as the ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs).

Minimum Qualifications:

Bachelor’s degree or higher in Computer Science or relevant work experience..

  • 3+ years’ experience in Site Reliability Engineering, DevOps, or System Engineering.

  • Must have public cloud operations experience (e.g., AWS, Azure, GCP, OCI).

  • Extensive experience with Major Incident Management in a cloud-based environment.

  • Demonstrate clear understanding of automation and orchestration principles.

  • Experience having worked in at least one modern object-oriented programming language.

  • Experience with professional software engineering standard methodologies such as Agile project management, coding standards, code reviews, source control management, build processes, testing, and operations.

  • Familiarity with infrastructure automation tools such as Chef, Ansible, Jenkins, Terraform

  • Excellent expertise with several of following technologies: Infrastructure-as-a-Service, CI/CD systems, Docker, RESTful APIs, log analysis tools, debugging tools

#LI-AH4

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $81,100 to $187,000 per annum. May be eligible for bonus and equity.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

  1. Medical, dental, and vision insurance, including expert medical opinion

  2. Short term disability and long term disability

  3. Life insurance and AD&D

  4. Supplemental life insurance (Employee/Spouse/Child)

  5. Health care and dependent care Flexible Spending Accounts

  6. Pre-tax commuter and parking benefits

  7. 401(k) Savings and Investment Plan with company match

  8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

  9. 11 paid holidays

  10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

  11. Paid parental leave

  12. Adoption assistance

  13. Employee Stock Purchase Plan

  14. Financial planning and group legal

  15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC3

About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on click.appcast.io or by calling View phone number on click.appcast.io in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Boston, MA vacancy
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Cambridge, MA
    3 days ago
  • $160k - $200k

     ...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    4 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    5 days ago
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Senior
    Work experience placement
    Work at office
    Remote work

    Axon

    Boston, MA
    1 day ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that...  ...maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering... 
    Senior
    Local area
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    5 days ago
  • $121.4k - $218.6k

     ...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner...  ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling... 
    Senior
    Work experience placement
    Work at office

    Akamai

    Boston, MA
    1 day ago
  •  ...commuting distance of one of our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team, you will operate the...  ...candidate is someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and automating cloud-based... 
    Senior
    Full time

    Federal Reserve Bank of Boston

    Boston, MA
    11 hours ago
  • $134.25k - $214.8k

     ...real change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability... 
    Senior
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Boston, MA
    3 days ago
  • $139k - $257.55k

     ...Individual Contributor The Challenge The Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe... 
    Senior
    Temporary work
    Local area
    Remote work
    Relocation

    Adobe

    Waltham, MA
    3 days ago
  • $135k - $165k

     ...tap into global manufacturing capacity.Xometry is seeking a Site Reliability Engineer II to join our Site Reliability Engineering (SRE)...  ...statements and drive them to completion with guidance from senior engineers.Write clean, efficient, and well-documented code... 
    Flexible hours

    Xometry

    Boston, MA
    3 days ago
  • $130k - $150k

     ...technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are...  ...career mentoring and performance coaching from an assigned senior colleague. Additional leadership and collaboration opportunities... 
    Work at office
    Work from home
    3 days per week

    CRA International

    Boston, MA
    2 days ago
  •  ...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the...  ...crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background.... 
    Senior
    Full time
    Remote work
    Worldwide

    Mongodb

    Boston, MA
    a month ago
  • $160k - $200k

    Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware... 
    Local area
    Remote work

    QuEra Computing

    Boston, MA
    4 days ago
  •  ...mission-critical industries, helping partners move more quickly and reliably from algorithm to silicon. Our platform accelerates deployment...  .... The Roles We are looking for an experienced software engineer to help us build a new generation of transpilation tools... 
    Senior
    Full time
    Remote work
    Relocation package
    Flexible hours

    Code Metal

    Boston, MA
    11 hours ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Cambridge, MA
    1 day ago
  •  ...Site Reliability Engineer Cambridge, MA About Watershed Our vision is to become the leading biocomputing platform. The future of biology is in big data analysis, and we are on a mission to accelerate digital drug discovery with the Watershed platform. Watershed... 

    Watershed Informatics

    Cambridge, MA
    5 days ago
  •  ...ISEE is seeking an experienced Senior Software Engineer to join our team. The ideal candidate has several years of work experience, and have worked on complex, performance-critical code-bases. Role responsibilities include: - Support full software development life... 
    Senior
    Full time
    Work experience placement

    Isee

    Cambridge, MA
    11 hours ago
  • $160k - $225k

     ...Staff Site Reliability Engineer Cambridge, MA Manifold is the AI platform for life sciences, accelerating life-changing medicines to patients. Our products speed up workflows in areas from target identification and clinical development to market access and precision... 

    Manifold

    Cambridge, MA
    2 days ago
  •  ...Distributed Systems engineers at Datadog design, implement and run in production the foundational platforms powering our applications. Your data pipelines will ingest, store, analyze and query in real-time billions of events per second from companies all over the globe... 
    Senior
    Full time
    Work at office

    Datadog

    Boston, MA
    11 hours ago
  • $108k - $209k

     ...seeking an experienced, creative, and talented Principal / Senior Software Engineer. The ideal candidate will have a strong background in software...  .... Leverage AWS cloud infrastructure to build scalable, reliable, and efficient applications and AI-powered services. Uphold... 
    Senior

    Seres Therapeutics

    Cambridge, MA
    2 days ago
  • $150k - $215k

     ...individuals optimize their health, fitness, and recovery. As a Senior Software Engineer on the AI team, you will play a key role in building and...  ...is ideal for an engineer who is passionate about building reliable, scalable applications and thrives in a fast-paced,... 
    Senior
    Full time
    Work at office
    Relocation

    Whoop

    Boston, MA
    11 hours ago
  • $191k - $253k

     ...we want you to join Anduril’s Maritime Division and help us build the future of defense capability. About the Job Senior Software Engineers independently drive the delivery of a variety of embedded and/or safety critical software integrated in to our products. This... 
    Senior
    Full time
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    Anduril Industries

    Boston, MA
    11 hours ago
  •  ...mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocation(s): Waltham, MA | HybridAbout the RoleSr Site Reliability Engineer- Guardian of the products to ensuring systems are reliable, scalable, and efficient... 
    Senior
    Ongoing contract
    Full time
    Temporary work
    Work experience placement
    Worldwide

    SS&C Technologies

    Waltham, MA
    3 days ago
  • $191k - $253k

     ...we want you to join Anduril’s Maritime Division and help us build the future of defense capability. ABOUT THE JOB Senior Software Engineers independently drive the delivery of a variety of software integrated in to our products. This includes autonomy, simulation... 
    Senior
    Full time
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    Anduril

    Boston, MA
    11 hours ago
  •  ...Company Description Zenith Talent specializes in staffing professional positions in Information Technology, Engineering, Marketing, Sales, Finance, HR and Operations. We have the knowledge and skills to supply candidates that fit perfectly in your organization. As... 
    Senior
    Full time

    Zenith Talent

    Boston, MA
    11 hours ago
  •  ...supports  Fortune 500 companies and platform engineers across 100+ countries . Crossplane has...  ...Learn more at  . Upbound is hiring a Senior Software Engineer to help us build and...  ...team, you will help us scale Upbound to reliably support thousands of control planes,... 
    Senior
    Full time
    Remote work
    Worldwide

    Upbound

    Boston, MA
    11 hours ago
  • $150k - $195k

     ...empowers members to perform at a higher level through a deeper understanding of their bodies and daily lives.WHOOP is seeking a Senior Reliability Engineer to lead the charge in ensuring our hardware products deliver a consistent, high-reliability experience for members. In... 
    Senior
    Full time
    Work at office
    Relocation

    WHOOP

    Boston, MA
    1 day ago
  •  ...ensuring our systems scale securely and reliably is core to this mission.The AI Devex...  ...underpinning business acceleration.As a Senior Platform Engineer on the AI Devex team, you will play a...  ...of experience in DevOps, Platform, Site Reliability, CloudEngineering, or Backend... 
    Senior
    Work at office
    Relocation

    WHOOP

    Boston, MA
    1 day ago
  • $191k - $253k

     ...Anduril is dedicated to building the products and systems that empower our engineering teams to develop, test, and deploy cutting-edge defense technology with unparalleled efficiency and reliability. We design and implement sophisticated automated systems that enhance... 
    Senior
    Full time
    Work experience placement
    Immediate start

    Anduril

    Boston, MA
    11 hours ago
  •  ...and as we absorb other repositories into the monorepo. As a senior software engineer on the team, you will own projects from start to finish,...  ...build, test and packaging tools that are simpler and more reliable to use. Push performance and cost efficiency at scale, raising... 
    Senior
    Full time
    Work at office

    Datadog

    Boston, MA
    11 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!