Senior Site Reliability Engineer
$91.4k - $187kOracle
Job Description
OCI Incident Response is the first line of defense in maintaining the high availability of Oracle's cloud. We minimize customer-impacting events by making them shorter, less frequent, and less impactful through large-scale incident management. We are at the forefront of reducing event duration by leveraging our operational experience, knowledge of best practices, and ability to develop tools that automate incident management.
Description
We are looking for a Senior Site Reliability Engineer to join our OCI team. This role is part of a globally distributed team responsible for detecting, triaging, and mitigating OCI service-impacting events as quickly as possible. You will be part of one of these regional teams and will be responsible for minimizing the downtime of OCI services. You will achieve this by delivering excellent major incident management and operating systems with high scalability, performance, and security that help prevent incidents from occurring.
Oracle's Cloud is state-of-the-art and constantly evolving. When issues arise, your team will respond within minutes to ensure customer impact is minimized. This role will expose you to the inner workings of OCI's systems and organization. You will interact with and influence leaders across Oracle and drive broad, cross-organization programs aimed at iteratively improving OCI-wide service availability. We are an agile team with significant impact. If you want to be part of a fast-moving team breaking new ground, we would love to speak with you!
We are looking for candidates who are flexible to work AMER shift hours (9:30 AM to 5:30 PM PST) on a rotating roster, including occasional weekends and public holidays.
Career Level - IC3
Responsibilities
Responsibilities:
Solve complex problems related to infrastructure cloud services and automate common tasks to ensure continuous availability with minimal human intervention.
Command and coordinate SMEs and service leaders to restore services as quickly as possible during major incidents, while keeping accurate and timely data on the progress of such incidents.
Utilize a deep understanding of cloud computing design patterns and their dependencies to mitigate complex major incidents.
Embed a methodical approach to troubleshoot large, complex, interconnected systems used in incident detection and orchestration.
Document pertinent information related to incidents that aids process improvement, identifies deviations, and enables the creation of an incident knowledge base.
Monitor and evaluate high-level service and infrastructure dashboards, taking action to address identified anomalies.
Identify opportunities and take ownership of automation and/or continuous improvement of incident management process steps and best practices.
Define and document the technical architecture of large-scale distributed systems.
Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services.
Be responsible for the design and delivery of the mission-critical stack, with a focus on security, resiliency, scalability, and performance.
Partner with development teams to define operational requirements for product roadmaps.
Articulate the technical characteristics of services and technology areas, and guide development teams to engineer and add premier capabilities to the Oracle Cloud service portfolio.
Act as the ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs).
Minimum Qualifications:
Bachelor's degree or higher in Computer Science or relevant work experience..
3+ years' experience in Site Reliability Engineering, DevOps, or System Engineering.
Must have public cloud operations experience (e.g., AWS, Azure, GCP, OCI).
Extensive experience with Major Incident Management in a cloud-based environment.
Demonstrate clear understanding of automation and orchestration principles.
Experience having worked in at least one modern object-oriented programming language.
Experience with professional software engineering standard methodologies such as Agile project management, coding standards, code reviews, source control management, build processes, testing, and operations.
Familiarity with infrastructure automation tools such as Chef, Ansible, Jenkins, Terraform
Excellent expertise with several of following technologies: Infrastructure-as-a-Service, CI/CD systems, Docker, RESTful APIs, log analysis tools, debugging tools
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $91,400 to $187,000 per annum. May be eligible for bonus and equity.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
Medical, dental, and vision insurance, including expert medical opinion
Short term disability and long term disability
Life insurance and AD&D
Supplemental life insurance (Employee/Spouse/Child)
Health care and dependent care Flexible Spending Accounts
Pre-tax commuter and parking benefits
401(k) Savings and Investment Plan with company match
Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
11 paid holidays
Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
Paid parental leave
Adoption assistance
Employee Stock Purchase Plan
Financial planning and group legal
Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC3
About Us
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on click.appcast.io or by calling View phone number on click.appcast.io in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
$160k - $200k
...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident...SeniorTemporary workWork at officeLocal areaFlexible hours3 days per week$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$134.25k - $214.8k
...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed... ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,...SeniorWork experience placementWork at officeRemote work$166k - $220k
...failure. As such, it is critical that Anduril services are reliable and maintainable. This means that all services &... ...ground systems & Kubernetes infrastructure.ABOUT THE JOBAs a Site Reliability Engineer on the Observability team, you will build & operate Anduril...SeniorFull timeWork experience placementImmediate start$127k - $249k
...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background....SeniorLocal areaRemote workWorldwideFlexible hours$168k - $200k
...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and...Senior$160k - $200k
...Senior Site Reliability Engineer This role is located in Somerville, MA - We are a hybrid work environment and are in the office 3+ days/per week. Tulip, the leader in AI-native frontline operations, is helping companies around the world equip their workforce with...SeniorTemporary workWork at officeLocal areaFlexible hours3 days per week$140k - $210.9k
...position will be primarily on-site with residency commutable to... ...DevOps backgrounds or software engineering backgrounds (e.g., Java... ...interest in operating and improving reliability of distributed production... ...Responsibilities As a Senior Engineer of the SRE / Production...SeniorFull timeTemporary workPart timeWork at officeShift work$121.4k - $218.6k
...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner... ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling...SeniorWork experience placementWork at office- ...commuting distance of one of our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team, you will operate the... ...candidate is someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and automating cloud-based...SeniorFull time
$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...SeniorTemporary workImmediate startFlexible hoursShift work- ...Information Technology group delivers secure, reliable technology solutions that enable... ...You Will Have in This RoleAs a Senior Application Support Engineer, you will help power DTCC's global... ...processing and settlement.Leveraging Site Reliability Engineering (SRE) principles...SeniorRemote workFlexible hours
$60 - $70 per hour
...Job Description- 100% REMOTE! This DevOps Automation Engineer role sits within a platform operations team and focuses on supporting... ...a global, regulated MedTech context. The position emphasizes site reliability engineering and platform operations over CI/CD-heavy...SeniorContract workTemporary workRemote work$139k - $257.55k
...Individual Contributor The Challenge The Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe...SeniorTemporary workLocal areaRemote workRelocation- The Depository Trust & Clearing Corporation (DTCC) seeks a Senior Application Support Engineer to ensure reliability and performance of its critical trade processing platforms. You will apply SRE principles, drive automation, and partner with global teams to support AWS...Senior
$130k - $150k
...technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are... ...career mentoring and performance coaching from an assigned senior colleague. Additional leadership and collaboration opportunities...Work at officeWork from home3 days per week$160k - $200k
Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware...Local areaRemote work- ...mission-critical industries, helping partners move more quickly and reliably from algorithm to silicon. Our platform accelerates deployment... .... The Roles We are looking for an experienced software engineer to help us build a new generation of transpilation tools...SeniorFull timeRemote workRelocation packageFlexible hours
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Site Reliability Engineer Cambridge, MA About Watershed Our vision is to become the leading biocomputing platform. The future of biology is in big data analysis, and we are on a mission to accelerate digital drug discovery with the Watershed platform. Watershed...
- ...efforts and improve strategic decision-making. As a member of our engineering team, you'll be working closely with other team members and our... ...engineering teams can work and the tools they useLocation: on-site in BostonWe believe that it takes a diverse team to build the...Senior
- ...ISEE is seeking an experienced Senior Software Engineer to join our team. The ideal candidate has several years of work experience, and have worked on complex, performance-critical code-bases. Role responsibilities include: - Support full software development life...SeniorFull timeWork experience placement
$108k - $209k
...seeking an experienced, creative, and talented Principal / Senior Software Engineer. The ideal candidate will have a strong background in software... .... Leverage AWS cloud infrastructure to build scalable, reliable, and efficient applications and AI-powered services. Uphold...Senior$148k - $185k
...the future together. The Crown Is Yours As a Lead Site Reliability Engineer, you'll set the reliability standard across our Infrastructure... ...into clear, actionable insights that help teams and senior leaders make better decisions about reliability, risk, and...Full timeImmediate start$160k - $225k
...Staff Site Reliability Engineer Manifold is the AI platform for life sciences, accelerating life-changing medicines to patients. Our products speed up workflows in areas from target identification and clinical development to market access and precision medicine in the...$150k - $215k
...individuals optimize their health, fitness, and recovery. As a Senior Software Engineer on the AI team, you will play a key role in building and... ...is ideal for an engineer who is passionate about building reliable, scalable applications and thrives in a fast-paced,...SeniorFull timeWork at officeRelocation$130k - $140k
...mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocation(s): Waltham, MA | HybridAbout the RoleSr Site Reliability Engineer- Guardian of the products to ensuring systems are reliable, scalable, and efficient...SeniorOngoing contractFull timeTemporary workWork experience placement$150k - $195k
...empowers members to perform at a higher level through a deeper understanding of their bodies and daily lives.WHOOP is seeking a Senior Reliability Engineer to lead the charge in ensuring our hardware products deliver a consistent, high-reliability experience for members. In...SeniorFull timeWork at officeRelocation$138k - $252k
...scaled autonomy Solicit and incorporate feedback from end users of the APIs and implementations, and collaborate with adjacent engineering teams to help make the product vision a reality Develop and improve our APIs for commanding and controlling teams of...SeniorFull timeWork experience placementLocal areaRelocation packageFlexible hours$140k - $160k
...About this role: Pickle is seeking a dynamic, driven Senior Software Engineer, Navigation, to enhance the speed and safety of our autonomous... ...algorithms and capable of optimizing for performance and reliability. Detail-oriented, but with a system-level mindset....SeniorFull timeWork at office3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer Boston, MA
- site reliability engineer sre Boston, MA
- srs distribution Boston, MA
- senior operations coordinator Boston, MA
- senior associate architect Boston, MA
- senior dynamics crm developer Boston, MA
- senior application security Boston, MA
- senior account director Boston, MA
- sr hr business partner Boston, MA
- senior supervisor Boston, MA


