Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer

Oracle

Principal Site Reliability Engineer (IC4)

As a Principal Site Reliability Engineer (IC4), you will be responsible for designing, building, and operating highly available, scalable, secure, and resilient cloud services. You will combine software engineering with infrastructure expertise to improve service reliability, operational efficiency, and developer productivity across large-scale distributed systems.

You will lead complex reliability initiatives, drive automation-first operational practices, and develop software solutions that eliminate manual toil. You will partner closely with software engineering, cloud infrastructure, security, and product teams to architect resilient platforms that meet aggressive availability, scalability, and performance objectives.

Success in this role requires deep expertise in distributed systems, cloud infrastructure, coding, automation, observability, incident management, and operational excellence. You will leverage modern AI technologies, machine learning, and intelligent automation to streamline operations, accelerate incident response, improve troubleshooting, and enable autonomous system management.

You are expected to be a technical leader who influences architecture, establishes engineering best practices, mentors other engineers, and drives continuous improvements across multiple services and organizations.

Reliability Engineering & Service Ownership
  • Design, build, and operate highly available, scalable, and fault-tolerant cloud services that meet defined Service Level Objectives (SLOs) and Service Level Agreements (SLAs).
  • Lead architecture reviews to improve resiliency, scalability, observability, and operational readiness.
  • Forecast infrastructure growth, capacity requirements, and resource utilization while proactively mitigating operational risks.
  • Continuously improve platform reliability through engineering solutions rather than manual operational processes.
  • Define reliability standards, operational best practices, and service readiness criteria across multiple engineering teams.
Software Engineering, Automation & AI
  • Design and develop production-quality software, automation frameworks, and internal platforms using Python, Java, Go, or similar programming languages.
  • Build scalable automation to eliminate repetitive operational work, reduce manual intervention, and improve engineering productivity.
  • Develop APIs, microservices, and tooling that simplify infrastructure management, deployment, monitoring, and operational workflows.
  • Leverage Generative AI, Large Language Models (LLMs), AI agents, and intelligent automation to:
    • Automate routine operational tasks and runbooks.
    • Accelerate incident triage and root cause analysis.
    • Improve log analysis and anomaly detection.
    • Generate operational insights and recommendations.
    • Automate knowledge management and operational documentation.
    • Enhance developer productivity and self-service capabilities.
  • Identify opportunities to incorporate AI-driven operational intelligence into existing systems to improve efficiency, reliability, and scalability.
Infrastructure Engineering
  • Design and optimize cloud infrastructure supporting distributed services across multiple regions and availability domains.
  • Improve system resiliency through redundancy, automation, and infrastructure-as-code.
  • Build and maintain deployment pipelines, provisioning frameworks, and configuration management solutions.
  • Drive infrastructure standardization and platform modernization initiatives.
Observability & Operational Excellence
  • Design comprehensive monitoring, logging, tracing, and alerting strategies.
  • Build meaningful dashboards, health reporting, and service performance metrics.
  • Improve alert quality, reduce operational noise, and enhance system visibility.
  • Define and measure Service Level Indicators (SLIs), SLOs, and error budgets.
  • Continuously optimize operational processes using data-driven insights.
Incident Management & Reliability
  • Lead critical production incident response and act as a senior escalation point during major service events.
  • Drive root cause analysis, corrective actions, and post-incident reviews with a focus on long-term engineering improvements.
  • Develop automated remediation solutions to reduce Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
  • Continuously improve operational readiness through disaster recovery testing, game days, and failure injection exercises.
Performance & Scalability
  • Identify performance bottlenecks across applications, infrastructure, storage, databases, and networking.
  • Drive optimization initiatives to improve system throughput, latency, efficiency, and cost.
  • Perform capacity planning and predictive scaling using historical trends and telemetry data.
  • Optimize resource utilization while maintaining service reliability and customer experience.
Technical Leadership
  • Serve as the technical leader for reliability engineering initiatives spanning multiple services and organizations.
  • Establish engineering standards, architectural patterns, and operational best practices.
  • Review system designs and influence technical decisions across engineering organizations.
  • Mentor engineers on software engineering, distributed systems, cloud technologies, operational excellence, and automation.
  • Drive engineering excellence through code reviews, design reviews, technical guidance, and knowledge sharing.
Innovation & Continuous Improvement
  • Evaluate emerging cloud technologies, AI capabilities, automation frameworks, and observability platforms.
  • Drive adoption of modern engineering practices including GitOps, Infrastructure as Code, continuous delivery, and policy-as-code.
  • Identify opportunities to simplify architecture, eliminate operational complexity, and improve platform reliability.
  • Champion engineering initiatives that increase service scalability, operational efficiency, and customer satisfaction.
Core Competencies
Planning & Execution
  • Lead complex, cross-functional technical initiatives from design through production deployment.
  • Balance reliability, scalability, performance, cost, and delivery priorities across multiple concurrent projects.
  • Drive execution with minimal direction while effectively managing technical risk and dependencies.
Collaboration & Influence
  • Partner with software engineering, cloud infrastructure, security, networking, and product organizations to deliver reliable services.
  • Influence technical direction across organizations through strong communication and technical leadership.
  • Build consensus among stakeholders with differing priorities and technical perspectives.
Problem Solving
  • Solve highly ambiguous and complex technical problems involving distributed systems at cloud scale.
  • Apply systematic debugging and data-driven analysis to identify root causes and implement long-term engineering solutions.
  • Make sound technical decisions using engineering judgment and operational experience.
Continuous Learning
  • Stay current with advances in cloud computing, distributed systems, AI, software engineering, cybersecurity, and Site Reliability Engineering.
  • Evaluate emerging technologies and drive adoption where they provide measurable operational value.
  • Foster a culture of continuous learning, experimentation, and technical excellence.
Talent Development
  • Mentor engineers and contribute to their technical growth through coaching and knowledge sharing.
  • Participate in hiring, interviewing, and technical assessments to build high-performing engineering teams.
  • Lead by example through engineering excellence, ownership, and customer-focused decision making.
Preferred Technical Skills
  • Strong software engineering experience with Python and/or Java (Go is a plus).
  • Experience building production automation, APIs, services, and developer tooling.
  • Cloud platforms (OCI, AWS, Azure, or GCP).
  • Kubernetes, Docker, container orchestration, and service mesh technologies.
  • Infrastructure as Code (Terraform, Ansible, Helm, Pulumi).
  • CI/CD platforms and DevOps practices.
  • Observability platforms such as Prometheus, Grafana, OpenTelemetry, ELK, Splunk, or Datadog.
  • Distributed systems, networking, Linux, storage, databases, and performance tuning.
  • AI-assisted operations, LLM integrations, automation frameworks, and intelligent operational tooling.
  • Experience operating mission-critical, internet-scale production systems with stringent availability and reliability requirements.
Qualifications

Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in Seattle, WA vacancy
  • $92.25k - $140k

    Job Title: Senior Site Reliability Engineer (SRE) Job Description We're Concentrix. The intelligent transformation partner. Solution-focused. Tech-powered. Intelligence-fueled. The global technology and services leader that powers the world’s best brands, today and... 
    Suggested
    Full time
    Work at office
    Immediate start
    3 days per week

    CNX

    Bellevue, WA
    3 hours ago
  •  ...services and event driven architectures, efficient and reliable message brokering systems become more crucial than ever...  ...grow and learn at multiple levels . Career Level - IC4 Principal Platform Software Engineer. Lead platform projects crossing multiple teams, evolving... 
    Principal
    Full time
    Flexible hours

    Oracle

    Seattle, WA
    2 days ago
  •  ...production at global scale. * Foundational Frameworks: Spearhead the engineering of new container runtimes and distributed frameworks to power...  ...and writes, regardless of scale. * A journaling service that reliably and efficiently records sequential, immutable logs (journal... 
    Principal
    Full time
    Worldwide
    Flexible hours

    Oracle

    Seattle, WA
    10 hours ago
  • $173.5k - $234.7k

     ...and real-time sync mechanisms that deliver the right data to each consumer within strict latency and reliability targets. The role requires applying seasoned engineering judgment to solve complex systems challenges and to make recommendations on architecture, technology... 
    Principal
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Flexible hours

    T-Mobile

    Bellevue, WA
    3 hours ago
  • $163.62k - $212.71k

     ...maintaining the tools, platforms, and processes that improve our engineering teams' productivity and streamline the software...  ...Responsibilities: We are seeking a seasoned and strategic Lead/Principal Site Reliability Engineer to drive the reliability, scalability, and... 
    Principal
    Full time
    Part time
    Work experience placement
    Work at office
    Local area
    Immediate start
    Remote work
    Work from home
    Flexible hours
    Shift work
    3 days per week
    1 day per week

    iSpot

    Bellevue, WA
    9 days ago
  • $163.62k - $212.71k

     ...to contribute to the success of something great. As a Principal Software Development Engineer, you will be a hands-on leader and strategic thinker,...  ...measurement and data processing platform, ensuring scalability, reliability, and high performance. ● Lead the design of core... 
    Principal
    Full time
    Part time
    Work experience placement
    Work at office
    Local area
    Remote work
    Work from home
    Flexible hours
    3 days per week
    1 day per week

    iSpot

    Bellevue, WA
    2 days ago
  •  ...for controlplanes and data planes. We are hoping to enhance engineering efficiency by concentrating our expertise on building low level...  ...deliver a database validation framework that will ensure the reliability of databases being used by critical tier-0 OCI services. This... 
    Principal
    Worldwide
    Flexible hours

    Ll Oefentherapie

    Seattle, WA
    4 days ago
  • $166k - $244k

    # Senior Software Engineer, Site Reliability EngineeringGoogle • onsite • 601 N 34th St, Seattle, WA 98103, USA • full\_timePay: USD 166000.00 - USD 244000.00 / unspecifiedBusinesses of all shapes and sizes rely on Google’s unparalleled advertising solutions to help them... 
    Temporary work

    Epic Games

    Seattle, WA
    2 days ago
  • $160k - $250k

     ...public clouds when the right fit. As we continue to commercialize our machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering for our customers. Our ideal candidate is someone who is able... 

    Hive

    Seattle, WA
    4 days ago
  • $127k - $249k

    THE TEAM Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational...  ..., alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Seattle, WA
    5 days ago
  •  ...Sr. Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA. We offer IT solutions across the disciplines of program/project management, applications development, infrastructure, Cyber security, and enterprise... 
    Local area

    Comtech LLC

    Seattle, WA
    4 days ago
  • $134.25k - $214.8k

     ...change. Constantly grow as you work hard for a mission that matters at a company where you matter. Your Impact As a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Seattle, WA
    5 days ago
  •  ...looking for people like you. As a Senior Principal Architect at JPMorganChase within...  ...reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain...  ...training or certification on software engineering concepts and 10+ years applied experience... 
    Principal

    JPMorgan Chase & Co.

    Seattle, WA
    4 days ago
  •  ...of the world's most influential companies. As a Senior Principal Software Engineer at JPMorganChase within the CDAO AI/ML Data Platforms Team,...  ...These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan,... 
    Principal

    Socket.dev

    Seattle, WA
    10 hours ago
  • $264.1k - $369.74k

    Sr Principal Software Engineering - Enterprise Technology page is loaded## Sr Principal Software Engineering - Enterprise Technologylocations: Greater...  ...for:WA applicants is $264,103.00 - $369,743.85**Other site ranges may differ****Culture Statement****Export Control Regulations... 
    Principal
    Permanent employment
    Temporary work
    Local area
    Relocation

    Blue Origin LLC

    Seattle, WA
    4 days ago
  •  ...to our people, and the incredible connections we get to make in every community we are in. About this team Site Reliability Engineering We are looking for a motivated engineer to join the Foundations team which is responsibility for observability and... 

    Kaav Inc.

    Seattle, WA
    3 days ago
  •  ...Site Reliability Engineer Comtech is a woman-owned small business founded in 1998 and headquartered in Reston, VA. We offer IT solutions across the disciplines of program/project management, applications development, infrastructure, Cyber security, and enterprise content... 

    Comtech LLC

    Seattle, WA
    4 days ago
  • $100k - $170k

     ...who wants to own systems, not just watch them. You'll take real surface area: the automation and tooling other engineers depend on, and the reliability of production services running AI and GPU workloads at scale. You'll sit in the incident rotation, and you'll... 
    Flexible hours
    Shift work

    Nscale

    Seattle, WA
    5 days ago
  • Technical/Functional Skills Windows Servers, Digital: Microsoft Azure Windows Powershell, Digital: DevOps Roles & Responsibilities Windows Server 2012 -2019 Administration Microsoft Azure Azure AAD DFSR, DHCP DNS, KMS, WSUS TCP/IP Hyper-V High Availability Clusters ...

    The Dignify Solutions, LLC

    Bellevue, WA
    3 days ago
  • Job Title Required Skills: CHEF experience - Must have most critical Azure Cloud – experience - Must have most critical AKS- Azure Kubernetes services - Must have most critical Kubernetes - Must have most critical NoSQL DB – Cassandra / Mongo DB ...

    Syntricate Technologies

    Seattle, WA
    5 days ago
  •  ...The Role Join us in revolutionizing the lending landscape. SoFi is seeking enthusiastic Principal Software Engineers who are ready to lead the technical and strategic evolution of our financial services platform in support of our goals that put our members in control... 
    Principal
    Full time
    Temporary work
    Work experience placement

    Sofi

    Seattle, WA
    10 hours ago
  •  ...Position Overview SingleStore is seeking a Site Reliability Engineer to help optimize and scale our managed service offering across all three major cloud providers. In this role, you will be at the intersection of leading technology trends - A highly performant distributed... 
    Worldwide

    SingleStore

    Seattle, WA
    1 day ago
  • $166k - $244k

    Senior Software Engineer, Site Reliability Engineering Google Seattle, WA, USA Mid Experience driving progress, solving problems, and mentoring more junior team members; deeper expertise and applied knowledge within relevant area. Minimum qualifications: Bachelor’s degree... 
    Full time

    Google

    Seattle, WA
    10 hours ago
  • $196k - $294k

     ...ABOUT THE TEAM ~ We are seeking a Principal Software Engineer to lead architectural and strategic initiatives for our ArsenalOS Forge platform. You'll be pivotal in evolving Forge's capabilities, driving integration and scalability, and maturing its architecture for... 
    Principal
    Full time
    Work experience placement
    Local area
    Relocation package

    Anduril Industries

    Seattle, WA
    10 hours ago
  • $194k - $267k

     ...something more than once, automate it" and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    Bellevue, WA
    1 day ago
  • Overview Tech Talent Specialist | USA | Around the whole SDLC Senior / Principal Software Engineer Some roles keep you fixing features no one notices. Others let you build tech that actually stops real cyber threats before they cause chaos. (This role is the second one)... 
    Principal
    Home office

    Tact.ai

    Seattle, WA
    4 days ago
  • $200k

     ...systems, focusing on services that ensure Synthesia is secure, reliable, and scalable for our largest customers. You will...  ...clear steps that can be delivered and validated iteratively. Engineers within Synthesia are empowered to contribute heavily to product... 
    Principal
    Full time
    For contractors
    Local area
    Visa sponsorship

    Synthesia Limited

    Seattle, WA
    4 days ago
  • $280k - $330k

     ...This role is data‑centric software engineering at a very high bar: you own substantial...  ...across the team. What You’ll Do As a Principal Software Development Engineer, you bring...  ...Institutionalize observability and reliability engineering for data: SLIs/SLOs where appropriate... 
    Principal

    Auger Services Inc

    Bellevue, WA
    4 days ago
  • $200k - $250k

     ...excel at building microservices and managing cloud infrastructure? If so, we have an exciting opportunity for you as a Principal Backend Software Engineer at Trase Systems. As a Principal Software Engineer specializing in Backend Engineering, you will play a pivotal role... 
    Principal
    Remote work

    CloudDevs

    Seattle, WA
    4 days ago
  •  ...Senior Director, Principal Gifts About the Company Philanthropic organization supporting Indigenous culture & individuals Industry Non-Profit Organization Management Type Non Profit Founded 2017 Employees 11-50 Categories... 
    Principal

    Confidential

    Seattle, WA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!