Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer

$84.9k - $209.5k

Oracle

Job Description

As a Principal Site Reliability Engineer (IC4), you will be responsible for designing, building, and operating highly available, scalable, secure, and resilient cloud services. You will combine software engineering with infrastructure expertise to improve service reliability, operational efficiency, and developer productivity across large-scale distributed systems.

You will lead complex reliability initiatives, drive automation-first operational practices, and develop software solutions that eliminate manual toil. You will partner closely with software engineering, cloud infrastructure, security, and product teams to architect resilient platforms that meet aggressive availability, scalability, and performance objectives.

Success in this role requires deep expertise in distributed systems, cloud infrastructure, coding, automation, observability, incident management, and operational excellence. You will leverage modern AI technologies, machine learning, and intelligent automation to streamline operations, accelerate incident response, improve troubleshooting, and enable autonomous system management.

You are expected to be a technical leader who influences architecture, establishes engineering best practices, mentors other engineers, and drives continuous improvements across multiple services and organizations.

Responsibilities

Reliability Engineering & Service Ownership

  • Design, build, and operate highly available, scalable, and fault-tolerant cloud services that meet defined Service Level Objectives (SLOs) and Service Level Agreements (SLAs).

  • Lead architecture reviews to improve resiliency, scalability, observability, and operational readiness.

  • Forecast infrastructure growth, capacity requirements, and resource utilization while proactively mitigating operational risks.

  • Continuously improve platform reliability through engineering solutions rather than manual operational processes.

  • Define reliability standards, operational best practices, and service readiness criteria across multiple engineering teams.

Software Engineering, Automation & AI

  • Design and develop production-quality software, automation frameworks, and internal platforms using Python, Java, Go, or similar programming languages .

  • Build scalable automation to eliminate repetitive operational work, reduce manual intervention, and improve engineering productivity.

  • Develop APIs, microservices, and tooling that simplify infrastructure management, deployment, monitoring, and operational workflows.

  • Leverage Generative AI, Large Language Models (LLMs), AI agents, and intelligent automation to:

  • Automate routine operational tasks and runbooks.

  • Accelerate incident triage and root cause analysis.

  • Improve log analysis and anomaly detection.

  • Generate operational insights and recommendations.

  • Automate knowledge management and operational documentation.

  • Enhance developer productivity and self-service capabilities.

  • Identify opportunities to incorporate AI-driven operational intelligence into existing systems to improve efficiency, reliability, and scalability.

Infrastructure Engineering

  • Design and optimize cloud infrastructure supporting distributed services across multiple regions and availability domains.

  • Improve system resiliency through redundancy, automation, and infrastructure-as-code.

  • Build and maintain deployment pipelines, provisioning frameworks, and configuration management solutions.

  • Drive infrastructure standardization and platform modernization initiatives.

Observability & Operational Excellence

  • Design comprehensive monitoring, logging, tracing, and alerting strategies.

  • Build meaningful dashboards, health reporting, and service performance metrics.

  • Improve alert quality, reduce operational noise, and enhance system visibility.

  • Define and measure Service Level Indicators (SLIs), SLOs, and error budgets.

  • Continuously optimize operational processes using data-driven insights.

Incident Management & Reliability

  • Lead critical production incident response and act as a senior escalation point during major service events.

  • Drive root cause analysis, corrective actions, and post-incident reviews with a focus on long-term engineering improvements.

  • Develop automated remediation solutions to reduce Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).

  • Continuously improve operational readiness through disaster recovery testing, game days, and failure injection exercises.

Performance & Scalability

  • Identify performance bottlenecks across applications, infrastructure, storage, databases, and networking.

  • Drive optimization initiatives to improve system throughput, latency, efficiency, and cost.

  • Perform capacity planning and predictive scaling using historical trends and telemetry data.

  • Optimize resource utilization while maintaining service reliability and customer experience.

Technical Leadership

  • Serve as the technical leader for reliability engineering initiatives spanning multiple services and organizations.

  • Establish engineering standards, architectural patterns, and operational best practices.

  • Review system designs and influence technical decisions across engineering organizations.

  • Mentor engineers on software engineering, distributed systems, cloud technologies, operational excellence, and automation.

  • Drive engineering excellence through code reviews, design reviews, technical guidance, and knowledge sharing.

Innovation & Continuous Improvement

  • Evaluate emerging cloud technologies, AI capabilities, automation frameworks, and observability platforms.

  • Drive adoption of modern engineering practices including GitOps, Infrastructure as Code, continuous delivery, and policy-as-code.

  • Identify opportunities to simplify architecture, eliminate operational complexity, and improve platform reliability.

  • Champion engineering initiatives that increase service scalability, operational efficiency, and customer satisfaction.

Core Competencies

Planning & Execution

  • Lead complex, cross-functional technical initiatives from design through production deployment.

  • Balance reliability, scalability, performance, cost, and delivery priorities across multiple concurrent projects.

  • Drive execution with minimal direction while effectively managing technical risk and dependencies.

Collaboration & Influence

  • Partner with software engineering, cloud infrastructure, security, networking, and product organizations to deliver reliable services.

  • Influence technical direction across organizations through strong communication and technical leadership.

  • Build consensus among stakeholders with differing priorities and technical perspectives.

Problem Solving

  • Solve highly ambiguous and complex technical problems involving distributed systems at cloud scale.

  • Apply systematic debugging and data-driven analysis to identify root causes and implement long-term engineering solutions.

  • Make sound technical decisions using engineering judgment and operational experience.

Continuous Learning

  • Stay current with advances in cloud computing, distributed systems, AI, software engineering, cybersecurity, and Site Reliability Engineering.

  • Evaluate emerging technologies and drive adoption where they provide measurable operational value.

  • Foster a culture of continuous learning, experimentation, and technical excellence.

Talent Development

  • Mentor engineers and contribute to their technical growth through coaching and knowledge sharing.

  • Participate in hiring, interviewing, and technical assessments to build high-performing engineering teams.

  • Lead by example through engineering excellence, ownership, and customer-focused decision making.

Preferred Technical Skills

  • Strong software engineering experience with Python and/or Java (Go is a plus).

  • Experience building production automation, APIs, services, and developer tooling.

  • Cloud platforms (OCI, AWS, Azure, or GCP).

  • Kubernetes, Docker, container orchestration, and service mesh technologies.

  • Infrastructure as Code (Terraform, Ansible, Helm, Pulumi).

  • CI/CD platforms and DevOps practices.

  • Observability platforms such as Prometheus, Grafana, OpenTelemetry, ELK, Splunk, or Datadog.

  • Distributed systems, networking, Linux, storage, databases, and performance tuning.

  • AI-assisted operations, LLM integrations, automation frameworks, and intelligent operational tooling.

  • Experience operating mission-critical, internet-scale production systems with stringent availability and reliability requirements.

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $84,900 to $209,500 per annum. May be eligible for bonus and equity.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

  1. Medical, dental, and vision insurance, including expert medical opinion

  2. Short term disability and long term disability

  3. Life insurance and AD&D

  4. Supplemental life insurance (Employee/Spouse/Child)

  5. Health care and dependent care Flexible Spending Accounts

  6. Pre-tax commuter and parking benefits

  7. 401(k) Savings and Investment Plan with company match

  8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

  9. 11 paid holidays

  10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

  11. Paid parental leave

  12. Adoption assistance

  13. Employee Stock Purchase Plan

  14. Financial planning and group legal

  15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC4

About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on click.appcast.io or by calling View phone number on click.appcast.io in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in Austin, TX vacancy
  • Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate...  ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,... 
    Suggested
    Work at office
    Local area

    Realtor.com

    Austin, TX
    1 day ago
  •  ...Description:About the Role: We are looking for a Senior SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering... 
    Suggested
    Full time

    Dimensional Fund Advisors

    Austin, TX
    17 hours ago
  •  ...across multiple clouds and regions while partnering with network engineers, systems architects, and game studio developers. This is an ownership role: driving technical direction, influencing reliability from architecture review through production operation, and closing... 
    Suggested

    Gearbox Software

    Austin, TX
    2 days ago
  •  ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS... 
    Suggested
    Temporary work
    Casual work
    Worldwide

    TeamViewer

    Austin, TX
    14 hours ago
  •  ...Dimensional leverages the rapidly evolving state of the art to engineer scalable, innovative, and research driven solutions to improve...  ...each of the developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including Python toolchains (... 
    Suggested
    Full time
    Local area

    Dimensional Fund Advisors

    Austin, TX
    1 day ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    2 days ago
  •  ...Principal Site Reliability Engineer About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we power shipping logic and checkout optimization... 
    Principal
    Full time
    Work at office

    ShipperHQ

    Austin, TX
    a month ago
  • $152k - $241.5k

     ...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (...  ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,... 
    Full time

    Nvidia

    Austin, TX
    5 days ago
  •  ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical... 
    Full time
    Work at office

    The Charles Schwab Corporation

    Austin, TX
    2 days ago
  • $165k - $241.4k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    Austin, TX
    4 days ago
  • $127k - $249k

     ...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas...  ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    3 days ago
  • $196k - $269.5k

    Senior Principal AI Agent EngineerThe Software Engineering team delivers next-generation software application enhancements and new products for a changing world. Working at the cutting edge, we design and develop software for platforms, peripherals, applications and diagnostics... 
    Principal

    Dell Technologies

    Austin, TX
    1 day ago
  • $98.58k - $138.02k

     ...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company...  ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,... 
    Full time
    Work at office

    Restaurant 365

    Austin, TX
    4 days ago
  •  ...Artificial Intelligence at Schwab. We are an integrated product, engineering, strategy and risk team, all based in San Francisco. We help...  ...the most exciting areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts for cutting-... 
    Full time

    The Charles Schwab Corporation

    Austin, TX
    4 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    5 days ago
  •  ...innovators who want to make an impact on the world of technology.Cadence Design Systems Inc. is looking for a motivated DevOps Sr Principal Software Engineer to work with us in Austin, Texas. At Cadence, we hire and develop leaders and innovators who want to impact the world of... 
    Principal
    Full time

    Cadence Design Systems

    Austin, TX
    4 days ago
  •  ...resources via automation to enhance efficiency and consistency. Our engineers collaborate closely with internal product teams and customer...  ...- 6+ years of experience in a DevOps Engineer or Site Reliability Engineer - 5+ years of hands-on Terraform experience for building... 
    Local area

    Insight Global

    Austin, TX
    2 days ago
  • $167.18k - $203.61k

     ...remotely part of the weekTravel %NoWork ShiftJob DescriptionCox Automotive Corporate Services, LLCLEAD SITE RELIABILITY ENGINEERJob Description: Lead Site Reliability Engineer positions offered by Cox Automotive Corporate Services, LLC (Austin, Texas). Lead the... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cox Enterprises

    Austin, TX
    3 days ago
  •  ...Senior Site Reliability Engineer Come join a growing bank at the heart of the innovation, technology, green tech and life sciences space. We continue to expand our global footprint and our banking technology is at the core of everything we do. As a Senior Site Reliability... 

    Professional Recruiters

    Austin, TX
    4 days ago
  •  ...Senior Site Reliability Engineer Onsite - Austin, TX Apptronik is a human-centered robotics company developing AI-powered robots to support humanity in every facet of life. Our flagship humanoid robot, Apollo, is built to collaborate thoughtfully with people, starting... 
    Full time
    Local area

    Apptronik

    Austin, TX
    4 days ago
  • $141k

     ...be a part of our journey! About the role We are committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building and leading processes to ensure the reliability... 
    Local area
    Remote work
    Home office
    Flexible hours

    GrabJobs

    Austin, TX
    1 day ago
  • $196k - $364k

     ...innovators who want to make an impact on the world of technology.Cadence Design Systems is looking for a highly motivated hardware engineer to work with the Modus R&D engineering team in the Design-For-Test (DFT) IP business unit.As a member of the DFT R&D team, you... 
    Principal
    Full time

    Cadence Design Systems

    Austin, TX
    1 day ago
  • $110.7k - $171.8k

     ...components Participation in on-call rotation as a platform reliability escalation point Incident response, post-incident reviews,...  ..., and internal control requirements. Collaborate with engineering teams across the organization to influence platform adoption,... 
    Work experience placement
    Work at office
    Local area

    Visa

    Austin, TX
    4 days ago
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working...  ...they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA... 
    Principal
    Full time
    Remote work
    Shift work

    Nvidia

    Austin, TX
    2 days ago
  • Senior Principal Software Engineer (ServiceNow Information Architect)Be a part of a team that’s ensuring Dell Technologies' product integrity and customer satisfaction. Our IT Software Engineer team turns business requirements into technology solutions by designing, coding... 
    Principal

    Dell Technologies

    Round Rock, TX
    14 hours ago
  • $198k - $247.5k

     ...of the open source community, Cloudera advances digital transformation for the world’s largest enterprises.About Forward Deployed Engineering (FDE) at ClouderaCloudera’s Forward Deployed Engineering (FDE) function sits within the Applied AI organization, and is a... 
    Principal
    Full time
    Work from home
    Relocation

    Cloudera

    Austin, TX
    4 days ago
  • $130k - $180k

     ...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and...  ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to... 
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    GrabJobs

    Austin, TX
    2 days ago
  •  ...Automation & AI)Be a part of a team that’s ensuring Dell Technologies' product integrity and customer satisfaction. Our IT Software Engineer team turns business requirements into technology solutions by designing, coding and testing/debugging applications, as well as... 
    Principal

    Dell Technologies

    Round Rock, TX
    14 hours ago
  • $182.75k - $236.5k

    Senior Principal Systems Development Engineer Our customers’ system requirements are usually highly complex. Bringing together hardware and software...  ...proactive solution design, helping customers achieve reliable, scalable, and high‑performance data ecosystems.... 
    Principal
    Full time

    Dell Technologies

    Austin, TX
    2 days ago
  •  ...METRIX IT SOLUTIONS INC is seeking a senior database engineer to design, deploy, and manage multi-region CockroachDB clusters in production. The role focuses on high availability, data consistency, and scalable capacity planning for global deployments. You will monitor... 

    METRIX IT SOLUTIONS INC

    Austin, TX
    17 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!