Lead Principal Site Reliability Engineer
$96.3k - $264.1kOracle Corporation
Job Description Serves as a consultant and leads the design and architecture of infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full ownership of forecasting of demands and responding to capacity needs. Owns collaborations with software development teams to develop reliable and scalable infrastructures. Recommends methods for performing data collection to maintain and optimize operations and reliability. Oversees incident response and/or maintenance tasks. Provides strategic, future-oriented health and performance reporting. Contributes to strategies for automation and reviews the development and implementation of automation. Provides expert-level communication about services and anticipates, analyzes, and explains the impact of changes, considering strategic goals. Serves as a role model in providing support for technology and reviews documentation for accuracy. Leads the implementation of innovative tools and provides expertise in site reliability trends. Responsibilities Key Responsibilities Reliability Strategy and Technical Leadership
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC5 About Us Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing or by calling 1- in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
- Define and drive the site reliability engineering strategy for large-scale, distributed, and business-critical platforms.
- Establish reliability standards, engineering practices, and operational readiness requirements across multiple teams.
- Serve as a senior technical authority for system reliability, scalability, resilience, performance, and production operations.
- Influence architecture and design decisions to ensure systems are supportable, observable, fault tolerant, and capable of meeting availability objectives.
- Identify systemic reliability risks and lead cross-functional initiatives to address them.
- Provide technical direction and mentorship to site reliability engineers, software engineers, platform engineers, and operations teams.
- Lead technical reviews and promote consistent engineering practices across the organization.
- Define and implement service-level indicators, service-level objectives, error budgets, and operational health metrics.
- Develop comprehensive monitoring, logging, tracing, alerting, and observability strategies.
- Improve the quality and actionability of alerts while reducing unnecessary operational noise.
- Establish dashboards and reporting mechanisms that clearly communicate service health, performance, capacity, and risk.
- Use production data and reliability trends to prioritize engineering investments and continuous-improvement initiatives.
- Design and implement automation that reduces manual intervention, operational toil, and human error.
- Build or enhance tools for deployment, configuration management, infrastructure provisioning, incident response, and service recovery.
- Promote infrastructure-as-code, policy-as-code, automated testing, and repeatable deployment practices.
- Partner with development teams to improve continuous integration and continuous delivery pipelines.
- Develop self-healing and automated remediation capabilities where appropriate.
- Contribute production-quality software and reusable platform components using modern programming and scripting languages.
- Provide technical leadership during complex, high-severity production incidents.
- Coordinate diagnosis, containment, recovery, and stakeholder communication during service disruptions.
- Lead blameless post-incident reviews and ensure that corrective actions address root causes rather than symptoms.
- Identify recurring failure patterns and develop long-term engineering solutions.
- Improve incident-management processes, escalation procedures, runbooks, and recovery playbooks.
- Participate in an on-call rotation or provide senior escalation support for critical services, as required.
- Lead capacity planning, performance analysis, load testing, and demand forecasting for critical platforms.
- Identify performance bottlenecks and recommend architectural or operational improvements.
- Design and validate high-availability, disaster-recovery, backup, and business-continuity capabilities.
- Lead resilience testing, failure-mode analysis, game days, and controlled fault-injection exercises.
- Ensure recovery-time and recovery-point objectives are defined, tested, and achievable.
- Partner with security and compliance teams to embed security into infrastructure, automation, and operational practices.
- Support vulnerability remediation, access-control improvements, audit readiness, and secure configuration management.
- Ensure production environments meet organizational standards for change management, data protection, and operational governance.
- Balance reliability, security, delivery speed, cost, and business priorities when recommending technical solutions.
- Extensive professional experience in site reliability engineering, software engineering, cloud infrastructure, platform engineering, systems engineering, or a related technical discipline.
- Demonstrated experience designing, operating, and improving highly available production systems at significant scale.
- Deep knowledge of distributed systems, cloud architecture, networking, operating systems, storage, databases, and service dependencies.
- Advanced experience with at least one major cloud platform, such as Oracle Cloud Infrastructure, Amazon Web Services, Microsoft Azure, or Google Cloud Platform.
- Strong experience with containerization and orchestration technologies, including Docker and Kubernetes.
- Proven expertise with infrastructure-as-code and configuration-management technologies such as Terraform, Ansible, Chef, Puppet, or equivalent tools.
- Experience implementing observability solutions using metrics, logs, traces, dashboards, and automated alerting.
- Strong programming or scripting skills in one or more languages such as Python, Go, Java, JavaScript, Bash, or similar.
- Experience with continuous integration, continuous delivery, automated testing, and modern release-management practices.
- Demonstrated leadership during critical production incidents and complex technical investigations.
- Ability to diagnose difficult system issues across applications, infrastructure, networks, databases, and cloud services.
- Strong written and verbal communication skills, including the ability to explain technical risk and recommendations to engineering leaders and business stakeholders.
- Proven ability to lead cross-functional technical initiatives without relying solely on formal authority.
- Experience supporting enterprise-scale cloud services, software-as-a-service platforms, or other high-availability customer-facing systems.
- Experience defining and operating service-level objectives, error budgets, and reliability scorecards.
- Knowledge of chaos engineering, resilience testing, and automated recovery techniques.
- Experience with multi-region, hybrid-cloud, or multi-cloud architectures.
- Familiarity with security frameworks, compliance requirements, and regulated operating environments.
- Experience improving cloud cost efficiency, capacity utilization, or infrastructure performance.
- Contributions to internal engineering standards, technical communities, open-source projects, or industry publications.
- Bachelor's or advanced degree in computer science, engineering, information systems, or a related field, or equivalent practical experience.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC5 About Us Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing or by calling 1- in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Lead Principal Site Reliability Engineer in Nashville, TN vacancy
$84.9k - $209.5k
...architecture with practical systems engineering, deployment, automation, patching, troubleshooting... ..., and compliance support. The Principal Site Reliability Engineer will work across Windows,... ...complex technical decisions, and lead improvements that reduce operational...PrincipalTemporary workFlexible hours$169.3k - $304.7k
...maintaining fast, efficient, scalable, and reliable routing software and infrastructure... ...of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible... ...Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K...PrincipalWork experience placementWork at office$142k - $194k
...stakeholders using technical expertise in a core engineering or technical domain. Experience shaping and leading technical programs and highly complex, cross-... ...standards with an emphasis on security, scalability, reliability, and best practices. Experience with program...PrincipalFull timeShift work$142k - $194k
...sequencing Strong background in leading engagements with colocation... ...multidisciplinary internal engineering and operations teams Deep... ...environments Review new site layouts and fit-out proposals... ...engineering approaches to improve reliability, efficiency, and innovation...PrincipalFull timeFor contractorsShift work$114.6k - $234.6k
...Job Description The Lead Principal Infrastructure Capacity Planner will be responsible... ...experience in network planning, backbone engineering, and infrastructure strategy, with a... ..., maintaining a high level of network reliability and performance. #LI-KR4 Qualifications...PrincipalTemporary workFlexible hours$146.3k - $306.4k
...and industry standards. Leads engagements with... ...influences the review of new site layouts and proposed... ...multidisciplinary engineering functions (e.g., Mechanical... ..., enhance system reliability, anticipate complex risks... ...impact. -Serves as the principal authority on mission-...PrincipalContract workTemporary workFor contractorsFlexible hoursShift work$119.2k - $264.1k
...service dependencies and customer experience needs Experience leading cross-functional teams on complex technology projects, defining... .... Our Dedicated Cloud Product Management team is seeking a Principal Product Manager to advance capacity management tooling for Dedicated...PrincipalFull timeFlexible hours$81.1k - $187k
...infrastructure and service to ensure reliability and functionality. Forecasts... ...and develops knowledge of site reliability trends.Only... ...your potential at a company leading the way in AI and cloud solutions... ...and mentorship to junior engineers.Communicate status, risks, blockers...Temporary workFlexible hours- Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week
$69.8k - $148.3k
...infrastructure and service to ensure reliability and functionality. Responds... ...working knowledge of site reliability trends. Responsibilities... ..., Information Technology, Engineering, or a related field, or... ...your potential at a company leading the way in AI and cloud solutions...Temporary workFlexible hours- ...Sr. Principal Member Of Technical Staff As a Sr. Principal Member of Technical Staff, you will work with other senior engineers and product management to define requirements for OCI's storage infrastructure services. Expertise in one or more Public Cloud offerings...Principal
- ...Robotics Platform Security Engineer You will be working on enabling authentication (Secure... ...incident response paradigms to maintain reliability and availability. Software... ...new software features and enhancements leading design specifications, ensuring accessibility...PrincipalTemporary workFlexible hours
$140k - $210k
...Our Mission As the world’s number 1 job site*, our mission is to help people get jobs. We strive to cultivate an... ...Comscore, Total Visits, March 2026) Day to Day As an Engineering Manager in Site Reliability Engineering at Indeed, you will manage and grow a team that...Work experience placementLocal area$105.79k - $141.05k
...NaaS) platform delivers on-demand networking at scale. As Lead SRE, you'll own the reliability of that platform — partnering with operations teams and... ...automation, and you'll coordinate across architecture, engineering, and systems development organizations to measurably...Temporary workRemote work$146.3k - $306.4k
...center deployability availability Mentor engineers participate in design reviews, and... ...production-grade firmware across teams. Lead the integration of system firmware with... ...best practices in systems integration, reliability, and operational excellence. Leads project...PrincipalTemporary workFlexible hours$143k - $193.05k
...Modernization business unit is seeking a Senior Principal Software Engineer to serve as a visionary and hands-on... ...deliver robust protection, highly reliable, performant, operationally simple,... ...differentiation Market leading (e.g. Gartner Magic Quadrant) Full...PrincipalFull timeRemote workWorldwide$96.8k - $306.4k
...have 12+ years of software engineering experience with a strong record... ...production systems and leading across multiple engineering... ...to design secure, scalable, reliable, and extensible platforms.... ...across the enterprise. This Lead Principal Platform Engineer role is a...PrincipalFull timeFlexible hoursShift work$135.2k - $306.4k
...data planes. We are hoping to enhance engineering efficiency by concentrating our... ...validation framework that will ensure the reliability of databases being used by critical... ...Responsibilities As a Senior Principal Engineer, you will lead the design and implementation of...PrincipalTemporary workWorldwideFlexible hours$180.35k - $270.52k
...improve efficiency, and support accurate payments. The Principal Software Engineer is responsible for shaping the technical direction of our products... ...strategy for our flagship initiatives. The position leads the design and development of complex, scalable, and high-performance...PrincipalFull timeRemote work$139.4k - $306.4k
...a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives... ..., measurement)• Strong technical background in software engineering, architecture, security, or platform operations• Strong knowledge...PrincipalTemporary workFlexible hours$160.8k - $165.8k
...at . Job Details The Principal Architect for Warehouse Management... ...stakeholders, product teams, engineering teams, vendors, and... ...leaders to design scalable, reliable, and innovative solutions that... ...Warehouse Management Systems. • Lead architecture decisions for WMS...PrincipalWork experience placementSeasonal work$135.2k - $306.4k
...Job Description As a Sr. Principal Member of Technical Staff, you will work with other senior engineers and product management to define requirements for OCI’s storage infrastructure services. Expertise in one or more Public Cloud offerings is a plus. You will be expected...PrincipalTemporary workFlexible hours- ...Nashville, TN and is an on-site role. Remote work is... ...applications for leading enterprises worldwide.... ...problems, and mentor engineers while continuing to deepen... ...expertise. As a Principal Software Developer, you... ..., and drive scalable, reliable solutions within and across...PrincipalWork at officeRemote workWorldwideRelocationRelocation package
$104.5k - $234.6k
...masters degree in Computer Science, Computer Engineering, or a related discipline, or equivalent... .... We value demonstrated experience leading or shaping architecture and technical... ...solutions, edge computing, and more. This Principal Platform Software Engineer position is...PrincipalFull timeWork at officeWorldwideRelocationRelocation packageFlexible hours- ...openly celebrates all cultures and affords personal and professional growth opportunities. Learn more at ( om. Verint’s Principal Software Engineer is responsible for developing, troubleshooting, resolving challenging problem and providing technical guidance to...PrincipalLocal areaShift work
$135.2k - $306.4k
...software delivery. Responsibilities Lead the design, implementation, and... ...support. Drive improvements to workflow reliability, deployment safety, scalability, observability... ...platform problems into durable engineering solutions. Set high standards for code...PrincipalTemporary workFlexible hours$142k - $194k
...year Requirements: ~ Bachelors degree in Computer Science, Engineering, or a related technical field, or equivalent practical... ...customer workflows. Troubleshoot performance, scalability, and reliability issues, then implement mitigations to reduce risk and downtime...PrincipalFull time$169.3k - $304.7k
...shape the future of our cloud-native and Kubernetes services We are looking for a Principal Software Engineer with deep expertise in Kubernetes and cloud-native architectures to lead the evolution of our managed services. You will be responsible for helping to...PrincipalWork experience placementWork at office$142k - $194k
...with deep expertise in network products engineering and cloud-scale infrastructure. We need... ..., automation, testing, monitoring, reliability, and capacity planning. You should be... ...processes to increase efficiency. We lead automation framework development, expand...PrincipalFull timeWorldwide$114.6k - $234.6k
...Infrastructure (OCI) is seeking a Senior Principal Technical Program Manager (IC5) to join... ...all. Discover your potential at a company leading the way in AI and cloud solutions that... ...timelinesDrive cross-functional coordination across engineering, construction, networking, and...PrincipalTemporary workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Principal Site Reliability Engineer. Be the first to apply!
Related searches
- lead infrastructure engineer Nashville, TN
- lead engineer Nashville, TN
- lead operating engineer Nashville, TN
- lead network engineer Nashville, TN
- senior chief engineer Nashville, TN
- general engineer Nashville, TN
- principal infrastructure engineer Nashville, TN
- chief engineer Nashville, TN
- principal developer Nashville, TN
- senior principal engineer Nashville, TN


