Lead Principal Site Reliability Engineer
$96.3k - $264.1kOracle Corporation
Serves as a consultant and leads the design and architecture of infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full ownership of forecasting of demands and responding to capacity needs. Owns collaborations with software development teams to develop reliable and scalable infrastructures. Recommends methods for performing data collection to maintain and optimize operations and reliability. Oversees incident response and/or maintenance tasks. Provides strategic, future-oriented health and performance reporting. Contributes to strategies for automation and reviews the development and implementation of automation. Provides expert-level communication about services and anticipates, analyzes, and explains the impact of changes, considering strategic goals. Serves as a role model in providing support for technology and reviews documentation for accuracy. Leads the implementation of innovative tools and provides expertise in site reliability trends.Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on click.appcast.io or by calling View phone number on click.appcast.io in the United States.Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.Disclaimer:Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.Range and benefit information provided in this posting are specific to the stated locations onlyUS: Hiring Range in USD from: $96,300 to $264,100 per annum. May be eligible for bonus, equity, and compensation deferral.Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.Oracle US offers a comprehensive benefits package which includes the following:1. Medical, dental, and vision insurance, including expert medical opinion2. Short term disability and long term disability3. Life insurance and AD&D4. Supplemental life insurance (Employee/Spouse/Child)5. Health care and dependent care Flexible Spending Accounts6. Pre-tax commuter and parking benefits7. 401(k) Savings and Investment Plan with company match8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.9. 11 paid holidays10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.11. Paid parental leave12. Adoption assistance13. Employee Stock Purchase Plan14. Financial planning and group legal15. Voluntary benefits including auto, homeowner and pet insuranceThe role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.Career Level - IC5Key ResponsibilitiesReliability Strategy and Technical LeadershipDefine and drive the site reliability engineering strategy for large-scale, distributed, and business-critical platforms. Establish reliability standards, engineering practices, and operational readiness requirements across multiple teams. Serve as a senior technical authority for system reliability, scalability, resilience, performance, and production operations. Influence architecture and design decisions to ensure systems are supportable, observable, fault tolerant, and capable of meeting availability objectives. Identify systemic reliability risks and lead cross-functional initiatives to address them. Provide technical direction and mentorship to site reliability engineers, software engineers, platform engineers, and operations teams. Lead technical reviews and promote consistent engineering practices across the organization. Service Reliability and ObservabilityDefine and implement service-level indicators, service-level objectives, error budgets, and operational health metrics. Develop comprehensive monitoring, logging, tracing, alerting, and observability strategies. Improve the quality and actionability of alerts while reducing unnecessary operational noise. Establish dashboards and reporting mechanisms that clearly communicate service health, performance, capacity, and risk. Use production data and reliability trends to prioritize engineering investments and continuous-improvement initiatives. Automation and Platform EngineeringDesign and implement automation that reduces manual intervention, operational toil, and human error. Build or enhance tools for deployment, configuration management, infrastructure provisioning, incident response, and service recovery. Promote infrastructure-as-code, policy-as-code, automated testing, and repeatable deployment practices. Partner with development teams to improve continuous integration and continuous delivery pipelines. Develop self-healing and automated remediation capabilities where appropriate. Contribute production-quality software and reusable platform components using modern programming and scripting languages. Incident Management and Problem ResolutionProvide technical leadership during complex, high-severity production incidents. Coordinate diagnosis, containment, recovery, and stakeholder communication during service disruptions. Lead blameless post-incident reviews and ensure that corrective actions address root causes rather than symptoms. Identify recurring failure patterns and develop long-term engineering solutions. Improve incident-management processes, escalation procedures, runbooks, and recovery playbooks. Participate in an on-call rotation or provide senior escalation support for critical services, as required. Capacity, Performance, and ResilienceLead capacity planning, performance analysis, load testing, and demand forecasting for critical platforms. Identify performance bottlenecks and recommend architectural or operational improvements. Design and validate high-availability, disaster-recovery, backup, and business-continuity capabilities. Lead resilience testing, failure-mode analysis, game days, and controlled fault-injection exercises. Ensure recovery-time and recovery-point objectives are defined, tested, and achievable. Security and Operational GovernancePartner with security and compliance teams to embed security into infrastructure, automation, and operational practices. Support vulnerability remediation, access-control improvements, audit readiness, and secure configuration management. Ensure production environments meet organizational standards for change management, data protection, and operational governance. Balance reliability, security, delivery speed, cost, and business priorities when recommending technical solutions. Required QualificationsExtensive professional experience in site reliability engineering, software engineering, cloud infrastructure, platform engineering, systems engineering, or a related technical discipline. Demonstrated experience designing, operating, and improving highly available production systems at significant scale. Deep knowledge of distributed systems, cloud architecture, networking, operating systems, storage, databases, and service dependencies. Advanced experience with at least one major cloud platform, such as Oracle Cloud Infrastructure, Amazon Web Services, Microsoft Azure, or Google Cloud Platform. Strong experience with containerization and orchestration technologies, including Docker and Kubernetes. Proven expertise with infrastructure-as-code and configuration-management technologies such as Terraform, Ansible, Chef, Puppet, or equivalent tools. Experience implementing observability solutions using metrics, logs, traces, dashboards, and automated alerting. Strong programming or scripting skills in one or more languages such as Python, Go, Java, JavaScript, Bash, or similar. Experience with continuous integration, continuous delivery, automated testing, and modern release-management practices. Demonstrated leadership during critical production incidents and complex technical investigations. Ability to diagnose difficult system issues across applications, infrastructure, networks, databases, and cloud services. Strong written and verbal communication skills, including the ability to explain technical risk and recommendations to engineering leaders and business stakeholders. Proven ability to lead cross-functional technical initiatives without relying solely on formal authority. Preferred QualificationsExperience supporting enterprise-scale cloud services, software-as-a-service platforms, or other high-availability customer-facing systems. Experience defining and operating service-level objectives, error budgets, and reliability scorecards. Knowledge of chaos engineering, resilience testing, and automated recovery techniques. Experience with multi-region, hybrid-cloud, or multi-cloud architectures. Familiarity with security frameworks, compliance requirements, and regulated operating environments. Experience improving cloud cost efficiency, capacity utilization, or infrastructure performance. Contributions to internal engineering standards, technical communities, open-source projects, or industry publications. Bachelor’s or advanced degree in computer science, engineering, information systems, or a related field, or equivalent practical experience. Full timePosting Date: 2026-07-16
$84.9k - $209.5k
...architecture with practical systems engineering, deployment, automation, patching, troubleshooting... ..., and compliance support. The Principal Site Reliability Engineer will work across Windows,... ...complex technical decisions, and lead improvements that reduce operational...PrincipalTemporary workFlexible hours$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity...SuggestedWork at officeLocal areaVisa sponsorshipFlexible hours3 days per week$139.4k - $306.4k
...a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives... ..., measurement)• Strong technical background in software engineering, architecture, security, or platform operations• Strong knowledge...PrincipalTemporary workFlexible hours$114.6k - $234.6k
...Infrastructure (OCI) is seeking a Senior Principal Technical Program Manager (IC5) to join... ...all. Discover your potential at a company leading the way in AI and cloud solutions that... ...timelinesDrive cross-functional coordination across engineering, construction, networking, and...PrincipalTemporary workFlexible hours$116.6k - $264.1k
...priorities and goals using technical know-how in a fundamental engineering area or technical domain. Establishes scope and milestones for... ...cross-functional teams to deliver programs or services. Shapes and leads technical programs and highly complex, cross-organizational...PrincipalTemporary workFlexible hours$119.2k - $264.1k
As a Lead Principal Product Manager, you will own a major strategic product area that reaches... ...customer executives, architects, and engineering leaders to understand priorities,... ...customer needs with industry-leading reliability, security, performance, and price-performance...PrincipalTemporary workFlexible hours$119.2k - $264.1k
...needs, ensuring alignment with Oracle's goals and strategies. Leads the engagement of key strategic customers, non-customers, partners... ...and decision-making.-Guides collaboration efforts with engineering and senior colleagues to execute innovative solutions tied to OKRs...PrincipalTemporary workFlexible hoursShift work- Music City Center in Nashville seeks an Elementary School Assistant Principal to support the principal in managing academic and administrative school operations. The role requires proven leadership skills and experience in educational administration to foster a positive...Principal
$116.6k - $264.1k
...Transformation & Delivery team as a Senior Principal Program Manager at our Nashville HQ.As... ...individual contributor role, you will lead highly visible programs that span multiple... ...strong coordination across product, engineering, operations, and corporate functions.You...PrincipalTemporary workFlexible hours$81.1k - $187k
...infrastructure and service to ensure reliability and functionality. Forecasts... ...and develops knowledge of site reliability trends.Only... ...your potential at a company leading the way in AI and cloud solutions... ...and mentorship to junior engineers. Communicate status, risks, blockers...Temporary workFlexible hours$81.1k - $187k
...infrastructure and service to ensure reliability and functionality. Forecasts... ...and develops knowledge of site reliability trends.Only... ...your potential at a company leading the way in AI and cloud solutions... ..., Information Technology, Engineering, or a related field, or equivalent...Temporary workFlexible hours- A forward-thinking CPA firm in Nashville is seeking an experienced Principal Accountant. This remote role offers an opportunity to manage client relationships, lead a tax team, and navigate complex tax issues. Requirements include CPA certification, 8+ years of experience...PrincipalPart timeRemote work
- Leads the design and delivery of AI-native developer platforms, services, and workflows that improve how OCI teams build and operate... ...full software development lifecycle. Applies strong software engineering and distributed-systems expertise to integrate AI-assisted development...PrincipalFull timeFlexible hours
- You will be a senior member of the engineering team that builds the Oracle Managed Kubernetes (OMK) platform at Oracle Cloud Infrastructure... ...broader OCI ecosystem. Typical activities will range from leading design and discovery discussions, reviewing the roadmap with...PrincipalFlexible hours
- This role will help evolve Oracle Cloud’s pre-deployment safety platform for production changes. The engineer will lead cross-team platform projects across OCI deployment safety systems, and service teams to build scalable release-safety capabilities for policy evaluation...PrincipalFull timeFlexible hours
$74.1k - $148.3k
...software performance analysis, and system tuning. As a Site Reliability Engineer, you will solve interesting technical challenges by defining... ...better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of...Temporary workImmediate startFlexible hours$100k - $130k
...creating transformative change in healthcare. We are seeking a Site Reliability Engineer to play a key role in designing, optimizing, and securing... .... What You Will Do Reliability & Technical Ownership Lead the design and implementation of reliability improvements...Flexible hours- ...Role: Site Reliability Engineer (SRE) Location: Brentwood, TN (Onsite) Contract Experience: 6-8+ years Role Description: Combines software engineering and IT operations to ensure the reliability, scalability, and performance of systems, with...Contract work
- ...to operations. In this role, you will lead complex, cross-functional infrastructure... ...~ Drive alignment across Construction, Engineering, Network, Hardware/GPU, Facilities, Security... ...commissioning readiness, design changes, site constraints, vendor performance). ~...PrincipalFull time
$135.2k - $306.4k
...handling, and unsafe credential storage patterns.Build reliable APIs, tooling, and workflows that help other engineering teams adopt secrets and key-management... ...safety for security-critical production systems.Lead threat modeling, security reviews, design hardening...PrincipalTemporary workFlexible hours$135.2k - $306.4k
.... Discover your potential at a company leading the way in AI and cloud solutions that... ...support.• Drive improvements to workflow reliability, deployment safety, scalability,... ...ambiguous platform problems into durable engineering solutions.• Set high standards for code...PrincipalTemporary workFlexible hours$135.2k - $306.4k
As a Sr. Principal Member of Technical Staff, you will work with other senior engineers and product management to define requirements for OCI’s storage infrastructure services... ...all. Discover your potential at a company leading the way in AI and cloud solutions that impact...PrincipalTemporary workFlexible hours$105.79k - $141.05k
...shape the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role...Temporary workRemote work$135.2k - $306.4k
...Oracle Health Infrastructure and Platform Services is looking for Engineers with experience in building cloud-native, AI-assisted platforms... ...a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives...PrincipalTemporary workFlexible hours$104.5k - $234.6k
...pre-deployment safety platform for production changes. The engineer will lead cross-team platform projects across OCI deployment safety systems... ...ensure service/product availability, health, support, and reliability.Core ResponsibilitiesPlanning & Execution:Manages and...PrincipalTemporary workFlexible hoursShift work$250.6k - $362.6k
...comprehensive security outcomes, as a Principal Engineer. The team delivers secure, scalable capabilities... ...networking, with a strong emphasis on reliability, interoperability, and long-term... ...insurance. Please see the Cisco careers site to discover more benefits and perks....PrincipalFull timeTemporary workLocal areaRemote workFlexible hours- ...JOB TITLE: Principal Software Engineer DEPARTMENT: Transportation REPORTS TO: Chief Technology Officer JOB LOCATION: Remote (U.S. based) TRAVEL: No SUMMARY OF POSITION: The Software Engineer...PrincipalFull timeRemote work
$96.8k - $223.4k
As a member of the software engineering division, you will apply intermediate to advanced knowledge of software architecture to perform... ...a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives...PrincipalTemporary workRelocationFlexible hours$114.6k - $234.6k
...platform optimized for live and linear streaming. This role will lead the design and development of advertising infrastructure that... ...deliveryWork with a highly technical, distributed systems-focused engineering teamOnly Oracle brings together the data, infrastructure,...PrincipalTemporary workFlexible hours$96.8k - $306.4k
...drive region build automation to next level. This role will lead the design and development of scalable software solutions... ...workplace where you’ll belong and be encouraged.As a Lead Principal Software Engineer, you will work with teams of software engineers responsible...PrincipalTemporary workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Principal Site Reliability Engineer. Be the first to apply!
- lead operating engineer Nashville, TN
- lead engineer Nashville, TN
- chief engineer Nashville, TN
- engineering director Nashville, TN
- principal network engineer Nashville, TN
- senior director engineering Nashville, TN
- director data engineering Nashville, TN
- senior chief engineer Nashville, TN
- hotel chief engineer Nashville, TN
- principal developer Nashville, TN


