Principal Site Reliability Engineer
$84.9k - $209.5kOracle
Job Description This role combines strategic architecture with practical systems engineering, deployment, automation, patching, troubleshooting, incident response, and compliance support. The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud Infrastructure, hybrid cloud, and legacy environments while partnering with engineering, operations, cybersecurity, networking, application, and client-facing teams. The successful candidate will serve as a senior technical authority, establish reliability standards, guide complex technical decisions, and lead improvements that reduce operational risk and manual effort. This individual must be comfortable moving between architecture and hands-on execution, including accessing deployed hosts, troubleshooting failed services, reviewing logs, correcting configurations, and validating production changes. Responsibilities Key Responsibilities
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC4 About Us Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on click.appcast.io or by calling View phone number on click.appcast.io in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
- Design and architect reliable, secure, scalable, and maintainable infrastructure and services. Take proactive steps to ensure solutions meet defined reliability and functionality requirements.
- Establish technical direction, engineering standards, and operational best practices across complex infrastructure and application environments.
- Identify system dependencies, operational risks, capacity constraints, performance issues, and potential failure points before they affect service.
- Translate business, client, security, and application requirements into practical infrastructure and reliability solutions.
- Lead the installation, configuration, deployment, and validation of applications across Windows Server and Linux environments.
- Oversee structured builds and deployments using runbooks, scripts, readiness assessments, change controls, and post-deployment validation.
- Troubleshoot complex operating system, service, application, installation, patching, permissions, certificate, and connectivity issues.
- Define and improve monitoring, alerting, logging, observability, capacity planning, and service-health practices.
- Develop and promote automation that reduces manual effort, improves consistency, and lowers operational risk.
- Lead operating system, middleware, and application patching initiatives, including change planning, rollback preparation, execution, and validation.
- Direct major incident response, root cause analysis, corrective-action planning, and prevention of recurring failures.
- Partner with cybersecurity teams on vulnerability remediation, system hardening, STIG compliance, and other security-driven changes.
- Evaluate emerging technologies and recommend solutions that improve reliability, resilience, security, and operational efficiency.
- Create and maintain technical standards, architecture documentation, runbooks, deployment procedures, and troubleshooting guides.
- Provide technical leadership, mentorship, and design guidance to engineers across multiple teams.
- Communicate technical risks, dependencies, decisions, and recommendations clearly to leadership and stakeholders.
- Extensive experience in site reliability engineering, systems engineering, infrastructure architecture, production operations, or application hosting.
- Demonstrated ability to design and support highly available, resilient, and secure enterprise systems.
- Experience leading complex technical initiatives across engineering, operations, security, networking, and application teams.
- Ability to make sound architectural decisions, evaluate tradeoffs, and communicate recommendations to technical and nontechnical stakeholders.
- Experience defining engineering standards, operational controls, and reliability practices.
- Advanced, hands-on experience administering Windows Server and/or Linux systems.
- Ability to access deployed hosts and perform post-deployment configuration, troubleshooting, and validation.
- Experience installing, configuring, and validating applications in Windows Server and Linux environments.
- Ability to resolve operating-system-level, service-level, and application-level issues.
- Strong knowledge of system services, permissions, configuration files, logs, processes, and resource utilization.
- Experience leading structured build and deployment activities using runbooks, deployment guides, scripts, and technical procedures.
- Ability to execute and troubleshoot scripts, validate outputs, and resolve build or configuration issues.
- Experience with build handoffs, environment-readiness assessments, deployment validation, and post-build verification.
- Ability to identify process gaps, document exceptions, and improve deployment procedures.
- Experience managing complex or high-risk production changes.
- Ability to investigate complex service failures, installation errors, patching failures, application startup problems, permissions issues, and connectivity incidents.
- Experience reviewing logs, event viewers, service status, configuration files, ports, certificates, and access controls.
- Strong analytical and problem-solving skills, with the ability to isolate root causes and implement sustainable solutions.
- Experience leading major incident response and coordinating technical teams during business-critical outages.
- Ability to document symptoms, findings, impact, corrective actions, and recommended next steps clearly.
- Extensive experience supporting production or other mission-critical environments.
- PowerShell
- Bash
- Python
- Ansible
- Chef
- Experience planning and executing operating system, middleware, and application patching.
- Ability to troubleshoot patch failures, compatibility issues, and post-patch application problems.
- Strong understanding of maintenance windows, change control, rollback planning, risk assessment, and post-change validation.
- Experience coordinating patching and remediation activities across application, infrastructure, cybersecurity, and client teams.
- Strong understanding of cloud-hosted and hybrid infrastructure.
- Experience with Oracle Cloud Infrastructure or another major cloud platform.
- Knowledge of cloud compute, storage, networking, identity, access management, load balancing, and environment provisioning.
- Experience designing or supporting reliable, secure, and scalable cloud environments.
- Familiarity with infrastructure-as-code and configuration-management practices.
- Working knowledge of DNS, firewalls, routing, load balancers, ports, certificates, and network communication between systems.
- Ability to identify and isolate host, application, certificate, firewall, DNS, and routing-related issues.
- Familiarity with standard connectivity and network diagnostic tools.
- Experience with vulnerability remediation, system hardening, secure configuration, and compliance-driven infrastructure changes.
- Familiarity with Security Technical Implementation Guides and federal cybersecurity requirements.
- Ability to implement security remediation without disrupting application functionality or service availability.
- Experience supporting federal, government-hosted, healthcare, or other regulated environments is highly valued.
- Ability to create and maintain architecture documentation, operational standards, technical procedures, runbooks, and change records.
- Strong attention to detail when documenting completed work, risks, exceptions, decisions, and validation results.
- Experience working in ticketing, incident-management, problem-management, and change-management systems.
- Excellent written and verbal communication skills.
- Ability to present complex technical issues, risks, and recommendations to engineers, clients, and senior leadership.
- Demonstrated ability to mentor engineers and influence technical direction across teams.
- Experience supporting Oracle Cloud Infrastructure environments.
- Experience supporting Oracle Health, Millennium, or Cerner applications and related infrastructure.
- Knowledge of federal cybersecurity workflows, STIG implementation, and compliance requirements.
- Experience supporting federal clients or government-hosted environments.
- Experience with Citrix technologies.
- Experience supporting legacy infrastructure and business-critical legacy applications.
- Advanced experience with infrastructure-as-code or configuration-management tools.
- Experience with production support, incident command, problem management, and SRE operational practices.
- Knowledge of service-level indicators, service-level objectives, error budgets, observability, and capacity planning.
- Experience designing high-availability, disaster-recovery, backup, and service-continuity solutions.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC4 About Us Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on click.appcast.io or by calling View phone number on click.appcast.io in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in Nashville, TN vacancy
- Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs... ...new tools and develops and maintains advanced knowledge of site reliability trends. Only Oracle brings together the data, infrastructure...PrincipalFull timeFlexible hours
- As a Principal Site Reliability Engineer (IC4), you will be responsible for designing, building, and operating highly available, scalable, secure, and resilient cloud services. You will combine software engineering with infrastructure expertise to improve service reliability...PrincipalFull timeFlexible hours
- We are looking for an experienced Senior Site Reliability Engineer to join our team and drive the reliability and functionality of our critical infrastructure and applications. In this role, you will be a key contributor, taking ownership of designing and architecting...SuggestedFull timeFlexible hours
- Software Engineer, Project Babylon (IC4) Group Description The Java Platform Group is responsible for advancing the Java platform through the OpenJDK community and delivering innovative capabilities to millions of developers worldwide. Within this organization, Project...PrincipalFull timeWorldwideFlexible hours
- Oracle Cloud Infrastructure (OCI) is seeking an experienced Strategic Sr Principal Technical Program Manager to drive high-impact, cross-organizational initiatives that enable OCI's continued growth and expansion. This role operates at the intersection of infrastructure...PrincipalFull timeFlexible hours
- ...Site Reliability Engineer - Compute Focus Duration: 6 Months to Hire Location: On-Site 2-3 days/week in Nashville, TN Job Description: We are seeking a skilled Site Reliability Engineer with a focus on compute infrastructure to join our dynamic team. The ideal...2 days per week3 days per week
- ...significantly reduces costs and improves the critically important 24x7 performance for building owners, developers and tenants. Site Reliability Engineer II The SRE II sits at the intersection of software engineering and platform operations. You will own the reliability,...Remote work
$75.7k - $136.3k
...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and...Work experience placementWork at office- ...Role: Site Reliability Engineer (SRE) Location: Brentwood, TN (Onsite) Contract Experience: 6-8+ years Role Description: Combines software engineering and IT operations to ensure the reliability, scalability, and performance of systems, with...Contract work
- ...JOB TITLE: Principal Software Engineer DEPARTMENT: Transportation REPORTS TO: Chief Technology Officer JOB LOCATION: Remote (U.S. based) TRAVEL: No SUMMARY OF POSITION: The Software Engineer...PrincipalFull timeRemote work
$51.9 per hour
...job is responsible for the reliability, availability, and performance... ...healthcare IT systems, principally in the Environment of Care (... .... This role blends software engineering, clinical engineering, and security... ...cross-functionally with AHN site leaders and teams to...For contractorsLocal area- Works with management and stakeholders to align priorities and goals using technical know-how in a fundamental engineering area or technical domain. Establishes scope and milestones for each aspect of a technical program, aligning to the broader program plan and company...PrincipalFull timeFlexible hours
- ...ambiguous technical problems, and mentor engineers while continuing to deepen your own technical expertise. Responsibilities As a Principal Software Developer, you will lead the... ...architectural decisions, and drive scalable, reliable solutions within and across teams. You...PrincipalWork at officeWorldwideRelocationRelocation package
- ...Senior Director, Principal Gifts About the Company Philanthropic organization supporting Indigenous culture & individuals Industry Non-Profit Organization Management Type Non Profit Founded 2017 Employees 11-50 Categories...Principal
- ...will collaborate closely with partner security teams (such as SOC, digital forensics, incident response, physical security, and engineering) and work cross-functionally with senior leaders from HR, Legal, crisis management, compliance, and other business units during...PrincipalFull timeFlexible hours
- ...Infrastructure is building developer tool products that help engineering teams create software with greater speed, quality, and confidence... ..., onboarding, and service comprehension. We are seeking a Principal Product Manager to lead product strategy and execution for...PrincipalFull timeLocal areaFlexible hours
$173.5k - $310k
...analytics with a focus on security, product experience, and scalability. Close Collaboration: Work in Audit’s collaborative model where engineers understand business problems as deeply as technical solutions, partnering closely with product managers who prototype their own...PrincipalWork at office- As a member of the software engineering division, you will apply basic to intermediate knowledge of software architecture to perform software development tasks associated with developing, debugging or designing software applications or operating systems according to...Full timeRelocationFlexible hours
$185k - $237.5k
...A leading financial technology company based in Nashville is seeking a Principal Product Operations and Risk Analyst. This role focuses on leveraging over 10 years of risk management experience to support the day-to-day operations of Product Risk Management. Responsibilities...Principal- A forward-thinking CPA firm in Nashville is seeking an experienced Principal Accountant. This remote role offers an opportunity to manage client relationships, lead a tax team, and navigate complex tax issues. Requirements include CPA certification, 8+ years of experience...PrincipalPart timeRemote work
- As a Senior Software Development Engineer, you will own the design and development of major components that improve the developer experience for software teams building Oracle Cloud Infrastructure. You should be a rock-solid coder and distributed systems generalist...Full timeFlexible hours
- Responsibilities Meet new business production goals and objectives as established. Prospects for new business including sales leads generated from referrals, networking, marketing, cold-calling, and lead databases. Grow sales revenue by utilizing phone, email...Principal
- Music City Center in Nashville seeks an Elementary School Assistant Principal to support the principal in managing academic and administrative school operations. The role requires proven leadership skills and experience in educational administration to foster a positive...Principal
- ...Vice President, Principal Owner Development & Growth About the Company Top-tier mutual life insurance company Industry Financial Services Type Privately Held, Private Equity-backed Founded 1860 Employees 5001-10,000 Categories...PrincipalHome office
- NoneFull time
$87k - $187k
Job Description An experienced consulting professional who has a broad understanding of solutions, industry best practices, multiple business processes or technology designs within a product/technology family. Operates independently to provide quality work products ...PrincipalTemporary workFlexible hours$115.3k - $264.1k
...national narrative and local engagement model for one of Oracle's most visible growth areas: data center and AI infrastructure. The Sr Principal Program Manager - Data Center Campaigns will own the operating rhythm for a national campaign that connects campaign strategy,...PrincipalTemporary workLocal areaFlexible hours$75k
...High School Assistant Principal of Instruction Location: Nashville, Tennessee Employment Type: Full-time, in-person 12-month position Starting Salary: $75,000 (Final salary is based on experience) Why Choose STEM Preparatory Academy? At STEM Prep, we are more...PrincipalFull timeLocal area- ...JOB TITLE: Assistant Principal of Instruction (API) REPORTS TO: Principal JOB OVERVIEW: The Assistant Principal of Instruction at LEAD Public Schools will drive outstanding academic results by managing instructional managers who in turn work with teachers to strengthen...PrincipalFull time
- ...fully integrate data across the enterprise. What You’ll Do As the Principal AI Architect for Teradata AI Studio, you will define the... ...development environment — the platform where data scientists, ML engineers, and AI developers build, test, deploy, and monitor AI and...PrincipalPermanent employmentFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
Related searches
- senior director engineering Nashville, TN
- chief engineer Nashville, TN
- senior principal engineer Nashville, TN
- engineering director Nashville, TN
- senior chief engineer Nashville, TN
- data center chief engineer Nashville, TN
- senior civil engineer project manager Nashville, TN
- general engineer Nashville, TN
- principal engineer Nashville, TN
- hotel chief engineer Nashville, TN



