Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer

$84.9k - $209.5k
Full-time

Oracle Corporation

This role combines strategic architecture with practical systems engineering, deployment, automation, patching, troubleshooting, incident response, and compliance support. The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud Infrastructure, hybrid cloud, and legacy environments while partnering with engineering, operations, cybersecurity, networking, application, and client-facing teams.The successful candidate will serve as a senior technical authority, establish reliability standards, guide complex technical decisions, and lead improvements that reduce operational risk and manual effort. This individual must be comfortable moving between architecture and hands-on execution, including accessing deployed hosts, troubleshooting failed services, reviewing logs, correcting configurations, and validating production changes.Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on job-api.jobget.com or by calling View phone number on job-api.jobget.com in the United States.Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.Disclaimer:Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.Range and benefit information provided in this posting are specific to the stated locations onlyUS: Hiring Range in USD from: $84,900 to $209,500 per annum. May be eligible for bonus and equity.Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.Oracle US offers a comprehensive benefits package which includes the following:1. Medical, dental, and vision insurance, including expert medical opinion2. Short term disability and long term disability3. Life insurance and AD&D4. Supplemental life insurance (Employee/Spouse/Child)5. Health care and dependent care Flexible Spending Accounts6. Pre-tax commuter and parking benefits7. 401(k) Savings and Investment Plan with company match8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.9. 11 paid holidays10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.11. Paid parental leave12. Adoption assistance13. Employee Stock Purchase Plan14. Financial planning and group legal15. Voluntary benefits including auto, homeowner and pet insuranceThe role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.Career Level - IC4Key ResponsibilitiesDesign and architect reliable, secure, scalable, and maintainable infrastructure and services. Take proactive steps to ensure solutions meet defined reliability and functionality requirements. Establish technical direction, engineering standards, and operational best practices across complex infrastructure and application environments. Identify system dependencies, operational risks, capacity constraints, performance issues, and potential failure points before they affect service. Translate business, client, security, and application requirements into practical infrastructure and reliability solutions. Lead the installation, configuration, deployment, and validation of applications across Windows Server and Linux environments. Oversee structured builds and deployments using runbooks, scripts, readiness assessments, change controls, and post-deployment validation. Troubleshoot complex operating system, service, application, installation, patching, permissions, certificate, and connectivity issues. Define and improve monitoring, alerting, logging, observability, capacity planning, and service-health practices. Develop and promote automation that reduces manual effort, improves consistency, and lowers operational risk. Lead operating system, middleware, and application patching initiatives, including change planning, rollback preparation, execution, and validation. Direct major incident response, root cause analysis, corrective-action planning, and prevention of recurring failures. Partner with cybersecurity teams on vulnerability remediation, system hardening, STIG compliance, and other security-driven changes. Evaluate emerging technologies and recommend solutions that improve reliability, resilience, security, and operational efficiency. Create and maintain technical standards, architecture documentation, runbooks, deployment procedures, and troubleshooting guides. Provide technical leadership, mentorship, and design guidance to engineers across multiple teams. Communicate technical risks, dependencies, decisions, and recommendations clearly to leadership and stakeholders. Core Skills and QualificationsTechnical Leadership and ArchitectureExtensive experience in site reliability engineering, systems engineering, infrastructure architecture, production operations, or application hosting. Demonstrated ability to design and support highly available, resilient, and secure enterprise systems. Experience leading complex technical initiatives across engineering, operations, security, networking, and application teams. Ability to make sound architectural decisions, evaluate tradeoffs, and communicate recommendations to technical and nontechnical stakeholders. Experience defining engineering standards, operational controls, and reliability practices. Windows and Linux System AdministrationAdvanced, hands-on experience administering Windows Server and/or Linux systems. Ability to access deployed hosts and perform post-deployment configuration, troubleshooting, and validation. Experience installing, configuring, and validating applications in Windows Server and Linux environments. Ability to resolve operating-system-level, service-level, and application-level issues. Strong knowledge of system services, permissions, configuration files, logs, processes, and resource utilization. Manual Build and Deployment ExperienceExperience leading structured build and deployment activities using runbooks, deployment guides, scripts, and technical procedures. Ability to execute and troubleshoot scripts, validate outputs, and resolve build or configuration issues. Experience with build handoffs, environment-readiness assessments, deployment validation, and post-build verification. Ability to identify process gaps, document exceptions, and improve deployment procedures. Experience managing complex or high-risk production changes. Troubleshooting and Operational SupportAbility to investigate complex service failures, installation errors, patching failures, application startup problems, permissions issues, and connectivity incidents. Experience reviewing logs, event viewers, service status, configuration files, ports, certificates, and access controls. Strong analytical and problem-solving skills, with the ability to isolate root causes and implement sustainable solutions. Experience leading major incident response and coordinating technical teams during business-critical outages. Ability to document symptoms, findings, impact, corrective actions, and recommended next steps clearly. Extensive experience supporting production or other mission-critical environments. Scripting and AutomationAdvanced hands-on experience with one or more of the following:PowerShell Bash Python Ansible Chef Candidates should be able to create, modify, validate, and troubleshoot scripts and automation workflows. Experience identifying automation opportunities and establishing safe, repeatable operational processes is essential.Patching and Software MaintenanceExperience planning and executing operating system, middleware, and application patching. Ability to troubleshoot patch failures, compatibility issues, and post-patch application problems. Strong understanding of maintenance windows, change control, rollback planning, risk assessment, and post-change validation. Experience coordinating patching and remediation activities across application, infrastructure, cybersecurity, and client teams. Cloud and Oracle Cloud InfrastructureStrong understanding of cloud-hosted and hybrid infrastructure. Experience with Oracle Cloud Infrastructure or another major cloud platform. Knowledge of cloud compute, storage, networking, identity, access management, load balancing, and environment provisioning. Experience designing or supporting reliable, secure, and scalable cloud environments. Familiarity with infrastructure-as-code and configuration-management practices. Network and Connectivity TroubleshootingWorking knowledge of DNS, firewalls, routing, load balancers, ports, certificates, and network communication between systems. Ability to identify and isolate host, application, certificate, firewall, DNS, and routing-related issues. Familiarity with standard connectivity and network diagnostic tools. Cybersecurity and ComplianceExperience with vulnerability remediation, system hardening, secure configuration, and compliance-driven infrastructure changes. Familiarity with Security Technical Implementation Guides and federal cybersecurity requirements. Ability to implement security remediation without disrupting application functionality or service availability. Experience supporting federal, government-hosted, healthcare, or other regulated environments is highly valued. Documentation and CommunicationAbility to create and maintain architecture documentation, operational standards, technical procedures, runbooks, and change records. Strong attention to detail when documenting completed work, risks, exceptions, decisions, and validation results. Experience working in ticketing, incident-management, problem-management, and change-management systems. Excellent written and verbal communication skills. Ability to present complex technical issues, risks, and recommendations to engineers, clients, and senior leadership. Demonstrated ability to mentor engineers and influence technical direction across teams. Preferred QualificationsExperience supporting Oracle Cloud Infrastructure environments. Experience supporting Oracle Health, Millennium, or Cerner applications and related infrastructure. Knowledge of federal cybersecurity workflows, STIG implementation, and compliance requirements. Experience supporting federal clients or government-hosted environments. Experience with Citrix technologies. Experience supporting legacy infrastructure and business-critical legacy applications. Advanced experience with infrastructure-as-code or configuration-management tools. Experience with production support, incident command, problem management, and SRE operational practices. Knowledge of service-level indicators, service-level objectives, error budgets, observability, and capacity planning. Experience designing high-availability, disaster-recovery, backup, and service-continuity solutions. Full timePosting Date: 2026-07-16

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in Nashville, TN vacancy
  • $84.9k - $209.5k

    This role combines strategic architecture with practical systems engineering, deployment, automation, patching, troubleshooting, incident response, and compliance support. The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud Infrastructure... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    1 day ago
  • $96.3k - $264.1k

     ...infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full...  ...tools and provides expertise in site reliability trends.Only Oracle brings...  ...LeadershipDefine and drive the site reliability engineering strategy for large-scale, distributed,... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    6 days ago
  • $81.1k - $187k

     ...architect infrastructure and service to ensure reliability and functionality. Forecasts demands and...  ...impact and develops knowledge of site reliability trends.Only Oracle brings together...  ...guidance and mentorship to junior engineers. Communicate status, risks, blockers, and... 
    Suggested
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    6 days ago
  • $81.1k - $187k

     ...architect infrastructure and service to ensure reliability and functionality. Forecasts demands and...  ...impact and develops knowledge of site reliability trends.Only Oracle brings together...  ...Science, Information Technology, Engineering, or a related field, or equivalent practical... 
    Suggested
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    1 day ago
  • $157k - $190k

    What We NeedCorpay is currently looking to hire an Site Reliability Engineer within our Prepaid card division. This position falls under our Corporate Payments line of business and is located in Brentwood, TN. This role will improve the reliability, scalability, and operational... 
    Suggested
    Currently hiring
    Local area

    Corpay

    Brentwood, TN
    17 hours ago
  • $121.5k - $264.1k

     ...infrastructure and service and shares guidance on practices for reliability and functionality. Provides direction to ensure accurate...  ...experimenting with new technology, executing improvements, building site reliability knowledge, and providing clear data.Only Oracle brings... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle Corporation

    Nashville, TN
    3 days ago
  • $110.1k - $234.6k

    This role combines ethical hacking, vulnerability research, and clean-room reverse-engineering practices to understand how systems work, identify security weaknesses, and help engineering teams remediate them responsibly along with future proofing. This role works closely... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    1 day ago
  •  ...Site Reliability Engineer - Compute Focus Duration: 6 Months to Hire Location: On-Site 2-3 days/week in Nashville, TN Job Description: We are seeking a skilled Site Reliability Engineer with a focus on compute infrastructure to join our dynamic team. The ideal... 
    2 days per week
    3 days per week

    United IT

    Nashville, TN
    3 days ago
  •  ...BRCityNashvilleJob TypeFull Time Key responsibilitiesUBS Business Solutions US LLC is seeking an Associate Director, Tech Site Reliability Engineer in Nashville, TNAre you an innovative thinker? Do you enjoy delivering enhanced change capabilities across a range of business... 
    Flexible hours

    UBS

    Nashville, TN
    1 day ago
  • $105.79k - $141.05k

     ...delivers on-demand networking at scale. As Lead SRE, you'll own the reliability of that platform — partnering with operations teams and...  ..., and automation, and you'll coordinate across architecture, engineering, and systems development organizations to measurably improve... 
    Temporary work
    Remote work

    Lumen Inc

    Nashville, TN
    3 days ago
  • $146.3k - $306.4k

     ...conformance, and rapid incident mitigation. Influences silicon/board/firmware roadmaps to optimize for reliability, performance, and cost at hyperscale. Champions engineering excellence: coding standards, threat modeling, resource management, concurrency, and fault... 
    Principal
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Nashville, TN
    4 days ago
  • Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    Hermitage, TN
    4 days ago
  • $69.8k - $148.3k

     ...a bachelors degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent hands-on experience. We look for three or more years of experience in site reliability engineering, systems administration, infrastructure operations,... 
    Full time
    Flexible hours

    Oracle

    Nashville, TN
    2 days ago
  •  ...As a Sr. Principal Member of Technical Staff, you will work with other senior engineers and product management to define requirements for OCI’s storage infrastructure services. Expertise in one or more Public Cloud offerings is a plus. You will be expected to make substantial... 
    Principal

    Ll Oefentherapie

    Nashville, TN
    17 hours ago
  • $135.2k - $306.4k

    As a Sr. Principal Member of Technical Staff, you will work with other senior engineers and product management to define requirements for OCI’s storage infrastructure services. Expertise in one or more Public Cloud offerings is a plus. You will be expected to make substantial... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    17 hours ago
  • $135.2k - $306.4k

     ...readiness, and production support.• Drive improvements to workflow reliability, deployment safety, scalability, observability, and developer...  ....• Break down ambiguous platform problems into durable engineering solutions.• Set high standards for code quality, system design... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    1 day ago
  •  ...Overview Oracle Cloud Infrastructure is seeking a Principal Software Engineer to help build and operate OCI Search Service with OpenSearch. This individual-contributor role owns complex distributed-systems design and service excellence for a managed search and analytics... 
    Principal

    Ll Oefentherapie

    Nashville, TN
    3 days ago
  • $104.5k - $234.6k

     ...development workflows, and AI-powered engineering tools that accelerate software delivery...  ...across the enterprise.We are seeking a Principal Platform Software Engineer to define the...  ...highly available services that operate reliably across large-scale, distributed cloud environments... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    17 hours ago
  • $104.5k - $234.6k

     ...role in OCI’s next phase of growth.As a Principal Software Platform Engineer, you will define and drive the technical direction for reliable, secure, and delightful deployment capabilities...  ...orchestration, or platform architecture.Site reliability engineering, production... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    4 days ago
  • $142k - $194k

     ...We expect experience influencing silicon, board, and firmware roadmaps with a focus on reliability, performance, and cost at hyperscale. We look for strong knowledge of engineering best practices, including coding standards, threat modeling, resource management,... 
    Principal
    Full time
    Temporary work

    Oracle

    Nashville, TN
    11 days ago
  • $104.5k - $234.6k

     ...software development lifecycle; provides guidance and coaching to engineers to drive improvements.Utilizes advanced knowledge to develop...  ...to ensure service/product availability, health, support, and reliability.Core ResponsibilitiesPlanning & Execution:Manages and... 
    Principal
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Nashville, TN
    3 days ago
  • $104.5k - $234.6k

     ...hybrid, and multicloud solutions, edge computing, and more.As a Principal Platform Software Engineer, you will shape the platform foundations that allow OCI services and their consumers to work together reliably at scale. You will lead cross-team initiatives involving APIs... 
    Principal
    Temporary work
    Work at office
    Worldwide
    Relocation
    Relocation package
    Flexible hours

    Oracle Corporation

    Nashville, TN
    4 days ago
  • $104.5k - $234.6k

    PRINCIPAL PLATFORM SOFTWARE ENGINEER (IC4)OCI Log Analytics - Query Language and Distributed ExecutionLocation...  ..., Tennessee Work Arrangement: On-site Monday through Friday On-Call Requirement...  ..., correctness, availability, and reliability requirements.• Develop load, stress,... 
    Principal
    Temporary work
    Live in
    Relocation
    Relocation package
    Monday to Friday
    Flexible hours

    Oracle Corporation

    Nashville, TN
    2 days ago
  • $104.5k - $234.6k

    Overview Oracle Cloud Infrastructure is seeking a Principal Software Engineer to help build and operate OCI Search Service with OpenSearch. This...  ...service/product availability, health, support, and reliability.Core ResponsibilitiesPlanning & Execution:Manages and coordinates... 
    Principal
    Temporary work
    Flexible hours
    Shift work

    Oracle Corporation

    Nashville, TN
    17 hours ago
  • $135.2k - $306.4k

     ...customers who are tackling some of the world’s biggest challenges.Oracle Health Infrastructure and Platform Services is looking for Engineers with experience in building cloud-native, AI-assisted platforms and systems. This role is suited for a hands-on engineer with... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    3 days ago
  • $96.8k - $306.4k

     ...with broad customer impact and a direct role in OCI’s next phase of growth.As a Senior Principal Software Development Engineer, you will set the long-term technical vision for reliable, secure, and delightful deployment capabilities across the OCI Developer Platform. You... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    17 hours ago
  • $250.6k - $362.6k

     ...comprehensive security outcomes, as a Principal Engineer. The team delivers secure, scalable capabilities...  ...networking, with a strong emphasis on reliability, interoperability, and long-term...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Principal
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Nashville, TN
    6 days ago
  • $114.6k - $234.6k

     ...to define monetization architecture for next-generation video deliveryWork with a highly technical, distributed systems-focused engineering teamOnly Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    2 days ago
  • $96.8k - $223.4k

    As a member of the software engineering division, you will apply intermediate to advanced knowledge of software architecture to perform software development tasks associated with developing, debugging, or designing software applications or operating systems according to... 
    Principal
    Temporary work
    Relocation
    Flexible hours

    Oracle Corporation

    Nashville, TN
    4 days ago
  • $96.8k - $306.4k

     ...and support to do your best work. It is a dynamic and flexible workplace where you’ll belong and be encouraged.As a Lead Principal Software Engineer, you will work with teams of software engineers responsible for the software design, development, and operations for our... 
    Principal
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!