Principal Site Reliability Engineer
$84.9k - $209.5kOracle Corporation
Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment when performing data collection to maintain and optimize operations and reliability. Leverages advanced knowledge to perform incident response and/or maintenance tasks. Provides comprehensive health and performance reporting. Identifies and recommends opportunities for automation. Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes. Provides comprehensive support for technology and documents incidents. Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends.Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on click.appcast.io or by calling View phone number on click.appcast.io in the United States.Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.Disclaimer:Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.Range and benefit information provided in this posting are specific to the stated locations onlyUS: Hiring Range in USD from: $84,900 to $209,500 per annum. May be eligible for bonus and equity.Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.Oracle US offers a comprehensive benefits package which includes the following:1. Medical, dental, and vision insurance, including expert medical opinion2. Short term disability and long term disability3. Life insurance and AD&D4. Supplemental life insurance (Employee/Spouse/Child)5. Health care and dependent care Flexible Spending Accounts6. Pre-tax commuter and parking benefits7. 401(k) Savings and Investment Plan with company match8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.9. 11 paid holidays10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.11. Paid parental leave12. Adoption assistance13. Employee Stock Purchase Plan14. Financial planning and group legal15. Voluntary benefits including auto, homeowner and pet insuranceThe role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment. /* Font Definitions */ @font-face {font-family:"Cambria Math"; } @font-face {font-family:Calibri; } @font-face {font-family:"Oracle Sans Light"; } /* Style Definitions */ p.MsoNormal, li.MsoNormal, div.MsoNormal {margin-top:0in; margin-right:0in; margin-bottom:8.0pt; margin-left:0in; line-height:107%; font-size:11.0pt; font-family:"Calibri",sans-serif;} p.MsoFooter, li.MsoFooter, div.MsoFooter { margin:0in; font-size:11.0pt; font-family:"Calibri",sans-serif;} span.FooterChar { } .MsoChpDefault {font-family:"Calibri",sans-serif;} .MsoPapDefault {margin-bottom:8.0pt; line-height:107%;} /* Page Definitions */ @page WordSection1 {size:8.5in 11.0in; margin:1.0in 1.0in 1.0in 1.0in;} div.WordSection1 {page:WordSection1;} /* List Definitions */ ol {margin-bottom:0in;} ul {margin-bottom:0in;} Key ResponsibilitiesCapacity Ingestion and Management:- Designs and architects infrastructure and/or service according to terms for reliability and functionality.- Forecasts demands for infrastructure and responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads and identifying resource gaps.- Collaborates with the software development team to develop infrastructures, ensuring features are reliable and scalable according to deployment requirements.- Proactively identifies opportunities for prototyping and drives prototyping initiatives (e.g., testing new applications or infrastructures, assisting in onboarding) to explore novel approaches.Incident and Service Lifecycle Management:- Exercises judgment when performing data collection, triage, technical analysis, and redirection to maintain and optimize operations and infrastructure reliability.- Takes proactive steps to monitor services, maintain up-to-date knowledge of their performance, and document their condition.- Leverages advanced knowledge to perform incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).- Provides comprehensive health and performance reporting and takes appropriate actions based on trends in data.- May perform provisioning to support infrastructure, applications, and services.- May experiment with new approaches for and performs decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.Automation:- Identifies and recommends opportunities for automation and assesses potential benefits to enhance operational efficiency.- Develops and implements design, automation tools, or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures.- Conducts testing on moderately complex automations to ensure they perform tasks correctly and produce expected results.Technical Communication and Guidance:- Writes release notes and/or communicates comprehensive information about the scale, capacity, security, performance attributes, and requirements of services and technology with customers and immediate and related teams.- Proactively anticipates and articulates the potential impact of infrastructure, feature, and tool changes, considering their impact across team operations.- Serves as a resource to team members on what information to communicate and how to communicate.Troubleshooting and Resolution:- Provides comprehensive operational support for technology, serving as a key escalation point for incidents and moderately complex issues arising within Oracle services.- Drives and actively participates in on-call shifts to address issues.- Executes the resolution of technical issues spanning multiple services, applying advanced investigation and debugging techniques to achieve SLOs (service level objectives).- Documents incidents according to reporting methods and performs root cause analyses, capturing essential information for analysis and future reference.- Performs post-mortem procedures to prevent incident reoccurrence.Innovation and Improvement:- Conducts advanced experiments and evaluations of cutting-edge tools and technologies to optimize infrastructure performance and reliability, taking proactive steps to adhere to security standards.- Identifies and seeks opportunities to execute improvements for performance bottlenecks and deployments, ensuring efficient resource usage, speed, and scalability.- Develops and maintains advanced knowledge of site reliability trends, sharing valuable insights and information with senior team members, management, and beyond to promote innovative building, testing, deploying, and running services.- Performs moderately complex analyses and provides clear data on production to drive business development decisions (e.g., design changes).Core ResponsibilitiesPlanning & Execution:- Manages and coordinates moderately complex tasks, monitoring timelines and deliverables to ensure timely completion and adherence to requirements for a moderately sized project or initiative. Efficiently delegates, monitors, and prioritizes work across multiple projects, providing technical oversight and adjusting plans to address shifts in resources or timelines.Collaboration & Partnership:- Collaborates across the organization to align on expectations and achieve shared objectives. Leverages understanding of business leaders, stakeholders, and/or customers to ensure proposed solutions meet their needs. Supports inclusivity by actively seeking and listening to diverse perspectives, ensuring others feel heard and respected.Problem Solving:- Identifies and addresses moderately complex issues by analyzing a wide range of data and/or information to identify solutions in accordance with standard practices. Proactively escalates unresolved or critical issues with a thorough assessment and suggests potential solutions. Reviews, contributes to, and documents problem solving strategies.Continuous Learning:- Pursues learning opportunities to expand knowledge and skills and/or tools in new areas and stays abreast of the latest industry trends and best practices. Proactively seeks and leverages ongoing feedback and training to improve skills. Coaches and mentors junior team members, fostering continuous learning and knowledge sharing within and across teams.Continuous Improvement:- Develops ideas, recommends updates, and/or collaborates on the implementation of process improvements to increase the efficiency and effectiveness of processes, protocols, and workflows across teams, and evaluates the impact on key stakeholders. Solicits feedback from others on ideas for alternative approaches and methods for continued improvement.Performance and Development:- Contributes to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.- Full timePosting Date: 2026-09-17
- ...Continuous Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives... ...equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation)...PrincipalFull time
- IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...SuggestedWork at officeImmediate start
$152.6k - $191.5k
...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include composing... ...improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior individual...SuggestedFull timeWork at officeDay shift- Role Profile:We are evolving our Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division.As a Senior SRE, you will be a senior hands‑on technical person help...SuggestedFull timeShift work
$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,...SuggestedFull time$118.6k - $195.68k
Job SummaryThe Red Hat IT OpenShift team is looking for a Senior Site Reliability Engineer (SRE) to design, develop, scale, and operate our Red Hat Hybrid OpenShift Platforms (on-prem & cloud). As a Senior Engineer, you will contribute to running Red Hat OpenShift at scale...Permanent employmentFull timeContract workWork experience placementWork at officeRemote workFlexible hours- ...We are seeking a Senior SRE / DevSecOps Engineer with strong experience in Kubernetes, AWS... .... The role will focus on platform reliability, incident management, SLO/SLI governance... ...Experience with SLO/SLI governance and site reliability practices. ~ Strong understanding...
- ...home!Where you’ll be:This position will be based at our Corporate Headquarters located in Charlotte, NC.About the Role:The Site Reliability Engineer plays a critical role in designing, building, and maintaining scalable, secure, and highly available cloud infrastructure...Full timeFlexible hours
- ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises...Permanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
- ...English (Required)Work Shift:1st Shift (United States of America)Please review the following job description:Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)We are seeking a Lead Site Reliability & Environment Monitoring Engineer to...Full timeTemporary workShift workDay shift
- ...some of the most important challenges in global education. Client is currently seeking a talented Software Engineer who is able to work into the Site Reliability Engineer role. This candidate is expected to work towards becoming a Subject Matter Expert in the cloud space...Remote work
$175k - $250k
...includes rapidly growing SaaS companies like OpenAI, Cursor, Perplexity, Vercel, Plaid, and hundreds of others. About The Site Reliability Engineering Team The Site Reliability Engineering (SRE) team ensures the WorkOS platform remains fast, reliable, and resilient at...Local areaRemote work- ...Collins Aerospace, part of RTX, seeks a Senior Principal Software Engineer for onsite work in Winston-Salem or Jamestown to lead the cargo software team and shape the next generation of aircraft cargo systems. You will coordinate with Systems Engineers, drive software...Principal
- ...NVIDIA seeks a Senior Staff Site Reliability Operations Technical Lead in Durham, NC to own local service delivery, lead critical incidents... ...across AD, Exchange, and compute platforms. You will mentor engineers, shape site runbooks, and partner with global teams on...Local area
$184k - $264.5k
...Senior Staff Site Reliability Operations Technical Lead For over 25 years, NVIDIA has been at the forefront of transforming computer graphics... ...teams, and provides technical leadership to site support engineers. Be responsible for the hardest issues across Active...Permanent employmentWork at officeLocal areaRelocation$60 - $65 per hour
...retail industries. Rate Range: $60-$65/Hr Job Description: The Client Document Generation team is seeking a Senior Software Engineer ( IT Onshore Band 4) to participate in the full system development lifecycle (SDLC) of enterprise applications that support high-...Immediate start- ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally...Contract workLocal area
- ...SoftPro has won this prestigious award 14 times since 2012! What are we looking for? SoftPro is seeking a well-rounded Site Reliability Engineer (SRE) to join our Cloud Operations Team in our Raleigh, NC office or as a remote employee. This team supports our...Hourly payWork at officeRemote work
- ...driven solutions through our 5 core specialist Practices; Software Engineering, Data & Analytics, IT Operations, Change & Transformation, and Risk, Regulation & Compliance. FDM is seeking a Site Reliability Engineer located in New York City, NY or Alpharetta, GA to...Work at officeRelocationVisa sponsorship3 days per week
- ...opportunities to advance your career while creating software that changes the world. Your role and responsibilities As a Site Reliability Engineer, you will work in an agile, collaborative environment to build, deploy, configure, and maintain systems for the IBM client...Full timeContract workPart timeFixed term contractInternshipWorldwideFlexible hoursShift work
- ...the world of technology.Role SummaryCadence is hiring a Principal AI Forward Deployment Engineer to embed with strategic semiconductor customers and... ...story from pilot to production rollout across BUs and sites.Trusted Technical Advisor. Be the senior technical face...PrincipalFull timeRelocation
- ...environment. Whether it’s building award-winning games or crafting engine technology that enables others to make visually stunning... ...computing users in the world.What You'll DoAs Technical Lead / Principal Engineer on the AI Engineering team, you'll provide technical leadership...Principal
$132.4k - $251.6k
...a rapidly evolving global market.The Secure Subsystems Engineering team seeks a Senior Principal Systems Engineer - Lead Systems Engineer with a passion... ...locations, regardless of whether the role is designated as on-site, hybrid or remote.The salary range for this role is 132...PrincipalTemporary workWork experience placementWork at officeRemote workRelocation packageFlexible hours$272k - $431.25k
...make a lasting impact on the world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software... ...-world deployment and integration needs. Establish reliability, security, validation, and left-shift strategies...PrincipalFull timeRemote workShift work- About this role:Wells Fargo is seeking a Principal Engineer, Platform, Cloud Engineering and DevOps... ...OpenShift, observability, and platform reliability. This role drives a reliability-first... ...Infrastructure Engineering, Cloud Engineering, Site Reliability Engineering (SRE), or...PrincipalFull timeWork experience placement
$145k - $185k
...maturity of data center and critical systems engineering by translating enterprise strategy into... ...structured cabling systems, ensuring reliability, efficiency, and compliance with... ...integration activities within a large or multi‑site environment. Experience leading teams or...PrincipalFull timeContract workWork at officeLocal areaWorldwideDay shift- ...Full Stack Engineer Duration: Long Term Contract Location: Westlake, TX / Durham, NC Job Description: Seeking a Principal Software Engineer to develop enterprise-wide data capabilities pertaining to our customers’ Communication Preferences & Profiles. The...PrincipalLong term contract
- ...systems, and policy-as-code enforcement engines that integrate self-service developer experiences... ...tolerance patterns that ensure platform reliability at enterprise scale. Applies modern... ...plans, please visit our Benefits site. Depending on the position and division,...PrincipalFull timePart timeShift workDay shift
$272k - $431.25k
...NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI Infrastructure. We expect you to have... ...health management capabilities that enable industry leading reliability, availability, and scalability of GPU assets. You will be harnessing...PrincipalFull time- ...development and designs meet scalability, reliability, security, and performance requirements... ...Bachelor’s degree in Computer Science, Engineering, Information Technology, Information... ...and five (5) years of experience as a Principal Software Engineer/Developer (or closely...PrincipalFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- general engineer North Carolina
- chief engineer North Carolina
- principal developer North Carolina
- engineering director North Carolina
- data center chief engineer North Carolina
- hotel chief engineer North Carolina
- principal engineer North Carolina
- director software engineering North Carolina
- senior principal cloud computing engineer North Carolina
- principal North Carolina


