Principal Site Reliability Engineer
$84.9k - $209.5kOracle
Job Description
Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services. Design and develop designs, architectures, standards, and methods for large-scale distributed systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning
You will provide cloud operations for Oracle National Security Realms. You'll be part of a dynamic team with a broad knowledge of how Oracle's cloud platform works. You'll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers.
Note - this role is not a Monday to Friday core hours role - it will involve working a 24/7 shift rotation with on-call duties, including nights, weekends and public holidays.
Responsibilities
Escalation points for junior site reliability engineers during complex or high-impact incidents.
Manage and execute complex manual Change Management tickets, by working closely with the service teams to ensure safety and minimal disruption to services.
Support the on-boarding of new services and tools, ensuring they are operationally ready and properly integrated.
Provide mentorship and training to SREs, helping build team capability and confidence.
Create and maintain clear, useful documentation for operational processes and system support.
Identify areas of manual work and drive automation to reduce toil and improve efficiency.
Automate tasks to enable continuous delivery and ensure continuous availability with minimal human overhead
Recognize unsafe or inefficient practices and work with teams to design safer, more effective solutions.
Complete change requests to enable new functionality and maintain realm compliance
Ensure timely resolution of incidents, service requests, and change requests
Collaborate with global service and engineering teams
Define and drive change management, continuous integration, and deployment best practices
Help create and maintain real-world production architectures, scalability, and system design
Use a methodical approach to troubleshoot, large, complex, interconnected systems
We also use...
Linux and Unix operating systems
Docker, Kubernetes, and Terraform
Scripting languages such as Bash, shells, Perl, or Python
Citizenship/location requirements - i.e. US Citizenship, U.S. Citizenship and possess and maintain TS/SCI w/Poly security clearance, reside in Austin, TX or Reston, VA
Technology related bachelor's degree and/or equivalent work experience
A desire to learn and keep up with modern technologies
Proficient with writing services/task automation in any modern development language (e.g. Python, Bash, Ruby, Perl, JavaScript, or Java)
Familiarity with core protocols and OSI model (DNS, DHCP, TCP/IP)
Deep knowledge of Linux or Unix OS internals and host-based networking
Familiarity with configuration management solutions such as Chef, Puppet, etc
Experience with devising, managing, and extending monitoring solutions for large scale environments.
Knowledge of cloud computing concepts
Experience working in a mission-critical environment (Operations, Technical Support, NOC etc)
Proficient with communication skills (writing, organization, learning exchange)
Experience executing tasks under change management procedures
Experience resolving auto-cut and manual alarms following runbooks
A focus on customer satisfaction
Specific experience working with deployment of AI infrastructure to include clustered GPUs, LLM deployment and maintenance, and understanding of model integration for customer solutions
Certifications in VMware or other hypervisors will stand out
Certifications in Cisco (e.g. CCNA, CCNP) will stand out
Certifications in CISSP or other security related will stand out
Certifications in Oracle Databases or RAC will stand out
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $84,900 to $209,500 per annum. May be eligible for bonus and equity.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
Medical, dental, and vision insurance, including expert medical opinion
Short term disability and long term disability
Life insurance and AD&D
Supplemental life insurance (Employee/Spouse/Child)
Health care and dependent care Flexible Spending Accounts
Pre-tax commuter and parking benefits
401(k) Savings and Investment Plan with company match
Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
11 paid holidays
Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
Paid parental leave
Adoption assistance
Employee Stock Purchase Plan
Financial planning and group legal
Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC4
About Us
Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing View email address on click.appcast.io or by calling View phone number on click.appcast.io in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
- ...Principal Site Reliability Engineer About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we power shipping logic and checkout optimization...PrincipalFull timeWork at office
- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,...SuggestedWork at officeLocal area
- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SuggestedTemporary workCasual workWorldwide
- ...Dimensional leverages the rapidly evolving state of the art to engineer scalable, innovative, and research driven solutions to improve... ...each of the developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including Python toolchains (...SuggestedFull timeLocal area
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SuggestedWork at officeLocal areaRemote workWorldwideFlexible hours- ...selected candidate for this role to work on site in the specified location(s).Schwab... ...their money by delivering innovative and reliable technology solutions that support investing... ...Within the Bank Platform Operations and Engineering organization, you will help ensure the...Full timeWork at office
$167.7k - $245.2k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$127k - $249k
...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This...Local areaRemote workWorldwideFlexible hours$196k - $269.5k
Senior Principal AI Agent EngineerThe Software Engineering team delivers next-generation software application enhancements and new products for a changing world. Working at the cutting edge, we design and develop software for platforms, peripherals, applications and diagnostics...Principal$109.65k - $182.76k
...encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication...Full timeLocal area3 days per week$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...Full timeWork at office$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...Local areaRemote workWorldwideFlexible hours- ...Senior Site Reliability Engineer Apptronik is a human-centered robotics company developing AI-powered robots to support humanity in every facet of life. Our flagship humanoid robot, Apollo, is built to collaborate thoughtfully with people, starting with critical industries...Full time
- ...Senior Site Reliability Engineer Come join a growing bank at the heart of the innovation, technology, green tech and life sciences space. We continue to expand our global footprint and our banking technology is at the core of everything we do. As a Senior Site Reliability...
$167.18k - $203.61k
...remotely part of the weekTravel %NoWork ShiftJob DescriptionCox Automotive Corporate Services, LLCLEAD SITE RELIABILITY ENGINEERJob Description: Lead Site Reliability Engineer positions offered by Cox Automotive Corporate Services, LLC (Austin, Texas). Lead the...Full timeWork at officeRemote workFlexible hours- Job Description:About the Role:We are looking for a Senior SRE to join our Platform Engineering team where you’ll own the reliability, scalability, and operational excellence of our workflow orchestration platforms - primarily Apache Airflow and Broadcom Automic/UC4. This...Full time
- ...innovators who want to make an impact on the world of technology.Cadence Design Systems Inc. is looking for a motivated DevOps Sr Principal Software Engineer to work with us in Austin, Texas. At Cadence, we hire and develop leaders and innovators who want to impact the world of...PrincipalFull time
$85 - $90 per hour
...hr - $90.00/hr Date Posted: 04/01/2025 Hiring Organization: Rose International Position Number: 480571 Job Title: Senior Principal Software Engineer Work Model: Onsite Employment Type: Temporary Min Hourly Rate($): 85.00 Max Hourly Rate($): 90.00 Job Description ***Only...PrincipalHourly payTemporary workFlexible hours$140k - $215k
...processing trillions of events per day. As a Principal SRE, you will operate at the intersection of our Core Platform and Embedded Reliability charters: building the foundational... ..., while embedding directly with product engineering teams and their leadership to drive reliability...Full timeWork experience placementWork at officeLocal area2 days per week3 days per week$196k - $364k
...innovators who want to make an impact on the world of technology.Cadence Design Systems is looking for a highly motivated hardware engineer to work with the Modus R&D engineering team in the Design-For-Test (DFT) IP business unit.As a member of the DFT R&D team, you...PrincipalFull time- ...debug methodologies, and serviceability requirements.Partner with silicon, platform, firmware, software, validation, and customer engineering teams.Influence future platform management architecture and industry standards adoption.REQUIRED EXPERIENCEStrong background in...Principal
- ...territory, and do work that genuinely matters, Future Secure AI is the place for you. About the Role We are looking for a Site Reliability Engineer to help design, build, and operate the platforms that power AI Co-Workers. This is a hands-on role for an engineer who...Flexible hours
- ...assisted developers or autonomous agents is reliable, secure, and maintainable.Integrating... ...descriptionAs a member of one of our engineering teams, you'll be a key player in making... ...team of Engineers (Cloud Engineers and Site Reliability Engineers), providing guidance...Full timeRelocationFlexible hours
$140k
...Title: Site Reliability Engineer SRE - ML platform Location: Austin, TX OR Sunnyvale, CA Type: FTE Salary/Rate : $140K Title: Site Reliability Engineer SRE - ML platform Responsibilities - Continuous Deployment using...$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Site Reliability Engineer (SRE) Location: Austin, TX We’re searching for a driven Site Reliability Engineer (SRE) to join our innovative team. As an SRE, you’ll be a cornerstone of our production software, ensuring our systems are uncompromisingly reliable, secure...
- ...2+ years related experience in managing Cloud Infrastructure ~ Bachelor's degree in an applicable field, such as CS, CIS or Engineering. Preferred Qualifications: Experience working with CI/CD pipelines Experience working with Containers or Kubernetes...
- ...fully intend for the selected candidate for this role to work on site 4-days per week during night shifts (2pm-10pm CST), and weekends as needed, in the specified location(s). As a Site Reliability Engineer, you will play a critical role in protecting the stability,...Permanent employmentWork at officeNight shiftWeekend work
- ...Role: Site Reliability Engineer Rate: Location: Austin, TX (Hybrid 2 days onsite in a week, locals to TX) Visa: USC/GC/EAD/OPT Duration: 12+ months Client: Must have public sector (state client) at least on 1 project & focus on industry exp...Local area2 days per week
$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working... ...they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA...PrincipalFull timeRemote workShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- principal network engineer Austin, TX
- principal application developer Austin, TX
- civil engineer project manager Austin, TX
- principal security engineer Austin, TX
- principal engineer Austin, TX
- principal battery engineer Austin, TX
- principal infrastructure engineer Austin, TX
- director quality engineering Austin, TX
- senior civil engineer project manager Austin, TX
- senior chief engineer Austin, TX


