Principal Site Reliability Engineer
hackajob
Job Title
Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence. Design, write, and deploy software to improve the availability, scalability, and efficiency of Oracle products and services. Design and develop designs, architectures, standards, and methods for large-scale distributed systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning You will provide cloud operations for Oracle National Security Realms.
Responsibilities
Escalation points for junior site reliability engineers during complex or high-impact incidents.
Manage and execute complex manual Change Management tickets, by working closely with the service teams to ensure safety and minimal disruption to services.
Support the on-boarding of new services and tools, ensuring they are operationally ready and properly integrated.
Provide mentorship and training to SREs, helping build team capability and confidence.
Create and maintain clear, useful documentation for operational processes and system support.
Identify areas of manual work and drive automation to reduce toil and improve efficiency.
Automate tasks to enable continuous delivery and ensure continuous availability with minimal human overhead
Recognize unsafe or inefficient practices and work with teams to design safer, more effective solutions.
Complete change requests to enable new functionality and maintain realm compliance
Ensure timely resolution of incidents, service requests, and change requests
Collaborate with global service and engineering teams
Define and drive change management, continuous integration, and deployment best practices
Help create and maintain real-world production architectures, scalability, and system design
Use a methodical approach to troubleshoot, large, complex, interconnected systems
Technologies
Linux and Unix operating systems
Docker, Kubernetes, and Terraform
Scripting languages such as Bash, shells, Perl, or Python
Citizenship/location requirements - i.e. US Citizenship, U.S. Citizenship and possess and maintain TS/SCI w/Poly security clearance, reside in Austin, TX or Reston, VA
Technology related bachelor's degree and/or equivalent work experience
A desire to learn and keep up with modern technologies
Proficient with writing services/task automation in any modern development language (e.g. Python, Bash, Ruby, Perl, JavaScript, or Java)
Familiarity with core protocols and OSI model (DNS, DHCP, TCP/IP)
Deep knowledge of Linux or Unix OS internals and host-based networking
Familiarity with configuration management solutions such as Chef, Puppet, etc
Experience with devising, managing, and extending monitoring solutions for large scale environments.
Knowledge of cloud computing concepts
Experience working in a mission-critical environment (Operations, Technical Support, NOC etc)
Proficient with communication skills (writing, organization, learning exchange)
Experience executing tasks under change management procedures
Experience resolving auto-cut and manual alarms following runbooks
A focus on customer satisfaction
Specific experience working with deployment of AI infrastructure to include clustered GPUs, LLM deployment and maintenance, and understanding of model integration for customer solutions
Certifications
Certifications in VMware or other hypervisors will stand out
Certifications in Cisco (e.g. CCNA, CCNP) will stand out
Certifications in CISSP or other security related will stand out
Certifications in Oracle Databases or RAC will stand out
Qualifications
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC4
- ...Principal Site Reliability Engineer About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we power shipping logic and checkout optimization...PrincipalFull timeWork at office
- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,...SuggestedWork at officeLocal area
- ...Description:About the Role: We are looking for a Senior SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering...SuggestedFull time
- ...across multiple clouds and regions while partnering with network engineers, systems architects, and game studio developers. This is an ownership role: driving technical direction, influencing reliability from architecture review through production operation, and closing...Suggested
- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SuggestedTemporary workCasual workWorldwide
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours- ...Dimensional leverages the rapidly evolving state of the art to engineer scalable, innovative, and research driven solutions to improve... ...each of the developer tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including Python toolchains (...Full timeLocal area
- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical...Full timeWork at office
$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,...Full time$165k - $241.4k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$127k - $249k
...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This...Local areaRemote workWorldwideFlexible hours$196k - $269.5k
Senior Principal AI Agent EngineerThe Software Engineering team delivers next-generation software application enhancements and new products for a changing world. Working at the cutting edge, we design and develop software for platforms, peripherals, applications and diagnostics...Principal$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...Full timeWork at office- ...encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication...Full timeLocal area3 days per week
- ...Artificial Intelligence at Schwab. We are an integrated product, engineering, strategy and risk team, all based in San Francisco. We help... ...the most exciting areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts for cutting-...Full time
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...Local areaRemote workWorldwideFlexible hours$152k - $195k
...Senior Site Reliability Engineer Austin, TX (Hybrid) SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Alex Yampolskiy and...- ...commercialization, and mass production to change the world for the better. JOB SUMMARY We are seeking an experienced Site Reliability Engineer to own and maintain the deployment of our cloud-based infrastructure to customer sites. In this role, you will work...Full timeLocal area
- ...Job Description Job Description Sr. Software Engineer - Site Reliability About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we...Full timeWork at office
$167.18k - $203.61k
...remotely part of the weekTravel %NoWork ShiftJob DescriptionCox Automotive Corporate Services, LLCLEAD SITE RELIABILITY ENGINEERJob Description: Lead Site Reliability Engineer positions offered by Cox Automotive Corporate Services, LLC (Austin, Texas). Lead the...Full timeWork at officeRemote workFlexible hours$110.7k - $171.8k
...components Participation in oncall rotation as a platform reliability escalation point Incident response, postincident reviews... ..., and internal control requirements. Collaborate with engineering teams across the organization to influence platform adoption,...Work experience placementWork at officeLocal area$121.4k - $218.6k
...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner with... ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling robust...Work experience placementWork at office- ...the layer where data becomes decisions, and decisions make the advantage. About the Role Gallatin is looking for a Site Reliability Engineer to keep our production systems running with the reliability our national security customers require. You'll work at the...Full timeLocal area
- ...foundation of success and bringing it to the digital space - ready to join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will...Remote workFlexible hoursNight shift
- ...innovators who want to make an impact on the world of technology.Cadence Design Systems Inc. is looking for a motivated DevOps Sr Principal Software Engineer to work with us in Austin, Texas. At Cadence, we hire and develop leaders and innovators who want to impact the world of...PrincipalFull time
$81.1k - $187k
...infrastructure and/or service according to terms for reliability and functionality. ~Assists team... .... ~Gains basic knowledge of site reliability trends and shares relevant information... ...are seeking a skilled Site Reliability Engineer to design, build, operate, and automate...Temporary workImmediate startFlexible hoursShift work$196k - $364k
...innovators who want to make an impact on the world of technology.Cadence Design Systems is looking for a highly motivated hardware engineer to work with the Modus R&D engineering team in the Design-For-Test (DFT) IP business unit.As a member of the DFT R&D team, you...PrincipalFull time$140k - $215k
...processing trillions of events per day. As a Principal SRE, you will operate at the intersection of our Core Platform and Embedded Reliability charters: building the foundational... ..., while embedding directly with product engineering teams and their leadership to drive reliability...Full timeWork experience placementWork at officeLocal area2 days per week3 days per week$272k - $431.25k
We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working... ...they can deploy, monitor, and operate these systems reliably at fleet scale. In this role, you will collaborate with NVIDIA...PrincipalFull timeRemote workShift work- Senior Principal Software Engineer (ServiceNow Information Architect)Be a part of a team that’s ensuring Dell Technologies' product integrity and customer satisfaction. Our IT Software Engineer team turns business requirements into technology solutions by designing, coding...Principal
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- principal network engineer Austin, TX
- senior director engineering Austin, TX
- civil engineer project manager Austin, TX
- principal developer Austin, TX
- chief design engineer Austin, TX
- principal test engineer Austin, TX
- principal engineer Austin, TX
- director data engineering Austin, TX
- principal infrastructure engineer Austin, TX
- technical director engineering Austin, TX


