Principal Site Reliability Engineer
Akamai Technologies
Principal Site Reliability Engineer
Are you passionate about enabling next-generation server hardware and Linux platforms while building reliable, high-performance infrastructure for a global cloud platform?
Are you excited about solving complex server reliability challenges and delivering hardware, firmware, and Linux solutions for Akamai's global cloud infrastructure?
Join our highly skilled Site Reliability Engineering team
Our HIVE Server Enablement & Qualification (SEQ) team designs, develops, and manages applications and infrastructure supporting Akamai's Compute products and services. HIVE is responsible for supplying our fleet with hardware, the host and guest Linux platforms. We make life better for billions of people, billions of times a day.
Partner with the best!
As a Principal Site Reliability Engineer, you will be at the forefront of Akamai Cloud hypervisor host technologies. SEQ works to enable new hardware via our provisioning service, provides seamless firmware upgrade models, and works with other HIVE teams. Together, we are providing cohesive and performant solutions to our internal infrastructure customers and operators. We ensure reliable server behavior cross all our datacenters.
As a Principal Site Reliability Engineer, you will be responsible for:
- Shaping compute-platform strategy and qualification: define direction and criteria for x64, ARM, accelerator, and inference hardware across firmware, OS, runtime, reliability.
- Leading provisioning and CI/CD improvements: designing reproducible, safe, and scalable bare-metal provisioning, imaging configuration, and infrastructure delivery pipelines.
- Driving reliability and observability: addressing systemic failures, enhance telemetry and diagnostics, and establishing safe rollout and recovery safeguards for production.
- Providing incident expertise: investigate complex HW/SW incidents, identify root causes, guide decision-making, and convert findings into engineering fixes.
- Acting as a senior technical force multiplier: reviewing designs, mentoring engineers, establishing standards, and resolving multi-domain issues across teams and vendors.
Do what you love
To be successful in this role you will:
- Have a Linux and server-platform expertise: knowledge of Linux systems, boot processes, storage, networking, hardware diagnostics, firmware, BIOS/UEFI, BMCs, server lifecycles.
- Have a provisioning & automation background: experience designing and operating bare-metal provisioning, configuration management, and scalable infrastructure automation.
- Demonstrate software & CI/CD engineering skills: fluency in Python and BASH, with experience building automation, APIs, deployment pipelines, and debugging multi-layer failures.
- Possess skills identifying whether hardware-test failures originate from firmware, BIOS/UEFI configuration, device firmware, kernel/driver behavior, or the physical test environment.
- Possess a Reliability & operations experience: expertise in observability, metrics, capacity analysis, incident root-cause investigation, and production readiness.
- Display practical knowledge of x86 and ARM platforms, accelerators, and inference infrastructure, including drivers, runtimes, and compatibility.
- Have hands-on experience with Linux virtualization technologies, including KVM, QEMU and libvirt.
- Have understanding of nested virtualization, CPU virtualization extensions, VM networking and storage, and Linux-level troubleshooting of virtualized environments.
About us
At Akamai, we make life better for billions of people, trillions of times a day. Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple: Cloud and Edge: Running apps closer to users for instant performance. Security: Neutralizing threats before they ever reach your data. Content Delivery: Scaling the world's biggest moments without a glitch. AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.
Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.
Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work. We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.
Connect with us on social and see what life at Akamai is like!
- ...Principal Site Reliability Engineer Location: New York, NY (Onsite) Job type: Contract Job Description: Job Requirements Must Have: - Site Reliability Engineering and system reliability optimization - GitLab CI/CD and HashiCorp Vault secrets management -...PrincipalContract work
- ...Setting the reliability strategy for the platform, the full-time Principal Site Reliability Engineer will define deployment and operational standards for distributed systems, ensuring reliability and automation across customer environments while working remotely. Key...PrincipalFull timeRemote work
$139.7k - $232.9k
...designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence...PrincipalFull timeWork experience placement$142.8k - $274.8k
...yearEmployment type: Full-TimeWork site: 0 days / week in-office - remoteRole... ...Software EngineeringDiscipline: Site Reliability EngineeringCompany:... ...world’s most demanding workloads. As a Principal Site Reliability Engineer, you will set technical and operational...PrincipalOngoing contractWork at officeLocal area- ...lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.We are seeking a Principal Site Reliability Engineer (SRE) to define and scale reliability practices across large-scale cloud platforms.This is a senior individual contributor...PrincipalMinimum wageFull timeWork experience placementWork at officeLocal areaRemote work
$134.6k - $230.8k
...Connecting. Growing together.Are you passionate about reimagining operations through AI? Optum Financial is seeking a Principal Site Reliability Engineer to lead the evolution of our reliability platform by combining modern SRE practices with AI-assisted operations. You'...PrincipalMinimum wageFull timeWork experience placementWork at officeLocal areaRemote work$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production... ...strengthen the reliability of production environments.As a Principal SRE, you will shape the technical direction of NVIDIA’s...PrincipalFull time- ...role will be supporting products and services, including Gen Insurance and Engine by Gen, delivering trusted digital and financial experiences to consumers at scale. As a Principal Site Reliability Engineer, you will define and lead the reliability, scalability, and...Principal
$190k - $220k
...anticipating our clients’ needs and exceeding their expectations. About the RoleWe are looking for a skilled and motivated Principal Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance of...PrincipalFull timeWorldwideFlexible hours- ...Continuous Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives... ...equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation)...PrincipalFull time
$84.9k - $209.5k
Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs... ...new tools and develops and maintains advanced knowledge of site reliability trends.Only Oracle brings together the data, infrastructure...PrincipalTemporary workImmediate startFlexible hoursShift work$84.9k - $209.5k
This role combines strategic architecture with practical systems engineering, deployment, automation, patching, troubleshooting, incident response, and compliance support. The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud Infrastructure...PrincipalTemporary workFlexible hours- ...Principal Site Reliability Engineer As a Principal Site Reliability Engineer, you set the reliability strategy for the platform. You will define how we build, deploy, observe, and operate a distributed system that runs both in our own cloud and inside customer-controlled...PrincipalRemote work
- ...Principal Site Reliability Engineer The Principal Site Reliability Engineer will be a critical technical leader responsible for driving the operational excellence, resilience, and security of our core systems for a key Randstad client in the Washington D.C. area. This...Principal
$200k - $250k
...build the future together. The Crown Is Yours As a Principal Site Reliability Engie r , you'll shape the long-term strategy for the infrastructure... ...direction of our cloud and on-premise platforms, helping engineering teams build, deploy, and operate highly reliable systems...PrincipalFull timeImmediate start$169.3k - $304.7k
...building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...the growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting...PrincipalWork experience placementWork at officeRemote work$84.9k - $209.5k
.... You’ll partner with customer support, service owners, and engineering teams around the globe to ensure high-quality service for customers... ...posted.Career Level - IC4Escalation points for junior site reliability engineers during complex or high-impact incidents.Manage and...PrincipalTemporary workMonday to FridayFlexible hoursShift workNight shift$198.24k - $272.58k
We’re looking for a Principal Site Reliability Engineer to join Procore’s Compute Division to work on our FedRAMP initiative. In this role, you’ll help build Procore’s next-generation construction compute platform for others to build upon, including Procore developers,...PrincipalFull timeContract workWork at officeLocal areaImmediate start- ...Infrastructure Code. Builds reliability into the ecosystem by applying... ...practices in resiliency engineering and observability by developing... ...engineering techniques with site reliability engineering... ...5) years of experience as a Principal Site Reliability Engineer (or...PrincipalFull time
$159k - $272k
...generosity. Join us for the opportunity to grow and make a difference in ways that matter to you. Role SummaryIn this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (...PrincipalFull timePrivate practiceLocal areaRemote workWork from home3 days per week- ...Principal Site Reliability Engineer Deimos is a cloud-native developer and security operations technology services company. We help companies of all sizes adopt the cloud for improved service delivery to their clients. We're a fully remote African-based team of engineers...PrincipalCurrently hiringRemote workWork from home
$96.3k - $264.1k
...infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full... ...tools and provides expertise in site reliability trends.Only Oracle brings... ...LeadershipDefine and drive the site reliability engineering strategy for large-scale, distributed,...PrincipalTemporary workFlexible hours$165k - $185k
...the work is personal, and we are committed to the cause. Learn more at tandemdiabetes.com A DAY IN THE LIFE: The Principal Site Reliability Engineer (SRE) is responsible for the reliability, availability, and performance of the company's production systems. This...PrincipalPermanent employmentContract workLocal areaRemote workFlexible hoursShift work$194k - $237k
## Principal Site Reliability EngineerApplylocations: Scottsdaletime type: Full timeposted on: Posted 5 Days Agojob requisition id: REQ2026426At... ....**Overall Purpose**The Principal Site Reliability Engineer partners with development teams by designing availability...PrincipalHourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours- ...Principal Site Reliability Engineer LivePerson transforms customer care from voice calls to mobile messaging. Our cloud-based software platform, LiveEngage, allows brands with millions of customers and tens of thousands of care agents to deliver digital experiences...PrincipalLocal areaRemote work
$132.6k - $214.5k
...As part of this role, you will collaborate closely with our engineering teams to develop innovative solutions that provide clear and... ...team to influence the operability of the product and ensure the reliability and availability of our services. Qualifications DevOps...PrincipalFull timeWork at officeVisa sponsorshipWork visa$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics...PrincipalFull timeWork experience placementWorldwide- ...Principal Site Reliability Engineer Your work will help clinicians and healthcare staff access applications consistently, reduce disruption from platform changes, and accelerate recovery when problems occur. You will help the team build a platform that is easier to...PrincipalTemporary workRemote workFlexible hours
- ...with software development teams to build reliable, scalable, secure, and cloud-native... ...influence scalable architecture patterns across engineering teams, helping ensure systems are... ...~8+ years of hands-on experience in Site Reliability Engineering, DevOps, cloud infrastructure...PrincipalRemote work
$152.3k - $246.4k
...Palo Alto Networks runs a large infrastructure and is one of the largest Google Cloud Platform customers. As a Principal Site Reliability Engineer for the ADEM (Autonomous Digital Experience Management) team, you will be part of a team supporting the services that...PrincipalFull timeWork at officeVisa sponsorshipWork visa
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- senior chief engineer United States
- principal reliability engineer United States
- associate director engineering United States
- aerospace engineering director United States
- director of hardware engineering United States
- general engineer United States
- project engineer assistant project manager United States
- principal devops engineer United States
- chief design engineer United States
- principal infrastructure engineer United States

