Senior Site Reliability Engineer
$121.4k - $218.6kAkamai Technologies
Do you enjoy collaborating with teams to solve complex challenges?
Do you enjoy solving large scale distributed content delivery challenges?
Join our critical AI Hardware SRE Team!
The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings.
Partner with the best
In this role, you'll play a part in pioneering the reliability an elite, high-density hardware and software infrastructure spanning the globe. You'll collaborate with product teams from the earliest stages of development to ensure the reliability, scalability, and performance of our systems. You'll define key performance indicators and defend them when they are breached.
As a Senior Site Reliability Engineer, you will be responsible for:
Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning.
Integrating automated workflows across disconnected corporate ticketing systems to optimize time-to-mitigate metrics for hardware and network break-fix events.
Leveraging advanced AI utilities and LLM-assisted development paradigms where appropriate to accelerate technical execution, script authorship, and system analysis
Working on cutting-edge private cloud and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments.
Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments.
Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows.
Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities.
Do what you love
To be successful in this role you will:
Have 5 years of relevant experience and a Bachelor's degree in Computer Engineering, Computer Science or equivalent
Possess tooling and coding ability in languages like Python to construct scalable operational tools, API integrations, and automation frameworks.
Show hands-on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki.
Possess a working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.
Have experience acting as a key designer for new service rollouts, including establishing operational readiness criteria, telemetry baselines, and alerting thresholds.
Demonstrate extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post-mortems.
Display a proven ability to take absolute ownership of ambiguous technical problems, coordinate cross-functional teams, and drive for production-grade solutions.
About us
At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security : Neutralizing threats before they ever reach your data.
Content Delivery : Scaling the world's biggest moments without a glitch.
AI : Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.
Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits. (
FlexBase adapts to your job's needs
Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.
Connect with us on social and see what life at Akamai is like!
Compensation
Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $121,400 - $218,600/year; a candidate’s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.
Equal Employment Opportunity Rights
Akamai Technologies is an Affirmative Action, Equal Opportunity Employer that values the strength that diversity brings to the workplace. All qualified applicants will receive consideration for employment and will not be discriminated against on the basis of gender, gender identity, sexual orientation, race/ethnicity, protected veteran status, disability, or other protected group status.
- Job Description:About the Role: We are looking for a Senior SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering...SeniorFull time
- ...the team takes that seriously!The RoleThe Senior SRE at 2K is a hands-on technical leader... ...regions while partnering with network engineers, systems architects, and game studio developers... ...technical direction, influencing reliability from architecture review through production...Senior
- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SeniorTemporary workCasual workWorldwide
- ...the rapidly evolving state of the art to engineer scalable, innovative, and research... ...business. About the Role:We are looking for a Senior SRE to serve as the operations owner for... ...tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including...SeniorFull timeLocal area
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$152k - $241.5k
...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and... ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data-... ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through...SeniorFull time- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical...SeniorFull timeWork at office
$127k - $249k
...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background....SeniorLocal areaRemote workWorldwideFlexible hours$210.6k - $305.1k
...Minimum Qualifications: You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...SeniorFull timeTemporary workLocal areaFlexible hours- ...Schwab. We are an integrated product, engineering, strategy and risk team, all based in San... ...how we serve our clients. As a Senior Engineer on AI.x, you will play a key role... ...areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts...SeniorFull time
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...SeniorLocal areaRemote workWorldwideFlexible hours$152k - $195k
...Senior Site Reliability Engineer Austin, TX (Hybrid) SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Alex Yampolskiy and...Senior$168k - $200k
...that is passionate about creating transformative change in healthcare. What We’re Looking For We’re looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You’ll be at the forefront of building and operating a resilient, observable,...Senior- ...commercialization, and mass production to change the world for the better. JOB SUMMARY We are seeking an experienced Site Reliability Engineer to own and maintain the deployment of our cloud-based infrastructure to customer sites. In this role, you will work...SeniorFull timeLocal area
$110.7k - $171.8k
...components Participation in on-call rotation as a platform reliability escalation point Incident response, post-incident reviews,... ..., and internal control requirements. Collaborate with engineering teams across the organization to influence platform adoption,...SeniorWork experience placementWork at officeLocal area- ...Senior Site Reliability Engineer At Schwab, you're empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us challenge the status quo and transform the finance industry together. We believe in the importance...SeniorFull timeWork at office
- ...the layer where data becomes decisions, and decisions make the advantage. About the Role Gallatin is looking for a Site Reliability Engineer to keep our production systems running with the reliability our national security customers require. You'll work at the...SeniorFull timeLocal area
$185k - $227k
...professionals. If the opportunity to build your career is compelling, read on for more details. ROLE AND RESPONSIBILITIES: A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance ofJuul’s hybrid cloud infrastructure (Nutanix, AWS/...SeniorRemote work- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization...SeniorWork at officeLocal area
- ...in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously improve Schwab's enterprise streaming platform ecosystem...SeniorFull timeWork at office
- ...Job Description Job Description Sr. Software Engineer - Site Reliability About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce... ...logistics. Position Overview: We’re seeking a Senior Site Reliability Engineer to join our fast-paced Engineering...SeniorFull timeWork at office
$140k - $215k
...intersection of our Core Platform and Embedded Reliability charters: building the foundational... ...while embedding directly with product engineering teams and their leadership to drive... ...eliminated manual deployment processes.At the Senior Engineer level, your influence is...SeniorFull timeWork experience placementWork at officeLocal area2 days per week3 days per week- Selby Jennings is seeking a Senior Site Reliability Engineer to scale and support critical workflow orchestration and automation platforms across the organization. The role sits in Platform Engineering, delivering highly available and resilient infrastructure for business...Senior
$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...Full timeWork at office$109.65k - $182.76k
...encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication...Full timeLocal area3 days per week$196k - $269.5k
Senior Principal AI Agent EngineerThe Software Engineering team delivers next-generation software application enhancements and new products for a changing world. Working at the cutting edge, we design and develop software for platforms, peripherals, applications and diagnostics...Senior- ...Principal Site Reliability Engineer About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we power shipping logic and checkout optimization...Full timeWork at office
$152k - $241.5k
...the world.Join the Simulation Software team at NVIDIA as a Senior System Software Engineer! This role offers an outstanding opportunity to work on... ...development by enabling Chips Simulation as a trusted and reliable virtual platform.What you will be doing:Drive early...SeniorFull time$167.18k - $203.61k
...remotely part of the weekTravel %NoWork ShiftJob DescriptionCox Automotive Corporate Services, LLCLEAD SITE RELIABILITY ENGINEERJob Description: Lead Site Reliability Engineer positions offered by Cox Automotive Corporate Services, LLC (Austin, Texas). Lead the...Full timeWork at officeRemote workFlexible hours- ...You will provide cloud operations for Oracle National Security Realms. Responsibilities Escalation points for junior site reliability engineers during complex or high-impact incidents. Manage and execute complex manual Change Management tickets, by working...Temporary workWork experience placementFlexible hoursNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer Austin, TX
- site reliability engineer remote Austin, TX
- site reliability engineer sre Austin, TX
- senior maintenance supervisor Austin, TX
- senior lead project manager Austin, TX
- senior robotics software engineer Austin, TX
- senior firewall engineer Austin, TX
- senior devops engineer remote Austin, TX
- senior sas administrator Austin, TX
- senior IT manager Austin, TX


