Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$121.4k
Full-time

Akamai

Do you enjoy collaborating with teams to solve complex challenges? Do you enjoy solving large scale distributed content delivery challenges? Join our critical AI Hardware SRE Team! The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner with the best In this role, you'll play a part in pioneering the reliability an elite, high-density hardware and software infrastructure spanning the globe. You'll collaborate with product teams from the earliest stages of development to ensure the reliability, scalability, and performance of our systems. You'll define key performance indicators and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: * Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning. * Integrating automated workflows across disconnected corporate ticketing systems to optimize time-to-mitigate metrics for hardware and network break-fix events. * Leveraging advanced AI utilities and LLM-assisted development paradigms where appropriate to accelerate technical execution, script authorship, and system analysis * Working on cutting-edge private cloud and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments. * Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments. * Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows. * Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities. Do what you love To be successful in this role you will: * Have 5 years of relevant experience and a Bachelor's degree in Computer Engineering, Computer Science or equivalent * Possess tooling and coding ability in languages like Python to construct scalable operational tools, API integrations, and automation frameworks. * Show hands-on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki. * Possess a working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks. * Have experience acting as a key designer for new service rollouts, including establishing operational readiness criteria, telemetry baselines, and alerting thresholds. * Demonstrate extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post-mortems. * Display a proven ability to take absolute ownership of ambiguous technical problems, coordinate cross-functional teams, and drive for production-grade solutions. About us At Akamai, we make life better for billions of people, trillions of times a day. Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone. Our focus is simple: Cloud and Edge: Running apps closer to users for instant performance. Security: Neutralizing threats before they ever reach your data. Content Delivery: Scaling the world's biggest moments without a glitch. AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform. At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works. Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits. [ FlexBase adapts to your job's needs Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work. We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both. Connect with us on social and see what life at Akamai is like!

    Compensation Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $121,400 - $218,600/year; a candidate’s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.

    Vacancy posted 1 day ago
    Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in United States vacancy
    • $96k - $163k

       ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer, Performance Engineering Senior Site Reliability Engineer, Performance Engineering Payment Optimization unifies... 
      Senior
      Full time
      Part time
      Worldwide
      Flexible hours

      Mastercard

      O Fallon, MO
      1 day ago
    • $96k - $163k

       ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
      Senior
      Full time
      Part time
      Worldwide
      Flexible hours

      Mastercard

      O Fallon, MO
      1 day ago
    • $163k

       ...and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview-The ProCOM team is looking for a Site Reliability Engineering (SRE) who can help us solve problems, build... 
      Senior
      Full time
      Part time
      Immediate start
      Worldwide
      Flexible hours

      Mastercard

      O Fallon, MO
      1 day ago
    • $120k - $175k

       ...level of sports fandom. Ready to reimagine the DFS industry together? We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting-edge solutions and pushing the boundaries of what's... 
      Senior
      Full time
      Remote work
      Work visa
      Flexible hours

      PrizePicks

      United States
      3 hours ago
    • $189k - $283.6k

       ...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics...  ...of accountability * A strong desire to perform and grow as an engineer * 5+ years of software development experience Technologies... 
      Senior
      Full time
      Local area
      Remote work
      Relocation package
      Flexible hours
      Shift work

      Block

      California
      4 days ago
    • $166k - $220k

       ...networking technology to the military in months, not years. ABOUT THE TEAM We are seeking a highly skilled and mission-driven Site Reliability Engineer (SRE) to join our Mission Autonomy team. In this critical role, you will be responsible for ensuring the reliability,... 
      Senior
      Full time
      Work experience placement
      Immediate start
      Remote work

      Anduril Industries

      Costa Mesa, CA
      3 days ago
    • $118.6k - $195.68k

      About the Job The Red Hat IT OpenShift team is looking for a Senior Site Reliability Engineer (SRE) to design, develop, scale, and operate our Red Hat Hybrid OpenShift Platforms (on-prem & cloud). As a Senior Engineer, you will contribute to running Red Hat OpenShift... 
      Senior
      Permanent employment
      Full time
      Contract work
      Work experience placement
      Work at office
      Remote work
      Flexible hours

      Red River

      North Carolina
      3 hours ago
    • Role Description Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure...  ...for billions of people, billions of times a day. As a Senior Site Reliability Engineer, you will be: ~Designing, developing... 
      Senior
      Full time
      Work at office

      Akamai

      Remote
      2 days ago
    • $110k - $137.49k

      Role Description The Sr Site Reliability Engineer, Release will prototype, write, maintain, and test code in multiple stages of the release process...  ...multiple environments, including production ~Work with senior team members to understand stakeholder requirements and... 
      Senior
      Full time
      Work experience placement
      Remote work

      Alkami Technology

      Remote
      7 days ago
    • Role Description Stack AV Site Reliability Engineers are responsible for enabling and ensuring our production systems meet their service-level objectives. Through the implementation of centralized observability and automation, the SRE team constantly ensures the health... 
      Senior
      Full time

      Stack AV

      Remote
      3 days ago
    • Role Description We are expanding our Site Reliability Engineering (SRE) team and seeking a highly skilled and passionate Senior SRE to join us. As a member of our growing SRE function, you will play a critical role in ensuring the reliability, scalability, and performance... 
      Senior
      Full time
      Temporary work

      QAD, Inc.

      Remote
      7 days ago
    • $160k - $180k

       ...Socure is seeking a Site Reliability Engineer in New York to enhance our identity trust infrastructure. In this role, you will take full ownership of AWS and Kubernetes platforms, ensuring high reliability and operability. The ideal candidate will possess extensive experience... 
      Senior

      Socure Inc

      New York, NY
      1 day ago
    •  ...Karsun Solutions, LLC is seeking a Site Reliability Manager to lead a multi-disciplinary team responsible for reliability, security, and platform lifecycle across AWS-based services. The role emphasizes collaboration, observability, and continuous improvement in a client... 
      Senior

      Karsun Solutions

      New York, NY
      1 day ago
    • $152.5k - $205k

      Role Description As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical digital-assets, AI, and application workloads. You will bring an engineering mindset to production... 
      Senior
      Full time

      Circle

      Remote
      15 hours ago
    • Role Description We’re looking for a Senior Site Reliability Engineer who takes ownership seriously — someone who designs for reliability, ships the automation, and stands behind it in production. You’ll work across cloud-native infrastructure on systems that process millions... 
      Senior
      Full time

      CertifyOS

      Remote
      7 days ago
    •  ...A leading technology firm is in search of a Senior Wireless Network Site Reliability Engineer to manage and enhance their wireless network infrastructure. The ideal candidate has over 8 years of experience in wireless network operations and a strong background in wireless... 
      Senior

      TechDigital Group

      Santa Clara, CA
      1 day ago
    •  ...The Depository Trust & Clearing Corporation (DTCC) is seeking a Senior Application Support Engineer (SRE) to enhance reliability, scalability, and performance of mission-critical applications. You will apply SRE principles across engineering, infrastructure, and operations... 
      Senior

      The Depository Trust & Clearing Corporation

      Dallas, TX
      1 day ago
    •  ...Istio) ~Defining and monitoring Service-Level Objectives (SLOs) and Service-Level Agreements (SLAs) to ensure that systems meet reliability and performance targets ~Monitoring Tools like New Relic, Prometheus, Grafana, and/or Datadog ~OpenTelemetry knowledge for... 
      Senior
      Full time
      Remote work

      Shippo

      Remote
      7 days ago
    • Role Description Join us as a Senior Site Reliability Engineer on our mission to turn payments into possibilities! The Site Reliability Engineering (SRE) team ensures the reliability, availability, scalability, and performance of a mission-critical payment orchestration... 
      Senior
      Full time
      Immediate start
      Remote work

      CellPoint Digital

      Remote
      1 day ago
    • $54k - $150k

      Role Description As Senior Site Reliability Engineer for Remote Build, you'll own the operational excellence and infrastructure strategy that makes Build's platform reliable, performant, and safe for customers. You'll report to the Engineering Manager and work closely... 
      Senior
      Full time
      Local area
      Remote work
      Home office
      Flexible hours

      Remote

      Remote
      1 day ago
    •  ...PVH (Tommy Hilfiger/Calvin Klein) seeks a Senior Software Engineer to own the reliability and performance of our Kubernetes-based data platform across multi-region deployments. You will design scalable infrastructure, optimize deployment pipelines, and strengthen security... 
      Senior

      PVH (Tommy Hilfiger/Calvin Klein)

      Livingston, NJ
      1 day ago
    •  ...optimize production infrastructure across CI/CD, cloud deployments, and security. You will collaborate with our internal product and engineering teams to keep services scalable, secure, and highly available. The role emphasizes GitHub Actions, Terraform, Vercel, AWS core... 
      Senior
      Remote work

      Remote Leverage

      New York, NY
      1 day ago
    • $145k - $165k

       ...A technology solutions firm in Sunnyvale, CA is looking for a highly experienced Site Reliability Engineer (SRE). This role involves maintaining uptime and performance across systems. Exceptional Linux expertise and automation skills in Bash and Python are crucial. Key... 
      Senior

      Bolt Graphics, Inc.

      Sunnyvale, CA
      1 day ago
    •  ...Karsun Solutions in the DMV area is seeking a Site Reliability Manager to ensure reliability, scalability, and performance of our systems. You will lead a team focusing on Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate has 1... 
      Senior

      Karsun Solutions

      New York, NY
      1 day ago
    •  ...A global leader in fast food is seeking a Senior Manager for Edge Operations/SRE in Chicago. This pivotal role involves leading edge...  ..., collaborating across teams to ensure high availability and reliability of the platform. Candidates should have 10+ years in infrastructure... 
      Senior

      McDonald's Corporation

      Chicago, IL
      1 day ago
    • Role Description The Senior Site Reliability Engineer is a technical leader responsible for architecting the reliability strategy for large-scale, distributed government systems. You will lead the implementation of the SRE framework, driving the adoption of SLO-based management... 
      Senior
      Contract work
      Remote work

      Arctiq

      Remote
      6 days ago
    • Role Description Versant's Sports & Entertainment Digital Products division is seeking a Senior Site Reliability Engineer to help drive the reliability, scalability, and usability of internal developer platforms, tooling, and engineering workflows across a portfolio of... 
      Senior
      Full time
      Local area
      Remote work
      Worldwide

      Versant

      Remote
      2 days ago
    •  ...The Home Depot is seeking a Senior Software Reliability Engineer to join the Platform Reliability Engineering team, ensuring the resilience, performance, and security of our enterprise Cloud Platform. You will mentor junior engineers, lead incident triage, root cause... 
      Senior

      Home Depot

      Atlanta, GA
      1 day ago
    • $135k - $145k

       ...A medical equipment manufacturing company is hiring a Senior Site Reliability Engineer in Carlsbad, CA. The role involves ensuring system reliability, automating operational tasks, and managing incident response. Candidates should have extensive experience in SRE, strong... 
      Senior

      ATEC Spine

      Carlsbad, CA
      1 day ago
    • $54k - $150k

      Role Description As Senior Site Reliability Engineer for Remote Build, you'll own the operational excellence and infrastructure strategy that makes Build's platform reliable, performant, and safe for customers. You'll report to the Engineering Manager and work closely... 
      Senior
      Full time
      Local area
      Immediate start
      Remote work
      Home office
      Flexible hours

      Referral Board

      Remote
      10 hours ago

    Do you want to receive more vacancies?

    Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!