Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$121.4k - $218.6k

Akamai Technologies

Job Description

Do you enjoy collaborating with teams to solve complex challenges?

Do you enjoy solving large scale distributed content delivery challenges?

Join our critical AI Hardware SRE Team!

The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings.

Partner with the best

In this role, you'll play a part in pioneering the reliability an elite, high-density hardware and software infrastructure spanning the globe. You'll collaborate with product teams from the earliest stages of development to ensure the reliability, scalability, and performance of our systems. You'll define key performance indicators and defend them when they are breached.

As a Senior Site Reliability Engineer, you will be responsible for:
  • Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning.
  • Integrating automated workflows across disconnected corporate ticketing systems to optimize time-to-mitigate metrics for hardware and network break-fix events.
  • Leveraging advanced AI utilities and LLM-assisted development paradigms where appropriate to accelerate technical execution, script authorship, and system analysis
  • Working on cutting-edge private cloud and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments.
  • Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments.
  • Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows.
  • Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities.
Do what you love

To be successful in this role you will:
  • Have 5 years of relevant experience and a Bachelor's degree in Computer Engineering, Computer Science or equivalent
  • Possess tooling and coding ability in languages like Python to construct scalable operational tools, API integrations, and automation frameworks.
  • Show hands-on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki.
  • Possess a working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.
  • Have experience acting as a key designer for new service rollouts, including establishing operational readiness criteria, telemetry baselines, and alerting thresholds.
  • Demonstrate extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post-mortems.
  • Display a proven ability to take absolute ownership of ambiguous technical problems, coordinate cross-functional teams, and drive for production-grade solutions.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.

Our focus is simple:

Cloud and Edge: Running apps closer to users for instant performance.

Security : Neutralizing threats before they ever reach your data.

Content Delivery : Scaling the world's biggest moments without a glitch.

AI : Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.

At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.

We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!


Compensation

Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $121,400 - $218,600/year; a candidate's salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Cambridge, MA vacancy
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Cambridge, MA
    2 days ago
  • $160k - $200k

     ...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    3 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    4 days ago
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Senior
    Work experience placement
    Work at office
    Remote work

    Axon

    Boston, MA
    15 hours ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that...  ...maintains our continuous delivery infrastructure, ensuring reliable code deployment from development through production for all engineering... 
    Senior
    Local area
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    4 days ago
  • $140k - $210.9k

     ...position will be primarily on-site with residency commutable to...  ...DevOps backgrounds or software engineering backgrounds (e.g., Java...  ...interest in operating and improving reliability of distributed production...  ...Responsibilities As a Senior Engineer of the SRE / Production... 
    Senior
    Full time
    Temporary work
    Part time
    Work at office
    Shift work

    Federal Reserve System

    Boston, MA
    3 days ago
  • $160k - $200k

     ...Senior Site Reliability Engineer This role is located in Somerville, MA – We are a hybrid work environment and are in the office 3+ days/per week. Tulip, the leader in AI-native frontline operations, is helping companies around the world equip their workforce with... 
    Senior
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Venturefizz Product Management Community

    Somerville, MA
    2 days ago
  •  ...commuting distance of one of our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team, you will operate the...  ...candidate is someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and automating cloud-based... 
    Senior
    Full time

    Federal Reserve Bank of Boston

    Boston, MA
    4 days ago
  • $165.75k - $224.45k

     ...are dedicated to solving complex problems and making a huge impact. Where You Fit We're looking for a skilled staff level Site Reliability Engineer focused on designing, building, and operating our hybrid cloud/on-prem environment. What You’ll Do Advancing the state... 
    Senior
    Hourly pay

    PathAI

    Boston, MA
    3 days ago
  • $134.25k - $214.8k

     ...real change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability... 
    Senior
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Boston, MA
    2 days ago
  • $132.23k - $176.31k

     ...shape the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem.... 
    Senior
    Temporary work
    Remote work

    Lumen Inc

    Boston, MA
    5 days ago
  •  ...Site Reliability Engineering (SRE) Team Lead The Site Reliability Engineering (SRE) team is foundational to the growth and scale of our platform...  ...execution of this roadmap, collaborating with SREs and senior engineers across the organization, while performing hands-... 
    Shift work

    Roberts Recruiting

    Boston, MA
    1 day ago
  • $135k - $165k

     ...tap into global manufacturing capacity.Xometry is seeking a Site Reliability Engineer II to join our Site Reliability Engineering (SRE)...  ...statements and drive them to completion with guidance from senior engineers.Write clean, efficient, and well-documented code... 
    Flexible hours

    Xometry

    Boston, MA
    2 days ago
  • $130k - $150k

     ...technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are...  ...career mentoring and performance coaching from an assigned senior colleague. Additional leadership and collaboration opportunities... 
    Work at office
    Work from home
    3 days per week

    CRA International

    Boston, MA
    1 day ago
  •  ...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the...  ...crucial workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background.... 
    Senior
    Full time
    Remote work
    Worldwide

    Mongodb

    Boston, MA
    a month ago
  • $160k - $200k

    Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware... 
    Local area
    Remote work

    QuEra Computing

    Boston, MA
    3 days ago
  •  ...mission-critical industries, helping partners move more quickly and reliably from algorithm to silicon. Our platform accelerates deployment...  .... The Roles We are looking for an experienced software engineer to help us build a new generation of transpilation tools... 
    Senior
    Full time
    Remote work
    Relocation package
    Flexible hours

    Code Metal

    Boston, MA
    15 hours ago
  •  ...Job Title: Site Reliability Engineer Location: Remote with Quarterly visits to Chennai, Tamil Nadu, India Duration: Full-Time bout BigRio: BigRio is a remote-based, technology consulting firm headquartered in Boston, MA. We deliver software solutions... 
    Full time
    Remote work

    Saviance

    Boston, MA
    3 days ago
  • $75.7k - $136.3k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Work experience placement
    Work at office

    Akamai

    Cambridge, MA
    5 days ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Boston, MA
    5 days ago
  •  ...Software Development Engineer We're creating a platform that will change the way organizations measure their software development efforts...  ...teams can work and the tools they use Location: on-site in Boston We believe that it takes a diverse team to build the... 
    Senior

    Roberts Recruiting

    Boston, MA
    1 day ago
  •  ...ISEE is seeking an experienced Senior Software Engineer to join our team. The ideal candidate has several years of work experience, and have worked on complex, performance-critical code-bases. Role responsibilities include: - Support full software development life... 
    Senior
    Full time
    Work experience placement

    Isee

    Cambridge, MA
    15 hours ago
  •  ...Distributed Systems engineers at Datadog design, implement and run in production the foundational platforms powering our applications. Your data pipelines will ingest, store, analyze and query in real-time billions of events per second from companies all over the globe... 
    Senior
    Full time
    Work at office

    Datadog

    Boston, MA
    15 hours ago
  • $160k - $225k

     ...tens of thousands of users across hundreds of organizations globally. About the Role Manifold is looking for a Staff Site Reliability Engineer (SRE) to work at the intersection of AI, data infrastructure, and life sciences. In this high-impact role, you will help... 

    Manifold AI

    Cambridge, MA
    3 days ago
  • $108k - $209k

     ...seeking an experienced, creative, and talented Principal / Senior Software Engineer. The ideal candidate will have a strong background in software...  .... Leverage AWS cloud infrastructure to build scalable, reliable, and efficient applications and AI-powered services. Uphold... 
    Senior

    Seres Therapeutics

    Cambridge, MA
    1 day ago
  • $150k - $215k

     ...individuals optimize their health, fitness, and recovery. As a Senior Software Engineer on the AI team, you will play a key role in building and...  ...is ideal for an engineer who is passionate about building reliable, scalable applications and thrives in a fast-paced,... 
    Senior
    Full time
    Work at office
    Relocation

    Whoop

    Boston, MA
    15 hours ago
  • $191k - $253k

     ...Senior Software Engineer, VMS Anduril’s Maritime Division has assembled a diverse team of experts in software, robotics, artificial intelligence, sensor fusion, and data analysis to create software and hardware solutions that radically evolve the capabilities of... 
    Senior
    Full time
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    Anduril

    Boston, MA
    15 hours ago
  •  ...Company Description Zenith Talent specializes in staffing professional positions in Information Technology, Engineering, Marketing, Sales, Finance, HR and Operations. We have the knowledge and skills to supply candidates that fit perfectly in your organization. As... 
    Senior
    Full time

    Zenith Talent

    Boston, MA
    15 hours ago
  • $191k - $253k

     ...we want you to join Anduril’s Maritime Division and help us build the future of defense capability. ABOUT THE JOB Senior Software Engineers independently drive the delivery of a variety of software integrated in to our products. This includes autonomy, simulation... 
    Senior
    Full time
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    Anduril

    Boston, MA
    15 hours ago
  •  ...mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocation(s): Waltham, MA | HybridAbout the RoleSr Site Reliability Engineer- Guardian of the products to ensuring systems are reliable, scalable, and efficient... 
    Senior
    Ongoing contract
    Full time
    Temporary work
    Work experience placement
    Worldwide

    SS&C Technologies

    Waltham, MA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!