Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

$95k - $171k

Akamai Technologies

Job Description

Are you passionate about cutting‑edge AI infrastructure? Do you want to build your SRE career on one of the most exciting platforms in cloud computing? Join the Akamai Inference Cloud Team

The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications.

Partner with the best

In this role, responsibilities will include automation, monitoring, incident response, and working collaboratively with skilled team members. Candidates should possess expertise in Linux systems, automation, and SRE practices. Daily activities involve coding, improving dashboards, enhancing alerts, and minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform.

As an Site Reliability Engineer II, you will be responsible for:
  • Building and maintaining dashboards, alerts, and monitoring for inference workloads using Akamai's existing observability platform
  • Writing automation and tooling in Python or Go to reduce operational toil and improve system reliability
  • Building and improving runbooks for inference‑specific operational procedures, integrating into Akamai's existing incident management processes
  • Contributing to SLO tracking and reporting, identifying trends and areas for improvement
  • Supporting CI/CD pipeline maintenance, deployment safety checks, and rollback procedures
  • Collaborating with product engineering teams to troubleshoot complex problems across the stack
  • Participating in on‑call rotations, responding to production incidents, and conducting blameless post‑mortems
To be successful in this role you will:
  • Have 2+ years of experience in Site Reliability Engineering and a Bachelor's Degree or its equivalent experience
  • Demonstrate coding ability in at least one programming language (Python or Go) with experience writing automation
  • Have experience with Linux systems administration and the ability to troubleshoot complex infrastructure issues
  • Show familiarity with Kubernetes and containerization concepts
  • Have experience with monitoring and observability tools such as Prometheus, Grafana, or similar
  • Have exposure to CI/CD pipelines and infrastructure‑as‑code tools (Terraform, SaltStack, or equivalent)
  • Show a willingness to learn and grow, with genuine curiosity about AI infrastructure and distributed systems
FlexBase: Work in a way that works for you

FlexBase, Akamai's Global Flexible Working Program, is based on the principles that are helping us create the best workplace in the world. When our colleagues said that flexible working was important to them, we listened. We also know flexible working is important to many of the incredible people considering joining Akamai. FlexBase gives 95% of employees the choice to work from their home, their office, or both (in the country advertised). This permanent workplace flexibility program is consistent and fair globally, to help us find incredible talent, virtually anywhere. We are happy to discuss working options for this role and encourage you to speak with your recruiter in more detail when you apply.

Benefits
  • Your health
  • Your finances
  • Your family
  • Your time at work
  • Your time pursuing other endeavors
About Us

Akamai powers and protects life online. Leading companies worldwide choose Akamai to build, deliver, and secure their digital experiences helping billions of people live, work, and play every day. With the world's most distributed compute platform from cloud to edge we make it easy for customers to develop and run applications, while we keep experiences closer to users and threats farther away.

Join us

Are you seeking an opportunity to make a real difference in a company with a global reach and exciting services and clients? Come join us and grow with a team of people who will energize and inspire you!

Compensation

Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $95,000 - $171,000/year; a candidate’s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry‑leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.

#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Cambridge, MA vacancy
  • $134.25k - $214.8k

     ...matters at a company where you matter.Your ImpactAre you an engineer who gets excited about the challenge of making complex distributed...  ...it.You will be part of the Observability team within Axon's Site Reliability organization — a focused team responsible for Axon's metrics,... 
    Suggested
    Work experience placement
    Work at office
    Remote work

    Axon

    Boston, MA
    5 days ago
  • $160k - $200k

     ...data, ideally using promQLKey Responsibilities:Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. Contributing to and maintaining Tulip's triage & remediation processes as a player / coachPerform incident... 
    Suggested
    Temporary work
    Work at office
    Local area
    Flexible hours
    3 days per week

    Tulip Interface

    Somerville, MA
    3 days ago
  • $134.25k - $214.8k

     ...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance... 
    Suggested
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Axon

    Boston, MA
    2 days ago
  • $135k - $165k

     ...while also making it easy for buyers at Fortune 1000 companies to tap into global manufacturing capacity.Xometry is seeking a Site Reliability Engineer II to join our Site Reliability Engineering (SRE) Organization. In this role as an individual contributor, you will guide... 
    Suggested
    Flexible hours

    Xometry

    Boston, MA
    2 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Suggested

    Google

    Cambridge, MA
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    4 days ago
  • $130k - $150k

     ...systems and hybrid infrastructure, meaning experience with cloud technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable, scalable, and performant across on-premises and cloud... 
    Work at office
    Work from home
    3 days per week

    CRA International

    Boston, MA
    6 days ago
  •  ...The Depository Trust & Clearing Corporation (DTCC) seeks a Senior Application Support Engineer to ensure reliability and performance of its critical trade processing platforms. You will apply SRE principles, drive automation, and partner with global teams to support AWS... 

    The Depository Trust & Clearing Corporation

    Boston, MA
    3 days ago
  •  ...commuting distance of one of our 12 Reserve Bank locations As a Senior Engineer of the SRE / Production Operations team, you will operate the...  ...ideal candidate is someone who loves building and maintaining reliable and scalable systems, CI/CD tooling, and automating cloud-based... 
    Full time

    Federal Reserve Bank of Boston

    Boston, MA
    5 days ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Cambridge, MA
    1 day ago
  • $166.3k - $238.3k

     ...and more intuitive with technology that simply works. The SRE Engineering Enablement org for our Network Platform team supports our CI...  ...customers are all engineers at Cisco.YOUR IMPACT As a Lead Site Reliability Engineer, you will architect, build, and evolve the... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Boston, MA
    4 days ago
  • $166k - $220k

     ...failure. As such, it is critical that Anduril services are reliable and maintainable. This means that all services &...  ...ground systems & Kubernetes infrastructure.ABOUT THE JOBAs a Site Reliability Engineer on the Observability team, you will build & operate Anduril... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Boston, MA
    4 days ago
  • $127k - $249k

     ...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas...  ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    5 days ago
  • $160k - $200k

    Role Overview We are looking for a Control System Engineer/Site Reliability Engineer (SRE) to integrate and maintain the hardware and software systems that enable QuEra’s quantum controls and software stack. You’ll work closely with software engineers, physicists, hardware... 
    Local area
    Remote work

    QuEra Computing

    Boston, MA
    3 days ago
  • Google Cambridge is hiring a Senior Software Engineer for Site Reliability Engineering. You will join a team responsible for reliability, latency, and capacity of Google's public services, spanning Search, Ads, Gmail, YouTube, and more. The role blends software and operations... 

    Google Inc.

    Cambridge, MA
    10 hours ago
  • $174k - $252k

    Google in Cambridge, MA is seeking a Site Reliability Engineer to design, build, and operate highly available services across a broad portfolio. The role demands strong background in software development, large-scale distributed systems, and leadership to guide projects... 

    Google

    Cambridge, MA
    4 days ago
  • $174k - $252k

    Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5 years of experience with...  ...s degree in Computer Science or Engineering. About The Job Site Reliability Engineering (SRE) is what you get when you treat operations... 

    Google

    Cambridge, MA
    4 days ago
  • $130k - $140k

     ...mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocation(s): Waltham, MA | HybridAbout the RoleSr Site Reliability Engineer- Guardian of the products to ensuring systems are reliable, scalable, and efficient... 
    Ongoing contract
    Full time
    Temporary work
    Work experience placement

    SS&C Technologies

    Waltham, MA
    2 days ago
  • $151k - $297k

     ...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB's cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Boston, MA
    1 day ago
  • $132.23k - $176.31k

     ...future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role... 
    Full time
    Temporary work
    Remote work

    Lumen

    Cambridge, MA
    3 days ago
  • $40 - $45.78 per hour

     ...Job Description Job Description Site Reliability Engineer 1 Job Details Site Reliability Engineer 1 (Contract) Location: Waltham, MA 02451 (Hybrid) Duration: 10/22/2025 to 4/03/2026 Team: Campaign Core RD US Key Responsibilities: Deploy and manage... 
    Hourly pay
    Contract work

    Cypress HCM

    Waltham, MA
    8 days ago
  •  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,...  ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    Boston, MA
    11 days ago
  •  ...Distributed Systems engineers at Datadog design, implement and run in production the foundational platforms powering our applications. Your data pipelines will ingest, store, analyze and query in real-time billions of events per second from companies all over the globe... 
    Full time
    Work at office

    Datadog

    Boston, MA
    10 hours ago
  • $191k - $253k

     ...you to join Anduril’s Maritime Division and help us build the future of defense capability. About the Job Senior Software Engineers independently drive the delivery of a variety of embedded and/or safety critical software integrated in to our products. This includes... 
    Full time
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    Anduril Industries

    Boston, MA
    10 hours ago
  • $125k - $175k

     ...of their bodies and daily lives. WHOOP is hiring a Software Engineer II to join the Business Systems team. In this role, you will design...  ...teams across WHOOP to ensure that our systems are scalable, reliable, and efficient, directly impacting how we serve members and... 
    Full time
    Work at office
    Relocation

    Whoop

    Boston, MA
    3 hours ago
  • $191k - $253k

     ...you to join Anduril’s Maritime Division and help us build the future of defense capability. ABOUT THE JOB Senior Software Engineers independently drive the delivery of a variety of software integrated in to our products. This includes autonomy, simulation, data... 
    Full time
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    Anduril

    Boston, MA
    10 hours ago
  • $100k - $120k

     ...team as we help shape a brighter way forward. JLL is seeking a Reliability Engineer to join our team!   ​In JLL Work Dynamics our most...  ...replacement decisions. Support all efforts in the execution of site level Asset Management & Reliability programs and processes to... 
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours

    Jones Lang LaSalle

    Boston, MA
    5 days ago
  • $52.7 - $62 per hour

     ...TitleReliability EngineerJob Description SummaryThis individual will provide support for Facility Operations/Engineering department sites as part of the Reliability Engineering team. The ideal candidate will be responsible for ensuring the reliability and performance of... 
    Minimum wage
    Full time
    Flexible hours

    Cushman & Wakefield

    Boston, MA
    2 days ago
  • $135k - $325k

    Job OverviewWe are seeking an experienced Engineer to join our Trading Systems team within the Front Office Systems group. This role...  ...engineering best practices.Identify opportunities to improve system reliability, performance, scalability, automation, and developer... 
    Full time
    Local area

    Arrowstreet Capital

    Boston, MA
    2 days ago
  • $150k - $195k

     ...members to perform at a higher level through a deeper understanding of their bodies and daily lives.WHOOP is seeking a Senior Reliability Engineer to lead the charge in ensuring our hardware products deliver a consistent, high-reliability experience for members. In this... 
    Full time
    Work at office
    Relocation

    WHOOP

    Boston, MA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!