Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Sporty Group

Site Reliability Engineer

EMEA

What You'll Be Doing
  • Work with a team of DevOps and DBA professionals
  • Improve existing infrastructure and processes across the countries we're deployed in, as well as streamlining processes to deploy to new countries in the future
  • Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices
  • Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
  • Own weekday on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
  • Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
  • Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
  • Take ownership and responsibility for our cloud operation activities
  • Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
  • Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
  • Mentoring less experienced team members
What You'll Bring
  • 3+ years DevOps / SRE / platform engineering experience
  • Must be based in Europe
  • Experience independently leading the planning and deployment of a project
  • Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
  • Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued
  • Experience with Infrastructure-as-Code, particularly Terraform
  • Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
  • Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry
  • Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus
  • Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
  • Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
  • Experience defining SLIs and SLOs and using them to inform reliability work
  • Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments
  • Solid networking knowledge, especially the TCP / IP stack and protocol
  • Experience handling high request volumes and designing systems for high availability and high traffic environments
  • A strong understanding of cache, including CDN, cache, Redis / Memcached
  • Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous
Our Stack
  • Languages: Java / Spring Boot, Node.js, Python, JavaScript
  • Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community
  • Cache: ElastiCache, Redis, Valkey
  • Messaging: Apache RocketMQ, AutoMQ, Kafka
  • Networking & Proxy: Nginx, Kong, Cilium, eBPF
  • Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm
  • Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3
  • CI/CD: Jenkins, GitHub Actions
  • Metrics: Prometheus, Mimir, Grafana, Alertmanager
  • Logs: Loki, Vector
  • Traces: Tempo, OpenTelemetry, Alloy
  • Profiling: Pyroscope
  • RUM: Grafana Faro, OpenTelemetry SDK
  • Infrastructure as Code: Terraform
  • CDN & Edge: Cloudflare, AWS CloudFront
  • AWS CloudWatch
What's In It For You
  • Sporty is a remote first company in pursuit of sustainability
  • A competitive salary + individual performance based bonuses every quarter
  • 28 days paid annual leave
  • Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
  • Referral bonuses & flash bonuses
  • Top of the line equipment
  • Annual company retreats to provide great internal networking opportunities
Interview Process
  • Remote video screening with our Talent Acquisition Team
  • Online assessment via Hackerrank
  • Remote video interview with 3 x Team Members (45 mins each, not separate days)

If you're interested, we encourage you to apply! Every application is reviewed by a member of our team (AI is not used in our recruitment process), and we aim to respond within 48 hours.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in United States vacancy
  • $96k - $163k

     ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Overview The BizOps team is looking for a Senior Site Reliability Engineer who can help us solve problems and... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours
    Shift work

    Mastercard

    O Fallon, MO
    16 hours ago
  • $96k - $163k

     ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    16 hours ago
  • $76k - $127k

     ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    16 hours ago
  • $76k - $127k

     ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II The BizOps team at Mastercard is looking for a Site Reliability Engineer (SRE) who thrives on solving complex... 
    Suggested
    Full time
    Part time
    Worldwide
    Flexible hours
    Early shift

    Mastercard

    O Fallon, MO
    16 hours ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure... 
    Suggested
    Flexible hours

    Circle

    San Francisco, CA
    16 hours ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    16 hours ago
  • $138.4k - $173k

     ...infrastructure as well as help improve the reliability, quality of services and overall...  ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability...  ...about our locations by visiting our site.Compensation & BenefitsThe base salary that... 
    Full time
    Flexible hours

    AppFolio

    San Diego, CA
    16 hours ago
  • $112k - $179k

     ...delivery of system, network, software, and security solutions.About The RolePeraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems... 
    Contract work
    Worldwide
    Shift work

    Peraton Corporation

    Washington DC
    16 hours ago
  • $128k - $216k

     ...another millions of times a day - quickly, reliably, and securely. Any time you swipe your...  ...make a difference at Fiserv.Job TitleSr. Site Reliability EngineerAbout CloverClover is...  ...does a successful Senior Site Reliability Engineer do at Fiserv?As a Senior Site Reliability... 
    Full time
    Worldwide

    Fiserv

    Sunnyvale, CA
    16 hours ago
  • The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams... 
    Full time
    Work at office
    Local area

    Castleton Commodities International

    Houston, TX
    16 hours ago
  • $130k - $140k

     ...mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionSite Reliability EngineerLocation(s): Waltham, MA | HybridAbout the RoleSr Site Reliability Engineer- Guardian of the products to ensuring systems are reliable, scalable, and efficient... 
    Ongoing contract
    Full time
    Temporary work
    Work experience placement

    SS&C Technologies

    Waltham, MA
    16 hours ago
  •  ...English (Required)Work Shift:1st Shift (United States of America)Please review the following job description:Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)We are seeking a Lead Site Reliability & Environment Monitoring Engineer to... 
    Full time
    Temporary work
    Shift work
    Day shift

    TIH

    Charlotte, NC
    16 hours ago
  •  ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 
    Temporary work

    Kong

    Washington DC
    16 hours ago
  •  ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises... 
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Charlotte, NC
    13 hours ago
  • $104.9k - $174.7k

    Are you passionate about improving reliability, scalability, and resilience in complex database...  ....Own prioritization of reliability engineering tasks within team backlogs.Lead incident...  ...a Service (IaaS).Background in DevOps, site reliability engineering practices, or related... 
    Full time
    Local area

    RELX Group

    Georgia
    2 days ago
  • $78k - $124.75k

     ...annually + bonus + benefitsJob Function: Engineering & ArchitectureSchedule: Full timeShift:...  ...TechnologyCompany: American ExpressDescriptionSite Reliability Engineer I enhances system resilience...  ...cloud platformsKnowledge of cloud‑based Site Reliability Engineering (SRE) practices... 
    Remote work

    American Express

    Sunrise, FL
    1 day ago
  • $152.6k - $191.5k

     ...is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include composing...  ...improvement.Position Summary:The Senior Azure Site Reliability Engineer acts as an advanced senior... 
    Full time
    Work at office
    Day shift

    Bank of America

    Plano, TX
    13 hours ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank, Healthcare Payments team, you will solve complex and broad... 

    JP Morgan Chase

    Irvine, CA
    3 days ago
  •  ...communities.This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job family responsible...  ...across the business.Job SummaryWe are looking for a Site Reliability Engineer with a minimum of 5 years of industry... 
    Flexible hours
    Weekend work

    Morgan Stanley

    Alpharetta, GA
    4 days ago
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Full time

    Vanguard

    Dallas, TX
    3 days ago
  • $98.58k - $138.02k

     ...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company...  ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,... 
    Full time
    Work at office

    Restaurant 365

    Irvine, CA
    4 days ago
  • $243.29k - $295.25k

     ...challenges at scale, and helping to create safer, more civil shared experiences for everyone.The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    3 days ago
  • $69.8k - $148.3k

     ...of infrastructure and service to ensure reliability and functionality. Responds to infrastructure...  ...tools and develops working knowledge of site reliability trends.Only Oracle brings...  ...Science, Information Technology, Engineering, or a related field, or equivalent practical... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Nashville, TN
    2 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $102.1k - $202.2k

     ...per yearEmployment type: Full-TimeWork site: Fully on-siteRole type: Individual ContributorTravel...  ...: Software EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...at the intersection of large-scale cloud engineering, service reliability, and operational excellence... 
    Ongoing contract
    Local area
    Worldwide

    Microsoft

    Reston, VA
    4 days ago
  • $145.7k - $218.5k

     ...synonymous with entertainment excellence and creativity.Service Reliability EngineerDo you want to use transformative technologies to...  ...scalability and efficiency? Do you want a career that combines your engineering skills and your passion for video gaming? Are you fascinated... 
    Work experience placement
    Shift work

    Sony Interactive Entertainment America

    Aliso Viejo, CA
    1 day ago
  • $117k - $209.33k

    Job Requisition ID #26WD99276Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Full time
    For contractors
    Remote work

    Autodesk

    Plano, TX
    1 day ago
  • Job ID: 18719802Reference Number: 23-00164Title: site reliability engineerLocation: Iselin, NJ, 08830Posted Date: 2023-01-20Contact: Shyam MaramContact Email: ****@*****.*** Phone: (***) ***-****Company: HAN Staffing Devops/SRE/Python Role Malvern PA - hybrid... 

    HAN Staffing

    Iselin, NJ
    1 day ago
  • $134k - $170k

    The Senior Site Reliability Engineer (SRE) is a hands-on engineering role responsible for improving the reliability, observability, performance, and operational efficiency of ISO New England's IT services. The SRE works across infrastructure, platform, cyber security, and... 
    Permanent employment
    H1b
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours
    3 days per week

    ISO New England

    Holyoke, MA
    1 day ago
  • $86.8k - $198k

    Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development—if you have... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    McLean, VA
    13 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!