Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Site Reliability Engineer

FreedomPay

Job Description

Job Description

The FreedomPay Commerce Platform is the technology of choice for many of the largest companies across the globe in retail, hospitality, lodging, gaming, sports and entertainment, foodservice, education, healthcare and financial services.  FreedomPay’s technology has been purposely built to deliver rock solid performance in the highly complex environment of global commerce. The company maintains a world-class security environment and was first to earn the coveted validation by the PCI Security Standards Council against Point-to-Point Encryption with EMV standard in North America. FreedomPay’s robust solutions across payments, security, identity and data analytics are available in-store, online and on-mobile and are supported by rapid API adoption. The award winning FreedomPay Commerce Platform operates on a single, unified technology stack across multiple continents allowing enterprises to deliver a consistent, repeatable experience on a global scale.  FreedomPay is a fast paced, high growth company with a great culture with competitive benefits and compensation with a business casual atmosphere.

FreedomPay is seeking an experienced Senior Site Reliability Engineer to help ensure the highest possible availability and resiliency of a rapidly growing global payment platform. This full-time salaried position builds on a strong foundation of observability, incident response, and support experience across the development lifecycle — and pushes it forward with AI-driven operations and automation at its core. The right candidate finds real satisfaction in eliminating manual toil, treats every recurring task as an automation opportunity, and is eager to apply modern AI tooling to detect, diagnose, and resolve issues faster than ever before. 

About the Role

You’ll join a team of SREs who work closely with other teams of world-class engineers to tenaciously and creatively solve problems and reduce manual toil wherever possible. We expect AI and automation to be a force multiplier in everything you do — from accelerating root-cause analysis and enriching alerts, to generating runbooks and codifying remediation so that the platform increasingly heals itself. 

Successful candidates are heavily results-driven, bring well-established expertise across both traditional and bleeding-edge technology, and have a strong desire to continuously grow and improve themselves and our platform. This is a global operation spanning multiple regions and time zones, and the role demands the flexibility and commitment that a 24/7 payment platform requires. 

  • This position participates in an engineering on-call rotation and provides after-hours support for production issue escalations on a rotational basis. 

  • This position is based in the Philadelphia area with a hybrid schedule. Remote arrangements may be considered for exceptional candidates, with occasional travel to Philadelphia required.

Primary Responsibilities:

  • Build and maintain a comprehensive understanding of the platform and custom application stack.
  • Implement, maintain, and continuously improve observability strategies and metrics that ensure complete system health for numerous complex products throughout all stages of the development lifecycle, up to and including production.
  • Continuously identify automation opportunities and follow through to successful implementation, applying AI-assisted tooling to accelerate development and reduce manual effort.
  • Design, build, and maintain automated remediation and self-healing workflows that detect, triage, and resolve common failure modes with minimal human intervention.
  • Leverage AI/ML-driven observability — anomaly detection, alert correlation, and intelligent noise reduction — to surface issues earlier and shorten time to detection.
  • Use AI-assisted analysis to accelerate root-cause investigation, enrich incident context, and generate first-draft postmortems and runbooks for human review.
  • Handle escalations and collaborate effectively with other team members to quickly determine the root cause of any type of service degradation.
  • Implement, maintain, and continuously improve incident response procedures and other operational documentation, automating documentation generation and upkeep wherever practical.
  • Assist with troubleshooting and remediation of failed scheduled jobs and data-related concerns.
  • Champion responsible, secure adoption of AI tooling across the SRE function — sharing patterns, prompts, and automations that raise the productivity of the whole team
AI Enablement & Automation

AI and automation are central to how this team operates. We are looking for someone who will not only use these tools but help define how the SRE function applies them. In this role you will: 

  • Apply AI-assisted development and operations tools — including Anthropic (Claude), OpenAI (Codex), and Azure AI services (Foundry, Azure SRE Agent) and the agentic workflows built on them — to write, review, and accelerate automation and infrastructure code. 
  • Build and integrate automation that turns repetitive operational work into codified, repeatable, and self-service workflows. 

  • Use AIOps and ML-driven observability capabilities within the APM stack for anomaly detection, predictive alerting, and alert correlation. 

  • Develop and refine prompts, agents, and integrations that connect monitoring, ticketing, and remediation systems into faster end-to-end response loops. 

  • Evaluate emerging AI tooling for reliability and operations use cases, and advocate for adoption where it delivers measurable improvements in toil reduction, MTTR, or availability. 

  • Ensure all AI and automation usage adheres to FreedomPay’s security, privacy, and PCI obligations — keeping sensitive data appropriately protected and human review in place for high-impact actions. 
Required Background and Experience

  • BS degree in Computer Science or equivalent, or equivalent years of relevant experience.
  • Minimum of 5 years of hands-on technical experience in highly available, high-throughput, web-based technology environments.
  • Demonstrated history of self-directed learning — someone who independently seeks out knowledge, builds new skills without being told to, and doesn’t wait for formal training to close gaps.
  • Next-level problem-solving abilities and a strong bias toward practical, proven solutions.
  • A track record of identifying and eliminating manual toil through automation.
  • Excellent communication and organizational skills, with a strong sense of ownership and service. 
Required Technical Skills

  • Expert-level proficiency in an enterprise APM platform and its AI/ML-driven (AIOps) capabilities; Dynatrace experience strongly preferred, though deep expertise in comparable tools such as Datadog or New Relic where readily transferable.
  • Hands-on experience with AI-assisted development and automation tools — such as Anthropic (Claude), OpenAI (Codex), and Azure AI services (Foundry, Azure SRE Agent) — and a demonstrated ability to apply them to real operational and engineering work.
  • Proficiency in scripting and automation — PowerShell and/or Python — to build tooling and remediation workflows.
  • Strong SQL / T-SQL skills.
  • Solid understanding of core networking concepts: DNS, load balancing, and TCP/IP routing and switching.
  • Working knowledge of modern technology infrastructure including container orchestration, IaaS/PaaS cloud services, Azure, and VMware.
  • Working knowledge of application development processes. 
Preferred Technical Skills and Experience

  • Proven track record of successfully implementing SLI/SLOs and fostering their adoption across an organization.
  • Experience implementing enterprise incident management practices.
  • Experience building AIOps or ML-driven automation into production observability and incident response.
  • Azure Kubernetes Service (AKS) and broader container orchestration experience.
  • Windows Server (IIS) administration.
  • PagerDuty Process Automation (formerly Rundeck) or comparable runbook automation platforms.
  • Comprehensive experience supporting real-time transaction processing applications.
  • PCI policies and best practices. 
Additional Experience, a Plus

  • AI/ML model deployment, evaluation, or operations (MLOps). 

  • Documentation automation and self-service tooling / service catalog implementation. 

  • Experience integrating QA test automation into CI/CD pipelines. 

 

 

As the fastest growing commerce company in the industry, we offer the opportunity for tremendous upward mobility within the company as well as development and professional growth opportunities. FreedomPay's fulltime roles provide exceptional benefits including medical, prescription, dental and vision coverage, Life Insurance, Retirement Plans with company match, commission sharing plan, flexible hybrid working environment, and great parental and other leave programs. All positions must be able to successfully pass a background check as well as a credit check.

 

FreedomPay is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Vacancy posted 22 days ago
Similar jobs that could be interesting for youBased on the Sr. Site Reliability Engineer in Remote vacancy
  •  ...Senior Site Reliability Engineer The FreedomPay Commerce Platform is the technology of choice for many of the largest companies across the globe in retail, hospitality, lodging, gaming, sports and entertainment, foodservice, education, healthcare and financial services... 
    Senior
    Full time
    Casual work
    Remote work
    Flexible hours

    FreedomPay

    Philadelphia, PA
    4 days ago
  • $178.13k - $205.4k

     ...telecommuting. Salary Range: $178,131 - $205,400 About You Basic Qualification ​Bachelor’s degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5) years of progressive, post‑baccalaureate experience in job offered... 
    Senior
    Work at office
    Remote work
    Flexible hours

    Workday

    Atlanta, GA
    1 day ago
  •  ...Recognized as the No. 1 site trusted by real estate professionals, Realtor.com® has been at the forefront of online real estate...  ...confidence through expert guidance. We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,... 
    Senior
    Work at office
    Local area
    Immediate start
    Flexible hours

    realtor.com

    Austin, TX
    3 days ago
  •  ...SLOs, SLIs, and error budgets in partnership with product and engineering teams. Lead incident response: triage, coordinate remediation...  ...habits. Collaborate with software engineers to establish reliability‑first design patterns and review architectures for operational... 
    Senior
    Remote work

    DOMA Technologies

    Leesburg, VA
    5 days ago
  • $110k - $145k

    Senior Site Reliability Engineer $110,000 - $145,000 / year + Bonus The insurance industry runs on Vertafore. We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data‑driven foundation to eliminate distribution drag across the insurance... 
    Senior
    Temporary work
    Work at office
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    2 days ago
  • Role Description Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services. Our SRE teams solve reliability, security, and usability at scale for... 
    Senior
    Full time
    Work at office

    Akamai

    Remote
    4 days ago
  • Role Description Stack AV Site Reliability Engineers are responsible for enabling and ensuring our production systems meet their service-level objectives. Through the implementation of centralized observability and automation, the SRE team constantly ensures the health... 
    Senior
    Full time

    Stack AV

    Remote
    4 days ago
  • Role Description As a Senior Site Reliability Engineer you will champion all things pertaining to reliability at Okta for Auth0. Working closely with the Product Engineers, Quality Engineers, Platform Engineers and Architecture teams, your primary focus will be on ensuring... 
    Senior
    Full time
    Remote work

    Okta

    Remote
    7 days ago
  •  ...Istio) ~Defining and monitoring Service-Level Objectives (SLOs) and Service-Level Agreements (SLAs) to ensure that systems meet reliability and performance targets ~Monitoring Tools like New Relic, Prometheus, Grafana, and/or Datadog ~OpenTelemetry knowledge for... 
    Senior
    Full time
    Remote work

    Shippo

    Remote
    8 days ago
  • Role Description We’re looking for a Senior Site Reliability Engineer who takes ownership seriously — someone who designs for reliability, ships the automation, and stands behind it in production. You’ll work across cloud-native infrastructure on systems that process millions... 
    Senior
    Full time

    CertifyOS

    Remote
    4 hours ago
  • $54k - $150k

    Role Description As Senior Site Reliability Engineer for Remote Build, you'll own the operational excellence and infrastructure strategy that makes Build's platform reliable, performant, and safe for customers. You'll report to the Engineering Manager and work closely with... 
    Senior
    Full time
    Local area
    Remote work
    Home office
    Flexible hours

    Remote - Referral Board

    Remote
    2 days ago
  •  ...through intelligent automation and modern engineering. We are seeking a Senior SRE Engineer...  ...efficient delivery, observability, and reliability across Sleek’s products and internal operations...  ...~6+ years of progressive experience in Site Reliability Engineering (SRE). ~6+... 
    Senior
    Full time
    Remote work
    Flexible hours

    Sleek

    Remote
    4 days ago
  • $54k - $150k

    Role Description As Senior Site Reliability Engineer for Remote Build, you'll own the operational excellence and infrastructure strategy that makes Build's platform reliable, performant, and safe for customers. You'll report to the Engineering Manager and work closely with... 
    Senior
    Full time
    Local area
    Immediate start
    Remote work
    Home office
    Flexible hours

    Referral Board

    Remote
    4 hours ago
  • Role Description Versant's Sports & Entertainment Digital Products division is seeking a Senior Site Reliability Engineer to help drive the reliability, scalability, and usability of internal developer platforms, tooling, and engineering workflows across a portfolio of... 
    Senior
    Full time
    Local area
    Remote work
    Worldwide

    Versant

    Remote
    3 days ago
  • Role Description We are seeking a Site Reliability Engineer (SRE) with deep expertise in monitoring, observability, and reliability engineering to support systems running across on-premises infrastructure and Google Cloud Platform (GCP). This role is primarily responsible... 
    Senior
    Long term contract
    Full time
    Remote work
    Flexible hours

    Devsu

    Remote
    1 day ago
  • Role Description We are looking for a talented and driven Sr. Site Reliability Engineering (SRE) to support our engineering team, which manages the infrastructure and services that power our Waystar products. This role is ideal for an experienced engineer who thrives in... 
    Senior
    Full time
    Live out
    Flexible hours

    Waystar

    Remote
    8 days ago
  • Role Description We’re looking for a Senior Platform Engineer to design, build, and operate the core services that power Optura’s AI...  ...systems end-to-end, from model and agent orchestration to routing, reliability, and observability. You will partner closely with product and... 
    Senior
    Full time
    Remote work

    Optura

    Remote
    8 days ago
  • $135k - $170k

    Role Description Climavision is seeking a Senior Site Reliability Engineer to contribute towards reliability, operational excellence, and production resilience for our customer-facing platform and weather data services. This role is focused on ensuring our systems consistently... 
    Senior
    Full time
    Temporary work
    Flexible hours

    Climavision

    Remote
    4 days ago
  • $137.9k - $221.4k

     ...for someone to lead development aspects of the Infrastructure engineering team at ServiceTitan. You must have a strong background in...  ...leadership and strong architectural thought process. Our Site Reliability and Infrastructure Engineering team is an investment by Cloud... 
    Senior
    Full time
    Immediate start
    Flexible hours

    ServiceTitan

    Remote
    6 days ago
  • $165.5k - $289.6k

    Role Description We are seeking an exceptional Sr Staff Site Reliability Engineer to lead critical infrastructure initiatives and drive innovation across our organization. You'll architect scalable solutions, navigate complex technical challenges independently, and deliver... 
    Senior
    Full time
    Flexible hours

    ServiceNow

    Remote
    2 days ago
  • $125k - $135k

    Role Description Vultr is seeking a highly skilled and experienced Senior Site Reliability Engineer to build and own the observability pipeline for the physical and provisioning infrastructure that powers Vultr's global datacenter footprint. The ideal candidate is a builder... 
    Senior
    Full time
    Work at office
    Immediate start
    Remote work

    Vultr

    Remote
    8 days ago
  •  ...The Role We're looking for a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of the production systems that power Nectar's platform. We run high-volume data ingestion pipelines and real-time AI agents on top of a fast... 
    Senior
    Remote work

    Nectar Social

    Palo Alto, CA
    5 days ago
  •  ...Partner with software developers, platform engineers, and IT staff to improve system design,...  ...requirements, service quality, reliability, security, and compliance needs. Drive...  ...Required: ~8+ years of experience in Site Reliability Engineering, DevOps, Platform... 
    Senior
    Work at office
    Remote work

    Applied Research Associates

    Albuquerque, NM
    4 days ago
  •  .... The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi-cloud...  ...AI-assisted development workflows Partner closely with engineering on reliability reviews and architecture decisions Requirements... 
    Senior
    Remote work

    Satsuma

    United States
    4 days ago
  •  ...Seeking a full-time Senior Site Reliability Engineer with expertise in C# and .NET to ensure production reliability for customer-facing platforms and weather data services in a remote setting, focusing on high availability, incident response, and operational excellence... 
    Senior
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    1 day ago
  •  ...To support a growing infrastructure team, the full-time Senior Site Reliability Engineer II - Infrastructure (AI Native) will design and maintain scalable platforms for over 200 backend services, utilizing AI tools to enhance operational efficiency in a remote work environment... 
    Senior
    Full time
    Remote work

    Virtual Vocations Inc

    United States
    3 days ago
  • $135k - $150k

    Senior Site Reliability Engineer Job number: 884 This is a remote position. Ad Hoc is a technology company that empowers organizations to deliver scalable, impactful digital services. Using modern, agile methods, our team creates products that meet people's... 
    Senior
    Remote work
    Flexible hours

    Ad Hoc LLC

    Silver Spring, MD
    4 days ago
  •  ...networking, serverless, and AI services, ensuring global scalability, reliability, and performance. The ideal candidate is a strategic thinker...  ...deep technical expertise in cloud infrastructure, platform engineering, and AI systems, capable of bridging architecture vision with... 
    Senior
    Full time
    Contract work

    Bitdeer Technologies Group

    Remote
    3 days ago
  •  ...Senior Site Reliability Engineer We are seeking a full-time Senior Site Reliability Engineer at Garmin's U.S. headquarters in the Greater Kansas City area. In this role, you will be responsible for ensuring the integrity of Garmin's production environment is maintained... 
    Senior
    Full time
    Work experience placement
    Local area
    Remote work

    Garmin

    Olathe, KS
    4 days ago
  • $160k - $180k

     ...Senior Site Reliability Engineer Location: Remote Compensation: $160,000 - 180,000 per year, depending on experience and qualifications. Employment Type: Full-Time What you can expect as the Senior Site Reliability Engineer at Fortress… The Senior Site Reliability... 
    Senior
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    Fortress Information Security

    Orlando, FL
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Site Reliability Engineer. Be the first to apply!