Sr. Site Reliability Engineer
FreedomPay
Job Description
Job Description
The FreedomPay Commerce Platform is the technology of choice for many of the largest companies across the globe in retail, hospitality, lodging, gaming, sports and entertainment, foodservice, education, healthcare and financial services. FreedomPay’s technology has been purposely built to deliver rock solid performance in the highly complex environment of global commerce. The company maintains a world-class security environment and was first to earn the coveted validation by the PCI Security Standards Council against Point-to-Point Encryption with EMV standard in North America. FreedomPay’s robust solutions across payments, security, identity and data analytics are available in-store, online and on-mobile and are supported by rapid API adoption. The award winning FreedomPay Commerce Platform operates on a single, unified technology stack across multiple continents allowing enterprises to deliver a consistent, repeatable experience on a global scale. FreedomPay is a fast paced, high growth company with a great culture with competitive benefits and compensation with a business casual atmosphere.
FreedomPay is seeking an experienced Senior Site Reliability Engineer to help ensure the highest possible availability and resiliency of a rapidly growing global payment platform. This full-time salaried position builds on a strong foundation of observability, incident response, and support experience across the development lifecycle — and pushes it forward with AI-driven operations and automation at its core. The right candidate finds real satisfaction in eliminating manual toil, treats every recurring task as an automation opportunity, and is eager to apply modern AI tooling to detect, diagnose, and resolve issues faster than ever before.
About the RoleYou’ll join a team of SREs who work closely with other teams of world-class engineers to tenaciously and creatively solve problems and reduce manual toil wherever possible. We expect AI and automation to be a force multiplier in everything you do — from accelerating root-cause analysis and enriching alerts, to generating runbooks and codifying remediation so that the platform increasingly heals itself.
Successful candidates are heavily results-driven, bring well-established expertise across both traditional and bleeding-edge technology, and have a strong desire to continuously grow and improve themselves and our platform. This is a global operation spanning multiple regions and time zones, and the role demands the flexibility and commitment that a 24/7 payment platform requires.
This position participates in an engineering on-call rotation and provides after-hours support for production issue escalations on a rotational basis.
This position is based in the Philadelphia area with a hybrid schedule. Remote arrangements may be considered for exceptional candidates, with occasional travel to Philadelphia required.
- Build and maintain a comprehensive understanding of the platform and custom application stack.
- Implement, maintain, and continuously improve observability strategies and metrics that ensure complete system health for numerous complex products throughout all stages of the development lifecycle, up to and including production.
- Continuously identify automation opportunities and follow through to successful implementation, applying AI-assisted tooling to accelerate development and reduce manual effort.
- Design, build, and maintain automated remediation and self-healing workflows that detect, triage, and resolve common failure modes with minimal human intervention.
- Leverage AI/ML-driven observability — anomaly detection, alert correlation, and intelligent noise reduction — to surface issues earlier and shorten time to detection.
- Use AI-assisted analysis to accelerate root-cause investigation, enrich incident context, and generate first-draft postmortems and runbooks for human review.
- Handle escalations and collaborate effectively with other team members to quickly determine the root cause of any type of service degradation.
- Implement, maintain, and continuously improve incident response procedures and other operational documentation, automating documentation generation and upkeep wherever practical.
- Assist with troubleshooting and remediation of failed scheduled jobs and data-related concerns.
- Champion responsible, secure adoption of AI tooling across the SRE function — sharing patterns, prompts, and automations that raise the productivity of the whole team
AI and automation are central to how this team operates. We are looking for someone who will not only use these tools but help define how the SRE function applies them. In this role you will:
- Apply AI-assisted development and operations tools — including Anthropic (Claude), OpenAI (Codex), and Azure AI services (Foundry, Azure SRE Agent) and the agentic workflows built on them — to write, review, and accelerate automation and infrastructure code.
- Build and integrate automation that turns repetitive operational work into codified, repeatable, and self-service workflows.
- Use AIOps and ML-driven observability capabilities within the APM stack for anomaly detection, predictive alerting, and alert correlation.
- Develop and refine prompts, agents, and integrations that connect monitoring, ticketing, and remediation systems into faster end-to-end response loops.
- Evaluate emerging AI tooling for reliability and operations use cases, and advocate for adoption where it delivers measurable improvements in toil reduction, MTTR, or availability.
- Ensure all AI and automation usage adheres to FreedomPay’s security, privacy, and PCI obligations — keeping sensitive data appropriately protected and human review in place for high-impact actions.
- BS degree in Computer Science or equivalent, or equivalent years of relevant experience.
- Minimum of 5 years of hands-on technical experience in highly available, high-throughput, web-based technology environments.
- Demonstrated history of self-directed learning — someone who independently seeks out knowledge, builds new skills without being told to, and doesn’t wait for formal training to close gaps.
- Next-level problem-solving abilities and a strong bias toward practical, proven solutions.
- A track record of identifying and eliminating manual toil through automation.
- Excellent communication and organizational skills, with a strong sense of ownership and service.
- Expert-level proficiency in an enterprise APM platform and its AI/ML-driven (AIOps) capabilities; Dynatrace experience strongly preferred, though deep expertise in comparable tools such as Datadog or New Relic where readily transferable.
- Hands-on experience with AI-assisted development and automation tools — such as Anthropic (Claude), OpenAI (Codex), and Azure AI services (Foundry, Azure SRE Agent) — and a demonstrated ability to apply them to real operational and engineering work.
- Proficiency in scripting and automation — PowerShell and/or Python — to build tooling and remediation workflows.
- Strong SQL / T-SQL skills.
- Solid understanding of core networking concepts: DNS, load balancing, and TCP/IP routing and switching.
- Working knowledge of modern technology infrastructure including container orchestration, IaaS/PaaS cloud services, Azure, and VMware.
- Working knowledge of application development processes.
- Proven track record of successfully implementing SLI/SLOs and fostering their adoption across an organization.
- Experience implementing enterprise incident management practices.
- Experience building AIOps or ML-driven automation into production observability and incident response.
- Azure Kubernetes Service (AKS) and broader container orchestration experience.
- Windows Server (IIS) administration.
- PagerDuty Process Automation (formerly Rundeck) or comparable runbook automation platforms.
- Comprehensive experience supporting real-time transaction processing applications.
- PCI policies and best practices.
AI/ML model deployment, evaluation, or operations (MLOps).
Documentation automation and self-service tooling / service catalog implementation.
Experience integrating QA test automation into CI/CD pipelines.
As the fastest growing commerce company in the industry, we offer the opportunity for tremendous upward mobility within the company as well as development and professional growth opportunities. FreedomPay's fulltime roles provide exceptional benefits including medical, prescription, dental and vision coverage, Life Insurance, Retirement Plans with company match, commission sharing plan, flexible hybrid working environment, and great parental and other leave programs. All positions must be able to successfully pass a background check as well as a credit check.
FreedomPay is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.
- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization,...SeniorWork at officeLocal area
$134.25k - $214.8k
...change. Constantly grow as you work hard for a mission that matters at a company where you matter.Your ImpactAs a Senior Site Reliability Engineer within the APX SRE organization, you’ll focus on delivering practical, scalable solutions to support the reliability and performance...SeniorWork at officeRemote workFlexible hours$87.12k - $151.25k
...to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a Digital Site Reliability Sr Engineer - Remote to join our team in Memphis, Tennessee (US-TN), United States (US).Digital Site Reliability Senior EngineerWe are...SeniorTemporary workWork at officeRemote workFlexible hours- ...consumers and companies, alikeKlover’s engineering team powers one of the fastest-growing fintech... ...-grade systems that prioritize reliability, security, and performance, and that integrate... ...the RoleAs a Senior/Staff Site Reliability Engineer, you will play a critical...SeniorWork at officeImmediate startRemote work
$232k - $263k
...scaling rapidly toward long-term growth and IPO readiness.Sr. Staff Site Reliability EngineerAs a Sr. Staff SRE at Obsidian, you will define... ...will operate as a strategic partner to DevOps and Platform Engineering leadership, shaping a unified reliability strategy that...SeniorWork from home$160k - $185k
...in their fitness journey and revolutionized the industry along the way. And we’re just getting started!OverviewThe Sr. Manager, Site Reliability Engineering (SRE) leads the strategy, execution, and continuous improvement of reliability, availability, and performance...SeniorWork at officeLocal areaRemote workWork from home- ...capabilities successfully. Provide expert-level guidance to engineering and product teams, contributing to high-level architecture... ...and recommending improvements for scalability, performance, reliability, and operational readiness. Partner with application teams...SeniorRemote workFlexible hours
- ...Pismo Platform Squad Engineer Join Pismo's Platform squad within the SRE Tribe, dedicated to owning and evolving the containerized... ...workloads. You'll work cross-functionally to ensure our platform is reliable, scalable, secure, and easy to operate, focusing on Kubernetes...SeniorWork at officeLocal areaRemote work
$165k - $225k
...demanding AI workloads with enterprise-grade reliability and compliance. Your Role: You will... ...core. Working closely with our systems engineers, network engineers, and platform... ...custom Kubernetes networking solutions with SR-IOV for high-performance GPU interconnects...SeniorRemote workFlexible hours- ...Senior Site Reliability Engineer We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms...SeniorRemote work
- ...Sr Site Reliability Engineer (SRE) SigNoz is an open-source observability platform that helps modern engineering teams monitor, debug, and optimize their applications with deep visibility into metrics, traces, and logs — all in one place. We're built natively on OpenTelemetry...SeniorRemote work
- ...Site Reliability Engineer As a Site Reliability Engineer, you will play a critical role in ensuring the reliability, availability, and performance of our systems. You will be responsible for designing, implementing, and maintaining scalable infrastructure solutions...SeniorRemote workFlexible hours
$139.76k - $287.75k
...Salary: $139,764 - 287,749 per year Requirements: We bring at least 4 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure. We have strong hands-on experience running AWS in production. We possess deep expertise...SeniorFull timeRemote workRelocation package- ...straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements: As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data...SeniorRemote work
$178.13k - $205.4k
...Bachelor's degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5)... ...websites that are not Workday Careers. Please be aware of sites that may ask for you to input your data in connection with a job...SeniorWork at officeRemote workFlexible hours- Role Description We're looking for an SRE to own the reliability, scalability, and operability of the SigNoz cloud platform. You'll keep... ...Benefits ~Work on a globally used open-source project that engineers actually love. ~Huge scope and ownership — your work directly...SeniorFull timeRemote work
$110k - $155k
...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role is...SeniorContract workWork at officeWork from homeFlexible hours- ...grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers...SeniorWork experience placementRemote workFlexible hours
$125k - $145k
...General information Press space or enter keys to toggle section visibility Job Title Sr. Site Reliability Engineer Functional Area Teammate - Information Technology City Remote Work Location Type...SeniorFull timeWork experience placementLocal areaRemote workFlexible hoursShift work- ...Job Description The Opportunity: Versant's Sports & Entertainment Digital Products division is seeking a Senior Site Reliability Engineer to help drive the reliability, scalability, and usability of internal developer platforms, tooling, and engineering workflows...SeniorLocal areaRemote workWorldwide
$140k - $170k
..., and intelligence. If you want to push yourself and reshape a $200B+ market, we're excited to talk to you! What will the Site Reliability Engineer do? We're looking for a Senior Site Reliability Engineer who's passionate about building and maintaining reliable, scalable...SeniorFull timeImmediate startRemote workVisa sponsorshipFlexible hours$151.5k - $252.5k
...artifacts rather than getting direct access to environments from day one. This is a ground-up role - you'll help define how reliability engineering works here by mapping systems, writing runbooks, setting baselines, and building the practices this team will run on going...SeniorBase plus commissionLocal areaRemote workWorldwide$120k - $170k
Sr. Manager/Manager Site Reliability Engineering Join to apply for the Sr. Manager/Manager Site Reliability Engineering role at Aritzia Sr. Manager/Manager Site Reliability Engineering 1 day ago Be among the first 25 applicants Join to apply for the Sr. Manager/Manager...SeniorFull timeWork at officeRemote workFlexible hours- Sr Site Reliability Engineer (Linux, UNIX, Reliability Engineering, Python, C, C++, Java, DevOps) in New York City C, C++, DevOps Engineer, Java, Linux, Perl, Python, Reliability Engineering, SQL, Unix Location: New York Job Function: Reliability Engineering Date Of...SeniorPermanent employmentFull timeRemote work
$65 - $75 per hour
DescriptionKforce has a client seeking a remote Senior Site Reliability Engineer to be a l be a leading member of the team working with a diverse range of technologies. You will enjoy working in a friendly environment and benefit from our investment in staff. The role also...SeniorRemote work- ...: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8 to... ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will have...SeniorRemote work
- ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...networking teams to improve service reliability and deployment... ...rotationYouHave 5+ years of experience in Site Reliability Engineering,... ...virtualization technologies, SR-IOV, and DPDKUnderstanding of...SeniorWork at officeLocal areaWork from homeFlexible hours
- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...SeniorWork at officeLocal areaWork from homeFlexible hours
$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...SeniorFull timeWork at officeLocal areaRemote workWork from home$210k - $230k
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation...SeniorCurrently hiringRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. Site Reliability Engineer. Be the first to apply!
- site reliability engineer remote Remote
- site reliability engineer sre Remote
- site reliability engineer Remote
- senior operations associate Remote
- senior safety specialist Remote
- senior technology project manager Remote
- senior c++ developer Remote
- remote senior business analyst Remote
- senior manager clinical operations Remote
- senior supervisor Remote



