Manager, Site Reliability Engineering - Paylo Platform
PDI Technologies
Job Description
Job Description
At PDI Technologies, we empower some of the world's leading convenience retail and petroleum brands with cutting-edge technology solutions that drive growth and operational efficiency. By “Connecting Convenience” across the globe, we empower businesses to increase productivity, make more informed decisions, and engage faster with customers through loyalty programs, shopper insights, and unmatched real-time market intelligence via mobile applications, such as GasBuddy. We’re a global team committed to excellence, collaboration, and driving real impact. Explore our opportunities and become part of a company that values diversity, integrity, and growth.
Role Overview
PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This role owns the reliability, infrastructure, and operational strategy for a portfolio of high-traffic, customer- and partner-facing platforms that power payment transactions, fuel pricing, loyalty and rewards, and offer/coupon redemption for convenience retail and fuel customers around the world.
This is a hands-on, leadership-first role. You will manage a team of three SRE Managers/Leads who together lead approximately 20 engineers, while staying technically engaged yourself — reviewing architecture, unblocking hard infrastructure problems, and setting the technical bar across the organization. You will bring strong, current, hands-on expertise across AWS, Azure, Kubernetes, Helm, Argo CD, Terraform/OpenTofu, Jenkins, and Datadog, and you will be a strong, visible people leader who can coach managers and represent SRE to senior engineering and business stakeholders.
Key ResponsibilitiesDirectly manage and develop 3 SRE Managers/Leads and own the overall health, growth, and performance of an ~20-person SRE organization supporting the Paylo product suite.
Set the vision, priorities, and operating cadence for the SRE function; translate business and product priorities into a reliability roadmap your managers can execute against.
Build a strong bench by hiring, coaching, and developing managers and senior engineers while creating clear career paths and succession plans.
Foster a blameless, learning-oriented culture around incidents, on-call, and operational excellence.
Partner closely with engineering directors, product managers, and business stakeholders across the Paylo organization to align reliability investments with business risk and customer impact.
Stay technically engaged day to day by participating in architecture and design reviews, troubleshooting complex production issues, and directly contributing to infrastructure-as-code, Kubernetes manifests/Helm charts, and CI/CD pipelines when needed.
Set and enforce engineering standards for multi-cloud infrastructure across AWS and Azure and for container orchestration on Kubernetes at scale.
Own adoption and standards for GitOps-based continuous delivery using Argo CD/Argo Workflows, including deployment strategy, rollout policy, and multi-cluster promotion.
Own the Infrastructure-as-Code strategy across teams (Terraform, OpenTofu), including module standards, state management, drift detection, and remediation.
Own CI/CD pipeline architecture and standards built on Jenkins, driving build/deploy automation, pipeline reliability, and progressive delivery practices such as blue-green/canary deployments and automated rollback.
Evaluate and guide adoption of new infrastructure tooling and patterns as the platform evolves across AWS and Azure.
Own the observability strategy across all supported products, with deep, hands-on expertise in Datadog (APM, infrastructure monitoring, log management, dashboards, and alerting) as the standard platform for metrics, tracing, and alerting.
Define and drive adoption of SLIs/SLOs, error budgets, and reliability KPIs across the organization, holding managers and teams accountable to them.
Own the incident management program end to end, including on-call structure, escalation paths, severity definitions, postmortems, and follow-through on remediation actions.
Drive root-cause analysis and long-term reliability investments that reduce Sev1/Sev2 frequency and recurrence.
Ensure appropriate resilience, disaster recovery, and capacity planning practices are in place given the sensitivity of payment- and transaction-related systems.
Partner with Security and Compliance to maintain awareness of PCI DSS and related compliance requirements and ensure the SRE organization supports audit and compliance readiness.
Track and report cost, capacity, and operational KPIs to senior leadership.
8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure/Platform Engineering, including 4+ years in a people-leadership role.
Proven experience managing managers — you have directly led team leads/managers, not just individual contributors, and are comfortable operating at the scale of ~20 total reports.
Strong, hands-on expertise across AWS and Azure — you can architect, troubleshoot, and operate multi-cloud infrastructure yourself, not just direct others to do so.
Strong, hands-on expertise with Kubernetes and Helm — cluster operations, troubleshooting at scale, and chart design/maintenance.
Strong, hands-on expertise with Argo CD/Argo Workflows for GitOps-based continuous delivery.
Strong, hands-on expertise with Infrastructure as Code (Terraform, OpenTofu), including module design and state management.
Strong, hands-on expertise with Jenkins for CI/CD pipeline design, administration, and automation.
Strong, hands-on expertise with Datadog (or equivalent enterprise observability platform), including designing monitoring/alerting strategy, dashboards, and APM/tracing at scale.
Demonstrated track record of driving incident management, on-call, and postmortem programs for high-traffic, customer-facing systems.
Excellent communication and stakeholder-management skills; able to represent SRE to engineering leadership and business partners with equal credibility.
A strong, visible leadership style — someone who sets clear direction, holds teams accountable, and builds trust across the organization.
- Applicants must be legally authorized to work in the United States without the need for employer sponsorship, now or in the future. PDI Technologies is unable to offer visa sponsorship for this role.
Experience supporting payments, fuel/retail, or loyalty platforms, or other systems with PCI DSS or similar compliance obligations.
Relevant certifications such as CKA/CKAD, AWS Certified Solutions Architect, Microsoft Certified: Azure Solutions Architect, or HashiCorp Terraform Associate.
Experience with messaging systems (Kafka/SQS/SNS), PagerDuty (or similar), and multi-region/multi-AZ resilience patterns.
Prior experience consolidating or standardizing SRE and DevOps practices across multiple product lines or recently- integrated/acquired teams.
Experience partnering with product and business stakeholders to translate reliability investments into business outcomes.
A stable, well-led SRE organization with clear ownership, career paths, and low regrettable attrition among your managers and their teams.
Consistent, Datadog-driven observability and SLOs in place across the organization, with measurable reduction in Sev1/Sev2 incidents and mean time to detect/resolve.
Modern, standardized infrastructure practices — GitOps delivery via Argo, IaC via Terraform/OpenTofu, and reliable CI/CD via Jenkins — adopted consistently across teams and clouds.
A mature, blameless incident-management culture with strong postmortem follow-through.
Strong cross-functional trust with engineering, product, and security/compliance stakeholders.
PDI is committed to offering a well-rounded benefits program, designed to support and care for you, and your family throughout your life and career. This includes a competitive salary, market-competitive benefits, and a quarterly perks program. We encourage a good work-life balance with ample time off [time away] and, where appropriate, hybrid working arrangements. Employees have access to continuous learning, professional certifications, and leadership development opportunities. Our global culture fosters diversity, inclusion, and values authenticity, trust, curiosity, and diversity of thought, ensuring a supportive environment for all.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
- About the RoleThe Enterprise Platforms team within our Data & Agentic Platform Solutions... ...and DevOps capabilities. As an IT Site Reliability Engineer within the Enterprise Platforms team,... ...— the backbone of TI's API management and integration strategy — while also...SuggestedLocal area
- Site Reliability Engineer - Vice PresidentSite Reliability Engineering (SRE) is an engineering discipline... ...of the firm’s most critical platform services and ensures they meet the requirements... ...the enterprise.Complex Incident Management & Post-Mortem Analysis: Lead critical...Suggested
$147k - $210k
...product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e.g., style... ..., and troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when you treat operations...Suggested$174k - $252k
...consulting, developing software platforms and frameworks, capacity... ...pushing for changes that improve reliability and velocity.Practice... ...degree in Computer Science, Engineering, a related field, or equivalent... ...Science or Engineering.Site Reliability Engineering (SRE...Suggested$128.6k - $184.9k
...that powers our global cloud platform. As a team of six engineers distributed across the US,... ...focus on automation, reliability, and operational excellence... ...than 500,000 customers and manages over 18 million devices worldwide... ...7+ years of experience in Site Reliability Engineering,...SuggestedPermanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours$138.4k - $173k
...well as help improve the reliability, quality of services... ..., configuration management, DDoS protection, infrastructure... ...or embed with engineering teams, helping them to... ...AppFolio Real Estate Platform. You’ll help build the... ...locations by visiting our site.Compensation & BenefitsThe...Full timeFlexible hours$104.9k - $174.7k
...Credit Risk mitigation and Customer Data Management. You can learn more about LexisNexis... ...Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and... ...operating monitoring and uptime platforms such as Grafana, Pingdom, and UptrendsStrong...Full timeWork at officeLocal areaRemote workWork from home- Core Responsibilities:Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that... ...tooling.Actively participate in reliability engineering and resilience communities of practice, contributing...Full time
- Qualifications: 8+ years of Software Engineering experience, or... ...maintain scalable and reliable infrastructure on Google Cloud Platform (“GCP”) for Snowflake... ...effectively with the client, IT management and staff, and other... ...Willingness to work on-site at stated location in...Contract workFor contractorsWork experience placement
$120.6k - $150.9k
...highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our team as a technical... ...transformative impact across WEX’s platform reliability and operational excellence... ...observability, automation, incident management, problem management, capacity planning...Full timeFlexible hours- ...client-facing – internal & external); Managing and continually improving platform infrastructure and applications with high reliability, resiliency, performance & quality, and... ...runbooks/playbooks; and, Using Chaos Engineering to test the robustness of the systems and...
- ...Senior Site Reliability Engineer This role will require someone onsite at our client office in Cleveland, OH, Pittsburgh, PA, or Dallas, TX... ...who will work within the Production support and Performance Management practice to drive excellence and innovation to support...Work at officeFlexible hoursShift workWeekend work
$48 per hour
...Site Reliability Engineer Trident Consulting is seeking a Site Reliability Engineer for one of our... ...for Application Performance management to manage Transaction journeys. Experience... ...databases · Experience in transitioning platforms to the cloud and Containerization –...Contract work- ...Senior Site Reliability Engineer (Permanent Role) Cleveland, OH, Pittsburgh, PA, or Dallas, TX Your future duties and responsibilities... ...issues. . Assisting with troubleshooting on call. . Managing and tracking incidents such as outages. . Facilitating...Permanent employmentFull timeLocal areaFlexible hoursShift workWeekend work
- ...are highly available, reliable, and performant at a global... ..., measurability, and manageability are integrated into... ...application/service/platform uptime with minimal human... ...Science or related Engineering field required. Master... ...of lead experience of site reliability...Contract workWork at office
- ...Site Reliability Engineer (SRE) The successful applicant may be performing work in FedRAMP High or IL-5 environments, and therefore, must... ...Experience operating large-scale, globally distributed SaaS platforms. Familiarity with hybrid cloud environments and multi-region...Permanent employmentWorldwideShift work
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly... ...Analytics Office (CDAO) AI/ML & Data Platforms team, you will solve complex and broad... ...problems, coordinate incident management coverage, and proactively address issues...Work at office
- ...Site Reliability Engineer TXSE is building the next-generation exchange infrastructure to support transparent, efficient, and resilient capital... ...expertise. Monitoring & Incident Response: Monitor platform and infrastructure health, respond to incidents, troubleshoot...Currently hiring
- ...From generative AI and cloud-native platforms to advanced release engineering practices, our teams are... ...accelerate development and improve reliability. Your work will directly influence... ...performing Root Cause Analysis and Problem Management Experience working in Agile...Full timeH1bWork at officeRemote workVisa sponsorshipFlexible hours2 days per week3 days per week
- ...Job Title: Site Reliability Engineer (SRE) Work Location: RichardsonTX 75082 **FULL ONSITE WORK*... ...expertise in Kubernetes (AKS), cluster management, Helm Charts, Istio, and production... ...optimize cloud infrastructure, container platforms, and deployment pipelines....Contract work
- ...Senior Site Reliability Engineer Our client is seeking a Senior Site Reliability Engineer for a... ...New Relic or similar APM/observability platforms. ~ Experience using additional tools... ..., Service Now. ~ Proven experience managing high-severity incidents and driving...Full timeContract workTemporary workWork experience placementFlexible hours
$72.1k - $158.62k
...seeking a highly skilled Software Development Engineer, Site Reliability Engineering (SRE), for Retail and Pharmacy platforms to drive reliability, scalability, and... ...analytics. Exposure to AIOps-driven incident management and self-healing architectures . Strong...Hourly payFull timeTemporary workLocal area- ...EngineeringWe are Compliance Engineering, a global team of more than... ...build and operate a suite of platforms and applications that... ...Responsibilities:Proactive management of our production services by... ...changes that improve capacity and reliability.Practicing sustainable incident...
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied... ...of high-visibility products and platforms and the environments they run in. Your... ...privilege/RBAC, deploy approvals, secrets management) in partnership with security and...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week- Compliance Engineering, Site Reliability Engineering, Vice President, Dallas location_on Dallas, TX, United... .... We build and operate a suite of platforms and applications that prevent,... ...SRE. Job Responsibilities: Proactive management of our production services by measuring...Full timeTemporary workWork at office
- ...can immediately lead engineers, influence senior stakeholders... ...of AI-enabled and platform-driven ways of working... ...combines change management, AI adoption, strategic... ...Services, Platform Support, Reliability Engineering & Agentic... ..., Platform Support, Site Reliability...Full timeWork experience placementWork at officeImmediate startWork from homeVisa sponsorship2 days per week3 days per week
- Engineering - SRE Platforms - Software Engineer - Vice President - Dallas Job Description Goldman Sachs is seeking a talented and motivated Site Reliability Engineering Manager to join our team. As a leader within the firm's Technology division, you will be responsible...
$156.18k
...Technology Consulting Senior Manager to join our growing Microsoft... ...integrations within the Dynamics 365 platform. Advanced proficiency in... ..., Computer Science, Software Engineering, or related field). 7+ years... ...offices and on client sites, which can include local or out...Full timeTemporary workWork at officeLocal areaRemote workFlexible hours- ...Pay Rate: $40/Hr. W2 Experience: 3-5 Years Overview We are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing in a DevOps‑driven environment...Long term contractContract workInternshipRemote work
$124k - $280k
...PwC, our people in data and analytics engineering focus on leveraging advanced technologies... ...Sets You Apart- Certification in Cloud Platforms [e.g., AWS Certified Solutions... ...solutions using cloud services- Designing and managing data warehouses and data lakes- Implementing...Full timeH1b
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Manager, Site Reliability Engineering - Paylo Platform. Be the first to apply!
- site reliability engineer Dallas, TX
- site reliability engineer sre Dallas, TX
- senior platform engineer Dallas, TX
- platform developer Dallas, TX
- platform engineer Dallas, TX
- client platform engineer Dallas, TX
- platform engineering manager Dallas, TX
- data platform engineer Dallas, TX
- junior website developer Dallas, TX
- website content developer Dallas, TX


