Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Manager, Site Reliability Engineering - Paylo Platform

PDI Technologies

Job Description

Job Description

At PDI Technologies, we empower some of the world's leading convenience retail and petroleum brands with cutting-edge technology solutions that drive growth and operational efficiency. By “Connecting Convenience” across the globe, we empower businesses to increase productivity, make more informed decisions, and engage faster with customers through loyalty programs, shopper insights, and unmatched real-time market intelligence via mobile applications, such as GasBuddy.  We’re a global team committed to excellence, collaboration, and driving real impact. Explore our opportunities and become part of a company that values diversity, integrity, and growth.

Role Overview

PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This role owns the reliability, infrastructure, and operational strategy for a portfolio of high-traffic, customer- and partner-facing platforms that power payment transactions, fuel pricing, loyalty and rewards, and offer/coupon redemption for convenience retail and fuel customers around the world. 

This is a hands-on, leadership-first role. You will manage a team of three SRE Managers/Leads who together lead approximately 20 engineers, while staying technically engaged yourself — reviewing architecture, unblocking hard infrastructure problems, and setting the technical bar across the organization. You will bring strong, current, hands-on expertise across AWS, Azure, Kubernetes, Helm, Argo CD, Terraform/OpenTofu, Jenkins, and Datadog, and you will be a strong, visible people leader who can coach managers and represent SRE to senior engineering and business stakeholders.

Key Responsibilities

  • Directly manage and develop 3 SRE Managers/Leads and own the overall health, growth, and performance of an ~20-person SRE organization supporting the Paylo product suite.

  • Set the vision, priorities, and operating cadence for the SRE function; translate business and product priorities into a reliability roadmap your managers can execute against.

  • Build a strong bench by hiring, coaching, and developing managers and senior engineers while creating clear career paths and succession plans.

  • Foster a blameless, learning-oriented culture around incidents, on-call, and operational excellence.

  • Partner closely with engineering directors, product managers, and business stakeholders across the Paylo organization to align reliability investments with business risk and customer impact.

  • Stay technically engaged day to day by participating in architecture and design reviews, troubleshooting complex production issues, and directly contributing to infrastructure-as-code, Kubernetes manifests/Helm charts, and CI/CD pipelines when needed.

  • Set and enforce engineering standards for multi-cloud infrastructure across AWS and Azure and for container orchestration on Kubernetes at scale.

  • Own adoption and standards for GitOps-based continuous delivery using Argo CD/Argo Workflows, including deployment strategy, rollout policy, and multi-cluster promotion.

  • Own the Infrastructure-as-Code strategy across teams (Terraform, OpenTofu), including module standards, state management, drift detection, and remediation.

  • Own CI/CD pipeline architecture and standards built on Jenkins, driving build/deploy automation, pipeline reliability, and progressive delivery practices such as blue-green/canary deployments and automated rollback.

  • Evaluate and guide adoption of new infrastructure tooling and patterns as the platform evolves across AWS and Azure.

  • Own the observability strategy across all supported products, with deep, hands-on expertise in Datadog (APM, infrastructure monitoring, log management, dashboards, and alerting) as the standard platform for metrics, tracing, and alerting.

  • Define and drive adoption of SLIs/SLOs, error budgets, and reliability KPIs across the organization, holding managers and teams accountable to them.

  • Own the incident management program end to end, including on-call structure, escalation paths, severity definitions, postmortems, and follow-through on remediation actions.

  • Drive root-cause analysis and long-term reliability investments that reduce Sev1/Sev2 frequency and recurrence.

  • Ensure appropriate resilience, disaster recovery, and capacity planning practices are in place given the sensitivity of payment- and transaction-related systems.

  • Partner with Security and Compliance to maintain awareness of PCI DSS and related compliance requirements and ensure the SRE organization supports audit and compliance readiness.

  • Track and report cost, capacity, and operational KPIs to senior leadership.

Required Qualifications

  • 8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure/Platform Engineering, including 4+ years in a people-leadership role. 

  • Proven experience managing managers — you have directly led team leads/managers, not just individual contributors, and are comfortable operating at the scale of ~20 total reports. 

  • Strong, hands-on expertise across AWS and Azure — you can architect, troubleshoot, and operate multi-cloud infrastructure yourself, not just direct others to do so. 

  • Strong, hands-on expertise with Kubernetes and Helm — cluster operations, troubleshooting at scale, and chart design/maintenance.  

  • Strong, hands-on expertise with Argo CD/Argo Workflows for GitOps-based continuous delivery. 

  • Strong, hands-on expertise with Infrastructure as Code (Terraform, OpenTofu), including module design and state management. 

  • Strong, hands-on expertise with Jenkins for CI/CD pipeline design, administration, and automation. 

  • Strong, hands-on expertise with Datadog (or equivalent enterprise observability platform), including designing monitoring/alerting strategy, dashboards, and APM/tracing at scale. 

  • Demonstrated track record of driving incident management, on-call, and postmortem programs for high-traffic, customer-facing systems. 

  • Excellent communication and stakeholder-management skills; able to represent SRE to engineering leadership and business partners with equal credibility. 

  • A strong, visible leadership style — someone who sets clear direction, holds teams accountable, and builds trust across the organization. 

  • Applicants must be legally authorized to work in the United States without the need for employer sponsorship, now or in the future. PDI Technologies is unable to offer visa sponsorship for this role.
Preferred Qualifications

  • Experience supporting payments, fuel/retail, or loyalty platforms, or other systems with PCI DSS or similar compliance obligations. 

  • Relevant certifications such as CKA/CKAD, AWS Certified Solutions Architect, Microsoft Certified: Azure Solutions Architect, or HashiCorp Terraform Associate. 

  • Experience with messaging systems (Kafka/SQS/SNS), PagerDuty (or similar), and multi-region/multi-AZ resilience patterns. 

  • Prior experience consolidating or standardizing SRE and DevOps practices across multiple product lines or recently- integrated/acquired teams. 

  • Experience partnering with product and business stakeholders to translate reliability investments into business outcomes. 

What Success Looks Like

  • A stable, well-led SRE organization with clear ownership, career paths, and low regrettable attrition among your managers and their teams. 

  • Consistent, Datadog-driven observability and SLOs in place across the organization, with measurable reduction in Sev1/Sev2 incidents and mean time to detect/resolve. 

  • Modern, standardized infrastructure practices — GitOps delivery via Argo, IaC via Terraform/OpenTofu, and reliable CI/CD via Jenkins — adopted consistently across teams and clouds. 

  • A mature, blameless incident-management culture with strong postmortem follow-through. 

  • Strong cross-functional trust with engineering, product, and security/compliance stakeholders. 

PDI is committed to offering a well-rounded benefits program, designed to support and care for you, and your family throughout your life and career.  This includes a competitive salary, market-competitive benefits, and a quarterly perks program. We encourage a good work-life balance with ample time off [time away] and, where appropriate, hybrid working arrangements.  Employees have access to continuous learning, professional certifications, and leadership development opportunities. Our global culture fosters diversity, inclusion, and values authenticity, trust, curiosity, and diversity of thought, ensuring a supportive environment for all.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 14 days ago
Similar jobs that could be interesting for youBased on the Manager, Site Reliability Engineering - Paylo Platform in Houston, TX vacancy
  •  ...Overview PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing...  ...high-traffic, customer- and partner-facing platforms that power payment transactions, fuel pricing... 
    Suggested

    PDI Technologies

    Houston, TX
    14 days ago
  • As an Entry-Level DevOps Site Reliability Engineer, you will join a team responsible for continuous improvement and support of customer facing...  ..., you will work closely with IT, Development, and Product Management teams across the business.#LI-DNIBenefitsWe strive to offer... 
    Suggested
    Work from home
    2 days per week

    Reynolds & Reynolds

    Houston, TX
    4 days ago
  •  ...EngineeringPosition Summary: DevOps / Site Reliability Engineer to implement and evolve the...  ...Python, or Go)Experience with incident management, SLObased reliability practices, and...  ...similar)Experience with virtualization platforms including VM provisioning, storage, networking... 
    Suggested
    Full time
    Local area

    VoltaGrid

    Houston, TX
    2 days ago
  •  ...complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the...  ...and scalability of your application or platform.Job ResponsibilitiesGuides and...  ...financial transaction processing and asset management. We offer a competitive total... 
    Suggested

    JP Morgan Chase

    Houston, TX
    2 days ago
  • Reliability Engineering Design, implement, and operate scalable, resilient, and...  ...systems on Google Cloud Platform. Improve service availability...  ...Observability and Incident Management Develop actionable alerts that...  ...years of experience in Site Reliability Engineering, platform... 
    Suggested
    Remote work

    Patterson-UTI

    Houston, TX
    2 days ago
  • As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate Know Your Customer (KYC)...  ...stability of your team’s applications and platforms using data-driven analytics to...  ...financial transaction processing and asset management.We offer a competitive total rewards... 

    JP Morgan Chase

    Houston, TX
    1 day ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly...  ...and scalability of your application or platform. Job Responsibilities Guides...  ...transaction processing and asset management. We offer a competitive total rewards... 

    Chase

    Houston, TX
    1 day ago
  • $104.9k - $174.7k

     ...passionate about improving reliability, scalability, and...  ...mitigation and Customer Data Management. You can learn more...  ...-based container platforms, leading vulnerability...  ...prioritization of reliability engineering tasks within team...  ...Background in DevOps, site reliability... 
    Full time
    Local area

    RELX

    Houston, TX
    3 days ago
  •  ...Site Reliability Engineer II About PROS: PROS, Inc. is the leading offer management provider to the airline industry, helping airlines deliver seamless retail experiences...  ...margin growth. Powered by AI, the PROS Platform enables commercial teams to align... 
    Flexible hours

    PROS Holdings, Inc.

    Houston, TX
    5 days ago
  •  ...Senior Site Reliability Engineer The Senior Site Reliability Engineer is responsible for improving...  ...of our critical infrastructure platforms and services. This role partners closely...  ...across compute, storage, networking, and managed services (e.g., autoscaling, load... 
    Work at office
    Local area

    Castleton Commodities International

    Houston, TX
    1 day ago
  •  ...Site Reliability Engineer As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident Management Maintain and monitor production systems for availability...  ...-on experience with cloud platforms (AWS, Azure, or GCP). ~ Strong... 
    Permanent employment
    Full time
    Shift work

    NOV

    Houston, TX
    4 days ago
  • $74.1k - $148.3k

     ...Site Reliability Engineer Solve complex problems related to infrastructure cloud services and build automation to prevent problem recurrence...  ..., deploying, and solving key Oracle Cloud services, platforms, and infrastructure, always thinking about reliability, scalability... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle

    Houston, TX
    4 days ago
  • $118k - $162k

     ...on Senior Mechanical Engineer & Team Lead to define...  ...Direction: Lead, task, and manage the mechanical...  ...anchorages comply with site-specific seismic and wind...  ..., supply chain, and platform solutions, Celestica enables...  ..., regulated and high-reliability markets such as... 
    Work at office
    Local area
    Remote work
    Worldwide
    Shift work

    Celestica

    Houston, TX
    5 days ago
  • DLA Piper is seeking a Director of Engineering to lead design, development, delivery, modernization, and ongoing support of enterprise technology solutions aligned with firm priorities. You will direct software engineering, application development, and systems integration... 

    DLA Piper

    Houston, TX
    3 days ago
  •  ...Services.Role OverviewThe Release Engineer is responsible for the...  ...applications and services across platforms such as Azure and AWS .This...  ...and security teams to deliver reliable and scalable cloud solutions....  ....Provision configure and manage cloud resources using Infrastructure... 

    LTM

    Houston, TX
    2 days ago
  • $156.18k

     ...Technology Consulting Senior Manager to join our growing Microsoft...  ...integrations within the Dynamics 365 platform. Advanced proficiency in...  ..., Computer Science, Software Engineering, or related field). 7+ years...  ...offices and on client sites, which can include local or out... 
    Full time
    Temporary work
    Work at office
    Local area
    Remote work
    Flexible hours

    Protiviti

    Houston, TX
    4 days ago
  • $99k - $232k

     ...PwC, our people in data and analytics engineering focus on leveraging advanced technologies...  ...team member’s unique strengths, and managing performance to deliver on client expectations...  ...Sets You Apart- Certification in Cloud Platforms [e.g., AWS Certified Solutions... 
    Full time
    H1b

    PwC

    Houston, TX
    2 days ago
  • $124k - $280k

     ...PwC, our people in data and analytics engineering focus on leveraging advanced technologies...  ...Sets You Apart- Certification in Cloud Platforms [e.g., AWS Certified Solutions...  ...solutions using cloud services- Designing and managing data warehouses and data lakes- Implementing... 
    Full time
    H1b

    PwC

    Houston, TX
    3 days ago
  •  ...in Advisory.KPMG is currently seeking a Manager, AI Engineer to join our Advisory Services practice....  ...certifications in cloud AI/ML platforms (for example: Azure AI Engineer Associate...  ...towards the bottom of our KPMG US Careers site at Benefits & How We Work.Follow this link... 
    H1b
    Local area

    KPMG

    Houston, TX
    1 day ago
  •  ...You AreManagers are the hands-on delivery engine of the Secure AI practice. They lead day...  ...into capable practitioners. Each Manager hire will be expected to operate across...  ...and Governance offerings powered by LLM platforms, frontier grade models, agentic frameworks... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Houston, TX
    1 day ago
  • $99k - $232k

     ...Description & SummaryThe OpportunityAs a SAP BRIM Consultant, Manager, you will lead our clients in their customer transformation journey...  ...in finance, supply chain, customer, human capital, and engineering.As a Manager, you will develop new skills outside your comfort... 
    Full time
    H1b

    PwC

    Houston, TX
    5 days ago
  • $99k - $232k

     ...for coaching, leveraging team member’s unique strengths, and managing performance to deliver on client expectations. With your growing...  ...requirements.The OpportunityAs part of the Data and Analytics Engineering team, you will serve as both a technical leader and a trusted... 
    Full time
    H1b

    PwC

    Houston, TX
    2 days ago
  • $91k - $321.5k

    Industry/SectorNot ApplicableSpecialismIFS - Information Technology (IT)Management LevelSenior ManagerJob Description & SummaryThe OpportunityAs a CTIO - AI Engineer- Senior Manager, you will play a pivotal role in transforming raw data into actionable insights, enabling... 
    Full time
    H1b

    PwC

    Houston, TX
    4 days ago
  •  ...cyber defense, application security, and managed service solutions to rethink the entire...  ...engagement. A Cybersecurity Forward Deployed Engineer is a production engineer who works...  ...multi-system integration risk across cloud platforms (AWS, Azure, or GCP)Lead AI governance... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Houston, TX
    4 days ago
  • $96 - $103 per hour

    DescriptionKforce has a client that is seeking a AWS DevOps/Platform Engineer III in Houston, TX.Summary:We are looking for a Senior (Level III) DevOps/Platform Engineer to design, build and operate scalable AWS Infrastructure and related DevOps for AI, Data Science, and... 
    Shift work

    KForce

    Houston, TX
    2 days ago
  •  ...Release Engineer Visa status: U.S. Citizens and those authorized to work in the U.S. are encouraged to apply. Tax Terms: W2, 1099 Corp...  ...is also preferred. Must be able to communicate effectively with Testers, Developers, DBA's, Infrastructure teams and Managers.... 

    Keylent Inc

    Houston, TX
    3 days ago
  •  ...Release Train Engineer 4 Months- Contract To Hire Pay- $65-$7...  ...deliver value, removes barriers, manages risk and dependencies, and...  ...Customer & Commercial business and platform partners keeping planning,...  ...and metrics are visible and reliable. Coach teams and Scrum... 
    Contract work
    Work at office

    Anveta

    Houston, TX
    1 day ago
  • $85.1k - $161.7k

     ...exclusively to serving the cloud engineering needs of our clients. This...  ..., implementation and managed solutions. As part of this...  ...infrastructure automation, and site reliability engineering (SRE) best...  ...solutions as code on client cloud platforms to include, to Amazon Web... 
    Full time
    Work experience placement
    Internship
    Local area

    RSM US LLP

    Houston, TX
    2 days ago
  •  ...Please extend your support for this role. Local candidate will get 1st preference. Job Title: SRE Engineer Location: Houston, TX and Jersey City, NJ - 3 Days Onsite Role FTE role with Mphasis Client: Mphasis H1B transfer will work... 
    Work experience placement
    H1b
    Local area

    Trinity Technology Solutions

    Houston, TX
    5 days ago
  •  ...Advanced Technology Centers (ATCs) are the engine for reinvention in our clients’...  ...services on the Databricks Intelligence Platform. You develop executive relationships with...  ...with APIs, model monitoring, and prompt management.Operate MLOps / LLMOps pipelines with CI... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Houston, TX
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Manager, Site Reliability Engineering - Paylo Platform. Be the first to apply!