Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$146.03k - $162.26k

CardWorks

Become an everyday champion — and build a career where your impact fuels financial progress.


What We Do

CardWorks Financial Group is a diversified financial services platform building ethical solutions across credit, lending, and the full customer lifecycle. Through our family of companies, CardWorks Financial Group tackles the complex challenges that larger financial institutions leave behind. We’re embedded throughout the credit card ecosystem as a lender, servicer, and merchant acquirer.

Who We Are

  • Merrick Bank: The bank that builds
  • CardWorks Servicing: One partner, total performance
  • Carson Smithfield: Resolution with respect

With nearly 40 years of operating history, our track record is solid: disciplined in downturns and built to accelerate in recovery. The CardWorks Financial Group companies take precise approach in complex markets, as a top three non-prime focused general purpose card issuer and a top fifteen U.S. merchant acquirer. 

Our team tackles the industry’s most complex credit and payment challenges. And we believe that excellent work starts with a team that feels supported, respected, and empowered to grow.

CardWorks Servicing, LLC provides end-to end operational servicing functions for credit cards, secured cards, and installment loans.  We service consumer and small business loans across the credit spectrum and offers backup servicing and due diligence services to capital providers and trustees.


Founded in 1997, Merrick Bank is an FDIC®-insured financial institution headquartered in South Jordan, Utah, with over $10 billion in assets. A wholly owned subsidiary of CardWorks Financial Group, Merrick Bank serves roughly five million cardmembers and more than 100,000 merchant customers, offering credit cards, recreational loans, deposit accounts, merchant services and bank sponsorships to consumers and businesses.

Carson Smithfield, LLC provides a variety of post-charge-off debt recovery services, including digital self-service, IVR, live agent, and external agency management.

Essential Functions:

  • Establish the SRE operating model (service onboarding, engagement model, governance, reliability reviews, production readiness standards, and quarterly planning) and ensure it is adopted across teams.

  • Identify, pilot, and operationalize AI-enabled reliability use cases (e.g., alert noise reduction, incident summarization, correlation/root-cause hypothesis generation, runbook assistance, and auto-remediation with human approval) with appropriate guardrails.

  • Define, implement, and operationalize reliability metrics by establishing and managing SLIs, SLOs, and error budgets to quantify and continuously improve service reliability, supporting engineering and business decisions.

  • Own the centralized SRE service engagement model by defining service tiers, onboarding criteria, reliability standards, and a transparent intake/prioritization process aligned to business criticality.

  • Define and enforce error budget policies (including escalation paths and release risk decisions) in partnership with Product/Engineering, using SLO attainment to guide trade-offs between feature velocity and reliability

  • Establish and maintain centralized “paved road” reliability standards and assets (instrumentation conventions, golden signals, alerting standards, runbook templates, SLO dashboards) that product teams can adopt with minimal friction.

  • Design the on-call and escalation model for a centralized SRE team (e.g., SRE overlay for major incidents, defined handoffs with service owners, and clear ownership boundaries) to improve response quality without creating single-team dependency.

  • Design and engineer automation and observability solutions by developing tooling, dashboards, and systems to reduce operational toil (measure, report, and drive toil down over time), enhance system visibility, and accelerate delivery.

  • Participates in incident and problem management by serving as incident coordinator for high-severity events, driving cross-functional responses, conducting blameless root cause analysis, running post-incident reviews (postmortems) with clear owners and due dates, ensuring remedial actions drive reliability improvements.

  • Oversee operational readiness and performance by managing capacity planning, validating disaster recovery, conducting production readiness reviews, and ensuring systems meet availability, scalability, and recovery expectations.

  • Partner with security, risk, and compliance teams to align reliability goals with governance and compliance requirements, ensuring secure, auditable, and well-documented practices.

  • Collaborate across the organization by working closely with end users, product management, development, architecture, and IT Operational teams to embed reliability principles throughout the software development lifecycle, including service onboarding, reliability reviews, and shared SLO ownership.

  • Champion reliability as a core product feature by promoting reliability throughout all phases of development, advocating for continuous improvement, and communicating key metrics and potential customer impact to stakeholders.

  • Train, mentor, and upskill engineering teams by coaching engineers in SRE practices, supporting junior team members, and fostering a culture of shared ownership and accountability for reliability, including influencing teams without direct authority through standards, data, and executive-aligned priorities. Remain current on the latest SRE trends and best practices, including observability, AI-enabled operations (AIOps), and SLO management, and implement these methodologies to effectively support desired business outcomes. Evaluate AI tools for reliability with security/privacy/compliance guardrails (e.g., data handling, prompt/content controls, auditability) and measure impact.

  • Participate in on-call rotations and operational support for SRE-supported systems and products.

Summary of Qualifications:

  • Experience in Site Reliability Engineering with a track record of delivering measurable improvements in uptime, scalability, release stability, and overall reliability in complex enterprise environments.

  • Demonstrated experience standing up or significantly maturing an SRE practice (operating model, SRE/service engagement, production readiness, incident/postmortem program, and reliability roadmap).

  • Hands-on experience applying AI/ML to operations (AIOps) or GenAI in production support workflows, with a focus on measurable outcomes (MTTD/MTTR, alert fatigue reduction, change failure rate) and responsible use controls.

  • Proven ability to establish Service Level Indicators (SLIs) and SLOs in production environments, including hands-on definition and implementation.

  • Demonstrated background in production incident response, leading resolution efforts, conducting blameless post-incident reviews, and implementing actionable remediation strategies.

  • Strong observability and telemetry expertise in designing instrumentation, building actionable dashboards and alerts, and delivering proactive reliability insights using metrics, logs, and traces.

  • Infrastructure engineering experience with strong Infrastructure as Code skills using tools such as Terraform and Ansible.

  • Thorough understanding and practical experience in CI/CD pipeline design, optimization, and troubleshooting using modern tooling and platforms such as Azure DevOps, GitHub Actions, Jenkins, or GitLab CI, with an emphasis on speed, reliability, and security.

  • Practical knowledge of containerization and platform modernization, including architecting and operating containerized workloads with Docker, VMware, and Kubernetes (or comparable orchestration platforms) to modernize legacy applications and improve fault tolerance.

  • Knowledge of emerging reliability practices, including SLO automation platforms, AIOps, or predictive operations to advance proactive reliability management.

  • Preferred certifications include AWS Professional, Terraform, Ansible, Azure DevOps, Octopus Deploy or other automation-focused credentials that demonstrate continuous technical development.

Education and Experience:

  • Master’s degree in computer science, Engineering, or equivalent practical experience designing and operating production systems at scale.

  • 7+ years of experience in Site Reliability Engineering.

Ideally, the qualified candidate will work at the following location(s): Woodbury, NY; Pittsburgh, PA, Orlando, Fl, South Jordan, UT. A hybrid work model or fully remote model can be considered based on hiring manager decision and priorities of the role.

The salary range for this position, if located in NY Metro/NY State is $146,032 to $162,257. However, please note that the salary range will vary for other geographic areas.

#INDHP

Our Employee Value Proposition

  • Competitive Pay, including a Bonus Target or Variable Pay Incentive Program 
  • Benefits Package -Medical, Dental, and Vision (plus much more) 
  • 401(k) Plan with Company Match 
  • Short- & Long-Term Disability 
  • Wellness Programs 
  • Group Life and AD&D Insurance 
  • Paid Vacation, Sick Days and bank Holidays 
  • Employee Engagement Activities including Employee Appreciation Day, DEI Employee Resource Groups, Corporate Social Responsibility, Service Recognition

We offer a total rewards package comprised of a competitive base rate of pay, variable pay incentive programs based on the role, and a comprehensive benefit suite.  Offered rates of pay are determined based on job-related knowledge, relevant experience, skills, certifications, and geographic location.

We are proud to be an equal opportunity employer. All qualified applicants will receive consideration without regard to age, race, color, sex, or gender identity/expression (including pregnancy, childbirth, transgender status, or sexual orientation), religion or creed, ancestry, citizenship, national origin, disability, military or veteran status, marital status, genetic information, or any other characteristic protected by applicable law.

We do not tolerate discrimination, harassment, or retaliation. Employment decisions are based solely on qualifications, merit, and business needs. Everyone is welcome here, and we hire based on your ability to do the job, not any protected characteristics.

If you need help or reasonable accommodation during the application or hiring process, please let your TA Partner know.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in United States vacancy
  • $90k - $130k

     ...Credence has an immediate opening for a Site Reliability SME who has hands-on experience working as a Cloud Operations Engineer with experience in IT operations to join our...  ...like the Operations Manager/TOPM, technical leads, and environmental engineers to prioritize,... 
    Suggested
    Temporary work
    Work experience placement
    Immediate start
    Worldwide

    Credence

    McLean, VA
    4 days ago
  •  ...build a successful career with opportunities to learn, grow, and make an impact. Join us! Position Summary: The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the... 
    Suggested
    Work at office
    Flexible hours
    Shift work
    Day shift

    Bank of America Corporation

    Charlotte, NC
    23 days ago
  •  ...Role Overview As a ServiceNow Lead Applied AI Site Reliability Engineer II , you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity of high-visibility products and platforms and the environments... 
    Suggested
    Work at office
    Local area
    Flexible hours
    3 days per week

    Deloitte LLP

    Tennessee
    2 days ago
  • $153k - $210k

     ...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient...  ...to resolution with very infrequent after-hours support. Lead blameless postmortems and implement long-term improvements... 
    Suggested
    Full time

    Ridgeline

    New York, NY
    8 hours ago
  •  ...EIT) organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this...  ...SLOs and error budgets. • Modernization & Migration: Lead the technical execution of re-architecting and redeploying... 
    Suggested
    Permanent employment
    Full time
    H1b
    Local area
    Remote work
    Shift work

    Jack Henry & Associates

    New York, NY
    1 hour ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development...  ...deployment processes, and drive automation initiatives. Lead incident resolution: You will conduct thorough root cause... 
    Full time

    Zoox

    Remote
    8 hours ago
  •  ...only provider of enterprise-scale context engines capable of analyzing trillions of real-...  ...seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team....  ...issues before they impact end-users. Lead troubleshooting efforts for complex production... 
    Full time

    Lovelace Ai

    Pittsburgh, PA
    8 hours ago
  • $99k - $225k

    Site Reliability Engineer, LeadThe Opportunity:  As a Lead Site Reliability Engineer (SRE) on our team, you’ll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Chantilly, Loudoun County, VA
    4 days ago
  • $76k - $127k

     ...realize their greatest potential. Title and Summary Site Reliability Engineer II The BizOps team at Mastercard is looking for a Site...  ...complex problems, building robust CI/CD pipelines, and leading the organization in DevOps automation and best practices.... 
    Full time
    Part time
    Worldwide
    Flexible hours
    Early shift

    Mastercard

    O Fallon, MO
    2 days ago
  • $96k - $163k

     ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer, Performance Engineering Senior Site Reliability Engineer, Performance Engineering Payment Optimization unifies... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    2 days ago
  • $100.1k - $180.2k

     ...'s top brands, offering comprehensive engineering, supply chain, and manufacturing solutions...  ...and a vast network of over 100 sites worldwide, Jabil combines global reach...  ...communities around the globe.Jabil is seeking a Lead Site Reliability Infrastructure and Security Engineer... 
    Temporary work
    Work at office
    Local area
    Remote work
    Worldwide

    Jabil Circuit

    Austin, TX
    4 days ago
  • $96k - $163k

     ...their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work...  ...design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    2 days ago
  • $96k - $163k

     ...their greatest potential. Title and Summary Senior Site Reliability Engineer Overview The BizOps team is looking for a Senior Site...  ...through validation and operational gating, and lead Mastercard in DevOps automation and best practices. • Practice... 
    Full time
    Part time
    Worldwide
    Flexible hours
    Shift work

    Mastercard

    O Fallon, MO
    2 days ago
  • $76k - $127k

     ...realize their greatest potential. Title and Summary Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we...  ...design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big picture... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Mastercard

    O Fallon, MO
    2 days ago
  •  ...Job Title:  Site Reliability Engineer (Azure Government & Infrastructure) Pay Type : SALARIED EXEMPT  Location:  Remote Citizenship Requirement: U.S. Citizen (Required) Summary of Position Role/Responsibilities The Site Reliability Engineer (SRE) for... 
    Full time
    Remote work
    Monday to Friday

    Quzara LLC

    United States
    4 days ago
  •  ...Senior Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: AWS,...  ...Requirements: Senior SRE with 5+ years AWS infra, 3+ years in senior/lead roles; strong automation with Terraform, Ansible,... 
    Full time
    Remote work

    CyberArk

    United States
    4 days ago
  • $7.5k

     ...manager, and we have ambitious goals for the future. As a Site Reliability Engineer (SRE), you will work at the intersection of production...  ...and trading systems Diagnose and fix bugs in code Lead complex deployments Automate manual workflows... 
    Local area
    Remote work

    The Voleon Group

    United States
    4 days ago
  •  ...Site Reliability Engineer OXIO is the first NeoTelco. We arebuilding the world’s largest, most accessible, and insightful Telecom network. Our platform empowers anyone to spin up their own carrier from a browser, scaling and supporting you as you scale your network... 
    Remote work

    OXIO

    United States
    4 days ago
  •  ...Site Reliability Engineer Company: GitLab Work Type: Remote Employment: Full Time Location: CA, US Seniority: Senior Level Technologies: Terraform, Ansible, Kubernetes, Go, Ruby, Jsonnet, Prometheus, ELK, Grafana Requirements: Senior-level SRE with strong Terraform/IaC... 
    Full time
    Remote work

    GitLab

    United States
    4 days ago
  •  ...customers rely on us in the moments that matter. Engineering delivers on that promise.   The Senior Site Reliability Engineer is responsible for ensuring our SaaS...  ...•    Participate in on-call duties 365/24/7 and lead the triage and RCA of production incidents... 
    Work experience placement
    Remote work
    Flexible hours

    Donnelley Financial Solutions

    United States
    2 days ago
  •  ...encourage you to apply. The Role  As a Senior Platform Engineer, you are a champion for DevOps and SRE culture and industry...  ...met. \n What You Will Be Doing Improving production reliability and system resilience within an SRE scoped team Championing... 
    Remote work
    Flexible hours

    Megaport

    United States
    2 days ago
  • $141.8k - $195k

     ...We’re one of the fastest‑growing private companies and a leading player in a massive, fast‑moving market. With a global workforce...  ...You’ll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all... 
    Temporary work
    Remote work

    Cribl

    United States
    4 days ago
  • $152k - $195k

     ...and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and...  ...observability — define SLOs, alerts, and dashboards. Lead incident response and postmortems, focusing on root cause and... 
    Remote work

    SecurityScorecard

    United States
    5 days ago
  •  ...organizations maintain accurate, compliant, and reliable provider networks at scale. Our...  ...provider data. We're backed by leading investors and built by a team with...  ...Role We're looking for a Senior Site Reliability Engineer who takes ownership seriously -... 
    Remote work

    CertifyOS

    United States
    4 days ago
  •  ...have come to expect, and help raise the reliability bar as we grow. What you would do:...  ...operate the shared platform foundations engineers ship on every day: GCP infrastructure, Kubernetes...  ...technologies. There are many roads leading up to being an SRE. Our team is already... 
    Remote work
    Worldwide
    Flexible hours

    Sanity

    United States
    4 days ago
  • $147k - $168k

     ...Inc. as one of the most innovative and fastest-growing technology companies in the country. Role Summary As a Site Reliability Engineer at Filevine, you will improve the reliability, scalability, and operational maturity of the Filevine platform. You’ll... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    Remote work
    2 days per week
    3 days per week

    Filevine

    United States
    4 days ago
  •  ...Site Reliability Engineer Company: Milestone Systems Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Golang, Python, Linux, Shell scripting, Kubernetes, Docker, Terraform, CI/CD, GitOps, ArgoCD, Spinnaker, Prometheus, Datadog,... 
    Full time
    Remote work

    Milestone Systems Inc

    United States
    4 days ago
  •  ...Cloud Operations Engineer III Function: Engineering  Reports...  ...team, responsible for the reliability, observability, performance,...  ...monitors, and alert routing — and leads the migration off our legacy...  ...in cloud operations, site reliability, platform, or DevOps... 
    Temporary work
    Casual work
    Local area
    Remote work
    Worldwide
    Flexible hours

    On-Board Services

    United States
    4 days ago
  • $160k - $208k

     ...conditions. We are looking for an experienced Site Reliability and Infrastructure Engineer to join our engineering team. You will support...  ...issues as they arise. You will collaborate with technical leads across engineering disciplines, as well as data scientists... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Clover Health

    United States
    4 days ago
  • $182.8k - $247.3k

     ...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems... 
    Work experience placement

    Socket

    Eastern, KY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!