Lead, Site Reliability Engineer (Infrastructure operations)
Full-time
Mastercard
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary
Mastercard powers economies and empowers people across more than 200 countries and territories worldwide.
We are committed to building an inclusive, digital economy that benefits everyone, everywhere—by making transactions safe, simple, smart, and accessible. Through secure data, trusted networks, strong partnerships, and relentless innovation, we help individuals, financial institutions, governments, and businesses unlock their greatest potential. About the Role:
Mastercard’s Program aligned Site Reliability Engineering (SRE) teams are dedicated to delivering a seamless experience for our customers. We achieve this by maintaining every aspect of our Programs infrastructure and technology ecosystem to the highest standards, ensuring compliance with rigorous security requirements.
Within Mastercard, SRE focuses on the reliability and performance of core infrastructure, networks, and foundational services that power our applications. Our mission is to ensure these components operate with excellence, enabling applications to deliver an outstanding customer experience.
In this role, you will join our Payments Network SRE team and take ownership of continuously assessing and elevating the end to end service quality of our platform. You will leverage data to drive root cause analysis and deliver strategic insights to key stakeholders on resource utilization, capacity forecasting, and performance trends—ensuring the availability, scalability, and resilience of our network. Key Responsibilities: Lead continuous assessments of the application infrastructure supporting critical Mastercard applications, focusing on health, performance, monitoring and alerting, and capacity analysis. Collaborate with Product and Development teams to forecast growth requirements and ensure scalability and resiliency. Champion observability as a core principle for infrastructure services by assessing environments and technologies to uncover gaps in monitoring and alerting. Design and implement strategies to close these gaps, ensuring all infrastructure telemetry is integrated into a unified, single-pane-of-glass view. Build custom dashboards to investigate and perform root cause analysis on complex issues. Lead regular incident reviews with internal support teams to ensure root causes are identified. When patterns of failure or compatibility issues between software and infrastructure emerge, develop and implement strategies to remediate or mitigate risks. Leverage automation and AI technologies to enhance proactive issue detection, enable self-healing capabilities, reducing Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM). Develop testing and validation plans for new environment builds, disaster recovery exercises and post-maintenance activities to certify environment readiness before customer traffic is routed to it. Champion continuous learning, development, and knowledge sharing across networking and other infrastructure disciplines to strengthen multi-disciplinary SRE team capabilities. Lead training initiatives for team members and Product and Development on networking aspects of the platforms. Evaluate vendor hardware, firmware, and software upgrade roadmaps, and conduct proof-of-concept (POC) testing to identify potential risks and opportunities for improvement in upcoming releases. All about you: • 5–10 years of experience in an SRE or SRE related operations role, including 3+ years supporting e commerce, financial services, or large scale SaaS platforms.
• Excellent infrastructure troubleshooting and analytical problem solving skills.
• Strong hands on experience with observability and monitoring tools such as Splunk, Dynatrace, or equivalent, with a proven ability to triage and investigate complex issues.
• Familiarity with network telemetry tools such as SolarWinds and NetScout.
• Proficiency in packet level debugging, including capturing traffic with tools like tcpdump and analyzing packets using Wireshark.
• Broad understanding of end to end infrastructure supporting payment platforms—spanning platform services, networking, databases, and storage.
• Experience with automation and Infrastructure as Code tools such as Chef, Ansible, and Terraform, as well as structured data formats (JSON/YAML).
• Excellent communication skills with the ability to coordinate cross functional troubleshooting efforts and lead RCA processes to closure.
• Demonstrated ability to troubleshoot complex production issues, perform root cause analysis, and drive long term corrective actions.
• Experience partnering with development teams to shape architecture, define SLIs/SLOs, and embed reliability into services from design through operation.
• Strong understanding of monitoring and observability ecosystems, including Prometheus, Grafana, ELK/EFK, Splunk, Dyantrace, and OpenTelemetry.
• Effective incident management skills with a structured, analytical approach to problem solving. The Payments Network SRE team is responsible for the runtime availability of some of Mastercard’s most critical core payment systems, which support national infrastructure and operate 24/7 year‑round. As a result, this role will include periodic on‑call responsibilities when required. Corporate Security Responsibility All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must: • Abide by Mastercard’s security policies and practices;
• Ensure the confidentiality and integrity of the information being accessed;
• Report any suspected information security violation or breach, and
• Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines. Corporate Security Responsibility
Lead, Site Reliability Engineer (Infrastructure operations)
Lead SRE Engineer, Site Reliability Engineering
Our Purpose:Mastercard powers economies and empowers people across more than 200 countries and territories worldwide.
We are committed to building an inclusive, digital economy that benefits everyone, everywhere—by making transactions safe, simple, smart, and accessible. Through secure data, trusted networks, strong partnerships, and relentless innovation, we help individuals, financial institutions, governments, and businesses unlock their greatest potential. About the Role:
Mastercard’s Program aligned Site Reliability Engineering (SRE) teams are dedicated to delivering a seamless experience for our customers. We achieve this by maintaining every aspect of our Programs infrastructure and technology ecosystem to the highest standards, ensuring compliance with rigorous security requirements.
Within Mastercard, SRE focuses on the reliability and performance of core infrastructure, networks, and foundational services that power our applications. Our mission is to ensure these components operate with excellence, enabling applications to deliver an outstanding customer experience.
In this role, you will join our Payments Network SRE team and take ownership of continuously assessing and elevating the end to end service quality of our platform. You will leverage data to drive root cause analysis and deliver strategic insights to key stakeholders on resource utilization, capacity forecasting, and performance trends—ensuring the availability, scalability, and resilience of our network. Key Responsibilities: Lead continuous assessments of the application infrastructure supporting critical Mastercard applications, focusing on health, performance, monitoring and alerting, and capacity analysis. Collaborate with Product and Development teams to forecast growth requirements and ensure scalability and resiliency. Champion observability as a core principle for infrastructure services by assessing environments and technologies to uncover gaps in monitoring and alerting. Design and implement strategies to close these gaps, ensuring all infrastructure telemetry is integrated into a unified, single-pane-of-glass view. Build custom dashboards to investigate and perform root cause analysis on complex issues. Lead regular incident reviews with internal support teams to ensure root causes are identified. When patterns of failure or compatibility issues between software and infrastructure emerge, develop and implement strategies to remediate or mitigate risks. Leverage automation and AI technologies to enhance proactive issue detection, enable self-healing capabilities, reducing Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM). Develop testing and validation plans for new environment builds, disaster recovery exercises and post-maintenance activities to certify environment readiness before customer traffic is routed to it. Champion continuous learning, development, and knowledge sharing across networking and other infrastructure disciplines to strengthen multi-disciplinary SRE team capabilities. Lead training initiatives for team members and Product and Development on networking aspects of the platforms. Evaluate vendor hardware, firmware, and software upgrade roadmaps, and conduct proof-of-concept (POC) testing to identify potential risks and opportunities for improvement in upcoming releases. All about you: • 5–10 years of experience in an SRE or SRE related operations role, including 3+ years supporting e commerce, financial services, or large scale SaaS platforms.
• Excellent infrastructure troubleshooting and analytical problem solving skills.
• Strong hands on experience with observability and monitoring tools such as Splunk, Dynatrace, or equivalent, with a proven ability to triage and investigate complex issues.
• Familiarity with network telemetry tools such as SolarWinds and NetScout.
• Proficiency in packet level debugging, including capturing traffic with tools like tcpdump and analyzing packets using Wireshark.
• Broad understanding of end to end infrastructure supporting payment platforms—spanning platform services, networking, databases, and storage.
• Experience with automation and Infrastructure as Code tools such as Chef, Ansible, and Terraform, as well as structured data formats (JSON/YAML).
• Excellent communication skills with the ability to coordinate cross functional troubleshooting efforts and lead RCA processes to closure.
• Demonstrated ability to troubleshoot complex production issues, perform root cause analysis, and drive long term corrective actions.
• Experience partnering with development teams to shape architecture, define SLIs/SLOs, and embed reliability into services from design through operation.
• Strong understanding of monitoring and observability ecosystems, including Prometheus, Grafana, ELK/EFK, Splunk, Dyantrace, and OpenTelemetry.
• Effective incident management skills with a structured, analytical approach to problem solving. The Payments Network SRE team is responsible for the runtime availability of some of Mastercard’s most critical core payment systems, which support national infrastructure and operate 24/7 year‑round. As a result, this role will include periodic on‑call responsibilities when required. Corporate Security Responsibility All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must: • Abide by Mastercard’s security policies and practices;
• Ensure the confidentiality and integrity of the information being accessed;
• Report any suspected information security violation or breach, and
• Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines. Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard’s security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.
Vacancy posted 8 days ago
Similar jobs that could be interesting for youBased on the Lead, Site Reliability Engineer (Infrastructure operations) in Ireland vacancy
- ...realize their greatest potential. Title and Summary Site Reliability Engineer II RiskPS BizOps team is looking for a Site Reliability... ...and look to automate everything you can? Business Operations is leading the Site Reliability Engineering (SRE) transformation at...OperationsFull timeWork experience placementWorldwideShift work
- ...greatest potential. Title and Summary Lead Site Reliability Engineer Who is Mastercard? At... ...About the Role The Business Operations team is seeking a highly motivated... ...reliability. • Cloud Computing and Infrastructure - Ability to design, deploy, and manage...OperationsFull timeWorldwide
- ...largest multidisciplinary consultancy, engineering and operations firms in the country. Today we have... ...some of Ireland’s most significant infrastructure programmes. From operating the... ...include but are not limited to: Leading the coordination effort with the design...OperationsFull time
- ...ABOUT GREYSTAR Greystar is a leading, fully integrated global real estate platform offering... ..., South Carolina, Greystar manages and operates over $350 billion of real estate in more... ...may be eligible to participate in on-site bonus programs. JOB DESCRIPTION Key...OperationsFor contractorsWork at officeLocal areaImmediate startFlexible hours
- ...Title and Summary Senior Software Engineer Overview ~ Mastercard... ...that accelerate business value and operational efficiency. Role Lead the planning, design, and delivery... ...continuously improve platform performance, reliability, and customer experience. All...OperationsFull timeWorldwide
- ...vertically integrated AI infrastructure company built from... ...up, we own and operate each layer of the stack... ...Crusoe Cloud Network Engineering team is actively looking... ...and improve network reliability. Manage, optimize... ...meets business needs Lead operational...OperationsHourly payFull timeWork experience placementImmediate start
- ...only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from... ...from the ground up. As a Software Engineer on the Cloud Availability... ...thousands of accelerators across sites into one cohesive, logical system...OperationsHourly payFull time
- ...greatest potential. Title and Summary Lead Software Engineer - Commercial Connect API Team... ...compliance with enterprise security, operations, and architecture standards. • Optimize application performance and reliability for large-scale, high-traffic systems....OperationsFull timeWorldwide3 days per week
- ...potential. Title and Summary Platform Engineer II Platform Engineer II, Generative... ...Platforms Overview The Enterprise Operations team is seeking a Platform Engineer II... ...engineering best practices for security, reliability, testing, and operational readiness. 12...OperationsFull timeWorldwide
- ...intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to... ...low-cost GPU compute power. As a Cloud Support Engineer, you'll play a crucial role in empowering our customers...OperationsHourly payFull time
- ...Summary Manager, Software Engineering Overview The InControl... ...Manager, Software Engineering to lead a team responsible for... ...development, delivery, and operation of business-critical applications... ..., Architecture, Security, Reliability Engineering, and Operations...OperationsFull timeWorldwide
- ...Virtu is an industry-leading financial technology firm that operates both proprietary trading and client-facing businesses... ..., fully stocked kitchen, on-site gym & workout classes, games room,... ...talented group of versatile software engineers. These engineers are responsible...Full timeSummer workInternship
- ...potential. Title and Summary Senior Software Engineer Mastercard Foundry R&D is seeking a... ...of trade-offs (cost, latency, reliability, capability) Build scalable, production... ..., not a side interest Comfortable operating in R&D and early-stage environments...Full timeWorldwide
- ...Title and Summary Senior Software Engineer What is the Services... ...largest organizations. You will: Lead team level system design, including... ...requirements such as scalability, reliability, performance, security, and operability. Partner with lead engineers and...Full timeWorldwide
- ...realize their greatest potential. Title and Summary Senior Software Engineer Overview We are the global technology company behind the... ...looking for an innovative software engineering manager who will lead the team responsible for the design and build of a full stack...Full timeWorldwideFlexible hours
- ...governments realize their greatest potential. Title and Summary Software Engineer II Job Description Summary Overview: Mastercard software engineering teams are obsessed with security, reliability, and performance and driven to deliver solutions that delight our...Full timeWorldwideFlexible hours
- ...technical leadership and operational support for a large... ...the organization's DB2 infrastructure. You will work... ...Ensure high availability, reliability, and performance of DB... ...maintenance activities. Lead technical... ...DATA offices or client sites. This ensures we can provide...OperationsWork at officeRemote workFlexible hours
- ...QCP is Asia's leading digital asset partner, empowering clients to seamlessly integrate digital assets into their portfolios. We offer... ..., and feeds to fulfil business requirements from the trading, operations and risk teams of the company ● To be responsible for...OperationsRemote jobFull timeCasual workFlexible hours
- ...greatest potential. Title and Summary Lead Software Engineer - Frontend (UI) Overview Who We... ...company in the payments industry, operating in over 210+ countries and territories... ...architecture reviews, design standards, reliability practices, and secure coding patterns...Full timeWorldwide
- ...Summary Senior Software Engineer Senior Software... ..., testing, and operating highly scalable, resilient... ..., capacity planning, reliability, and continuous improvement... ..., network, and infrastructure performance characteristics... ..., scalability, or site reliability. •...OperationsFull timeWorldwide
- ...greatest potential. Title and Summary Agentic AI - Senior Software Engineer Overview Mastercard Foundry R&D is seeking a Senior... ...models, frameworks, and tools across quality, cost, latency, reliability, and maintainability • Contribute to new product and...Full timeWorldwide
- ...greatest potential. Title and Summary Lead Software Engineer Position Overview The Business... ..., performance, security, and reliability—ensuring our solutions meet the highest... ..., data scientists, QA, security, and operations teams to translate business and functional...OperationsFull timeImmediate startWorldwide
- ...Summary Manager, Software Engineering Overview Authentication... ...Token Provisioning use cases, operating at Mastercard scale with... ...Software Engineering, you will lead a team of engineers responsible... ...through strong ownership of reliability, performance, availability...OperationsFull timeWorldwide
- ...has been one of the area's leading healthcare providers since 1... ...leaders to ensure that clinical operations are efficient and effective,... ...staff. Builds a strong infrastructure with designated charge nurses... ...Coast and Hawaii with over 440 sites of care, including 27 acute...OperationsWork experience placementShift work
- ...organization, apply now. Banking Practice Lead -AI Location: Dublin, Ireland... ...one of the world's leading AI and digital infrastructure providers, with unmatched capabilities... ...hire locally to NTT DATA offices or client sites. This ensures we can provide timely and...Work at officeRemote workFlexible hours
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview The InControl platform is a fundamental part of Mastercard's Commercial Solutions, enabling some of our largest...Full timeWork experience placementWorldwide
- ...Title and Summary Software Engineer II L8 ENGINEER REQ Wednesday... ..., Architects, Technical Leads, and fellow Engineers to define... ..., testing, deployment, and operational excellence. ○ Troubleshoot... ...Mindset □ Passion for building reliable, maintainable, and scalable...Full timeWorldwide
- ...realize their greatest potential. Title and Summary Software Engineer I Overview: Mastercard works to connect and power a... ...Mastercard software engineering teams are obsessed with security, reliability, and performance and driven to deliver solutions that delight...Full timeWorldwideFlexible hours
- ...and Summary Senior Software Engineer Senior Software Engineer... ...single secure virtual card infrastructure. Overview The Virtual... ..., performance, and operational excellence during platform... ...to ensure software quality, reliability, and security. • Collaborate...Full timeWorldwide
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Software Engineer Overview • The Decision Management program enables intelligent decision based products through streaming analytics with the...Full timeWork at officeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead, Site Reliability Engineer (Infrastructure operations). Be the first to apply!
Related searches
- on-site clinical research associate (traveling/remote) Ireland
- business operations intern Ireland
- operations tech Ireland
- vice president of field operations Ireland
- senior vice president of operations Ireland
- senior operations technician Ireland
- travel operations Ireland
- mainframe lead developer
- lead telecom engineer
- lead operating engineer


