Manager, Site Reliability Engineer
Mastercard
Manager, Site Reliability Engineer
Who is Mastercard?
At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits everyone, everywhere, by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team – one that makes better decisions, drives innovation, and delivers better business results.
What we create today will define tomorrow. Revolutionary technologies that reshape the digital economy to be more connected and inclusive than ever before. Safer, faster, more sustainable.
And we need the best people to do it. Technologists who are energized by the challenges of a truly global network. With the talent and vision to create the critical systems and products that power global commerce and connect people everywhere to the vital goods and services they need every day.
Working at Mastercard means being part of a unique culture. Inclusive and diverse, a rich collaboration of ideas and perspectives. A place that celebrates your strengths, values your experiences, and offers you the flexibility to shape a career across disciplines and continents. And the opportunity to work alongside experts and leaders at every level of the business, improving what exists, and inventing what’s next. About the Role
The Business Operations team is seeking a highly motivated and experienced Manager, Site Reliability Engineer (SRE) to join our team. You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power Mastercard's global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor. The role of the Business Operations Site Reliability Engineer is to be the production readiness steward for Mastercard products. As Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to running our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principles that include operational design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture. We support daily operations with a hyper focus on triage, root cause by understanding the business impact of our products and subsequently performing blameless post-mortems. The goal of every Business Operations team is to engage early in the development lifecycle to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications. Business Operations teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. Ultimately, the role of Business Operations is to align Product and Customer Focused priorities with Operational needs by providing continuous feedback throughout the lifecycle. As part of the Business Operations team, you will:
• Oversee a team of individual contributors, supporting the execution of strategic initiatives by providing technical expertise and leadership within the Site Reliability Engineering discipline to analyze complex problems and provide novel solutions and/or improvements.
• Guide the team in automating routine tasks, troubleshooting complex issues, and optimizing system performance.
• Collaborate with cross-functional teams to develop strategies for system scalability and resilience, training team members on technical skills, operational best practices, and incident management.
• Oversee incident response efforts, ensuring timely resolution and comprehensive root cause analysis.
• Cultivate a culture of continuous improvement by promoting best practices, innovation, and proactive risk management.
• Support the implementation and maintenance of high-availability systems to ensure operational stability.
• Contribute to documentation, knowledge sharing, and best practices to improve team operational procedures.
• Lead automation and scripting efforts to streamline operational processes and incident response workflows.
• Manage a team of individual contributors(s) and/or technical lead(s), directing area processes and work to ensure that they align with functional best practices and organizational standards; conduct goal setting and performance appraisal processes to coach team members and support their professional development. Role qualifications:
The ideal candidate will apply leadership skills independently and consistently in complex or nuanced situations to support broader goals. Recognized as a key contributor and may coach or support others informally. As a leader, you will:
• Build diverse, high performing teams with a customer-focused mindset. Attract, grow, and develop exceptional, future-ready talent.
• Inspire teams to look beyond their function, connect their work to enterprise impact, think end to end, and act in the best interests of the whole company.
• Anticipate market shifts and use curiosity, innovation, and technology to turn insights into strategies that drive growth and competitive advantage.
• Lead through ambiguity across diverse markets and regulatory environments, connecting insights and stakeholders to create clarity with sound judgment and cross cultural awareness.
• Inspire and mobilize people and teams to act with speed, agility, and accountability in driving ambitious business outcomes, with a relentless focus on the customer.
• Explore new ideas, ways of working and technology. Set clear direction, aligns stakeholders, and remove barriers to progress. Guide teams through uncertainty with clarity, empathy, and resilience.
As this is a player/coach role, the ideal candidate will also apply the following skills independently and consistently in complex or nuanced situations, begin using the skills to support broader goals, and be recognized as a key contributor who may coach or support others informally. • Observability - Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement.
• Programming and Scripting - Ability to write and maintain code and scripts to automate tasks, build operational tools, and support monitoring, deployment, and incident response using languages such as Python, Go, Bash, or similar.
• Systems and Network Administration - Ability to configure, operate, and troubleshoot Linux/Unix systems and network components, applying knowledge of networking concepts, protocols, security, and system reliability.
• Cloud Computing and Infrastructure - Ability to design, deploy, and manage applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP), ensuring scalability, security, availability, and operational efficiency.
• Reliability and Scalability - Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand
• DevOps Practices - Ability to apply DevOps principles and practices, including CI/CD pipelines, containerization, and orchestration, to enable faster, more reliable software delivery and operations.
• Troubleshooting - Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations.
• Capacity Planning and Performance Optimization - Ability to monitor resource utilization, forecast future capacity needs, and optimize system performance to support growth, scalability, and efficient infrastructure usage.
• IT Service Management - Ability to apply IT service management principles to incident, problem, and change management, ensuring reliable service delivery, effective incident response, and continuous service improvement aligned to business needs.
• Proactive Monitoring and Improvement (SRE Applications) - The ability to use application reliability signals to anticipate issues, identify risks, and drive preventative improvements that enhance application performance and availability. Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard’s security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.
- ...realize their greatest potential. Title and Summary Senior Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work... ...and upfront in the development process, and to proactively manage production and change activities to maximize customer...SuggestedFull timeWorldwide
- ...governments realize their greatest potential. Title and Summary Site Reliability Lead Engineer Lead Site Reliability Engineer Who is Mastercard?... ...upfront in the development process, and to proactively manage production and change activities to maximize customer...SuggestedFull timeWorldwide
- ...realize their greatest potential. Title and Summary Director, Site Reliability Engineering Director, Site Reliability Engineering Our Purpose:... ...will join our Payments Network SRE team, where you will manage a team of highly skilled SRE infrastructure engineers with...SuggestedFull timeWorldwide
- ...realize their greatest potential. Title and Summary Lead, Site Reliability Engineer (Infrastructure operations) Lead SRE Engineer, Site... ...Splunk, Dyantrace, and OpenTelemetry. • Effective incident management skills with a structured, analytical approach to problem solving...SuggestedFull timeWorldwide
- ...potential. Title and Summary Senior Platform Engineer Overview: As a Senior Platform... ...play a crucial role in ensuring the reliability, security, and efficiency of our z/OS communication... ...and related technologies. Role: • Manage and administer z/OS Communication Server...SuggestedFull timeWorldwideWeekend work
- ...their greatest potential. Title and Summary Senior Software Engineer Role Overview We are seeking a highly capable Senior Software... ..., ensuring high standards of code quality and system reliability Design and implement scalable, maintainable test automation...Full timeWorldwideFlexible hoursShift work
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Software Engineer II Who is Mastercard? Mastercard is a global technology company in the payments industry. Our mission is to connect and power...Full timeWorldwide
- ...potential. Title and Summary Software Engineer II The Business Experimentation and Optimization... ...skills while helping the team deliver reliable, high-quality software. Our teams are... ...customer base. We combine traditional management consulting with rich data assets and in-...Full timeImmediate startWorldwide
- ...their greatest potential. Title and Summary Senior Software Engineer Overview The Mastercard Fraud Scoring and Analytics Platform... ...new capabilities and frameworks on the MasterCard Fraud Management Platform. This platform provides sophisticated business solutions...Full timeWork experience placementWorldwide
- ...potential. Title and Summary Software Engineer – DevOps / SRE Overview The... ...Software Engineer II with emphasis on site reliability to support and evolve our Authentication... ...performance of the Mastercard Fraud Decision Management Platform (DMP), which provides...Full timeWorldwide
- At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences. Our dedication to remote...Full timeRemote workWorldwide
- ...that help people, businesses and governments realize their greatest potential. Title and Summary Software Engineer II Overview The Virtual Card Management (VCM) team is part of the Commercial Transaction Management & Controls (CTMC) program within Mastercard's...Full timeWorldwide
- ...potential. Title and Summary Senior Software Engineer Overview The Program Modernization... ...compliance portal for inventorying, managing, testing, and evidencing technology... ...standards for security, performance, and reliability. Role & Responsibilities Design, develop...Full timeWorldwide
- ...potential. Title and Summary Software Engineer II Senior Software Engineer Overview... ...emerging technologies into secure, reliable, and reusable capabilities that create measurable... ..., tool use, prompt and context management, memory patterns, evaluation, and AI system...Full timeWorldwide
- ...their greatest potential. Title and Summary Senior Software Engineer Overview Be part of the Operations & Technology Fraud Products... ...team developing new capabilities for MasterCard's Decision Management Platform, which serves as the core for multiple business...Full timeWorldwide
- ...the team as Twilio’s next Senior Software Engineer. About the job This position will... ...core business as it is responsible for managing the billing lifecycle and payment experience... ...design (architecture, design patterns, reliability and scaling) of new and current systems....Full timeWork experience placementRemote workWorldwide
- ...potential. Title and Summary Senior Software Engineer Overview Mastercard is a global... ...and system requirements into scalable, reliable, and maintainable solutions Define... ...vulnerabilities Troubleshooting & Incident Management Serve as a senior escalation point...Full timeWorldwide
- ...potential. Title and Summary Lead Software Engineer - Distributed Microservices Platform... ...You will be part of the Virtual Card Management domain within CTMC, working on... ...a strong focus on platform resilience, reliability, and safe evolution while continuing to...Full timeWorldwide
- ...their greatest potential. Title and Summary Lead, Platform Engineer Who is Mastercard? Mastercard is a global technology company... ...for internal processes in Server Builds, Operating System management, Patch Management, and/or Configuration Management on physical...Full timeWork experience placementWorldwide
- ...their greatest potential. Title and Summary Senior Platform Engineer - Linux Overview: Linux Systems Administrator, Platform... ...implementations, troubleshooting of systems • Use configuration management and infrastructure automation for repeatable tasks • Design...Full timeWorldwide
- ...greatest potential. Title and Summary Principal DevOps Engineer - Decision Management Platform Overview Join Mastercard’s Services... ...Terraform and CloudFormation. • Drive observability and reliability through monitoring, logging, and alerting systems (Prometheus...Full timeWorldwide
- About the role 🎯 Your goal will be to build and lead a world-class, AI-native Solutions Engineering organisation that enables n8n to win, deploy, and expand increasingly complex enterprise customers. To make that happen, you’ll set the strategy, strengthen our technical...Full time
- ...will be to build and lead a world-class, AI-native Solutions Engineering organisation that enables n8n to win, deploy, and expand... ...existing frontline leaders, helping them become stronger people managers, technical leaders, and strategic business partners. Define...Full timeTemporary workImmediate start
- About the Role: As a DevOps Engineer at Sardine, you'll play a critical role in evolving... ...developer productivity, and cloud cost management. You’ll work closely with application engineers... ...and Finance to ensure our systems are reliable, scalable, and cost-efficient. The role...Full time
€115k - €130k per year
...there. About the Role: As a DevOps Engineer at Sardine, you'll play a critical role... ...developer productivity, and cloud cost management. You’ll work closely with application... ...Security and Finance to ensure our systems are reliable, scalable, and cost-efficient. The role...Remote jobFull timeWorldwideHome officeFlexible hours- ...realize their greatest potential. Title and Summary Lead Software Engineer Overview Be part of the Operations & Technology Fraud... ...team developing new capabilities for MasterCard's Decision Management Platform, which serves as the core for multiple business solutions...Full timeWorldwide
- The Role We are looking for Android engineers to build the native Android experience for KIRA's AI Neobank App. This role is for someone... ...user flows, fast performance, product quality and shipping reliable mobile features used by real customers. What You'll Own Build...Full time
- ...The Role We are looking for Android engineers to build the native Android experience for KIRA's AI Neobank App. This role is for... ...clean user flows, fast performance, product quality and shipping reliable mobile features used by real customers. What You'll Own...Full timeRemote work
- ...potential. Title and Summary Lead Software Engineer Overview: Mastercard is seeking a... ...empowers businesses of all sizes to manage payments more efficiently when buying or... ...• Optimize application performance and reliability for large-scale, high-traffic systems....Full timeWorldwide3 days per week
- ...potential. Title and Summary Lead Software Engineer Lead Software Engineer Overview:... ...and delivery of secure, scalable, and reliable agentic applications that can reason,... ...native environments using Kubernetes and managed cloud services on AWS, Azure, or GCP •...Full timeTemporary workWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Manager, Site Reliability Engineer. Be the first to apply!




