Lead Site Reliability Engineer
Mastercard
Lead Site Reliability Engineer
Who is Mastercard?
At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits everyone, everywhere, by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team – one that makes better decisions, drives innovation, and delivers better business results.
The Business Operations team is seeking a highly motivated and experienced Lead Site Reliability Engineer (SRE) to join our team. You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power Mastercard's global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor. The role of the Business Operations Site Reliability Engineer is to be the production readiness steward for Mastercard products. As Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to running our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principles that include operational design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture. We support daily operations with a hyper focus on triage, root cause by understanding the business impact of our products and subsequently performing blameless post-mortems. The goal of every Business Operations team is to engage early in the development lifecycle to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications. Business Operations teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. Ultimately, the role of Business Operations is to align Product and Customer Focused priorities with Operational needs by providing continuous feedback throughout the lifecycle. As part of the Business Operations team, you will:
• Be a developing subject matter expert in the Site Reliability Engineering area, influencing stakeholders and applying advanced knowledge to drive achievement of area goals and initiatives by contributing to solution development and improvements for existing products, services, and/or processes.
• Implement and maintain high-availability system solutions, ensuring stability, performance, and operational continuity.
• Evaluate operational requirements to develop effective technical solutions within existing frameworks.
• Lead automation and scripting efforts to streamline operational processes and incident response workflows.
• Troubleshoot and resolve complex system issues, escalating as necessary to maintain system health and proactively address risks.
• Contribute to documentation, knowledge sharing, and best practices to improve team operational procedures.
• Conduct reviews and quality assurance activities to uphold organizational standards for system stability.
• Keep current with industry trends and emerging technologies relevant to system reliability and operational automation.
• Guide and mentor junior team members through on-the-job experiences, reviewing work and fostering a culture of continuous improvement to grow expertise around their discipline. Role qualifications:
The ideal candidate will apply the following skills independently and consistently in complex or nuanced situations, begin using the skills to support broader goals, and be recognized as a key contributor who may coach or support others informally. • Observability - Ability to use scripting and tooling to implement observability solutions, enabling the collection, analysis, and visualization of metrics, logs, and traces to support incident detection, diagnosis, and continuous service improvement.
• Programming and Scripting - Ability to write and maintain code and scripts to automate tasks, build operational tools, and support monitoring, deployment, and incident response using languages such as Python, Go, Bash, or similar.
• Systems and Network Administration - Ability to configure, operate, and troubleshoot Linux/Unix systems and network components, applying knowledge of networking concepts, protocols, security, and system reliability.
• Cloud Computing and Infrastructure - Ability to design, deploy, and manage applications and infrastructure on cloud platforms (e.g., AWS, Azure, GCP), ensuring scalability, security, availability, and operational efficiency.
• Reliability and Scalability - Ability to design and operate systems for high availability, fault tolerance, and disaster recovery, while ensuring systems can scale to meet current and future demand
• DevOps Practices - Ability to apply DevOps principles and practices, including CI/CD pipelines, containerization, and orchestration, to enable faster, more reliable software delivery and operations.
• Troubleshooting - Capability to systematically identify, diagnose, and resolve technical issues across systems, applications, and networks, using analytical methods and tools to restore functionality, minimize disruption, and ensure stable operations.
• Capacity Planning and Performance Optimization - Ability to monitor resource utilization, forecast future capacity needs, and optimize system performance to support growth, scalability, and efficient infrastructure usage.
• IT Service Management - Ability to apply IT service management principles to incident, problem, and change management, ensuring reliable service delivery, effective incident response, and continuous service improvement aligned to business needs.
• Proactive Monitoring and Improvement (SRE Applications) - The ability to use application reliability signals to anticipate issues, identify risks, and drive preventative improvements that enhance application performance and availability. Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard’s security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.
- Your main goal will be to build and scale integrated campaigns that generate measurable pipeline for our sales-led growth motion and help the sales team engage and nurture high-value accounts. To make that happen, you’ll own campaigns end to end and bring teams together...SuggestedRemote work
- ...ABOUT GREYSTAR Greystar is a leading, fully integrated global real estate platform offering... ...may be eligible to participate in on-site bonus programs. JOB DESCRIPTION... ...qualification in electrical/mechanical engineering or plumbing ( i.e. NVQ, City Guilds or...SuggestedContract workFor contractorsImmediate startFlexible hours
- ...the team as Twilio’s next Senior Software Engineer. About the job This position will... ...simultaneously delivering industry leading availability. We do this by leveraging technologies... ...design (architecture, design patterns, reliability and scaling) of new and current systems....SuggestedFull timeWork experience placementRemote workWorldwide
- ...looking for a collaborative Senior Software Engineer to join our Revenue Engineering team.... ...implementation partners to deliver reliable, scalable solutions. As a senior member of... ...architecture of Muck Rack's integration platform, lead complex technical initiatives, and ensure...SuggestedFull timeRemote work
- ...governments realize their greatest potential. Title and Summary Lead Software Engineer Position Overview The Business Experimentation and... ...—such as scalability, performance, security, and reliability—ensuring our solutions meet the highest engineering standards...SuggestedFull timeImmediate startWorldwide
- ...The Role We are looking for Android engineers to build the native Android experience for KIRA's AI Neobank App. This role is for... ...clean user flows, fast performance, product quality and shipping reliable mobile features used by real customers. What You'll Own...Full timeRemote work
- ...that help people, businesses and governments realize their greatest potential. Title and Summary Director, Software Engineering Job Summary: Leads major projects and delivers quality software solutions in a timely and cost effective manner. Researches and...Full timeWorldwide
- ...and forward-thinking organization, apply now. Banking Practice Lead -AI Location: Dublin, Ireland Experience: 15+ years... ...Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored...Work at officeRemote workFlexible hours
- ...ABOUT GREYSTAR Greystar is a leading, fully integrated global real estate platform offering expertise in property management, investment... ...the base salary, this role may be eligible to participate in on-site bonus programs. JOB DESCRIPTION Key Role...For contractorsWork at officeLocal areaImmediate startFlexible hours
- ...Expectations Ensure high availability, reliability, and performance of DB2 for z/OS... ...incidents and planned maintenance activities. Lead technical troubleshooting and root cause... ...hire locally to NTT DATA offices or client sites. This ensures we can provide timely and...Work at officeRemote workFlexible hours
- ...DATA is currently looking for a Lead MS Fabric Architect in Dublin... ...data architects, data engineers, platform specialists and consultants... ...opportunities, leadership and reliable delivery. Demonstrate and... ...to NTT DATA offices or client sites. This ensures we can provide timely...Work at officeRemote workFlexible hours
- ...As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development... ...work closely with Product, Design, and other engineers to ship reliable agentic experiences. What you’ll do here Design and build...Full timeRemote work
- ...potential. Title and Summary Senior Software Engineer Position Overview The Business... ...organization are building industry-leading software that empowers business users around... ...scalability, performance, security, and reliability, ensuring our solutions meet the highest...Full timeWorldwide
- ...potential. Title and Summary Senior Software Engineer Overview ~ Mastercard Emerging Data... ...operational efficiency. Role Lead the planning, design, and delivery of... ...improve platform performance, reliability, and customer experience. All About You...Full timeWorldwide
- ...About the role Your goal will be to build and lead a world-class, AI-native Solutions Engineering organisation that enables n8n to win, deploy, and expand increasingly complex enterprise customers. To make that happen, you’ll set the strategy, strengthen our technical...Full timeTemporary workImmediate start
- About the role 🎯 Your goal will be to build and lead a world-class, AI-native Solutions Engineering organisation that enables n8n to win, deploy, and expand increasingly complex enterprise customers. To make that happen, you’ll set the strategy, strengthen our technical...Full time
- ...society through responsible innovation. We are one of the world's leading AI and digital infrastructure providers, with unmatched... ...Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored...Work at officeRemote workFlexible hours
- The Role We are looking for Android engineers to build the native Android experience for KIRA's AI Neobank App. This role is for someone... ...user flows, fast performance, product quality and shipping reliable mobile features used by real customers. What You'll Own Build...Full time
- ...InfluxData is the creator of InfluxDB, the leading time series platform used to collect, store, and analyze all time series data at... ...design and implement with respect and kindness across multiple engineering teams. We embrace an empathetic, supportive, and communicative...Full timeRemote work
- ...architecture as models and capabilities change. Lead complex projects from early ambiguity through production. Raise the engineering bar through design reviews, code reviews,... ...useful agent experiences. Improve the reliability, performance, and operability of agentic...Full time
- About the Role: As a DevOps Engineer at Sardine, you'll play a critical role in evolving our infrastructure and platform tooling to support... ...like Security and Finance to ensure our systems are reliable, scalable, and cost-efficient. The role involves building and maintaining...Full time
- ...help shape it. We power card issuing for leading banks, fintechs and digital brands... ...scale. We're looking for a Data Platform Engineer to join our Data Engineering team and help... ..., playing a critical role in enabling reliable, high-performance, and secure data systems...Full timeLocal areaRemote work
- ...mobile developers to build BJAK's AI Neobank App across Android and iOS. This role is for strong mobile builders who can ship clean, reliable and intuitive mobile experiences, whether their main strength is Android, iOS or both. What You'll Own Build and ship mobile...Full time
- ...coderd . Improve database migration safety and upgrade reliability through schema compatibility, background migrations, and safe... ...What we're looking for ~5+ years of professional software engineering experience, including significant production experience with...Full time
- ...SoSafe has the ambition to become the leading human risk management provider in Europe. Our award-winning awareness platform triggers... ...make a difference: Develop, monitor and maintain scalable and reliable fullstack software systems using Typescript (React, Node) in...Full timeWork at officeLocal areaWorldwideFlexible hours
- ...responsibilities will include working closely with members of the backend team and Integrations team to develop features together. Engineers at ClickUp are also responsible for the quality of their own code, so you will work with a QA counterpart to ensure all edge cases...Full time
- As a Principal Cloud Security Engineer at LastPass, you will partner with DevOps and CI/CD engineers and our Architects team to ensure security best practices are embedded across our cloud infrastructure. We are looking for an experienced security engineering leader with...Full time
- ...the steward of how well our BI tool thinks. Partner with data engineering to optimise performance, and build what's missing Own and... ...Work with GTM stakeholders, and Primer’s AI Strategy & Enablement Lead to spot processes where in-house AI workflows, skills and...Full time
- .... We're looking for a senior full stack engineer to anchor this product, owning the realtime... ..., we are on a mission to become the leading travel platform globally, powering Hopper... ...integrating our fintech products on their sites or powering end-to-end travel portals. Today...Contract workWork from home
- ...Shopify Full Stack Developer with deep expertise in the Shopify ecosystem, a strong grasp of headless architecture, and experience leading enterprise-scale development initiatives. You'll be responsible for building scalable, high-performance eCommerce solutions using Shopify...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!












