Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Manager, Site Reliability Engineering

$121k - $169k
Full-time

Mastercard

Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary

Manager, Site Reliability Engineering

Who is Mastercard?
At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits everyone, everywhere, by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team – one that makes better decisions, drives innovation, and delivers better business results.

Technology at Mastercard
What we create today will define tomorrow. Revolutionary technologies that reshape the digital economy to be more connected and inclusive than ever before. Safer, faster, more sustainable and we need the best people to do it. Technologists who are energized by the challenges of a truly global network. With the talent and vision to create the critical systems and products that power global commerce and connect people everywhere to the vital goods and services they need every day.

About the Role
The Business Operations team is seeking a highly motivated and experienced Manager, Site Reliability Engineering (SRE) to join our team. You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power Mastercard's global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor.

The role of the Business Operations Site Reliability Engineer is to be the production readiness steward for Mastercard products. As Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to running our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principles that include operational design, automation, capacity planning, and monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture.

As part of the Business Operations team, you will:
• Oversee a team of individual contributors, supporting the execution of strategic initiatives by providing technical expertise and leadership within the Site Reliability Engineering discipline to analyze complex problems and provide novel solutions and/or improvements. ​
• Guide the team in automating routine tasks, troubleshooting complex issues, and optimizing system performance. ​
• Collaborate with cross-functional teams to develop strategies for system scalability and resilience, training team members on technical skills, operational best practices, and incident management. ​
• Oversee incident response efforts, ensuring timely resolution and comprehensive root cause analysis.
• Cultivate a culture of continuous improvement by promoting best practices, innovation, and proactive risk management. ​
• Support the implementation and maintenance of high-availability systems to ensure operational stability. ​
• Contribute to documentation, knowledge sharing, and best practices to improve team operational procedures. ​
• Lead automation and scripting efforts to streamline operational processes and incident response workflows. ​
• Manage a team of individual contributors(s) and/or technical lead(s), directing area processes and work to ensure that they align with functional best practices and organizational standards; conduct goal setting and performance appraisal processes to coach team members and support their professional development.

Role:
• Serve as the primary contact responsible for the overall application health, performance, and capacity
• Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
• Partner with the development and product team of a new application to establish the right monitoring and alerting strategy and create the framework to achieve zero downtime during deployment.
• Serve as the primary contact responsible for ensuring application scalability, performance, and resilience.
• Practice sustainable incident response and blameless post-mortems while taking a holistic approach to problem-solving and optimizing time to recover.
• Automate data-driven alerts to proactively escalate issues. Work with development teams to establish SLOs and improve reliability.
• Tackle complex development, automation, and business process problems. Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation, and refinement.
• Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead Mastercard in DevOps automation and best practices.
• Increase automation and tooling to reduce toil and manual interventiono Analyses ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns

All about you:
The ideal candidate will have experience with leading teams in many of these areas:
• BS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics), or equivalent practical experience.
• Coding or scripting exposure.
• Appetite for change and pushing the boundaries of what can be done with automation. Be curious about new technology, infrastructure, and practices to scale our architecture and prepare for future growth.
• Experience with algorithms, data structures, scripting, pipeline management, and software design• Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive.
• Interest in designing, analyzing, and troubleshooting large-scale distributed systems.
• Willingness and ability to learn and take on challenging opportunities and to work as a member of a matrix-based, diverse and geographically distributed project team.
• Ability to balance doing things right with fixing things quickly. Flexible and pragmatic, while working towards improving the long-term health of the system.
• Comfortable collaborating with cross-functional teams to ensure that expected system behaviour is understood and monitoring exists to detect anomalies.Experience with DevOps practices and tools such as Chef, Ansible, Artifactory, GitHub, Bitbucket, Jenkins, XLR, and Remedy

Mastercard is a merit-based, inclusive, equal opportunity employer that considers applicants without regard to gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law. We hire the most qualified candidate for the role. In the US or Canada, if you require accommodations or assistance to complete the online application process or during the recruitment process, please contact View email address on decentrajobs.com and identify the type of accommodation or assistance you are requesting. Do not include any medical or health information in this email. The Reasonable Accommodations team will respond to your email promptly.

Corporate Security Responsibility

All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:

  • Abide by Mastercard’s security policies and practices;

  • Ensure the confidentiality and integrity of the information being accessed;

  • Report any suspected information security violation or breach, and

  • Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.

In line with Mastercard’s total compensation philosophy and assuming that the job will be performed in Canada, the successful candidate will be offered a competitive pay based on location, experience and other qualifications for the role and may be eligible to participate in a discretionary annual incentive program.

Pay Ranges

Vancouver, Canada: $121,000 - $169,000 CAD

Vacancy posted 7 days ago
Similar jobs that could be interesting for youBased on the Manager, Site Reliability Engineering in Canada vacancy
  •  ...SUMMARY We're hiring a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on-premise infrastructure and CI/CD pipelines...  ...best practices , and high-availability standards . Manage and optimize Kubernetes cluster environments (Helm, ArgoCD... 
    Suggested
    Full time
    Local area

    Blackpoint Cyber

    Canada
    9 days ago
  •  ...digital experiences, and identity and access management. You will: Manage multiple cloud...  ...health—utilization, performance, and reliability—across our infrastructure Understand...  ...code review—that automates reliability engineering work: Deployment tooling Fault-... 
    Suggested
    Full time
    Casual work
    Local area
    Worldwide
    Flexible hours
    Shift work

    Ping Identity

    Canada
    21 days ago
  •  ...Reserve]( Learn more at [chain.link]( About The Role The Engineering Team As Chainlink Labs’ engineering org scales, the DevEx...  ...define how engineering teams build, test, and deploy, shaping the reliability and scalability of the developer experience org-wide.... 
    Suggested
    Full time
    Remote work

    Chainlink Labs

    Canada
    10 days ago
  • $121k - $169k

     ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Manager, Site Reliability Engineering Overview: Ethoca and Mastercard are seeking a Manager, Site Reliability Engineering, to lead globally... 
    Suggested
    Full time
    Worldwide

    Mastercard

    Canada
    10 days ago
  •  ...innovated the market-leading enterprise password manager and pioneered Unified Access Management, a new...  ...Claude, Codex , and other tools selected for reliability, security, and fit. This is an internal product engineering role. You’ll treat Sales and Marketing teams... 
    Suggested
    Full time
    Immediate start
    Remote work

    1Password

    Canada
    8 days ago
  •  ...About the AI Platform Foundations Team AI Platform Foundations is a team within Platform Engineering on a mission to make AI a reliable, governed, and productivity-multiplying force across Wealthsimple's engineering organization. We build and maintain the shared AI tooling... 
    Full time
    Immediate start

    Wealthsimple

    Canada
    a month ago
  •  ...is looking for an experienced AWS Cloud Engineer – Platform Operations to support and...  ...This role combines AWS Cloud Operations, Site Reliability Engineering (SRE), Infrastructure Automation...  ...and related platform services. Manage Infrastructure as Code (IaC) using Terraform... 
    Full time
    Internship
    Remote work
    Relocation

    Miratech

    Canada
    23 days ago
  •  ...Summary Platform Infrastructure Engineering builds and operates Menlo Security's Infrastructure...  ...of experienced engineers building and managing the company's core infrastructure...  ...Outcomes & KPIs Key Outcome(s) Owned: Reliable, secure, and scalable infrastructure across... 
    Full time

    Menlo Security

    Canada
    28 days ago
  • $164.49k - $197.39k

     ...open culture. Grafana Cloud, our fully managed observability platform, is flexible and...  ...Salesforce – trust Grafana Labs to ensure reliability of their applications and systems,...  ...tool includes a rule-based recommendation engine that provides useful contextual recommendations... 
    Full time
    Local area
    Remote work
    Flexible hours

    Grafanalabs

    Canada
    a month ago
  •  ...of how work gets done. We are looking for a Senior Solution Engineer who is accustomed to solving customer’s most complex problems...  ...position Snowflake in relation to them. Collaborate with Product Management, Engineering, and Marketing to continuously improve Snowflake’... 
    Full time
    Remote work

    Snowflake

    Canada
    4 days ago
  •  ...The Forward Deployment Engineer (FDE) drives the on-site deployment, integration, and scaling of our enterprise Generative AI solutions. This role...  ...● Leverage the Agent Developer Kit (ADK) to build and manage multi-agent systems that collaborate to solve end-to-end... 
    Full time

    Tiger Analytics

    Canada
    a month ago
  •  ...of this role As a Lead People Systems Engineer, you'll drive how GitLab's People systems...  ...GitLab's People automation ecosystem running reliably at scale across Workato, Google Cloud,...  ..., code reviews, and disciplined change management practices. Develop and manage Slack... 
    Full time
    Remote work

    GitLab

    Canada
    a month ago
  •  ...provision, deploy, observe, and optimize infrastructure. We give engineering teams unprecedented velocity without ever compromising on...  ...# Screening with Marie, Head of HR(45 min to 1hr) # Hiring Manager interview to deep dive into your tech skills (45 min to 1hr)... 
    Remote work
    Canada
    more than 2 months ago
  •  ...AWS Cloud Solutions Engineer responsible for configuration, migration, and management of AIRS systems in AWS Cloud. On-site required first 2 weeks, optional remote after. Requirements: AWS expertise System migration experience Cloud configuration Large system... 
    Temporary work
    Remote work

    Department of Finance and Administration

    Canada
    29 days ago
  •  ...early diagnosis and longitudinal care management of chronic conditions. We value diversity...  ...the future of healthcare. Clover's engineering team is empathetic, caring, and supportive...  ...process changes and ensure pods adopt new reliability standards without friction. You will... 
    Full time
    Work at office
    Remote work
    Work from home
    Shift work
    Afternoon shift
    Early shift

    Cloverhealth

    Canada
    28 days ago
  •  ...signal, at low latency and high reliability in a trading environment...  ...harnesses — tools, context management, evals, guardrails — that do...  ...including what broke. ~ Strong engineering fundamentals in Python, Go,...  ..., monthly wellness plan, on-site weekly massages, and games... 
    Full time
    Local area

    Drw

    Canada
    10 days ago
  •  ...SLAs end-to-end, ensuring data moves reliably between systems with clear expectations and well-defined metrics Define and manage data contracts that prevent upstream changes...  ...across the company, advocating for engineering-specific roadmap items and integrations... 
    Full time

    Wrapbook

    Canada
    10 days ago
  •  ...BeyondTrust is seeking a Software Development Engineer to help build the NHI Governance domain within our Pathfinder products. You...  ...Investigate and fix bugs, and help improve the quality and reliability of the services your team owns. Collaborate with product, design... 
    Full time

    Beyondtrust

    Canada
    23 days ago
  •  ...About the Role: As a frontend-focused Senior Software Engineer , you will own the design and delivery of the user-facing experiences at the heart of our AI products. Your primary focus is building polished, performant, production-grade interfaces in React and TypeScript... 
    Full time
    Casual work

    Sardine

    Canada
    a month ago
  •  ...financial infrastructure with hands-on support so founders can operate efficiently and scale with confidence. We’re hiring a Software Engineer who wants real ownership in a fast-moving startup environment. What You’ll Do Build full-stack web features (front-end,... 
    Full time

    Meow Technologies

    Canada
    20 days ago
  •  ...ownership of quality, testing, security, and reliability from the earliest stages of the...  ...workflows. Collaborate with product managers, designers, and other developers to define...  ...Bachelor’s degree in computer science, Engineering, or a related field (or equivalent... 
    Full time
    Local area
    Remote work
    Shift work

    Fortive

    Canada
    23 days ago
  •  ...Under the supervision of Architecture and Development Management, the Senior Software Engineer is accountable for working with business stakeholders...  ...in modernizing our technology ecosystem, improving the reliability and scalability of our applications, and building... 
    Full time

    WelbeHealth

    Canada
    7 days ago
  •  ...and lending. You'll join a small, high-trust group of senior engineers and own meaningful scope from early on - designing and building...  ...end-to-end, from technical design through to running your work reliably in production. Solve genuinely hard problems around real-... 
    Full time

    Wealthsimple

    Canada
    23 days ago
  • $160.65k - $217.35k

     ...supportive, and versatile group of software engineers and product designers. We love finding...  ...applications, with a focus on memory management and rendering efficiency. AI Usage...  ...such as Playwright or Jest to ensure code reliability and performance. Operational... 
    Full time
    Remote work

    Mapbox

    Canada
    a month ago
  •  ...are always curious , and build with engineering excellence as a first principle. About...  ...and Exchange platform, contributing to reliability, scalability, and performance targets....  ...decision to an engineer or a product manager without losing either of them. ~ Openness... 
    Full time
    Flexible hours

    Life360

    Canada
    9 days ago
  • $111k - $160k

     ...potential. Title and Summary Senior Software Engineer The AI & Decision Engineering...  ..., governed decision-making, and outcome management with business agility at global scale....  ...based in Vancouver, requiring three days on-site per week. Overview • The Decision... 
    Full time
    Work at office
    Worldwide
    Flexible hours
    3 days per week

    Mastercard

    Canada
    22 days ago
  • $91k - $140k

     ...realize their greatest potential. Title and Summary Software Engineer II Overview Ethoca, a Mastercard company, is transforming...  ...peers, supports less experienced team members, and helps deliver reliable, scalable software that meets business and customer needs.... 
    Full time
    Worldwide

    Mastercard

    Canada
    3 days ago
  •  ...Role Overview We’re looking for a talented Software Engineer who excels in technical problem-solving. You’re skilled at developing and shipping high-quality code, with a hunger to learn, and a desire to build things that make a difference. As a Software Engineer... 
    Full time

    Magnet Forensics

    Canada
    3 days ago
  •  ...detect errors and measure model capabilities across the software engineering lifecycle. Key Responsibilities Curate and author code...  ...generated code for correctness, efficiency, scalability, and reliability, and provide structured feedback and improvements.... 
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    Canada
    more than 2 months ago
  • $61.6k - $113.9k

     ...both development and production settings. Assist in release management, version control, and ongoing improvement initiatives....  ...resource to stay updated with new technologies, frameworks, and engineering best practices. Technologies: AI Cloud Copilot... 
    Full time
    Work at office
    2 days per week
    3 days per week

    BMO Financial

    Canada
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Manager, Site Reliability Engineering. Be the first to apply!