Principal Site Reliability Engineer
Mastercard
Principal Site Reliability Engineer
Who is Mastercard?
At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits everyone, everywhere, by making transactions safe, simple, smart, and accessible. Using secure data and networks, partnerships, and passion, our innovations and solutions help individuals, financial institutions, governments, and businesses realize their greatest potential. Our decency quotient, or DQ, drives our culture and everything we do inside and outside of our company. We cultivate a culture of inclusion for all employees that respects their individual strengths, views, and experiences. We believe that our differences enable us to be a better team – one that makes better decisions, drives innovation, and delivers better business results.
What we create today will define tomorrow. Revolutionary technologies that reshape the digital economy to be more connected and inclusive than ever before. Safer, faster, more sustainable. And we need the best people to do it. Technologists who are energized by the challenges of a truly global network. With the talent and vision to create the critical systems and products that power global commerce and connect people everywhere to the vital goods and services they need every day. Working at Mastercard means being part of a unique culture. Inclusive and diverse, a rich collaboration of ideas and perspectives. A place that celebrates your strengths, values your experiences, and offers you the flexibility to shape a career across disciplines and continents. And the opportunity to work alongside experts and leaders at every level of the business, improving what exists, and inventing what’s next. About the Role
The Business Operations team is seeking a Principal Site Reliability Engineer.
The role of Business Operations Organization is to be the production readiness steward for Mastercard products. As a Business Operations SRE, we are responsible for ensuring that our platform is stable and healthy. We break down barriers to run our products by fostering developer run ownership and empowering developers to build resilient products. We support our developers during the application build phase in software run principals that includes operational design, automation, capacity planning, monitoring that leads to fault-tolerant, scalable products. We see the big picture and help create and enforce operations standards while facilitating an agile and learning culture. We are seeking a highly motivated and experienced Principal Site Reliability Engineer to join our growing team. You will play a critical role in ensuring the reliability, scalability, and performance of our applications, supporting essential services that power Mastercard's global operations. As a thought leader in your field, you will bring technical expertise, a passion for automation, and the ability to mentor. We support daily operations with a hyper focus on triage, root cause by understanding the business impact of our products and subsequently performing blameless post-mortems. The goal of every Business Operations team is to engage early in the development lifecycle to be more proactive and upfront in the development process, and to proactively manage production and change activities to maximize customer experience and increase the overall value of supported applications. Business Operations teams also focus on risk management by tying all our activities together with an overarching responsibility for compliance and risk mitigation across all our environments. Ultimately, the role of Business Operations is to align Product and Customer Focused priorities with Operational needs by providing continuous feedback throughout the lifecycle. Team Specific Skills:
It is not expected that any single candidate would have expertise across all these areas, but a Site Reliability
engineer will spend time throughout their career with all of these aspects of the role:
• Operational Readiness Architect:
o Serve as the primary contact responsible for the overall application health,
performance, and capacity
o Support services before they go live through activities such as system design
consulting, capacity planning and launch reviews.
o Partner with the development and product team of a new application to establish
the right monitoring and alerting strategy and create the framework to achieve zero
downtime during deployment.
• Site Reliability Engineering:
o Performs operability and resilience design and implements and maintains highly
reliable and scalable infrastructure.
o Perform root cause analysis of incidents and collaborate with development teams to
resolve issues.
o Stay up to date with the latest technologies and trends in SRE and cloud computing.
o Participate in on-call rotations and be available to respond to critical incidents.
o Complete end-to-end run ownership of the product.
Practice sustainable incident response and blameless post-mortems while taking a
holistic approach to problem solving and optimizing time to recover.
o Automate data-driven alerts to proactively escalate issues. Work with development
teams to establish SLOs and improve reliability.
• DevOps/Automation:
o Tackle complex development, automation, and business process problems. o Support the application CI/CD pipeline for promoting software into higher
environments through validation and operational gating, and lead Mastercard in
DevOps automation and best practices.
o Performs operational and resilience Design and implements solutions for capacity
planning and performance optimization.
o Increase automation and tooling to reduce toil and manual intervention
• ITSM Practices:
o Analyses ITSM activities of the platform and provide feedback loop to development
teams on operational gaps or resiliency concerns Role qualifications:
The ideal candidate will have experience in many of these areas:
• BS degree in Computer Science or related technical field involving coding (e.g., physics or
mathematics), or equivalent practical experience.
• Strong understanding of DevOps principles, practices along with configuration management.
• Experience in operational and resilience designing, building, and operating large-scale,
distributed systems.
• Appetite for change and pushing the boundaries of what can be done with automation. Be
curious about new technology, infrastructure, and practices to scale our architecture and
prepare for future growth.
• Experience with algorithms, data structures, scripting, pipeline management, and software
design.
• Systematic problem-solving approach, analytical, coupled with strong communication skills
and a sense of ownership and drive.
• Interest in designing, analysing, and troubleshooting large-scale distributed systems.
• Strong leadership and mentoring skills.
• A passion for observability, automation and continuous improvement.
• Willingness and ability to learn and take on challenging opportunities and to work as a
member of matrix based diverse and geographically distributed project team.
• Ability to balance doing things right with fixing things quickly. Flexible and pragmatic, while
working towards improving the long-term health of the system.
• Comfortable collaborating with cross-functional teams to ensure that expected system
behaviour is understood and monitoring exists to detect anomalies. Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard’s security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and
- Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.
$151.6k - $245.3k
...Site Reliability Engineer Palo Alto Networks runs a large hybrid infrastructure and is one of the largest GCP customers. As a Site Reliability Engineer, you will be part of a team supporting the services running on this infrastructure. This includes automation, architecture...Principal$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics...PrincipalFull timeWork experience placementWorldwide$174k - $252k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...Suggested$90k - $180k
...nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We are...SuggestedRemote work$160k - $240k
...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit... ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in...SuggestedFull time$165k - $280k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most...Permanent employmentTemporary workWorldwideWeekend work- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Principal Site Reliability Engineer Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that...PrincipalFull timeWorldwideFlexible hours
$222k - $300.5k
...possible.Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps... ...hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency...WorldwideShift work$210.6k - $305.1k
...Minimum Qualifications: You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...Full timeTemporary workLocal areaFlexible hours- Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms...
$165k - $190k
...DevOps / SRE TeamThe DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-... ...security platformAddress complex challenges around scalability, reliability, observability, and cost efficiencyCollaborate with Engineering...Work from home$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,...Full timeWorldwide$147k - $210k
...product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e.g., style... ..., and troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when you treat operations...- ...keep the world running. Location: 5 on-site days a week in Sunnyvale, CA Headquarters. Our Team's Vision: Our Engineering team is shaping the future of... ...are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background in...Work experience placementImmediate start
$189k - $232k
...Site Reliability Engineer Mountain View, US About EarnIn As one of the first pioneers of earned wage access, our passion at EarnIn is building products that deliver real-time financial flexibility for those with the unique needs of living paycheck to paycheck....Full timeWork at office2 days per week$145k - $165k
...Your Ego : Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to...Work at officeImmediate start- ...Senior Site Reliability Engineer Location: Remote Duration: 12 month contract to start IV Process: 1-3 Round IV process International Tech Top Skills: Java Python NodeJS -DevOps Engineer should work here too Main Responsibilities:...Contract workLocal areaRemote work
$100k - $200k
...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...Full time- ...Palo Alto Networks in Santa Clara seeks a visionary Senior Principal Engineer/Architect to serve as the technical authority for global SRE and Platform Engineering initiatives. You will be the primary architect and driver of our AI-driven Autonomous SRE transformation...Principal
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability...Work at office
- ...Site Reliability Engineer Sunnyvale, CA Site Reliability Eng. Must have LinkedIn profile. • t least 8+ years in a Reliability Engineering, DevOps or infrastructure focused role • Strong AWS and Linux Operating System, standard networking protocols, component...
- ...Overview: *Must have Apple experience* • At least 8+ years in a Reliability Engineering, DevOps or infrastructure focused role • Advanced experience with programming languages (Python, Java) • Passion for designing and building reliable systems • Strong sense...
$119.8k - $234.7k
...per yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual... ...the Role We’re building AI‑first engineering systems that power growth at Microsoft —... ...applications Proficiency in designing scalable, reliable systems that support rapid iteration and...PrincipalOngoing contractLocal area3 days per week$261.5k - $353.5k
...already powers the QuickBooks Connector on Claude and ChatGPT, and our own Intuit Intelligence capability.We're hiring a Principal Software Engineer to own the multi-year technology vision and strategy for agent-readiness across Fintech — both bets, end to end. This is...PrincipalWorldwide$174k - $252k
Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical experience. 5 years of experience with... ...s degree in Computer Science or Engineering. About The Job Site Reliability Engineering (SRE) is what you get when you treat operations...- Google is hiring Site Reliability Engineers (SRE) in Sunnyvale, CA, to ensure reliability and performance across Google’s services. The role blends software and systems engineering, allowing code fixes to improve systems while maintaining production reliability at scale...
- ..., and the challenges of building in a high-growth startup, we’d love to talk. This is more than a job—it’s a journey. Site Reliability Engineers (SREs) are responsible for the overall performance and reliability of ASAPP's infrastructure and products. The team owns...Remote work
$207k - $300k
Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability... ...Science or Engineering.1 year of people management experience. Site Reliability Engineering (SRE) combines software and systems engineering...$174k - $252k
Senior Software Engineer, Site Reliability Engineering X Applicants in San Francisco: Qualified applications with arrest or conviction records will be considered for employment in accordance with the San Francisco Fair Chance Ordinance for Employers and the California...Full time$200k
...accuracy, governance, and scale are non-negotiable. The Principal Software Engineer role exists to help us continue raising the engineering bar... ...issue resolutionOwn the quality, performance, and reliability of the application surfaces you build — including on-call...PrincipalWork at officeShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- principal network engineer Mountain View, CA
- principal engineer Mountain View, CA
- principal infrastructure engineer Mountain View, CA
- senior civil engineer project manager Mountain View, CA
- senior chief engineer Mountain View, CA
- engineering director Mountain View, CA
- senior director engineering Mountain View, CA
- director systems engineering Mountain View, CA
- principal developer Mountain View, CA
- general engineer Mountain View, CA


