Site Reliability Engineer
NTT Data
NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com
Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change based on client requirements. For employees near an NTT DATA office or client site, in-office attendance may be required for meetings or events, depending on business needs. At NTT DATA, we are committed to staying flexible and meeting the evolving needs of both our clients and employees. NTT DATA recruiters will never ask for payment or banking information and will only use @nttdata.com and @talent.nttdataservices.com email addresses. If you are requested to provide payment or disclose banking information, please submit a contact us.
Role Summary
We are looking for a strong, hands-on Senior Site Reliability Engineer to support and improve cloud operations for microservice-based platforms. This role requires a senior engineer who can independently manage production reliability, incident response, cloud infrastructure, automation, observability, Kubernetes operations, and CI/CD workflows across AWS and Azure environments.
The ideal candidate should be technically strong, proactive, comfortable in production support, and able to reduce operational toil through automation while improving service availability, performance, scalability, and resilience.
Key Responsibilities
- Own and improve reliability of cloud-based services and supporting infrastructure.
- Participate in on-call rotation and support production systems outside normal business hours.
- Lead incident response activities including triage, escalation, mitigation, and service restoration.
- Drive blameless postmortems and ensure corrective actions are tracked to closure.
- Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis .
- Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools.
- Support and improve cloud/container platforms across AWS and Azure .
- Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems.
- Build automation to reduce manual effort and improve operational efficiency.
- Configure and improve monitoring, alerting, logging, diagnostics, and observability.
- Analyze performance and capacity trends to identify bottlenecks and improve scalability.
- Troubleshoot complex infrastructure, networking, application runtime, and cloud platform issues.
- Support disaster recovery planning, validation, and recovery readiness.
- Create and maintain operational runbooks, support procedures, and engineering documentation.
- Coach and guide other engineers on SRE best practices, reliability, automation, and operational excellence.
Mandatory Skills
- 6–8+ years of experience as an SRE, DevOps Engineer, Infrastructure Engineer, Cloud Engineer, or Platform Engineer.
- Strong hands-on experience with AWS and Azure cloud platforms.
- Strong experience with Terraform for Infrastructure as Code.
- Experience with Atlantis, ArgoCD , or similar infrastructure/deployment automation tools.
- Strong hands-on experience with Docker and Kubernetes .
- Experience designing, maintaining, and troubleshooting complex CI/CD pipelines .
- Strong production support experience including incident management, RCA, postmortems, and runbook creation.
- Strong observability experience: monitoring, alerting, logging, diagnostics, and performance analysis.
- Good understanding of cloud networking, security, access controls, and InfoSec practices.
- Experience with version control, branching, merging, pull requests, and conflict resolution.
- Understanding of cloud cost optimization and resource utilization.
- Strong communication skills and ability to work with DevOps, Engineering, Product, and Delivery teams.
Good-to-Have Skills
- Experience with microservice-based platforms.
- Experience with tools such as Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics , or similar.
- Scripting/programming experience using Python, Bash, Go, or Java .
- Experience with SLI/SLO/SLA, error budgets, capacity planning, and resilience engineering.
- Experience with disaster recovery testing and production readiness reviews.
- Experience supporting customer-facing, high-availability platforms.
- Prior experience mentoring junior engineers or leading technical troubleshooting.
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.
We are currently seeking a Site Reliability Engineer to join our team in Guadalajara, Jalisco (MX-JAL), Mexico (MX).
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits...SuggestedRemote jobFull timeWorldwide
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Site Reliability Engineer The Xborder team is looking for a Senior Site Reliability Engineer who can help us solve problems, implement automation...SuggestedFull timeWorldwide
- ...you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now. We are currently seeking a Site Reliability Engineer to join our team in Guadalajara, Jalisco (MX-JAL), Mexico (MX). SRE – Site Reliability Engineer We are currently...SuggestedWork at officeRemote workMonday to FridayFlexible hoursRotating shiftDay shift
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Site Reliability Engineer Overview: The role of Business Operations Organization is to be the production readiness steward for Mastercard...SuggestedFull timeWorldwideShift work
- ...Overview WeWork Reforma Latino (97001), Mexico, Ciudad de Mexico, Ciudad de Mexico Lead Site Reliability Engineer We're building a Site Reliability Engineering center in Mexico City, and we're hiring a Manager-level Backend Engineer to own the reliability and...SuggestedInternshipLocal area
- ...Our trading platform powers every customer interaction, making reliability a first-class product concern. You will be responsible for... ...observable, and resilient. You'll collaborate closely with software engineers to improve monitoring, deployment safety, automation, fault...Full time
- ...services that help people, businesses and governments realize their greatest potential. Title and Summary Business Operations Site Reliability Engineer Overview: The role of Business Operations Organization is to be the production readiness steward for Mastercard...Full timeWorldwideShift work
- ...help people, businesses and governments realize their greatest potential. Title and Summary Director, Infrastructure & Site Reliability Engineering Who is Mastercard? Mastercard is a global technology company in the payments industry. Our mission is to connect...Full timeWorldwide
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineering Manager The Xborder team is looking for a Site Reliability Engineering Manager who can help us solve problems,...Full timeWorldwideShift work
- ...innovative SportsTech & EdTech company to find a Senior AI Systems Engineer who will play a key role in designing, optimizing, and scaling... ...(LLMs), processing structured data, and developing scalable, reliable systems with a strong focus on quality, automation, and...Contract workPart timeRemote workMonday to Friday
$15 per hour
...Summary The Wikimedia Foundation is seeking a Senior Software Engineer to join the team supporting the Wikidata Platform — the... ...-scale, production-grade services while ensuring performance, reliability, and maintainability. Working closely with the technical and product...Full timeRemote work- ...their greatest potential. Title and Summary Senior Software Engineer Overview: The Mastercard Developer Workbench, part of the... ...the development lifecycle, with a focus on usability and reliability. Contribute to the governance of the "birthright and optional...Full timeWorldwide
- ...realize their greatest potential. Title and Summary Software Engineer II Overview The CNPF Data & AI organization is looking for... ...platform innovation—turning emerging technologies into secure, reliable, and reusable capabilities that create measurable business...Full timeWorldwide
- ..., Mexico, Ciudad de Mexico, Ciudad de Mexico Senior Software Engineer - Full Stack Do you love building and pioneering in the technology... ..., educational tools or other information available through this site. Capital One Financial is made up of several different...InternshipLocal area
- ...hire locally to NTT DATA offices or client sites. This ensures we can provide timely and... ...Data is looking for a Senior DevOps Engineer with strong experience in infrastructure... ..., monitoring, security, and application reliability. The candidate should be hands-on with...Work at officeRemote workFlexible hours
- ...1), Mexico, Ciudad de Mexico, Ciudad de Mexico Lead Software Engineer - Full Stack Do you love building and pioneering in the technology... ..., educational tools or other information available through this site. Capital One Financial is made up of several different...InternshipLocal area
- ...their greatest potential. Title and Summary Principal Software Engineer, SE Guild Mastercard is a global technology company in the... ...the technical solutions that enable SE Guild programs to scale reliably and deliver value across the engineering community. This is...Full timeWorldwide
- ...We are looking for a skilled and technically driven Senior Software Engineer (Python) to join a fast-paced cybersecurity environment. You will own and evolve a critical layer of our software ecosystem — including microservices, security tool integrations, and an in-house...Full timeRemote workFlexible hours
- ...About the Team: We’re a small DevOps team supporting the whole Engineering organization, building applications on top of EC2 and AWS. We own all aspects of the SDLC but strive to automate self-service wherever possible. Being a small team, we also practice SRE, continuously...Full timeRemote work
- ...realize their greatest potential. Title and Summary Lead Software Engineer Overview The CNPF Data & AI organization is looking for a... ...You will lead the design and delivery of secure, scalable, and reliable agentic applications that can reason, orchestrate tools,...Full timeTemporary workWorldwide
- ...part of NTT Group, which invests over $3 billion each year in R&D. Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or...Work at officeRemote workFlexible hours
- ...flexibility. What We Need Symbotic is seeking a System Engineer to support the day-to-day reliability, availability, and performance of our automated... ...health, equipment reliability, and technical execution on site What we do The System Engineer is part of the...Full timeFlexible hours
- ...part of NTT Group, which invests over $3 billion each year in R&D. Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or...Work experience placementWork at officeImmediate startRemote workFlexible hours
- NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now. We are currently seeking a Sr. Salesforce Technical Lead to join ...
- ...will collaborate closely with business users, architects, data engineers, and leadership to transform complex data into meaningful dashboards... ...possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored...Full timeWork at officeRemote workFlexible hours
- ...Guadalajara, Jalisco (MX-JAL), Mexico (MX). 1. L2 Production Support Engineer: Job Description Mandatory Skills: · 3+ years of relevant... ...possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored...Work at officeRemote workFlexible hours
- ...application performance, scalability, and reliability Write clean, maintainable, well‑... ...Experience with ETL pipelines or data engineering concepts Who You Are Energetic, proactive... ...locally to NTT DATA offices or client sites. This ensures we can provide timely and...Work at officeRemote workFlexible hours
$2,500 per month
...test, and continuously tune live offer funnels and e-commerce sites that take real orders from real customers every hour of the... ...job demands. ● A degree (or equivalent) in computer science, engineering, or a related technical field. ● A genuine foundation in programming...Permanent employmentFull timeWorldwideTrial period- ...Florida-based client and an English-speaking, distributed team, so reliable working-hours overlap and clear communication are important.... ...with an established business and an experienced international engineering team. Opportunity to work on a distinctive product that...Long term contractFull timeRemote workTrial period
- ...About the Role: We are looking for an Engineering Manager to lead and grow an existing RPA team while contributing directly as a hands-on backend engineer at an AI-native healthcare platform. This is a dual role where you will manage people and write production-quality...Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Mexico
- junior website developer Mexico
- on-site clinical research associate (traveling/remote) Mexico
- site reliability engineer sre
- site reliability engineering manager
- junior site reliability engineer
- site reliability engineer
- site reliability engineer remote
- lead site reliability engineer
- website qa testing







