Site Reliability Engineer
NTT Data
NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com
Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change based on client requirements. For employees near an NTT DATA office or client site, in-office attendance may be required for meetings or events, depending on business needs. At NTT DATA, we are committed to staying flexible and meeting the evolving needs of both our clients and employees. NTT DATA recruiters will never ask for payment or banking information and will only use @nttdata.com and @talent.nttdataservices.com email addresses. If you are requested to provide payment or disclose banking information, please submit a contact us.
Role Summary
We are looking for a strong, hands-on Senior Site Reliability Engineer to support and improve cloud operations for microservice-based platforms. This role requires a senior engineer who can independently manage production reliability, incident response, cloud infrastructure, automation, observability, Kubernetes operations, and CI/CD workflows across AWS and Azure environments.
The ideal candidate should be technically strong, proactive, comfortable in production support, and able to reduce operational toil through automation while improving service availability, performance, scalability, and resilience.
Key Responsibilities
- Own and improve reliability of cloud-based services and supporting infrastructure.
- Participate in on-call rotation and support production systems outside normal business hours.
- Lead incident response activities including triage, escalation, mitigation, and service restoration.
- Drive blameless postmortems and ensure corrective actions are tracked to closure.
- Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis .
- Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools.
- Support and improve cloud/container platforms across AWS and Azure .
- Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems.
- Build automation to reduce manual effort and improve operational efficiency.
- Configure and improve monitoring, alerting, logging, diagnostics, and observability.
- Analyze performance and capacity trends to identify bottlenecks and improve scalability.
- Troubleshoot complex infrastructure, networking, application runtime, and cloud platform issues.
- Support disaster recovery planning, validation, and recovery readiness.
- Create and maintain operational runbooks, support procedures, and engineering documentation.
- Coach and guide other engineers on SRE best practices, reliability, automation, and operational excellence.
Mandatory Skills
- 6–8+ years of experience as an SRE, DevOps Engineer, Infrastructure Engineer, Cloud Engineer, or Platform Engineer.
- Strong hands-on experience with AWS and Azure cloud platforms.
- Strong experience with Terraform for Infrastructure as Code.
- Experience with Atlantis, ArgoCD , or similar infrastructure/deployment automation tools.
- Strong hands-on experience with Docker and Kubernetes .
- Experience designing, maintaining, and troubleshooting complex CI/CD pipelines .
- Strong production support experience including incident management, RCA, postmortems, and runbook creation.
- Strong observability experience: monitoring, alerting, logging, diagnostics, and performance analysis.
- Good understanding of cloud networking, security, access controls, and InfoSec practices.
- Experience with version control, branching, merging, pull requests, and conflict resolution.
- Understanding of cloud cost optimization and resource utilization.
- Strong communication skills and ability to work with DevOps, Engineering, Product, and Delivery teams.
Good-to-Have Skills
- Experience with microservice-based platforms.
- Experience with tools such as Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics , or similar.
- Scripting/programming experience using Python, Bash, Go, or Java .
- Experience with SLI/SLO/SLA, error budgets, capacity planning, and resilience engineering.
- Experience with disaster recovery testing and production readiness reviews.
- Experience supporting customer-facing, high-availability platforms.
- Prior experience mentoring junior engineers or leading technical troubleshooting.
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.
We are currently seeking a Site Reliability Engineer to join our team in Guadalajara, Jalisco (MX-JAL), Mexico (MX).
- ...Overview WeWork Reforma Latino (97001), Mexico, Ciudad de Mexico, Ciudad de Mexico Lead Site Reliability Engineer We're building a Site Reliability Engineering center in Mexico City, and we're hiring a Manager-level Backend Engineer to own the reliability and...SuggestedInternshipLocal area
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer, Data Protection Mastercard is a global technology company in the payments industry. Our mission is to connect and power...SuggestedFull timeWorldwide
- ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Cloud Platform Engineer Overview The Internal Cloud Platforms team is responsible for the daily maintenance, operation, and continuous improvement of...SuggestedFull timeWorldwideAfternoon shift
- ...This is a platform engineer, not a traditional SRE. We want a strong software engineer who specializes in infrastructure - someone who... ...) and compliance frameworks like SOC 1 and SOC 2 SRE and reliability practices ( SLOs , error budgets, incident response, runbooks...SuggestedFull time
- ...writing via AWS Bedrock What You'll Do: Design and build reliable, scalable frontend and backend systems that serve both... ...decisions as the portal grows in scope and complexity Mentor engineers and contribute to TrueML's engineering culture Who You Are...SuggestedFull timeLocal area
- ..., Mexico, Ciudad de Mexico, Ciudad de Mexico Senior Software Engineer - Full Stack Do you love building and pioneering in the technology... ..., educational tools or other information available through this site. Capital One Financial is made up of several different...InternshipLocal area
- ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Platform Engineer Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital...Full timeWorldwide
- ...We are looking for a skilled and technically driven Senior Software Engineer (Python) to join a fast-paced cybersecurity environment. You will own and evolve a critical layer of our software ecosystem — including microservices, security tool integrations, and an in-house...Full timeRemote workFlexible hours
- ...their greatest potential. Title and Summary Principal Software Engineer, SE Guild Mastercard is a global technology company in the... ...the technical solutions that enable SE Guild programs to scale reliably and deliver value across the engineering community. This is...Full timeWorldwide
- ...part of NTT Group, which invests over $3 billion each year in R&D. Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or...Work at officeRemote workFlexible hours
- ...to auditors and regulators. Works at the intersection of data engineering, security and legal or compliance functions, translating regulatory... ...possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored...Work at officeLocal areaRemote workFlexible hours
- ...Jalisco (MX-JAL), Mexico (MX). Job Description: Lex Support Engineer (JAVA and Linux) Role Overview: We are seeking an experienced... ...possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored...Permanent employmentWork experience placementWork at officeRemote workFlexible hours
- ...team in Guadalajara, Jalisco (MX-JAL), Mexico (MX). The SRE Engineer will be able to own the cloud infrastructure, build an infrastructure... ...possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored...Work experience placementWork at officeRemote workFlexible hours
- ...test, and continuously tune live offer funnels and e-commerce sites that take real orders from real customers every hour of the... ...job demands. ● A degree (or equivalent) in computer science, engineering, or a related technical field. ● A genuine foundation in programming...Permanent employmentFull timeWorldwideTrial period
- ..., Mexico, Ciudad de Mexico, Ciudad de Mexico Senior Software Engineer - Backend Do you love building and pioneering in the technology... ..., educational tools or other information available through this site. Capital One Financial is made up of several different...InternshipLocal area
- ...Florida-based client and an English-speaking, distributed team, so reliable working-hours overlap and clear communication are important.... ...with an established business and an experienced international engineering team. Opportunity to work on a distinctive product that...Long term contractFull timeRemote workTrial period
- ...Job Summary We are looking for an experienced and passionate Senior Angular Developer to join our client's engineering team. In this role, you will be a key player in designing and building the user interface for our core applications. You will tackle complex technical...Full timeWork experience placementFlexible hours
- ...and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Power Platform Engineer Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital...Full timeWorldwide
- ...organization, apply now. We are currently seeking a Cloud Migration Engineer to join our team in Guadalajara, Jalisco (MX-JAL), Mexico (MX).... ...possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support...Work at officeRemote workFlexible hours
- ...to global teams and clients. Required Qualifications ~8+ years of experience in Cloud Infrastructure, Architecture, or Cloud Engineering. ~ Professional Google Cloud Architect Certification (required). ~ Strong hands-on experience with Google Kubernetes Engine (...Remote work
- ...services that help people, businesses and governments realize their greatest potential. Title and Summary Lead, AI Security Operations Engineer As an Information Security Engineer – AI Security Operations, you will be at the forefront of protecting the organization’s AI...Full timeWorldwide
$99k
...Ciudad de Mexico, Ciudad de Mexico Senior Manager, Software Engineering (People Leader) Do you love building and pioneering in the... ...educational tools or other information available through this site. Capital One Financial is made up of several different entities...InternshipLocal area- ...their greatest potential. Title and Summary Director, Software Engineering Overview The CNPF Data & AI organization is looking for... ...can translate emerging AI and agentic concepts into secure, reliable, observable, and production-grade systems, while also accelerating...Full timeWorldwide
- ...Mexico, Ciudad de Mexico, Ciudad de Mexico Director, Software Engineering Capital One is seeking a Director of Software... ...experience in Agile practices ~7+ years of experience with Site Reliability Engineering (SRE) At Capital One, we respect individual...Local area
- ...Ciudad de Mexico, Ciudad de Mexico Senior Director, Software Engineering Capital One is seeking an experienced software engineering... ...Software Engineering, to help us build and grow our Technology Site in Mexico City. Based in Mexico City, the Senior Director of...Local areaShift work
- ...Key Responsibilities Develop and refine strategies for acquiring research data from scientific laboratory instruments. Engineer and operationalize custom file parsers to handle diverse instrument output formats—including .xlsx , .pdf , .txt , .raw , .fid...Full timeWork experience placement
- ...innovative SportsTech & EdTech company to find a Senior AI Systems Engineer who will play a key role in designing, optimizing, and scaling... ...(LLMs), processing structured data, and developing scalable, reliable systems with a strong focus on quality, automation, and...Part time
$15 per hour
Summary The Wikimedia Foundation is seeking a Senior Software Engineer to join the team supporting the Wikidata Platform — the structured... ...-scale, production-grade services while ensuring performance, reliability, and maintainability. Working closely with the technical and...Full time- ...This is a platform engineer, not a traditional SRE. We want a strong software engineer who specializes in infrastructure - someone who... ...) and compliance frameworks like SOC 1 and SOC 2 SRE and reliability practices ( SLOs , error budgets, incident response, runbooks...Full time
- ...organization, apply now. We are currently seeking a Cloud Support Engineer/Sr Specialist to join our team in guadalajara, Jalisco (MX-JAL)... ...possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored...Permanent employmentWork at officeRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- on-site clinical research associate (traveling/remote) Mexico
- junior website developer Mexico
- site reliability engineer
- site reliability engineering manager
- junior site reliability engineer
- site reliability engineer sre
- site reliability engineer remote
- lead site reliability engineer
- savannah river site
- site inspector




