Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

NTT DATA, Inc.

NTT DATA is a $30 billion trusted global innovator of business and technology services. We serve 75% of the Fortune Global 100 and are committed to helping clients innovate, optimize and transform for long term success. As a Global Top Employer, we have diverse experts in more than 50 countries and a robust partner ecosystem of established and start-up companies. Our services include business and technology consulting, data and artificial intelligence, industry solutions, as well as the development, implementation and management of applications, infrastructure and connectivity. We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a part of NTT Group, which invests over $3.6 billion each year in R&D to help organizations and society move confidently and sustainably into the digital future. Visit us at us.nttdata.com

Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or hybrid work options, these arrangements are subject to change based on client requirements. For employees near an NTT DATA office or client site, in-office attendance may be required for meetings or events, depending on business needs. At NTT DATA, we are committed to staying flexible and meeting the evolving needs of both our clients and employees. NTT DATA recruiters will never ask for payment or banking information and will only use @nttdata.com and @talent.nttdataservices.com email addresses. If you are requested to provide payment or disclose banking information, please submit a contact us.

Role Summary

We are looking for a strong, hands-on Senior Site Reliability Engineer to support and improve cloud operations for microservice-based platforms. This role requires a senior engineer who can independently manage production reliability, incident response, cloud infrastructure, automation, observability, Kubernetes operations, and CI/CD workflows across AWS and Azure environments.

The ideal candidate should be technically strong, proactive, comfortable in production support, and able to reduce operational toil through automation while improving service availability, performance, scalability, and resilience.

Key Responsibilities

  • Own and improve reliability of cloud-based services and supporting infrastructure.
  • Participate in on-call rotation and support production systems outside normal business hours.
  • Lead incident response activities including triage, escalation, mitigation, and service restoration.
  • Drive blameless postmortems and ensure corrective actions are tracked to closure.
  • Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis .
  • Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools.
  • Support and improve cloud/container platforms across AWS and Azure .
  • Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems.
  • Build automation to reduce manual effort and improve operational efficiency.
  • Configure and improve monitoring, alerting, logging, diagnostics, and observability.
  • Analyze performance and capacity trends to identify bottlenecks and improve scalability.
  • Troubleshoot complex infrastructure, networking, application runtime, and cloud platform issues.
  • Support disaster recovery planning, validation, and recovery readiness.
  • Create and maintain operational runbooks, support procedures, and engineering documentation.
  • Coach and guide other engineers on SRE best practices, reliability, automation, and operational excellence.

Mandatory Skills

  • 6–8+ years of experience as an SRE, DevOps Engineer, Infrastructure Engineer, Cloud Engineer, or Platform Engineer.
  • Strong hands-on experience with AWS and Azure cloud platforms.
  • Strong experience with Terraform for Infrastructure as Code.
  • Experience with Atlantis, ArgoCD , or similar infrastructure/deployment automation tools.
  • Strong hands-on experience with Docker and Kubernetes .
  • Experience designing, maintaining, and troubleshooting complex CI/CD pipelines .
  • Strong production support experience including incident management, RCA, postmortems, and runbook creation.
  • Strong observability experience: monitoring, alerting, logging, diagnostics, and performance analysis.
  • Good understanding of cloud networking, security, access controls, and InfoSec practices.
  • Experience with version control, branching, merging, pull requests, and conflict resolution.
  • Understanding of cloud cost optimization and resource utilization.
  • Strong communication skills and ability to work with DevOps, Engineering, Product, and Delivery teams.

Good-to-Have Skills

  • Experience with microservice-based platforms.
  • Experience with tools such as Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics , or similar.
  • Scripting/programming experience using Python, Bash, Go, or Java .
  • Experience with SLI/SLO/SLA, error budgets, capacity planning, and resilience engineering.
  • Experience with disaster recovery testing and production readiness reviews.
  • Experience supporting customer-facing, high-availability platforms.
  • Prior experience mentoring junior engineers or leading technical troubleshooting.

NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.

We are currently seeking a Site Reliability Engineer to join our team in Guadalajara, Jalisco (MX-JAL), Mexico (MX).

Vacancy posted 22 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Mexico vacancy
  •  ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer II Who is Mastercard? At Mastercard technology, we work to connect and power an inclusive, digital economy that benefits... 
    Suggested
    Remote job
    Full time
    Worldwide

    Mastercard

    Mexico
    a month ago
  •  ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Site Reliability Engineer Overview: The role of Business Operations Organization is to be the production readiness steward for Mastercard... 
    Suggested
    Full time
    Worldwide
    Shift work

    Mastercard

    Mexico
    a month ago
  •  ...Overview WeWork Reforma Latino (97001), Mexico, Ciudad de Mexico, Ciudad de Mexico Lead Site Reliability Engineer We're building a Site Reliability Engineering center in Mexico City, and we're hiring a Manager-level Backend Engineer to own the reliability and... 
    Suggested
    Internship
    Local area

    Capital One

    Mexico
    a month ago
  •  ...services that help people, businesses and governments realize their greatest potential. Title and Summary Business Operations Site Reliability Engineer Overview: The role of Business Operations Organization is to be the production readiness steward for Mastercard... 
    Suggested
    Full time
    Worldwide
    Shift work

    Mastercard

    Mexico
    a month ago
  •  ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Site Reliability Engineer, Data Protection Mastercard is a global technology company in the payments industry. Our mission is to connect and power... 
    Suggested
    Full time
    Worldwide

    Mastercard

    Mexico
    6 days ago
  •  ...innovative SportsTech & EdTech company to find a Senior AI Systems Engineer who will play a key role in designing, optimizing, and scaling...  ...(LLMs), processing structured data, and developing scalable, reliable systems with a strong focus on quality, automation, and... 
    Contract work
    Part time
    Remote work
    Monday to Friday

    Rehire

    Mexico
    27 days ago
  • Summary The Wikimedia Foundation is seeking a Senior Software Engineer to join the team supporting the Wikidata Platform — the structured...  ...-scale, production-grade services while ensuring performance, reliability, and maintainability. Working closely with the technical and... 
    Full time

    Wikimedia Foundation

    Mexico
    7 days ago
  •  ...services that help people, businesses and governments realize their greatest potential. Title and Summary Senior Cloud Platform Engineer Overview The Internal Cloud Platforms team is responsible for the daily maintenance, operation, and continuous improvement of... 
    Full time
    Worldwide
    Afternoon shift

    Mastercard

    Mexico
    11 days ago
  •  ...realize their greatest potential. Title and Summary Software Engineer II Overview The CNPF Data & AI organization is looking for...  ...platform innovation—turning emerging technologies into secure, reliable, and reusable capabilities that create measurable business... 
    Full time
    Worldwide

    Mastercard

    Mexico
    a month ago
  •  ...their greatest potential. Title and Summary Senior Software Engineer Overview: The Mastercard Developer Workbench, part of the...  ...the development lifecycle, with a focus on usability and reliability. Contribute to the governance of the "birthright and optional... 
    Full time
    Worldwide

    Mastercard

    Mexico
    a month ago
  •  ..., Mexico, Ciudad de Mexico, Ciudad de Mexico Senior Software Engineer - Full Stack Do you love building and pioneering in the technology...  ..., educational tools or other information available through this site. Capital One Financial is made up of several different... 
    Internship
    Local area

    Capital One

    Mexico
    more than 2 months ago
  •  ...products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Platform Engineer Mastercard is a global technology company in the payments industry. Our mission is to connect and power an inclusive, digital... 
    Full time
    Worldwide

    Mastercard

    Mexico
    7 days ago
  •  ...experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale. ElevenCreative...  ...doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to... 
    Remote job
    Full time
    Immediate start

    Valor Capital Group

    Mexico
    5 days ago
  •  ...We’re looking for an accomplished DevOps Engineer to help shape and evolve our cloud,...  ...automation, and infrastructure that support reliable, secure, and scalable service delivery across...  ...infrastructure, automation, CI/CD, and site reliability engineering. What you’ll... 
    Full time

    Cision

    Mexico
    14 days ago
  •  ...We are looking for a skilled and technically driven Senior Software Engineer (Python) to join a fast-paced cybersecurity environment. You will own and evolve a critical layer of our software ecosystem — including microservices, security tool integrations, and an in-house... 
    Full time
    Remote work
    Flexible hours

    FusionHit

    Mexico
    26 days ago
  • About the Team: We’re a small DevOps team supporting the whole Engineering organization, building applications on top of EC2 and AWS. We own all aspects of the SDLC but strive to automate self-service wherever possible. Being a small team, we also practice SRE, continuously... 
    Full time
    Remote work

    Peek

    Mexico
    7 days ago
  •  ...to ensure the delivery of high-quality, reliable software products. Core Technical Skills...  ...DevOps) Nice to Have 1) Software engineering  - SQL, Python, ETL, ELT tools proficiency...  ...locally to NTT DATA offices or client sites. This ensures we can provide timely and... 
    Work at office
    Remote work
    Flexible hours

    NTT DATA, Inc.

    Mexico
    6 days ago
  • About Sezzle: With a mission to financially empower the next generation, Sezzle is revolutionizing the shopping experience beyond payments, blending cutting-edge tech with seamless, interest-free installment plans that make shopping smarter and more accessible. We’re not...
    Full time

    Sezzle

    Mexico
    7 days ago
  •  ...realize their greatest potential. Title and Summary Lead Software Engineer Overview The CNPF Data & AI organization is looking for a...  ...You will lead the design and delivery of secure, scalable, and reliable agentic applications that can reason, orchestrate tools,... 
    Full time
    Temporary work
    Worldwide

    Mastercard

    Mexico
    a month ago
  •  ...their greatest potential. Title and Summary Principal Software Engineer, SE Guild Mastercard is a global technology company in the...  ...the technical solutions that enable SE Guild programs to scale reliably and deliver value across the engineering community. This is... 
    Full time
    Worldwide

    Mastercard

    Mexico
    28 days ago
  • We’re looking for an accomplished DevOps Engineer to help shape and evolve our cloud, DevOps...  ..., and infrastructure that support reliable, secure, and scalable service delivery across...  ...cloud infrastructure, automation, CI/CD, and site reliability engineering. What you’ll do... 
    Full time

    Brandwatch

    Mexico
    7 days ago
  • We are looking for a skilled and technically driven Senior Software Engineer (Python) to join a fast-paced cybersecurity environment. You will own and evolve a critical layer of our software ecosystem — including microservices, security tool integrations, and an in-house... 
    Full time

    FusionHit

    Mexico
    7 days ago
  • At Cision, we believe in empowering every individual to make an impact. Here, your voice is heard, your ideas are valued, and your unique perspective fuels our collective success. As part of our global team, you'll thrive in an environment that champions curiosity, collaboration...
    Full time

    Cision

    Mexico
    7 days ago
  • Strong proficiency in Java (Java 8 or above) Hands-on experience with Spring Boot, Spring MVC, Spring Core Experience with JPA/Hibernate or Spring Data Strong knowledge of REST APIs, JSON, and protocols Experience working with databases such as: ...
    Full time

    Siri InfoSolutions Inc

    Mexico
    1 day ago
  •  ...seeking a talented and motivated best-in-class Principal Software Engineer to own and lead the evolution of authentication and...  ...microservices architecture while maintaining the performance, reliability, and latency budgets expected of an edge gateway. Partner with... 
    Full time
    Immediate start

    Sezzle

    Mexico
    13 days ago
  •  ...team in Guadalajara, Jalisco (MX-JAL), Mexico (MX). The SRE Engineer will be able to own the cloud infrastructure, build an infrastructure...  ...possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    NTT DATA, Inc.

    Mexico
    20 days ago
  •  ...part of NTT Group, which invests over $3 billion each year in R&D. Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored to each client’s needs. While many positions offer remote or... 
    Work at office
    Remote work
    Flexible hours

    NTT DATA, Inc.

    Mexico
    14 days ago
  •  ...projects and operational issues for existing and proposed Oracle Engineered Systems (Exadata, Exalogic and Super Clusters) Ability to...  ...Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support tailored... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    NTT DATA, Inc.

    Mexico
    12 days ago
  •  ...About the Role: We are looking for an Engineering Manager to lead and grow an existing RPA team while contributing directly as a hands-on backend engineer at an AI-native healthcare platform. This is a dual role where you will manage people and write production-quality... 
    Full time
    Remote work

    Hireoverseas

    Mexico
    a month ago
  • Job Summary We are looking for an experienced and passionate Senior Angular Developer to join our client's engineering team. In this role, you will be a key player in designing and building the user interface for our core applications. You will tackle complex technical... 
    Full time

    KUNAI

    Mexico
    7 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!