SRE
LTM
Role description We are seeking an experienced Site Reliability Engineer SRE Lead to drive platform reliability observability and operational excellence across the API Services ecosystem
Description Role description We are seeking an experienced Site Reliability Engineer SRE Lead to drive platform reliability observability and operational excellence across the API Services ecosystem
This role combines
Production engineering and reliability leadership
Platform security and vulnerability remediation
Ownership of largescale distributed runtime environments
Key Responsibilities Include
- Leading reliability engineering for highscale API platforms 40K runtimes
- Driving EOL remediation and platform stabilization efforts
- Implementing SRE best practices
- SLIs SLOs error budgets
- Incident management and postmortem culture
- Enhancing observability monitoring and proactive fault detection
- Building resilient platforms capable of handling AIdriven usage patterns and threat models
- Supporting global production environments with oncall and escalation coverage
Required Skills
- Strong experience in Site Reliability Engineering Production Engineering
- Handson expertise with
- MuleSoft TIBCO or similar middleware platforms
- Largescale distributed systems and runtime management
- Deep understanding of
- System reliability scalability and high availability design
- Incident management root cause analysis and problem management
- Experience with
- Observability tools eg Dynatrace Splunk Prometheus Grafana
- CICD pipelines Jenkins Git Ansible
- Strong scriptingautomation skills
- Shell Python PowerShell
- Experience managing LinuxUnix and Windows production environments
- Knowledge of
- Microservices API platforms and cloudbased architectures
- Understanding of
- Platform security vulnerability remediation and risk mitigation in production systems
- Excellent troubleshooting skills in highpressure realtime environments
Desired Skills
- Experience implementing SRE frameworks SLIs SLOs error budgets
- Familiarity with
- Kubernetes containerized platforms
- Infrastructure as Code Terraform Ansible
- Exposure to
- AIdriven operational monitoring or security tooling
- Largescale platform modernization or migration programs
- Middleware certifications MuleSoft or equivalent
- Experience in regulated environments eg financial services
Other Details Actual compensation within the range will be dependent upon the individual's skills, experience, performance and internal equity.
Benefits/perks listed below may vary depending on the nature of your employment with LTIMindtree (“LTIM”):
Benefits And Perks
- Comprehensive Medical Plan Covering Medical, Dental, Vision
- Short Term and Long-Term Disability Coverage
- 401(k) Plan with Company match
- Life Insurance
- Vacation Time, Sick Leave, Paid Holidays
- Paid Paternity and Maternity Leave
Disclaimer : The compensation and benefits information provided herein is accurate as of the date of this posting.
LTIMindtree is an equal opportunity employer that is committed to diversity in the workplace. Our employment decisions are made without regard to race, color, creed, religion, sex (including pregnancy, childbirth or related medical conditions), gender identity or expression, national origin, ancestry, age, family-care status, veteran status, marital status, civil union status, domestic partnership status, military service, handicap or disability or history of handicap or disability, genetic information, atypical hereditary cellular or blood trait, union affiliation, affectional or sexual orientation or preference, or any other characteristic protected by applicable federal, state, or local law, except where such considerations are bona fide occupational qualifications permitted by law. #J-18808-Ljbffr
- JPMorgan Chase & Co. is recruiting a Lead Site Reliability Engineer to define the future of reliability on the AI ML and Data platform. You will lead a team, drive resiliency reviews, and mentor engineers while advocating for secure, observable, scalable systems across...Suggested
- ...collaborating with J.P. Morgan to connect them with exceptional professionals for this role. JOB DESCRIPTION We are seeking a Delivery SRE leader who will ensure security applications are delivered with strong SDLC discipline and measurable reliability. This role partners...Suggested
- JPMorgan Chase & Co. is seeking a Lead Software Engineer within the Consumer and Community Banking technology -Deposits Platform to lead resiliency and DR efforts across the product portfolio. The role involves cross-team collaboration, development of failover frameworks...Suggested
- ...SRE Production Support Engineer Location: Plano, TX Duration: 6 Months (Contract to hire) Interview Process: 1st round - Zoom 2nd round – In Person Role Overview: Position is part of the Central Site Reliability Engineering (SRE) Team. Looking...SuggestedContract workShift work
- ...SRE MAHIN-JOB-31492 Location: Plano TX Skill: Web application designing basics-1 8+ years of professional experience processing a culture of learning through the development and sharing of skills, knowledge, process and tools 2. A driving passion for...Suggested
$229.9k - $262.4k
...Overview Full-Stack Engineer 5 (Java, Python, SRE, AWS, AI) (Cloud Operations Resilience Engineering) Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative, inclusive and iterative...Full timePart timeInternshipLocal area- ...Role : SRE Engineer Location : Plano, Texas Job Summary: We are looking for a highly motivated Site Reliability Engineer (SRE) to improve system reliability, scalability, and performance of mission-critical applications. The ideal candidate should have strong...
$87.5k - $125k
Observability Engineer DISH is transforming the future of connectivity. We're doing it by building the country's first virtualized, standalone 5G wireless network from scratch. The foundation of a connected world, it's a network free of the limitations of the past, ...Flexible hoursNight shift- ...integration, and end-to-end tests for control plane components; participate in code reviews Collaborate with infrastructure, platform, and SRE teams to define scheduling policies, resource quotas, and placement constraints Document architecture decisions, APIs, and...
$147.25k - $190k
...storage, load balancing, DNS, certificates/TLS), managing scope, dependencies, risk, and status. Partner with security, IAM, network, SRE/operations, and vendors to architect and implement scalable, resilient solutions and platform modernization. Automate and...Full timeWork at office$229.9k - $262.4k
...incident response efforts and ensure root‑cause analysis drives durable reliability improvements Partner with Cloud, Security, and SRE teams to define cross‑platform observability, telemetry, and self‑healing automation Mentor and upskill engineers across the...Full timePart timeH1bLocal area- ...Cloud-Native Architecture & Platform Modernization* DevOps, CI/CD & Engineering Productivity Platforms* Site Reliability Engineering (SRE) & Operational Excellence* Security, Governance, Compliance & Engineering Controls* LLM Workflows, Prompt Engineering & Context...Hourly payContract workWork experience placement
- ...limits, N+1 mitigation (DataLoader). Proven delivery of API/schema governance, versioning/deprecation, and CI policy gates. Strong SRE practices: SLIs/SLOs, error budgets, OpenTelemetry, data-driven post-incident improvements. Developer productivity: time-to-first-hello...Contract work
$54.68 - $64.68 per hour
...OpenSSH,Oracle Enterprise Linux,performance tuning,performance optimization,postfix,Python,system reliability,scripting,sendmail,SMTP,SRE Practices,vulnerability remediation,high availability,TCP/IP,analytical,communication,organizational skills,Leadership,Troubleshoot,...Hourly payContract workTemporary workWork experience placement- ...reliability, and scalability; follow Agile practices such as Scrum and Continuous Delivery. Support Site Reliability Engineering (SRE) practices to ensure excellent user experience and system performance. Required qualifications, capabilities, and skills: Formal...
- ...ServiceNow - Kafka - API design/integration and troubleshooting - Incident, problem, and change management (ITIL-aligned) - SRE/operations metrics (availability, reliability, MTTR, SLA/SLO reporting) - Root cause analysis and post-incident governance -...
- ...recommendations before use, escalating when uncertain and following data handling expectations. Must have a background in development or SRE Advanced expertise in stakeholder management, with the ability to establish productive working relationships and influence decision-...
- ...Toyota Financial Services Technology Operations Center is looking for a passionate and highly motivated Senior Site Reliability Engineer (SRE) - Backup Infrastructure. In this role, you will apply software engineering principles to ensure the reliability, availability, and...H1bRelocation package
- ...cloud, and colocation environments. You will own end‑to‑end connectivity, security, and automation, partnering with Cloud, Security, and SRE teams to improve reliability and performance. The role requires 8+ years in network engineering, hands‑on experience with BGP, L2/...
- ...and reliable technology environment across banking operations. The ideal candidate brings 10+ years of IT operations leadership in regulated financial services, expertise in DevOps/SRE, and deep experience with hybrid cloud (Azure/AWS). #J-18808-Ljbffr Jobleads-US
- ...MongoDB, and Redis. Experience building and maintaining event streaming solutions using Kafka or RabbitMQ. Well-versed in DevOps and SRE practices, including CI/CD with Jenkins and observability using Splunk and Grafana. Hands-on experience using enterprise-...
- ...latency, scalability, and disaster recovery requirements. ~ Monitor system performance, troubleshoot production issues, and apply SRE and reliability engineering practices. ~ Identify opportunities for performance optimization and cloud cost efficiency. ~ Participate...Local area
- ...scalability, and security for mission-critical workloads. Champion the shift toward next-generation operating models. Drive the adoption of SRE, DevOps, and intelligent automation to reduce toil, increase speed-to-market, and optimize unit economics. Move beyond basic SLA...Local areaFlexible hoursShift work
- ...Toyota Financial Services is seeking a Senior Site Reliability Engineer (SRE) focused on Backup Infrastructure in Plano, TX. You will ensure reliability, availability, and performance of enterprise backup ecosystems, and automate backup workflows for efficiency. You...
- ...incident management, disaster recovery, and business continuity ~ Drive automation, observability, and platform stability through DevOps/SRE principles ~ Ensure robust cybersecurity posture in alignment with FFIEC and NIST frameworks ~ Manage vendor relationships and...Full timeLocal areaRelocation
- ...Description Job Description Responsibilities: \t3-4 years of experience in production engineering and site reliability engineering (SRE) to design, implement, and maintain highly available, scalable, and resilient systems. \tOwn end-to-end operational...
- ...to $66.00/hr. w2 Responsibilities: Build and maintain CI/CD pipelines using Jenkins. Implement and refine observability for SRE, including metrics, logs, and traces. Participate in on-call rotations and ensure service readiness. Drive incident management...Hourly payFor contractorsLocal area
- ...assisted practices and ensuring security and auditability. You will manage cross-functional workstreams, collaborate with security, IAM, and SRE teams, and deliver scalable, resilient infrastructure solutions with IaC and robust runbooks. #J-18808-Ljbffr Jobleads-US
- ...and familiarity with Infrastructure as Code (IaC) tools like Terraform or CloudFormation. ~5+ years in Site Reliability Engineering (SRE), Disaster Recovery Planning, or Distributed Systems Engineering. ~ Demonstrated experience leading effective use of approved AI-...
- ...and networking. Design highly available, scalable, and disaster-resilient solutions. Troubleshoot production issues and apply SRE/reliability practices. Participate in code reviews, technical design discussions, and engineering best practices. Provide technical...Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!


