Senior Associate, Enterprise Technology Command Center, Major Incident Manager, Site Reliability Engineering
DBS Bank Ltd
Group Technology enables and empowers the bank with an efficient, nimble and resilient infrastructure through a strategic focus on productivity, quality & control, technology, people capability and innovation. In Group Technology, we manage the majority of the Bank's operational processes and inspire to delight our business partners through our multiple banking delivery channels.
Role Purpose:Sr Associate, Major Incident Manager for ETCC SRE Operations, is responsible for the end-to-end management of critical and major incidents impacting DBS's technology services. This role focuses on minimizing service disruption, restoring service rapidly, and driving continuous improvement in incident response processes. The Major Incident Manager will lead cross-functional teams during incidents, ensuring clear communication, effective problem resolution, and adherence to SRE principles for reliability and operational excellence.
Key Responsibilities:- Major Incident Management:
- Act as Incident manager, to manage major incidents from detection through resolution, ensuring rapid service restoration and minimal business impact.
- Act as the primary communication point during major incidents, providing timely and accurate updates to senior management, stakeholders, and affected business units. Own incident escalation to senior management
- Coordinate and drive technical teams (including SRE, development, infrastructure, and operations) to diagnose, troubleshoot, and resolve complex production issues.
- Ensure all incidents are thoroughly documented, including timelines, actions taken, and resolution steps.
- Ensure compliance with MAS, local in-country regulatory process during incident reporting requirements
- Facilitate Blameless-incident reviews (BIRs) to identify root causes, preventive measures, and opportunities for process and system improvements.
- Experience in generating management reports on incidents, problem trends, and thematic analysis.
- Champion SRE best practices within incident management, focusing on automation, observability, and proactive problem prevention.
- Collaborate with SRE teams to develop and implement incident response playbooks, runbooks, and automation tools to enhance efficiency and effectiveness.
- Contribute to the definition and monitoring of Service Level Objectives (SLOs) and Service Level Indicators (SLIs) related to incident response and system reliability.
- Operational Excellence and Continuous Improvement:
- Analyze incident trends and data to identify systemic issues and areas for improvement in IT systems and processes.
- Develop and implement strategies to reduce Mean Time To Detect (MTTD), Mean Time To Acknowledge (MTTA), and Mean Time To Restore (MTTR).
- Drive initiatives to enhance operational resilience, stability, and performance across critical applications and infrastructure.
- Contribute to the continuous improvement of ETCC incident management policies, procedures, and tools.
- Stakeholder Management and Communication:
- Build strong relationships with key stakeholders across technology and business units.
- Effectively manage expectations and provide transparent communication during periods of service disruption.
- Represent ETCC SRE Operations in various forums and provide updates on incident management performance and initiatives.
- Experience:
- Minimum of 6-10 years of experience in IT Operations, Incident Management, or Site Reliability Engineering within a large-scale enterprise environment, preferably in the financial services industry.
- Demonstrated experience in leading major incidents and coordinating cross-functional technical teams under pressure.
- Strong understanding of SRE principles and practices.
- Technical Proficiency:
- Solid understanding of IT infrastructure (servers, storage, networking), cloud platforms (e.g., AWS, Azure, GCP), and enterprise applications.
- Familiarity with incident management tools (e.g., ITSM, Remedy, ServiceNow, PagerDuty) and monitoring tools (e.g., Grafana, Splunk, Dynatrace, Elk).
- Experience with scripting and automation (e.g., Python, Shell) is a plus and added advantage.
- Experience in insurance, banking or regulated financial services environment (mandatory)
- Leadership & Communication:
- Excellent leadership, communication, and interpersonal skills, with the ability to influence and collaborate effectively at all levels.
- Proven ability to remain calm and decisive during critical incidents, with strong problem-solving capabilities.
- Exceptional written and verbal communication skills for technical and non-technical audiences.
- Certifications (Good to Have):
- ITIL V3/V4 Foundation or higher certification.
- Relevant SRE or DevOps certifications.
- Proficiency in using GenAI and Microsoft office
Hyderabad - DTI Skyview SEZ
Job:Technology
Schedule:Regular
Employee Status:Full time
- ...global data and technology company,... ...(EGSO), the Senior Enterprise Security Incident Manager (ESIM) serves... ...Incident Commander for significant... ...Fusion Center. As an individual... ...potentially major security... ...Science, Computer Engineering, Information... ...our Careers Site to...SeniorFull timeWork at officeLocal areaRemote workFlexible hoursNight shiftWeekend work
$112.7k - $193.2k
...seeking an experienced Senior Manager to lead enterprise Site Reliability Engineering (SRE), DevOps, IT... ...automation, observability, incident management, and... ...improvement while ensuring technology services meet business... ...Management processesLead major incident management...SeniorMinimum wageFull timeWork experience placementLocal areaRemote work- ...intelligence-driven technology services... ...professional and managed services across... ...Intelligence, and Enterprise Service... ...Overview: The Senior Site Reliability Engineer is a technical... ...system scaling. Incident Command: Act as the Incident... ...Commander for major system outages,...SeniorContract workRemote work
$166k - $220k
...is a defense technology company with a... ...realtime, 3D command and control center. As the world... ...mission-driven Site Reliability Engineer (SRE) to join... ...ll Do Manage and expand specialized... ...and associated tools. Experience... ...in the majority of full time offers...SeniorFull timeWork experience placementImmediate startRemote work- ...leader of our Majors Air sales... ...response technology.Axon is rapidly... ...and command centers. By delivering... ...team of ~11 enterprise-level Account... ...-response incident awareness.... ...including product, engineering, strategy,... ...experience managing sales teams... ...conditions associated with this...SuggestedImmediate startRemote work
- NBCUniversal is seeking a Senior Major Incident Manager to own global Major Incident Management within Enterprise Technology. You will lead governance, oversee services, and drive automation in ITSM, partnering with Change, Problem, and AIOps to improve incident delivery...Remote job
$80 - $90 per hour
...hiring for a Senior Site Reliability Engineer (Storage Platforms... ...-focused, enterprise-leading... ...dedicated team managing software-defined... ...Provide incident response and... ..., Marketing, Technology, Supply Chain... ...Cycle, Call Center, Human Resources... ...awards from major publications...SeniorHourly payContract workTemporary workRemote work$102.69k - $287.49k
...rock-solid reliability is fundamental... ...experienced Site Reliability Engineers at the Senior level and... ...maintaining enterprise-grade security... ...error budget management Help... ...pipelines Lead incident response... ...cutting-edge technology company building... ...across all major platforms...SeniorFull timeRemote workWorldwideFlexible hours$132k - $170k
...the role: The Senior Account Executive... ...an effective, enterprise-wide strategy... ...do:Account management with an outcome... ...territory of major client accountsMastery... ...in high technology (services,... ...grown to 20,000 associates globally who support... ...at the very center of the AI...SeniorFull timeContract workRemote workWorldwide$105k
...enables global technology companies to push... ...immediate opening for a Senior Account Manager with a proven... .... As a Senior Enterprise Account Manager,... ...across product, engineering, community... ...across more than 50 major markets in the... ...Washington Post, and the Associated Press. Forbes...SeniorHourly payLocal areaImmediate startShift work$110k - $130k
...connect technologies to help protect... ...Command Center Software... ...Regional, and major Municipal... ...Records Management, Crime Analysis... ...a Senior Federal Software... ...-scale, enterprise software... ...customer sites across... ...the risks associated with a... ...Science, Engineering, IT, or Business...SeniorFull timeFor subcontractorRemote workRelocation$166k - $220k
...is a defense technology company with a... ...realtime, 3D command and control center. As the world... ...looking for a Site Reliability Engineer to join the Imaging... ...service management and observability... ...Familiarity with incident tooling (... ...included in the majority of full time offers...Full timeWork experience placementImmediate startRemote workWeekend workDay shift- ...the world's largest enterprises. We solve real... ...with cutting-edge technology and a strong sense... ...and impact-focused Senior Account Executive... ...Partner with Solutions Engineers to run workshops/... .... ~ Ability to manage complex, multi-... ...German (or another major European language...SeniorLive inWork at officeRemote work
- ...Description We are seeking a senior technology leader, program manager, or technical advisor... ..., organizations, or enterprise services. Relevant experience... ...deep expertise within a major technical domain,... ...administration ~Software engineering ~Systems integration...SeniorTemporary workFor contractorsFor subcontractorRemote work
$90k - $180k
...portfolio of life-changing technologies spans the spectrum... ...About the RoleThis Senior Site Reliability Engineer position works on-... ...the Cardiac Rhythm Management Division.We are... ...following incidents and drive systematic... ...Monitor, or similar enterprise-grade monitoring and...SeniorRemote work$250k
...communities recover after major catastrophes. For... ...platform, and enterprise deals ranging... ...leaders, technology stakeholders, and... ...relationships with senior stakeholders across... ...Sales, Product, Engineering, and Leadership.... ...Talent Partner Manager interview with our...SeniorWork at officeImmediate startRemote workFlexible hoursNight shift2 days per week3 days per week$158.5k - $172k
...innovative restaurant technology, easy-to-use... ...a Senior Engineer on the Runtime... ...responsible for managing our centralized Enterprise Logging Platform... ...driving continuous reliability, deep system... ...as an incident responder, leading... ...Expertise: Solid command over administering...SeniorFull timeTemporary workWork at officeFlexible hours3 days per week$118.6k - $195.68k
...looking for a Senior Site Reliability Engineer (SRE) to design... ...to Red Hat IT managed cloud platform... ...using the latest technologies from Red Hat and... ...This empowers our associates to focus on... ...tenantsDrive sustainable incident response and... ...provider of enterprise open source...SeniorPermanent employmentFull timeContract workWork experience placementWork at officeRemote workFlexible hours$147.1k - $167.9k
Senior Software Engineer, Full Stack (Enterprise Platform Technology) Do you love building and... ...forefront of driving a major transformation... ...digital product managers, and deliver... ...to the pay range associated with that location... ...available through this site. Capital...SeniorFull timePart timeInternshipH1bLocal area- ...Resilience Engineering is a subset of the Site Reliability Engineering... ...improvement through incident analysis,... ...make our technology more reliable... ...how we manage unexpected outages... ...or Incident Commander) rotation... ...resolution of major incidents... ...Applicants. Seniority level Mid‑Senior...Full timeSummer workWork at officeLocal areaRemote workMonday to FridayShift work
- ...Major Incident Manager The Major Incident Manager is a critical technology leadership role responsible for ensuring the swift, coordinated restoration of essential... ...the heart of Digital Services, ensuring the reliability and resilience of critical technology systems...Immediate startRemote work
$90k - $128.5k
...environments. This role engineers and governs... ...Vulnerability Management · Research... ..., including incident response, root... ...with enterprise monitoring to... ...Partner with technology team to execute... ...readiness, and major OS releases delivered... ...periodic on-site participation...SeniorLocal areaImmediate start$149.4k - $202k
...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA... ...discipline at Noctua Technology, LLC is a strategic... ...Our SREs don’t just manage infrastructure; they... ...system stability and incident response. This role... ...and championing major automation projects...SeniorRemote work$141.8k - $195k
...world's biggest enterprises, including half... ...flexibility to manage and analyze... ...Inc is seeking a Senior Site Reliability Engineer to join our mission... ...changing a technology. But here at Cribl... ...truly at the center of the wheel helping... ...sustainable incident response in a...SeniorTemporary workRemote work- ...Senior Site Reliability Engineer As a Senior Site Reliability... ..., strengthen incident response capabilities... ...a global technology organization.... ...role supporting enterprise-scale platforms... ...knowledge of incident management, root cause... ...SciAps is the Center of Excellence for...SeniorRemote work
$185k - $227k
...Senior Site Reliability Engineer Remote - United States; United... ...by leading technology investors, we are... ..., product managers, operations experts... ...point for critical incidents to ensure the... ..., and maintain enterprise-scale Nutanix... ...Session Manager, Run Command, State Manager,...SeniorRemote work$96k - $163k
...and accessible. Our technology and innovation,... ...Title and Summary Senior Site Reliability Engineer, Performance Engineering... ...and lifecycle management, leverages AI-driven... ...Review production incidents, identify contributing... ...and alignment with enterprise security and compliance...SeniorFull timePart timeWorldwideFlexible hours- A leading technology company is seeking a Customer Success Manager to enhance relationships with enterprise customers. This remote position involves guiding customers to realize business value, managing a team, and driving strategic initiatives. The ideal candidate will...SeniorRemote job
- ...Information Technology Division (ITD... ...seeking a Senior Director of Enterprise Application... ...to lead and manage its application... ...software engineering teams that deliver... ...scalable, reliable software... ...monitoring, and incident response... ...contracts, and associated budgets to...SeniorFull timeWork experience placementWork at officeRemote work2 days per week
- ...Startup Ai Agent Technology We are a fast-growing startup that is helping enterprise businesses to unlock the power of AI to replace work and transform costs.... ...businesses to automate manual operations with zero engineering work or process change - with contractually...SeniorRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Associate, Enterprise Technology Command Center, Major Incident Manager, Site Reliability Engineering. Be the first to apply!



