Incident Management Lead
Forward
About the Role Incidents are inevitable. How fast you detect them, how quickly you act, and whether the organization actually learns from them - that is what separates payments companies that scale from ones that spiral. Forward processes payments for thousands of merchants across dozens of partner platforms. When something breaks - a submission failure blocks merchants mid-onboarding, a processing outage hits a partner's book, a compliance flag freezes accounts at scale - the impact lands on real businesses in real time. The question is not whether incidents will happen. It is whether Forward detects them in minutes or hours, resolves them with coordination or chaos, and fixes the root cause or patches the symptom. The Incident Management Lead owns that answer. This is a modern role for a modern problem. You will build an AI-assisted incident intelligence layer that gives Forward signal before issues become incidents, run coordinated response when they do, and drive the post-incident work that makes the organization genuinely more resilient - not just less embarrassed. You will own the closed loop between reactive resolution and proactive prevention: the governance function that ensures Forward gets faster and smarter after every incident rather than repeating the same failures. This role sits at the intersection of Engineering, GTM, Support, and Operations, and is directly accountable to the CRCO. When things go wrong at Forward, you are the named owner - before, during, and after. Key Responsibilities Build the Detection Layer Design and operate a proactive monitoring and alerting infrastructure: SLO burn-rate alerting, synthetic health checks, deployment risk scoring, and real-time anomaly detection across submission, processing, and compliance pipelines. Build and maintain AI-assisted signal intelligence: use AIOps platforms (PagerDuty,Incident.io, or equivalent) to correlate alerts, suppress noise, and surface high-confidence incident precursors before they manifest as partner escalations. Own the governance loop over Support: review support ticket themes and escalation patterns on a weekly cadence to identify systemic issues before they cross into incident territory. Establish alert-to-noise discipline: define what a true signal looks like for each incident type, tune alerting thresholds, and drive Alert-to-Noise Ratio above 80% - the team acts on signals, not volume. Build and maintain the runbook library: pre-written, AI-augmented playbooks for the most common incident classes - submission failures, processing outages, ACH return spikes, TM system failures, compliance freezes - so the first 15 minutes of every incident are not spent figuring out who does what. Run Incident Response Serve as the named incident owner when an incident is declared - responsible for coordinating Engineering, Support, GTM, and Operations from detection through resolution. Declare incidents using a consistent severity framework (Sev-1 through Sev-3) with defined, documented SLAs for each tier. Drive MTTD (Mean Time to Detect) and MTTA (Mean Time to Acknowledge) toward P1 targets: detection under 5 minutes, acknowledgment under 15 minutes for Sev-1 and Sev-2. Manage communications during incidents on a defined cadence: internal stakeholder updates, partner-facing status, and merchant-level communications where required - proactive, not reactive. Classify merchant and partner impact in real time: GPV-at-risk, number of affected merchants, partner SLA implications, and any regulatory reporting obligations under DORA or card network rules. Use AI-assisted investigation tooling to compress diagnosis time: automated root cause hypothesis generation, timeline reconstruction, and runbook suggestion reduce the first 30 minutes of investigation to seconds. Drive Post‑Incident Learning Own the post-incident review (PIR) for every Sev-1 and Sev-2: complete root cause analysis, contributing factor mapping, timeline reconstruction, and remediation item ownership - delivered within 48 hours. Track remediation commitments to closure - architectural fixes, tooling gaps, process changes, and partner education. Not just documented: done. Verified. Closed. Produce partner-facing incident summaries for high-impact events: clear, factual, and accountable. Where DORA or card network reporting obligations apply, own those submissions on deadline. Build and maintain the incident knowledge base: a searchable, AI-indexed record of every incident, RCA, and remediation action that the full team can learn from and reference. Track Incident Recurrence Rate as a primary quality signal. A repeated incident is a failed post-incident review. Quantify Partner and Merchant Impact Develop and maintain a GPV-at-risk classification framework: when an incident fires, the team knows immediately which partners and merchants are affected, what volume is at risk per hour, and what the SLA clock looks like. Build per-partner SLA attainment reporting: monthly scorecards showing incident frequency, MTTR, and resolution quality by partner - inputs into GTM conversations and partner health reviews. Own merchant reachability during incidents: ensure that communication channels, escalation paths, and support routing remain operational when the systems they depend on are not. Prepare and maintain required regulatory reporting: DORA major incident classification (4-hour reporting threshold), card network operational incident disclosures, and FCA Operational Resilience documentation where applicable. Build the Tooling and AI Layer Own and evolve the incident management tooling stack: AIOps platform (PagerDuty AIOps,Incident.io, FireHydrant, or equivalent), incident communication tooling, and integration into Forward's engineering observability layer. Build AI-assisted incident workflows: automated triage, runbook execution, stakeholder notification, and post-incident report generation - reducing manual coordination overhead measurably within 90 days. Partner with Engineering to integrate deployment risk scoring and change failure rate monitoring into the release process - so high-risk deployments trigger elevated alerting before they reach production. Maintain SLO dashboards and error budget tracking for Forward's core merchant-facing surfaces: submission flow, payment processing, bank linking, and TM decisioning. Required Qualifications 4+ years in incident management, site reliability engineering, technical program management, or engineering operations - with direct ownership of production incident response. Experience running cross-functional incident response: coordinating Engineering, Support, and business stakeholders under pressure with clear, structured communication. Hands-on experience with modern incident management platforms: PagerDuty,Incident.io, FireHydrant, Rootly, Blameless, or equivalent AIOps tooling. Strong analytical mindset: comfortable with dashboards, error logs, SQL-based data pulls, and identifying patterns across support ticket data to find systemic signals. Strong written communication: produces clear, concise incident summaries, RCAs, and partner communications on a tight timeline. Experience building incident management infrastructure from scratch - severity frameworks, runbooks, post-mortem templates, SLO definitions. Familiarity with SLO/SLA frameworks, error budget concepts, and alert fatigue management. Preferred Qualifications Payments, fintech, or financial services domain experience: processing failures, compliance flags, card network reporting, and partner escalation dynamics. Familiarity with DORA regulatory incident reporting obligations or equivalent financial services operational resilience frameworks. Experience integrating AI tooling into incident workflows: automated RCA, alert correlation, or runbook execution. SQL proficiency for pulling operational data to support incident diagnosis and post-incident analysis. Background in platform engineering or SRE at a payments or financial services company. Experience managing merchant or partner communications during production outages at scale. What Success Looks Like Detection and Response MTTD (Mean Time to Detect): under 5 minutes for Sev-1, under 15 minutes for Sev-2. MTTA (Mean Time to Acknowledge): under 15 minutes for Sev-1, under 30 minutes for Sev-2. MTTR (Mean Time to Resolve): under 1 hour for Sev-1, under 3 hours for Sev-2. Incidents are declared within 15 minutes of first confirmed signal. Quality and Recurrence Incident Recurrence Rate below 10%: repeated incidents are a process failure, not a norm. Change Failure Rate below 5%: incidents caused by deployments trend toward DORA Elite tier. Alert-to-Noise Ratio above 80%: every alert the team responds to is a real signal. Every Sev-1 and Sev-2 has a completed post-incident review within 48 hours. Remediation items close on schedule. Partner and Merchant Impact GPV-at-risk is quantified in real time during every incident. Per-partner MTTR trends down quarter-over-quarter. Partner communications go out within 15 minutes of incident declaration. DORA and card network reporting obligations are met with zero missed deadlines. Prevention Support ticket pattern reviews happen weekly. At least one systemic prevention initiative is in-flight at all times. The incident log shows a declining trend in repeat incident types quarter-over-quarter. AI tooling is embedded in at least two core incident workflows within 90 days, with measurable reduction in coordination overhead. What We Offer Competitive salary and equity package. Comprehensive health, dental, and vision benefits. Flexible work arrangements and generous PTO. Learning & development budget for conferences, courses, and certifications. A direct line to building the incident management function at a high-growth payments company from the ground up. #J-18808-Ljbffr Forward
- Forward in Austin, TX is seeking an Incident Management Lead to design and own an AI-assisted incident intelligence layer, coordinating across Engineering, Support, GTM, and Ops to detect, respond, and prevent incidents fast. You will drive post-incident learning, build...Suggested
$207k - $301k
Google is seeking a seasoned cybersecurity professional for its Incident Response team in Austin, TX. The role involves managing enterprise incident response operations and conducting forensics to combat cybersecurity threats. With a focus on creating a safe environment...Suggested- Fluidstack is seeking a seasoned Incident Response Lead to secure our frontier compute infrastructure across corporate, cloud, and data center environments. You will own end-to-end IR from detection to eradication, drive cross‑team investigations, and define severity models...Suggested
- Q2 is seeking a Senior Fraud Response Analyst to act as the central coordinator for major fraud incidents impacting customers. You will manage end-to-end response, including intake, investigation support, escalation, root cause analysis, and remediation tracking. Partner...Suggested
- Allied Universal is seeking a Global Security Operations Center (GSOC) Shift Lead to oversee day-to-day GSOC operations, manage shift activities, and ensure compliance with laws and client regulations. The role requires coordinating the team, updating schedules, and driving...SuggestedShift workDay shift
- Fluidstack is seeking a senior incident commander to lead 24/7 on-call operations across the US region. You will own the coverage model, rotation discipline, and SLAs, and grow a team of senior responders while coordinating with legal, customer, and executive stakeholders...
- TeleTech Holdings, Inc. is seeking an Incident Response Manager to lead our security operations from a remote location in the United States. You will oversee detection, containment, and remediation of cybersecurity threats while guiding a skilled team of analysts. You’ll...Remote work
- TTEC is seeking an Incident Response Manager to lead the cybersecurity incident response team from a fully remote position in the United States. You will manage detection, containment, and remediation of threats while guiding analysts, developing IR playbooks, and coordinating...Remote job
- A leading entertainment company in Austin, Texas, is looking for a Manager of the Technical Operations Center (TOC) to oversee globally distributed teams. The successful candidate will lead efforts for the operational performance and availability of critical business platforms...
$105k - $125k
A leading data center infrastructure firm is seeking an experienced Environmental, Health & Safety (EHS) Manager in Texas. The role focuses on implementing safety programs on construction sites and ensuring compliance with local regulations. Candidates must have a bachelor...Full timeLocal area- ...(MaxJob DescriptionWe are seeking a Tech Lead for a support related role for deployed ML... ...ensuring timely resolution of support incidents, implementation and deployment of change... ...Databricks for building, deploying, and managing machine learning workflows.Hands on with...Hourly payShift workWeekend work
- ...a global technology company, building the best way to move and manage the world’s money.Min fees. Max ease. Full speed.Whether people... ...etc.Experienced in handling privacy inquiries, complaints and incidents.Experience in privacy by design and data protection impact assessments...Local area
- ...Custodial Lead page is loaded## Custodial Leadlocations: Austin, TXtime type: Full timeposted... ...service, and safety inspections* Report incidents and hazardous conditions to supervisor*... ..., please visit SBM's website at: Management Services, LP and its affiliates are proud...Hourly payImmediate startMonday to FridayShift work
$160k - $225k
The Lead Architect is a hands on technical leader responsible for defining and delivering... ..., resilience, automated recovery, and incident response readiness.· Support Practice Growth... ...Sub, and Cloud Storage.· Experience with managed database options on GCP, including Cloud...Permanent employmentFull timeTemporary workWork experience placementRemote work- ...opportunities, a world-class training facility, and leading market tools, we help our people... ...Lead Specialist, Oracle SCM to join our Managed Services practice.Responsibilities:... ...specialized investigation and diagnosis of all incidents to identify problems and escalate...H1bLocal area
$114.1k - $268.18k
...opportunities, a world-class training facility, and leading market tools, we help our people... ...Specialist, Cloud Security to join our Managed Services practice.Responsibilities:... ...governance managed services, including incident, problem, and service request management...H1bLocal area$79.4k
...Position Overview The Field Office Support Lead manages field IT support operations to ensure end‑user devices, connectivity, and local... ...while serving as the primary escalation point for complex incidents. This leader aligns field practices with enterprise service‑management...Contract workWork experience placementWork at officeLocal areaRemote work- ...To do this, we provide enterprise risk management services and programs specifically designed... .... The Readiness & Resilience Lead, assigned to a specific client, will be... ...provides leadership and coordination during incidents, partners with various stakeholders, and...Work at officeLocal area
- ...full-time team members.Recreation Shift Lead:The Recreation Shift Lead is responsible... ...coordinate staff, events, and activities, manage emergencies, settle disputes, answer inquiries... ...doors, setting alarms, and documenting incidents or concerns.Enforces all Community Center...Full timePart timeSeasonal workFlexible hoursShift workAfternoon shiftWeekday work
- ...Description Total Safety is looking for a Lead Rescue Technician to join their safety... ...systems design, and materials management. Our Core Values are People, Safety & Wellbeing... ...always maintained. Completes daily ICS (Incident Command System) reports. In case of accident...Full timeWork at officeLocal area
- ...Job Description Job Description Position Description: The Lead Front Desk Clerk is responsible for the support and guidance of... ...the lease and community policies, preparing documentation such as incident reports and shift reports, monitoring video surveillance and guest...Work at officeShift work
$16.5 per hour
...consistent with the requirements of the job. RESPONSIBILITIES Management and maintenance of the equipment and supplies used for events... ...to your direct supervisor as soon as possible following an incident resulting in an injury. QUALIFICATIONS Education/...Hourly payPart timeLocal areaImmediate start- Fluidstack is seeking a Construction Project Manager to lead day-to-day site execution for utility-scale solar and BESS projects. You will... ...Primavera P6, MS Project) across multiple scopes, enforce a zero-incident safety program, and identify long-lead equipment risks early...For contractors
- ...a Security Supervisor - Night to oversee overnight safety and security across properties. You will lead Loss Prevention Officers, manage patrols, and respond to incidents while maintaining guest and associate safety. The role requires strong leadership, familiarity with...Night shift
- About Fullscript We’re an industry-leading health technology company on a mission to help people get better. We started in 2011 with one... ...of new workflows as the program scales. Lead privacy incident response activities, including intake, triage, coordination with...
- ROBOTIC PROCESS AUTOMATION LLC is seeking the ID DataWeb Lead to ensure smooth operations, incident resolution, and service continuity for the ID DataWeb platform. You will lead hands-on troubleshooting, platform monitoring, and collaboration with L3 engineering and business...
- Q2 is seeking a Senior Fraud Response Analyst to lead the Fraud Response Team and coordinate major fraud incidents end-to-end, from intake to remediation. You will partner with Product, Security, Engineering, and Customer Support to minimize customer impact and strengthen...
$89k - $100k
Sr Technical Services Team Lead at OTAVA is responsible for overseeing a team of technical... ...professionals who provide high-quality managed cloud and infrastructure services to... ...growth. Manage daily operations, including incident resolution, service request fulfillment,...Flexible hours$15 per hour
...Burlington Stores, Inc. as a Shortage Control Lead ! As a Shortage Control Lead you will be... ..., identifying and reporting theft incidents, and driving shortage education and awareness... ...or repeat theft incidents Support store manager by providing internal controls and...Hourly payFull timeLocal areaFlexible hoursNight shift- Tract Capital Management, LP is seeking a Data Center Construction Security Manager in Austin, TX. This role oversees security operations... ...programs in high-stakes environments, with a strong focus on incident response, risk assessment, and cross-functional collaboration....
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Incident Management Lead. Be the first to apply!
- focus workforce management Austin, TX
- director workforce management Austin, TX
- event management intern Austin, TX
- director of knowledge management Austin, TX
- management fast track program Austin, TX
- materials management associate Austin, TX
- director change management Austin, TX
- director management consulting Austin, TX
- property management specialist Austin, TX
- change management specialist Austin, TX


