Senior Staff Engineer - SRE - Incident Prevention / Post Incident Correction of Errors
$110k - $260kGEICO Insurance Agent
Why Join GEICO?
At GEICO, we offer a rewarding career where your ambitions are met with endless possibilities.
Every day we honor our iconic brand by offering quality coverage to millions of customers and being there when they need us most. We thrive on relentless innovation to exceed our customers' expectations while making a real impact on local communities nationwide.
Founded in 1936, GEICO is a member of the Berkshire Hathaway family of companies and one of the largest auto insurers in the United States. When you join our company, we want you to feel valued, supported, and proud to work here. That's why we offer the GEICO Pledge: Great Company, Great Culture, Great Rewards, and Great Careers.
Senior Staff Engineer - SRE - Incident Prevention / Post Incident Correction of Errors
Position Summary
Correction or Error’s / Post Incident Review is at the core of GEICO’s culture for improving the reliability and availability of our applications. We are rethinking how we do CoEs at scale and the tooling we need to enable GEICO’s engineering teams to learn from the incidents and identify and eliminate patterns leading to them.
GEICO is seeking an experienced SRE Software Engineer with a passion for building high-performance, low-maintenance, zero-downtime complex distributed platforms and applications. You will help drive our transformation to a tech organization with engineering excellence and site reliability as its mission, while co-creating a culture of psychological safety and continuous improvement.
This role focuses on improving the Correction of Error (COE) tooling and process across GEICO. It is a hands-on technical leadership role focused on better COEs, stronger root cause analysis, fewer repeat incidents, and building the tools and automation that help all GEICO engineering teams learn from incidents and act on that learning.
Position Description
Success in this role requires strong technical depth and equally strong process leadership. The right candidate can go deep on incident analysis, system behavior, and architecture, while also improving how teams run COEs, learn from incidents, and turn those lessons into engineering improvements.
Why This Role Is Different
- You are shaping how GEICO learns from incidents and turns that learning into better engineering practices.
- Your work directly impacts root cause quality, repeat incidents, availability, MTTR, and operational confidence.
- This role blends deep technical understanding, hands on execution and ability to build SW with real time incident leadership, platform and process
Position Responsibilities
As a Senior Staff Engineer, you will:
- Design, develop and operate automation, self-service tools, dashboards, and data pipelines that automate and scale COE workflows and reduce manual tracking and follow-ups.
- Run and moderate weekly GEICO-wide COE presentation sessions for qualified high-severity incidents, ensuring the right owners are prepared to present.
- Lead and improve the COE process across the engineering organization.
- Provide technical leadership in system design, architecture, and hands-on engineering for COE improvements, automation, and incident tooling.
- Coach application engineering teams so they can identify the true root cause, explain incidents clearly, and produce clear, complete, high-quality COEs.
- Provide technical leadership for root cause analysis across distributed systems using logs, metrics, traces, and observability data.
- Identify gaps exposed by incidents, such as missing alerts, weak monitoring, incomplete runbooks, poor testing, or bypassed deployment controls.
- Look across incidents to find repeat patterns, connect lessons across COEs, and share insights that improve prevention and reliability.
- Partner with application engineering teams, platform teams, SRE, and operational stakeholders to drive accountability and continuous improvement, by driving improvements in their technology strategy and roadmaps.
We have adopted “You Build it You Run it strategy.
- All our senior technologists take an active role in leading and managing high-severity incidents requiring strong technical judgment, clear communication, and calm execution under pressure.
- All our engineers have on-call responsibilities as part of a 24x7 rotation supporting incident response and production support for mission-critical platforms and processes they build and operate.
Qualifications
- Hands on proficiency in multiple languages, including Go, Java, Python, C# for building production-grade full stack applications on Kubernetes and serverless technologies (KNative) in Azure and AWS.
- Experience with SQL and NoSQL technologies
- Experience with building and using data pipelines analytics, and dashboards for operational metrics, trends, and KPIs using technologies such as Spark, Trino, Grafana, Superset, PowerBI
- Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, Azure Monitor.
- Experience with incident management platforms such as PagerDuty.
- Proficiency with AI assisted development processes and tools such as Claude Code, Cursor and GitHub Copilot.
- Experience improving incident, post incident review, or reliability processes at scale through automation, data, and cross-team influence.
- Deep incident forensics and root cause analysis skills, with the ability to raise COE quality through clear action items and follow-through across teams.
- Strong understanding of observability, reliability engineering, incident management, and post-incident improvement practices.
- Experience supporting incident response and high-severity production incidents in complex environments.
- Strong software engineering fundamentals and system design skills, with experience building reliable production systems at scale.
- Ability to lead technical design and architecture decisions in complex distributed systems.
- Strong communication skills and the ability to coach engineering teams and present findings clearly to leadership.
Experience
- 10+ years of professional software engineering experience, preferably in platform engineering, reliability engineering, backend engineering, distributed systems, or operational tooling.
- 8+ years of experience with architecture, design, system reliability, scalability, and technical leadership for production systems.
- 6+ years of experience with open-source frameworks, modern engineering practices, or platform technologies.
- 4+ years of experience with Azure, AWS, GCP, or another cloud service provider, or equivalent experience in complex hybrid environments.
- Demonstrated ownership of mission-critical systems operating in 24x7 production environments.
Education
- Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience.
Additional Job Requirements
- Ability to influence engineering outcomes across teams in complex organizations.
- Must be able to communicate in a clear, concise, professional oral and written manner with customers, clients, co-workers, leadership, and other employees of the organization.
- Must be able to perform effectively under pressure and in stressful situations, including during production support and high-severity incident response.
- Must be able to participate in a 24x7 on-call rotation for incident response and production support of mission-critical platforms.
Annual Salary
$110,000.00 - $260,000.00The above annual salary range is a general guideline. Multiple factors are taken into consideration to arrive at the final hourly rate/ annual salary to be offered to the selected candidate. Factors include, but are not limited to, the scope and responsibilities of the role, the selected candidate’s work experience, education and training, the work location as well as market and business considerations.
At this time, GEICO will not sponsor a new applicant for employment authorization for this position.The GEICO Pledge:
Great Company: Protecting customers through life’s twists and turns with innovation and integrity.
Great Careers: Personalized development programs, mentorship, and certification assistance.
Great Culture: Inclusive and collaborative culture rooted in shared success.
Great Rewards: Competitive pay, benefits, and flexibility to support your well-being and future.
The equal employment opportunity policy of the GEICO Companies provides for a fair and equal employment opportunity for all associates and job applicants regardless of race, color, religious creed, national origin, ancestry, age, gender, pregnancy, sexual orientation, gender identity, marital status, familial status, disability or genetic information, in compliance with applicable federal, state and local law. GEICO hires and promotes individuals solely on the basis of their qualifications for the job to be filled.
GEICO reasonably accommodates qualified individuals with disabilities to enable them to receive equal employment opportunity and/or perform the essential functions of the job, unless the accommodation would impose an undue hardship to the Company. This applies to all applicants and associates. GEICO also provides a work environment in which each associate is able to be productive and work to the best of their ability. We do not condone or tolerate an atmosphere of intimidation or harassment. We expect and require the cooperation of all associates in maintaining an atmosphere free from discrimination and harassment with mutual respect by and for all associates and applicants.
- Box is seeking a Senior Technical Duty Officer (Senior Incident Commander) to join the GTOC team in Redwood City, CA. You will lead high-severity incidents... ...critical services. You will collaborate with SRE and engineering teams, drive automation, and mentor others while...Senior
- Box is seeking a Senior Technical Duty Officer to lead live-site incidents and drive rapid resolution for critical cloud services in Redwood City, CA. You will own... ...bridges, design automated tooling, and partner with SRE teams to reduce MTTR while improving observability...Senior
$105k
...Quality Assurance; Engineering / Science;... ...including through incident reporting and investigations... ...Power Generation Corrective Action Program (... ...corrective or preventive measures to... ...time of the job posting. This compensation... ...consensus with multiple senior leaders to stay...SeniorWork experience placementWork at officeRemote workFlexible hours$140k
...Quality Assurance; Engineering / Science;... ...including through incident reporting and investigations... ...Power Generation Corrective Action Program (... ...corrective or preventive measures to... ...time of the job posting. This compensation... ...gaining consensus with senior and executive...SuggestedWork experience placementWork at officeRemote workFlexible hours- ...of San Mateo seeks an IS Manager II - Security to lead a critical cybersecurity program and a team of engineers. You will shape security architecture, drive incident response, and ensure compliance with industry standards. You will manage a supervisor and four security...Senior
$160k - $200k
An innovative software services startup in Redwood City is seeking an Incident Response Lead to tackle complex problems and engage directly with customers. You will lead investigations, develop playbooks, and ensure compliance with SLAs. The ideal candidate has 6+ years...Senior- ...HIRING: Senior / Staff Full Stack Engineer – AI Location: Mountain View, CA – 3 Days Hybrid ⏰ Interview Panel Availability: 7 AM – 12 Noon CST... ...✅ Production Observability – Logs, Metrics, Alerts & Incident Response ✅ Distributed Backend Services & Microservices...SeniorTemporary work
$175.75k - $260k
...possible. As a Senior Lead Software Engineer at JPMorgan Chase... ...Management (SRE) and Central SRE... ...RCA generation, incident analysis, log analytics... ...analysis, and corrective-action... ...triage through prevention. ~ Proven ability... ...reliability metrics, error budget...SeniorFull time$160k - $200k
...ownership and demonstrate a drive to tackle complex problems, conduct thorough analysis, work with AI workflows, and effectively triage incidents. The role will involve direct, hands‑on engagement with customers to spearhead the response and resolution efforts for critical...Senior- ...with Moveworks’ Reasoning Engine and natural language... ...looking for a hands-on Staff Engineer who can move machine... ..., SLOs, alerts, and error budgets across... ..., deployment, on-call, incident response, blameless postmortems... ...platform engineering, SRE, production engineering...Full timeWork at officeImmediate startRemote workFlexible hours
$130k - $220k
Santa Clara, CASoftware Engineering - Motion Planning /Full-time /HybridPlusAI is a Physical... ...improve performance through systematic error analysis and targeted experimentation.Collaborate... ...testing, and monitoring to ensure correctness, safety, and production reliability....SeniorFull time$298k - $350k
...build, evaluate, and improve their own products. As a Senior Staff Machine Learning Engineer, you will define and uphold the quality bar for ML... ...engineering teams to ensure systems meet clear standards for correctness, safety, latency, and user satisfaction. Your work...SeniorWork at officeFlexible hours3 days per week- ...-Meta product and engineering leaders, we've raised... ...'re looking for a Senior Software Engineer,... ...need a seasoned SRE to help us scale... ...drive SLOs, SLIs, and error budgets, and build... ...support them Lead incident response and... ...improvements that prevent recurrence Improve...SeniorWork at officeRemote workFlexible hoursShift work
- ...We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority... ...and operationalise SLIs, SLOs, and error budgets. Strengthen observability across... ..., and predictability of SLAs. Lead incident response for complex cross-system...Shift work
- ...will work with global PL marketing and engineering teams to drive new product introductory.... ...occasionally Support product quality incidents and coordinate engineering analysis occasionally... ...category protected by federal, state or local law. Job Posted by ApplicantPro...SeniorLocal area
$225k - $275k
Senior Staff Network Deployment Engineer Crusoe Cloud is seeking a Senior Staff Network Deployment Engineer to serve as the technical owner of how... ...Senior network deployment engineers. Lead design reviews, post-incident reviews, and drive systemic improvements to...SeniorTemporary workRemote work- ...You will be part of our Global Network engineering team responsible for architecting, implementing... ...Matter Expert to respond to network incidents/service requests according to Service... ...deadline will be 10/22/2026 (3 months from posting), although we reserve the right to close...SeniorTemporary workLocal areaRemote workFlexible hoursShift work
$152k - $241.5k
...workloads. We are looking for Software Engineers with SRE or Production Engineering experience who... ....Take part in on-call duties, incident response, root-cause analysis, and follow... ...accepted at least until October 3, 2026.This posting is for an existing vacancy. NVIDIA uses...SeniorPermanent employmentFull time- ...Senior Principal Software Engineer - Fraud Prevention Engineering Job Information Job Identification 210791287 Job Category Software Engineering Business Unit Commercial & Investment Bank Posting Date 09/17/2026, 05:44 PM Locations 3223 Hanover St, Palo...SeniorFull timeFor contractors
- ...influential companies. As a Senior Principal Software Engineer at JPMorganChase within... ...Bank Trust & Safety Fraud Prevention team, you provide deep... ...release readiness gating, incident triage/root-cause acceleration... ...platform engineering, and SRE teams to productionize...Senior
- ...Synopsys is the leader in engineering solutions from silicon... ..., and what would prevent the same issue from resurfacing... ...planning, and incident response. Automate... ...DevOps, Cloud Engineering, SRE, Platform Engineering,... ...to drive lasting corrective actions. You communicate...SeniorWork at officeRelocation
$130k - $260k
...offer the GEICO Pledge: Great Company, Great Culture, Great Rewards, and Great Careers. The Role: We are seeking a Senior Staff Software Engineer to provide technical leadership for critical software development systems while establishing a strong culture of...SeniorHourly payFull timeWork experience placementLocal area$182k - $242k
...We're seeking a Senior Technical... ...failure rates, MTTR, incident trends,... ...reliability OKRs across engineering and operations... ...Use data and post-incident learnings... ...and drive corrective action Minimum... ..., hardware, or SRE) ~ Demonstrated... ...as SLIs/SLOs, error budgets, incident...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$115k - $230k
...Position Summary GEICO's Pulse team is looking for a Staff Software Engineer who is passionate about building data platforms that give engineering... ...and adjacent data sources (AI spend, DORA, ADO work items, incident data), ensuring coherent, joined datasets for leadership...Hourly payWork experience placementLocal area- ...technology products. As a Senior Lead Software Engineer at JPMorganChase within... ...acceleration, release readiness, incident/root-cause analysis),... ...validating AI outputs for correctness, performance, and security... ...and how they work in SRE environments to solve reliability...SeniorFor contractors
$155k - $230k
...model or preparing for the post-quantum computing era, we... ...We are looking for a Senior/Staff Infrastructure & Platform Engineer to help architect, build,... ...platform issues, participate in incident response and root-cause... ..., software engineering, SRE, DevOps, or related...SeniorTemporary workH1bWorldwide$186k - $225k
...Proactive Collaboration About the Role Mainspring is seeking a Senior Staff Mechanical Engineer to shape technical direction across multiple teams,... ...limit. Thank You! Does your experience not meet all of our posted requirements? Studies have shown that some people are...SeniorLocal areaFlexible hours$175k - $265k
...possibilities of AI. Role Overview d-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation... ...systems end-to-end, from provisioning through live incident response, partnering with hardware and software teams...SeniorFull time$197k - $285k
...Senior Staff Electrical Engineer Aurora is delivering the benefits of self-driving technology safely, quickly, and broadly to make transportation safer, increasingly accessible, and more reliable and efficient than ever before. Aurora hires talented people with...Senior$197k - $285k
...get crucial goods where they need to go, and make mobility more efficient and accessible for all. We’re searching for a Senior Staff Electrical Engineer.In this role, you willJoin an exceptional team of electrical engineers responsible for the development of electronics...SeniorWork at officeLocal area3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Staff Engineer - SRE - Incident Prevention / Post Incident Correction of Errors. Be the first to apply!
- technology administrator Palo Alto, CA
- assistant engineer Palo Alto, CA
- staff engineer Palo Alto, CA
- senior staff systems engineer Palo Alto, CA
- engineering aide Palo Alto, CA
- senior manager customer operations Palo Alto, CA
- senior software engineer ruby on rails Palo Alto, CA
- sr marketing manager Palo Alto, CA
- senior customer service Palo Alto, CA
- senior business manager Palo Alto, CA




