Site Reliability Engineer Incident Manager
$111.3k - $156.5kNetApp
:
If you run toward knowledge and problem-solving, join us
In a world of cloud complexity, NetApp simplifies. Our customers are looking for a more unified and secure multicloud experience, and we provide the services, infrastructure and expertise they need to achieve it.
If you want to have a real impact, NetApp is the place for you. You'll make a differencewhile still maintaining a healthy work-life balance. Who are we? Forward-thinking technology people with a heart. Join us.
Site Reliability Engineer Incident Manager
Remote - US
Job category: University
Job ID: 126103-en_US
About NetApp
We're forward-thinking technology people with heart. We make our own rules, drive our own opportunities, and try to approach every challenge with fresh eyes. Of course, we can't do it alone. We know when to ask for help, collaborate with others, and partner with smart people. We embrace diversity and openness because it's in our DNA. We push limits and reward great ideas. What is your great idea?
"At NetApp, we fully embrace and advance a diverse, inclusive global workforce with a culture of belonging that leverages the backgrounds and perspectives of all employees, customers, partners, and communities to foster a higher performing organization." -George Kurian, CEO
Job summary
At NetApp, we have an amazing opportunity to transform industries with cutting-edge data services. We are developing a broad portfolio of solutions that harness the power of data. NetApp public/private cloud offerings, high performance & highly scalable storage products are stretching what's possible for our customers & partners. These solutions change how data is stored, consumed & interpreted unleashing innovation. Join our team & push the boundaries of what's possible.
The SRE team within NetApp Public Cloud Service group is responsible for the scaling and support of our multi-region, multi-cloud application. This team is made up of a group of software engineers, site reliability engineers, and security experts that own the deployment architecture and strive to improve our infrastructure through deep partnership with other teams across the organization.
We are a customer focused team continuously improving our services to meet their needs and reduce toil on our team. We meet regularly to mentor and challenge each other as we collaborate across projects. We're only able to accomplish our mission through diverse teams innovating, together.
As an SRE Incident Manager, you will work in a command-and-control role focusing on uptime and Time to Recovery (TTR). To fulfill this role, you will collaborate with multiple NetApp teams to bring about safe and rapid mitigation of incidents that impact customers on a global scale. You will partner with Site Reliability Engineers (SREs) and lead by example by being more of a contributor than a delegator.
Job requirements
- Define and refine incident management, change management and problem management-related workflows.
- Reduce operational inefficiencies in the incident management process to ensure the fastest path to SREs through automation and continuous process improvement. Identify when escalation is required and trigger such escalation accordingly.
- Manage proactive notification for planned events and broad communication associated with critical incidents. This requires composure under pressure, broad analytical, and problem-solving expertise, and the ability to confidently collaborate with varied partners. These skills would be applied in producing both written and verbal communication to update customers, partners, and senior leadership.
- Create and maintain recovery playbooks for commonly occurring customer patterns and issues.
- Drive down resolution times by improving alert coverage and accuracy. Deflect customer incident submission by promoting supportability tools (e.g. documentation, self-service workflows).
- Lead after action reviews and root cause analysis. Complete postmortems on a timely basis that identify repair items preventing future customer impact. Ensure resolution of product/service defects, process improvements and documentation enhancement to address live site or customer reported incidents.
- Present monthly incident availability and operability metrics to cross functional leadership teams. Build dashboards to provide insights and visibility into critical business metrics for a variety of audiences.
- You must be able to work outside of normal business hours (weekend shifts, holidays, & evenings) as needed.
Education & experience
- Typically requires a minimum 2 years of related experience with a bachelor's degree; or 2 years and a master's degree; or equivalent work experience.
- Experience managing incidents and running incident management programs, preferably in large-scale environments.
- Experience working with service owners running a DevOps team in public cloud platforms such as AWS, Azure, or Google Cloud is big plus.
- A basic understanding of public cloud vendors such as AWS, Azure, Google Cloud, or others.
- Outstanding communication and presentation skills, written and verbal. Excellent listening skills and a high degree of empathy.
- You are great at solving problems, sorting meaningful information from noise, and taking action.
Equal Opportunity Employer:
NetApp is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination based on age, race, color, gender, sexual orientation, gender identity, national origin, religion, disability or genetic information, pregnancy, protected veteran status, and any other protected classification.
Did you know Statistics show women apply to jobs only when they're 100% qualified. But no one is 100% qualified. We encourage you to shift the trend and apply anyway! We look forward to hearing from you.
Why NetApp?
In a world full of generalists, NetApp is a specialist. No one knows how to elevate the world's biggest clouds like NetApp. We are data-driven and empowered to innovate. Trust, integrity, and teamwork all combine to make a difference for our customers, partners, and communities. We expect a healthy work-life balance. Our volunteer time off program is best in class, offering employees 40 hours of paid time off per year to volunteer with their favorite organizations. We provide comprehensive medical, dental, wellness, and vision plans for you and your family. We offer educational assistance, legal services, and access to discounts. We also offer financial savings programs to help you plan for your future. If you run toward knowledge and problem-solving, join us.
USA and Canada Residents Only:
The base salary hiring wage range for this position which the Company reasonably and in good faith expects to pay for the position in the specified geographic areas or locations, is $111,300 - $156,500. Final compensation will be dependent on various factors relevant to the position and candidate such as geographical location, candidate qualifications, certifications, relevant job-related work experience, education, skillset and other relevant business and organizational factors, consistent with applicable law. In addition, the position may include some of the following comprehensive benefits such Medical, Dental, Vision, Life, 401(K), Paid Time off (PTO), sick time, leave of absence as per the FMLA and other relevant leave laws, Company bonus/commission, employee stock purchase plan, and/or restricted stocks (RSU's).
$160k - $185k
...the way. And we’re just getting started!OverviewThe Sr. Manager, Site Reliability Engineering (SRE) leads the strategy, execution, and continuous... ...at scale. This position will lead teams responsible for incident management, observability, platform reliability, and end...SuggestedWork at officeLocal areaRemote workWork from home$221.2k - $387.1k
...DescriptionIt all started when engineer Fred Luddy wrote code that... ....Job DescriptionTeam:Our Site Reliability Engineering (SRE) team consists... ...to reduce the number of incidents and minimize Mean Time to Recovery... ...organization of engineering managers, technical leaders, and SREs...SuggestedWork at officeImmediate startRemote workFlexible hours- ...Job Description Job Description We are hiring Site Reliability Engineer Manager- Hybrid for a Contract To Hire position in santa clara, CA... ...• Serve as the senior technical escalation for critical incidents guiding cross-team triage, driving RCA, and ensuring systemic...SuggestedContract workRemote work
- ...engagement solutions. We partner with managed care organizations to provide... ...US-based candidates only) Manager, Site Reliability Engineering (SRE) Position Overview We are seeking... ...while actively participating in major incident response, reliability initiatives, and...SuggestedFull timeRemote workFlexible hours
$140k - $230k
...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience... ...automation initiatives. Lead incident resolution: You will conduct thorough... ...with a strong, objective background in managing large-scale distributed systems....SuggestedFull time$104.9k - $174.7k
...Credit Risk mitigation and Customer Data Management. You can learn more about LexisNexis... ...Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and... ..., and responding to real production incidents.If you live near one of our offices,...Full timeWork at officeLocal areaRemote workWork from home- ...platforms to advanced release engineering practices, our teams are... ...Subscriptions, Resource Groups and Management GroupsStrong automation... ...The Role: The Site Reliability Engineer under the general... ...through proper response to incidents and requestsWork with management...Work experience placementH1bWork at officeRemote workVisa sponsorshipFlexible hoursShift work2 days per week
$115.5k - $164.8k
...you matter.Your ImpactAs an engineer on the APX SRE CloudOps team,... ...required human intervention with reliable, tested automation. You will... ...in on-call rotations and incident response. Understanding production... ...experience.Experience managing cloud platforms such as Azure...Work experience placementWork at officeRemote work$105.6k - $145.2k
Architect the Future as our Site Reliability Engineer!Are you ready to take your skills to the next level... ...alerting.Perform code deployments and manage CI/CD pipelines using Azure DevOps,... ...ensuring best practices are followed.Lead incident response efforts and conduct deep-dive...Ongoing contractFull timeWork at officeLocal areaWorldwide$174k - $252k
...pushing for changes that improve reliability and velocity.Practice sustainable incident response and blameless postmortems... ...’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...Computer Science or Engineering.Site Reliability Engineering (SRE) is...- ...at an MUFG office or client sites four days per week and work... ...motivated Certified Sr. Cloud Site Reliability Engineer to build a robust, scalable,... ...environments for security incidents and ensuring rapid response... ...mechanisms.Develop and manage infrastructure as code using...Full timeWork at officeLocal areaRemote work
$108.08k - $172.5k
Work with development and platform engineering teams to migrate and maintain applications in Google Cloud. Apply Observability concepts... ...-call rotation support for production systems, facilitate incident management and conduct post-incident reviews. Drive, contribute and...Full timeRemote workWorldwide- ...apply now.We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (... ...infrastructure provisioning, configuration management, and operational workflows.Support... ...practices including observability, incident management, capacity planning, and...Temporary workWork at officeRemote workFlexible hours
- ...ServicesSelling Points Contribute to the reliability of a high-transaction... ...Reliability Engineer OverviewThe Site Reliability Engineer ensures... ...improvements.Participate in incident response, maintaining runbooks... ...to reduce operational toil.Manage RDS SQL Server deployments,...Remote work
$102.1k - $202.2k
...yearEmployment type: Full-TimeWork site: 3 days / week in-... ...: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft... ..., and easier to manage for customers... ...will collaborate with engineers across disciplines to... ..., assisting with incident response, and building...Ongoing contractWork experience placementLocal areaRemote work3 days per week- ...SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX... ...skilled Senior Site Reliability Engineer (SRE) to join our dynamic... ...or other scripting languages.Manage and troubleshoot Kubernetes... ...CloudFormation.Participate in incident response and post mortem...Remote work
$130k - $150k
...for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business... ...observability, and strengthen incident response. The SRE partners closely with... ...VMware vSphere, VMware Site Recovery Manager (SRM), SAN technologies, and the Rubrik...Work at officeWork from home3 days per week- ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...for system deployment, management and maintenance.What You’ll... ...degradation, recovery, resizing, and incident response using fleet... ...services, workloads, and platform reliability.You6+ years of experience in...Work at officeLocal areaWork from homeFlexible hours
- The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve... ...You will also play a key role in responding to production incidents, troubleshooting complex issues across the infrastructure...Full timeWork experience placementRemote work
- ...is currently Tuesday.Engineering at Lambda is responsible... ...system deployment, management and maintenance.What You... ...teams to improve service reliability and deployment... ...through observability, incident management, capacity planning... ...of experience in Site Reliability Engineering...Work at officeLocal areaWork from homeFlexible hours
$118.6k - $195.68k
...OpenShift team is looking for a Senior Site Reliability Engineer (SRE) to design, develop, scale, and... ...scale which are unique to Red Hat IT managed cloud platform services, while using... ...needs of our tenantsDrive sustainable incident response and lead blameless...Permanent employmentFull timeContract workWork experience placementWork at officeRemote workFlexible hours$130k - $180k
...collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer... ...a key voice in observability, change management, and service scalability, providing... ...during critical events. Leading incident management and post-incident...Work at officeLocal areaRemote workWorldwideMonday to FridayFlexible hours$139k - $257.55k
...organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through... ...operational ownership. On-call and incident response come with the territory, but... ..., and golden-image lifecycle management across the fleet — triage, remediate,...Full timeTemporary workLocal areaRemote workWorldwide- Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and... ...reliability initiatives, and manage operational risk.Design and maintain... ...failover, and disaster recovery.Lead incident response efforts, act as an escalation...Local areaRemote workFlexible hoursShift work
$72.8k - $130k
...technology systems in accordance with modern design standardsThe Site Reliability Engineer will architect, develop, and maintain Optum Serve's... ...with centralized logging, monitoring, metrics, and incident management systemsConfigure and maintain observability tools (dashboards...Minimum wageFull timeWork experience placementWork at officeLocal areaRemote work- Reliability Engineering Design, implement, and operate scalable, resilient, and highly available... ...software delivery.Observability and Incident Management Develop actionable alerts that identify... ...or more years of experience in Site Reliability Engineering, platform engineering...Remote work
$96.8k - $145.2k
...apply now.We are currently seeking a Site Reliability Engineer (Onsite Hybrid) to join our team in Plano... ...Job Responsibilities Include: Own and manage observability using New Relic (APM,... ...SLIs/SLOs and alerting strategiesDrive incident response, root-cause analysis (RCA), and...Full timeTemporary workWork at officeRemote workFlexible hours$62 - $80 per hour
...LocalContract$62/hr - $80/hrA senior Site Reliability Engineer will join an established infrastructure... ...provisioning and lifecycle management. The position blends hands-on production... ...infrastructure, including participation in shared incident response or on-call rotations.Desired...Full timeTemporary workRemote workFlexible hours- Edmond, OKYouVersion - YouVersion Engineering /Full-Time/ Salary /On-siteThe YouVersion Senior Site Reliability Engineer is responsible for ensuring the integrity, performance... ...that objectives are met.Respond and initiate incident responses for high-severity site reliability...Full timeContract workTemporary workWork experience placementCasual workInternshipLocal areaWorldwide
$150k - $180k
...seeking an experienced SeniorSite Reliability Engineer to help design, build,... ...organization.This position is based on-site in either our Arlington, VA... ...monitoring and effective incident response.Develop and promote... ...-functional teams, product managers, and stakeholders to align...Permanent employmentFull timeWork at officeLocal areaRemote workWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer Incident Manager. Be the first to apply!
- site reliability engineer sre Remote
- site reliability engineer Remote
- site reliability engineer remote Remote
- IT site lead Remote
- site safety Remote
- website content developer Remote
- site leader Remote
- on-site clinical research associate (traveling/remote) Remote
- junior website developer Remote
- historic site Remote


