Director, Data & Storage Reliability Engineering
$221.2k - $387.1kJobleads-US
Director, Data & Storage Reliability Engineering
- Full-time
- Employee Type: Regular
- Region: AMS - North America and Canada
- Work Persona: Flexible or Remote
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.
Join us to put AI to work for people.
What you get to do in this role:
Team Management
The successful candidate will lead the Data & Storage Reliability Engineering organization responsible for improving reliability, resilience, performance, scalability, observability, and customer experience across ServiceNow's database, storage, and supporting platform infrastructure.
This leader will be responsible for building, developing, and scaling high-performing engineering teams focused on reliability engineering, observability, performance engineering, diagnostics, automation, production analytics, migration readiness, resilience engineering, and prevention engineering.
Responsibilities include talent acquisition, performance management, career development, succession planning, objective setting, coaching, and prioritization of strategic initiatives.
The role will establish a strong engineering-first culture centered on data-driven decision making, continuous improvement, operational excellence, customer experience, and systemic risk reduction.
This position is accountable for identifying recurring failure patterns, reliability risks, performance bottlenecks, scalability constraints, migration challenges, and operational inefficiencies across database services, storage platforms, cloud infrastructure, and distributed application environments, and driving engineering improvements that eliminate entire classes of issues before they impact customers.
The Director will partner closely with Product Engineering, Database Engineering, Cloud Infrastructure, Architecture, Storage Engineering, Support, and Operations teams to ensure reliability, observability, performance, and resilience considerations are incorporated throughout the software development lifecycle.
The successful candidate will also partner closely with SWAT and Customer & Production Engineering teams to establish a continuous feedback loop between production operations and platform improvement. SWAT remains responsible for customer escalations, production operations, incident response, and service restoration, while this organization is responsible for identifying systemic opportunities, defining engineering priorities, and driving platform improvements that reduce future customer impact.
The successful candidate will serve as the senior technical leader for complex reliability investigations, customer-critical escalation reviews, migration readiness assessments, and platform improvement initiatives, transforming production insights into long-term engineering outcomes.
They will influence architectural decisions and technology investments by providing reliability expertise, observability insights, performance guidance, and production-based evidence that improve platform resilience, scalability, efficiency, and customer outcomes.
This role requires a strong product mindset. The leader will treat reliability, observability, resilience, performance, and automation capabilities as products with roadmaps, priorities, adoption goals, and measurable outcomes. They will be responsible for identifying the highest-value engineering opportunities, prioritizing investments, and driving adoption across multiple product and infrastructure organizations.
Process and Procedures
The successful candidate will establish scalable reliability engineering practices, standards, governance processes, and operating models across the organization.
They will drive adoption of observability standards, reliability engineering frameworks, resiliency assessments, migration readiness practices, diagnostics capabilities, engineering guardrails, and automation strategies.
This leader will continuously evaluate incidents, customer escalations, migration outcomes, platform telemetry, performance trends, capacity signals, and operational data to identify systemic risks and drive long-term engineering improvements.
The role will establish a formal review process with SWAT and Customer & Production Engineering teams to evaluate major incidents, recurring operational challenges, migration learnings, customer-impacting events, and emerging platform risks. These insights will be used to prioritize engineering investments and platform improvements.
The successful candidate will establish meaningful KPIs and engineering metrics that provide visibility into platform reliability, resiliency, performance, operational efficiency, customer experience, engineering productivity, and risk reduction.
The successful candidate will leverage AI-powered tools, analytics, automation frameworks, and production intelligence to identify emerging risks, improve detection coverage, accelerate engineering insights, reduce operational toil, and improve engineering productivity.
They will use production telemetry, incident learnings, customer escalations, migration outcomes, observability data, and operational trends to drive architectural improvements, reliability investments, platform standards, and long-term engineering evolution.
The Director will maintain a portfolio of reliability investments spanning observability, performance, diagnostics, resilience, automation, and prevention, balancing immediate customer needs with long-term platform strategy.
The Director will champion a proactive reliability engineering model that shifts the organization from reactive issueresponse toward predictive analysis, prevention, resilience, and continuous optimization.
To be successful in this role you have:
- Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
Strong product mindset with demonstrated experience treating technical capabilities as products with roadmaps, priorities, customers, adoption goals, and measurable business outcomes.
Experience translating production insights, customer pain points, operational challenges, reliability risks, and platform telemetry into prioritized engineering investments and long-term roadmaps.
Experience partnering closely with production operations, customer escalation teams, reliability organizations, and software engineering teams to drive systemic improvements based on operational learnings.
Experience defining product strategies, developing roadmaps, prioritizing investments, and aligning stakeholders across multiple organizations without direct authority.
Experience operating a portfolio of engineering investments, balancing short-term customer needs with long-term reliability, performance, scalability, and resilience objectives.
15+ years of experience in software engineering, platform engineering, reliability engineering, infrastructure engineering, database engineering, distributed systems, product management, or large-scale SaaS environments.
8+ years of engineering leadership experience, including leading managers and globally distributed teams.
Extensive experience leading Reliability Engineering, Platform Engineering, Database Engineering, Infrastructure Engineering, Production Engineering, Performance Engineering, or related technical organizations.
Deep expertise in distributed systems, databases, storage technologies, cloud infrastructure, and large-scale SaaS architectures.
Strong understanding of reliability engineering principles, observability, scalability, resiliency, operational excellence, and performance engineering.
Experience building and operating observability, telemetry, diagnostics, reliability, or performance capabilities at scale.
Proven experience identifying systemic issues and converting operational insights into strategic engineering improvements.
Experience partnering closely with Product Management organizations to influence roadmaps and deliver customer-centric outcomes.
Experience driving engineering initiatives through data, metrics, customer impact analysis, and measurable business outcomes.
Experience leveraging AI technologies to improve decision-making, analytics, engineering workflows, operational efficiency, reliability insights, automation, or customer outcomes.
Exceptional communication, stakeholder management, and leadership skills.
Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Desired Skills
Previous Product Management experience in a platform, infrastructure, cloud, database, storage, or SaaS environment.
Experience applying product management disciplines such as roadmap planning, prioritization, customer-centric thinking, outcome measurement, and portfolio management to engineering organizations.
Experience operating large-scale enterprise database and storage platforms supporting mission-critical workloads.
Experience building and scaling Reliability Engineering, Performance Engineering, Platform Engineering, SRE, or Production Engineering organizations.
Experience with observability platforms, telemetry systems, diagnostics frameworks, and production analytics.
Experience with migration readiness, resiliency validation, reliability testing, operational risk reduction, and large-scale cloud transformations.
Experience leveraging AI technologies to improve anomaly detection, forecasting, incident analysis, prioritization, and engineering productivity.
Strong understanding of distributed systems architecture, cloud platform operations, and hyperscale environments.
Experience developing executive-facing reliability scorecards, engineering metrics, and business impact reporting.
Experience influencing platform architecture, database strategy, storage strategy, and long-term engineering roadmaps.
Experience with Linux-based production environments and large-scale cloud infrastructure.
Experience supporting enterprise database technologies such as MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native database platforms.
Familiarity with ServiceNow platform architecture and large-scale SaaS operations.
JV20
For positions in this location, we offer a base pay of $221,200 - $387,100 , plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.
We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact View email address on click.appcast.io for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply
#J-18808-Ljbffr Jobleads-US- ...ServiceNow is seeking a Director of Data & Storage Reliability Engineering to lead a large, global team focused on reliability, observability, and platform performance. You will drive engineering priorities, build scalable processes, and partner with product, architecture...SuggestedRemote workFlexible hours
$136.6k - $184.8k
...Senior Supplier Quality Engineer, (Infrastructure Reliability & Quality Engineer) Job ID: 10569119 | Amazon Data Services, Inc. AWS Infrastructure Services owns the design... ...AWS data centers and all of the servers, storage, networking, power, and cooling equipment...SuggestedFlexible hours$170k
...industry-leading safety standards and years of proven deployment data, we're pioneering a new era of automation that enhances human... ...tasks in demanding environments. We are seeking a Senior Reliability Engineer with 5–8+ years of hands‑on experience in mechatronics and...SuggestedFull timeTemporary workWork at officeRelocation packageFlexible hours- ...learning and growth.The location in Hawesville, (Kentucky, United States), is seeking talent to fill the position of Electrical Reliability Engineer. This job is full-time permanent.Electrical Reliability EngineerThe Electrical Engineer is responsible for design...SuggestedPermanent employmentFull timeFor contractorsFlexible hours
- ...JOB DESCRIPTION We are seeking a Director, Data Management to lead large-scale Analytics & Data Management programs for marquee clients in our Sports & Gaming practice. This role combines delivery excellence, stakeholder management, team leadership, and account growth...Suggested
$135.38k - $194.46k
...healthcare experience. Here, industry depth meets media scale, where data becomes direction, where creativity and storytelling bring truth... ..., advantage follows. Go deeper. Be found. Overview The Director, Applied Data Science plays a key role in delivering advanced...Temporary workFreelanceFlexible hours- ...grounding meet the standards required for reliable charger and vehicle operation. Assess... ...with Highland's Construction Electrical Engineering and Procurement teams to build field... ...upstream assets—including battery energy storage systems (BESS), solar canopies, and other...Temporary workRemote work
- ## Director of Data and AnalyticsApply: Remote: USA\\_GA\\_Remote: USA\\_TX\\_Remote: USA\\_PA... ...enterprise applications to deliver reliable, consistent, and timely information.*... ...information systems, business analytics, engineering, or a related field; an advanced degree...Full timeH1bLocal areaRemote work
$95k - $142.6k
...Position Summary: ~ The Electrical Reliability Engineer is responsible for electrical reliability projects supporting the Maintenance and Reliability... ...by applying Reliability Engineering principles, statistical data analysis and supporting work process. This is a fast-paced...Hourly payFor contractorsWork at office- ...The Director of Data Governance is a senior leadership role that owns NVA’s enterprise data governance strategy, operating model, and framework... ...root‑cause resolution with the appropriate data owners and engineering teams. Establish and chair NVA’s Data Governance Council,...Local areaRemote work
$269.1k - $307.2k
## Director, Data ScienceApply: McLean, VA: San Francisco, CA: Cambridge, MA: Richmond, VA: New York, NY: Full time: Posted Today: R100... ...portfolio of agentic AI products to transform Capital One's engineering organizations and lines of business – autonomous coding agents...Full timePart timeLocal areaFlexible hours$175k - $210k
...Job Title: Director, Data Science Job Description About the Position People Inc. is... ...while working with Data Operations and Engineering to ensure our models, measurement systems... ...Code/Codex and how they can be used reliably (and where they fail) in a data science...Temporary workWork at officeLocal areaRemote workFlexible hours- ...communities thrive. That begins with how we use our science, data, and unmatched technical expertise to develop market-leading... ...depend on Chemours chemistry.Chemours is seeking a Senior Reliability Engineer to join our Asset Care team. This position will be located at...Full timeLocal area
- ...Description/ResponsibilitiesLeads maintenance reliability across Manufacturing while supporting... .... Works with Maintenance Manager, Plant Engineer, bottling and processing staff along... ...downtime based on supervisory input and data. Provide support for the structure of PM...Hourly payFull timeWork at officeNight shift
- ...contacts internal and external experts as required.• Utilizes reliability tools such as reliability analytics, failure evaluations, and... ....MINIMUM QUALIFICATIONS:• Bachelor’s Degree in Mechanical Engineering required.• Zero (0) years or more of experience required.As an...Full timeLocal area
- ...Ashland has an exciting opportunity for a Reliability Instrument / PM Engineer to join our Ashland Inc., business... ...calibrations of material storage vessels governed by the 2002 Sarbanes... ...instrumentation installation drawings, data sheets and specifications, in full compliance...Hourly payFull timeContract workFor contractors
- ...Kentucky, United States), is seeking talent to fill the position of Reliability Engineer. This job is full-time permanent.The Reliability Engineer is... ...engineering principles, advanced maintenance strategies, and data-driven decision-making.Working closely with maintenance,...Permanent employmentFull time
- ...plays a key part in supporting equipment reliability, reducing unplanned downtime, and... ...compressors, etc.).Conduct route-based data collection, analysis, and reporting using... ...corrective action.Collaborate with maintenance, engineering, and operations teams to optimize...Night shift
- ...team.Broadridge is hiring! We are seeking a Director of Data Science to lead Finance Transformation,... ...workflows.Work with financial systems and engineering teams to integrate data from relevant platforms and deploy reliable solutions into production.Establish standards...Full timeLocal area
- ...decarbonize the economy, please visit Position Summary We are hiring a Reliability Engineer with at least 10 years of experience to work in our Operations team. This position reports to Director -BioGas Operations and Delivery will work remotely from home, with travel...Full timeTemporary workLocal areaRemote workWork from homeMonday to FridayWeekend workDay shift
$120k - $140k
...we evolve to win in self-care. Description Overview The Reliability Engineer serves as the site reliability leader and technical subject... ...recurring equipment issues, analyze maintenance and performance data, and implement sustainable corrective actions to improve...For contractors- ...wear and lubrication Who is it for Engineers, researchers, students and industry... ...contributing: Drives effort to ensure desired reliability and maintainability of equipment,... ...Computerized Maintenance Management System (CMMS) data analysis Facilitates team approach in...Contract workFor contractors
- ...Responsibilities SUMMARY The Reliability Engineer at STRATTEC Corporation is responsible for improving the performance, safety, and longevity... ...manufacturing equipment by applying engineering principles, data analysis, and proactive maintenance strategies. This role...Work at officeLocal area
- ...Location: Yankton, SD, US, 57078 Career area: Engineering Department: Engineering - YTX Job Type:... ...you will be doing Job Summary: The Reliability Engineer is responsible for ensuring the... ..., control and analysis of all data, records and equipment histories to continually...Permanent employmentLocal areaWorldwide
- ## Reliability EngineerApply: On-site: Moncks Corner, South Carolina: Full time: Posted Today... ...DuPont Careers**Mechanical Reliability Engineer****Drive Reliability. Power Performance.... ...downtime and improve availability· Leverage data for impact: Analyze performance trends...Full timeWorldwide
- ...JOB DESCRIPTION: Hansen Agri-PLACEMENT is representing an established agricultural company who are searching for a Reliability Engineer. The Reliability Engineer/Coordinator will perform routine maintenance practices, resource management and tools and processes...Work experience placement
- ...as a company. Our success comes from finding, retaining, and supporting the highest quality talent by offering: Position Reliability Engineer Job Category Maintenance Industry Type Reliability Engineer Location of Job Sayreville, NJ, US, 08872 Salary Negotiable...Immediate start
$100k - $115k
...Senior Reliability Engineer – Electrical Engineer Salary $100,000 - $115,000 + Benefits + Paid Relocation to Kentucky where it’s a wonderfulplace to raise a family! City amenities with a small-town feel. History, fun music & food festivals with a charming downtown....For contractorsRelocation package$165k - $210k
Keep the systems we operate running, and push what you learn on call back into how we design the next one. Team Operate Type Full-time Band $165k - $210k Location San Francisco or remote (UTC-8 to UTC+2) What we look for Production on-call experience...Full timeRemote work$69k - $88.66k
...City, IL, USA On-site Maintenance/ Reliability Full-Time Requisition #: RELIA0016... ..., prepares and analyzes reports and data to identify trends, root cause analysis... ...interface between operations, maintenance, and engineering to ensure reliability principles are...Full timeWork at officeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Director, Data & Storage Reliability Engineering. Be the first to apply!
- director data analytics Kentucky
- data integration manager Kentucky
- data manager Kentucky
- director data architecture Kentucky
- director data engineering Kentucky
- senior manager data science Kentucky
- senior clinical data manager Kentucky
- senior data manager Kentucky
- director data management Kentucky
- lead clinical data manager Kentucky


