Director, Data & Storage Reliability Engineering
$221.2k - $387.1kServiceNow
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.
Join us to put AI to work for people.
Job Description
What you get to do in this role:
Team Management
The successful candidate will lead the Data & Storage Reliability Engineering organization responsible for improving reliability, resilience, performance, scalability, observability, and customer experience across ServiceNow's database, storage, and supporting platform infrastructure.
This leader will be responsible for building, developing, and scaling high-performing engineering teams focused on reliability engineering, observability, performance engineering, diagnostics, automation, production analytics, migration readiness, resilience engineering, and prevention engineering.
Responsibilities include talent acquisition, performance management, career development, succession planning, objective setting, coaching, and prioritization of strategic initiatives.
The role will establish a strong engineering-first culture centered on data-driven decision making, continuous improvement, operational excellence, customer experience, and systemic risk reduction.
This position is accountable for identifying recurring failure patterns, reliability risks, performance bottlenecks, scalability constraints, migration challenges, and operational inefficiencies across database services, storage platforms, cloud infrastructure, and distributed application environments, and driving engineering improvements that eliminate entire classes of issues before they impact customers.
The Director will partner closely with Product Engineering, Database Engineering, Cloud Infrastructure, Architecture, Storage Engineering, Support, and Operations teams to ensure reliability, observability, performance, and resilience considerations are incorporated throughout the software development lifecycle.
The successful candidate will also partner closely with SWAT and Customer & Production Engineering teams to establish a continuous feedback loop between production operations and platform improvement. SWAT remains responsible for customer escalations, production operations, incident response, and service restoration, while this organization is responsible for identifying systemic opportunities, defining engineering priorities, and driving platform improvements that reduce future customer impact.
The successful candidate will serve as the senior technical leader for complex reliability investigations, customer-critical escalation reviews, migration readiness assessments, and platform improvement initiatives, transforming production insights into long-term engineering outcomes.
They will influence architectural decisions and technology investments by providing reliability expertise, observability insights, performance guidance, and production-based evidence that improve platform resilience, scalability, efficiency, and customer outcomes.
This role requires a strong product mindset. The leader will treat reliability, observability, resilience, performance, and automation capabilities as products with roadmaps, priorities, adoption goals, and measurable outcomes. They will be responsible for identifying the highest-value engineering opportunities, prioritizing investments, and driving adoption across multiple product and infrastructure organizations.
Process and Procedures
The successful candidate will establish scalable reliability engineering practices, standards, governance processes, and operating models across the organization.
They will drive adoption of observability standards, reliability engineering frameworks, resiliency assessments, migration readiness practices, diagnostics capabilities, engineering guardrails, and automation strategies.
This leader will continuously evaluate incidents, customer escalations, migration outcomes, platform telemetry, performance trends, capacity signals, and operational data to identify systemic risks and drive long-term engineering improvements.
The role will establish a formal review process with SWAT and Customer & Production Engineering teams to evaluate major incidents, recurring operational challenges, migration learnings, customer-impacting events, and emerging platform risks. These insights will be used to prioritize engineering investments and platform improvements.
The successful candidate will establish meaningful KPIs and engineering metrics that provide visibility into platform reliability, resiliency, performance, operational efficiency, customer experience, engineering productivity, and risk reduction.
The successful candidate will leverage AI-powered tools, analytics, automation frameworks, and production intelligence to identify emerging risks, improve detection coverage, accelerate engineering insights, reduce operational toil, and improve engineering productivity.
They will use production telemetry, incident learnings, customer escalations, migration outcomes, observability data, and operational trends to drive architectural improvements, reliability investments, platform standards, and long-term engineering evolution.
The Director will maintain a portfolio of reliability investments spanning observability, performance, diagnostics, resilience, automation, and prevention, balancing immediate customer needs with long-term platform strategy.
The Director will champion a proactive reliability engineering model that shifts the organization from reactive issue response toward predictive analysis, prevention, resilience, and continuous optimization.
Qualifications
To be successful in this role you have:
- Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
Strong product mindset with demonstrated experience treating technical capabilities as products with roadmaps, priorities, customers, adoption goals, and measurable business outcomes.
Experience translating production insights, customer pain points, operational challenges, reliability risks, and platform telemetry into prioritized engineering investments and long-term roadmaps.
Experience partnering closely with production operations, customer escalation teams, reliability organizations, and software engineering teams to drive systemic improvements based on operational learnings.
Experience defining product strategies, developing roadmaps, prioritizing investments, and aligning stakeholders across multiple organizations without direct authority.
Experience operating a portfolio of engineering investments, balancing short-term customer needs with long-term reliability, performance, scalability, and resilience objectives.
15+ years of experience in software engineering, platform engineering, reliability engineering, infrastructure engineering, database engineering, distributed systems, product management, or large-scale SaaS environments.
8+ years of engineering leadership experience, including leading managers and globally distributed teams.
Extensive experience leading Reliability Engineering, Platform Engineering, Database Engineering, Infrastructure Engineering, Production Engineering, Performance Engineering, or related technical organizations.
Deep expertise in distributed systems, databases, storage technologies, cloud infrastructure, and large-scale SaaS architectures.
Strong understanding of reliability engineering principles, observability, scalability, resiliency, operational excellence, and performance engineering.
Experience building and operating observability, telemetry, diagnostics, reliability, or performance capabilities at scale.
Proven experience identifying systemic issues and converting operational insights into strategic engineering improvements.
Experience partnering closely with Product Management organizations to influence roadmaps and deliver customer-centric outcomes.
Experience driving engineering initiatives through data, metrics, customer impact analysis, and measurable business outcomes.
Experience leveraging AI technologies to improve decision-making, analytics, engineering workflows, operational efficiency, reliability insights, automation, or customer outcomes.
Exceptional communication, stakeholder management, and leadership skills.
Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Desired Skills
Previous Product Management experience in a platform, infrastructure, cloud, database, storage, or SaaS environment.
Experience applying product management disciplines such as roadmap planning, prioritization, customer-centric thinking, outcome measurement, and portfolio management to engineering organizations.
Experience operating large-scale enterprise database and storage platforms supporting mission-critical workloads.
Experience building and scaling Reliability Engineering, Performance Engineering, Platform Engineering, SRE, or Production Engineering organizations.
Experience with observability platforms, telemetry systems, diagnostics frameworks, and production analytics.
Experience with migration readiness, resiliency validation, reliability testing, operational risk reduction, and large-scale cloud transformations.
Experience leveraging AI technologies to improve anomaly detection, forecasting, incident analysis, prioritization, and engineering productivity.
Strong understanding of distributed systems architecture, cloud platform operations, and hyperscale environments.
Experience developing executive-facing reliability scorecards, engineering metrics, and business impact reporting.
Experience influencing platform architecture, database strategy, storage strategy, and long-term engineering roadmaps.
Experience with Linux-based production environments and large-scale cloud infrastructure.
Experience supporting enterprise database technologies such as MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native database platforms.
Familiarity with ServiceNow platform architecture and large-scale SaaS operations.
JV20
For positions in this location, we offer a base pay of $221,200 - $387,100 , plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as
qualifications
, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.Additional Information
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact View email address on us.fitly.work for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.
$139.52k - $255k
...vision and mission to be the go-to partner for optimized data storage solutions. You can be part of the takeoff of an innovative... ...Description NAND array algorithm development and performance/reliability engineering. Key Responsibilities: Design, plan, and execute...SuggestedFull timeTemporary workFlexible hours3 days per week$267k - $356k
...home day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...We own the full spectrum of Lambda's data platform services—from low-level storage... ...in the industry, which means reliability and performance aren't just goals—they...SuggestedPart timeWork experience placementWork at officeLocal areaWork from homeFlexible hours$305k
...AIML - Head of Data Science and Insights Cupertino, California, United States Software and Services Apple is where individual... ...Leading a large team of data scientists and machine learning engineers — recruiting, developing, and retaining strong technical talent...SuggestedRelocation package- ...Evaluation team is looking for a seasoned, technical leader to lead our Data Science and Insights team. The organization leads Evaluation... ...organizations of 50+ data scientists and/or machine learning engineers ~ Deep experience in human evaluation methodology, logging,...Suggested
$229.8k - $344.8k
...to prioritize use cases while working closely with customers, engineering, sales, and business development teams. You will also function... ...Drive executive decision making for new investments in cloud data center solutions, including competitive analysis, program goals...SuggestedWork experience placementWork at officeWork from home- ...accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded... ...Together, we advance your career. THE ROLE The Sr Director of Silicon Design Engineering is responsible for owning end to end silicon performance...
$229.8k - $344.8k
...Overview Join Qualcomm's Data Center Business Unit to lead the strategy, definition, and execution of custom silicon and chiplet... ...ramp. Working closely with customers and Qualcomm's architecture, engineering, operations, and business organizations, you will help define...Work experience placementWork at office- ...cloud infrastructure and managed data services, including... ...database, operating system, storage, and network layers. 3+... ...experience in Linux systems engineering (performance tuning, memory management... ...Responsibilities As a Senior Database Reliability Engineer, you will help make...Full time
- ...We are seeking a Senior Database Reliability Engineer (DBRE) to design, operate, and improve reliable... ..., and highly available database and data platforms. The role combines database engineering... ..., database, operating system, storage, and network layers. Design and maintain...Full time
- ...Analog Devices, Inc. is seeking a Senior Director to lead the global Product Marketing and Applications Engineering organization for Data Center Power Delivery. You will connect market insight, semiconductor expertise, and customer execution to accelerate profitable...
- ...Analog Devices, Inc. is seeking a Senior Director to lead the global Product Marketing and Applications Engineering for Data Center Power Delivery. You will own the portfolio, translate hyperscaler needs into silicon solutions, and drive a strategic roadmap to accelerate...
$168k - $270.25k
...intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and... ...scalable and efficient Storage solutions tailored for data-intensive applications, optimizing performance and cost...Full time- ...production methods. Job Summary: The primary purpose of the Reliability Engineer (focused on L10/L11 execution) is to own the physical,... ...cause analysis of product and component failures through test data review, physical inspection, and material evaluation to identify...Work at officeLocal area
$150k - $230k
...Arista Networks is an industry leader in data-driven, client-to-cloud networking for... ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation... ...We are seeking an Optical Transceiver Reliability and Qualification Engineer to lead the...Full timeContract work- ...everything from smartphones to data centers. As a global leader... ...the Team: SK HMS Systems engineering / Quality assurance group is... ...motivated individual to join our reliability testing team. Job... ...focusing on SSDs, enterprise storage devices, data servers, or embedded...Contract workFor contractors
- ...well as cable assemblies for diverse applications including server, storage, data center, mobile, RF, networking, industrial, business equipment, and automotive. Position: Entry Level Electrical Engineer – Class of 2025/2026 Location: Santa Clara, CA Amphenol High...
$140k
...Head Of Data & Analytics Bay FC is the first NWSL team in the Bay Area. Co-founded... ...candidate will work closely with the Global Director of Data & Technology at Bay Collective,... ...effectively, enabling clean, connected, and reliable data flows across departments. At club...Ongoing contractLocal areaAfternoon shift$100k - $120k
...enterprise datacenter SSD development. The Reliability System Engineering team is seeking a talented and detail... ..., and reliability of our enterprise data center SSD products. You will work... ...in QA testing, preferably with SSDs, storage devices, data servers, or embedded...$120k - $145k
...Supermicro® is a Top Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise IT, Hadoop/ Big... ...community. We seek talented, passionate, and committed engineers, technologists, and business leaders to join us. Job...Worldwide$137k - $156k
...Tier provider of advanced server, storage, and networking solutions for Data Center, Cloud Computing, Enterprise... ..., passionate, and committed engineers, technologists, and business leaders... ...compatibility, performance, stress, and reliability testing, leveraging proprietary in...Worldwide- ...Supermicro is seeking an experienced Sr. Director of Solutions Architecture to lead Data Center Solutions with a primary focus on second-level service management and operational excellence. This leader will define how Supermicro’s data center offerings are positioned...
- ...size, lower power, and better reliability. With more than 4 billion... ...product leader to drive SiTime's Data Center & AI Infrastructure... ...work closely with customers, engineering, sales, applications engineering... ...semiconductor industry. ~ Director-level experience driving product...
$194.27k - $323.78k
...FLASH™, is shaping the future of storage in high-density applications,... ..., PCs, SSDs, automotive and data centers. Job Description... ...Americas, Inc. is seeking a Director of Data Center SSD Product Management... ...strong partnerships across engineering, marketing, sales, operations...Local areaWorldwideFlexible hours- ...ability to multi-task in fast-paced environments. Global Collaboration: Comfortable working cross-functionally with product/engineering units across multiple time zones. Documentation: High care in creating detailed design specifications and presenting...
- ...Site Reliability Engineer Foxconn Industrial Internet (Fii), is a world leading professional design and manufacturing service provider of communication network equipment, cloud service equipment, precision tools and industrial robots. FII provides customers with...Permanent employmentFull timeWork at officeLocal area
$170k - $200k
...Site Reliability Engineer We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible... ...such as firewall, WAF, switch etc,. Experience with storage service management such as Ceph. Knowledge of...Full timeWorldwide- ...Experienced with the setup/configuration/operation of CI/CD pipelines and source code control (e.g. - GitHub) Ability to communicate data, facts, and analysis of technical subject matter. Good to have knowledge of AWS (EB, EC2, RDS, Lambda, S3, CloudFront,...Flexible hours
$160k - $240k
...Senior Site Reliability Engineer Calling all innovators - find your future at Fiserv. We're Fiserv, a global leader in Fintech and payments, and we move money and information in a way that moves the world. We connect financial institutions, corporations, merchants...$145k - $175k
...Site Reliability Engineer (SRE) Bolt Graphics is a semiconductor startup based in Sunnyvale, CA building the fastest and most efficient graphics... ..., performance, and operational excellence across compute, storage, and networking environments. Exceptional Linux expertise...Work at officeImmediate startWork from home$110k - $130k
...with the World's leading AI-first Quality Engineering Company? Ready to advance your career,... ...QualityAI! We are looking for a Site Reliability Engineer to join our growing team in Riverwoods... .... SRE Skillsets - Expectations from Data Platform team: Expertise in Message...Casual workLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Director, Data & Storage Reliability Engineering. Be the first to apply!
- data services manager Santa Clara, CA
- data integration manager Santa Clara, CA
- director data center Santa Clara, CA
- senior manager data analytics Santa Clara, CA
- senior data manager Santa Clara, CA
- senior clinical data manager Santa Clara, CA
- senior manager data science Santa Clara, CA
- director data architecture Santa Clara, CA
- director data management Santa Clara, CA
- director data analytics Santa Clara, CA




