Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director, Engineering - Data Infrastructure & Reliability

$195k - $290k
Full-time

Crowdstrike

As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.

About the Role:

CrowdStrike's Data Platform is the foundation beneath Falcon and Next-Gen SIEM: the ingestion, streaming, storage, and query systems that take in almost 3 trillion events per day, at millions of events per second, and make them searchable in seconds. It spans multiple clouds and regions, including sovereign and regulated environments, and it holds petabytes of security telemetry that customers rely on in the middle of active investigations.

We set a high bar for that platform. Reliability, cost efficiency, and provable compliance are engineering requirements here rather than afterthoughts, and we are raising that bar again as we move the organization toward domain ownership, data as a product, self-serve tooling, and federated governance.

We support planetary scale and our systems are stateful and order-dependent, which puts them outside what generic cloud build automation is designed for. Validating a change means tracing real data from the sensor through ingestion to query rather than exercising a service in isolation. The largest cost levers are spread across every team that queries or stores data. Each new sovereign region has to prove residency and controls on its own evidence. Work like this is worth far more solved once, well, and for everyone than solved separately on six teams.

We are forming a centralized Data Platform Enablement team to own that category of work, and we are hiring its founding engineering leader.

You will start by hiring and forming a small, senior team, then grow it into a multi-team organization as the charter expands. The team works as a horizontal function outside individual product team scope, in three modes. It builds platform-wide automation, tooling, and test infrastructure. It partners hands-on with individual data platform teams. And it enables the rest of the organization through guardrails, best practices, and self-service tooling. Every hour your team spends on cross-cutting reliability, testing, and cost work is an hour another data platform team gets back for work in its own domain.

The reach of this role is wider than its headcount. You will not own the product domains, but you will answer for reliability, cost, and compliance outcomes across all of them, which means most of your influence has to come from somewhere other than your reporting line. This one suits someone who would rather build a function than inherit one.

This role is hybrid, linked to one of the posted locations: Austin, TX; Sunnyvale, CA; Redmond, WA; or NY, NY.

What You'll Do:

Organizational Leadership & Team Building

  • Stand up the team: define its operating model, hire net-new engineers, and bring in senior operators from existing data platform teams without leaving holes behind them.

  • Hire, develop, and lead engineering managers and technical leads as the team grows from one group into several domain-aligned ones.

  • Set the technical bar. Your team's output is tooling, harnesses, and guardrails that engineers outside your org have to actually want to adopt.

  • Own org design as the charter matures: forming, rechartering, and sequencing teams against where the platform's risk actually sits.

Cloud Expansion Automation

  • Build the automation that makes stateful data platform systems first-class citizens in our cloud build tooling, including the ordering, state awareness, and dependency sequencing that generic SRE automation cannot handle.

  • Take regional and sovereign cloud build-outs to full automation, so standing up a new environment is a repeatable, hands-off sequence.

  • Own time-to-launch per cloud and region as a headline metric, alongside the teams that own the destination architecture.

End-to-End Validation & Continuous Health Checks

  • Build a testing system that generates and traces data from the CrowdStrike sensor through ingestion to query validation. The same harness does pre-launch validation for every new cloud and region, then continuous health checks in steady state.

  • Raise the fidelity of generated test data so it exercises the same paths production traffic takes, and a passing run means what it claims.

  • Hold end-to-end (i..e sensor to query) coverage and escaped-defect rate as first-class metrics for every launch.

Observability & Proactive Incident Response

  • Standardize reliability and efficiency metrics across every data platform team, so the whole platform reports from one instrumented view.

  • Build detection that catches problems before customers report them, with automated routing to the right responders plus first-line triage and guidance for owning teams.

  • Set and drive down targets for mean time to detect, respond, and recover, backed by instrumentation that proves the trend.

  • Push the incident rate between releases and operations down by addressing systemic causes.

Cost Engineering

  • Run fast detection and remediation for cost anomalies, so an efficiency regression surfaces in days rather than quarters.

  • Find and execute optimization work across the ecosystem, either directly with partner teams or by handing them tooling. Some of it is straightforward compute migration. Some of it is pipeline and architecture rework.

  • Move cost accountability to the layers that own the levers, so that application teams see what their query patterns, storage choices, and pipeline hops actually cost. This is as much an organizational change as a technical one, and you will lead it.

Resilience, Scale & Capacity Planning

  • Establish a recurring game day practice that stresses systems deliberately to establish their real limits.

  • Run scale and stress testing ahead of projected demand.

  • Work with data platform teams on forward-looking capacity planning, and hold forecast accuracy as a real metric.

Data Residency & Compliance Automation

  • Turn residency verification and audit evidence into reusable platform capabilities that every new region inherits. GRC defines the controls and other teams own the sovereign architecture; your team automates the data platform's side of proving both.

  • Build automated validation that data lands, is processed, and is queried inside its declared residency boundary, wired into the same harness as the rest of the testing rather than a parallel one.

  • Express the data platform's share of sovereign and certification controls as executable checks that run continuously, and generate attestation evidence from telemetry so certification and re-certification draw on evidence that is always current.

  • Treat a control that has quietly stopped holding as an incident, detected and routed under the same targets as any reliability event.

Executive Narrative & Cross-Functional Influence

  • Turn your team's delivery into a clear executive story on reliability, cost, and compliance posture, and carry it into senior forums with candor about risk.

  • Earn adoption from engineering leaders who do not report to you. A horizontal team runs on that trust.

  • Negotiate scope deliberately. This charter will attract more requests than any team can absorb, and protecting your team's focus is part of the job.

What You'll Need:

  • 12+ years of software engineering experience, including 5+ years leading engineering teams and at least 2 years leading through other managers.

  • A track record of building, growing, and retaining high-performing platform, SRE, or infrastructure teams in a fast-paced, high-growth environment, including hiring senior engineers onto a team that did not exist yet.

  • Hands-on grounding in SRE practice at scale: SLOs, SLIs, error budgets, incident command, blameless postmortems, and capacity planning for high-throughput distributed systems.

  • Experience owning reliability for large-scale stateful distributed systems, where ordering, data state, and recovery semantics rule out generic automation. This is the hardest technical part of the job.

  • Working fluency with distributed data infrastructure: streaming platforms, OLAP and search engines, object storage, and large-scale query systems. Enough to hold a credible design conversation and to recognize a bad proposal.

  • Ownership of a substantial infrastructure cost portfolio, including driving optimization across organizational boundaries and shifting cost accountability to the teams that control the spend.

  • Proven ability to influence without authority. You have changed how engineering teams outside your reporting line operate, and you can explain how you earned that adoption instead of mandating it.

  • Experience operating under regulatory, residency, or certification constraints, and comfort turning control requirements into engineering work.

  • Strong executive communication. You can compress a complicated reliability or cost story into something a leadership team can decide on, and you do not let your team's work go unseen.

  • A specific point of view on applying AI to reliability and operations: where it pays off, where it does not, and what you would build first.

  • Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.

  • Bachelor's degree in Computer Science or related field, or equivalent work experience.

Bonus Points:

  • Experience standing up sovereign, air-gapped, or regulated cloud environments, and automating the evidence that proves they comply.

  • Hands-on depth with Kafka, Flink, Spark, Cassandra, OpenSearch, Pinot, Trino, or comparable systems at petabyte scale.

  • Python and/or Golang, Infrastructure as Code (Terraform, Ansible, Pulumi), Kubernetes at fleet scale, and GitOps workflows.

  • Advanced observability practice with Prometheus, Grafana, OpenTelemetry, distributed tracing, and large-scale log aggregation, weighted toward SLO dashboards and reliability scorecards rather than vanity metrics.

  • Having built and run a chaos engineering or game day practice that other teams joined voluntarily.

  • FinOp, or having led a cost optimization and/or re-attribution program to completion.

  • Having shipped LLM-native or agentic tooling for incident prevention, triage, or remediation.

#LI-MP2

This role will require the candidate to periodically undergo and pass additional background and fingerprint check(s) consistent with government customer requirements.

Benefits of Working at CrowdStrike:

  • Market leader in compensation and equity awards

  • Comprehensive physical and mental wellness programs

  • Competitive vacation and holidays for recharge

  • Paid parental and adoption leaves

  • Professional development opportunities for all employees regardless of level or role

  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections

  • Vibrant office culture with world class amenities

  • Great Place to Work Certified™ across the globe

CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program.

CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, genetic information, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions--including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs--on valid job requirements.

If you need assistance accessing or reviewing the information on this website or need help submitting an application for employment or requesting an accommodation, please contact us at View email address on aiapply.co for further assistance.

Find out more about your rights as an applicant.

CrowdStrike participates in the E-Verify program.

Notice of E-Verify Participation

Right to Work

CrowdStrike, Inc. is committed to fair and equitable compensation practices. Placement within the pay range is dependent on a variety of factors including, but not limited to, relevant work experience, skills, certifications, job level, supervisory status, and location. The base salary range for this position for all U.S. candidates is $195,000 - $290,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off.

For detailed information about the U.S. benefits package, please click here.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Director, Engineering - Data Infrastructure & Reliability in Sunnyvale, CA vacancy
  • $200k - $245k

     ...We are seeking a strategic and technically accomplished Director, Data Products Engineering to lead the architecture, engineering, and delivery of AI...  ..., designing scalable data architectures, and enabling reliable enterprise data consumption at scale. The current ecosystem... 
    Suggested
    Work from home
    Flexible hours

    PayNearMe, Inc.

    Santa Clara, CA
    17 days ago
  • $195k - $285k

     ...manufactures purpose-built AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we build. This role...  ...of 1–3 SRE engineers. Direct a dedicated Data Center & Lab Technician team, setting work... 
    Suggested
    Full time
    Remote work

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $200k - $391k

     ...NVIDIA is looking for an exceptional Engineering Manager to lead, scale, and innovate our core Data Labeling Platform. This is a highly visible, high-impact role...  ..., and alerting built on it against defined reliability and latency targets Front-End & Annotation Interfaces... 
    Suggested
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $224k - $321k

     ...organization of scientists, engineers, and physicians and we are using...  ...-art computer science and data science to overcome one of...  ...visionary and execution-focused Director of Software Engineering to...  ...in big data, cloud infrastructure, and AI-driven workflows (including... 
    Suggested
    Full time
    Work at office
    Local area
    Flexible hours

    GRAIL

    Sunnyvale, CA
    6 days ago
  • $185k - $230k

     ...The Opportunity We are looking for a Senior Data Engineer to join our Data Platform team and build the core data foundations that power...  ...and key business metrics Design, operate, and scale reliable data pipelines and platforms Partner with stakeholders to... 
    Suggested

    Cacheflow

    Mountain View, CA
    3 days ago
  • $160.36k - $240.54k

     ...diversity of its training and evaluation data. The team plays a crucial role in...  ...systems by creating a scalable and reliable data infrastructure. This infrastructure is designed to...  ...team collaborates closely with system engineers to thoroughly validate the autonomous... 
    Work experience placement

    Kindredventures

    Mountain View, CA
    3 days ago
  •  ...legal entity.Atlassian’s Trusted Data Platform (TDP) provides secure, reliable, and scalable data-storage...  ...relational database platform. It gives engineering teams a supported way to build...  ...having to manage the underlying infrastructure, availability, security, scaling... 
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    17 hours ago
  • $253k - $416k

     ...Mountain View, CA OR Bellevue, WA. LinkedIn’s Data Infrastructure team builds and operates some of the...  ...constantly pushing the boundaries of scale, reliability and user experience. We are seeking a Distinguished Engineer with deep expertise in Online Infrastructure... 
    Full time
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    a month ago
  • $165k - $242k

     ...CoreWeave combines superior infrastructure performance with deep technical...  .... About the role: The Data Platforms Team serves as the...  ...We are seeking a senior engineer with specialization in database...  ...solve for scalability and reliability.  Improve the performance,... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Core Weave

    Sunnyvale, CA
    more than 2 months ago
  • $213k - $263k

     ...Platform team, builds tools and infrastructure to realize the ML flywheel...  ...ML workflows manageable and reliable. This team also partners...  ...Develop and contribute to Waymo's data infrastructure platform to...  ...in the field of software engineering ~ Experience programming in... 
    Full time
    Remote work

    Waymo

    Mountain View, CA
    more than 2 months ago
  •  ...The Role This role offers a mix of data infrastructure, distributed systems, financial data, and product engineering. You will work on systems that ingest and normalize...  ...metrics at scale, and turn that data into reliable products and experiences for our clients.... 
    Full time
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Arta Finance

    Mountain View, CA
    a month ago
  • $272k - $431.25k

     ...massive superchip. We are looking for expert engineers to come and help design rack level...  ...manageability architecture for these products in data centers. You will work with various...  ...center health management workflow.Drive reliability and optimization in firmware... 
    Work at office

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $193.93k - $291.15k

     ...investors. About the Role We are a team of high-output generalists where ML and systems engineering converge to push autonomy performance forward. As a Senior Perception ML Data Infrastructure Engineer, you will own the critical bridge between our autonomous vehicle hardware... 
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    11 days ago
  •  ...technologies that enable AI, data centers, telecom, industrial...  ...that keep data moving reliably, efficiently, and at massive...  ...products support AI and compute infrastructure, cloud and DCI environments,...  ...locations. Senior Manager, Data Engineering Responsibilities... 

    Lumentum Operations LLC

    San Jose, CA
    1 day ago
  • $180k - $220k

     ...deployments, application security, reliability, compliance, and cost optimization....  ...seeking a talented and experienced Data Platform Engineer to join our team as the Technical Lead...  ...maintaining our data pipelines and infrastructure. They will work closely with cross-functional... 
    Flexible hours
    Shift work
    3 days per week

    B Capital

    Mountain View, CA
    3 days ago
  • $83.98 - $111.27 per hour

     ...Care is seeking a dynamic and experienced Senior Manager of Data Engineering to lead our data engineering teams. This role is critical to...  ...Governance and Operational Excellence: Ensure the integrity, reliability, and performance of all data platforms and pipelines.... 
    Hourly pay
    Full time

    Stanford Health Care

    Palo Alto, CA
    1 day ago
  • $130.6k - $285k

     ...DoorDash is building the world's most reliable on-demand, logistics engine for delivery! We're looking for...  ...to help us develop a 24x7, global infrastructure system that powers DoorDash's three...  ...and dashers. About the Team Data Platform’s Data User Experience Engineering... 
    Hourly pay
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours

    DoorDash USA

    Sunnyvale, CA
    1 day ago
  • $149.8k - $262.2k

     ...Description It all started when engineer Fred Luddy wrote code that...  ...brings together any AI, any data, and any workflow— helping 8...  ..., with a strong focus on reliability, scalability, and operability...  ...as operators, controllers, infrastructure automation, and platform services... 
    Permanent employment
    Full time
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours
    2 days per week

    ServiceNow

    Santa Clara, CA
    3 days ago
  • $149.8k - $262.2k

     ...Description It all started when engineer Fred Luddy wrote code that...  ...brings together any AI, any data, and any workflow— helping 8...  ...You will contribute to the reliability, scalability, and...  ...changes the way software and infrastructure are built. ~8+ years of software... 
    Permanent employment
    Full time
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours
    2 days per week

    ServiceNow

    Santa Clara, CA
    3 days ago
  • $148.32k - $203.94k

     ...higher performance, smaller size, lower power, and better reliability. With more than 4 billion devices shipped, SiTime is...  ...: Job Summary We are seeking a hands-on Principal Infrastructure Hardware Engineer to architect, design, and deliver system platforms supporting... 

    SiTime Corporation

    Santa Clara, CA
    27 days ago
  • $240k - $320k

     ...senior electrical and system integration engineers to define, build, and scale our vehicle...  ...the Senior Principal Engineer – ADAS & AV Data Loop & AI Flywheel, you will spearhead...  ...functional leads (data engineering, cloud infrastructure, embedded runtime SOC teams) to define,... 
    Full time
    Work experience placement
    Local area
    Flexible hours

    Bosch Group

    Sunnyvale, CA
    1 day ago
  • $258k - $387k

     ...looking for a seasoned engineering leader to lead the Eval Platform as a Director‑level role. This...  ...simulation, evaluation infrastructure, and compute and storage...  ...platform brings together data, simulation, metrics,...  ...leadership, ensuring it is reliable, scalable, and trusted... 
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    11 days ago
  • $116.8k - $160k

     ...Description AWS Infrastructure Services owns the design, planning,...  ...running. We support all AWS data centers and all of the servers...  ...software, hardware, and network engineers, supply chain specialists,...  ...maintain consistency and reliability in services delivered. Work... 
    Work at office
    Local area
    Flexible hours

    Amazon

    Mountain View, CA
    2 days ago
  • $175k - $265k

     ...Overview d-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on —...  ...of that team, responsible for reliability, automation, and observability across...  ...used across the global SRE and data center services teams. Build and... 
    Full time

    d-Matrix

    Santa Clara, CA
    1 day ago
  • $164k - $246k

     ...Job Description Job Description As Engineering Manager for the Data Platform team, you'll lead a team of engineers building and operating the infrastructure that powers data ingestion, governance, storage, processing, and access across FloQast's product and analytics... 
    Day shift

    FloQast

    San Jose, CA
    7 days ago
  • $182k - $242k

     ...CoreWeave combines superior infrastructure performance with deep technical...  ...CoreWeave is looking for an Engineering Manager to lead a team...  ...and is responsible for the reliability, scalability, and operational...  ...each day in our office and data center locations ~ A casual... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    16 days ago
  • $140k - $260k

     ...and Cloud & AI, the digital infrastructure powering our collaborative foundation...  ...infrastructure that enables engineers and researchers to train,...  ...workloads, simulation, and data processing are critical to...  ...ensure that our platform is reliable, efficient, secure, and easy... 
    Temporary work
    Work at office
    Immediate start
    Flexible hours

    Woven by Toyota

    Palo Alto, CA
    2 days ago
  • $170k - $240.8k

     ...driven expert in ML Training Infrastructure with a strong ability to execute...  ...and building scalable, reliable, and high-performance AI/ML platform...  ...initiatives. As a Senior ML Engineer, you will collaborate closely...  ...and optimizing training and data loading performance.... 
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $140k - $200k

     ...office. These include frontend and backend engineers, AI research scientists, and others from...  ...Overview We're looking to hire for our Data side of our AI team at Speechify. This...  ...low cost through a tight integration of infrastructure, engineering, and research work. We are... 
    Full time
    Work at office
    Shift work

    Speechify

    Cupertino, CA
    more than 2 months ago
  •  ...real-world experience, and the data platform is what turns raw...  ...happens next: ingesting that data reliably, validating and curating it...  ...the role: As a Software Engineer working on the data platform,...  ...workflows, computer vision, or data infrastructure Bias for ownership: you'... 
    Immediate start

    Mind Robotics Inc.

    Palo Alto, CA
    12 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director, Engineering - Data Infrastructure & Reliability. Be the first to apply!