Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director, AI Platform Reliability

$247.5k - $275k
Full-time

LogicMonitor

About Us:

We love going to work and think you should too. Our team is dedicated to trust, customer obsession, agility, and striving to be better everyday. These values serve as the foundation of our culture, guiding our actions and driving us towards excellence. We foster a culture of performance and recognition, allowing us to transform growth as we enable our employees to do the best work of their careers.

This role is open to candidates based in or near San Francisco, CA. At LogicMonitor, we hire within our Centers of Energy—vibrant locations where our teams connect, collaborate, and innovate.

To learn more about life at LogicMonitor, check out our Careers Page.

What You'll Do:

LogicMonitor® is the AI-first hybrid observability platform powering the next generation of digital infrastructure. LogicMonitor delivers complete visibility and actionable intelligence across on-premises, cloud, and edge environments. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise.

Our customers love LogicMonitor's ability to bring cloud and traditional IT together into one view, as seen in minimal churn rates, expansion business, and exciting new customer references. In fact, LogicMonitor has received the highest Net Promoter Score of any IT Infrastructure Management provider. LogicMonitor also boasts high employee satisfaction. We have been certified as a Great Place To Work®, and named one of BuiltIn's Best Places to Work for the seventh year in a row!

We are looking for an accomplished and hands-on Director of AI Platform Reliability to lead the architecture, development, and operation of highly scalable, distributed software platforms.

This leader will be responsible for systems that process hundreds of millions/billions of transactions and events , manage terabytes to petabytes of data , and deliver reliable, low-latency services to enterprise customers. The ideal candidate combines strong engineering depth in Java, Kafka, distributed systems, and cloud-native microservices with a demonstrated ability to build and lead high-performing engineering organizations.

This is a strategic leadership role, but it requires a leader who can remain close to the technology, participate in architecture reviews, challenge design decisions, guide teams through complex production problems, and establish the engineering practices required to operate mission-critical platforms at scale.

Here's a closer look at this key role:

  • Lead and scale multiple engineering teams responsible for high-volume, business-critical distributed systems and data platforms.
  • Define the technical strategy and architecture for platforms processing hundreds of millions of transactions and terabytes of data.
  • Guide the development of Java-based microservices, APIs, Kafka streaming pipelines, batch-processing workflows, and cloud-native services.
  • Build and evolve scalable data lake and Data Lakehouse platforms supporting real-time, near-real-time, and batch analytics workloads.
  • Establish reliable data ingestion, transformation, storage, governance, lineage, retention, and data-quality practices across streaming and batch pipelines.
  • Build low-latency, highly available, fault-tolerant systems with strong scalability, resiliency, and disaster-recovery capabilities.
  • Define and own operational SLAs, SLOs, availability targets, recovery objectives, and performance metrics for critical services and data pipelines.
  • Drive capacity planning, load testing, throughput optimization, and improvements to p95 and p99 latency.
  • Ensure effective Kafka design, including partitioning, consumer groups, ordering, schema evolution, replay, and lag management.
  • Establish engineering standards for architecture, coding, testing, security, observability, and production readiness.
  • Partner with Product, Architecture, SRE, Security, Data, and Infrastructure teams to deliver strategic platform initiatives.
  • Strengthen operational excellence through monitoring, incident management, on-call practices, root-cause analysis, and continuous reliability improvements.
  • Recruit, mentor, and develop engineering managers, architects, and senior technical leaders.
  • Improve developer productivity, CI/CD automation, deployment safety, and release predictability.
  • Manage technical debt, platform modernization, cloud costs, and long-term scalability investments.

What You'll Need:

  • 10+ years of professional software-engineering experience, including significant experience building large-scale distributed systems.
  • Experience leading engineering teams, architects, and staff engineers.
  • Demonstrated success delivering and operating platforms that process hundreds of millions of transactions, requests, or events.
  • Deep technical expertise in Java, JVM performance, concurrency, multithreading, memory management, and application profiling.
  • Strong experience designing and operating microservice-based and event-driven architectures.
  • Extensive production experience with Apache Kafka or a comparable distributed streaming platform.
  • Strong understanding of Kafka partitioning, replication, consumer groups, offset management, ordering, delivery semantics, schema evolution, and reprocessing.
  • Experience designing low-latency, highly available APIs and backend services.
  • Experience managing terabyte- or petabyte-scale datasets across relational, NoSQL, streaming, and object-storage technologies.
  • Strong understanding of distributed-systems concepts, including consensus, replication, partitioning, consistency models, idempotency, backpressure, and fault tolerance.
  • Experience operating cloud-native applications using Kubernetes, containers, infrastructure as code, and automated CI/CD pipelines.
  • Experience with at least one major cloud platform, such as AWS, Google Cloud, or Microsoft Azure.
  • Strong knowledge of observability practices involving metrics, logs, traces, profiling, alerting, dashboards, and service-level objectives.
  • Demonstrated experience improving system reliability, scalability, latency, cost efficiency, and engineering productivity.
  • Strong written and verbal communication skills, including the ability to explain complex technical decisions to engineering teams, executives, and business stakeholders.
  • Proven ability to build inclusive, accountable, and high-performing engineering organizations.

Residents of California, click Here to view our California Applicant Privacy Notice.

Anticipated Application Close Date: 10/26/26

LogicMonitor is an Equal Opportunity Employer
At LogicMonitor, we believe that innovation thrives when every voice is heard and each individual is empowered to bring their unique perspective. We’re committed to creating a workplace where diversity is celebrated, and all employees feel inspired and supported to contribute their best.

For us, equal opportunity means fostering a truly inclusive culture where everyone has the chance to grow and succeed. We don’t just open doors; we invite you to step through and be part of something bigger. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or status as a protected veteran.

Work Authorization:
At this time, we are able to consider candidates who are authorized to work in the United States on a full-time, permanent basis without requiring new or initial employer-sponsored work authorization.
Candidates who currently hold valid U.S. work authorization that can be transferred to a new employer (such as certain H-1B statuses) may be considered on a case-by-case basis.
We are not able to provide new sponsorship for employment-based visas that require an initial petition or application by the employer.

#LI-JP1 #LI-Hybrid #BI-Hybrid

LogicMonitor is dedicated to fostering a culture of transparency and fairness, including our commitment to pay transparency. We provide the base salary ranges for all positions posted within the United States.

Compensation packages at LogicMonitor for eligible roles include base salary, a variable plan depending on role, along with comprehensive benefits. The range displayed on each job posting reflects the minimum and maximum base salary target for new hires in the position, determined by work location and additional factors, including job-related skills, experience, interview performance, and relevant education or training. As part of our holistic compensation philosophy, your package will also include, but is not limited to: Comprehensive health, dental and vision coverage, generous parental leave policies, access to our Employee Assistance Program and various Wellness programs, a 401K with company matching, a Lifestyle Spending Account, and an unlimited vacation policy. For more information on our benefits, see our careers page.

The Base Salary range for this role is:

$247,500—$275,000 USD

Our goal is to ensure an accessible and inclusive experience for every candidate.

If you need a reasonable accommodation during the application or interview process under applicable local law, please submit a request via this Accommodation Request Form.

Know your rights: workplace discrimination is illegal. Please click here to review LogicMonitor’s U.S. Pay Transparency Nondiscrimination Provision.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Director, AI Platform Reliability in San Francisco, CA vacancy
  • $189.8k - $256.16k

     ...this by building and operating the world's best data and AI infrastructure platform, enabling our customers to leverage deep data insights and...  ...Senior Staff Technical Program Manager (TPM) for Reliability to lead the strategy, execution, and continuous improvement... 
    Platform
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  • $258k - $348k

     ...design accessible to all. Figma’s platform helps teams bring ideas to...  ...into code, or iterating with AI. From idea to product, Figma...  ...scale, we're looking for a Director of IT to lead and grow three...  ...IT infrastructure, ensuring reliability, scalability, and security across... 
    Platform
    Minimum wage
    Full time
    Contract work
    Work at office
    Local area
    Remote work
    Flexible hours

    Figma

    San Francisco, CA
    4 days ago
  • $180.2k - $252.3k

     ...work ensures Reddit’s ads products function reliably in-market while supporting GTM readiness...  ...API setupsExperience leveraging AI tools, machine learning workflows, or AI-...  ...discrepanciesExperience with a data management platform (ex. Hex, Tableau, Looker, Sense)Strong cross... 
    Platform
    For contractors
    Work experience placement
    Work at office
    2 days per week
    1 day per week

    Reddit

    San Francisco, CA
    1 day ago
  •  ...a hands-on technical leader with over 7 years of experience in AI and robotics, capable of managing teams and driving strategic ownership...  ...will bridge AI and physical systems, ensuring safety and reliability in autonomous operations. This is a full-time role within a... 
    Platform
    Full time

    Strativ Group

    San Francisco, CA
    5 hours ago
  • MaintainX is the world's leading AI-powered maintenance and asset management platform, serving 14,000+ customers including Duracell, Shell, Cintas, and Brenntag...  ...our applications. You’ll help teams deliver fast, reliable, and relevant search without reinventing ingestion... 
    Platform
    Immediate start
    Worldwide
    Flexible hours

    MaintainX

    San Francisco, CA
    4 days ago
  • $230k - $340k

     ...built. As the team behind Next.js, v0, and AI SDK, we create products that help...  ...and operated by agents.We are building the platform for that future, trusted by companies like...  ...infrastructure, performance, and production reliability.We are looking for an Engineering Team Lead... 
    Platform
    Work from home
    Worldwide
    Flexible hours

    Vercel

    San Francisco, CA
    3 days ago
  • $104.9k - $199.07k

     ...seeking a Technical Product Manager with a platform engineering focus to join our team of...  ...robust, scalable healthcare analytics and AI products. This role will help shape the strategy...  ...MedInsight solutions more efficiently, reliably, and securely.The Technical Product... 
    Platform
    Full time
    Work experience placement
    Remote work
    Flexible hours

    Milliman

    San Francisco, CA
    1 day ago
  • $230k - $287.5k

     ...About the Director, AI Architect at Headspace At Headspace, our mission is to transform mental...  ...across every layer of our product and platform. Reporting directly to the Chief...  ...personalization, with a relentless focus on reliability, safety, and member impact. Partner directly... 
    Platform
    Full time
    Work at office
    Local area
    3 days per week

    Headspace

    San Francisco, CA
    4 days ago
  • $180.8k - $226k

     ...in a hyper-growth, demanding AI environment, you will translate...  ...our engineering teams deliver reliable, high-value solutions at scale...  ...communication across core teams (e.g., Platform, Forward Deployed Engineering,...  ..., subject to Board of Director approval. Your recruiter can... 
    Platform
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  •  ...information.As a Product Manager on the AI Foundations team, you will drive Plaid’s...  ...who thrives at the intersection of AI and platform products. You are technically fluent, strategic...  ...machine-learning capabilities into reliable, trusted infrastructure that scales across... 
    Platform
    Work experience placement
    Local area

    Plaid Financial

    San Francisco, CA
    3 days ago
  • $172.2k - $258.4k

     ...Databricks is building an Agentic Enterprise applications Platform a scalable, governed AI application platform built on Databricks that enables...  ...while maintaining enterprise grade quality, security, and reliability. You will partner closely with Application Engineering,... 
    Platform
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • $185k - $317k

     ...to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether...  ...translating designs into code, or iterating with AI. From idea to product, Figma empowers...  ...inference latency, throughput, cost, and reliability)Desktop, browser, native mobile (React... 
    Platform
    Minimum wage
    Full time
    Local area
    Remote work
    Flexible hours

    Figma

    San Francisco, CA
    1 day ago
  • $365k

     ...Manager Anthropic is investing in shared platform capabilities that power Claude across our...  ...shipping more easily, faster, and more reliably because they're built on a common...  ...interested in the challenges of bringing AI capabilities to users safely and reliably... 
    Platform
    Work at office
    Visa sponsorship
    Flexible hours

    Colorwave Inc

    San Francisco, CA
    3 days ago
  • $202k - $253k

     ...hyperscaler for the edge, delivering modular AI infrastructure from first deployment to...  ..., data centers, managed compute, AI platforms, or high‑performance infrastructure. Experience...  ...account notes, sharp qualification, reliable follow‑through. Executive presence and the... 
    Platform
    Flexible hours

    Armada

    San Francisco, CA
    5 days ago
  • $195k - $209k

     ...actionable insights across sectors like software, AI, cloud, e-commerce, ridesharing, and...  ...shape the evolution of our Central Data Platform. In this role, you'll define and...  ...platform capabilities that enable scalable, reliable, and reusable data solutions supporting both... 
    Platform
    Work at office
    Remote work
    Flexible hours

    GrabJobs

    San Francisco, CA
    4 days ago
  • $145.75k - $300.07k

     ...Millions of people around the world come to our platform to find creative ideas, dream about new...  ...you love? It’s Possible.At Pinterest, AI isn't just a feature, it's a powerful...  ...that powers company-wide decisions fast and reliable.Own measurement of GenAI-driven... 
    Platform
    Work at office
    Local area
    Remote work
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    3 days ago
  • $200k - $240k

     ...intelligence. As the only vertically integrated AI infrastructure company built from the...  ...compromising on sustainability or reliability. Crusoe Cloud is 1,400 people and growing...  ...how it already works.The Managed Inference platform is where customers run production LLM workloads... 
    Platform
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  •  ...work ensures Reddit’s ads products function reliably in-market while supporting GTM readiness...  ...Experience with a data management platform (ex. Hex, Tableau, Looker, Sense) 3+ years...  ...support role or similar Experience leveraging AI tools, machine learning workflows, or AI-... 
    Platform
    Home office
    Flexible hours

    Reddit

    San Francisco, CA
    4 days ago
  • $173k - $215k

     ...customers, with particular depth on the data platform that the entire product portfolio is...  ...quality checkpoints that improved delivery reliability across the data platform.Earned trust as...  ...executive-level reporting.Expectations of AI Use in this role (required):Expected to... 
    Platform
    Contract work
    For contractors
    Work experience placement
    Work at office
    Local area
    Flexible hours

    Komodo Health

    San Francisco, CA
    4 days ago
  • $185k - $317k

     ...to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether...  ...translating designs into code, or iterating with AI. From idea to product, Figma empowers...  ...infrastructure initiatives, including reliability, storage, distributed systems, cloud-native... 
    Platform
    Minimum wage
    Full time
    Local area
    Remote work
    Flexible hours

    Figma

    San Francisco, CA
    4 days ago
  • $220k - $302.5k

     ...About FaireFaire is a technology wholesale platform built on the belief that the future is...  ...nearly every product decision we make, and AI data agents are how that insight shows up...  ...them, and hold the bar on quality, reliability, and adoption. You will also establish the... 
    Platform
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    2 days ago
  •  ...StripeStripe is a financial infrastructure platform for businesses. Millions of companies—...  ...infrastructure for Stripe to build secure, reliable, and differentiated products, while...  ...working in a fast-changing environment as the AI tool chain continues to evolve.Experience... 
    Platform
    Flexible hours

    Stripe

    San Francisco, CA
    1 day ago
  • $124k - $155k

     ...focused on building the foundation for an AI-native future across our three core...  ...connecting frontier model research, internal platform investments, operator teams, and strategic...  ...custom foundation models that ensure safety, reliability, and long-term leverage.Operationalizing... 
    Platform
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    4 days ago
  • $365k

     ...functional programs Have experience with API platforms, distributed systems, and core...  ...including the trade‑offs across performance, reliability, and cost Have coordinated programs...  ...interested in the challenges of operating AI infrastructure safely and reliably Compensation... 
    Platform
    Work at office
    Flexible hours

    Anthropic

    San Francisco, CA
    3 days ago
  • $150.9k - $226.3k

     ...will lead Harvey's response to security and reliability incidents, coordinating across...  ...detect, contain, and remediate threats to the platform that serves law firms and enterprises. What...  ...growth and reflect the realities of operating AI‑native SaaS infrastructure. Lead post‑... 
    Platform

    Harvey

    San Francisco, CA
    4 days ago
  •  ...ecosystem powering the next generation of AI products. We build the infrastructure,...  ...just possible, but practical: a unified platform where high-performance inference, orchestration...  ...and open escalations are tracked in one reliable system. Invoice and capacity... 
    Platform
    Contract work

    features and labels

    San Francisco, CA
    5 days ago
  • $196k - $242k

     ...applied to a range of vehicle platforms and product use cases. The...  ...reports and you will report to a Director of Technical Program...  ...loop, establishing a seamless, reliable path for model training, offline...  ...technical field, ideally related to AI/ML The expected base... 
    Platform
    Full time
    Remote work

    Waymo

    San Francisco, CA
    5 days ago
  • $290k - $365k

    About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users...  ...managing decommissions across cloud providers and hardware platforms. Partner with engineering and research leadership to... 
    Platform
    Visa sponsorship

    Anthropic

    San Francisco, CA
    6 days ago
  • $365k

     ...the Role Anthropic is investing in shared platform capabilities that power Claude across our...  ...shipping more easily, faster, and more reliably because they are built on a common foundation...  ...interested in the challenges of bringing AI capabilities to users safely and reliably... 
    Platform
    Visa sponsorship

    Anthropic

    San Francisco, CA
    3 days ago
  • $10 per hour

     ...every shipment. Tally is the first billing platform that reasons, remembers, and understands...  ...operations, enterprise systems, and AI infrastructure. We’re backed by the best...  ...holding the team to a high bar on accuracy and reliability Go-to-market collaboration with the... 
    Platform
    Contract work

    Tally

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director, AI Platform Reliability. Be the first to apply!