Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Engineering Manager, Cloud Monitoring Services Platform

$215k - $260k

Crusoe

Job Description

Job Description

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

About the Role:

Crusoe builds cloud infrastructure for AI workloads. Cloud Monitoring Services owns observability across Crusoe Cloud: metrics, logs, alerting, and the telemetry agent that runs on every node in the fleet. Our customers run large, demanding AI training and inference workloads, and they depend on us for a clear, trustworthy view of what their infrastructure is doing.

We are hiring an Engineering Manager to lead the Platform team: the time series and log storage systems, and the query layer that serves every dashboard, API call, and investigation on top of them. This is a first-line management role reporting to the Engineering Manager for Cloud Monitoring Services. You will take direct people management responsibility for a team of 4 to 6 engineers, growing, and own how telemetry is stored, retained, and read back at fleet scale.

This is a people-first leadership role with real delivery stakes. You will bring the technical depth to guide hard calls and earn your team's trust, but your success is measured through what your team accomplishes. Query latency, retention, and storage cost are live tradeoffs your team will be making continuously, and customers feel all three directly. The role sits alongside peer managers who own telemetry collection and the ingestion pipeline, and together you run a platform that has to work end to end. This is a full-time position.

What You'll Be Working On:

  • Grow and develop your team. Manage 4 to 6 engineers directly: 1:1s, career growth, performance, and team health, with the team growing over time.

  • Own storage and query. Be accountable for the time series and log storage systems and the query layer on top of them, including how they behave under load and how much they cost to run.

  • Own delivery. Plan and sequence work across a roadmap that mixes customer-facing feature work with storage and query infrastructure, and make honest calls early when a plan is at risk.

  • Set the technical direction for your area. Partner closely with the Staff engineers who own technical direction across the platform, so they can focus on engineering rather than absorbing delivery and handoff work alone. Ask the hard questions and make sound tradeoff calls alongside your engineers.

  • Keep the operational bar high. Query is on the critical path when something is wrong in the fleet, so availability, query performance, and correctness of what comes back are core to the job, not afterthoughts.

  • Own the team's oncall rotation and the operational health of the storage and query stack.

  • Manage cost and scale together. Retention policy, downsampling, cardinality, and storage tiering are product decisions as much as engineering ones, and you will be in those conversations.

  • Coordinate across team boundaries. Your team consumes what the collection and ingestion teams produce and serves internal platform teams as well as external customers, so sequencing dependencies and escalating early is a core part of the job.

  • Shape the team over time. Work with recruiting on sourcing, run a high-quality interview loop, close strong candidates, and onboard them well.

  • Collaborate across functions. Partner with product, neighboring infrastructure teams, peer managers, and leadership to keep priorities aligned as the observability product expands.

What You'll Bring to the Team:

We know great engineers and leaders come from many different paths. If you're excited about this work but don't match every point below, we'd still love to hear from you.

  • You care about people. You want the engineers around you to grow, you give feedback that's both direct and kind, and you measure your own success through your team's.

  • Technical depth. 5+ years of hands-on engineering experience, ideally in backend or distributed systems: databases, storage engines, query engines, streaming pipelines, or other data-intensive services (Go, Rust, Java, or C++ environments), and hands-on familiarity with Kubernetes. This role stays close to the technical work.

  • Experience managing engineers. 2+ years managing software engineers directly, including owning performance cycles and career conversations. This team has an active roadmap and live customer commitments, so we are looking for someone who has run a team through delivery before.

  • Experience coaching senior and staff-level engineers, and comfort partnering with strong technical leads rather than competing with them.

  • A track record of shipping. You've delivered multi-phase projects against fixed deadlines, and you know how to keep infrastructure work moving alongside feature work rather than letting one crowd out the other.

  • Operational judgment. You've run a team that owns a system customers depend on during an incident, and you know what it takes to keep that system trustworthy.

  • Strong communication and judgment. You can align people across functions, explain tradeoffs to technical and non-technical partners alike, and give your team clarity about what matters and why.

Bonus Points

  • Experience with time series databases or columnar storage at scale: Prometheus, Thanos, Mimir, VictoriaMetrics, InfluxDB, ClickHouse, or similar

  • Experience with query engines and query cost control: planning, pushdown, caching, concurrency limits, and protecting a cluster from expensive queries

  • Experience with retention, compaction, downsampling, and storage tiering for large telemetry or event datasets

  • Familiarity with log storage and search systems: Loki, Elasticsearch, OpenSearch, or similar

  • Observability domain experience: OpenTelemetry, Grafana, metrics and log pipelines, cardinality management

  • Experience owning multi-tenant systems with per-tenant isolation, quotas, and fairness

  • Experience running software in GPU or accelerated computing environments

  • Experience managing or scaling a team through growth

Benefits:

  • Competitive compensation and equity packages

  • Restricted Stock Units

  • Paid time off, paid holidays & leave of absence programs

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

  • Global travel insurance & emergency assistance

  • Daily meals allowance

  • Additional perks & programs specific to location

Compensation Range

Compensation will be paid in the range of up to $215,000 -$260,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Vacancy posted 8 days ago
Similar jobs that could be interesting for youBased on the Engineering Manager, Cloud Monitoring Services Platform in San Francisco, CA vacancy
  • $215k - $260k

     ...data center construction, and cloud services. If you want to do the...  ...seeking a Staff Software Engineer to lead distributed systems...  ...within Crusoe Cloud's Cloud Monitoring Service. This team builds the...  ...product, compute, networking, and platform teams to make sure... 
    Platform
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    a month ago
  • $148.5k - $223.9k

     ...Category Software Engineering About Salesforce...  ...The Experience Cloud Atlas is a globally...  ...between Salesforce services. We're looking for an engineering manager to lead a friendly...  ...APIs), logging and monitoring, performance...  ...and orchestration platforms, such as Docker and... 
    Platform
    Temporary work
    Remote work

    Jobleads-US

    San Francisco, CA
    14 hours ago
  •  ...software is the moment all of engineering's work becomes real for customers...  .... We're looking for a Senior Manager to build and drive Atlassian Government Cloud platform components, especially building...  ...product teams and hundreds of service teams, own the go/no-go decision... 
    Platform
    Work at office
    Local area
    Shift work

    Atlassian

    San Francisco, CA
    4 days ago
  •  ...Salesforce seeks an engineering manager to lead a high-impact team responsible for the Experience Cloud Atlas authentication infrastructure. You will guide engineers delivering...  ..., logging, and highly available datastore services, using Java, REST/RPC, and cloud-native... 
    Platform

    Jobleads-US

    San Francisco, CA
    14 hours ago
  •  ...sustainable growth. This role The Senior Engineering Manager, Cloud Infrastructure leads the engineering team responsible...  ...the company’s cloud infrastructure, developer platform, and foundational engineering services. This role sets the technical and organizational... 
    Platform

    PayJoy

    San Francisco, CA
    a month ago
  • $150.1k - $227k

     ...investment. As a Customer Success Manager aligned to our Technology,...  ...strategic priorities, platform adoption, and long-term value...  ...to driving adoption of Sales Cloud and Service Cloud - so customers see measurable...  ...'s goals and roadmap. Monitor platform health, track... 
    Platform

    100 Salesforce, Inc.

    San Francisco, CA
    15 hours ago
  • $175k - $210k

     ..., data center construction, and cloud services. If you want to do the most...  ...We are seeking a Senior Software Engineer to build and own customer-facing...  ...within Crusoe Cloud's Cloud Monitoring Service. This team delivers the managed logs experience and self-service... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    6 days ago
  • $124k - $280k

     ...ManagerJob Description & SummaryThe OpportunityAs a Senior Engineering Manager, Google Cloud AI/ML you will lead engineering efforts that design,...  ...reliable AI behavior- Shaping LLMOps practices for versioning, monitoring, release coordination, and lifecycle management-... 
    Full time
    H1b

    PwC

    San Francisco, CA
    1 day ago
  •  ...Alo in San Ramon, CA seeks an Engineering Manager, Platform Engineering to lead the core platforms, cloud infrastructure, and developer tooling for a high-scale eCommerce ecosystem. You will balance technical leadership with hands-on engineering to guide architecture and... 
    Platform

    Jobleads-US

    San Francisco, CA
    1 day ago
  •  ...time ever, you can manage and automate every...  ...experience is a platform that allows teams...  ...looking for a Senior Engineering Manager to help...  ...for Rippling's core cloud infrastructure and...  ...and platform self-service capabilitiesEstablish...  ...observability, monitoring, incident response... 
    Platform
    Work at office
    3 days per week

    Rippling

    San Francisco, CA
    2 days ago
  • $188k - $235k

     ...seeking anEngineering Manager, Platform Engineeringto lead...  ...operating thecore platforms, cloud infrastructure, and...  ..., and enable engineering teams to deliver reliable...  ...practices, including monitoring, logging, alerting, and...  ..., reliability, and service ownership practices.... 
    Platform

    Alo

    San Francisco, CA
    1 day ago
  • $237.6k - $285k

     ...manufacturing, data center construction, and cloud services. If you want to do the most...  ...high value workloads. As a Senior Engineering Manager, you will lead the team responsible for...  ...next generation, bespoke cloud storage platform optimized for AI and HPC workloads.... 
    Platform
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    a month ago
  • $181.22k - $217.46k

     ...they love. Fastly’s edge cloud platform enables customers to...  ...applicants. Senior Engineer - Production Cloud and Container Services Fleet Operations and...  ...helping to scale and manage Fastly’s Kubernetes based...  ...includes effective monitoring and resilient operations... 
    Platform
    Full time
    Work at office
    Local area
    Remote work
    Flexible hours
    Night shift

    Fastly

    San Francisco, CA
    9 days ago
  •  ...time ever, you can manage and automate every...  ...every product and service across the company...  ...more than 1,000 engineers and millions of end...  ...internal tooling, platform self-service capabilities...  ..., and error-monitoring infrastructure...  ...distributed systems and cloud-based production... 
    Platform
    Work at office
    3 days per week

    Rippling

    San Francisco, CA
    1 day ago
  • $138.1k - $196.5k

     ...TeamJoin Cisco's AI Platform Group to...  ...generative AI and cloud-native SaaS solutions...  ...building an AI-native engineering culture that...  ...leaders, product managers, and applied AI researchers...  ...reliability and monitoring requirements...  ...practices, cloud services, and APIs to... 
    Platform
    Full time
    Temporary work
    Internship
    Local area
    Flexible hours

    CISCO Systems

    San Francisco, CA
    18 hours ago
  •  ...Salesforce, Inc. is seeking a Manager of Software Engineering for Cloud Atlas in a hybrid role across SF and Seattle-area locations. You will lead...  ...impact team responsible for APIs, authentication datastore services, and AI-driven development practices. The role... 

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $171.6k - $245k

     ...on-premises, hybrid-cloud, and multi-cloud...  ...application performance monitoring and full-stack...  ...As a Senior Product Manager, you will work with engineering, design, sales, marketing...  ..., the Splunk platform, Splunk Observability...  ...Cloud, Splunk IT Service Intelligence, and relevant... 
    Platform
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    2 days ago
  • $236.7k - $319.5k

     ...a Sr Director - AI & Platform Engineering you are accountable for...  ...procedural knowledge management, and the data access...  ..., and context/memory services that enable reliable,...  ...gates, performance monitoring, and continuous improvement...  ...Governance, Google Cloud, and executive... 
    Platform
    Minimum wage

    Gap

    San Francisco, CA
    2 days ago
  • $188k - $235k

     ...Engineering Manager, Performance Engineering San Ramon, California...  ...of our web and mobile platforms. This role partners closely...  ...across applications, services, infrastructure, and cloud platforms, ensuring performance...  ..., benchmarking, and monitoring tools. ~ Ability to... 
    Platform

    Alo

    San Francisco, CA
    31 minutes ago
  • $198k - $264k

     ...CoreWeave is The Essential Cloud for AI™. Built for...  ...CoreWeave delivers a platform of technology, tools,...  ...The Technical Support Engineering - Infrastructure team...  ...the Role As the Senior Manager of Technical Support Engineering...  ...~ Flexible, full-service childcare support with... 
    Platform
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    25 days ago
  • $285k - $335k

     ...center construction, and cloud services. If you want to...  ...As the Director of Engineering - Cloud Compute, you...  ...reducing, AI focused cloud platform. Leveraging your...  ...the team, including managing engineer onboarding,...  ...wide observability and monitoring to drive reliability,... 
    Platform
    Temporary work

    Crusoe

    San Francisco, CA
    23 days ago
  • $220k

     ...is today’s application monitoring standard and our team...  ...Sentry's Infrastructure Engineering team is what makes...  ...the internal control platforms, configuration systems...  ...product engineers operate services safely at scale...  ...As the Engineering Manager for Infrastructure Engineering... 
    Platform
    Hourly pay
    Remote work

    Sentry

    San Francisco, CA
    3 days ago
  •  ...orchestration and payments platform that helps large enterprises...  ...DEUNA is looking for an Engineering Manager to lead our AI/ML engineering...  ...depth in ML systems, backend services, and AI/LLM workflows to...  ...work (training, evaluation, monitoring, retraining) and LLM-powered... 
    Platform
    Full time
    Shift work

    Deuna

    San Francisco, CA
    4 days ago
  •  ...Rippling, Inc. is seeking a Senior Engineering Manager to lead a team within Platform. You will shape the technical direction of foundational systems used across Rippling and partner with engineering and product leaders to balance long-term platform investments with near... 
    Platform

    Jobleads-US

    San Francisco, CA
    14 hours ago
  •  ...Native across iOS / Android / Web but will include more platforms and frameworks in the future. Build a theming / "...  ...migration systems, business logic, deployments, ranking services, ACL / access control, monitoring etc Extend the capabilities of the Node.js "... 
    Platform

    Hired Recruiters

    San Francisco, CA
    3 days ago
  •  ...accounts to the apps and services they want to use....  ...consumers understand and manage their finances, to how...  ...AI applications and platforms in this evolving ecosystem...  ...will lead a team of 4 engineers, ranging from junior...  ...~ Evaluation and monitoring framework of open-ended... 
    Platform
    Full time
    Work experience placement
    Local area
    Shift work

    Plaid

    San Francisco, CA
    6 days ago
  • Senior Cloud, AI & Data Security Engineer We are seeking an enthusiastic and...  ...for systems and services across AWS, Azure, and AI/ML platforms. We need someone who...  ...provide assurance to management and auditors, and ensure...  ...responsibility of monitoring, detecting, protecting... 
    Platform

    Bmo

    San Francisco, CA
    5 days ago
  •  ...and AI inference service, making high-performance...  ...a Executive Engineering hire (SF, Hybrid)...  ...Infrastructure, Platform, and SRE...  ...evolution of the AI cloud platform architecture...  ...optimization, and capacity management Personally...  ..., observability, monitoring, and reliability... 
    Platform

    Hyphen Connect

    San Francisco, CA
    a month ago
  • $182k - $242k

     ...powerful end-to-end platform to develop, deploy, and...  ...’s industry-leading cloud infrastructure with the...  ...products and services enabling faster and more...  ...enterprises to trace, monitor, and evaluate their AI...  ...the Role As an Engineering Manager on Weave’s Enterprise... 
    Platform
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    Weights & Biases

    San Francisco, CA
    4 days ago
  •  ...Job Description VP of Engineering Applications, Artificial...  ...backend, mobile, and desktop platforms. The VP of Engineering Applications...  ...readiness: observability, monitoring, retries, fallbacks, privacy...  ...system design skills across services, APIs, clients, and data... 
    Platform
    Full time
    Remote work
    Work from home

    Next Step Systems

    San Francisco, CA
    12 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Engineering Manager, Cloud Monitoring Services Platform. Be the first to apply!