Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Customer Cluster Engineer

$220k - $320k

Jobleads-US

Customer engineering at iframe.ai is not solutions architecture and is not customer support. You will be embedded with three to five reserved-capacity accounts at a time, owning their training and inference performance end-to-end. The same engineers who tune the kernel pick up the page.

Stack PyTorch FSDP / Megatron-LM CUDA / Triton NCCL wandb / aim Linux perf InfiniBand familiarity

Hiring manager replies within 5 business days.

The team

About the team

Customer engineering is a small team of senior engineers — most have prior runtime, SRE, or research experience. You're paired with a single account manager who handles commercial; you handle the technical relationship from kickoff through renewal.

Reports to the head of customer engineering. Paired one-to-one with a sales engineer / AE for commercial.

The role

What you'll do

  • 01 Own three to five reserved-capacity accounts as their named customer engineer. Most are AI labs or AI-native scale-ups running 256–2048-GPU jobs.

  • 02 Profile distributed training jobs: NCCL collectives, gradient overlap, checkpoint cost, MFU. Make and defend specific recommendations.

  • 03 Tune kernels and configurations alongside the runtime team — your changes go through the same review process and ship in the same release train.

  • 04 Run the technical pre-sales for your accounts' expansions and renewals: pilot design, scope, success criteria, post-mortem.

  • 05 Lead the post-mortem on any P1 affecting your accounts; escalate root-cause work into the runtime, cluster, or platform teams.

  • 06 Carry the customer-engineering on-call rotation alongside runtime and cluster SRE — about one week per six.

The bar

What we're looking for

Five-plus years of distributed-systems or ML-systems engineering — production experience with 64+ GPU jobs is required.

Strong PyTorch / FSDP / Megatron-LM debugging skills. You can read a wandb run and tell us where time is going.

Working knowledge of CUDA and at least one of Triton or CUTLASS. You don't need to write them daily; you need to read them.

Strong communicator. You'll write four post-mortems a quarter and present at customer all-hands.

Comfort with executive-level conversations on technical scope, timelines, and trade-offs.

Nice to have, not required

Prior frontier-lab or hyperscaler training-team experience.

Experience running benchmarks at customer sites (MLPerf or similar).

Open-source contributions to PyTorch, FSDP, NCCL, or the major training frameworks.

Comfort with Slack-based customer relationships and high-tempo asynchronous communication.

Compensation

In writing, like everything else

We publish bands. We meet them. The number you see on the offer is the same number your future peers got at the same level. We do not negotiate; we level.

Base

$220,000 – $320,000 USD (US, IC4 / IC5).

Equity

Meaningful early-stage equity, refreshed on tenure milestones.

Notes

No revenue-tied compensation. Customer engineers are not compensated on the size or renewal of the accounts they own; they are compensated on technical impact.

How to apply

One email is enough

Send a short note to View email address on click.appcast.io with the role title in the subject line. Include your CV or LinkedIn, one or two links to work you're proud of, and a sentence on why this role specifically. Hiring managers reply within five business days, regardless of outcome.

01

Application

A hiring manager reads every email. Reply within five business days.

02

30–45 minutes. Scope, role, mutual fit. We share the comp band on this call.

03

Technical loop

3–4 sessions on the same day. Real problems, no homework, no whiteboard riddles.

04

Offer

Same-week offer at the published band for your level. Start dates are flexible.

Equal opportunity

We hire on the work. Race, gender, age, nationality, religion, sexual orientation, disability, and veteran status do not factor into our decisions. We sponsor visas for senior roles in the US, UK, and EU — bring it up on the manager call.

Need an accommodation for the interview process? Mention it in your application or write View email address on click.appcast.io .

If this role isn't quite right but you'd be a fit at iframe.ai, write anyway.

Senior engineers and researchers can apply outside the listed roles. The bar is the same. The reply window is the same.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Customer Cluster Engineer in San Francisco, CA vacancy
  • $170k - $245k

     ...ML application from their laptop to the cluster without needing to be a distributed systems...  ...the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations...  ...be used by open-source Ray users and customers of AnyscaleWork across the stack integrating... 
    Customer
    Work at office

    Anyscale

    San Francisco, CA
    1 day ago
  • $140k - $180k

     ...actuators, electronics, and software. Our custom-built supervision stack helps handle the...  ...AI. Getting nerd-sniped by an engineering metric is way less important than solving...  ...cable tester, adapter PCB, or ToF sensor cluster, then bring it up, debug it, and make sure... 
    Customer
    Internship
    Shift work

    Reflex Robotics

    San Francisco, CA
    2 days ago
  • $168.75k - $270k

     ...seeking out diverse perspectives from our customers, communities and each other.Life at Axon...  ...ImpactAs a PrincipalSecurity Operations Engineer, you will play a key role in building...  ...Kubernetes and container ecosystems, including cluster hardening, workload security, RBAC,... 
    Customer
    Work experience placement
    Work at office

    Axon

    San Francisco, CA
    3 days ago
  • $206.3k - $388k

     ...THE ROLE We’re looking for a Principal ML Engineer to architect and scale the multimodal...  ...and storage choices, job scheduling, GPU cluster utilization that let the platform scale...  ...creativity, productivity and personalized customer experiences. Adobe’s industry-leading... 
    Customer
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    14 hours ago
  • $152.5k - $205k

     ...responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design...  ...cost-effectiveness of the systems our customers depend on.What you'll work on: Design,...  ..., and troubleshooting production clusters and containerized workloads at scale.Strong... 
    Customer
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical...  ...build and ship products to delight our customers. We manage the end-to-end lifecycle of our...  ...the critical components that ensure cluster reliability and security (e.g., CoreDNS,... 
    Customer
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    14 hours ago
  •  ...Title: Software Engineer (C++ Systems) Location: San Francisco, CA, United States (Onsite...  ...platform and serving a rapidly growing customer base. What You'll Do Performance optimization...  ..., checkpointing, and distributed GPU clusters. Support new architectures with deep... 
    Customer
    Work experience placement
    Relocation
    Weekend work

    SK HR Consultants.com

    San Francisco, CA
    4 days ago
  • $194k - $267k

     ...reducing costs, and doing more for your customers.If you like to be challenged and have a...  ....Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and...  ...-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads,... 
    Customer
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  •  ...knowledge into action. We’re looking for engineers who want to build the operating system...  ...and implement core components of our cluster scheduler to improve resource utilization...  ...entire product development process, from customer research to implementation This... 
    Customer

    Tensorlake Inc.

    San Francisco, CA
    14 hours ago
  •  ...are seeking a skilled Site Reliability Engineer to join the GMI Global Infrastructure team...  ...the large-scale high performance AI/ML clusters in our data center. The ideal candidate...  ...GPU, CPU, Storage, Network). # Handle customer provisioning requests for GPU resources,... 
    Customer

    GMI Cloud

    San Francisco, CA
    20 hours ago
  •  ...team at UCSF is seeking an HPC Systems Engineer to play a key role in the development, maintenance...  ...-day operations of the Institute's HPC clusters. The HPC Systems Engineer will: Apply...  ...in support of onboarding research customers, or making systems improvements. Department... 
    Customer
    Work experience placement
    Work at office
    Worldwide

    University of California , San Francisco

    San Francisco, CA
    2 days ago
  • $250k

     ...The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering...  ..., observability, and monitoring frameworks for GPU compute clusters Collaborate with ML, data, and platform engineering teams to... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $300 per month

     ...meaningful work of your career, help our customers and partners advance their AI...  ...Crusoe Energy Systems, our Production Engineering (PE) team plays a mission-critical role...  ...planes) that support large-scale AI compute clusters.Your responsibilities will also include... 
    Customer
    Temporary work

    Crusoe

    San Francisco, CA
    6 hours ago
  • $153k - $191.3k

     ...company and data company all rolled into one.Customers and users across the globe use Planet's...  ..., data processing, and software engineering, our office is a truly inspiring mix of...  ...resource optimization, management, and cluster tuning in a constrained environmentAbility... 
    Customer
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Remote work
    Home office
    3 days per week

    Planet Labs PBC

    San Francisco, CA
    1 day ago
  •  ...Description Akka is seeking a Forward Deployed Engineer (FDE) who perfectly balances sharp...  .... Your success is measured by the customer's technical self-sufficiency and the...  ...Specific hands-on experience with Akka (Clustering, Persistence, and Streams) is highly desirable... 
    Customer
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Akka

    San Francisco, CA
    18 days ago
  •  ...Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco...  ...Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel...  ...and operate Kubernetes-based clusters for customers across multiple providers. Build... 
    Customer
    Full time
    Remote work

    Andromeda Cluster

    San Francisco, CA
    1 day ago
  •  ...standard. Role Scope Own reliability for named customer workloads: their clusters, their SLAs, their escalations. Debug across the...  ...depth and no spin. Turn recurring customer pain into engineering fixes with the production teams. What We’re Looking... 
    Customer

    Fluidstack

    San Francisco, CA
    4 days ago
  • $150k - $175k

    As a Customer Success Engineer at vCluster, you aren't just managing accounts — you are the primary architect of customer outcomes after the deal...  ...adoption into business outcomes. You connect Tenant Cluster growth, provisioning time savings, and developer self-service... 
    Customer
    Remote work
    Flexible hours
    Day shift

    Vcluster

    San Francisco, CA
    5 days ago
  • $209k - $253k

     ...meaningful work of your career, help our customers and partners advance their AI strategies...  ..., and our Compute-focused Production Engineers are the backbone of that mission. This role...  ...and workload orchestration across GPU clusters.Contributions to Linux kernel, KVM, or other... 
    Customer
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $170k - $205k

     ...meaningful work of your career, help our customers and partners advance their AI strategies...  ...Role:As a Senior Deployment Automation Engineer for the Compute Team, you will be responsible...  ...of large-scale, multi-node GPU clusters. You will develop the CI/CD infrastructure... 
    Customer
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  •  ...Hamilton Barnes is seeking a Senior Storage Engineer to own the high-performance storage layer for large-scale GPU clusters and AI workloads. You will design, deploy, and operate storage platforms, collaborating with infrastructure, compute, and networking teams to scale... 
    Remote job

    Jobleads-US

    San Francisco, CA
    2 days ago
  • $200k - $260k

     ...This enables the creation and operation of customer instances in our ecosystem in a...  ...reliability practices across teams, mentor senior engineers, and be a primary escalation point for...  ...production-scale multi-tenant or multi-cluster environments ~10+ years coding... 
    Customer
    Casual work
    Work at office
    Remote work
    Flexible hours

    Sight Machine

    San Francisco, CA
    4 days ago
  • $145k - $160k

     ...specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and...  ...provisioning and automation Agent & Cluster Agent deployment/management Monitors,...  ...inclusion in creating success for our customers, business partners, shareholders, employees... 
    Customer
    Temporary work
    Remote work
    Flexible hours

    EPAM Systems Inc

    San Francisco, CA
    3 days ago
  •  ...in assessment of their operation. Provides Pre-Market Quality Engineering support to New Product Development (NPD) working in partnership...  ...services, prevent defects, and allow Healthcare to provide customers with the highest quality and reliable products. Provides technical... 
    Customer

    Katalyst Healthcares & Life Sciences

    San Francisco, CA
    14 hours ago
  • # Staff Mechanical Engineer· San Francisco Office (Fremont St)Critical FacilitiesFullTimeremotePosted...  ...serving tens of thousands of customers. Our customers range from AI researchers...  ...schedule. - Translate GPU, rack, and cluster thermal requirements into facility cooling... 
    Customer
    Work at office
    Local area
    Work from home
    Flexible hours

    DataCenterGuidelines

    San Francisco, CA
    2 days ago
  •  ...and out of the office. We are looking for a Senior CMF Quality Engineer to join our Supply Chain team in San Francisco, California. In...  ...quality and aesthetic consistency on behalf of Oura and our customers.You'll collaborate closely with our CMF Design, Industrial Design... 
    Customer
    Work at office
    Local area
    Remote work
    Flexible hours
    Shift work
    3 days per week

    Oura

    San Francisco, CA
    1 day ago
  • $144.5k - $180.6k

     ...are both a space company and data company all rolled into one.Customers and users across the globe use Planet's data to develop new technologies...  ...hardware design, manufacturing, data processing, and software engineering, our office is a truly inspiring mix of experts from a variety... 
    Customer
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Immediate start
    Remote work
    Home office

    Planet Labs

    San Francisco, CA
    4 days ago
  • $165k - $200k

     ...meaningful work of your career, help our customers and partners advance their AI strategies...  ...a Senior Network Production Operations Engineer to support production reliability across...  ..., backbone, data center fabric, and GPU cluster interconnects. This is a hands-on production... 
    Customer
    Temporary work
    Worldwide

    Crusoe

    San Francisco, CA
    14 hours ago
  •  ...Automated Operations, our managed platform for customer workloads on dedicated, BYOC, and BYOK8s...  ..., working alongside the senior engineers already on the team. You'll contribute to...  ...and a custom resource's status. Customer clusters on AWS, Azure, and GCP are provisioned with... 
    Customer
    Remote work
    Flexible hours

    Akka

    San Francisco, CA
    a month ago
  • $188k - $275k

     ...Learn more at  What You'll Do: The Field Engineering organization at CoreWeave is dedicated to ensuring every customer running AI workloads at scale has a seamless,...  ...entire customer lifecycle: leading new GPU cluster bring-up and acceptance, driving InfiniBand/... 
    Customer
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Customer Cluster Engineer. Be the first to apply!