Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Cluster Infrastructure Engineer

$224k - $284k
Full-time

ATOMS Careers page

Who we are

Atoms is building the machines that power the next era of progress.

Over the last decade, software has transformed the digital world. But the physical world, where food is made, minerals are mined, goods are moved, and industries are run, remains far less intelligent, far less efficient, and far more constrained. We’re changing that.

Atoms builds Physical AI - real-world robots for the industries that move civilization forward, starting with food, mining, and transport. Our systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive.

This work requires more than robotics. It requires deep integration across hardware, software, AI, operations, manufacturing, and real estate. We don’t just build machines in a lab. We deploy them into real environments, operate them, learn from them, and improve them until they work at scale.

We are roboticists, engineers, operators, and builders. We believe the next great technology companies will not only transform information, but the physical systems that shape everyday life.

If you want to work on hard problems with real-world impact, join us.

What you’ll do We’re seeking a Staff Cluster Infrastructure Engineer to join our founding team who will own the GPU compute fabric that trains our foundation models - optimizing the machines we have today, automating how we manage them, and laying the groundwork to scale as we grow.

  • Manage and automate our GPU training clusters, including provisioning, bootstrapping, and lifecycle management.
  • Automate bare-metal bring-up so new machines come online quickly and reliably as we add capacity.
  • Build software abstractions that present a clean, unified interface to our training and simulation workloads.
  • Work at the hardware/software boundary, where speed and reliability are critical, continuously raising the bar for automation and uptime.
  • Run day-to-day operations: diagnose and resolve issues quickly when systems are under pressure.
  • Design our infrastructure to scale smoothly as we grow from a smaller cluster of machines toward a larger fleet.

What we’re looking for

  • 6+ years experience operating GPU compute on Kubernetes (or similar orchestration), with the judgment to scale it as demand grows.
  • Strong programming and scripting skills in Python, Go, or similar.
  • Familiarity with Infrastructure-as-Code tools such as Terraform or CloudFormation.
  • Comfort with bare-metal Linux environments, GPU hardware, and networking.
  • A bias toward automation, reliability, and operating critical systems well.

Why join us

At Atoms, you’ll work on one of the defining challenges of our time — bringing automation into the physical world to drive real, lasting impact. We exist to uncover valuable unknown truths and turn them into progress, which means constantly pushing beyond what’s known and building what doesn’t yet exist. The work is ambitious and often challenging, but it’s grounded in a shared sense of purpose and a team committed to seeing it through together. Our work only matters if it serves others, and we know that meaningful progress depends on the trust of the people we serve and the strength of our team — so we invest in both, creating an environment where you can do your best work and grow.

What else you need to know

This role is based in our San Franciscooffice. Atoms is a company driven by invention and continuous change - we are constantly reimagining our industries, building new products, and refining how we operate. We do our best work together. That’s why all of our office-based teams work onsite, five days a week.

The base salary range for this role is $224,000 - $284,000 per year.

Actual compensation will be determined on an individual basis and may vary depending on experience, skills, and qualifications.

Base salary is just one part of your total rewards package. You may also be eligible for equity awards.

Benefits Summary (USA Full-Time Exempt Employees):

  • Medical, Dental, Vision, Disability, and Life Insurance
  • Flexible Spending Account / Health Savings Account Options
  • 401(k)
  • Equity
  • Sick Time, Unlimited Flexible Time Off, and Paid Holidays
  • Paid Parental Leave
  • Pre-Tax Commuter Benefit Plan
  • Team lunch in our SoMa office every Tuesday and Thursday

Benefits are subject to change at the company's discretion.
Atoms accepts applications on an ongoing basis.

Ready to join us as we serve those who serve others?

#LI-Onsite

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Staff Cluster Infrastructure Engineer in San Francisco, CA vacancy
  • $224k - $284k

     ...them until they work at scale. We are roboticists, engineers, operators, and builders. We believe the next great technology...  ...impact, join us. What you’ll do We’re seeking a Staff Cluster Infrastructure Engineer to join our founding team who will own the GPU... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Atoms

    San Francisco, CA
    3 days ago
  •  ...models. About the Role We are looking for engineers to operate the next generation of compute clusters that power OpenAI’s frontier research. This...  ...blends distributed systems engineering with hands-on infrastructure work on our largest datacenters. You will scale... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • StratITech is hiring a senior network engineer to architect and operate secure, high-performance data center and edge networks supporting AI and distributed compute workloads. You will translate loosely defined needs into concrete requirements and plans, and own the end... 
    Suggested

    StratITech

    San Francisco, CA
    4 days ago
  •  ...meaningful impact. As a Senior Lead Software Engineer - Windows Server Engineering at...  ...Chase within the Corporate Sector Compute Infrastructure Platform (CIP) organization, you...  ...budgets and experience in Windows Server Clusters. Experience in infrastructure virtualization... 
    Suggested

    J.P. Morgan

    San Francisco, CA
    1 day ago
  • OpenAI is seeking an experienced engineer for its Consumer Devices team in San Francisco to design and operate CI/CD pipelines and...  ...decisions for a growing platform while collaborating with product, systems, release, quality, and infrastructure teams. #J-18808-Ljbffr OpenAI
    Suggested

    OpenAI

    San Francisco, CA
    4 days ago
  • Discord is seeking a strategic Staff Product Marketing Manager to lead the go-to-market for Discord's Unified Developer Platform, helping...  ...programs, and cross-functional partnership across Product, Engineering, Business Development, Developer Relations, Marketing, and... 

    Discord Inc.

    San Francisco, CA
    5 days ago
  • $155.4k - $273.7k

     ...platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting...  ...Experience running containerisation in production (a real cluster, not a lab), with experience in Helm and Terraform or... 
    Full time
    Work at office
    Local area
    Shift work

    Writer

    San Francisco, CA
    3 days ago
  • $207k - $230k

     ...latest Whatnot updates on our news and engineering blogs and join us as we enable anyone...  .... Role We're looking for an Infrastructure Engineer to own the path between Whatnot...  ...traffic between our services, accounts, and clusters, and keep it secure as the topology... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Home office

    Whatnot

    San Francisco, CA
    2 days ago
  •  ...About the Role A well-funded early-stage Kubernetes infrastructure company is hiring a Frontend Engineer to design and build the interface for their...  ...visualizations, and real-time UIs that translate complex cluster state into clear, actionable experiences. This is a... 
    Remote work

    Clera

    San Francisco, CA
    26 days ago
  • A technology company is seeking a Cloud Infrastructure Engineer to build and maintain scalable infrastructure and provide reliable support for customers across AWS, Azure, and GCP. The ideal candidate will have 5+ years of experience in DevOps or Infrastructure Engineering... 
    Flexible hours

    Brain Trust Inc

    San Francisco, CA
    1 day ago
  •  ...collaboration through cutting-edge platforms that empower enterprises to evolve intelligently. The team is hiring a Head of Platform/AI Cluster Management to oversee the strategic development, integration, and optimization of AI and platform initiatives. The role will focus... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    1 day ago
  •  ...Capabilities, and Skills · 5+ years of hands-on experience in infrastructure engineering with demonstrated depth and breadth of knowledge across...  ...in Microsoft Hyper-V at enterprise scale: failover clustering, live migration, storage spaces / S2D, virtual networking... 

    PB consulting

    Colma, CA
    4 days ago
  • $240k

    Convex in San Francisco is seeking a Product Lead to shape the developer experience and guide product direction. This role requires a strong blend of product management skills, technical understanding, and customer engagement. Ideal candidates will have a track record of...

    Convex

    San Francisco, CA
    4 days ago
  • LiveKit is seeking an experienced Product Manager to lead a Growth squad focused on Product-Led Growth. You’ll own end-to-end growth, cultivate a large population of developers, and drive onboarding, activation, and production adoption. You’ll work with product, design,...
    Remote work

    LiveKit

    San Francisco, CA
    1 day ago
  •  ...(***) ***-****Job Title: Machine Learning Infrastructure EngineerLocation: San Francisco, CA Metro...  ...a Machine Learning Infrastructure Engineer to help architect the compute, training...  ...distributed systems across massive hardware clusters while collaborating directly with core... 
    Full time
    Work at office
    Flexible hours

    Objective Paradigm

    San Francisco, CA
    6 hours ago
  • $224k - $284k

     ...work at scale. We are roboticists, engineers, operators, and builders. We believe...  ...and the architecture to scale it as our cluster grows. Design, optimize, and scale...  .... Collaborate across the infrastructure team to solve cross-discipline problems... 
    Full time
    Work at office
    Immediate start
    Flexible hours

    ATOMS Careers page

    San Francisco, CA
    3 days ago
  • $190k - $280k

     ...RoleTogether AI is looking for a Senior Network Engineer to design, deploy, and operate the global network infrastructure supporting our production services and high-...  ...InfiniBand fabrics.Experience supporting GPU clusters, HPC environments, distributed storage, or other... 
    Full time

    Together AI

    San Francisco, CA
    5 days ago
  •  ...About the Team OpenAI’s Infrastructure organization builds the systems that power frontier AI...  ...required to bring compute online: server and cluster activation, storage platforms, Points...  ...availability, vendor execution, and engineering dependencies required to turn... 

    Neura Market

    San Francisco, CA
    3 days ago
  • $350k

     ...will and judgment. About the Role We're looking for a network engineer to own the lowest layers of the network stack that our large-scale...  ...(we use Python or Rust). Experience operating large-scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).... 
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab Inc.

    San Francisco, CA
    2 days ago
  • $150k - $200k

     ...mid-migration: we are moving off a single Redshift cluster onto an Apache Iceberg lakehouse on S3, with Flink CDC...  ...work the runway it deserves. Right now, our engineers wear two hats — owning the infrastructure and the business datasets running on top of it — and... 
    Full time
    Worldwide

    Aircall.io, Inc.

    San Francisco, CA
    1 day ago
  • $224k - $280k

     ...at scale. We are roboticists, engineers, operators, and builders. We believe...  ...We are seeking a foundational Staff Machine Learning Infrastructure Engineer to design and build the large...  ...jobs concurrently across large GPU clusters. Experiment Tracking & MLOps:... 
    Full time
    Work at office
    Flexible hours

    Atoms

    San Francisco, CA
    3 days ago
  •  ...ML Infrastructure EngineerSan FranciscoCompany OverviewEcho Neurotechnologies is an exciting...  ...driving innovation through advanced hardware engineering and AI solutions. Our mission is to...  ...model training infrastructure or multi-cluster computing environmentsWhat We OfferAn... 
    Flexible hours

    Echo Neurotechnologies

    San Francisco, CA
    4 days ago
  • $225k

     ...opportunity?  Join a rapidly scaling AI cloud infrastructure provider building next-generation GPU...  ...is looking for a Senior Network Engineer to design, deploy, and support ultra-low...  ...throughput, and resilience across GPU clusters Troubleshoot complex Layer 2/Layer 3... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $162k - $202k

     ...goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands...  ...trains on. Fabrics for 100k+ accelerator clusters, backbones between gigawatt campuses,...  ...could live with. You work with controls engineers as partners, not adversaries. Bonus:... 

    FluidStack

    San Francisco, CA
    5 days ago
  • Causal is building a Large Physics foundation Model and seeks an infrastructure engineer to design, deploy, and operate large distributed GPU clusters. You will extend schedulers, build self-serve interfaces, and own storage and lineage for checkpoints and logs. You will... 

    causal

    San Francisco, CA
    1 day ago
  • Join to apply for the HPC Engineer (Biohub Network) role at Jobright.ai 2 days ago Be among...  ...the Biohub. Responsibilities: • Manage cluster-level services via the SLURM scheduler...  ...• Experience building on-prem HPC infrastructure and capacity planning • Experience and... 
    Full time
    H1b

    jobright.com

    San Francisco, CA
    4 days ago
  • $175k - $240k

     ...ambitious team run by scientists and engineers from leading institutions across biology...  ...The Role As a Member of Technical Staff, Infrastructure Engineer, you'll play a key role in designing...  ..., implement, and operate Kubernetes clusters that support thousands of concurrent,... 
    Full time
    Work at office

    Edison Scientific

    San Francisco, CA
    3 days ago
  • $156.75k - $200k

     ...We're looking for a talented, senior engineering professional ready to take their career...  ...influential companies. As a Senior Lead Infrastructure Engineering at JPMorgan Chase within Enterprise...  ...Responsibilities Architect Hyper-V cluster designs across global pools, including... 
    Full time

    JPMorgan Chase & Co.

    San Francisco, CA
    2 days ago
  • $109k - $186k

     ...performance and scale, as well as executed flawlessly. As a Deployment Engineer, you will be responsible for translating customer...  ...Meter to scale and meet the growing demand for reliable internet infrastructure.You will have an impact byDesigning performant and resilient... 

    Meter

    San Francisco, CA
    4 days ago
  •  ...About the Team The Applied Engineering team works across research, engineering, product,...  ...team responsible for running the core infrastructure that supports products like ChatGPT and...  ...systems we support include our kubernetes clusters, infrastructure deployment, our... 
    Full time
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Cluster Infrastructure Engineer. Be the first to apply!