Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Engineer - Infrastructure

Flexion Robotics

Job Description

Job Description

About Flexion

At Flexion, we are building the autonomy stack for humanoid robots. Our mission is to drive the transition from fragile prototypes to real-world deployments of humanoids. We were founded by leading scientists in robot reinforcement learning (ex-Nvidia, ex-ETH Zürich) and backed by leading international VC firms. In just months, we went from our first line of code to deploying real humanoid capabilities with our customers, leveraging simulation and reinforcement learning. Today, we are rapidly expanding the capabilities of our autonomy stack, our customer base, and our team.

The role

We are looking for an experienced ML engineer to join Flexion’s experienced infrastructure team and take ownership of Flexion’s GPU compute platforms. This is a senior, on-site role with significant scope. 

At Flexion, we are building the brain for humanoid robots, which involves training foundation models with vast amounts of data on large GPU clusters. You will own the design, bring-up, operation and optimization of performant clusters. You will work with AI engineers to help them optimize their training speed and hardware utilization. You will also influence strategic compute planning and contribute to new tools and platforms for iterating on our AI models efficiently. This will put you at the heart of Flexion’s AI development and allow you to directly impact the execution of our ambitious roadmap. You will closely collaborate with the company’s leadership, engineers of the infrastructure team and AI engineers across the company.

Key responsibilities
  • Architect, run and continuously improve existing and future cloud-based GPU clusters. Select the best frameworks and tooling to run our clusters efficiently. Work on cluster provisioning, job schedulers and monitoring systems.
  • Help AI engineers optimize their training workloads and maximize hardware utilization using profilers, contributing to our core ML libraries.
  • Contribute to short- and long-term GPU compute strategies in collaboration with our AI engineering teams and help execute on them. 
  • Optimize capacity and cost by exploring multi-cloud strategies and evaluating trade-offs.
  • Raise the bar on engineering practices, including testing, code quality, documentation, and system reliability.

Requirements

  • Degree in Computer Science, Electrical Engineering or Software Engineering (or equivalent practical experience) plus significant industry experience.
  • Hands-on experience with the training or inference of large models (billions of parameters) on distributed multi-node GPU hardware. This can include bringing up and running the cluster, writing and optimizing training/inference code, building ML pipelines, etc.
  • Proficiency in Python and working knowledge of PyTorch.
  • Deep understanding of distributed training concepts (DDP, FSDP, NCCL).
  • Experience with at least one cloud platform (AWS, GCP, Azure or neoclouds) or large-scale on-premises GPU infrastructure.
  • Experience with job scheduling and orchestration tools: Slurm and/or Kubernetes/KubeRay.

Nice-to-haves

  • Familiarity with profilers (e.g., PyTorch Profiler, Dynolog, HTA, Nsight).
  • Experience with high-performance or parallel file systems (e.g., Lustre).
  • Experience provisioning compute on multiple cloud providers.
  • Experience with infrastructure-as-code and configuration management (Terraform, Ansible).

Benefits

  • Competitive Compensation
  • Joining a leading robotics team & exposure to never-done-before research
  • Energetic, collaborative culture with a bias for action and regular community events

Zurich

  • Enhanced pension plan
  • Relocation & permit sponsorship
  • Enhanced holiday & paid leave perks
  • Central Zürich office with top-tier robotics testing facilities and infrastructure

San Franciso

  • 401(k) with company contributions
  • Health, dental & vision coverage with the flexibility to choose your own plan
  • Open PTO policy & paid company holidays

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the ML Engineer - Infrastructure in San Francisco, CA vacancy
  • $292k - $417.2k

     ...ads optimization that shape the future of streaming.We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    1 day ago
  •  ...oprecruiting.comPhone: (***) ***-****Job Title: Machine Learning Infrastructure EngineerLocation: San Francisco, CA Metro Area (100% On-Site)...  ...company is seeking a Machine Learning Infrastructure Engineer to help architect the compute, training, and execution frameworks... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Objective Paradigm

    San Francisco, CA
    2 days ago
  • $200k - $280k

     ...Engineering San Francisco Full-time $200,000 - $280,000 About the Role Join our ML Infrastructure team to build the systems that train, deploy, and serve our AI models at scale. You'll work at the intersection of machine learning and systems engineering. What... 
    Suggested
    Full time
    Work at office

    Lattice

    San Francisco, CA
    2 days ago
  • $92k - $138k

     ...The opportunity Unity Vector builds an offline ML platform that powers insight, experimentation, attribution...  ...production ML systems. We’re looking for a Machine Learning Engineer to join our Offline Infrastructure team. This is an ideal role for a recent university... 
    Suggested
    Work at office
    Worldwide
    Relocation package

    UNITY

    San Francisco, CA
    1 day ago
  • $180k - $250k

     ...generation of AI. We are developing the context engine layer that solves a fundamental...  ...wave of AI progress will come from better infrastructure around models: Better Memory &...  ...Khan, CEO: ex-Amazon; PhD in Robotics and ML. Clark Zhang, CTO: ex-Meta; PhD in... 
    Suggested

    GraphOn

    San Francisco, CA
    4 days ago
  • $250k

     ...ML Infrastructure Engineer ML Infrastructure Engineer – Open Source ML Infra – Up to $250K Total Comp: $500K – Hybrid We are working with one of the leading ML Infra companies supporting hundreds of custom LLMs You will join a talented Infrastructure engineering... 
    Flexible hours

    Adapt Talent

    San Francisco, CA
    3 days ago
  • $180k - $230k

     ...ML Infrastructure Engineer San Francisco Company Overview Echo Neurotechnologies is an exciting new startup in the Brain-Computer Interface (BCI) space, driving innovation through advanced hardware engineering and AI solutions. Our mission is to deliver cutting... 
    Flexible hours

    Echo Neurotechnologies

    San Francisco, CA
    3 days ago
  •  ...Trajectory is seeking an ML Infrastructure Engineer to build the platform that powers training, inference, and kernels for next‑gen AI systems. You will own distributed training, low‑latency serving, and GPU‑kernel optimizations, shaping reusable systems that scale in... 

    Jobleads-US

    San Francisco, CA
    1 day ago
  • The role At Mach9, ML infrastructure engineers build and maintain the systems that power production AI models for civil engineering and surveying. Our ML pipeline spans 10,000+ miles of labeled survey data, image segmentation networks, and 3D prediction models serving real... 
    Work experience placement

    Jobleads-US

    San Francisco, CA
    1 day ago
  •  ...Mach9 Robotics Inc. is seeking an ML infrastructure engineer to build and maintain systems powering production AI models for civil engineering and surveying. You will manage training pipelines, data generation, and inference that integrates with CAD software. The role... 

    Jobleads-US

    San Francisco, CA
    2 days ago
  • $160k - $250k

     ...enjoy facing adversity, and can do the impossible at record breaking speeds. About You and The Role  As an ML Training & Inference Infrastructure Engineer on the Data Platform team you will be building and scaling the systems powering our data flywheel.  This person... 
    Local area

    Zipline

    San Francisco, CA
    a month ago
  • $190k - $210k

     ...backed by climate-tech and Silicon Valley investors. For more information, please visit  Role Description As a Senior ML Infrastructure Engineer, you will work directly in the Automation org with the core ML, Ops, and Analytics teams to help improve and build out... 
    Local area

    Gridware

    San Francisco, CA
    2 days ago
  • $200k - $400k

     ...Job Description Job Description Machine Learning Engineer - Infrastructure Company: Causal Labs Location: San Francisco, CA (South Park...  ...performance for large models. If you have built large-scale ML infrastructure for language, vision, robotics or biology models... 
    Full time
    Work at office
    Relocation
    Visa sponsorship

    Transparent Search Group

    San Francisco, CA
    16 days ago
  •  ...Job Description Job Description We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world . You’ll serve as the bridge between our researchers and... 
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    27 days ago
  •  ...Francisco is seeking a skilled Machine Learning Engineer to join our Risk and Trust team. This...  ...and maintaining model development infrastructure to support a decentralized financial ecosystem...  ...understanding of Python, advanced ML tools, and strong problem-solving skills... 

    Block Inc

    San Francisco, CA
    1 day ago
  •  ...Senior ML Platform Engineer We're on a mission to unleash the power of content… you in? We've got the brands, we've got the stars, we've got the power to achieve our mission to entertain the planet – now all we're missing is… YOU! Becoming a part of Paramount means... 

    VH1

    San Francisco, CA
    2 days ago
  • Title: ML Engineer Location: San Francisco, CA (Onsite) Direct HireCompany Mission Our client’s mission is to scale medical knowledge...  ...influence on the development of next-generation predictive infrastructure for healthcare.Opportunities to publish, co-author patents,... 

    Spectraforce Technologies

    San Francisco, CA
    18 hours ago
  • $170.1k - $258.3k

     ...export, kernel development, and performance engineering so that every cycle on our accelerators...  ...that sit at the heart of our on‑vehicle ML inference for ADAS and autonomous...  ...workloads.Build and improve tooling and infrastructure that make it easier to profile, debug, and... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    2 days ago
  •  ...000s of developers and enterprise users.ML performance, quality, and systems acumen...  ...Qualifications 7+ years of software engineering experience building and operating enterprise...  ...building the next layer of enterprise AI infrastructure.Job SummaryCategory: Engineering
    Work at office

    Atlassian

    San Francisco, CA
    18 hours ago
  •  ...Job Description Job Description The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to...  ...video streams. As a Machine Learning Engineer in ML Runtime & Optimization , you will develop technologies... 

    Zensors

    San Francisco, CA
    28 days ago
  •  ...OpenAI seeks a software engineer to build and evolve its classification and evaluation platform. You will design services, APIs, and pipelines...  ...performance, and improve observability to support fast iteration in a demanding ML environment. #J-18808-Ljbffr Jobleads-US

    Jobleads-US

    San Francisco, CA
    2 days ago
  •  ...Spring Health is seeking a Senior Machine Learning Engineer to help build and scale a centralized AI platform that powers care capabilities...  ...ideal candidate has 4-6 years of Python experience with GenAI/ML libraries, experience with Kubernetes, AWS, and Azure, and a track... 
    Work at office
    Relocation

    Jobleads-US

    San Francisco, CA
    5 days ago
  •  ...into the physical world. We are a group of engineers, scientists, roboticists, and company...  ...role sits at the intersection of our AI, infrastructure, and external partners. You’ll work...  ...in production. You do not need to be an ML researcher. Strong Python skills and... 
    Remote work

    Jobleads-US

    San Francisco, CA
    1 day ago
  •  ...Proteus Bio in San Francisco is seeking an ML Engineer to join our Engineering team on-site at Fort Mason. You will help build our Biological Intelligence Platform for autonomous discovery of personalised medicine, taking ideas from concept to production. You should... 

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $166k - $210.25k

    RDQ127R59SummaryAs a Senior Applied ML Engineer on the Applied AI team at Databricks, you will use machine learning, scheduling, and...  ...algorithms to maximize the efficiency and performance of our infrastructure. Your work will span the entire stack—from cluster... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  •  ...Machine Learning Engineer Bucket Robotics is hiring a Machine Learning Engineer to push the frontier of CAD-native computer vision for manufacturing. You'll work on the core ML systems that turn 3D geometry and synthetic data into reliable, production-grade vision... 
    Shift work

    Bucket Robotics

    San Francisco, CA
    18 hours ago
  •  ...Senior Client Engineer SAN FRANCISCO, CA ENGINEERING FULL-TIME What Will You Be Doing? Training machine learning models over billions of data points. Quantifying predictive uncertainty using probabilistic and Bayesian methods. Creating models that quickly... 
    Full time
    Work experience placement

    1872 Consulting

    San Francisco, CA
    18 hours ago
  • $165k - $200k

     ...We're seeking an exceptional Machine Learning Engineer to be a founding member of our cutting‑edge team. As a pioneer in this dynamic field...  ...of work for over 5 million shift workers. As a founding ML Engineer, you'll be responsible for developing and deploying state... 
    Home office
    Flexible hours
    Shift work

    Instawork

    San Francisco, CA
    1 day ago
  • THE GLOBAL LEADER IN DATA & ANALYTICS RECRUITMENTHarnham Search and Selection Company Number: 05723485Harnham Search and Selection is a registered company in England and Wales. Reg no. 05723485Harnham Europe Limited Company Number: 09956940Harnham GmbH HRB: 196954Harnham...
    Work at office

    Harnham

    San Francisco, CA
    2 days ago
  •  ...environments. What to expect This role is for a machine learning engineer who wants to work on the models that give LeLamp its...  ...for training and deploying machine learning models Integrate ML systems with robotic hardware and embedded systems Improve robot... 
    Immediate start

    Human Computer Lab

    San Francisco, CA
    18 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Engineer - Infrastructure. Be the first to apply!