ML Engineer - Infrastructure
Flexion Robotics
Job Description
Job Description
About Flexion
At Flexion, we are building the autonomy stack for humanoid robots. Our mission is to drive the transition from fragile prototypes to real-world deployments of humanoids. We were founded by leading scientists in robot reinforcement learning (ex-Nvidia, ex-ETH Zürich) and backed by leading international VC firms. In just months, we went from our first line of code to deploying real humanoid capabilities with our customers, leveraging simulation and reinforcement learning. Today, we are rapidly expanding the capabilities of our autonomy stack, our customer base, and our team.
The roleWe are looking for an experienced ML engineer to join Flexion’s experienced infrastructure team and take ownership of Flexion’s GPU compute platforms. This is a senior, on-site role with significant scope.
At Flexion, we are building the brain for humanoid robots, which involves training foundation models with vast amounts of data on large GPU clusters. You will own the design, bring-up, operation and optimization of performant clusters. You will work with AI engineers to help them optimize their training speed and hardware utilization. You will also influence strategic compute planning and contribute to new tools and platforms for iterating on our AI models efficiently. This will put you at the heart of Flexion’s AI development and allow you to directly impact the execution of our ambitious roadmap. You will closely collaborate with the company’s leadership, engineers of the infrastructure team and AI engineers across the company.
Key responsibilities- Architect, run and continuously improve existing and future cloud-based GPU clusters. Select the best frameworks and tooling to run our clusters efficiently. Work on cluster provisioning, job schedulers and monitoring systems.
- Help AI engineers optimize their training workloads and maximize hardware utilization using profilers, contributing to our core ML libraries.
- Contribute to short- and long-term GPU compute strategies in collaboration with our AI engineering teams and help execute on them.
- Optimize capacity and cost by exploring multi-cloud strategies and evaluating trade-offs.
- Raise the bar on engineering practices, including testing, code quality, documentation, and system reliability.
Requirements
- Degree in Computer Science, Electrical Engineering or Software Engineering (or equivalent practical experience) plus significant industry experience.
- Hands-on experience with the training or inference of large models (billions of parameters) on distributed multi-node GPU hardware. This can include bringing up and running the cluster, writing and optimizing training/inference code, building ML pipelines, etc.
- Proficiency in Python and working knowledge of PyTorch.
- Deep understanding of distributed training concepts (DDP, FSDP, NCCL).
- Experience with at least one cloud platform (AWS, GCP, Azure or neoclouds) or large-scale on-premises GPU infrastructure.
- Experience with job scheduling and orchestration tools: Slurm and/or Kubernetes/KubeRay.
Nice-to-haves
- Familiarity with profilers (e.g., PyTorch Profiler, Dynolog, HTA, Nsight).
- Experience with high-performance or parallel file systems (e.g., Lustre).
- Experience provisioning compute on multiple cloud providers.
- Experience with infrastructure-as-code and configuration management (Terraform, Ansible).
Benefits
- Competitive Compensation
- Joining a leading robotics team & exposure to never-done-before research
- Energetic, collaborative culture with a bias for action and regular community events
Zurich
- Enhanced pension plan
- Relocation & permit sponsorship
- Enhanced holiday & paid leave perks
- Central Zürich office with top-tier robotics testing facilities and infrastructure
San Franciso
- 401(k) with company contributions
- Health, dental & vision coverage with the flexibility to choose your own plan
- Open PTO policy & paid company holidays
$292k - $417.2k
...ads optimization that shape the future of streaming.We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction...SuggestedFull timeTemporary workLocal areaFlexible hours- ...oprecruiting.comPhone: (***) ***-****Job Title: Machine Learning Infrastructure EngineerLocation: San Francisco, CA Metro Area (100% On-Site)... ...company is seeking a Machine Learning Infrastructure Engineer to help architect the compute, training, and execution frameworks...SuggestedFull timeWork at officeFlexible hours
$200k - $280k
...Engineering San Francisco Full-time $200,000 - $280,000 About the Role Join our ML Infrastructure team to build the systems that train, deploy, and serve our AI models at scale. You'll work at the intersection of machine learning and systems engineering. What...SuggestedFull timeWork at office$92k - $138k
...The opportunity Unity Vector builds an offline ML platform that powers insight, experimentation, attribution... ...production ML systems. We’re looking for a Machine Learning Engineer to join our Offline Infrastructure team. This is an ideal role for a recent university...SuggestedWork at officeWorldwideRelocation package$180k - $250k
...generation of AI. We are developing the context engine layer that solves a fundamental... ...wave of AI progress will come from better infrastructure around models: Better Memory &... ...Khan, CEO: ex-Amazon; PhD in Robotics and ML. Clark Zhang, CTO: ex-Meta; PhD in...Suggested$250k
...ML Infrastructure Engineer ML Infrastructure Engineer – Open Source ML Infra – Up to $250K Total Comp: $500K – Hybrid We are working with one of the leading ML Infra companies supporting hundreds of custom LLMs You will join a talented Infrastructure engineering...Flexible hours$180k - $230k
...ML Infrastructure Engineer San Francisco Company Overview Echo Neurotechnologies is an exciting new startup in the Brain-Computer Interface (BCI) space, driving innovation through advanced hardware engineering and AI solutions. Our mission is to deliver cutting...Flexible hours- ...Trajectory is seeking an ML Infrastructure Engineer to build the platform that powers training, inference, and kernels for next‑gen AI systems. You will own distributed training, low‑latency serving, and GPU‑kernel optimizations, shaping reusable systems that scale in...
- The role At Mach9, ML infrastructure engineers build and maintain the systems that power production AI models for civil engineering and surveying. Our ML pipeline spans 10,000+ miles of labeled survey data, image segmentation networks, and 3D prediction models serving real...Work experience placement
- ...Mach9 Robotics Inc. is seeking an ML infrastructure engineer to build and maintain systems powering production AI models for civil engineering and surveying. You will manage training pipelines, data generation, and inference that integrates with CAD software. The role...
$160k - $250k
...enjoy facing adversity, and can do the impossible at record breaking speeds. About You and The Role As an ML Training & Inference Infrastructure Engineer on the Data Platform team you will be building and scaling the systems powering our data flywheel. This person...Local area$190k - $210k
...backed by climate-tech and Silicon Valley investors. For more information, please visit Role Description As a Senior ML Infrastructure Engineer, you will work directly in the Automation org with the core ML, Ops, and Analytics teams to help improve and build out...Local area$200k - $400k
...Job Description Job Description Machine Learning Engineer - Infrastructure Company: Causal Labs Location: San Francisco, CA (South Park... ...performance for large models. If you have built large-scale ML infrastructure for language, vision, robotics or biology models...Full timeWork at officeRelocationVisa sponsorship- ...Job Description Job Description We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world . You’ll serve as the bridge between our researchers and...Work at officeVisa sponsorship
- ...Francisco is seeking a skilled Machine Learning Engineer to join our Risk and Trust team. This... ...and maintaining model development infrastructure to support a decentralized financial ecosystem... ...understanding of Python, advanced ML tools, and strong problem-solving skills...
- ...Senior ML Platform Engineer We're on a mission to unleash the power of content… you in? We've got the brands, we've got the stars, we've got the power to achieve our mission to entertain the planet – now all we're missing is… YOU! Becoming a part of Paramount means...
- Title: ML Engineer Location: San Francisco, CA (Onsite) Direct HireCompany Mission Our client’s mission is to scale medical knowledge... ...influence on the development of next-generation predictive infrastructure for healthcare.Opportunities to publish, co-author patents,...
$170.1k - $258.3k
...export, kernel development, and performance engineering so that every cycle on our accelerators... ...that sit at the heart of our on‑vehicle ML inference for ADAS and autonomous... ...workloads.Build and improve tooling and infrastructure that make it easier to profile, debug, and...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...000s of developers and enterprise users.ML performance, quality, and systems acumen... ...Qualifications 7+ years of software engineering experience building and operating enterprise... ...building the next layer of enterprise AI infrastructure.Job SummaryCategory: EngineeringWork at office
- ...Job Description Job Description The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to... ...video streams. As a Machine Learning Engineer in ML Runtime & Optimization , you will develop technologies...
- ...OpenAI seeks a software engineer to build and evolve its classification and evaluation platform. You will design services, APIs, and pipelines... ...performance, and improve observability to support fast iteration in a demanding ML environment. #J-18808-Ljbffr Jobleads-US
- ...Spring Health is seeking a Senior Machine Learning Engineer to help build and scale a centralized AI platform that powers care capabilities... ...ideal candidate has 4-6 years of Python experience with GenAI/ML libraries, experience with Kubernetes, AWS, and Azure, and a track...Work at officeRelocation
- ...into the physical world. We are a group of engineers, scientists, roboticists, and company... ...role sits at the intersection of our AI, infrastructure, and external partners. You’ll work... ...in production. You do not need to be an ML researcher. Strong Python skills and...Remote work
- ...Proteus Bio in San Francisco is seeking an ML Engineer to join our Engineering team on-site at Fort Mason. You will help build our Biological Intelligence Platform for autonomous discovery of personalised medicine, taking ideas from concept to production. You should...
$166k - $210.25k
RDQ127R59SummaryAs a Senior Applied ML Engineer on the Applied AI team at Databricks, you will use machine learning, scheduling, and... ...algorithms to maximize the efficiency and performance of our infrastructure. Your work will span the entire stack—from cluster...Local areaWorldwide- ...Machine Learning Engineer Bucket Robotics is hiring a Machine Learning Engineer to push the frontier of CAD-native computer vision for manufacturing. You'll work on the core ML systems that turn 3D geometry and synthetic data into reliable, production-grade vision...Shift work
- ...Senior Client Engineer SAN FRANCISCO, CA ENGINEERING FULL-TIME What Will You Be Doing? Training machine learning models over billions of data points. Quantifying predictive uncertainty using probabilistic and Bayesian methods. Creating models that quickly...Full timeWork experience placement
$165k - $200k
...We're seeking an exceptional Machine Learning Engineer to be a founding member of our cutting‑edge team. As a pioneer in this dynamic field... ...of work for over 5 million shift workers. As a founding ML Engineer, you'll be responsible for developing and deploying state...Home officeFlexible hoursShift work- THE GLOBAL LEADER IN DATA & ANALYTICS RECRUITMENTHarnham Search and Selection Company Number: 05723485Harnham Search and Selection is a registered company in England and Wales. Reg no. 05723485Harnham Europe Limited Company Number: 09956940Harnham GmbH HRB: 196954Harnham...Work at office
- ...environments. What to expect This role is for a machine learning engineer who wants to work on the models that give LeLamp its... ...for training and deploying machine learning models Integrate ML systems with robotic hardware and embedded systems Improve robot...Immediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Engineer - Infrastructure. Be the first to apply!
- computer vision machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- senior ml engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- machine learning engineer San Francisco, CA
- remote infrastructure engineer San Francisco, CA
- infrastructure engineer San Francisco, CA
- principal infrastructure engineer San Francisco, CA
- senior infrastructure engineer San Francisco, CA



