Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Helix AI Engineer, Training Performance

$200k - $400k
Full-time

Figure

Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure is headquartered in San Jose, CA.

Figure's vision is to deploy autonomous humanoids at a global scale. Our Helix team is looking for an experienced AI Training Performance Engineer to take our model training to the next level. This role is focused on improving distributed training frameworks for large scale model training, optimizing GPU kernels, exploring the relative gains of different accelerator types and co-designing our models to maximize utilization of our hardware.

Responsibilities

  • Optimize training performance for a 100B+ parameter models across 100k+ GPUs.
  • Collaborate with the broader team on accelerator choice, cluster topology, scheduling, and hardware procurement decisions to inform future scaling.
  • Write and optimize custom kernels (Triton/CUDA)
  • Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis across training jobs
  • Optimize data loading and preprocessing pipelines so I/O never gates the accelerators
  • Improve checkpointing, fault tolerance, and elastic restart so large jobs recover quickly from node failures without losing significant wall-clock time
  • Partner with researchers to co-design model architectures and training recipes that are performant at scale (e.g., activation checkpointing strategies, mixed precision, sequence packing)
  • Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators
  • Build and extend agentic systems that automatically generate, benchmark, and iterate on custom kernels
  • Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware) for fit with our training workloads, and lead proof-of-concept ports/benchmarks
  • Explore different model/data parallelisms (FSDP, context parallel, expert parallel, etc.) to determine optimal configuration per model size.

Requirements

  • Bachelor's or Master's degree in Computer Science, Computer/Electrical Engineering, or a related field
  • 3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects
  • Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy)
  • Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ability to translate traces into concrete optimizations
  • Solid grasp of collective communication (NCCL) and modern networking concepts (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement).
  • Strong Python and CUDA/C++ skills; comfortable reading and modifying framework internals
  • Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence)
  • Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) and using them to drive optimization priorities

Bonus Qualifications

  • Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration
  • Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.)
  • Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management.

The US base salary range for this full-time position is between $200,000 - $400,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Helix AI Engineer, Training Performance in California vacancy
  • $200k - $400k

     ...Helix AI Engineer, Pretraining Figure is an AI robotics company developing autonomous general-purpose humanoid robots. Our goal is to build...  ..., planning, and action. Responsibilities Design and train large-scale foundation models across multimodal data (e.g.,... 
    Training
    Full time
    Work at office

    Figure

    San Jose, CA
    4 days ago
  • $200k - $350k

     ...Figure is an AI Robotics company developing a general purpose humanoid. Our humanoid robot...  ...humanoids at a global scale. Our Helix team is looking for Perception Engineers to empower Figure humanoid robots to perform highly dynamic operations in demanding real... 
    Performance
    Full time
    Work at office

    Figure

    California
    12 days ago
  • $150k - $350k

     ...Figure AI is an AI robotics company developing autonomous general...  ...intelligence. Its robots are engineered to perform a variety of tasks in the...  .... We are looking for a Helix AI Engineer, Agentic Systems...  ...Responsibilities Design, train, and deploy multimodal agents... 
    Training
    Full time

    Figure

    California
    a month ago
  • $180k - $300k

     ...Role You'll own the core AI systems that power Gamma: the...  ...existing foundation models, not training new ones. You'll focus on...  ...and fine-tuning for maximum performance across Gamma's product surface...  ...stack. You'll work closely with engineering and product to ship... 
    Training
    Performance
    Full time
    Work at office
    Immediate start
    Work from home

    Gamma

    San Francisco, CA
    1 day ago
  •  ...revolutionizing software development with AI-powered formal verification....  ...Join our team as an AI Engineer and help us push the...  ...pipelines for scaled distributed training. You'll be at the forefront...  ...and tools critical for high-performance computing in Python/C++ and machine... 
    Training
    Performance
    Full time
    Contract work

    Logical Intelligence

    San Francisco, CA
    1 day ago
  •  ...want to work at the intersection of AI, product, engineering, and real -world operations? \ \ Join...  .... You will join a small, high -performing team of exceptional builders and work...  ...implementations from onboarding to rollout and training \ Convert learnings from... 
    Training
    Performance
    Full time
    Immediate start

    Eth Juniors

    San Francisco, CA
    1 day ago
  • $155k - $180k

     ...disciplinary skills (embedded AI, high-tech manufacturing automation...  ...will assume the role of AI engineer, with the following main...  ...data processing architectures Perform business and technical analysis...  ...approaches, including the use of pre-trained networks or training custom... 
    Training
    Performance
    Full time

    Kickmaker

    San Francisco, CA
    1 day ago
  •  ...next few decades. For phase 1, we're training foundational AI models that create building code-...  ...foundational models, geometric and physics engines from scratch to transform how the...  ...pipelines that measure not just model performance, but constructibility and real-world... 
    Training
    Performance
    Full time
    Visa sponsorship

    Augrade Private Limited

    San Francisco, CA
    1 day ago
  • $50k - $120k

     ...Mission Altimate AI, founded in 2022 in San Francisco...  ...like VSCode, Git, and Slack, performing tasks ranging from data documentation...  ...of the AI-powered data engineering revolution. You can read more...  ..., Prompt Tuning, and Adapter Training Why you should join Altimate... 
    Training
    Performance
    Full time
    Worldwide

    Pa Early Stage Partners

    Sunnyvale, CA
    1 day ago
  • $200k - $400k

     ...machines at scale. At Scout AI, we’re developing Fury, the first...  ...for a Senior or Staff AI Engineer to join the Fury Orchestration...  ...contribute across the stack: model training and evaluation, reinforcement...  ...to validate system performance e under real-world constraints... 
    Training
    Performance
    Full time
    Relocation package

    Scout Ai

    Sunnyvale, CA
    1 day ago
  •  ...building the next generation of AI agents for the insurance...  ...looking for an exceptional AI Engineer to join our agent team and help...  ...to measure & improve agent performance. Experiment with new...  ...Note that this is not a model-training role - you’ll be building orchestration... 
    Training
    Performance
    Full time
    Work at office

    Further AI

    San Francisco, CA
    1 day ago
  • $100k

     ...leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and...  ...Speed Interconnect / Signal Integrity Engineer to design and validate high-...  ...next-generation AI inference and training clusters. This role is on-site in... 
    Training
    Performance
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    4 days ago
  •  ...role Our client is a well-funded AI startup building production-grade ML...  ...They are looking for a Senior AI/ML Engineer to own model training pipelines, evaluation systems, and inference...  ...evaluation, deployment Own model performance, latency, and cost trade-offs in... 
    Training
    Performance
    Full time

    Clera

    San Francisco, CA
    1 day ago
  •  ...Systems builds the world's largest AI chip, 56 times larger than...  ...to deliver industry-leading training and inference speeds; over 10...  ...You'll own model quality and performance for Cerebras' inference...  ...loop." You'll sit between engineering, product, and customer-facing... 
    Training
    Performance
    Full time

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will...  ...techniques and tools Experience investigating, and resolving, training & inference performance end to endDebugging and... 
    Training
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    5 days ago
  •  ...Cooperidge Consulting Firm is seeking an AI Agent Software Engineer for a high-momentum AI platform...  ...debuggable, and continuously improving. Performance Engineering: Analyze large-scale...  ...the team will provide comprehensive training on agentic reasoning and LLM... 
    Training
    Performance
    Full time
    Apprenticeship
    Flexible hours

    Cooperidge Consulting Firm

    San Francisco, CA
    1 day ago
  •  ...generation computing experiences—from AI and data centers, to PCs,...  .... THE ROLEWe are hiring AI Engineers to build recursive self-...  ...intersection of AI systems, performance engineering, hardware-aware optimization...  ...Reinforcement learning, post-training, reward modeling, and... 
    Training
    Performance

    AMD

    Santa Clara, CA
    5 days ago
  • $225k - $250k

     ...Suki Assistant, uses cutting-edge AI to automate clinical...  ...data quality and relevance for training and evaluation. You'll monitor and analyze system performance metrics, identifying areas for...  ...: We’re former Googlers, Apple engineers, Stanford docs, and healthcare... 
    Training
    Performance

    Suki

    Redwood City, CA
    5 days ago
  • $152k - $241.5k

    We are looking for outstanding Senior High Performance AI Engineers to build the next generation of agentic AI systems for the CUDA ecosystem. Our team works across the full agentic AI stack—from training and improving models, to designing agent architectures and multi-... 
    Training
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...At Viven, we're building AI-powered Digital Twins for businesses...  ...AI version of an employee — trained on their decisions, knowledge...  ...We are looking for an AI/ML Engineer with hands-on experience in...  ...’s products, with a focus on performance, scalability, and reliability... 
    Training
    Performance
    Full time

    Careers

    Santa Clara, CA
    1 day ago
  •  ...generation computing experiences—from AI and data centers, to PCs,...  ...a Forward Deployed Research Engineer to build, evaluate, and...  ...agentic workflows, and post-training techniques.THE PERSON:You'll...  ...Diagnose and optimize AI system performance across models, data, infrastructure... 
    Training
    Performance

    AMD

    Santa Clara, CA
    1 day ago
  • $186.5k - $328.5k

     ...leading enterprises orchestrate AI-powered work. Our vision is...  ...AI. About the roleAs an AI engineer at WRITER, you'll be at the forefront...  ...will directly impact the performance, scalability, and ethical...  ...massage/chiropractor, personal training, etc.Learning and development... 
    Training
    Performance
    Full time
    Work at office
    Local area

    Writer

    San Francisco, CA
    1 day ago
  •  ...revolutionizing software development with AI-powered formal verification....  ...Join our team as an AI Engineer and help us push the...  ...pipelines for scaled distributed training and validation of ML models....  ...and tools critical for high-performance computing in Python/C++ and machine... 
    Training
    Performance
    Full time
    Contract work

    Logical Intelligence

    San Francisco, CA
    1 day ago
  • $149.75k - $275.58k

     ...Details:Job Description: Join Intel's AI Frameworks PyTorch team to shape...  ...Intelligence (AI) and High-Performance Computing (HPC). As an AI Frameworks Engineer, you will design, build, and optimize...  ...software, enabling efficient training, inference, and deployment of modern... 
    Training
    Performance
    Full time
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    1 day ago
  •  ...About Orum Orum ’s AI-powered suite frees salespeople to do...  ...ML best practices across the engineering org Partner with Product and...  ...loops to improve model performance Mentor engineers and raise...  ...and deploying ML pipelines for training, inference, monitoring, and continuous... 
    Training
    Performance
    Remote job
    Full time

    Orum

    San Francisco, CA
    1 day ago
  • $110.7k - $372.9k

     ...they need. Deloitte has a new AI-first effort, backed by $1B...  ...into a lab. As an Agentic AI Engineer, you will design, build, and...  ...Partner with our modeling and post-training engineers to improve model...  ...role carries a substantial performance-based incentive opportunity designed... 
    Training
    Performance
    Local area
    Visa sponsorship

    Deloitte

    Costa Mesa, CA
    4 days ago
  • $120k - $220k

     ...information powered by advanced AI, recommendation systems, and...  ...every cycle.We're hiring the engineer who owns this agent end-to-end...  ...loop — Extend our LLM critic + performance-feedback regeneration from images...  .../ DreamBooth checkpoint — Train and ship on past CPI-winning ads... 
    Training
    Performance
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    5 days ago
  • $125k - $145k

     ...ultimate goal of enabling human life on Mars.AI SOFTWARE ENGINEER (VEHICLE ENGINEERING)Be a member of...  .... You will be responsible for training internal engineering models on SpaceX...  ...systems and multi-agent workflows to perform engineering tasksBuild and optimize large... 
    Training
    Performance
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    5 days ago
  • $229.9k - $262.4k

    {"description": "Senior Lead AI Engineer (MLX, Agentic AI, Gen AI platform Services)...  ...product experiences and scalable, high-performance AI infrastructure. At Capital One, you...  ...components including foundation model training, large language model inference, similarity... 
    Training
    Performance
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    San Jose, CA
    1 day ago
  • $91.1k - $179.5k

    Position Summary Agentic AI is moving from...  ...scale. We're growing a team of engineers who want to work at the center...  ...help clients improve financial performance, accelerate new digital ventures...  ...to skill sets; experience and training; licensure and certifications... 
    Training
    Performance
    Work at office
    Local area
    Visa sponsorship
    Shift work

    Deloitte

    Los Angeles, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Helix AI Engineer, Training Performance. Be the first to apply!