Staff, Pre-Training Infra — Distributed ML Training
B Capital
B Capital is seeking a talented engineer in San Francisco to build and scale distributed training systems for machine learning models. The ideal candidate will have strong experience in distributed training frameworks, debug GPU compute systems, and optimize training throughput. This role offers top-tier compensation, comprehensive health benefits, and the opportunity to work in a collaborative environment with daily meals and team celebrations. #J-18808-Ljbffr B Capital
$227.2k - $417k
...Software Engineer on the ML Infrastructure team,... ...have multiple openings;Staff Software EngineerPrincipal... ..., and low latency distributed systems using ScalaBuild... ...and efficiency of our infra. Lead large scale cross... ...Feast), ElastiCache, model training orchestration, etc....TrainingFull timeTemporary workLocal areaFlexible hours- ...Anthropic and beyond. About the Role Build and scale distributed training systems that power frontier model pre-training. Work closely with research teams to... ...distributed workloads. Experience working closely with ML researchers to productionize experimental training...TrainingFull timeRelocation package
- ...and accessible to all. About the Role Build and scale distributed training systems that power frontier model pre-training. Work closely with research teams to... ...distributed workloads. Experience working closely with ML researchers to productionize experimental training workflows...TrainingWork at officeVisa sponsorship
$204k - $300k
...science and electrical engineering, such as AI/ML, algorithms, digital signal processing,... ...vision, data science & analytics, distributed systems, cloud, edge & mobile computing,... ...internal parity, and relevant education or training. Your recruiter can share more about the...TrainingFull timeLocal areaWorldwideFlexible hours- ...San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize... ...memory and throughput, and contribute to open-source ML infrastructure. You should have strong expertise in PyTorch and...Training
- ...to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization... ...Ideal candidates have over 5 years of experience in ML infrastructure and a strong background in...Training
$117.2k - $313.7k
...new and exciting components/frameworks in distributed filesystems in an ever-growing and... ...Design patterns & Experience with Big-Data/ML and S3 Hands-on experience with Streaming... ...assignment, compensation, promotion, benefits, training, assessment of job performance,...TrainingFull timeImmediate startRemote work$260k - $300k
...the whole patient journey: pre-visit intake, care navigation... ...About this role As a Staff Backend/Infra Engineer at Amigo, you'll build... ...re strong with concurrency, distributed systems, and service design... ...Experience with production AI or ML systems Benefits (...Full timeFlexible hours- ...About the Team The Post-Training Frontiers team creates the frontier... ...(3) building the research and infra for horizontal integrations, such... ..., orchestration, scaling, and distributed infrastructure. - Solve hard... ...experience in some layer of ML infrastructure. - Have...TrainingFull time
- ...world business problems. We’re training and deploying frontier models... ...role. As a Member of Technical Staff, you will: Design and write... ...Proficiency in Python and related ML frameworks such as JAX,... ...and XLA/MLIR. Experience with distributed training infrastructures (Kubernetes...TrainingFull timeWork at officeLocal areaRemote workHome office
$281k - $356k
...and work with partners to scale eval and ML development by exploring new methods and delivering... ...hybrid role, you will report to a Senior Staff manager. You will: Develop tools for... ...exact work location, experience, relevant training and education, and skill level. Your...TrainingFull timeWork experience placementRemote work$190k - $265k
...improve their business.Training and customizing state-of... ...-tuning open models to pre-training frontier-scale... ...in the world.As a Staff Software Engineer for AI... ...scheduling and capacity, distributed training performance, fault... ...computing, or ML systems.Hands-on experience...TrainingLocal areaWorldwide$150k - $300k
...from frontier agentic models to the infra that enables anyone to create, train, and deploy them. We aggregate and... ...including SLURM and Kubernetes for distributed workloads Implement high‑performance... ...FSDP, DeepSpeed, Megatron‑LM) ML framework optimization and profiling...Training$192k - $260k
...build the most trusted data analytics and ML platform in the world. We’re looking to... ...language (preferably Python) ~ Experience with distributed data processing systems like Spark and... ...experience, relevant certifications and training, and specific work location. Based on the...TrainingRemote jobLocal areaWorldwide- ...interventions. We value a relentless problem-solving approach and rapid execution. You will work across data, model, eval, and infrastructure, taking ideas from prototype to scaled training runs with a focus on uncertainty and decision quality. #J-18808-Ljbffr KindredventuresTraining
- ...and act through imagined futures. As a Staff AI Researcher you will help set the technical... ...platform, including core modeling, training, evaluation, and deployment decisions Design... ...systems across large datasets or distributed training environments Publications at leading...Training
- ...will help scale and optimize our training systems and core model code.... ...role at the intersection of ML, software engineering, and scalable... ..., and metrics/logging. Scale distributed training: Work with... ...Translate research needs into infra capabilities and guide best practices...TrainingFull time
$200k - $280k
...architectures, engines) and post-training / RL systems. We build... ...large neural nets. Distributed systems / high-performance computing for ML. Are comfortable working... ...collaborating with infra, research, and product teams... ...technical leadership (Staff level) Set technical direction...TrainingFull time$150k - $300k
...from frontier agentic models to the infra that enables anyone to create, train, and deploy them. We aggregate and... ...meet throughput/latency SLOs. Model Distribution: Optimize model distribution and... ...Requirements Required Experience Building ML Systems at Scale: 3+ years building...TrainingWork at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work$180k - $350k
...from frontier agentic models to the infra that enables anyone to create, train, and deploy them. We aggregate and... ...: the hosted RL training platform, distributed GPU infrastructure, liquid compute... ...Experience securing GPU infrastructure or ML training pipelines Background in...TrainingWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours$150k - $300k
...from frontier agentic models to the infra that enables anyone to create, train, and deploy them. We aggregate and... ...workload management You will work on a distributed system with performance engineering... ...Experience with GPU computing and ML infrastructure Knowledge of AI/ML...TrainingWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours$190k - $250k
...comprehensive platform to manage heart disease. As a Staff Data Architect, you will lead the data... ...systems through curated analytical and ML-ready datasets. Advance the Semantic... ...Heartflow, including recruitment, hiring, training, relocation, promotion, and termination....TrainingWork experience placementLocal areaWorldwideRelocation- ...secure sandboxes, high-performance training, and deployment into one full-stack... ...roles Product management for AI, infra, devtools, or enterprise software ML engineering, applied research, or AI... ...systems without needing every detail pre-digested Excellent written and...TrainingRemote workVisa sponsorshipRelocation packageFlexible hours
$20 - $25 per hour
...transactions recorded by other staff members by performing various... ...records. Compile, review and distribute various reports to appropriate... ...Certification of previous training in computers Experience with... ...discounted hotel stays and more. Pre-employment background screening...TrainingFull time- ...About the Role Generalist trains very large robot foundation models. This requires utilizing... ...infrastructure (currently Nvidia) to run distributed training jobs and researcher experiments.... ...utilized Optimizing and improving ML data loading transport and storage in highly...TrainingFull time
$220k - $320k
...Help us build the systems that train specialized AI models for the fastest-growing companies... ...world. If you love taking cutting-edge ML techniques and turning them into products... ...and multimodal models Experience with distributed training at scale Contributions to...TrainingFull timeWork at office- ...tasks, apply to join us.Role MissionPost-training is the critical bridge between raw model... ...DemonstratedA track record of advancing ML systems through post-training, alignment,... ...industry resultsExperience with large-scale distributed training and the debugging that comes...TrainingShift work
- ...engineering: managing the large scale distributed systems that orchestrate... ...evals for our frontier model training Own eval throughput and cost... ...cut wall-clock time on the pre-train eval suite Designing... ...policy: Currently, we expect all staff to be in one of our offices...TrainingWork at officeVisa sponsorshipFlexible hours
- ...Language Models (LLMs) pipelines for scaled distributed training. You'll be at the forefront of... ...approaches, including latent space reasoning Pre-train, fine-tune, and modify the State-... ...~3+ years of production experience in ML Infra, DataOps, distributed training....TrainingFull timeContract work
- Magic AI, Inc. is seeking a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model execution meets distributed infrastructure, influencing latency,...Training
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff, Pre-Training Infra — Distributed ML Training. Be the first to apply!
- machine learning scientist San Francisco, CA
- machine learning part time San Francisco, CA
- machine learning intern San Francisco, CA
- machine learning remote San Francisco, CA
- machine learning researcher San Francisco, CA
- machine learning San Francisco, CA
- machine learning research scientist San Francisco, CA
- internship machine learning San Francisco, CA
- artificial intelligence - machine learning intern San Francisco, CA
- data engineer machine learning San Francisco, CA



