Distributed Training Infra Engineer for Large Models
Kindredventures
Kindredventures in San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize systems that run thousands of GPUs and accelerate research progress. Collaborate with researchers to bring prototype models to full scale, optimize memory and throughput, and contribute to open-source ML infrastructure. You should have strong expertise in PyTorch and JAX and a track record of performance profiling. #J-18808-Ljbffr Kindredventures
$227.2k - $417k
...the Role:As a Software Engineer on the ML... ...maintaining low-latency ML model serving systems that... ...throughput, and low latency distributed systems using... ...and efficiency of our infra. Lead large scale cross functional... ..., ElastiCache, model training orchestration, etc.Understanding...TrainingFull timeTemporary workLocal areaFlexible hours- ...technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating...Training
- Magic AI, Inc. is seeking a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model execution meets distributed infrastructure, influencing latency,...Training
- ...Protocol Learning : multi-participant training of foundation models where no single participant has, or can ever... .... We’re looking for Senior/Staff engineers with 5+ years of experience in distributed systems and ML large‑scale training. You’ll be implementing a...TrainingRemote workVisa sponsorship
- ...the capabilities of foundational models to support general-purpose robotics... ...About the Role As a Research Engineer, Distributed Data Systems, you will design and... ...scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage...TrainingFull timeWork at officeRelocation package
$190k - $205k
...noise characteristics. Design models that are robust,... ...and visual signals Production Engineering Write clean, scalable, well... ...code that integrates into a large shared codebase. Build end... ...processing, feature extraction, training, evaluation, and deployment....TrainingFull timeLive in- ...great server internals engineer to help maintain and... ...disclosed vulnerabilities in large software systems and... ...CVSS scoring, threat modeling, and security risk... ...work with a friendly, distributed team following open-source... ...also offer access to training on leading-edge...TrainingFull timeRemote workFlexible hours
$350k
Mirendil is looking for engineers to build infrastructure for frontier reasoning models at their San Francisco location. This role focuses on large-scale reinforcement learning (RL) model training and requires a solid understanding of engineering principles. The ideal...Training$160k - $194k
...manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in... ...application interface and database model changes, migrate and evolve schemas,... ...sets, market conditions, experience and training, licensure and certifications, and...TrainingFull timeRemote workFlexible hours- ...seeking an experienced backend/infrastructure engineer to build platforms powering AI workloads, including model training, serving, and vector search. You will join a high... .... You will collaborate across platform, infra, and ML teams to deliver end-to-end experiences...Training
$180k - $275k
...helping define and evolve the core data model and storage systems powering Gamma's business... ...rapid shipping velocity. As Software Engineer on the Platform team, you'll collaborate... ...Design and implement scalable APIs, distributed systems, and data infrastructure that...Full timeWork at officeWork from home$117.2k - $223.9k
...integrating Google's Gemini AI models into Salesforce's Agentforce... .... Join our team of talented engineers and help us advance the integration... ..., and the security of distributed and scalable distributed systems... ..., promotion, benefits, training, assessment of job performance...TrainingFull time$200.8k - $251k
...Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position...TrainingFull time- A leading tech company in San Francisco seeks a Machine Learning Engineer to build and maintain infrastructure for large-scale model training. In this hands-on role, you will design systems, work closely with researchers, and optimize training processes. Candidates should...Training
- ...San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will... ..., monitoring tools, and robust performance improvements for large‑scale runs. #J-18808-Ljbffr GenesisAITraining
$179.4k - $224.25k
...building upon our prior model evaluation work with... ...EngineOur Generative AI Data Engine powers the world’s most... ...on everything from large-scale system architecture... ...data processing and distributed systems.Familiarity with... ...relevant education or training. Scale employees in eligible...TrainingFull time$293k - $385k
...work closely with hardware, modeling, and architecture teams to... ...Workload Porting & Performance Engineer to evaluate new hardware... ....Experience working in large-scale or distributed system environments.Preferred... ...AI/ML workloads, including training or inference systems.Familiarity...TrainingWork at officeLocal areaRelocation packageFlexible hours$325k - $405k
...to learn from deployment and distribute the benefits of AI, while... ...an experienced Performance Engineer to help us scale the performance... ...building core services, training models, and developing real-time user... ...performance.Collaborate closely with infra, platform, training, and...TrainingWork at officeLocal areaRemote workFlexible hours- Senior ML Systems Engineer, Frameworks & Tooling at Cohere... ...serve humanity. We’re training and deploying frontier models for developers and... ...the intersection of large‑scale training, distributed systems, and HPC infrastructure... ...closely with infra teams to ensure Slurm...TrainingFull timeWork at officeRemote workFlexible hours
$293k - $385k
...building and applying performance modeling frameworks to understand... ...seeking Performance Modeling Engineers to develop and apply modeling... ...SkillsExposure to AI/ML workloads or distributed systems.Experience with... ...center infrastructure or large-scale systems.Experience working...Work at officeLocal areaRelocation packageFlexible hours- A leading technology firm in San Francisco is seeking a candidate to build and scale distributed training systems for large model pre-training. You will collaborate with research teams to design and operate training runs and enhance performance across distributed training...Training
$180k - $225k
...ever. New foundation models, reasoning techniques,... ...remains one of the hardest engineering challenges.As a... ...the latest advances in large language models, reasoning... ...building distributed production systems.Experience... ...relevant education or training. Scale employees in eligible...TrainingFull time- ...seeks a systems-focused ML infrastructure engineer to own scheduling, placement, and cluster management for large-scale model training. You will design multi-tenant schedulers,... ...engineers to translate workloads into scalable infra, optimize utilization, and build...Training
$300 per month
...Role:At Crusoe, our Production Engineering team ensures the reliability... ...with a strong background in distributed systems and hands-on experience with large language models to help us build and operate... ...teams to optimize large-scale training and inference clustersAutomate...TrainingTemporary work- ...on the front lines of innovation, supporting the engineering and research required to train large-scale AI models of unprecedented capability. About the Role... ...tools (e.g., shell scripting). ~ Experience with distributed systems to efficiently aggregate and analyze...TrainingFull time
$110.7k - $379.2k
Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare AI Three hundred... ...clean, synthesize, and evaluate large-scale instruction, preference,... ...based approaches; build scalable distributed training using DeepSpeed, FSDP, Megatron...TrainingLocal areaVisa sponsorship$100k - $120k
...generation robotic foundation models. As training and inference workloads... ...team of kernel and system engineers focused on performance-critical... ...kernel optimizations into distributed ML frameworks (e.g.,... ...validate improvements and de‑risk large‑scale rollouts Champion a...Training$167.2k - $209k
...DigitalOcean is seeking a Senior Engineer 2 to play a key technical... ...the world’s most advanced large models. As an IC leader, you will act... ...strategies across distributed GPU environments. Hardware Fluency... ...reimbursement for relevant conferences, training, and education. All...TrainingLocal areaRemote workWorldwideFlexible hours$285k - $315k
...looking for a Founding GPU Kernel Engineer who lives right at the... ...optimization passes that help every model we compile. What You'll Do... ...execution Experience with distributed training systems: collective ops like... ...background: experience with large-scale scientific computing,...TrainingFull timeWork at officeRelocation package$285k - $315k
...partnering with researchers, engineers, and organizations who... ...compiler. That means taking models from PyTorch, JAX, and... ...highly optimized binaries for large-scale AI pre-training. You'll own the entire... ...Nice to Have Background in distributed systems or multi-device compilation...TrainingFull timeWork at officeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Distributed Training Infra Engineer for Large Models. Be the first to apply!


