Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Distributed Training Infra Engineer for Large Models

Kindredventures

Kindredventures in San Francisco is seeking an infrastructure engineer to scale distributed training for Large Physics models. You will design, implement, and optimize systems that run thousands of GPUs and accelerate research progress. Collaborate with researchers to bring prototype models to full scale, optimize memory and throughput, and contribute to open-source ML infrastructure. You should have strong expertise in PyTorch and JAX and a track record of performance profiling. #J-18808-Ljbffr Kindredventures

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Distributed Training Infra Engineer for Large Models in San Francisco, CA vacancy
  • $227.2k - $417k

     ...the Role:As a Software Engineer on the ML...  ...maintaining low-latency ML model serving systems that...  ...throughput, and low latency distributed systems using...  ...and efficiency of our infra. Lead large scale cross functional...  ..., ElastiCache, model training orchestration, etc.Understanding... 
    Training
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    1 day ago
  •  ...technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation models. You will design distributed training systems and optimize GPU utilization while collaborating... 
    Training

    Baseten

    San Francisco, CA
    3 days ago
  • Magic AI, Inc. is seeking a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model execution meets distributed infrastructure, influencing latency,... 
    Training

    Magic AI Corp.

    San Francisco, CA
    3 days ago
  •  ...Protocol Learning : multi-participant training of foundation models where no single participant has, or can ever...  .... We’re looking for Senior/Staff engineers with 5+ years of experience in distributed systems and ML large‑scale training. You’ll be implementing a... 
    Training
    Remote work
    Visa sponsorship

    Pluralis Research

    San Francisco, CA
    1 day ago
  •  ...the capabilities of foundational models to support general-purpose robotics...  ...About the Role As a Research Engineer, Distributed Data Systems, you will design and...  ...scale the infrastructure that powers large-scale multimodal training and evaluation at OpenAI. You’ll manage... 
    Training
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    21 hours ago
  • $190k - $205k

     ...noise characteristics. Design models that are robust,...  ...and visual signals Production Engineering Write clean, scalable, well...  ...code that integrates into a large shared codebase. Build end...  ...processing, feature extraction, training, evaluation, and deployment.... 
    Training
    Full time
    Live in

    Gridware

    San Francisco, CA
    21 hours ago
  •  ...great server internals engineer to help maintain and...  ...disclosed vulnerabilities in large software systems and...  ...CVSS scoring, threat modeling, and security risk...  ...work with a friendly, distributed team following open-source...  ...also offer access to training on leading-edge... 
    Training
    Full time
    Remote work
    Flexible hours

    Altinity

    San Francisco, CA
    3 days ago
  • $350k

    Mirendil is looking for engineers to build infrastructure for frontier reasoning models at their San Francisco location. This role focuses on large-scale reinforcement learning (RL) model training and requires a solid understanding of engineering principles. The ideal... 
    Training

    Mirendil

    San Francisco, CA
    3 days ago
  • $160k - $194k

     ...manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in...  ...application interface and database model changes, migrate and evolve schemas,...  ...sets, market conditions, experience and training, licensure and certifications, and... 
    Training
    Full time
    Remote work
    Flexible hours

    Ad Verse

    San Francisco, CA
    21 hours ago
  •  ...seeking an experienced backend/infrastructure engineer to build platforms powering AI workloads, including model training, serving, and vector search. You will join a high...  .... You will collaborate across platform, infra, and ML teams to deliver end-to-end experiences... 
    Training

    Databricks

    San Francisco, CA
    21 hours ago
  • $180k - $275k

     ...helping define and evolve the core data model and storage systems powering Gamma's business...  ...rapid shipping velocity. As Software Engineer on the Platform team, you'll collaborate...  ...Design and implement scalable APIs, distributed systems, and data infrastructure that... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    21 hours ago
  • $117.2k - $223.9k

     ...integrating Google's Gemini AI models into Salesforce's Agentforce...  .... Join our team of talented engineers and help us advance the integration...  ..., and the security of distributed and scalable distributed systems...  ..., promotion, benefits, training, assessment of job performance... 
    Training
    Full time

    Salesforce

    San Francisco, CA
    1 day ago
  • $200.8k - $251k

     ...Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • A leading tech company in San Francisco seeks a Machine Learning Engineer to build and maintain infrastructure for large-scale model training. In this hands-on role, you will design systems, work closely with researchers, and optimize training processes. Candidates should... 
    Training

    Monograph

    San Francisco, CA
    3 days ago
  •  ...San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will...  ..., monitoring tools, and robust performance improvements for large‑scale runs. #J-18808-Ljbffr GenesisAI
    Training

    GenesisAI

    San Francisco, CA
    4 days ago
  • $179.4k - $224.25k

     ...building upon our prior model evaluation work with...  ...EngineOur Generative AI Data Engine powers the world’s most...  ...on everything from large-scale system architecture...  ...data processing and distributed systems.Familiarity with...  ...relevant education or training. Scale employees in eligible... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $293k - $385k

     ...work closely with hardware, modeling, and architecture teams to...  ...Workload Porting & Performance Engineer to evaluate new hardware...  ....Experience working in large-scale or distributed system environments.Preferred...  ...AI/ML workloads, including training or inference systems.Familiarity... 
    Training
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    1 day ago
  • $325k - $405k

     ...to learn from deployment and distribute the benefits of AI, while...  ...an experienced Performance Engineer to help us scale the performance...  ...building core services, training models, and developing real-time user...  ...performance.Collaborate closely with infra, platform, training, and... 
    Training
    Work at office
    Local area
    Remote work
    Flexible hours

    OpenAI

    San Francisco, CA
    1 day ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere...  ...serve humanity. We’re training and deploying frontier models for developers and...  ...the intersection of large‑scale training, distributed systems, and HPC infrastructure...  ...closely with infra teams to ensure Slurm... 
    Training
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    1 day ago
  • $293k - $385k

     ...building and applying performance modeling frameworks to understand...  ...seeking Performance Modeling Engineers to develop and apply modeling...  ...SkillsExposure to AI/ML workloads or distributed systems.Experience with...  ...center infrastructure or large-scale systems.Experience working... 
    Work at office
    Local area
    Relocation package
    Flexible hours

    OpenAI

    San Francisco, CA
    1 day ago
  • A leading technology firm in San Francisco is seeking a candidate to build and scale distributed training systems for large model pre-training. You will collaborate with research teams to design and operate training runs and enhance performance across distributed training... 
    Training

    Reflection

    San Francisco, CA
    1 day ago
  • $180k - $225k

     ...ever. New foundation models, reasoning techniques,...  ...remains one of the hardest engineering challenges.As a...  ...the latest advances in large language models, reasoning...  ...building distributed production systems.Experience...  ...relevant education or training. Scale employees in eligible... 
    Training
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  •  ...seeks a systems-focused ML infrastructure engineer to own scheduling, placement, and cluster management for large-scale model training. You will design multi-tenant schedulers,...  ...engineers to translate workloads into scalable infra, optimize utilization, and build... 
    Training

    Physical Intelligence

    San Francisco, CA
    2 days ago
  • $300 per month

     ...Role:At Crusoe, our Production Engineering team ensures the reliability...  ...with a strong background in distributed systems and hands-on experience with large language models to help us build and operate...  ...teams to optimize large-scale training and inference clustersAutomate... 
    Training
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  •  ...on the front lines of innovation, supporting the engineering and research required to train large-scale AI models of unprecedented capability. About the Role...  ...tools (e.g., shell scripting). ~ Experience with distributed systems to efficiently aggregate and analyze... 
    Training
    Full time

    OpenAI

    San Francisco, CA
    21 hours ago
  • $110.7k - $379.2k

    Position Summary Research Engineer — Post-Training & Small Language Models (SLMs), Healthcare AI Three hundred...  ...clean, synthesize, and evaluate large-scale instruction, preference,...  ...based approaches; build scalable distributed training using DeepSpeed, FSDP, Megatron... 
    Training
    Local area
    Visa sponsorship

    Deloitte

    San Francisco, CA
    21 hours ago
  • $100k - $120k

     ...generation robotic foundation models. As training and inference workloads...  ...team of kernel and system engineers focused on performance-critical...  ...kernel optimizations into distributed ML frameworks (e.g.,...  ...validate improvements and de‑risk large‑scale rollouts Champion a... 
    Training

    Coda Robotics

    San Francisco, CA
    21 hours ago
  • $167.2k - $209k

     ...DigitalOcean is seeking a Senior Engineer 2 to play a key technical...  ...the world’s most advanced large models. As an IC leader, you will act...  ...strategies across distributed GPU environments. Hardware Fluency...  ...reimbursement for relevant conferences, training, and education. All... 
    Training
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    11 hours ago
  • $285k - $315k

     ...looking for a Founding GPU Kernel Engineer who lives right at the...  ...optimization passes that help every model we compile. What You'll Do...  ...execution Experience with distributed training systems: collective ops like...  ...background: experience with large-scale scientific computing,... 
    Training
    Full time
    Work at office
    Relocation package

    SF Tensor

    San Francisco, CA
    11 hours ago
  • $285k - $315k

     ...partnering with researchers, engineers, and organizations who...  ...compiler. That means taking models from PyTorch, JAX, and...  ...highly optimized binaries for large-scale AI pre-training. You'll own the entire...  ...Nice to Have Background in distributed systems or multi-device compilation... 
    Training
    Full time
    Work at office
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    11 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Distributed Training Infra Engineer for Large Models. Be the first to apply!