Member of Technical Staff — Training Infrastructure
causal
Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it. To achieve this breakthrough, we are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect. We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather. Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. We look for infrastructure engineers who are excited to tackle unsolved problems. Training an LPM means scaling novel architectures over multimodal physical data — a problem where the playbooks from language and vision only partially apply. Your mission is to make large-scale training fast, efficient, and reliable, so that every GPU cycle accelerates research progress. Responsibilities Design, implement, and optimize distributed training systems that scale across thousands of GPUs Research and test parallelization strategies and numerical precision trade-offs across model scales, including for architectures that don't map cleanly onto existing LLM training stacks Analyze, profile, and debug low-level GPU operations to maximize throughput and hardware utilization Build reusable frameworks for checkpointing, fault tolerance, and reproducibility that stay robust under rapid research iteration Collaborate with researchers to bring novel model architectures from prototype to full scale Stay up-to-date on research to bring new ideas to work What we're looking for We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains. Demonstrated proficiency with distributed training frameworks and techniques (e.g. FSDP, DeepSpeed, Megatron, Pytorch, JAX/XLA) to train large foundation models Strong grasp of state‑of-the‑art techniques for optimizing training workloads: parallelism strategies, memory optimization, mixed precision, communication overlap Ability to profile and debug performance in complex codebases, from framework internals down to kernels and collectives Deep understanding of deep learning frameworks (e.g. PyTorch, JAX) and their underlying system architectures Bonus: contributions to open‑source ML infrastructure (e.g. PyTorch, Megatron‑LM, DeepSpeed, XLA) #J-18808-Ljbffr causal
- ...intelligence to serve humanity. We’re training and deploying frontier models for... ...embrace being remote-friendly! As a Member of Technical Staff, you will: Design and write high-performant... ...ideas on our supercompute and data infrastructure. Learn from and work with the best...Technical trainingFull timeWork at officeRemote workFlexible hours
- ...real-world business problems. We’re training and deploying frontier models for... ...embrace being remote-friendly! As a Member of Technical Staff, you will: Design and write high-performant... ...ideas on our supercompute and data infrastructure. Learn from and work with the best...Technical trainingFull timeWork at officeLocal areaRemote workHome office
$150k - $300k
Building Open Superintelligence Infrastructure Prime Intellect is building... ...that lets anyone create, train, and deploy them. We aggregate... ...that runs the jobs. Core Technical Responsibilities Hosted Training... ...and encourage team members to contribute to the broader...Technical trainingWork at officeLocal areaRemote workVisa sponsorshipRelocation packageFlexible hours$250k
...compute platform building the next generation of agentic infrastructure for GPU-intensive workloads. Operating across the full technology... .... This opportunity offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's growth. You'll help...SuggestedFull time- ...interactive applications.We develop the infrastructure that enables intelligent systems to... ...InfrastructurePhysical AI The Role We are hiring a Member of Technical Staff to lead reinforcement learning infrastructure and model post-training.You will work closely with Qi and the...Technical training
- Member of Technical Staff - Infrastructure Security We're partnering with a frontier AI research company that is building next-generation open-weight foundation models with the mission of making advanced AI broadly accessible. Their team includes researchers, engineers...
- ...DeepMind, OpenAI, Google Brain, Meta, Character.AI, Anthropic and beyond. Role Overview Reflection.AI is looking for a Member of Technical Staff - Infrastructure Security to secure our geographically diverse multi-cloud Kubernetes and cloud environments. In this role, you’ll...Relocation package
- # Founding Member of Technical Staff, AI Infrastructure**Location:** San Francisco / Bay Area preferred. Remote exceptional for the right person.We look for fast learners with high agency, AI-native workflows, clear technical communication, and evidence-backed judgment...Full timeRemote work
$200k - $400k
...to an algorithm. We're building the infrastructure to understand human behavior at scale... ...simulations for F100 enterprises, trained a first-of-its-kind confidence model... ...possible to run. About the Role As a Member of Technical Staff in Research Infrastructure, you will...Live inFlexible hours- ...people from OpenAI, xAI and DoorDash. The Token Company trains machine learning models to compress raw LLM inputs... ...is 5 people with a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own the cloud systems that serve our...Visa sponsorship
$200k
Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important... ...alone. Our approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and...RelocationVisa sponsorship- ...observe their code. We are responsible for designing, building, and scaling core infrastructure that powers a high-volume data platform for AI applications. We are looking for team members who love building enabling systems that empower our engineers and power our rapidly...Work at office
$350k
...possible. We are building a frontier AI research company and training our own models end‑to‑end. Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large‑scale experiments. Our team includes researchers and...Technical training$350k
...possible. We are building a frontier AI research company and training our own models end-to-end. Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments. Our team includes researchers and...Technical training$350k
...possible. We are building a frontier AI research company and training our own models end-to-end. Our work spans areas such as model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments. Our team includes researchers and...Technical training$225k
...than humans can alone. Our approach combines frontier-scale pre-training, domain-specific RL, ultra-long context, and inference-time... ...Systems team, you will design and operate the distributed infrastructure that trains Magic’s long-context models at scale. This role...Technical trainingRelocationVisa sponsorship- ...world deployment. You’ll own applied post-training work end-to-end for some of the world’s... ...between customer needs and internal technical teams, and push back when needed. The... ...shared or general-purpose post-training infrastructure. Prior exposure to customer-facing or...Technical training
- ...tasks. The present bottleneck is the lack of high-quality RL training environments. Our first step is to build RL environments that... ...has previous experience on Anthropic’s data team building data infrastructure, and datasets behind Claude. We are partnering with leading...Technical trainingVisa sponsorshipRelocation package
- ...Anthropic and beyond. About the Role Build and scale distributed training systems that power frontier model pre-training. Work closely... ...large-scale training runs for foundation models. Develop infrastructure that enables efficient training across thousands of GPUs...Technical trainingRelocation package
- About Us: AI needs a new infrastructure layer. We're building it at Modal. Every era of computing brought new workloads that previous infrastructure... ...re building a platform that covers the whole life of an LLM: training it, deploying it, and observing it in production. We already...Technical trainingWork at office
- ...Meta, Character.AI, Anthropic and beyond. About the Role Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads. Develop systems that power synthetic data generation and reinforcement learning pipelines...Technical trainingRelocation package
- ...deep learning literature. Lead small research projects independently while collaborating on larger initiatives Optimize the training infrastructure for efficient scaling. Contribute across the entire stack, from low-level optimizations to high-level model design. About...Technical trainingFull timeRelocation package
$250k - $300k
...career? Join one of the most interesting infrastructure companies in the AI space right now.... ...that the world's leading AI labs train on. Today they manage tens of thousands... ...Have: A track record of impressive technical work you can speak to in depth; the years...Full timeRemote work- ...people take ownership, grow together, and share both the challenges and the wins. What You'll Do Build the supercomputing infrastructure that runs our agents. Our agents tackle long-horizon, high-performance workloads, and you'll design the cloud compute,...Work at officeRemote workFlexible hours
$250k - $300k
...development? Join one of the most exciting AI infrastructure companies in the market, building a... ...powering next-generation AI training and inference at scale. This role offers... ...Have: ~ A track record of impressive technical work you can speak to in depth, the years...Full timeRemote work- ...Horowitz, GIC, Goldman Sachs, KKR, Visa, and others. Technical Skills Develop and maintain infrastructure that powers digital asset custody, trading, staking,... ...to solve problems, and assist or teach other team members when possible. You may be a fit for this role if you...Worldwide
$150k - $300k
Building Open Superintelligence Infrastructure Prime Intellect is building the open superintelligence... ...infra that enables anyone to create, train, and deploy them. We aggregate and... ...for GPU Infrastructure, you'll be the technical expert who transforms customer requirements...- ...continuously, in many formats, at a scale that dwarfs what is used to train today's LLMs. Your mission is to build the data platform... ...that data and research teams build their checks on Scale infrastructure to improve engineering velocity and ensure reliability, with monitoring...Immediate start
- About Us Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates... ...also have Experience building or operating AI inference, training, HPC, or neocloud infrastructure. Experience with bare‑metal...
- ...users create characters, worlds, stories, and relationships with AI, and making that feel fast, reliable, and alive takes serious infrastructure. We are looking for an engineer who wants to help own that whole stack. We run more of our own than most companies our size....
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff — Training Infrastructure. Be the first to apply!
- work from home technical support specialist San Francisco, CA
- product support technician San Francisco, CA
- helpdesk support technician San Francisco, CA
- help desk assistant San Francisco, CA
- senior technical associate San Francisco, CA
- IT help desk technician San Francisco, CA
- technical solutions specialist San Francisco, CA
- desktop support analyst San Francisco, CA
- trade support analyst San Francisco, CA
- senior IT support technician San Francisco, CA


