Senior ML Systems Engineer, Frameworks & Tooling (San Francisco)
Cohere
Senior ML Systems Engineer, Frameworks & Tooling at Cohere
Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents. We believe that our work is instrumental to the widespread adoption of AI.
Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft. Each person is one of the best in the world at what they do. We believe that a diverse range of perspectives is a requirement for building great products. We obsess over what we build and work hard and move fast to do what’s best for our customers. Join us on our mission and shape the future!
We’re looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier-scale language models. This role sits at the intersection of large‑scale training, distributed systems, and HPC infrastructure. You will design and maintain the core components that enable fast, reliable, and scalable model training and build the tooling that connects research ideas to thousands of GPUs.
Responsibilities
- Build and own the training framework responsible for large-scale LLM training.
- Design distributed training abstractions (data, tensor, and pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing).
- Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100).
- Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics.
- Collaborate closely with infra teams to ensure Slurm setups, container environments, and hardware configurations support high-performance training.
- Investigate and resolve performance bottlenecks across the ML systems stack.
- Build robust systems that ensure reproducible, debuggable, large-scale runs.
You Might Be a Good Fit If You Have
- Strong engineering experience in large-scale distributed training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops.
- Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).
- Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines.
- Experience working with containerized environments (Docker, Singularity/Apptainer).
- A track record of building tools that increase developer velocity for ML teams.
- Excellent judgment around trade-offs: performance vs complexity, research velocity vs maintainability.
- Strong collaboration skills — you’ll work closely with infra, research, and deployment teams.
Nice to Have
- Experience with training LLMs or other large transformer architectures.
- Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.).
- Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches).
- Experience with data pipeline optimization, sharded datasets, or caching strategies.
- Background in performance engineering, profiling, or low-level systems.
Why Join Us
- You’ll work on some of the most challenging and consequential ML systems problems today.
- You’ll collaborate with a world‑class team working fast and at scale.
- You’ll have end-to-end ownership over critical components of the training stack.
- You’ll shape the next generation of infrastructure for frontier-scale models.
- You’ll build tools and systems that directly accelerate research and model quality.
Sample Projects
- Build a high-performance data loading and caching pipeline.
- Implement performance profiling across the ML systems stack.
- Develop internal metrics and monitoring for training runs.
- Build reproducibility and regression testing infrastructure.
- Develop a performant fault-tolerant distributed checkpointing system.
We value and celebrate diversity and strive to create an inclusive work environment for all. We welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.
Full‑Time Employees At Cohere Enjoy These Perks
- An open and inclusive culture and work environment
- Work closely with a team on the cutting edge of AI research
- Weekly lunch stipend, in‑office lunches & snacks
- Full health and dental benefits, including a separate budget to take care of your mental health
- 100% Parental Leave top‑up for up to 6 months
- Personal enrichment benefits towards arts and culture, fitness and well‑being, quality time, and workspace improvement
- Remote‑flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co‑working stipend
- ✈️ 6 weeks of vacation (30 working days!)
- A leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across...SeniorFull timeFlexible hours
$160k - $185k
...Senior Agentic AI Engineer A frontier AI company is building systems that can act in the physical world... ...next‑gen LLM tool‑use Join early... ...calling flows using frameworks like LangGraph... ...Partner with ML, infra, and systems... ...Location: San Francisco, CA Salary: $...SeniorFull time- ...A leading technology firm is seeking a Senior Machine Learning Engineer to join their team in San Francisco. This hybrid role requires expertise in the entire machine... ...skills in Python, and familiarity with modern ML tools. Join to make an impact on how users connect with...SeniorFull time
- An innovative AI infrastructure startup in San Francisco is looking for a Senior ML Engineer to lead the development and operation of LLMs in production. The... ...be proficient in modern programming and monitoring tools, with strong collaboration skills. This hybrid role allows...SeniorFull timeWork at officeFlexible hours2 days per week3 days per week
- ...a small team of engineers is working on what... ...rack-scale systems worth millions of... ...hats (building ML platforms, MLOps tools, data/LLM infrastructure... .... As a Senior ML Engineer, you... ...from our downtown San Francisco office 2 to 3 days... ...with modern ML frameworks/libraries such as...SeniorFull timeWork at officeFlexible hours2 days per week3 days per week
$133.5k - $212k
...including offices in San Francisco, New York, Denver,... ...are looking for a Senior Machine Learning Engineer to build the core... ...: retrieval systems, evaluation frameworks, and model integration... ...Implement abstractions, tooling, and reusable... ...teams to build ML- and LLM-powered experiences...SeniorFull timeContract workLocal areaImmediate startRemote workWorldwideHome officeFlexible hours$175k - $250k
...A dynamic AI data science platform in San Francisco is looking for a Senior Engineer with a strong background in Python software development and machine learning... ...involves building and optimizing high-performance tools for data analysis within a diverse team. Successful...SeniorFull time- ...Senior Staff Machine Learning Engineer, Community Support Engineering... ...guests to their San Francisco home, and has... ...Marketing we rely on ML to ensure that... ...services and tools including LLM... ...Machine Learning systems, enable fast... ...robust testing frameworks for agent behavior...SeniorFull timeWork experience placementCasual workLive inWork at officeRemote work
$200k - $300k
...the job poster from Acceler8 Talent Senior Neuro-Symbolic Systems Engineer - San Francisco, CA A company building AI systems... ...into real workflows Develop tools for evaluating correctness, consistency... ...ensure symbolic layers work alongside ML, RL, and systems architecture...SeniorFull timeImmediate start- ...A tech recruiting firm in San Francisco is seeking a candidate to build robust agent systems using cutting-edge frameworks. The role involves designing workflows with complex logic and rigorous validation while having a significant influence within a technical environment...SeniorFull time
- ...Palo Alto or San Francisco offices and will... ...you’ll be the senior technical... ...doubling down on ML as the future... ...foundational systems surrounded by... ...collaborating with engineering, data science... ...integrate emerging AI tools and techniques... ...popular ML frameworks. A scrappy...Full timeCasual workWork at officeImmediate startFlexible hours
$164k - $312k
A leading technology company in San Francisco seeks a Senior ML Engineer to drive advanced prediction models for ad ranking systems. You will design and implement efficient systems that optimize engagement predictions. The ideal candidate holds an advanced degree and has...SeniorFull time- ...A growing healthtech startup in San Francisco is seeking a Member of Technical Staff – Machine Learning to build and scale machine learning systems for their AI-powered platform. You will have 5+ years of experience with ML systems in production and strong skills in Python...SeniorFull time
- ...A leading educational technology company in San Francisco is seeking a Machine Learning Engineer to drive AI initiatives. This role requires extensive experience in Python, ML libraries, and a solid understanding of NLP. The successful candidate will collaborate with various...SeniorFull timeRemote workFlexible hours
- A leading social media platform in San Francisco seeks a Machine Learning Engineer to develop personalized experiences using innovative ML techniques. The role demands expertise in recommendation systems and data processing, providing a unique opportunity to impact a user...SeniorFull time
- A leading AI evaluation platform in San Francisco is looking for a Senior Software Engineer specializing in ML infrastructure. The successful candidate will design and develop robust real-time data and API systems, enabling insights for researchers and developers. Ideal...SeniorFull time
- A leading AI platform company in San Francisco is seeking a talented individual to optimize model inference for their advanced visual AI product. The ideal candidate will engage in building efficient AI models and tackling complex challenges. The role requires a strong...SeniorFull time
- ...A biotechnology firm in San Francisco is searching for a Founding Machine Learning Engineer to develop scalable ML systems focused on RNA biology. This entry-level full-time role... ..., along with expertise in deep learning frameworks. Competitive salary package available....Full time
- A technology startup is seeking a Founding Engineer (Systems + ML) to develop GPU-accelerated engines and build end-to-end pipelines. The ideal... ...and a chance to shape the future of chip design. This is a full-time position based in San Francisco, California. #J-18808-LjbffrFull time
- A leading AI technology firm in San Francisco seeks a Machine Learning Engineer to join their AI/ML team. The role involves creating innovative AI experiences, training models, and developing unique features to wow customers. Candidates should have extensive experience...SeniorFull time
$175k - $250k
...yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote... ...empowerment, their tools help users experiment... ...distributed systems that power AI workloads... ...infrastructure, ML pipelines, or GPU... ...languages or frameworks Skills: prometheus...SeniorFull timeRemote workRelocationRelocation package- ...Senior / Staff / Principal Machine Learning Engineer Location: Onsite San Francisco (5 days onsite AND hybrid options) We... ...Designing and implementing ML algorithms and... ...and evaluation. System Integration: Collaborating... ...languages. ML Frameworks: Experience with...SeniorFull time
$200k - $300k
...A leading AI company based in San Francisco is seeking a Senior Neuro-Symbolic Systems Engineer to enhance AI systems interacting with the physical world. Candidates... ...cross-functional teams, and bridge symbolic AI with ML techniques. The role offers a competitive salary...SeniorFull time- ...A DeFi-native company in San Francisco is seeking a systems-focused engineer to join their small team. You will take ownership of significant projects, such as designing onchain event ingestion systems and optimizing performance-sensitive components. Ideal candidates are...SeniorPart time
$220k - $270k
...cryptography, mobile engineering, and global... ...Organization at Tools for Humanity is responsible... ...is looking for a Senior Software... ...engineers at our new San Francisco office. This is a... ...and Android system components, including... ...internals, HAL framework, and Android system...SeniorFull timeWork at officeOverseasFlexible hours$200k
...A tech startup in San Francisco is seeking senior or staff-level engineers to design and maintain their global cloud infrastructure. Ideal candidates will have... ...engineering experience, a strong understanding of systems at scale, and a passion for optimizing performance...SeniorFull time- ...Do Build agent planners, tool-call flows, and structured state... ...using LangGraph or similar frameworks Design schemas, action adapters... ...experience building agent systems, LLM tool-calling pipelines,... ...Work on complex, high-impact engineering workflows with real...SeniorFull time
- ...A leading hospitality platform in San Francisco is seeking a Staff Machine Learning Engineer to enhance guest and host experiences through cutting-edge machine learning... ...contribute to improving product experiences using ML. Ideal candidates should have extensive experience...Full time
- ...leading financial technology company in San Francisco is seeking a Senior Software Engineer to enhance its Machine Learning... ...role involves designing scalable systems, collaborating with cross-... ...with AWS, distributed systems, and ML workflows. Competitive compensation...SeniorFull time
- A leading customer engagement platform in San Francisco seeks a Senior Machine Learning Engineer to build core ML foundations. Responsibilities include designing and... ...candidate has over 5 years of experience in production systems, strong engineering skills with Python or...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior ML Systems Engineer, Frameworks & Tooling (San Francisco). Be the first to apply!
- senior ml engineer San Francisco, CA
- machine learning engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- system performance engineer San Francisco, CA
- software system engineer San Francisco, CA
- ground systems engineer San Francisco, CA
- application system engineer San Francisco, CA
























