Staff Engineer - ML Inference & Model Efficiency
Cohere
A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have over 5 years of coding experience in C++ or Python and a solid understanding of the LLM inference environment. This position offers a remote-friendly work model, a competitive salary, and extensive benefits including a generous vacation policy.#J-18808-Ljbffr
- Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across... ...C++ or Python and insights into the LLM inference ecosystem. A commitment to diversity and...SuggestedRemote job
- ...About the Team Our Inference team brings OpenAI’s most... ...our start-of-the-art AI models, allowing them to do things... ...on performant and efficient model inference, as well... ...We are looking for an engineer who wants to take the world... ...of modern ML architectures and an intuition...SuggestedFull time
$298k - $368k
...diverse set of sensors, enabling engineers like you to (1) develop methods for efficiently and continuously learning from... ...real-world data, to (2) develop models and model training at scale, to... ...expertise in low-latency on-device inference techniques and a deep...SuggestedFull timeRemote work- ...Member of Technical Staff, Model Efficiency Who are we? Our mission is to scale... ...Cohere is a team of researchers, engineers, designers, and more, who are... ...focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop...SuggestedFull timeWork at officeRemote workFlexible hours
$190k - $265k
...their business. Founded by engineers — and customer-obsessed... ...data apps, AI agents, model training, model serving... ...real-time and batch inference, powering model inference... ..., latency, and efficiency of distributed AI workloadsCollaborate... ...platform, infra, and ML teams to deliver...SuggestedLocal areaWorldwide- ...seeks candidates with expertise in AI simulation development. The role emphasizes optimizing training efficiency, enhancing GPU performance, and ensuring low-latency inference. Applicants should be proficient in methodologies for gradient checkpointing, Nsight profiling,...
$203.5k - $299.3k
...are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash grows... ...production causal systems: uplift models, heterogeneous treatment effect models... ...practical experience with causal inference, econometrics, experimentation, or causal...Hourly payWork at officeLocal areaRemote workFlexible hours$203.5k - $299.3k
...hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash grows... ...production causal systems: uplift models, heterogeneous treatment effect models... ...practical experience with causal inference, econometrics, experimentation, or...Hourly payFull timeWork at officeLocal areaRemote workFlexible hours- ...Token Company in San Francisco is seeking a Member of Technical Staff for their infrastructure team. In this role, you will own the... ...API and build global low-latency, high-throughput GPU ML inference infrastructure. The ideal candidate will have solid experience...Visa sponsorship
$192k - $260k
...A leading data and AI company is seeking a Staff Engineer to design and implement core systems for Foundation Model Serving. The ideal candidate will have over 10 years of experience in building large-scale distributed systems and will collaborate closely across teams...$170k - $216k
...set of sensors, enabling engineers like you to (1) develop methods for efficiently and continuously... ...world data, to (2) develop models and model training at scale... ...model training and model inference through model architecture... ...~ Experience with ML frameworks like PyTorch...Full timeRemote work- ...infrastructure layer for AI and is seeking strong engineers to optimize ML systems for performance at scale. You... ..., pushing language and diffusion models toward higher throughput and lower... ...NVIDIA GPU architectures to maximize efficiency, while exploring low-level OS...
$250k - $300k
...your time making large language models run faster, cheaper, and more... .... That means owning the inference stack end to end: profiling where... ...work directly with customer engineering teams to tailor deployments... ...methods across many kinds of ML models, with an emphasis on large...Temporary work$225k
Dormont Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and ensuring high reliability for RL and post-training workflows. The ideal candidate...- Parallel Bio in San Francisco is looking for a candidate to own the training pipeline behind models essential for both their search stack and agents. You will be responsible for building pathways from real product usage to high-quality training data while rigorously fine...
- Parallel is seeking a professional who will own the training pipeline behind models that support both the search stack and agents. Responsibilities include building pathways from product usage to high-quality training data, rigorously fine-tuning models, and shipping them...
- Parallel Web Systems in Palo Alto is seeking a professional to own the training pipeline behind models that power both their search stack and agents. The role involves building connections from product usage to training data, fine-tuning models, and ensuring safe deployment...
- ...Together AI is building the best inference infrastructure for voice applications. We seek a Staff ML Engineer to own the model serving stack and optimize latency and throughput for real-time voice workloads. You'll work with state-of-the-art accelerators and collaborate...
- ...powers mission-critical inference for the world's most... ...to bring cutting-edge models into production. We're... ...help build the platform engineers turn to to ship AI products... ..., reliable, and cost‑efficient. As part of this team,... ...fundamentals and curiosity. ML experience is a plus,...Full timeFlexible hours
$137.1k - $201.6k
...leverage AI and advanced ML to power decision making... ...our next big bet to help efficiently grow our subscriber base... ...We’re looking for a Staff Machine Learning Engineer to drive the design and development... ...Contribute to Causal inference modeling to measure the incremental...Hourly payFull timeWork at officeLocal areaRemote workFlexible hours$166k - $225k
...their business. Databricks’ Model Serving product provides... ...platform to deploy and manage AI/ML models — from traditional... ...real-time, low-latency inference, governance, monitoring, and... ...with strong SLAs and cost efficiency.As a Senior Engineer, you’ll play a critical role...Local areaWorldwide$175k - $215k
...learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust... ...the petabyte-scale data systems and ML pipelines at the heart of Waymo's foundation... ...leaps in the speed, reliability, and efficiency of the end-to-end ML development...Full timeRemote work- ...experiences, running generative models on-device, right where... ...Machine Learning Engineer for On-Device & Mobile... ...parts of the inference stack — from a trained... ...-bandwidth level.Apply efficiency techniques — dynamic resolution... ...integration between the ML runtime and the game engine...Full timeWork at officeRemote workWorldwide
$325k
...leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate... ...engineering experience, strong familiarity with ML architectures, and experience with distributed systems....$250k - $300k
...of experience in storage engineering, including operating distributed... ..., reliability, or cost efficiency at scale We need deep... ...to-have familiarity with ML/AI storage patterns such as model weights, checkpointing,... ...for training and inference workloads We will tune...Full timeRemote work$192k - $260k
...business. Foundation Model Serving is the API... ...frontier AI model inference for open source... ...this role, no prior ML or AI experience... ...We’re looking for engineers who have owned high... ...at scale.As a Staff Engineer, you’ll play... ..., and operational efficiency for GPU serving workloads...Local areaWorldwide$220k - $320k
...Help us make inference blazingly fast. If you love squeezing... ...specialized language models for companies that need... ...ten‑person team of engineers who work in‑person in downtown... ...stack as fast and efficient as possible. Your work... ...Collaborate with applied ML engineers to ensure...Work at office$213k - $263k
...learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust... ...operate the petabyte-scale data systems and ML pipelines at the heart of Waymo's... ...infrastructure to ensure the dependable and efficient rollout of data solutions at the...Full timeTemporary workRemote work$220k
We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust... ...laid out for you. Good if you touched any of ML compilers and framework internals: PyTorch...- Together AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms, architectures, and scheduling for low-latency, high-throughput inference. You will implement changes in production-grade inference engines, including kernel backends...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Engineer - ML Inference & Model Efficiency. Be the first to apply!
- assistant engineering manager San Francisco, CA
- assistant civil engineer San Francisco, CA
- assistant mechanical engineer San Francisco, CA
- assistant engineer San Francisco, CA
- staff engineer San Francisco, CA
- staff data engineer San Francisco, CA
- software engineer staff San Francisco, CA
- assistant electrical engineer San Francisco, CA
- staff design engineer San Francisco, CA
- senior staff engineer San Francisco, CA




