Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff ML Inference Engineer Model Efficiency (Remote)

Jaide Health

Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model execution. You'll work with advanced performance techniques such as GPU/CUDA optimizations and collaborate closely with modeling and systems teams. Ideal candidates will have over 5 years of experience in high-performance coding, plus strong skills in C++ or Python and insights into the LLM inference ecosystem. A commitment to diversity and inclusive work culture is celebrated. #J-18808-Ljbffr Jaide Health

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff ML Inference Engineer Model Efficiency (Remote) in San Francisco, CA vacancy
  • $251k - $310k

     ...machine learning (ML) engineers, software engineers...  ...ML algorithms, to model the real world,...  ...report to a Senior Staff Engineering Manager...  ...bottlenecks in training and inference performance (e.g.,...  ...distillation, and efficient attention...  ...can be performed remote, the specific salary... 
    Remote work
    Full time

    Waymo

    Remote
    22 hours ago
  •  ...cutting-edge foundation AI models and end-to-end products that...  ...is a team of researchers, engineers, designers, and more, who are...  ...systems and optimize audio inference serving efficiency using innovative techniques...  ...and London. We embrace a remote-friendly environment, and as... 
    Remote work
    Work at office
    Local area
    Home office

    Cohere

    New York, NY
    1 day ago
  • $195.2k - $262.2k

     ...enterprises from data and model training through to...  ...large in-house AI/ML infrastructure. Built by engineers, for engineers....  ...orchestration to inference optimization, we own...  ...programs in efficient LLM and VLM inference...  ...secondary caregivers. Remote work reimbursement:... 
    Remote work
    Temporary work
    Immediate start

    Nebius

    Palo Alto, CA
    12 hours ago
  • $117.7k - $221.4k

     ..., practical, and cost efficient for embodied AI systems...  ...not only on stronger models, but also on better infrastructure...  ...reflects how Cola engineers think: build durable...  ..., featurization, and inference foundations that power...  ...This role is based remotely, but if the selected... 
    Remote work
    Full time
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    1 day ago
  • $175k - $215k

     ...learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust...  ...research setting developing recipes for ML models We prefer: - Track record of...  ...or, if the role can be performed remote, the specific salary range for your preferred... 
    Remote work
    Full time

    Waymo

    New York, NY
    22 hours ago
  • $170k - $216k

     ...sensors, enabling engineers like you to (1) develop methods for efficiently and continuously learning...  ..., to (2) develop models and model training...  ...and model inference through model architecture...  ...~ Experience with ML frameworks like...  ...can be performed remote, the specific... 
    Remote work
    Full time

    Waymo

    Remote
    22 hours ago
  • $298k - $368k

     ...set of sensors, enabling engineers like you to (1) develop methods for efficiently and continuously...  ...world data, to (2) develop models and model training at scale...  ...low-latency on-device inference techniques and a deep...  ...role can be performed remote, the specific salary range... 
    Remote work
    Full time

    Waymo

    San Francisco, CA
    22 hours ago
  •  ...and deploying frontier models for developers and...  ...team of researchers, engineers, designers, and more,...  ...on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques...  ...London. We embrace a remote‑friendly environment,... 
    Remote work
    Full time
    Work at office
    Flexible hours

    Cohere

    New York, NY
    4 days ago
  • Apple Inc. in the San Francisco Bay Area is looking for a Staff/Sr. Machine Learning Engineer to optimize inference for large foundation models and deliver production-grade solutions that serve millions of users in real time. You will collaborate closely with research and... 

    Apple Inc.

    Seattle, WA
    3 days ago
  •  ...technology company in Seattle is seeking a Machine Learning Engineer for Model Serving Infrastructure. The ideal candidate will have at least...  ...in C/C++/CUDA. You will design and implement distributed inference infrastructure and collaborate with product teams. This position... 

    ByteDance

    Seattle, WA
    1 day ago
  •  ...Overview: The Principal AI/ML Engineer will support the development...  ...learning, and large language models. We offer generous relocation...  ...of applications within remote sensing such as tasking collections...  ...engineering techniques / Inference time techniques (e.g. chain of... 
    Remote work
    Full time
    Temporary work
    Work at office
    Local area
    Visa sponsorship
    Relocation package
    Flexible hours

    Arka Group, Lp

    Remote
    22 hours ago
  • $114.6k - $252.1k

    Job Title: Principal AI/ML Engineer (Large Language Model)Job Category: ScienceTime Type: Full timeMinimum...  ...to a variety of applications within remote sensing such as tasking collections,...  ...factory)Prompt engineering techniques / Inference time techniques (e.g. chain of... 
    Remote work
    Contract work
    Work experience placement
    Local area
    Flexible hours

    CACI International

    Philadelphia, PA
    1 day ago
  •  ...include occasional remote work, starting the...  ...for a performance engineer who specializes in...  ...fast and cost-efficient in the datacenter....  ...-throughput batch inference sweeping petabytes...  ...of accelerators, ML frameworks, and large...  ...roofline and performance models for our workloads,... 
    Remote work
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Day shift

    Applied Intuition

    Sunnyvale, CA
    2 days ago
  •  ...Mirantis empowers platform engineering teams to deliver composable,...  ...technical Product Manager to own AI inference and model serving for k0rdent AI, our...  ...technical role owning AI/ML and inference product(s) ~...  ...job opportunities. #remote We are a Leader for Container... 
    Remote work

    Mirantis

    Austin, TX
    more than 2 months ago
  • Bright Vision Technologies is seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance inference platforms for serving large ML models in production. The role emphasizes systems engineering, including request routing, batching, autoscaling... 
    Remote job

    Bright Vision Technologies

    Santa Clara, CA
    12 hours ago
  • $85 per hour

     ...realistic machine learning engineering tasks and assessing model outputs. You will perform...  ...assessments that reflect production ML workflows, helping a...  ...involving model training, inference systems, MLOps, and LLM...  ...Work Terms Location: Remote. Employment type: Hourly... 
    Remote work
    Hourly pay

    SaidGig

    United States
    17 days ago
  •  ...About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of...  ...acceleration, GPU parallelism, advanced model deployment, and video generation...  ...significant improvements to model speed and efficiency, ensuring our creative AI systems... 
    Full time
    Work at office
    3 days per week

    Pika

    Remote
    22 hours ago
  • $200k - $250k

     ...accuracy, and trust. Our ML systems sit at the...  ...a Senior MLOps Engineer to help us run them reliably and efficiently in production....  ...across a custom-built inference platform powering...  ..., and extraction models), each with...  ...holidays ~ Fully remote work within the United... 
    Remote work
    Full time
    Flexible hours

    Wizard

    United States
    22 hours ago
  •  ...security, transparency, and efficiency. About the Role We’re hiring an ML Engineer to join Kodex’s...  ...into production-grade models, pipelines, and decision...  ...evaluation, batch/streaming inference, backfills, and...  ...impact. Benefits ~ Remote-first within the U.S.... 
    Remote job
    Full time
    Flexible hours

    Kodex

    United States
    22 hours ago
  •  ...Machine Learning Engineer (Llama AI Platform) Location: Remote (Preferred U.S. Time...  ...improve profitability, efficiency, visibility,...  ...source large language models, agentic workflows,...  ...architectures. Optimize inference performance,...  ...-source LLMs. ML Engineering and... 
    Remote work
    Full time

    Performacentric

    Remote
    22 hours ago
  • $213k - $263k

     ...learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust...  ...for perception models. Implement efficient training & inference optimization methods...  ...models. Disclosure for WA Based & Remote Roles: In accordance with... 
    Remote work
    Full time
    Temporary work

    Waymo

    New York, NY
    22 hours ago
  • $170.1k - $258.3k

     ...software that can run efficiently and reliably on real vehicles...  ...new approaches to model export, kernel development, and performance engineering so that every cycle on...  ...of our on‑vehicle ML inference for ADAS and autonomous...  ...United States of America; Remote - Washington; Austin,... 
    Remote work
    Full time
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    1 day ago
  • $128.7k - $261.3k

     ...software that can run efficiently and reliably on...  ...new approaches to model export, kernel...  ..., and performance engineering so that every cycle...  ...into fast, reliable inference across GPUs powering...  ...effortless for ML engineers across the...  ...of America; Remote - Washington; Austin... 
    Remote work
    Full time
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    4 days ago
  • $175k - $280k

    Sesame, located in Bellevue, Washington, is seeking a talented engineer to join our team focused on revolutionizing the way computers interact...  ...with humans. The role involves optimizing machine learning models for a new consumer product category, working with state-of-the-... 

    SESAME

    Bellevue, WA
    2 days ago
  • $100.4k - $180.7k

     ...TitleML and Optimization Engineer.LocationCO - Golden....  ...computing, AI/ML, modeling and simulation, and visualization...  ...modeling, large-scale inference, foundation models,...  ...writing clean, efficient, and maintainable code...  ...SummaryLocation: Golden, CO; Remote SiteType: Full time
    Remote work
    Full time
    Fixed term contract
    Live in
    Local area
    Relocation
    Shift work

    National Renewable Energy Laboratory

    Golden, CO
    3 days ago
  • $175k - $280k

     ...New York is seeking an expert in optimizing machine learning models to turbocharge their serving layer, integrating LLM, speech, and...  ...significant experience in systems programming and performance engineering, aiming to improve high-throughput, low-latency serving. Join... 

    SESAME

    New York, NY
    1 day ago
  •  ...About the Role Reddit is building a dedicated Ads ML Efficiency function to make model training and inference materially faster, cheaper, safer, and more scalable. This person will be a key senior engineer on that team, owning meaningful efficiency work across training... 
    Full time
    Flexible hours

    Reddit

    Remote
    22 days ago
  • $155k - $180k

     ...using computer vision models. Today, over 1M+ developers...  ...(not only product and engineering), so Roboflow employs...  ...center of all of this is inference — one of our most...  ...latest computer vision and ML models to our users....  ...while also supporting remote team members. You can work... 
    Remote work
    Full time
    Second job
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    22 hours ago
  •  ...Senior Machine Learning Engineer We are looking for a Senior ML Engineer to lead our ML efforts for AI Agents...  ...team to build systems that efficiently and deploy the latest ML technologies...  ...infrastructure team to build tuning and inference layers that can scale to handle... 
    Remote work
    Flexible hours

    Lastmile AI

    United States
    1 day ago
  • $100k - $150k

     ...potential. Job Title: Edge ML Engineer Location: 100% Remote (U.S.) Position Type:  Full-time...  ..., and deploy machine learning models that run efficiently on resource-constrained edge devices...  ...with at least one major edge inference framework. Solid understanding... 
    Remote work
    Full time
    Local area
    Immediate start

    Bright Vision Technologies

    Apex, NC
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff ML Inference Engineer Model Efficiency (Remote). Be the first to apply!