Staff ML Inference Engineer Model Efficiency (Remote)
Jaide Health
- Remote job
Jaide Health is seeking an engineer for their Model Efficiency team in San Francisco. The role focuses on building reliable ML systems while enhancing core performance metrics across model execution. You'll work with advanced performance techniques such as GPU/CUDA optimizations and collaborate closely with modeling and systems teams. Ideal candidates will have over 5 years of experience in high-performance coding, plus strong skills in C++ or Python and insights into the LLM inference ecosystem. A commitment to diversity and inclusive work culture is celebrated. #J-18808-Ljbffr Jaide Health
$251k - $310k
...machine learning (ML) engineers, software engineers... ...ML algorithms, to model the real world,... ...report to a Senior Staff Engineering Manager... ...bottlenecks in training and inference performance (e.g.,... ...distillation, and efficient attention... ...can be performed remote, the specific salary...Remote workFull time- ...cutting-edge foundation AI models and end-to-end products that... ...is a team of researchers, engineers, designers, and more, who are... ...systems and optimize audio inference serving efficiency using innovative techniques... ...and London. We embrace a remote-friendly environment, and as...Remote workWork at officeLocal areaHome office
$195.2k - $262.2k
...enterprises from data and model training through to... ...large in-house AI/ML infrastructure. Built by engineers, for engineers.... ...orchestration to inference optimization, we own... ...programs in efficient LLM and VLM inference... ...secondary caregivers. Remote work reimbursement:...Remote workTemporary workImmediate start$117.7k - $221.4k
..., practical, and cost efficient for embodied AI systems... ...not only on stronger models, but also on better infrastructure... ...reflects how Cola engineers think: build durable... ..., featurization, and inference foundations that power... ...This role is based remotely, but if the selected...Remote workFull timeLocal areaWork from homeRelocation packageFlexible hours$175k - $215k
...learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust... ...research setting developing recipes for ML models We prefer: - Track record of... ...or, if the role can be performed remote, the specific salary range for your preferred...Remote workFull time$170k - $216k
...sensors, enabling engineers like you to (1) develop methods for efficiently and continuously learning... ..., to (2) develop models and model training... ...and model inference through model architecture... ...~ Experience with ML frameworks like... ...can be performed remote, the specific...Remote workFull time$298k - $368k
...set of sensors, enabling engineers like you to (1) develop methods for efficiently and continuously... ...world data, to (2) develop models and model training at scale... ...low-latency on-device inference techniques and a deep... ...role can be performed remote, the specific salary range...Remote workFull time- ...and deploying frontier models for developers and... ...team of researchers, engineers, designers, and more,... ...on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques... ...London. We embrace a remote‑friendly environment,...Remote workFull timeWork at officeFlexible hours
- Apple Inc. in the San Francisco Bay Area is looking for a Staff/Sr. Machine Learning Engineer to optimize inference for large foundation models and deliver production-grade solutions that serve millions of users in real time. You will collaborate closely with research and...
- ...technology company in Seattle is seeking a Machine Learning Engineer for Model Serving Infrastructure. The ideal candidate will have at least... ...in C/C++/CUDA. You will design and implement distributed inference infrastructure and collaborate with product teams. This position...
- ...Overview: The Principal AI/ML Engineer will support the development... ...learning, and large language models. We offer generous relocation... ...of applications within remote sensing such as tasking collections... ...engineering techniques / Inference time techniques (e.g. chain of...Remote workFull timeTemporary workWork at officeLocal areaVisa sponsorshipRelocation packageFlexible hours
$114.6k - $252.1k
Job Title: Principal AI/ML Engineer (Large Language Model)Job Category: ScienceTime Type: Full timeMinimum... ...to a variety of applications within remote sensing such as tasking collections,... ...factory)Prompt engineering techniques / Inference time techniques (e.g. chain of...Remote workContract workWork experience placementLocal areaFlexible hours- ...include occasional remote work, starting the... ...for a performance engineer who specializes in... ...fast and cost-efficient in the datacenter.... ...-throughput batch inference sweeping petabytes... ...of accelerators, ML frameworks, and large... ...roofline and performance models for our workloads,...Remote workFull timeFor contractorsFor subcontractorCasual workWork at officeDay shift
- ...Mirantis empowers platform engineering teams to deliver composable,... ...technical Product Manager to own AI inference and model serving for k0rdent AI, our... ...technical role owning AI/ML and inference product(s) ~... ...job opportunities. #remote We are a Leader for Container...Remote work
- Bright Vision Technologies is seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance inference platforms for serving large ML models in production. The role emphasizes systems engineering, including request routing, batching, autoscaling...Remote job
$85 per hour
...realistic machine learning engineering tasks and assessing model outputs. You will perform... ...assessments that reflect production ML workflows, helping a... ...involving model training, inference systems, MLOps, and LLM... ...Work Terms Location: Remote. Employment type: Hourly...Remote workHourly pay- ...About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of... ...acceleration, GPU parallelism, advanced model deployment, and video generation... ...significant improvements to model speed and efficiency, ensuring our creative AI systems...Full timeWork at office3 days per week
$200k - $250k
...accuracy, and trust. Our ML systems sit at the... ...a Senior MLOps Engineer to help us run them reliably and efficiently in production.... ...across a custom-built inference platform powering... ..., and extraction models), each with... ...holidays ~ Fully remote work within the United...Remote workFull timeFlexible hours- ...security, transparency, and efficiency. About the Role We’re hiring an ML Engineer to join Kodex’s... ...into production-grade models, pipelines, and decision... ...evaluation, batch/streaming inference, backfills, and... ...impact. Benefits ~ Remote-first within the U.S....Remote jobFull timeFlexible hours
- ...Machine Learning Engineer (Llama AI Platform) Location: Remote (Preferred U.S. Time... ...improve profitability, efficiency, visibility,... ...source large language models, agentic workflows,... ...architectures. Optimize inference performance,... ...-source LLMs. ML Engineering and...Remote workFull time
$213k - $263k
...learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust... ...for perception models. Implement efficient training & inference optimization methods... ...models. Disclosure for WA Based & Remote Roles: In accordance with...Remote workFull timeTemporary work$170.1k - $258.3k
...software that can run efficiently and reliably on real vehicles... ...new approaches to model export, kernel development, and performance engineering so that every cycle on... ...of our on‑vehicle ML inference for ADAS and autonomous... ...United States of America; Remote - Washington; Austin,...Remote workFull timeLocal areaWork from homeRelocation packageFlexible hours$128.7k - $261.3k
...software that can run efficiently and reliably on... ...new approaches to model export, kernel... ..., and performance engineering so that every cycle... ...into fast, reliable inference across GPUs powering... ...effortless for ML engineers across the... ...of America; Remote - Washington; Austin...Remote workFull timeLocal areaWork from homeRelocation packageFlexible hours$175k - $280k
Sesame, located in Bellevue, Washington, is seeking a talented engineer to join our team focused on revolutionizing the way computers interact... ...with humans. The role involves optimizing machine learning models for a new consumer product category, working with state-of-the-...$100.4k - $180.7k
...TitleML and Optimization Engineer.LocationCO - Golden.... ...computing, AI/ML, modeling and simulation, and visualization... ...modeling, large-scale inference, foundation models,... ...writing clean, efficient, and maintainable code... ...SummaryLocation: Golden, CO; Remote SiteType: Full timeRemote workFull timeFixed term contractLive inLocal areaRelocationShift work$175k - $280k
...New York is seeking an expert in optimizing machine learning models to turbocharge their serving layer, integrating LLM, speech, and... ...significant experience in systems programming and performance engineering, aiming to improve high-throughput, low-latency serving. Join...- ...About the Role Reddit is building a dedicated Ads ML Efficiency function to make model training and inference materially faster, cheaper, safer, and more scalable. This person will be a key senior engineer on that team, owning meaningful efficiency work across training...Full timeFlexible hours
$155k - $180k
...using computer vision models. Today, over 1M+ developers... ...(not only product and engineering), so Roboflow employs... ...center of all of this is inference — one of our most... ...latest computer vision and ML models to our users.... ...while also supporting remote team members. You can work...Remote workFull timeSecond jobWork from homeRelocation packageFlexible hoursNight shift- ...Senior Machine Learning Engineer We are looking for a Senior ML Engineer to lead our ML efforts for AI Agents... ...team to build systems that efficiently and deploy the latest ML technologies... ...infrastructure team to build tuning and inference layers that can scale to handle...Remote workFlexible hours
$100k - $150k
...potential. Job Title: Edge ML Engineer Location: 100% Remote (U.S.) Position Type: Full-time... ..., and deploy machine learning models that run efficiently on resource-constrained edge devices... ...with at least one major edge inference framework. Solid understanding...Remote workFull timeLocal areaImmediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff ML Inference Engineer Model Efficiency (Remote). Be the first to apply!
- software engineer staff San Francisco, CA
- assistant engineer San Francisco, CA
- engineering aide San Francisco, CA
- staff engineer San Francisco, CA
- staff security engineer San Francisco, CA
- assistant mechanical engineer San Francisco, CA
- assistant engineering manager San Francisco, CA
- senior staff systems engineer San Francisco, CA
- technology administrator San Francisco, CA
- project engineer assistant project manager San Francisco, CA



