Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer — Inference Optimization [Remote]

Full-time

jobgether

United States
  • Remote job

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Machine Learning Engineer — Inference Optimization based in Australia.

This role offers the opportunity to optimize the performance of advanced machine learning systems used in real-world production environments.
You will work at the intersection of research and engineering, transforming cutting-edge models into fast, reliable, and cost-efficient solutions.
Your work will directly impact model scalability, user experience, and the efficiency of AI-powered products.
You will dive deep into performance optimization, from model architecture and GPU execution to large-scale inference infrastructure.
Working with talented research, infrastructure, and product teams, you will help push the boundaries of what AI systems can achieve.
This position is ideal for an engineer who enjoys solving complex technical challenges and building high-performance ML systems from the ground up.

Accountabilities

As a Machine Learning Engineer specializing in inference optimization, you will own the performance and scalability of machine learning models in production. You will combine deep ML expertise, systems engineering, and performance analysis to deliver faster, more efficient AI experiences.

  • Optimize machine learning inference systems to improve latency, throughput, scalability, and operational cost.
  • Profile and identify bottlenecks across GPU and CPU inference pipelines, including memory usage, kernels, batching strategies, and data flow.
  • Implement advanced optimization techniques such as quantization, KV-cache optimization, speculative decoding, batching, streaming, and model simplification.
  • Collaborate with research engineers to productionize new model architectures and translate experimental results into reliable systems.
  • Build, improve, and maintain inference-serving infrastructure using modern frameworks, custom runtimes, or specialized serving solutions.
  • Benchmark model performance across different hardware environments, including GPUs, CPUs, and cloud-based systems.
  • Improve system reliability, monitoring, observability, and cost efficiency under real production workloads.
  • Contribute to engineering practices that improve the quality, scalability, and maintainability of ML infrastructure.

Requirements:

The ideal candidate is a technically strong machine learning engineer with experience optimizing production inference systems and a passion for high-performance AI engineering. You should enjoy working on complex technical problems, experimenting with new approaches, and taking ownership of critical systems.

  • Strong professional experience in ML inference optimization, high-performance machine learning systems, or related areas.
  • Deep understanding of machine learning fundamentals, including neural network architectures, attention mechanisms, memory optimization, and compute graphs.
  • Hands-on experience with PyTorch or similar deep learning frameworks and deploying models into production environments.
  • Experience with GPU performance optimization, including technologies such as CUDA, ROCm, Triton, or kernel-level tuning.
  • Proven experience scaling inference systems for real users beyond research prototypes or benchmarks.
  • Strong programming skills and the ability to work across machine learning and systems engineering domains.
  • Ability to operate effectively in fast-paced environments with ownership, autonomy, and evolving priorities.
  • Experience with inference frameworks such as TensorRT, ONNX Runtime, vLLM, or Triton is a plus.
  • Familiarity with large language models, long-context inference, distributed systems, low-latency services, or hardware optimization is considered an advantage.
  • Contributions to open-source ML systems or inference tooling are a plus.

Benefits:

  • Competitive compensation package with meaningful equity participation.
  • Opportunity to work on performance-critical AI systems with direct product impact.
  • High level of ownership over infrastructure that shapes scalability and efficiency.
  • Close collaboration with research, infrastructure, and product teams.
  • Opportunity to work on advanced machine learning technologies and real-world AI applications.
  • Engineering-focused culture that values technical excellence, experimentation, and quality.
  • Flexible remote work environment.
  • Opportunity to contribute to the growth of an innovative AI-focused organization.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 3 hours ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer — Inference Optimization [Remote] in United States vacancy
  •  ...Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika'...  ...at scale.   You will design and optimize inference pipelines, implement state-of...  ..., attention acceleration, and deep learning compiler stacks. GPU & Parallelism... 
    Suggested
    Full time
    Work at office
    3 days per week

    Pika

    Remote
    7 hours ago
  • We are looking for a Machine Learning Runtime Optimization Engineer to work on an innovative project that redefines software. In this role, you will focus...  ...ML runtimes, including hardware acceleration, and inference speed improvements. Experience with ML inference engines... 
    Suggested
    Remote job
    Full time
    Worldwide

    Interop Labs

    Remote
    7 hours ago
  •  ...explore, create, play, learn, and connect with friends...  ...a billion people with optimism and civility, and...  ...experiences for everyone. Our engine’s resource management...  ...the application of machine learning in real-time engine...  ...Design ML models that infer player and interaction... 
    Suggested
    Full time

    Roblox

    Remote
    7 hours ago
  •  ...Community You Will Join:  Machine Learning and Artificial Intelligence...  ...fine-tuning, alignment and optimization, RAG/Search, LLM evaluation...  ...principal machine learning engineer, you will be responsible...  ...tuning, optimizing models and inference run-time ~ Post-training... 
    Suggested
    Remote job
    Full time
    Casual work
    Live in
    Work at office

    Airbnb, Inc.

    United States
    7 hours ago
  • $170k - $216k

     ...builds the system which learns the spatial-temporal...  ...downstream teams on the optimization and integration into...  ...of sensors, enabling engineers like you to (1) develop...  ...model training and model inference through model...  ...+ years experience in Machine Learning and/or Computer... 
    Suggested
    Full time
    Remote work

    Waymo

    Remote
    7 hours ago
  • $298k - $368k

     ...builds the system which learns the spatial-temporal...  ...downstream teams on the optimization and integration into...  ...of sensors, enabling engineers like you to (1) develop...  ...years of experience in Machine Learning, with a focus...  ...low-latency on-device inference techniques and a deep... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    7 hours ago
  • $213k - $263k

     ...support and automate the lifecycle of the machine learning workflow, including feature and experiment management, model development, optimization and monitoring. These efforts have...  ...and Simulation. We are looking for engineers with ML software or ML systems expertise... 
    Full time
    Remote work

    Waymo

    Remote
    7 hours ago
  • $137.1k - $201.6k

     ...The mission of the Marketplace Optimization team is to ensure we maintain a healthy...  ...artificial intelligence and advanced ML, deep learning techniques to power decision-making...  ...the Role We’re looking for a Machine Learning Engineer to help design, build, optimize and scale... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash Usa

    San Francisco, CA
    7 hours ago
  • $216.7k - $303.4k

    Role Description We are hiring Machine Learning Engineers (IC4) to build and evolve the auction, bidding and budgeting systems that power Reddit Ads. In this role, you will: ~Design and implement optimization algorithms for auctions, bidding strategies, and pacing that... 
    Full time
    Work experience placement
    Flexible hours

    Reddit

    Remote
    6 days ago
  • $180k - $270k

     ...highest standards of data security and privacy protection. To learn more about Plaud, please visit and follow along on...  ...experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models.... 
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    7 hours ago
  • $185.8k - $303.4k

     ...Description This role sits in the Ads Optimization and Ads Marketplace Quality (AMQ)...  .... You’ll join a set of tight-knit engineers working on high-impact, internet-scale...  ...Role Description We are hiring Machine Learning Engineers (IC3 and IC4) to build and... 
    Full time
    For contractors
    Work experience placement
    Flexible hours

    Reddit

    United States
    7 hours ago
  • $198k - $286k

    Role Description At Modular, we optimize inference from kernel to cloud on one unified stack. We...  ...one, then keeps getting better. As we learn the shape and patterns of each customer...  ...optimizations across kernels, the inference engine, and distributed systems so that... 
    Full time
    Local area
    Flexible hours

    Modular

    Remote
    2 days ago
  • $141k - $249k

     ...adopted into Waabi’s training and inference frameworks. Examples include designing...  .... ~Work with researchers and ML engineers on best-practices for optimal resource usage. ~Create and...  ...C++ or Rust. ~Experience in deep learning frameworks such as PyTorch or Jax.... 
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    Remote
    11 hours ago
  • $170.5k - $315.49k

    ## Inference Optimization Engineer (local / edge runtime)Applylocations: US, California, Santa Clara: US,...  ...efficient models run directly on the user's machine (AI PC, edge, on-prem, and beyond),...  ...where it helps us# What you’ll learn / grow into*Curiosity is required. You... 
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    3 days ago
  • $170.5k - $315.49k

     ...efficient models run directly on the user’s machine, keeping data private and token costs...  ...the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained...  ...engines where it helps us What you’ll learn / grow into Understanding the internals... 
    Local area

    Intel

    Phoenix, AZ
    3 days ago
  • $155k - $180k

     ...half the Fortune 100, use Roboflow’s machine learning open source and hosted tools. That includes...  ...on all roles (not only product and engineering), so Roboflow employs developers...  ...At the center of all of this is inference — one of our most important open source... 
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    7 hours ago
  • Role Description Inference is growing fast — and so is the volume of contributions, increasingly authored with the help of AI agents. That...  ...the (genuinely fun) work of bringing new models into the engine. Qualifications ~5+ years of hands-on experience building and... 
    Full time
    Remote work
    Flexible hours
    Night shift

    Roboflow

    Remote
    6 days ago
  •  ...Job Description: As a Machine Learning Engineer, you will play a pivotal role in driving the development...  ...to develop success criteria and optimize new products, features, policies, and...  ...like containers, batch vs real time inference endpoints, application security... 
    Full time
    Work experience placement

    5 Star Recruitment

    Remote
    7 hours ago
  • $145k - $165k

     ...Profile, Chat, Growth, and Revenue optimization. Our mission is to apply machine learning to enhance user experiences,...  ...distinct roles: Machine Learning Engineers (this role) who focus on...  ...recommendation systems or casual inference Familiarity with big data or... 
    Full time
    Work experience placement
    Casual work
    Work at office

    Match Group

    Remote
    7 hours ago
  •  ...We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises...  ...our San Francisco office ~ Eager to learn and adapt quickly ~ Prior startup or...  ..., and active learning pipelines Optimize inference, batching, and quantization on GPU... 
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    Pulse

    San Francisco, CA
    7 hours ago
  •  ...About the Role Pangram Labs is hiring strong Machine Learning Engineers at all levels to join our team. In this role, you will...  ...infrastructure for multi-GPU LLM training Profiling and optimizing training and inference code Deploy efficient inference pipelines for... 
    Full time
    Work at office

    Pangram Labs

    Remote
    7 hours ago
  • $151k - $257k

     ...Do you enjoy applying machine learning to complex, real-world problems in autonomous...  ...are looking for a hands-on ML Engineer to integrate, implement, and optimize our next-generation AV scenario...  ...optimization techniques for low-latency inference systems Familiarity with... 
    Full time
    Temporary work
    Immediate start
    Relocation package

    Zoox

    Remote
    7 hours ago
  • $100k

     ...positions within our Algorithms Engineering group. Based on your...  ...a passionate and talented Machine Learning Engineer to join our Algorithms...  ..., real-time systems. Optimize the performance and scalability...  ...processing, or causal inference Contributions to open-source... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    Remote
    7 hours ago
  • $150k - $180k

     ...being part of the solution. As a Machine Learning at BetterHelp, you’ll join a diverse team of licensed clinicians, engineers, product pros, creatives, marketers, and...  ...quality, implementing guardrails, and optimizing inference systems for real-world applications.... 
    Full time
    Work experience placement
    Work at office
    Immediate start
    Remote work

    Betterhelp

    United States
    7 hours ago
  • $150k - $200k

     ...materials with dynamic digital learning. Through unparalleled...  ...departments, including Product, Engineering, Machine Learning and Analytics, to...  ...skills, with a knack for optimizing performance and ensuring data...  ...code, especially for model inference and lightweight processing... 
    Permanent employment
    Full time
    Work at office
    Local area
    Remote work

    Kiddom

    San Francisco, CA
    7 hours ago
  •  ...are seeing. We are looking for a Machine Learning Engineer to build creative, practical, and...  ...for model training, evaluation, and inference, both in the cloud and on edge devices...  ...intelligent active sampling infrastructure to optimize data collection and improve model... 
    Full time
    Work at office
    Flexible hours
    Weekend work

    Orchard Robotics

    San Francisco, CA
    7 hours ago
  •  ...Python, and applying software engineering and design principles (OOP,...  ...configuration, Spark optimization techniques and best practices...  ...operationalizing and deploying machine learning models using production-...  ...vector stores, low-latency inference pipelines. Benefits... 
    Full time
    Local area

    Tiger Analytics

    United States
    7 hours ago
  • $150k - $200k

     ...Build reliable, high-speed robot autonomy software stack optimized for inference performance ● Advance SOTA dexterous manipulation architecture...  ...Qualifications ● PhD or MS degree in Computer Science, Machine Learning, Robotics, or equivalent technical discipline ● Deep... 
    Full time

    Deft Ai, Inc.

    San Francisco, CA
    7 hours ago
  • $100k - $300k

     ...scale through data-driven machine learning is the key to unlocking these...  ...for a Machine Learning Engineer to be responsible for designing...  ...experiments, and optimizing these models to perform efficiently...  ...Communicate effectively with inference, application, and... 
    Full time

    Skild AI

    San Mateo, CA
    7 hours ago
  • $200k - $400k

     ...interpretability researchers and engineers from organizations like...  ...role We’re looking for Machine Learning Engineers to help build our...  ...production ready tools. Optimize pipelines and infrastructure...  ...interpretability, training, and inference. Integrate new machine... 
    Full time

    Goodfire

    San Francisco, CA
    7 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer — Inference Optimization [Remote]. Be the first to apply!