Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer - Inference / Serving

Yobi AI

Machine Learning Engineer - Inference / Serving

Join to apply for the Machine Learning Engineer - Inference / Serving role at Yobi AI

Overview

Yobi is a rapidly growing Behavioral AI company on a mission to ethically democratize the benefits of data and AI. Since 2019, we have built one of the largest consented behavioral datasets in the United States, extending far beyond the walled gardens of Big Tech. Unlike traditional LLM companies, Yobi builds foundation models of human behavior grounded in real‑world actions such as purchases and store visits. Our private‑by‑design modeling enables state‑of‑the‑art personalization and decisioning for leading brands and agencies while protecting privacy, safety, and ethics.

Today, we are focused on bringing the performance of closed‑web user acquisition to the open web and connected TV, giving brands walled‑garden results without the walls. At our core, Yobi is building the behavioral intelligence layer for any system that makes a personalization decision.

Working at Yobi

We’re at an inflection point—customer adoption is accelerating, but there’s still room to shape the architecture and culture from the ground up. Engineers here own major surface areas, build 0→1 systems in large‑scale data and model infrastructure, and help define how Behavioral AI scales ethically and effectively.

Highlights

  • Well‑funded with 5+ years of runway. We are scaling revenue quickly and project to be breakeven in 2026.
  • Partnerships with Microsoft and Databricks.
  • Fully remote or hybrid from hubs in SF Bay Area, Seattle, NYC.
  • World‑class team of Machine Learning experts with experience at Amazon, Uber, Twitter, Meta, etc.
  • Product and Go‑to‑Market teams that have taken ideas from concept to nine‑figure revenue streams.

Benefits

  • Competitive base salary.
  • Meaningful equity and financial upside.
  • Annual bonus target based on personal and company performance.
  • Health, dental, vision plans with low out‑of‑pocket costs.
  • Unlimited PTO.
  • 401(k) with company match.

About the Role

As a Machine Learning Engineer focused on inference and serving at Yobi, you’ll design, optimize, and operate the systems that bring our Behavioral AI models to life in real time. You’ll work at the core of our production environment, turning trained models into performant, reliable, and continuously improving services that power our open‑web and CTV products.

This is an applied ML systems role—equal parts engineering depth, deployment craft, and model intuition. You’ll shape how models are packaged, versioned, rolled out, and observed across environments, ensuring every prediction is fast, accurate, and accountable.

Responsibilities & Expectations

  • Build and scale production ML serving systems—handle versioning, rollouts, rollback strategies, and live experimentation.
  • Ensure low‑latency inference by optimizing model graphs, quantizing, batching, caching, and efficient feature retrieval.
  • Write robust, high‑performance code in Go, Rust, C++, or Java and bridge to Python for model integration and analysis.
  • Treat inference as a living system—monitor drift, track model lineage, and ensure observability from input to outcome.
  • Make serving systems reproducible and portable without over‑engineering—for instance, custom runtime design, model registries, or lightweight orchestration.
  • Reason about model performance and trade‑offs, and work with researchers to deploy more practical models.

Qualifications

  • Deep expertise in model deployment and production ML serving.
  • Strong low‑latency mindset and knowledge of inference optimization techniques.
  • Systems fluency: comfortable writing high‑performance code and bridging to Python.
  • Operational maturity: experienced with monitoring, drift detection, and observability.
  • Infrastructure intuition: understanding of custom runtimes, registries, and orchestration.
  • Applied ML understanding: can interpret performance, reasoning about trade‑offs, and collaborate with researchers.

Seniority Level

Mid‑Senior level.

Employment Type

Full‑time.

Job Function

Engineering and Information Technology. Software Development industry.

#J-18808-Ljbffr
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer - Inference / Serving in New York, NY vacancy
  • $209k - $313k

     ..., live in the moment, learn about the world, and have...  ...digital services.Snap Engineering teams build fun and...  ...forefront.We’re looking for a Machine Learning Engineer to...  ...of causal inference and modern approaches...  ...reinforce our values, and serve our community, customers... 
    Suggested
    Full time
    Live in
    Work at office
    Local area

    Snap

    New York, NY
    1 day ago
  • $229k - $343k

     ...themselves, live in the moment, learn about the world, and have...  ...on-device and server-side inference. Our team creates intuitive...  ....We’re looking for a Machine Learning Engineer to join Snap Inc!What you’ll...  ...technology and products that serve millions of SnapchattersBuild... 
    Suggested
    Full time
    Live in
    Work at office
    Local area
    Worldwide

    Snap

    New York, NY
    1 day ago
  •  ...global trading firm is building out its machine learning function and hiring engineers to develop the distributed training and inference systems that take models from research into...  ...— multi-GPU/multi-node training, model serving A track record of hands-on building (this... 
    Suggested
    Flexible hours

    Jobzhr

    New York, NY
    3 days ago
  • $180k - $230k

     ...the core of the business, and it runs on machine learning at serious scale. We're hiring ML Engineer to build and own the infrastructure...  ...of millions of rows of claims data, to serving real-time quotes in seconds against inference-time datasets that run into the trillions... 
    Suggested

    Arlo Corporation

    New York, NY
    1 day ago
  • $184.05k - $262.93k

     .... By combining cutting-edge machine learning, recommendation systems, and...  ...As a Senior Machine Learning Engineer, you will help shape the...  ...personalized recommendations that serve millions of Spotify...  ...have worked with large-scale inference systems and understand the challenges... 
    Suggested
    Remote job
    Full time
    Flexible hours

    Spotify

    New York, NY
    1 day ago
  • $162k - $210k

     ...people go on great dates! We are hiring Machine Learning Engineers to help us build the foundations of...  ..., both offline and real-time inference endpoints that directly impact the experience...  ...them, and make improvements to our serving infrastructure. What We're Looking... 
    Full time
    Work experience placement

    Match Group

    New York, NY
    1 day ago
  • $190k - $260k

    *Machine Learning Engineer – Search, Ranking & Personalization* *Stage:* Seed *Founded:* 2022...  ...and personalization across a platform serving hundreds of millions of items daily....  ...-scale data processing for real-time inference. - Strong backend integration experience... 
    Full time
    H1b
    Remote work
    Relocation
    Visa sponsorship

    Fuku

    New York, NY
    1 day ago
  • $147.6k - $274k

     ...Prescient Design), we are building the machine learning platforms that enable researchers and engineers to move models from...  ...role will contribute to our model-serving platform, and to the broader infrastructure...  ...including real-time and batch inference workloads, GPU-backed services,... 
    Full time
    Local area
    Immediate start
    Worldwide
    Relocation package

    Genentech

    New York, NY
    1 day ago
  • $209k - $313k

     ...themselves, live in the moment, learn about the world, and have...  ...on-device and server-side inference. Our team creates intuitive...  ....We're looking for a Machine Learning Engineer to join our Generative ML team...  ...technology and products that serve millions of... 
    Full time
    Live in
    Work at office
    Local area
    Worldwide

    Snap

    New York, NY
    3 days ago
  •  ...Our TeamThe Decisioning & Optimization engineering team owns the systems that determine...  ...areas:ML infrastructure for model serving: real-time inference at 1M+ QPS, multi-model parallel evaluation...  ...is a unique culture and environment. Learn more here.Inclusion is a Netflix... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours
    Shift work

    Netflix

    New York, NY
    5 hours ago
  •  ...data (the DT User Card).As a Senior Machine Learning Engineer, you'll design and scale the ML systems...  .... You'll build the pipelines and serving infrastructure that turn our data advantage...  ...features.Optimize training and inference for compute efficiency (CPU/GPU utilization... 
    Full time
    Local area
    Shift work

    Fyber

    New York, NY
    2 days ago
  • $244k - $293k

     ...the Role: We are hiring a Staff Machine Learning Engineer to drive the design, development,...  ...facing products. Experience with causal inference, uplift modeling, and interventional...  ..., AWS, or Azure. Familiarity with ML serving solutions like Ray, KubeFlow, or Weights... 
    Full time
    Work experience placement
    Remote work

    Match Group

    New York, NY
    1 day ago
  • $240k - $249.5k

     ...OpportunityGrubhub is looking for a Senior Staff Machine Learning Engineer to help lead the machine learning...  ...for cold start, and real-time inference. Assess rigorously what actually...  ...TensorFlow or PyTorch, including training, serving, and tuning runtime models on... 
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    New York, NY
    4 days ago
  • $158.1k - $213.8k

    The AppStar Data Analytics & Engineering (DNA) team within Amazon's...  ...organization is seeking a Machine Learning Engineer II to build, deploy...  ..., model training, scoring, serving) using AWS services such as...  ...for model training and inference, leveraging Spark/PySpark and... 
    Internship
    Flexible hours

    Amazon

    New York, NY
    5 hours ago
  • $186k - $245k

     ...the Role   Join Hinge as a Senior Machine Learning Engineer, where you'll lead the application of...  ...Kubernetes) to preprocess data, run inference, and manage post-processing pipelines...  ...store, model training environment, model serving environment, observability, workflow... 
    Full time
    Work experience placement

    Match Group

    New York, NY
    1 day ago
  •  ...We are looking for an engineer with robust experience in machine learning and strong mathematical foundations to join...  ...ever-evolving trading environment serves as a unique, rapid-feedback platform...  ...and maintaining training and inference infrastructure, with an understanding... 

    Jane Street

    New York, NY
    2 days ago
  • $160k - $235k

     ...Senior Machine Learning Engineer Affinity stitches together billions of data points from massive datasets to create a...  ...clustering and decision trees ~ Experience with serving ML models for streaming and batch inference at scale. ~ Experience with vector or graph... 
    Work at office
    Remote work
    Worldwide
    Flexible hours
    2 days per week

    Affinity Inc

    New York, NY
    3 days ago
  •  ...Sr Machine Learning Engineer Technology is at the heart of Disney's past, present, and future. Disney...  ...and consumer media touch points serving millions of people around the world....  ...infrastructure for scalable learning, inference, and monitoring, conduct in-depth data... 
    Work experience placement
    Local area
    Day shift

    Disney Entertainment and ESPN Product & Technology

    New York, NY
    4 days ago
  •  ...The Role We're looking for a Machine Learning Engineer to join our Engineering team. You'll...  ...build data pipelines for training and inference. Develop a robust set of tools for...  ...PyTorch or TensorFlow. ~ Experience with serving models for inference (FastAPI) ~... 

    Soris

    New York, NY
    5 days ago
  •  ...growth and superior returns, as we deliver rare value and impact across our businesses. The Role As a Senior ML Engineer for AWS and Real-Time Inference, you'll own the fast path: ingesting live trading data and scoring it in near real time. It's a systems-heavy role... 
    Full time

    TWG Global AI

    New York, NY
    14 days ago
  • $165k - $225k

     ...Career Renew is recruiting for one of its clients a Senior Machine Learning Engineer - this is a fully remote role for US/Canada based...  ...including CUDA kernel engineering, TensorRT/ONNX export, and inference serving frameworks such as Triton • Experience with hosting... 
    Remote work
    Worldwide

    Career Renew

    New York, NY
    2 days ago
  • $200k - $250k

     ...Description Principal Machine Learning Engineer Full-time New York City, NY, US Exclusive...  ...training, evaluation, and production serving. You will partner closely with...  ...model serving in production: low-latency inference, batching, optimization, autoscaling,... 
    Full time

    NextDeavor

    New York, NY
    1 day ago
  • $200k - $250k

     ...Principal ML Engineer New York, NY (Hybrid) About the Company We're building...  ...infrastructure that proves they work, and the serving stack that runs them at scale. This...  ...serving in production: low-latency inference, batching, optimization, autoscaling, and... 
    Work at office

    Blaze Talent

    New York, NY
    2 days ago
  • We are looking for an engineer with experience in low-level systems...  ...join our growing ML team. Machine learning is a critical pillar of Jane...  ...evolving trading environment serves as a unique, rapid-feedback...  ...models - both training and inference. We care about efficient large... 
    Work at office

    Trading Interview

    New York, NY
    2 days ago
  • Machine Learning Engineer — AI Investment Research Lab Location : New York, NY (on-site) Level : Mid...  ...end — from raw data through training, serving, and monitoring Build and maintain...  ...online experiments (A/B testing, causal inference) to validate model improvements... 
    Full time

    Riviera Partners

    New York, NY
    2 days ago
  • $200k - $300k

     ...Machine Learning Research Engineer Tower Research Capital is a leading quantitative trading firm founded...  ...& Infrastructure Benchmarking: Serve as the primary feedback loop for the...  ...comprehensively test both the training and inference environments. Validate the... 
    Casual work
    Work at office

    Tower Research Capital LLC

    New York, NY
    5 days ago
  • $175k - $280k

    Sesame in New York is seeking an expert in optimizing machine learning models to turbocharge their serving layer, integrating LLM, speech, and vision models....  ...experience in systems programming and performance engineering, aiming to improve high-throughput, low-latency... 

    SESAME

    New York, NY
    2 days ago
  •  ...Whatnot updates on our news and engineering blogs and join us as we...  ...core infrastructure that powers machine learning and self-hosted large language...  ...low-latency, large model serving to distributed training & high-throughput GPU inference. What you'll do: Own the infrastructure... 
    Work experience placement
    Work at office
    Local area
    Remote work
    Work from home
    Home office
    Flexible hours

    Whatnot

    New York, NY
    2 days ago
  •  ...and endless opportunity to serve the varied needs of our community...  ...and fulfillment. We use machine learning and Internet-scale data to...  ..., and general causal inference. Search & Discovery ML :...  ...works alongside world-class engineers, data scientists, and product... 
    Remote job
    Permanent employment
    Work experience placement
    Internship
    Work at office
    Work from home
    Flexible hours

    Instacart

    New York, NY
    2 days ago
  • $95k - $170k

     ...science and digital technologies to better serve their most vulnerable members and...  ...two-way street - we’ll invest in your learning and growth, just as you’ll advance the...  ...We are seeking a skilled and motivated Machine Learning Engineer to join our Software Engineering team.... 
    Full time

    N1 Health

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer - Inference / Serving. Be the first to apply!