Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Machine Learning Engineer - Inference

$160k - $230k
Full-time

Together Ai

About the Role


Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models models and ensuring they run efficiently and effectively at scale. If you are passionate about AI inference, PyTorch, and developing high-performance systems, we want to hear from you. This position offers the chance to collaborate closely with AI researchers and engineers to create cutting-edge AI solutions. Join us in shaping the future at Together AI!

Responsibilities



  • Design and build the production systems that power the Together AI inference engine, enabling reliability and performance at scale.

  • Develop and optimize runtime inference services for large-scale AI applications.

  • Collaborate with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.

  • Conduct design and code reviews to ensure high standards of quality.

  • Create services, tools, and developer documentation to support the inference engine.

  • Implement robust and fault-tolerant systems for data ingestion and processing.

Requirements



  • 3+ years of experience writing high-performance, well-tested, production-quality code.

  • Proficiency with Python and PyTorch.

  • Demonstrated experience in building high performance libraries and tooling.

  • Excellent understanding of low-level operating systems concepts including multi-threading, memory management, networking, storage, performance, and scale.

  • Preferred: Knowledge of existing AI inference systems such as TGI, vLLM, TensorRT-LLM, Optimum

  • Preferred: Knowledge of AI inference techniques such as speculative decoding.

  • Preferred: Knowledge of CUDA/Triton programming.

  • Nice to have: Knowledge of Rust, Cython and compilers.

About Together AI


Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society. Together, we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI. Our team has been behind technological advancements such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey to build the next-generation AI infrastructure.

Compensation


We offer competitive compensation, startup equity, health insurance, and other competitive benefits. The US base salary range for this full-time position is $160,000 - $230,000 + equity + benefits. Our salary ranges are determined by location, level, and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity


Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunities to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.




Please see our privacy policy at 

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Machine Learning Engineer - Inference in San Francisco, CA vacancy
  • $180k - $270k

     ...highest standards of data security and privacy protection. To learn more about Plaud, please visit and follow along on...  ...experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models.... 
    Suggested
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    more than 2 months ago
  • $203.5k - $299.3k

     ...creates a causal question.About the RoleWe are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash...  ...because you have…Deep practical experience with causal inference, econometrics, experimentation, or causal ML.Experience... 
    Suggested
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    3 days ago
  • $155k - $180k

     ...half the Fortune 100, use Roboflow’s machine learning open source and hosted tools. That includes...  ...on all roles (not only product and engineering), so Roboflow employs developers...  ...At the center of all of this is inference — one of our most important open source... 
    Suggested
    Full time
    Second job
    Remote work
    Work from home
    Relocation package
    Flexible hours
    Night shift

    Roboflow

    San Francisco, CA
    more than 2 months ago
  • $150k - $200k

     ...reliable, high-speed robot autonomy software stack optimized for inference performance ● Advance SOTA dexterous manipulation...  ...Required Qualifications ● PhD or MS degree in Computer Science, Machine Learning, Robotics, or equivalent technical discipline ● Deep... 
    Suggested
    Full time

    Deft Ai, Inc.

    San Francisco, CA
    more than 2 months ago
  • $150k - $190k

     ...-driven simulation software stack for engineering and manufacturing across advanced industries...  ..., multi-physics simulation through AI inference across the entire engineering...  ...goals. Who We're Looking For As a Machine Learning Engineer in Delivery, you are a... 
    Suggested
    Remote job
    Full time
    Flexible hours

    Physicsx

    San Francisco, CA
    more than 2 months ago
  • $165k - $230k

     ...like. About the role We're looking for exceptional Machine Learning Engineers focused on Ads to help take Higgsfield's advertising...  ...reliably at significant scale, from experimentation through inference and serving. Work closely with Product, Research, Engineering... 
    Full time
    Work at office
    Remote work
    Worldwide
    3 days per week

    Higgsfield

    San Francisco, CA
    a month ago
  •  ...Francisco, NYC, or London offices. About the Role As a Machine Learning Engineer on the Marketplace team, you will build the models and...  ...not just top-of-funnel engagement • Real-time and batch inference systems embedded in product-critical workflows Example... 
    Full time
    Work at office
    Relocation package

    Mercor

    San Francisco, CA
    more than 2 months ago
  •  ...We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises...  ...our San Francisco office ~ Eager to learn and adapt quickly ~ Prior startup or...  ...active learning pipelines Optimize inference, batching, and quantization on GPU Productionize... 
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    The Pulse

    San Francisco, CA
    a month ago
  •  ...that runs the real economy. Learn more about our vision in our...  ...Collaborate with product and engineering teams to integrate and deploy...  ...Have Strong experience in machine learning, deep learning, and...  ...generative AI, or real-time inference systems. Hands-on experience... 
    Full time
    Worldwide
    Shift work

    HappyRobot

    San Francisco, CA
    a month ago
  •  ...our growing team. About the Role We're looking for a Machine Learning Engineer to design, build, and deploy production-grade ML systems...  ...scalable ML pipelines for training, evaluation, monitoring, and inference Build intelligent services using modern NLP, LLM,... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    a month ago
  • $200k - $400k

     ...interpretability to understand, learn from, and design AI systems....  ...what models learn. Every engineering discipline has been gated by...  ...role We’re looking for Machine Learning Engineers to help build...  ...interpretability, training, and inference. Integrate new machine... 
    Full time
    Work at office
    Remote work

    Заявка На Вакансию «machine Learning Engineer» В Компании «g...

    San Francisco, CA
    more than 2 months ago
  •  ...is to reinvent the way people learn, starting with language....  ...role We’re hiring an ML Engineer, Assessments to help build best...  ...(Content/Learning Design) , Machine Learning, Product, and Engineering...  ...→ model training → inference → feedback generation) Own... 
    Full time
    Live in
    Immediate start

    Speak

    San Francisco, CA
    more than 2 months ago
  • $140k - $200k

     ....About the RoleWe’re looking for an ML Engineer to build the production systems that train...  ..., monitor, retrain, and serve our machine-learning models reliably. You sit between software...  ...stay working.You will build training and inference pipelines, serve predictions through... 
    Temporary work
    Work at office
    Monday to Friday
    Monday to Thursday

    Sprinter Health

    San Francisco, CA
    10 hours ago
  •  ...hope to be a match for you too. About the Role As a Machine Learning Engineer on the AI Platform team, you will develop tailored user experiences...  ...professional experience in developing ETL pipelines and inference services that use large language models (LLMs) and text... 
    Full time

    Workday

    San Francisco, CA
    4 days ago
  •  ...the place for you. The Role We're looking for a Machine Learning Engineer who loves getting close to the metal. This is a hands-on...  ...scheduling, and squeezing performance out of complex training and inference workloads. They should be just as comfortable optimizing... 
    Work at office

    Relace Inc

    San Francisco, CA
    3 days ago
  • $118k - $176k

     ...(*Comscore, Total Visits, March 2025) Day to Day The Machine Learning Engineer I role partners closely with business partners across various...  ...train and optimize models, and maintain and improve model inference services. You will learn and apply new techniques from... 
    Work experience placement
    Local area

    Indeed

    San Francisco, CA
    2 days ago
  •  ...Machine Learning Lead At Nudge, our mission is to develop the best technology for interfacing...  ...computer vision and real-time inference systems that track brain motion and dynamically...  ...Partner closely with mechanical engineers, electrical engineers, ultrasound engineers... 

    Nudge Inc.

    San Francisco, CA
    16 hours ago
  • $150k - $300k

     ...Founding ML Engineer Location: San Francisco, CA Company Stage: Early-Stage (YC-backed...  ...on pushing the boundaries of applied machine learning to power the next generation of AI-native...  ...GPU clusters Experience scaling LLM inference pipelines in production Research... 
    Visa sponsorship

    Recruiting from Scratch

    San Francisco, CA
    2 days ago
  •  ...Machine Learning Engineer San Francisco Who are we? RZR Global is an AI-driven company specializing in mobile advertising solutions designed...  ...including C++ and Rust is a plus. Exposure to online inference systems, gRPC/REST model endpoints, or streaming features... 

    RZR Global Inc.

    San Francisco, CA
    2 days ago
  • $120k - $160k

     ...Machine Learning Engineer Huntington Beach, California, United States; San Francisco, California, United States About Mach Industries...  ...Engineer, you will own and scale the training, data, and edge-inference backbone that every vision and multi-sensor model on our... 
    Permanent employment
    Work at office
    Shift work

    Mach Industries

    San Francisco, CA
    4 days ago
  •  ...connect and drive people forward. We are looking for a Machine Learning Engineer to join the growing AI and Machine Learning team at Strava...  ...to shipping production code to scaling and optimizing inference and deployment Shape AI at Strava : Be a strong voice... 
    Full time
    Work at office
    Worldwide
    Flexible hours
    3 days per week

    Strava

    San Francisco, CA
    a month ago
  •  ...that everyone else has simply learned to live with. We value...  ...You’ll help define how machine learning models run across Cloudflare...  ...accelerators. You’ll work with systems engineers, product teams, hardware...  ...role combines applied ML, inference optimization, evaluation, and... 
    Full time
    Local area

    Cloudflare

    San Francisco, CA
    15 days ago
  • $200k - $260k

     ...Role Together AI is building the best inference infrastructure for voice applications....  .... We're looking for a Senior ML Engineer to drive the model serving layer for voice...  ...strong plus but not required — you can learn this quickly if you have strong ML engineering... 
    Full time

    Together Ai

    San Francisco, CA
    more than 2 months ago
  • $211k - $290.5k

     ...re using the power of tech, data, and machine learning to connect this thriving community of...  ...data scientists and machine learning engineers, and you will help take us there!...  ...ranking, and/or experimentation and causal inference). ~ Strong programming skills.... 
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    1 day ago
  • $151.8k - $265.35k

     ...expanding rapidly into adjacent verticals. We are hiring a Senior Machine Learning Engineer to build the pipelines and services that turn Firefly...  ..., with significant ownership of production ML or inference services at scale. Strong Python and deep-learning engineering... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    2 days ago
  •  ...responsibilities and support their families.Atlassian is seeking a Machine Learning Engineer to join our Growth organization. Growth builds intelligent...  ...modeling, recommender systems, policy evaluation, causal inference, or other approaches for optimizing decisions under... 
    Temporary work
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    10 hours ago
  • $228.96k - $315.36k

     ...Plaid’s Fraud organization builds the machine learning systems that power Plaid’s fraud detection...  ...customers.As a Senior Machine Learning Engineer, you will own the development of high-...  ...feature computation and online inference pipelines that power production machine... 
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    2 days ago
  •  ...customers. Responsibilities What you’ll doAs a Senior Machine Learning Engineer, you’ll tackle the open-ended technical problems that...  ...prompting, fine-tuning, data strategy, system architecture, and inference; balancing quality with latency, reliability, safety, and... 
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    1 day ago
  • $148.7k - $199.4k

    Job Posting Title:Senior Machine Learning Engineer - ESPNReq ID:10150610Job Description:Disney Entertainment & ESPN TechnologyOn any given day...  ...EngineeringBuild and operate ML‑adjacent services such as inference inputs, feature APIs, and data access layers.Contribute to... 
    Full time
    Worldwide

    Hulu

    San Francisco, CA
    2 days ago
  •  ...analysis, and optimization at Haus using optimization, machine learning, and causal inference. We are looking for individuals who not only excel in problem...  ...working with applied scientists, data scientists, data engineers and other MLEs to deliver trustworthy results to our... 
    Full time
    Work at office
    Work from home
    Worldwide
    Flexible hours

    Haus Analytics

    San Francisco, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Machine Learning Engineer - Inference. Be the first to apply!