Machine Learning Engineer - Inference
$160k - $230kTogether Ai
About the Role
Together AI is seeking a Machine Learning Engineer to join our Inference Engine team, focusing on optimizing and enhancing the performance of our AI inference systems. This role involves working with state-of-the-art large language models models and ensuring they run efficiently and effectively at scale. If you are passionate about AI inference, PyTorch, and developing high-performance systems, we want to hear from you. This position offers the chance to collaborate closely with AI researchers and engineers to create cutting-edge AI solutions. Join us in shaping the future at Together AI!
Responsibilities
- Design and build the production systems that power the Together AI inference engine, enabling reliability and performance at scale.
- Develop and optimize runtime inference services for large-scale AI applications.
- Collaborate with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.
- Conduct design and code reviews to ensure high standards of quality.
- Create services, tools, and developer documentation to support the inference engine.
- Implement robust and fault-tolerant systems for data ingestion and processing.
Requirements
- 3+ years of experience writing high-performance, well-tested, production-quality code.
- Proficiency with Python and PyTorch.
- Demonstrated experience in building high performance libraries and tooling.
- Excellent understanding of low-level operating systems concepts including multi-threading, memory management, networking, storage, performance, and scale.
- Preferred: Knowledge of existing AI inference systems such as TGI, vLLM, TensorRT-LLM, Optimum
- Preferred: Knowledge of AI inference techniques such as speculative decoding.
- Preferred: Knowledge of CUDA/Triton programming.
- Nice to have: Knowledge of Rust, Cython and compilers.
About Together AI
Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society. Together, we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI. Our team has been behind technological advancements such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey to build the next-generation AI infrastructure.
Compensation
We offer competitive compensation, startup equity, health insurance, and other competitive benefits. The US base salary range for this full-time position is $160,000 - $230,000 + equity + benefits. Our salary ranges are determined by location, level, and role. Individual compensation will be determined by experience, skills, and job-related knowledge.
Equal Opportunity
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunities to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Please see our privacy policy at
$180k - $270k
...highest standards of data security and privacy protection. To learn more about Plaud, please visit and follow along on... ...experience building and deploying high-throughput, ultra-low-latency inference engines for large language models or foundational speech models....SuggestedFull timeWork at officeWorldwide$203.5k - $299.3k
...creates a causal question.About the RoleWe are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash... ...because you have…Deep practical experience with causal inference, econometrics, experimentation, or causal ML.Experience...SuggestedHourly payWork at officeLocal areaRemote workFlexible hours$155k - $180k
...half the Fortune 100, use Roboflow’s machine learning open source and hosted tools. That includes... ...on all roles (not only product and engineering), so Roboflow employs developers... ...At the center of all of this is inference — one of our most important open source...SuggestedFull timeSecond jobRemote workWork from homeRelocation packageFlexible hoursNight shift$203.5k - $299.3k
...a causal question. About the Role We are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash... ...you because you have… Deep practical experience with causal inference, econometrics, experimentation, or causal ML . Experience...SuggestedHourly payWork at officeLocal areaRemote workFlexible hours- ...a causal question. About the Role We are hiring a Causal Machine Learning Engineer to help build the causal ML foundation behind how DoorDash... ...because you have… Deep practical experience with causal inference, econometrics, experimentation, or causal ML . Experience...SuggestedHourly payWork at officeLocal areaFlexible hours
$150k - $200k
...instructional materials with dynamic digital learning. Through unparalleled curriculum... ...other departments, including Product, Engineering, Machine Learning and Analytics, to understand... ...execution of code, especially for model inference and lightweight processing tasks. ~...Permanent employmentFull timeWork at officeLocal areaRemote work$150k - $200k
...reliable, high-speed robot autonomy software stack optimized for inference performance ● Advance SOTA dexterous manipulation... ...Required Qualifications ● PhD or MS degree in Computer Science, Machine Learning, Robotics, or equivalent technical discipline ● Deep...Full time- ...connect and drive people forward. We are looking for a Machine Learning Engineer to join the growing AI and Machine Learning team at Strava... ...to shipping production code to scaling and optimizing inference and deployment Shape AI at Strava : Bring your voice and...Full timeWork at officeWorldwideFlexible hours3 days per week
$165k - $230k
...like. About the role We're looking for exceptional Machine Learning Engineers focused on Ads to help take Higgsfield's advertising... ...reliably at significant scale, from experimentation through inference and serving. Work closely with Product, Research, Engineering...Full timeWork at officeRemote workWorldwide3 days per week$150k - $190k
...-driven simulation software stack for engineering and manufacturing across advanced industries... ..., multi-physics simulation through AI inference across the entire engineering... ...goals. Who We're Looking For As a Machine Learning Engineer in Delivery, you are a...Remote jobFull timeFlexible hours- ...Francisco, NYC, or London offices. About the Role As a Machine Learning Engineer on the Marketplace team, you will build the models and... ...not just top-of-funnel engagement • Real-time and batch inference systems embedded in product-critical workflows Example...Full timeWork at officeRelocation package
- ...is to reinvent the way people learn, starting with language.... ...role We’re hiring an ML Engineer, Assessments to help build best... ...(Content/Learning Design) , Machine Learning, Product, and Engineering... ...→ model training → inference → feedback generation) Own...Full timeLive inImmediate start
- ...our growing team. About the Role We're looking for a Machine Learning Engineer to design, build, and deploy production-grade ML systems... ...scalable ML pipelines for training, evaluation, monitoring, and inference Build intelligent services using modern NLP, LLM,...Full timeWork at officeRemote workFlexible hours2 days per week
- ...that runs the real economy. Learn more about our vision in our... ...Collaborate with product and engineering teams to integrate and deploy... ...Have Strong experience in machine learning, deep learning, and... ...generative AI, or real-time inference systems. Hands-on experience...Full timeWorldwideShift work
- ...We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises... ...our San Francisco office ~ Eager to learn and adapt quickly ~ Prior startup or... ...active learning pipelines Optimize inference, batching, and quantization on GPU Productionize...Full timeWork at officeVisa sponsorshipRelocation package
$140k - $200k
....About the RoleWe’re looking for an ML Engineer to build the production systems that train... ..., monitor, retrain, and serve our machine-learning models reliably. You sit between software... ...stay working.You will build training and inference pipelines, serve predictions through...Temporary workWork at officeMonday to FridayMonday to Thursday$200k - $260k
...Role Together AI is building the best inference infrastructure for voice applications.... .... We're looking for a Senior ML Engineer to drive the model serving layer for voice... ...strong plus but not required — you can learn this quickly if you have strong ML engineering...Full time$200k - $265k
...us in building the future! About the Role As a Senior Machine Learning Engineer on the AI Image Generation (Imagine) team, you’ll design,... ...quality, character consistency, responsiveness to prompting, inference time, and incorporation of an ever‑increasing number of...Work at office- ...Harrison Clarke is seeking a technically talented engineer to join a well-funded AI startup in San Francisco, building cutting-edge... ...intersection of data pipelines, model training, and production inference. You will own the full ML lifecycle—from data generation and training...
$350k
...Machine Learning Infrastructure Engineer, Safeguards ResearchSan Francisco, CA | New York City, NYAbout AnthropicAnthropic's mission is to create reliable... ...the throughput, cost, and reliability of large-scale inference and scoring workloadsPartner closely with researchers...Work at officeVisa sponsorshipFlexible hoursShift work$200k
...All roles Machine Learning Engineer Scale the core ML platform behind our product, owning pipelines that ingest data, train models, and generate... ...to end, from data ingestion through training pipelines to inference and monitoring in production. It's a high-ownership role...Work at office- ...stage AI startup building cutting-edge machine learning systems at the intersection of large... ...to join a highly technical team where engineers work across the full machine learning... ...Profiling and optimizing model training and inference performance Deploying and maintaining...
- ...About the role A San Francisco startup is seeking a Senior Machine Learning Engineer to build and scale AI systems. You’ll work on impactful... ...Deploying models on cloud platforms (AWS, GCP) with scalable inference Why This Role is Exciting Lead ML strategy and...
- ...Member of Technical Staff, Machine Learning This is a unique opportunity to join an AI research... ...pipelines, including data preprocessing, inference, post‑processing, and quality... ...What Is Expected You are a strong Python engineer with hands‑on experience building production...Immediate start
- ...Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more cost-efficient, and more reliable. You will extend orchestration frameworks (Kubernetes...
- ...what gets built. About the Role As a ML engineer at Wispr, you’ll play a crucial role in... ...startup experience Experience optimizing ML inference or engineering systems for research... ...development Attention to detail and eagerness to learn Aptitude and clarity of thought...H1bWork at officeRemote workRelocationVisa sponsorshipFlexible hours
$230k - $322k
...with people who are likely to find them useful. As a Staff Machine Learning Engineer on Shopping Ads, you will lead the technical strategy and... ...feature engineering, training and evaluation pipelines, online inference, and experimentation. Record of delivering complex results...For contractorsWork experience placementFlexible hoursShift work- ...Seeking Founding Data Scientists and Machine Learning EngineersImagine multiplying your impact. You've unlocked major wins in your career... ...team is building core systems in behavioral modeling, causal inference, forecasting, agentic platforms, and more. You'll help...
- ...ML Infrastructure Engineer, Model InferenceAs an ML Infrastructure Engineer, Model Inference at Abridge, you'll play a pivotal role in building and optimizing the... ...core inference infrastructure that powers our machine learning models. Your work will be instrumental in...Hourly payFull timeFlexible hours
$264.8k - $331k
...Machine Learning Systems Research Engineer, Agent Post-training - Enterprise GenAIAI is becoming vitally important in every function of our society. At... ...You will:Build, profile and optimize our training and inference framework.Post-train state of the art models, developed...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Machine Learning Engineer - Inference. Be the first to apply!
- ai ml engineer San Francisco, CA
- graduate machine learning engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- senior ml engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- data scientist machine learning engineer San Francisco, CA
- machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- artificial intelligence - machine learning intern San Francisco, CA


