ML Systems Engineer — Production-Scale LLM Inference
$150k - $350kChipAgents
ChipAgents is looking for a skilled ML Systems Engineer in San Jose, California. In this technical role, you will optimize large language model inference for our agentic AI platform, impacting chip design efficiency. Your responsibilities include implementing performance optimizations, designing multi-node clusters, and collaborating with top semiconductor firms. With a competitive salary range of $150K–$350K, we also offer substantial benefits and equity options. #J-18808-Ljbffr ChipAgents
$150k - $350k
...generative AI to assist engineers in RTL design,... ...is deployed in production to companies that... ...We are seeking an ML Systems Engineer to optimize... ...large language model inference powering our... ...push the limits of LLM throughput and latency... ...with large‑scale ML systems, GPU computing...Suggested$174.72k - $295.68k
...strong foundation for LLM deployment and quality... ...PTQ, QAT, on-vehicle inference and related fields.Key... ...bit techniques.Develop production-quality Python code with... ...and software engineering skills.Ability to work... ...effectively across research, systems, infrastructure, and product...SuggestedFull time$184.7k - $324.8k
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS &... ...foundation models.Our systems serve billions of queries... ...ship them at Apple scale.This is a rare opportunity... ...of research and production, partnering closely... ...-on experience with LLM inference stacks. Working...SuggestedWorldwideRelocation- ...Applied Machine Learning Ark team in San Jose seeks engineers and researchers to advance LLM MaaS platforms. You will work across model... ...multimodal AI. You will contribute to end-to-end systems spanning training, inference, and evaluation in a fast-growing tech #J-18808...Suggested
$128k - $260.5k
...platform, its Zero Trust Engine, and the powerful... ...machine learning (ML) to protect the... ...deploy enterprise-scale AI solutions.... ...AI research into production-grade reality.Note... ...bleeding edge of LLM inference optimization, utilizing... ...production-grade systems.Architect High-Performance...Suggested$224k - $356.5k
...building agentic systems that can reason about... ...-layer of modern ML: the agents,... ...looking for exceptional engineers who are passionate... ...and researcher productivity.Create self-improving... ...and evolve large-scale Python and PyTorch... ....Strong agency in LLM-based systems, such...Full time$190.2k - $345.65k
...with creative production workflows and a... ...Machine Learning Engineer to architect... ...intelligence — the systems that turn... ...index media at scale, the hybrid and... ...that improve them.ML Engineering leadership... ...retrieval for LLM and agentic... ...models and the inference paths that...Full timeTemporary workLocal areaWorldwide$232k - $310k
...ground up, operating at scale across Azure and... ...of agentic talent systems.What sets Eightfold... ...standards. Our engineers, product leaders, and go-to-... ...world works.About AII/ML TeamOur AI/ML team... ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills &...Work experience placementWork at officeRemote workFlexible hours3 days per week$201.3k - $352.3k
...DescriptionIt all started when engineer Fred Luddy wrote code... ...AI and enterprise-scale search systems that power Now Assist, AI... ...design, build, and operate production-grade agentic AI systems... ...or retrieval. Exposure to LLM fine-tuning or inference optimization in production...Work experience placementWork at officeImmediate startRemote workFlexible hours- ...s Volcano Ark MaaS platform invites talented engineers to advance large-model inference systems across China and internal products. You will optimize inference performance, reduce... .... Join a team that blends cutting-edge ML with practical systems, working on disaggregated...
- ...Technologies is seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance inference platforms for serving large ML models in production. This remote U.S. role focuses on the systems engineering side of AI deployment, including routing,...Remote job
- ...hiring a PhD-level researcher to architect and build advanced LLM-based systems for processing high-dimensional data across diverse sources.... ...contextual prompting. You will optimize model serving for production with quantization and efficient attention, contributing to a...
- Rhoda AI in Mountain View is seeking a Staff / Principal ML Training Systems Engineer to lead the performance of large-scale multimodal training systems. This role involves improving training efficiency and collaborating closely with research teams to accelerate model iteration...
$162k - $387.6k
...Machine Learning Engineer Graduate (E-Commerce... ...& Logistics - LLM/Agent) - 2027 Start... ...learning, retrieval systems, and software engineering... ...at global scale.We are looking for... ...model compression, inference cost and latency optimization... ...series, events, product attributes, images...Temporary work- ...seeking a Staff Machine Learning Engineer in San Jose, CA to lead the development... ...and optimization of advanced ML models for integration into PayPal's software and systems. You will design, adapt, and deploy ML models across products, collaborate with cross-functional...
- ...infrastructure gap required to scale them efficiently,... ...cost-effectively in production. We bridge this exact gap by applying deep systems programming, software-... ...Austin and world-renowned ML systems researcher with... ...production experience engineering ML systems, OR a PhD...Shift work
$170k - $210k
...CORSAIR is building a production-grade AI/ML capability spanning intelligent... ...automation, agentic systems, predictive analytics, and data engineering. As AI/ML Engineering... ...systems using modern LLM orchestration frameworks... ...strategies at scale; apply LLM fine-tuning...For contractorsWork experience placement$151.8k - $265.35k
...paired with creative production workflows and a... ...Machine Learning Engineer to build the pipelines... ...them as services, scale those services to... .... Run production ML operationally - on... ...production ML or inference services at scale.... ...with multi-tenant systems and data isolation...Full timeTemporary workLocal areaWorldwide$150k - $230k
...recommendation systems, and adtech.Recognized... ...technologists, product innovators, and... ...challenges at scale.Together, we... ...Learning Engineer to drive the post... ...RequirementsHands-on LLM post-training... ...for ML. You can independently... ...DeepSpeed or FSDP, and inference engines like...Full timeLocal areaWork from home$206.4k - $379.1k
...paired with creative production workflows and a... ...Principal Machine Learning Engineer to serve as the... ...served at enterprise scale. You will set the inference architecture and technical... ...that makes those systems fast and cost-... ...architecture. Director, ML Engineering and ML Engineering...Full timeTemporary workLocal areaWorldwide$197.5k - $272k
...is already in production across more than... ...company with the scale and impact of... ...great Staff AI Engineer to join our seasoned... ..., including system logs, traces,... ...the end-to-end ML pipeline—from data... ...Transformers, LLM, CNN, LSTM,... ...(C++14/17 for inference).Deep proficiency...Work at officeWorldwideFlexible hoursShift work3 days per week$100k
...seeking an Physical Design Engineer to lead cross-functional... ...integrate, and deploy AI/ML-driven solutions into production physical design flows, creating... ...What you will learnHow to scale AI/ML-driven methodologies... ...access to information, systems, or technologies subject to...Permanent employment$155.5k - $315k
...generative AI, and production engineering, this role is... ...on experience with LLM-based agentic frameworks... ..., observable AI systems in a cloud-native... ...capabilities at production scale, including... ...production ML/AI systems, with at... ...clustering, causal inference, or related methods...Full timeWork experience placementWork at officeLocal areaImmediate start- ...the infrastructure for search features, impacting millions of users. You will handle large-scale data and be part of a collaborative team dedicated to creating innovative products. The ideal candidate has a computer science background, strong coding skills, and...
- ...Inc. is seeking a Sr. Machine Learning Engineer for the Foundation Models Inference team in Santa Clara, CA. You will... ...architectures from prototype to planetary-scale deployment, owning hard problems in inference efficiency and system design. You will build and optimize...
- ...commerce AI team. You will architect and develop sophisticated LLM-based systems to process high-dimensional unstructured data from diverse... .... Proficiency in C++, Python, Go, or Java, plus experience in ML/NLP, is required. Onboarding must occur by year end. Base pay,...
- ...is a Silicon Valley leader powering the future of physical AI. We seek a Performance Engineer to optimize large-scale ML workloads, focusing on distributed training, batch inference, and cost-effective data processing. You will own profiling across the stack, identify...
- NVIDIA’s Cosmos team seeks engineers to build agentic AI-native software and... ...codebases, design end-to-end ML pipelines, and create self-improving... ...We expect deep expertise in ML systems, robust software engineering, and hands-on work with LLM-based tooling in a fast-moving...
- ...Recommendation Algorithm team, where you will build and optimize production-grade recommendation models that shape user experience... ..., and platform security. You’ll deliver end-to-end ML solutions and own full-stack systems, collaborating with cross-functional teams to grow...
- PayPal in San Jose, CA seeks a Senior ML Engineer to lead development and optimization of advanced ML models, integrate them into existing software, and deploy solutions in production environments. You will collaborate with engineers, product management, and QA to ensure...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Systems Engineer — Production-Scale LLM Inference. Be the first to apply!
- machine learning ai engineer San Jose, CA
- computer vision machine learning engineer San Jose, CA
- machine learning engineer San Jose, CA
- ai ml engineer San Jose, CA
- graduate machine learning engineer San Jose, CA
- machine learning software engineer San Jose, CA
- senior ml engineer San Jose, CA
- system engineer remote San Jose, CA
- senior windows systems engineer San Jose, CA
- systems engineer intern San Jose, CA

