Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Systems Engineer — Production-Scale LLM Inference

$150k - $350k

ChipAgents

ChipAgents is looking for a skilled ML Systems Engineer in San Jose, California. In this technical role, you will optimize large language model inference for our agentic AI platform, impacting chip design efficiency. Your responsibilities include implementing performance optimizations, designing multi-node clusters, and collaborating with top semiconductor firms. With a competitive salary range of $150K–$350K, we also offer substantial benefits and equity options. #J-18808-Ljbffr ChipAgents

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the ML Systems Engineer — Production-Scale LLM Inference in San Jose, CA vacancy
  • $150k - $350k

     ...generative AI to assist engineers in RTL design,...  ...is deployed in production to companies that...  ...We are seeking an ML Systems Engineer to optimize...  ...large language model inference powering our...  ...push the limits of LLM throughput and latency...  ...with large‑scale ML systems, GPU computing... 
    Suggested

    ChipAgents

    San Jose, CA
    4 days ago
  • $174.72k - $295.68k

     ...strong foundation for LLM deployment and quality...  ...PTQ, QAT, on-vehicle inference and related fields.Key...  ...bit techniques.Develop production-quality Python code with...  ...and software engineering skills.Ability to work...  ...effectively across research, systems, infrastructure, and product... 
    Suggested
    Full time

    XPENG Motors

    Santa Clara, CA
    2 days ago
  • $184.7k - $324.8k

    Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS &...  ...foundation models.Our systems serve billions of queries...  ...ship them at Apple scale.This is a rare opportunity...  ...of research and production, partnering closely...  ...-on experience with LLM inference stacks. Working... 
    Suggested
    Worldwide
    Relocation

    Apple Inc.

    Santa Clara, CA
    3 days ago
  •  ...Applied Machine Learning Ark team in San Jose seeks engineers and researchers to advance LLM MaaS platforms. You will work across model...  ...multimodal AI. You will contribute to end-to-end systems spanning training, inference, and evaluation in a fast-growing tech #J-18808... 
    Suggested

    ByteDance

    San Jose, CA
    5 days ago
  • $128k - $260.5k

     ...platform, its Zero Trust Engine, and the powerful...  ...machine learning (ML) to protect the...  ...deploy enterprise-scale AI solutions....  ...AI research into production-grade reality.Note...  ...bleeding edge of LLM inference optimization, utilizing...  ...production-grade systems.Architect High-Performance... 
    Suggested

    Netskope

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...building agentic systems that can reason about...  ...-layer of modern ML: the agents,...  ...looking for exceptional engineers who are passionate...  ...and researcher productivity.Create self-improving...  ...and evolve large-scale Python and PyTorch...  ....Strong agency in LLM-based systems, such... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $190.2k - $345.65k

     ...with creative production workflows and a...  ...Machine Learning Engineer to architect...  ...intelligence — the systems that turn...  ...index media at scale, the hybrid and...  ...that improve them.ML Engineering leadership...  ...retrieval for LLM and agentic...  ...models and the inference paths that... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    2 days ago
  • $232k - $310k

     ...ground up, operating at scale across Azure and...  ...of agentic talent systems.What sets Eightfold...  ...standards. Our engineers, product leaders, and go-to-...  ...world works.About AII/ML TeamOur AI/ML team...  ...strategies (QLORA, DPO) and inference optimization (vLLM, TensorRT-LLM).Desired Skills &... 
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    3 days per week

    Eightfold

    Santa Clara, CA
    3 days ago
  • $201.3k - $352.3k

     ...DescriptionIt all started when engineer Fred Luddy wrote code...  ...AI and enterprise-scale search systems that power Now Assist, AI...  ...design, build, and operate production-grade agentic AI systems...  ...or retrieval. Exposure to LLM fine-tuning or inference optimization in production... 
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    2 days ago
  •  ...s Volcano Ark MaaS platform invites talented engineers to advance large-model inference systems across China and internal products. You will optimize inference performance, reduce...  .... Join a team that blends cutting-edge ML with practical systems, working on disaggregated... 

    ByteDance

    San Jose, CA
    2 days ago
  •  ...Technologies is seeking a Machine Learning Infrastructure Engineer to design, build, and operate high-performance inference platforms for serving large ML models in production. This remote U.S. role focuses on the systems engineering side of AI deployment, including routing,... 
    Remote job

    Bright Vision Technologies

    Mountain View, CA
    5 days ago
  •  ...hiring a PhD-level researcher to architect and build advanced LLM-based systems for processing high-dimensional data across diverse sources....  ...contextual prompting. You will optimize model serving for production with quantization and efficient attention, contributing to a... 

    TikTok USDS Joint Venture

    San Jose, CA
    19 hours ago
  • Rhoda AI in Mountain View is seeking a Staff / Principal ML Training Systems Engineer to lead the performance of large-scale multimodal training systems. This role involves improving training efficiency and collaborating closely with research teams to accelerate model iteration... 

    Rhoda AI

    Mountain View, CA
    5 days ago
  • $162k - $387.6k

     ...Machine Learning Engineer Graduate (E-Commerce...  ...& Logistics - LLM/Agent) - 2027 Start...  ...learning, retrieval systems, and software engineering...  ...at global scale.We are looking for...  ...model compression, inference cost and latency optimization...  ...series, events, product attributes, images... 
    Temporary work

    TikTok

    San Jose, CA
    3 days ago
  •  ...seeking a Staff Machine Learning Engineer in San Jose, CA to lead the development...  ...and optimization of advanced ML models for integration into PayPal's software and systems. You will design, adapt, and deploy ML models across products, collaborate with cross-functional... 

    PayPal

    San Jose, CA
    5 days ago
  •  ...infrastructure gap required to scale them efficiently,...  ...cost-effectively in production. We bridge this exact gap by applying deep systems programming, software-...  ...Austin and world-renowned ML systems researcher with...  ...production experience engineering ML systems, OR a PhD... 
    Shift work

    Success Matcher Recruitment

    Sunnyvale, CA
    3 days ago
  • $170k - $210k

     ...CORSAIR is building a production-grade AI/ML capability spanning intelligent...  ...automation, agentic systems, predictive analytics, and data engineering. As AI/ML Engineering...  ...systems using modern LLM orchestration frameworks...  ...strategies at scale; apply LLM fine-tuning... 
    For contractors
    Work experience placement

    Corsair

    Milpitas, CA
    4 days ago
  • $151.8k - $265.35k

     ...paired with creative production workflows and a...  ...Machine Learning Engineer to build the pipelines...  ...them as services, scale those services to...  .... Run production ML operationally - on...  ...production ML or inference services at scale....  ...with multi-tenant systems and data isolation... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    1 day ago
  • $150k - $230k

     ...recommendation systems, and adtech.Recognized...  ...technologists, product innovators, and...  ...challenges at scale.Together, we...  ...Learning Engineer to drive the post...  ...RequirementsHands-on LLM post-training...  ...for ML. You can independently...  ...DeepSpeed or FSDP, and inference engines like... 
    Full time
    Local area
    Work from home

    News Break

    Mountain View, CA
    2 days ago
  • $206.4k - $379.1k

     ...paired with creative production workflows and a...  ...Principal Machine Learning Engineer to serve as the...  ...served at enterprise scale. You will set the inference architecture and technical...  ...that makes those systems fast and cost-...  ...architecture. Director, ML Engineering and ML Engineering... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    2 days ago
  • $197.5k - $272k

     ...is already in production across more than...  ...company with the scale and impact of...  ...great Staff AI Engineer to join our seasoned...  ..., including system logs, traces,...  ...the end-to-end ML pipeline—from data...  ...Transformers, LLM, CNN, LSTM,...  ...(C++14/17 for inference).Deep proficiency... 
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    Sunnyvale, CA
    2 days ago
  • $100k

     ...seeking an Physical Design Engineer to lead cross-functional...  ...integrate, and deploy AI/ML-driven solutions into production physical design flows, creating...  ...What you will learnHow to scale AI/ML-driven methodologies...  ...access to information, systems, or technologies subject to... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    4 days ago
  • $155.5k - $315k

     ...generative AI, and production engineering, this role is...  ...on experience with LLM-based agentic frameworks...  ..., observable AI systems in a cloud-native...  ...capabilities at production scale, including...  ...production ML/AI systems, with at...  ...clustering, causal inference, or related methods... 
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start

    Hewlett Packard Enterprise

    San Jose, CA
    3 days ago
  •  ...the infrastructure for search features, impacting millions of users. You will handle large-scale data and be part of a collaborative team dedicated to creating innovative products. The ideal candidate has a computer science background, strong coding skills, and... 

    Apple

    Cupertino, CA
    2 days ago
  •  ...Inc. is seeking a Sr. Machine Learning Engineer for the Foundation Models Inference team in Santa Clara, CA. You will...  ...architectures from prototype to planetary-scale deployment, owning hard problems in inference efficiency and system design. You will build and optimize... 

    Apple Inc.

    Santa Clara, CA
    3 days ago
  •  ...commerce AI team. You will architect and develop sophisticated LLM-based systems to process high-dimensional unstructured data from diverse...  .... Proficiency in C++, Python, Go, or Java, plus experience in ML/NLP, is required. Onboarding must occur by year end. Base pay,... 

    TikTok USDS Joint Venture

    San Jose, CA
    3 days ago
  •  ...is a Silicon Valley leader powering the future of physical AI. We seek a Performance Engineer to optimize large-scale ML workloads, focusing on distributed training, batch inference, and cost-effective data processing. You will own profiling across the stack, identify... 

    Decisive Point

    Sunnyvale, CA
    4 days ago
  • NVIDIA’s Cosmos team seeks engineers to build agentic AI-native software and...  ...codebases, design end-to-end ML pipelines, and create self-improving...  ...We expect deep expertise in ML systems, robust software engineering, and hands-on work with LLM-based tooling in a fast-moving... 

    Thomas To

    Santa Clara, CA
    3 days ago
  •  ...Recommendation Algorithm team, where you will build and optimize production-grade recommendation models that shape user experience...  ..., and platform security. You’ll deliver end-to-end ML solutions and own full-stack systems, collaborating with cross-functional teams to grow... 

    TikTok

    San Jose, CA
    2 days ago
  • PayPal in San Jose, CA seeks a Senior ML Engineer to lead development and optimization of advanced ML models, integrate them into existing software, and deploy solutions in production environments. You will collaborate with engineers, product management, and QA to ensure... 
    Remote work

    Relha LLC

    San Jose, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Systems Engineer — Production-Scale LLM Inference. Be the first to apply!