Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior ML Systems Scientist: LLM/VLM Inference Expert

Nebius B.V.

Nebius Token Factory seeks a PhD-level researcher to lead focused ML research projects from hypothesis to production handoff. You will partner with MLEs to make prototypes production-ready and drive efficient LLM/VLM inference with measurable impact. The role emphasizes publishing results, sharing technical reports, and mentoring teams on rigorous experimental design and tradeoffs. You will operate at the intersection of research and production in Palo Alto. #J-18808-Ljbffr Nebius B.V.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior ML Systems Scientist: LLM/VLM Inference Expert in Palo Alto, CA vacancy
  • $195.2k - $262.2k

     ...large in-house AI/ML infrastructure. Built...  ...orchestration to inference optimization, we...  ...Factory needs scientists who can turn frontier...  ...capabilities. A Senior Applied Scientist...  ...programs in efficient LLM and VLM inference with...  ...rigor, and model/system tradeoffs. Must-haves... 
    Senior
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    20 hours ago
  • Nebius is seeking a Senior Applied Scientist to turn frontier inference bottlenecks into production-ready solutions. You...  ...and PyTorch, and collaborate with ML engineers to ship research into production...  ..., and cost per token across LLM/VLM inference workloads. #J-18808-... 
    Senior

    Nebius

    Palo Alto, CA
    5 days ago
  • Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput. Requires over... 
    Senior

    NVIDIA AI

    Santa Clara, CA
    2 days ago
  • d-Matrix in Santa Clara, CA is seeking a Sr. Staff ML Researcher to advance LLM algorithmic optimization on our DNN accelerators. You will design and implement efficient inference algorithms, collaborating with mathematicians, ML researchers and engineers on high-impact... 
    Senior
    3 days per week

    Entrada Ventures

    Santa Clara, CA
    2 days ago
  • Netskope is seeking a Senior Staff Machine Learning Scientist in Santa Clara to own the inference and optimization layer for AI in agentic workflows. You will fine-tune...  ...optimization, and hardware acceleration, partnering with systems and backend engineers to ship end-to-end... 
    Senior

    Netskope

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...engineers to join us and build AI inference systems that serve large-scale models...  ...decoding, data/tensor/expert/pipeline-parallelism, prefill...  ...pareto frontier for the field of ML Systems; survey recent...  ...crowdExperience building and optimizing LLM inference engines (e.g., vLLM... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:...  ...a team of system software experts to build out the deployment...  ...closely with other software (ML and compilers) and hardware...  ...(such as TensorRT-LLM, vLLM, SGLang, etc.)Experience... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...throughput, low-latency inference framework for serving...  ...accelerators feel like a single system at datacenter scale. As...  ...of cutting-edge LLM workloads.We are...  ...and memory pools.Mentor senior and junior engineers, set...  ...performance storage, or ML systems infrastructure... 
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $183.83k - $275.98k

     ...on bringing advancements in the field of ML and large-scale learning to the AV domain...  ...towards a more end-to-end autonomous driving system. This role requires working with and...  ...to improve model optimization and inference speeds. You will use your applied research... 
    Senior

    Nuro - Mountain View, CA, US

    Mountain View, CA
    5 days ago
  • $192.2k - $260k

     ...is assembling an elite team of world-class scientists and engineers to pioneer the next...  ...driven development tools. Join the Amazon Kiro LLM-Training team and help create groundbreaking...  ...and present your pioneering work at premier ML and NLP conferences (NeurIPS, ICML, ICLR ,... 
    Senior
    Work at office
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    2 days ago
  • Pinterest is seeking an experienced Data & Applied Scientist to advance ML measurement, feature understanding, and causal inference at scale. You will own end-to-end design of production ML systems and collaborate with cross-functional teams to turn research into durable... 
    Senior
    Work at office

    Pinterest

    Palo Alto, CA
    5 days ago
  • JPMorgan Chase & Co. in Palo Alto seeks an Executive Director and Senior Principal Engineer for Applied AI/ML to build production search, conversational AI, and agentic workflow systems. This hands-on role focuses on architecture and implementation with no people management... 
    Senior

    JPMorgan Chase & Co.

    Palo Alto, CA
    2 days ago
  •  ...headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to...  ...deployment software and collaborating with ML, compiler, and hardware experts. Required: strong background in system software... 
    Senior

    Jobleads-US

    Santa Clara, CA
    4 days ago
  •  ...low-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads. This role is in the... 
    Senior

    Sanas

    Palo Alto, CA
    5 days ago
  • A leader in AI technology in Palo Alto is seeking a Senior AI Systems Performance Engineer to optimize the latest foundation models on their innovative...  ...in deep learning, performance optimization, and major ML frameworks. This position offers competitive salary and... 
    Senior

    SambaNova

    Palo Alto, CA
    1 day ago
  • $184k - $230k

    Senior Applied Scientist, Inference (MCM) Join to apply for the Senior Applied Scientist, Inference (MCM) role...  ...built some of the most successful ad systems at Google, including YouTube's monetization...  ...: Guide and assist partner teams (ML, Product, and Sales) through the... 
    Senior
    Full time

    Moloco

    Redwood City, CA
    3 days ago
  • $124.5k - $272k

     ...Positions are available at Senior Staff and above....  ...Machine Learning Scientist, you own the inference and optimization layer...  ...matures.Partner with the systems and backend engineers...  ...4+ years hands-on in ML/AI (model development...  .../SGLang, TensorRT-LLM, ONNX Runtime, llama.... 
    Senior

    Netskope

    Santa Clara, CA
    3 days ago
  • $254k - $350k

     ...generation of autonomous system intelligence.As a...  ...for more efficient inference by sharing various...  ...We are looking for experts with hands-on...  ...you will optimize ML models, write custom...  ...driving paradigms (VLM/VLA models, Foundation...  ...technologies (e.g., TensorRT-LLM).$254,000 - $350,00... 
    Senior
    Full time
    Temporary work
    Relocation package

    Zoox

    Foster, CA
    5 days ago
  • $152k - $241.5k

     ...optimize containerized inference execution for the latest...  ...security patches)Contribute VLM-related features to Open...  ...-grade AI distributed systems, backend services,...  ...SGLang, Torch, TRT, TRT-LLM).Proficiency with Python...  ...already?Experience with ML model engineering: training... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...influential companies.As a Senior Principal Software...  ...frameworks to standardize ML Engineering services,...  ...optimization using model inference servers such as Triton...  ...and deploying LLM & GNN solutions on AWS...  ...optimization and distributed systems for large models focused... 
    Senior

    JP Morgan Chase

    Palo Alto, CA
    2 days ago
  •  ...with all of their business systems through natural language...  ...help build cutting edge ML infrastructure for building and serving LLM’s at Moveworks. This...  ...distributed training and inference pipeline for large language...  ..., machine learning experts, and product to build new... 
    Senior
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    2 days ago
  • d-Matrix inc. is looking for a Senior Staff ML Researcher to join our Algo team in Santa Clara, CA. This hybrid position involves working...  ...successful candidate will develop algorithms for optimizing LLM inference on our DNN accelerators. Ideal applicants should have a MSc... 
    Senior
    3 days per week

    d-Matrix inc.

    Santa Clara, CA
    20 hours ago
  • $192.2k - $260k

    We are looking for a Senior Applied Scientist to help drive the research and development...  ...building the post-training systems (reward modeling,...  ...roadmap, and work closely with inference engineers to ensure your models...  ...to open-source speech/audio ML systems or widely used... 
    Senior
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    5 days ago
  • $226k - $306k

     ...team develops novel AI/ML solutions that power intelligent...  ...modeling, causal inference, simulation-based...  ...agentic and multi-agent systems, neuro-symbolic AI, and LLM-based reasoning for business...  ...We are looking for a Senior Staff AI Research Scientist to shape and drive the... 
    Senior

    Intuit

    Mountain View, CA
    2 days ago
  • Voltai in California (Palo Alto) seeks a senior formal verification researcher to develop new...  ...collaborate with RTL, verification, and ML teams to scale AI-hardware verification, prototype...  ...RTL, and turn research into practical systems. You will work closely with cross-... 
    Senior

    Voltai

    Palo Alto, CA
    4 days ago
  •  ...Santa Clara, CA, headquarters 3 days per week. The role Senior Staff ML Researcher - LLM Algorithmic Optimization What You Will Do d-Matrix is...  ...algorithms that will be used to optimize large language model inference on DNN accelerators we develop. You would be part of a... 
    Senior
    3 days per week

    d-Matrix inc.

    Santa Clara, CA
    20 hours ago
  •  ...Member of Technical Staff for research on LLM agents in Palo Alto, California. You will...  ...collaborate with engineers to create impactful AI systems. Essential qualifications include a solid...  ...candidates also have familiarity with ML frameworks like PyTorch or TensorFlow.... 

    NeoCognition Inc.

    Palo Alto, CA
    5 days ago
  • Google Research is seeking a Senior Research Scientist to drive bold research initiatives in Gemini and related areas in Mountain View. You will...  ...contribute to the advancement of information retrieval, NLP, and AI systems while delivering tangible results for real-world products... 
    Senior

    Google Inc.

    Mountain View, CA
    20 hours ago
  • $119.8k - $234.7k

     ...representations, and heterogeneous event streams to infer user intent and advertiser value, even...  .... The team owns end-to-end ML systems, including large-scale data and label construction...  ...marketplace dynamics. Engineers and scientists on the team work at the intersection of... 
    Senior
    Ongoing contract
    Work at office
    Local area
    Shift work

    Microsoft Corporation

    Sunnyvale, CA
    3 days ago
  • Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI systems powering...  ...reliability. The ideal candidate has 10-12 years of software/ML infra experience, deep knowledge of PyTorch/JAX, and hands-on... 
    Senior

    Accellor

    Mountain View, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior ML Systems Scientist: LLM/VLM Inference Expert. Be the first to apply!