Senior ML Systems Scientist: LLM/VLM Inference Expert
Nebius B.V.
Nebius Token Factory seeks a PhD-level researcher to lead focused ML research projects from hypothesis to production handoff. You will partner with MLEs to make prototypes production-ready and drive efficient LLM/VLM inference with measurable impact. The role emphasizes publishing results, sharing technical reports, and mentoring teams on rigorous experimental design and tradeoffs. You will operate at the intersection of research and production in Palo Alto. #J-18808-Ljbffr Nebius B.V.
$195.2k - $262.2k
...large in-house AI/ML infrastructure. Built... ...orchestration to inference optimization, we... ...Factory needs scientists who can turn frontier... ...capabilities. A Senior Applied Scientist... ...programs in efficient LLM and VLM inference with... ...rigor, and model/system tradeoffs. Must-haves...SeniorTemporary workImmediate startRemote work- Nebius is seeking a Senior Applied Scientist to turn frontier inference bottlenecks into production-ready solutions. You... ...and PyTorch, and collaborate with ML engineers to ship research into production... ..., and cost per token across LLM/VLM inference workloads. #J-18808-...Senior
- Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput. Requires over...Senior
- d-Matrix in Santa Clara, CA is seeking a Sr. Staff ML Researcher to advance LLM algorithmic optimization on our DNN accelerators. You will design and implement efficient inference algorithms, collaborating with mathematicians, ML researchers and engineers on high-impact...Senior3 days per week
- Netskope is seeking a Senior Staff Machine Learning Scientist in Santa Clara to own the inference and optimization layer for AI in agentic workflows. You will fine-tune... ...optimization, and hardware acceleration, partnering with systems and backend engineers to ship end-to-end...Senior
$184k - $287.5k
...engineers to join us and build AI inference systems that serve large-scale models... ...decoding, data/tensor/expert/pipeline-parallelism, prefill... ...pareto frontier for the field of ML Systems; survey recent... ...crowdExperience building and optimizing LLM inference engines (e.g., vLLM...SeniorFull time- ...per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:... ...a team of system software experts to build out the deployment... ...closely with other software (ML and compilers) and hardware... ...(such as TensorRT-LLM, vLLM, SGLang, etc.)Experience...3 days per week
$272k - $431.25k
...throughput, low-latency inference framework for serving... ...accelerators feel like a single system at datacenter scale. As... ...of cutting-edge LLM workloads.We are... ...and memory pools.Mentor senior and junior engineers, set... ...performance storage, or ML systems infrastructure...Full timeLocal areaRemote work$183.83k - $275.98k
...on bringing advancements in the field of ML and large-scale learning to the AV domain... ...towards a more end-to-end autonomous driving system. This role requires working with and... ...to improve model optimization and inference speeds. You will use your applied research...Senior$192.2k - $260k
...is assembling an elite team of world-class scientists and engineers to pioneer the next... ...driven development tools. Join the Amazon Kiro LLM-Training team and help create groundbreaking... ...and present your pioneering work at premier ML and NLP conferences (NeurIPS, ICML, ICLR ,...SeniorWork at officeLocal areaWorldwideFlexible hours- Pinterest is seeking an experienced Data & Applied Scientist to advance ML measurement, feature understanding, and causal inference at scale. You will own end-to-end design of production ML systems and collaborate with cross-functional teams to turn research into durable...SeniorWork at office
- JPMorgan Chase & Co. in Palo Alto seeks an Executive Director and Senior Principal Engineer for Applied AI/ML to build production search, conversational AI, and agentic workflow systems. This hands-on role focuses on architecture and implementation with no people management...Senior
- ...headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to... ...deployment software and collaborating with ML, compiler, and hardware experts. Required: strong background in system software...Senior
- ...low-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads. This role is in the...Senior
- A leader in AI technology in Palo Alto is seeking a Senior AI Systems Performance Engineer to optimize the latest foundation models on their innovative... ...in deep learning, performance optimization, and major ML frameworks. This position offers competitive salary and...Senior
$184k - $230k
Senior Applied Scientist, Inference (MCM) Join to apply for the Senior Applied Scientist, Inference (MCM) role... ...built some of the most successful ad systems at Google, including YouTube's monetization... ...: Guide and assist partner teams (ML, Product, and Sales) through the...SeniorFull time$124.5k - $272k
...Positions are available at Senior Staff and above.... ...Machine Learning Scientist, you own the inference and optimization layer... ...matures.Partner with the systems and backend engineers... ...4+ years hands-on in ML/AI (model development... .../SGLang, TensorRT-LLM, ONNX Runtime, llama....Senior$254k - $350k
...generation of autonomous system intelligence.As a... ...for more efficient inference by sharing various... ...We are looking for experts with hands-on... ...you will optimize ML models, write custom... ...driving paradigms (VLM/VLA models, Foundation... ...technologies (e.g., TensorRT-LLM).$254,000 - $350,00...SeniorFull timeTemporary workRelocation package$152k - $241.5k
...optimize containerized inference execution for the latest... ...security patches)Contribute VLM-related features to Open... ...-grade AI distributed systems, backend services,... ...SGLang, Torch, TRT, TRT-LLM).Proficiency with Python... ...already?Experience with ML model engineering: training...SeniorFull time- ...influential companies.As a Senior Principal Software... ...frameworks to standardize ML Engineering services,... ...optimization using model inference servers such as Triton... ...and deploying LLM & GNN solutions on AWS... ...optimization and distributed systems for large models focused...Senior
- ...with all of their business systems through natural language... ...help build cutting edge ML infrastructure for building and serving LLM’s at Moveworks. This... ...distributed training and inference pipeline for large language... ..., machine learning experts, and product to build new...SeniorWork at officeRemote workFlexible hours
- d-Matrix inc. is looking for a Senior Staff ML Researcher to join our Algo team in Santa Clara, CA. This hybrid position involves working... ...successful candidate will develop algorithms for optimizing LLM inference on our DNN accelerators. Ideal applicants should have a MSc...Senior3 days per week
$192.2k - $260k
We are looking for a Senior Applied Scientist to help drive the research and development... ...building the post-training systems (reward modeling,... ...roadmap, and work closely with inference engineers to ensure your models... ...to open-source speech/audio ML systems or widely used...SeniorLocal areaFlexible hours$226k - $306k
...team develops novel AI/ML solutions that power intelligent... ...modeling, causal inference, simulation-based... ...agentic and multi-agent systems, neuro-symbolic AI, and LLM-based reasoning for business... ...We are looking for a Senior Staff AI Research Scientist to shape and drive the...Senior- Voltai in California (Palo Alto) seeks a senior formal verification researcher to develop new... ...collaborate with RTL, verification, and ML teams to scale AI-hardware verification, prototype... ...RTL, and turn research into practical systems. You will work closely with cross-...Senior
- ...Santa Clara, CA, headquarters 3 days per week. The role Senior Staff ML Researcher - LLM Algorithmic Optimization What You Will Do d-Matrix is... ...algorithms that will be used to optimize large language model inference on DNN accelerators we develop. You would be part of a...Senior3 days per week
- ...Member of Technical Staff for research on LLM agents in Palo Alto, California. You will... ...collaborate with engineers to create impactful AI systems. Essential qualifications include a solid... ...candidates also have familiarity with ML frameworks like PyTorch or TensorFlow....
- Google Research is seeking a Senior Research Scientist to drive bold research initiatives in Gemini and related areas in Mountain View. You will... ...contribute to the advancement of information retrieval, NLP, and AI systems while delivering tangible results for real-world products...Senior
$119.8k - $234.7k
...representations, and heterogeneous event streams to infer user intent and advertiser value, even... .... The team owns end-to-end ML systems, including large-scale data and label construction... ...marketplace dynamics. Engineers and scientists on the team work at the intersection of...SeniorOngoing contractWork at officeLocal areaShift work- Accellor is seeking a Technical Architect — AI Systems, Inference & Platform Internals to design, scale, and optimize internal AI systems powering... ...reliability. The ideal candidate has 10-12 years of software/ML infra experience, deep knowledge of PyTorch/JAX, and hands-on...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior ML Systems Scientist: LLM/VLM Inference Expert. Be the first to apply!
- materials scientist Palo Alto, CA
- scientist assay development Palo Alto, CA
- health scientist Palo Alto, CA
- quality control scientist Palo Alto, CA
- deep learning scientist Palo Alto, CA
- application scientist Palo Alto, CA
- decision scientist Palo Alto, CA
- scientist 1 Palo Alto, CA
- lab scientist Palo Alto, CA
- scientist Palo Alto, CA

