Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Applied Scientist — Efficient LLM Inference & Optimization

Nebius

Nebius is seeking a Senior Applied Scientist to turn frontier inference bottlenecks into production-ready solutions. You will design rigorous experiments, write high-quality code in Python and PyTorch, and collaborate with ML engineers to ship research into production. You will own well-scoped research projects, publish credible work, mentor engineers, and define evaluation methods for latency, throughput, and cost per token across LLM/VLM inference workloads. #J-18808-Ljbffr Nebius

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Applied Scientist — Efficient LLM Inference & Optimization in Palo Alto, CA vacancy
  • $195.2k - $262.2k

     ...GPU orchestration to inference optimization, we own the hard...  ...storage, networking and applied AI. Listed on...  ...Token Factory needs scientists who can turn frontier...  ...inference capabilities. A Senior Applied Scientist...  ...research programs in efficient LLM and VLM inference... 
    Senior
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    6 hours ago
  • $192k - $304.75k

     ...are now looking for an Applied Deep Learning Research Scientist, Efficiency!Join our ADLR - Efficiency...  ...and algorithms to optimize neural networks for training...  ...effect on neural network inference and training accuracy. This...  ..., optimizers and LLM training.Experience with... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • d-Matrix in Santa Clara, CA is seeking a Sr. Staff ML Researcher to advance LLM algorithmic optimization on our DNN accelerators. You will design and implement efficient inference algorithms, collaborating with mathematicians, ML researchers and engineers on high-impact... 
    Senior
    3 days per week

    Entrada Ventures

    Santa Clara, CA
    1 day ago
  • $139k - $229k

     ...for an ambitious data scientist to have an impact. We...  ...exhibit technical acumen on inference and algorithms, and...  ...centered on trust and optimized for culture,...  ...a thought partner to senior leaders to prioritize/...  ...Informatics, Engineering, Applied Mathematics, Economics... 
    Senior
    Full time
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    LinkedIn

    Mountain View, CA
    3 days ago
  •  ...projects from hypothesis to production handoff. You will partner with MLEs to make prototypes production-ready and drive efficient LLM/VLM inference with measurable impact. The role emphasizes publishing results, sharing technical reports, and mentoring teams on rigorous... 
    Senior

    Nebius B.V.

    Palo Alto, CA
    3 days ago
  • $192.2k - $260k

     ...practical experience to join the Modeling and Optimization (MOP) Routing Science team. Your main...  ...data structures, particularly as it applies to vehicle routing and related problems...  ..., and analysis that will improve the efficiency and cost effectiveness of global fulfillment... 
    Senior
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  • $119.8k - $234.7k

     ...end-to-end systems for AI inference and agent workflows on Windows...  ...We are looking for a Senior Applied Scientist to develop and ship machine...  ...model quality and the efficiency of AI workloads on a diverse...  ...solve challenges in model optimization, search and retrieval, inference... 
    Senior
    Ongoing contract
    Local area

    Microsoft Corporation

    Mountain View, CA
    2 days ago
  • $192.2k - $260k

     ...Amazon Advertising, we apply Machine Learning at massive scale to optimize programmatic...  ...looking for a talented Senior Applied Scientist to join our team of scientists...  ...) serving at inference latencies under 10ms...  ...stores) to deploy models efficiently at scaleMentor scientists... 
    Senior
    Local area
    Worldwide
    Flexible hours
    Shift work

    Amazon

    Palo Alto, CA
    1 day ago
  • $184k - $230k

    Senior Applied Scientist, Inference (MCM) Join to apply for the Senior Applied Scientist, Inference (MCM) role at Moloco . About Moloco Moloco builds...  ...to our ML models. Validate and quantify the efficiency and performance gain from hypotheses with advanced statistical... 
    Senior
    Full time

    Moloco

    Redwood City, CA
    2 days ago
  • $192.2k - $260k

     ...an elite team of world-class scientists and engineers to pioneer the...  ...tools. Join the Amazon Kiro LLM-Training team and help create...  ...any scale with unprecedented efficiency.Broadly, AWS Utility Computing...  ...Master's degree and 6+ years of applied research experience-... 
    Senior
    Work at office
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  •  ...days per week. The role Senior Staff ML Researcher - LLM Algorithmic Optimization What You Will Do d-...  ...design, and implement efficient algorithms that will be...  ...optimize large language model inference on DNN accelerators we...  ...who create and apply advanced algorithmic and... 
    Senior
    3 days per week

    d-Matrix inc.

    Santa Clara, CA
    6 hours ago
  • $182.5k - $260.5k

     ...Positions are available at Senior Staff and above....  ...Machine Learning Scientist, you own the inference and optimization layer that makes AI...  ...workflows fast, efficient, and production-grade...  .../SGLang, TensorRT-LLM, ONNX Runtime,...  ...applicable laws, the range applies to candidates in... 
    Senior

    Netskope

    Santa Clara, CA
    2 days ago
  •  ...: Sr. Staff, ML Researcher - LLM Algorithmic OptimizationWhat...  ...invent, design, and implement efficient algorithms that will be used to optimize large language model inference on DNN accelerators we develop...  ...ML engineers who create and apply advanced algorithmic and numerical... 
    Senior
    3 days per week

    d-Matrix

    Santa Clara, CA
    2 days ago
  • $192k - $304.75k

     ...looking for a passionate scientist at the intersection of...  .... As a Sr. Quantum Applied Research Scientist, you...  ...modeling, and co-optimized calibration-decoding pipelines...  ...and parameter inference without full experimental...  ...tuning—including parameter-efficient methods (LoRA, QLoRA,... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $192.2k - $260k

    We are looking for a Senior Applied Scientist to help drive the research and...  ...roadmap, and work closely with inference engineers to ensure your...  ...- Advance the scaling and efficiency of conversational models,...  ...SFT through RL alignment, optimized for real-time multimodal... 
    Senior
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    3 days ago
  • $192.2k - $260k

     ...discover new products they love, be the most efficient way for advertisers to meet their...  ...into specific plans for research and applied scientists, as well as engineering and product...  ...perform proof-of-concept, experiment, optimize, and deploy your models into production... 
    Senior
    Local area
    Flexible hours

    Amazon Science

    Palo Alto, CA
    2 days ago
  •  ...Sr. Staff, ML Researcher - LLM Algorithmic Optimization What You Will Do d-Matrix...  ...invent, design, and implement efficient algorithms that will be...  ...large language model inference on DNN accelerators we develop...  ...engineers who create and apply advanced algorithmic and numerical... 
    Senior
    3 days per week

    Entrada Ventures

    Santa Clara, CA
    1 day ago
  • $143.63k - $299.38k

     ...information. You will apply your insights on the data...  ...just follow the latest LLM trends; you understand...  ...architectures and how to optimize them for massive scale....  ...models using parameter‑efficient techniques (LoRA,...  ...distributed training and inference. Publications, patents... 
    Senior
    Work at office
    Flexible hours

    Yahoo Holdings Inc.

    Mountain View, CA
    1 day ago
  •  ...what’s possible with LLM inference on heterogeneous hardware...  ...patterns to deep optimization of inference kernels,...  ...computational fabric. We are an applied research and...  ...AI models run efficiently and cost-effectively...  ...real hardware.• Small, senior team with high autonomy... 
    Senior

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $228.7k - $309.4k

    As a Principal Applied Scientist for Full-Funnel Campaign optimization, you will invent the models that jointly allocate budget across sponsored ad products to maximize advertiser outcomes and long-term customer value.This is a rare charter to build foundational optimization... 
    Local area
    Flexible hours
    Day shift

    Amazon

    Palo Alto, CA
    2 days ago
  • $228.7k - $309.4k

     ...from ad creation and optimization to performance analysis...  ...You will be the Gen AI applied science leader that...  ...expertise in the area of ML, LLM and GenAI models. You...  ...our team of applied scientists and engineers.Key job...  ...of autonomy and efficiency. You'll be responsible... 
    Local area
    Flexible hours

    Amazon

    Palo Alto, CA
    4 days ago
  • $171.6k - $222.2k

     ...lifecycle from ad creation and optimization to performance analysis...  ...and motivated Applied Scientist with machine learning engineering...  ..., from training to inference, including emerging LLM-based systems, that deliver...  ..., automation, and efficiency of large-scale training... 
    Local area
    Worldwide
    Flexible hours

    Amazon Science

    Palo Alto, CA
    6 hours ago
  • $192.2k - $260k

     ...Science team is seeking an experienced Applied Scientist who will join a team of experts in the...  ...of Amazon's data to help automate and optimize key processes - Design, development...  ...feature creations - Establish scalable, efficient, automated processes for large scale data... 
    Senior
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    1 day ago
  • $192k - $304.75k

    NVIDIA is searching for an outstanding Senior Researcher working on efficient deep learning to join our learning...  ...methods for post-training model optimization (pruning, quantization, NAS),...  ...architecture design, adaptive/dynamic inference, resource-efficient training and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $165k - $238k

     ...About the role:As a Senior Research Scientist you will be a key architect...  ...and high-fidelity inference over extremely large...  ...the gap between applied research and...  ...building blocks, such as LLM-as-a-judge evaluation...  ...state management, and optimization for multi-step... 
    Senior
    Full time
    Work at office
    3 days per week

    X Company

    Mountain View, CA
    1 day ago
  • $183.83k - $275.98k

     ...with state-of-the-art architectures quickly and efficiently, collaborating with other teams to determine data...  ...infrastructure support needs, and working to improve model optimization and inference speeds. You will use your applied research skills to think through the creation and... 
    Senior
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    1 day ago
  • $119.8k - $234.7k

     ...power ad ranking, pricing, and optimization across large-scale consumer...  ...heterogeneous event streams to infer user intent and advertiser...  ...marketplace dynamics. Engineers and scientists on the team work at the...  ...from you. #MicrosoftAI Applied Sciences IC4 - The typical... 
    Senior
    Ongoing contract
    Work at office
    Local area
    Shift work

    Microsoft Corporation

    Sunnyvale, CA
    1 day ago
  • $192.2k - $260k

     ...Description Amazon is seeking an exceptional Sr. Applied Scientist to lead the development of perception systems that harness...  ...DENSE, Astyx, RADDet, Boreas) Experience with real-time inference, model optimization (TensorRT, ONNX), and edge deployment Experience... 
    Senior
    Local area
    Flexible hours
    Night shift

    Amazon

    Sunnyvale, CA
    2 days ago
  • $192.2k - $260k

     ...advertising lifecycle from ad creation and optimization to performance analysis and customer...  ...and motivated Machine Learning Applied Scientist who loves to innovate at the intersection...  ...discover new products they love, be the most efficient way for advertisers to meet their... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Amazon

    Palo Alto, CA
    4 days ago
  • d-Matrix inc. is looking for a Senior Staff ML Researcher to join our Algo team in Santa Clara, CA. This hybrid position involves...  ...week. The successful candidate will develop algorithms for optimizing LLM inference on our DNN accelerators. Ideal applicants should have a MSc... 
    Senior
    3 days per week

    d-Matrix inc.

    Santa Clara, CA
    6 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Applied Scientist — Efficient LLM Inference & Optimization. Be the first to apply!