Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior ML Scientist - Inference & Hardware Acceleration

Netskope

Netskope is seeking a Senior Staff Machine Learning Scientist in Santa Clara to own the inference and optimization layer for AI in agentic workflows. You will fine-tune models, push latency and throughput on real hardware, and build a runtime that executes bounded AI tasks with real customer data signals.

You will work on quantization, KV-cache optimization, and hardware acceleration, partnering with systems and backend engineers to ship end-to-end capabilities in production environments.

#J-18808-Ljbffr
Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Senior ML Scientist - Inference & Hardware Acceleration in Santa Clara, CA vacancy
  •  ...ByteDance Technology in San Jose is seeking a Senior Research Scientist to lead machine learning system development. You will help build and optimize large-scale distributed ML training and inference infrastructure, integrating GPU/NPU/RDMA and storage to run complex models... 
    Senior

    ByteDance

    San Jose, CA
    9 hours ago
  • $182.5k - $260.5k

     ...era. We secure and accelerate cloud, data, and AI...  ...Positions are available at Senior Staff and above....  ...Machine Learning Scientist, you own the inference and optimization layer...  ...throughput on real hardware, and build the runtime...  ...+ years hands-on in ML/AI (model... 
    Senior

    Netskope

    Santa Clara, CA
    3 days ago
  • $148k - $235.75k

    We are looking for a Senior Technical Product Marketing Manager....  ...business and pivotal in our inference marketing. You will be focused...  ...years of experience in LLM, AI/ML development in an engineering...  ...data center architectures, accelerated computing, distributed inference... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $195.2k - $262.2k

     ...building large in-house AI/ML infrastructure. Built...  ...GPU orchestration to inference optimization, we own...  ...deep expertise across hardware, software and AI R&D....  ...Nebius Token Factory needs scientists who can turn frontier inference...  ...capabilities. A Senior Applied Scientist owns... 
    Senior
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    6 days ago
  • $192.2k - $260k

    We are looking for a Senior Applied Scientist to help drive the research and development...  ..., and work closely with inference engineers to ensure your...  ...architectures informed by hardware constraints and inference...  ...to open-source speech/audio ML systems or widely used... 
    Senior
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    4 days ago
  • d-Matrix in Santa Clara, CA is seeking a Sr. Staff ML Researcher to advance LLM algorithmic optimization on our DNN accelerators. You will design and implement efficient inference algorithms, collaborating with mathematicians, ML researchers and engineers on high-impact... 
    Senior
    3 days per week

    Entrada Ventures

    Santa Clara, CA
    2 days ago
  • Voltai in California (Palo Alto) seeks a senior formal verification researcher to develop new methods for formal proofs of...  ...correctness. You will collaborate with RTL, verification, and ML teams to scale AI-hardware verification, prototype ideas on real RTL, and turn... 
    Senior

    Voltai

    Palo Alto, CA
    4 days ago
  •  ...A leading technology firm in Cupertino is seeking a Sr. Hardware Engineer to design and validate next-generation ML Chips. This role involves leading PCIe designs, collaborating with teams to enhance product performance, and applying innovative technologies. Ideal candidates... 
    Senior

    Amazon

    Cupertino, CA
    9 hours ago
  • $212.8k - $387.6k

     ...Senior Research Scientist - Machine Learning System Location: San Jose Team...  ...The Machine Learning (ML) System sub-team combines system...  ...ML training and Inference system/services around the...  ...LLM models, experience in accelerating LLM model optimization is preferred... 
    Senior
    Temporary work
    Local area

    ByteDance

    San Jose, CA
    9 hours ago
  • Nebius Token Factory seeks a PhD-level researcher to lead focused ML research projects from hypothesis to production handoff. You...  ...to make prototypes production-ready and drive efficient LLM/VLM inference with measurable impact. The role emphasizes publishing results,... 
    Senior

    Nebius B.V.

    Palo Alto, CA
    4 days ago
  •  ...industry-leading training and inference speeds; over 10 times faster than...  ...running on Cerebras hardware.Test and verify deployment infrastructure...  ...tools.Experience with ML inference infrastructure, model serving systems, or GPU-accelerated workloadsLocation: Toronto / SunnyvaleTeam... 
    Senior
    Work at office

    Cerebras Systems

    Sunnyvale, CA
    5 days ago
  • $332k

     ...graphics, PC gaming, and accelerated computing for more...  ....We are looking for a Senior leader to orchestrate...  ...NVIDIA platform, including hardware and software. "Embed...  ...a leadership Inference go-to-market strategy!...  ...pre-sales with an AI/ML focus.Passion for transforming... 
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...build great products that accelerate next-generation computing...  ...ROLE:AMD is seeking a Senior Product Manager to drive...  ...focus on large-scale model inference on AMD Instinct™ and Radeon™ hardware. This is a key, central role...  ...in the open-source AI/ML community: monitor GitHub... 
    Senior
    Remote work

    AMD

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

     ...engineers to join us and build AI inference systems that serve large-...  ...to push the frontier of accelerated computing for AI.What you’ll...  ...models with the latest NVIDIA GPU hardware features; profile and optimize...  ...frontier for the field of ML Systems; survey recent publications... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $193.3k - $261.5k

     ...machine learning training and inference clusters. Our organization...  ...life — drivers that expose the hardware to the OS, runtime libraries...  ...infrastructure to enable SoC validation, accelerate system software development,...  ...exploration. As part of the ML accelerator systems modeling... 
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    4 days ago
  • $192k - $304.75k

    We are now looking for a Senior Research Scientist for Human‑AI Perception & Interaction! NVIDIA has...  ...transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’...  ...multi-node, multi-GPU training and inference workflows.Familiarity with and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $192k - $304.75k

     ...problems with our unique approach to accelerated computing. We're looking for a passionate scientist at the intersection of quantum...  ...for fault-tolerant quantum hardware.At NVIDIA, we want to help...  ...performance prediction and parameter inference without full experimental... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $148.75k - $361k

     ...At the core of this is our Machine Learning, Experimentation and Inference Platform that powers the entire landscape which we continuously...  ...QualificationsExperience in the Advertising domainContributions to open-source ML projects #LI-DH2What's Roku's approach to hybrid working?Roku... 
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $183k - $247.6k

     ...designs silicon and software that accelerates innovation. Customers choose us...  ...the world.We are seeking a Hardware Design Engineer with role in the...  ...validation of AWS next generation ML Chips, Cards and server integration. As a senior member of our hardware team, you... 
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $192k - $304.75k

     ...our unique approach to accelerated computing. We're looking for a passionate scientist at the intersection of...  ...fault-tolerant quantum hardware.At NVIDIA, we want to help...  ...and parameter inference without full experimental...  ...quantum systems and AI/ML research.Hands-on expertise... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $192.2k - $260k

     ...talented, and inventive Applied Scientist with a strong machine...  ...scale computing resources to accelerate advances in machine learning...  ...Join our dynamic team of AI/ML practitioners and applied scientists...  ....Key job responsibilitiesThe Senior Applied Scientist will lead the... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    3 days ago
  • $173k

     ...build for travelers everywhere.Senior Machine Learning ScientistThe Senior Machine Learning Scientist is responsible for building and...  ...trip management. Owns end-to-end ML and GenAI projects—from problem...  ...technical challenges, from inference problems on long-tail traveler... 
    Senior
    Full time
    Worldwide

    Expedia

    San Jose, CA
    3 days ago
  •  ...to build great products that accelerate next-generation computing experiences...  ...seeking a Principal GenAI Inference Optimization Engineer to join...  ...working across the software-hardware stack.THE PERSONThe ideal...  ...architectures.- Experience with ML frameworks (PyTorch, JAX, or... 

    AMD

    San Jose, CA
    1 day ago
  •  ...to build great products that accelerate next-generation computing experiences...  ...Language Models (LLMs) and ML workloads on emerging...  ...infrastructure, and model-to-hardware optimization, with a strong focus...  ...movement optimization for ML inference workloads• Define and... 

    AMD

    San Jose, CA
    1 day ago
  •  ...industry-leading training and inference speeds and allows users to run large-scale ML applications with less hardware management. Cerebras’...  ...applications. About The Role Senior Director of Technical...  ...a focus on infrastructure (accelerators, clusters, interconnect,... 
    Senior

    Cerebras

    Sunnyvale, CA
    9 hours ago
  • $193.3k - $261.5k

     ...development kit used to accelerate deep learning and...  ...and Trainium ML accelerators. This...  ...enabling unparalleled ML inference and training...  ...PyTorch till the hardware-software boundary,...  ...functional team of applied scientists, system engineers,...  ...mentorship. Our senior members enjoy one-... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    19 hours ago
  •  ...Sr. Hardware Engineer - ML Acceleration, Annapurna Labs AWS Utility Computing (UC) provides product innovations — from foundational services such...  ...generation ML Chips, Cards and server integration. As a senior member of our hardware team, you will participate in the... 
    Senior
    Work from home

    Amazon

    Cupertino, CA
    9 hours ago
  • $192k - $304.75k

     ...problems with our unique approach to accelerated computing. We're looking for a passionate AI research scientist with deep quantum computing...  ...error-correcting codes and hardware platforms, while collaborating...  ...doing:Design and architect AI/ML models—including deep neural... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...-leading training and inference speeds; over 10 times...  ...This is one of the most senior IC roles on the team,...  ...partner closely with ML, Product and Infrastructure...  ...systems, or GPU-accelerated workloads is a plus. Why...  ...software make their own hardware. At Cerebras, we have... 

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $192.2k - $260k

     ...AWS) is assembling an elite team of world-class scientists and engineers to pioneer the next generation of...  ...and present your pioneering work at premier ML and NLP conferences (NeurIPS, ICML, ICLR , ACL, EMNLP)- Accelerate innovation by working directly with customers to... 
    Senior
    Work at office
    Local area
    Worldwide
    Flexible hours

    Amazon

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior ML Scientist - Inference & Hardware Acceleration. Be the first to apply!