Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior ML Researcher: Efficient Inference & Quantization

MakerMaker

MakerMaker in San Francisco seeks a senior research engineer to advance efficiency in ML models, focusing on quantization, speculative decoding, and efficient training-time techniques. The role blends model architecture with inference performance, delivering production-ready improvements. You will run large-scale experiments, co-design with inference engineers, and push findings to production while publishing where appropriate. Collaboration with researchers is essential. #J-18808-Ljbffr MakerMaker

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior ML Researcher: Efficient Inference & Quantization in San Francisco, CA vacancy
  • MLSys 2020 in San Francisco is looking for a Senior Researcher specialized in machine learning efficiency. The role involves designing and researching methods for quantization, speculative decoding, and other efficiency techniques, ensuring that research translates into... 
    Senior

    MLSys 2020

    San Francisco, CA
    2 days ago
  •  ...building autonomous research agents for recursive...  ...researching making models efficient: quantization, speculative decoding...  ..., mixture-of-experts inference, and the training-...  ...actually run. This is a senior research role with a...  ...across the efficient ML / efficient inference... 
    Suggested
    Shift work

    MakerMaker

    San Francisco, CA
    1 day ago
  • $84.13 - $91.34 per hour

    AI Researcher - Efficient AI (Contractor) Step into the innovative world of LG Electronics. As...  ...areas such as model compression, quantization, efficient inference, reasoning optimization, and next-generation...  ...or engineering experience in ML, efficient AI, model optimization,... 
    Suggested
    Full time
    Contract work
    Temporary work
    For contractors
    Local area
    Immediate start

    LG Electronics

    San Francisco, CA
    3 days ago
  • A leading AI research company in San Francisco is seeking an AI Researcher to drive performance and quality optimizations of AI models. As...  ...a Master’s or PhD in a relevant field and has experience in AI/ML, as well as familiarity with frameworks like PyTorch and TensorFlow... 
    Suggested

    Gimlet Labs

    San Francisco, CA
    2 days ago
  • $216.3k - $280.8k

     ..., we are leading frontier AI research across Cisco. Our mission is...  ...algorithms, evaluation science, inference optimization, and AI systems...  ...problems.Your ImpactAs a Senior AI Researcher you will operate...  ...reinforcement learning, reasoning, efficient inference, or distributed... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Francisco, CA
    4 days ago
  •  ...models better, faster and more efficient. Our Algorithm Discovery...  ...computational neuroscience, AI research and software engineering to develop...  ...imagined futures. As a Senior AI Researcher, you will own significant...  ..., robustness, latency, and inference cost in real robotic settings... 
    Senior

    The Biological Computing Co.

    San Francisco, CA
    1 day ago
  • $200k - $280k

     ...sits at the intersection of efficient inference (algorithms, architectures,...  ...high-performance computing for ML. Are comfortable working...  ...training stack. Have a solid research foundation in your area(s) of...  ...speculative decoding (e.g., ATLAS), quantization, etc. Profile and optimize... 
    Full time

    Together

    San Francisco, CA
    2 days ago
  • LG Electronics is seeking a Contract AI Researcher focusing on Efficient AI in Santa Clara, CA, hybrid work arrangement. You will explore model compression, quantization, efficient inference, and architectures to make LLMs/VLMs faster and more deployable on devices. You... 
    Contract work

    LG Electronics

    San Francisco, CA
    3 days ago
  • $250k

     ...that powers next-generation research and cloud platforms?...  ...now building a serverless inference platform, beginning with cost-efficient batch inference and expanding...  ...chance to join as a Senior Inference Platform Engineer...  ...tolerant distributed systems (ML inference, HPC, or... 
    Senior
    Full time
    San Francisco, CA
    more than 2 months ago
  •  ...the grid operates efficiently. The company is backed...  ...We are seeking a Senior Applied Scientist, On‑Device ML to design models...  ...blends applied research, model optimization...  ...resource‑constrained inference, and efficient...  ...with edge ML tools, quantization, model compression... 
    Senior
    Local area

    Gridware

    San Francisco, CA
    7 hours ago
  • $200.9k - $257.5k

     ...scale, to benefit the research. Pre‑train and fine‑...  ...Lead the design of efficient data loading strategies...  ...innovative AI/ML research at top‑tier...  ...optimizing large‑scale inference via quantization, distillation, or memory...  ...$226,200 - $290,000 Senior Scientist I, Machine... 
    Senior
    Contract work
    Work experience placement
    Local area

    Altoslabs

    San Francisco, CA
    3 days ago
  • $204k - $300k

     ...The Advanced Technology Group (ATG) is the research division of the company. ATG’s mission is...  ...and electrical engineering, such as AI/ML, algorithms, digital signal processing, audio...  ..., and IoT.What You Will AccomplishAs a senior research leader in the Multimodal... 
    Senior
    Full time
    Local area
    Worldwide
    Flexible hours

    Dolby

    San Francisco, CA
    3 days ago
  • $262.5k - $299.6k

     ...Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview...  ..., our applications of AI & ML are bringing humanity and...  ...in terms of training data and inference volumes. Experience in delivering...  ...: Model Sparsification, Quantization, Training Parallelism/Partitioning... 
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    17 hours ago
  •  ...time, our applications of AI & ML bring humanity and simplicity...  ...touches every aspect of the research life cycle, from partnering...  ...in terms of training data and inference volumes. Experience in delivering...  ...model sparsification, quantization, training parallelism, checkpointing... 
    Flexible hours

    Capital One

    San Francisco, CA
    3 days ago
  •  ...a Large Physics foundation Model focused on weather with a mission to enable verifiable cause and effect in AI systems. We seek researchers to tackle unsolved problems across language, vision, robotics, biology, physics, and weather, training ground-truth grounded models... 
    Senior

    causal

    San Francisco, CA
    1 day ago
  • University of California, San Francisco is looking for a postdoctoral researcher/data scientist. This position offers 100% funding for one year,...  ...will have a PhD in a relevant field, strong experience in ML algorithms, and an interest in analyzing Electronic Health Record... 
    Senior

    University of California , San Francisco

    San Francisco, CA
    1 day ago
  •  ...forefront of applying artificial intelligence and machine learning (AI/ML) to help improve outcomes in vulnerable and underserved...  ...and clinical notes from the electronic health record (EHR) system Research and assess the use of large language models to develop interpretable... 
    Senior
    Work experience placement
    Work at office
    Immediate start

    University of California , San Francisco

    San Francisco, CA
    17 hours ago
  •  ..., and accelerate delivery across the company. We also build the research platform that powers AI at scale and explore next-generation GenAI...  ...-technical audiences. Preferred Qualifications 7+ years in AI/ML research and development using Python. Familiarity with regulatory... 
    Senior
    Work at office

    Charles Schwab

    San Francisco, CA
    2 days ago
  •  ...immediately. No speculative research track here. If you want your...  ...scale training to production inference serving millions of calls a day...  ...audio codecs, compressing audio efficiently without losing quality...  ...direction, tooling, and (for senior hires) the team itself Work... 
    Permanent employment
    Full time
    Immediate start

    DeepRec.ai

    San Francisco, CA
    1 day ago
  • We are seeking an Edge AI Research Scientist to develop next-generation...  ...audio AI systems that run efficiently on smartphones, wearables,...  ...compression techniques, and optimize inference pipelines that enable real-...  ...frameworks such as: Core ML ExecuTorch ONNX Runtime LiteRT... 

    Huxley

    San Francisco, CA
    1 day ago
  •  ...infrastructure, influencing latency, throughput, and reliability of RL and training loops. You will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies, batching, and long-context workloads while collaborating with #J-18808-... 
    Senior

    Magic AI Corp.

    San Francisco, CA
    3 days ago
  • MakerMaker in San Francisco is seeking a Senior ML systems engineer to build and operate production inference systems for large models. You will own performance, profiling...  ...in production. You will collaborate with researchers to implement inference optimizations, design... 
    Senior

    MakerMaker

    San Francisco, CA
    1 day ago
  • $180k - $200k

     ...ensure the grid operates efficiently. The company is backed...  ...work closely with ML scientists and firmware...  ...algorithms and build ML inference pipelines into...  ...develop with algorithm / ML researchers to refine models for embedded...  ...targets (e.g., quantization, fixed-point, pruning)... 
    Senior
    Full time

    Gridware

    San Francisco, CA
    17 hours ago
  • $220k

     ...team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and...  ...3+ years of experience in software engineering with a focus on ML inference, familiarity with deep learning frameworks, and a strong... 
    Senior

    Perplexity

    San Francisco, CA
    2 days ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal candidate will have 3+ years of experience in production... 
    Senior

    MakerMaker.AI

    San Francisco, CA
    7 hours ago
  •  ...specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern inference frameworks and a solid... 
    Senior

    Reflection AI

    San Francisco, CA
    3 days ago
  •  ...workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and...  ...have strong software engineering skills and experience with ML inference systems, particularly in Python and C++. This position... 
    Senior

    Gimlet Labs

    San Francisco, CA
    1 day ago
  •  ...Inc. is seeking a Software Engineer to lead the evolution of its ML Platform within the Infrastructure team. You will build scalable...  ...systems for model experimentation, training, evaluation, inference, and retraining powering underwriting and ML-driven products. You... 
    Senior

    Parafin, Inc.

    San Francisco, CA
    1 day ago
  • $300k

     ...is a quickly growing group of committed researchers, engineers, policy experts, and...  ...systems. About the Role The Cloud Inference team scales and optimizes Claude to serve...  ..., the complexity of managing inference efficiently across providers with different hardware... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    17 hours ago
  •  ...Function of PositionAs a Senior Systems GPU Engineer -...  ...You will work alongside research, SW/ HW/ ML engineering, regulatory,...  ...models to real-time onboard inference—while serving as a core...  ...distillation, parameter-efficient fine-tuning and quantization and hardware... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior ML Researcher: Efficient Inference & Quantization. Be the first to apply!