Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

RESEARCHER, EFFICIENT INFERENCE

MLSys 2020

ABOUT THE COMPANY

We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site

ABOUT THE ROLE

You'll be researching making models efficient: quantization, speculative decoding, sparse and structured attention, distillation, mixture-of-experts inference, and the training-time techniques that make those methods possible. The work spans algorithm design, careful evaluation, and pushing methods to where they actually run. This is a senior research role with a clear engineering edge. You'll spend time at the intersection of model architecture and inference performance, designing methods that move accuracy/latency/cost trade-offs in our favor (then partnering with engineers to make those wins real in production).

WHAT YOU'LL DO

Research and develop quantization methods: post-training quantization, quantization-aware training, mixed-precision regimes, low-bit-width arithmetic Design and evaluate speculative decoding approaches: draft models, tree attention, parallel speculation, lookahead decoding Investigate training-time efficiency methods that compose well with inference: distillation, sparse attention, mixture-of-experts, low-rank adaptation, pruning Run controlled experiments at production scale; characterize what works on real workloads, not just toy benchmarks Co‑design methods with the inference engineering team: push results to where they actually run, not stop at the paper Read deeply across the efficient ML / efficient inference literature; translate the most useful ideas into our stack Publish when the work warrants it; share findings internally Partner with model and training researchers so efficiency choices align with model architecture and post‑training decisions

WHAT WE'RE LOOKING FOR

Strong track record of ML research on efficiency methods: quantization, speculative decoding, distillation, MoE, sparse attention, or adjacent 5+ years of hands‑on research experience Deep familiarity with both training and inference performance characteristics Fluent in PyTorch, Jax or equivalent; comfortable working at the kernel and serving‑framework level when methods require it Track record of moving efficiency research from prototype to production Strong statistical expertise: you'd notice a flawed comparison before someone else points it out Strong written communication Published research at NeurIPS, ICML, ICLR, MLSys, or comparable venues

NICE TO HAVE

PhD in ML, systems, or related field Open‑source contributions to quantization, speculative‑decoding, or efficient‑inference libraries Experience with hardware‑aware optimization and accelerator‑specific tooling Background in numerical methods, low‑precision arithmetic, or approximate computation

THIS ROLE IS PROBABLY NOT FOR YOU IF

You want to focus on pretraining large models from scratch (that's a different role) You prefer abstract algorithmic research without hands‑on implementation You want a fixed benchmark with stable targets (our targets shift with what our models actually need to do) #J-18808-Ljbffr MLSys 2020

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the RESEARCHER, EFFICIENT INFERENCE in San Francisco, CA vacancy
  • MakerMaker in San Francisco seeks a senior research engineer to advance efficiency in ML models, focusing on quantization, speculative decoding, and efficient...  ...techniques. The role blends model architecture with inference performance, delivering production-ready improvements.... 
    Suggested

    MakerMaker

    San Francisco, CA
    1 day ago
  • MLSys 2020 in San Francisco is looking for a Senior Researcher specialized in machine learning efficiency. The role involves designing and researching methods for quantization, speculative decoding, and other efficiency techniques, ensuring that research translates into... 
    Suggested

    MLSys 2020

    San Francisco, CA
    2 days ago
  • $84.13 - $91.34 per hour

    AI Researcher - Efficient AI (Contractor) Step into the innovative world of LG Electronics. As a global leader in technology, LG Electronics...  ...edge areas such as model compression, quantization, efficient inference, reasoning optimization, and next-generation AI... 
    Suggested
    Full time
    Contract work
    Temporary work
    For contractors
    Local area
    Immediate start

    LG Electronics

    San Francisco, CA
    3 days ago
  • LG Electronics is seeking a Contract AI Researcher focusing on Efficient AI in Santa Clara, CA, hybrid work arrangement. You will explore model compression, quantization, efficient inference, and architectures to make LLMs/VLMs faster and more deployable on devices. You... 
    Suggested
    Contract work

    LG Electronics

    San Francisco, CA
    3 days ago
  • A leading AI research company in San Francisco is seeking an AI Researcher to drive performance and quality optimizations of AI models. As part of the research team, you will explore new model architectures and experiment with techniques like KV caching and FlashAttention... 
    Suggested

    Gimlet Labs

    San Francisco, CA
    2 days ago
  • $216.3k - $280.8k

     ...Foundation AI, we are leading frontier AI research across Cisco. Our mission is to advance...  ...algorithms, evaluation science, inference optimization, and AI systems infrastructure...  ...AI, reinforcement learning, reasoning, efficient inference, or distributed training systems... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Francisco, CA
    4 days ago
  • Real-time interactivity can come from inference‑time methods applied to an existing model,...  ...generation of models is designed. Department: Research Location: San Francisco What You'll Do...  ...diffusion, diffusion distillation, efficient attention or state‑space models for... 
    Visa sponsorship
    Relocation package

    Reactor

    San Francisco, CA
    1 day ago
  • We are seeking an Edge AI Research Scientist to develop next-generation speech and audio AI systems that run efficiently on smartphones, wearables, and other resource-constrained...  ...compression techniques, and optimize inference pipelines that enable real-time speech AI... 

    Huxley

    San Francisco, CA
    1 day ago
  •  ...modern AI models better, faster and more efficient. Our Algorithm Discovery Platform brings...  ...biology, computational neuroscience, AI research and software engineering to develop new...  ...across fidelity, robustness, latency, and inference cost in real robotic settings Own major... 

    The Biological Computing Co.

    San Francisco, CA
    1 day ago
  •  ...customers almost immediately. No speculative research track here. If you want your work to...  ...large-scale training to production inference serving millions of calls a day, working...  ...neural audio codecs, compressing audio efficiently without losing quality Explore LLM-Audio... 
    Permanent employment
    Full time
    Immediate start

    DeepRec.ai

    San Francisco, CA
    1 day ago
  • $200k - $280k

    The Turbo team sits at the intersection of efficient inference (algorithms, architectures, engines) and post-training / RL systems. We build...  ...implementation in the engine and/or training stack. Have a solid research foundation in your area(s) of depth: Track record of... 
    Full time

    Together

    San Francisco, CA
    2 days ago
  • $150k - $250k

     ...manufacturing, consumer goods, and global social organizations.We research and deploy technologies that power AI-native operations — both...  ...trade-offs between generalization and specialization, data efficiency and robustness, capability and controllability. Their work... 
    Work at office
    3 days per week

    Distyl AI

    San Francisco, CA
    1 day ago
  • Modal is building an infrastructure layer for AI at scale in San Francisco. We are seeking a research-leaning engineer to own end-to-end inference research bets for LLM serving, including speculative decoding, quantization, and memory management. You will work with the... 

    Modal

    San Francisco, CA
    1 day ago
  • $218.7k - $249.6k

     ...at Capital One to life. Our work touches every aspect of the research life cycle, from partnering with Academia to building production...  ...delivering models at scale both in terms of training data and inference volumes. Experience in delivering libraries, platform level... 
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    2 days ago
  • $89.99k - $143.09k

     ...professional obligations.Track recurring editing issues and recommend solutions to improve overall billing quality, consistency, and efficiency.Support special projects related to billing compliance, client requirements, and invoicing process enhancements.Desired... 
    Work at office
    Remote work
    Relocation
    Visa sponsorship
    Relocation package

    Minnesota Jobs

    San Francisco, CA
    2 hours ago
  •  ...at Capital One to life. Our work touches every aspect of the research life cycle, from partnering with academia to building production...  ...delivering models at scale both in terms of training data and inference volumes. Experience in delivering libraries, platform‑level... 
    Flexible hours

    Capital One

    San Francisco, CA
    3 days ago
  • $150k

     ...– i.e. rigorously, proactively, and continuously fuzz-testing them. We are looking for Research Engineers to help develop our reliability platform, with a focus on: Data-efficient alignment of evaluation models Dynamic testing of AI applications Observability and anomaly... 
    Visa sponsorship

    Enboarder

    San Francisco, CA
    2 days ago
  • $262.5k - $299.6k

     ...Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview At Capital One, we are creating trustworthy and reliable AI systems...  ...delivering models at scale both in terms of training data and inference volumes. Experience in delivering libraries, platform level... 
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    13 hours ago
  •  ...you will work with a growing team comprised of quantitative researchers, software engineers, product managers, designers, and brokerage...  ...care deeply about the trade-offs between tracking error, tax efficiency, and transaction costs, continuously researching improvements... 
    Work at office
    Visa sponsorship
    Flexible hours

    Frec

    San Francisco, CA
    13 hours ago
  • $262.5k - $299.6k

    Applied Researcher II (AI Foundations) Overview: At Capital One, we are creating trustworthy and reliable AI systems, changing banking...  ...delivering models at scale both in terms of training data and inference volumes. Experience in delivering libraries, platform level... 
    Full time
    Part time
    Local area
    Flexible hours

    Capital One

    San Francisco, CA
    2 days ago
  • $295k

     ...leads the Automated Red Teaming (ART) effort: building scalable, research-driven systems that continuously discover failure modes in our...  ...- and driving alignment on what to fix first. Care about efficiency and prioritization, and you're happy to say "no" to low-leverage... 

    OpenAI

    San Francisco, CA
    3 days ago
  • About the job Reinforcement Learning Researcher (Humanoid) Location: San Francisco, CA (On-site at REK HQ) Reports to: CTO Company: REK...  ...simulation environments (Isaac Gym, MuJoCo, PyBullet, etc.) for efficient training and domain randomization. Sim-to-Real Transfer... 

    However, REK Inc

    San Francisco, CA
    2 days ago
  •  ...products behave safely across these experiences. We develop the research, training methods, and evaluations needed to make these...  ...layers, and modality fusion to cross-modal reasoning, scaling, and inference tradeoffs. Have improved frontier model behavior through post-... 
    Work at office
    Relocation package

    United States Digital Space LLC

    San Francisco, CA
    3 days ago
  • Postdoctoral Researcher, Computer Vision (PhD), New Grad Join to apply for the Postdoctoral Researcher, Computer Vision (PhD), New Grad...  ...Referrals increase your chances of interviewing at Jobright.ai by 2x Inferred from the description for this job Medical insurance Vision... 
    Full time
    Part time
    Work experience placement
    Internship

    Jobright.ai

    San Francisco, CA
    3 days ago
  •  ...invented State Space Models or SSMs, a new primitive for training efficient, large-scale foundation models. Our team combines deep...  ...generative audio models. This team is where customer needs meet research, and covers the full spectrum of modeling from ideation through... 
    Work at office
    Visa sponsorship
    Flexible hours

    Cartesia

    San Francisco, CA
    2 days ago
  • $200k - $300k

    Founding ML Researcher San Francisco In office Full-time We are hiring a Founding ML Researcher in San Francisco. We are building a...  ...Publications in top‑tier AI conferences Familiarity with model serving, inference optimization, or deployment at scale Compensation: $200k - $3... 
    Full time
    Work at office
    Visa sponsorship

    PassFort

    San Francisco, CA
    2 days ago
  • $200k - $300k

    Unsiloed AI — Founding ML Researcher Type: Full-time | On-site | San Francisco, CA Compensation: $200,000-$300,000 + 0.1%-1% equity Hiring...  ...parsing of unstructured data, and production model serving / inference optimization. Requirements Training and deploying state-of-... 
    Full time
    H1b
    Work at office
    Visa sponsorship
    Flexible hours
    Weekend work

    davidjoseph-co

    San Francisco, CA
    3 hours ago
  • The Token Company is seeking an ML Researcher to own a slice of open problems in applied AI, focusing on what information in an LLM context matters and how to represent it efficiently. This high-autonomy role requires running many experiments, reproducing papers, and shipping... 

    davidjoseph-co

    San Francisco, CA
    13 hours ago
  •  ...code is provably correct. About the role Join our team as an AI Research Engineer and help us push the boundaries of what's possible in...  ...Combine Reasoning algorithm and LLMs Build effective and efficient ML pipelines Collaborate with other teams to understand their... 
    Contract work

    Logical Intelligence

    San Francisco, CA
    3 days ago
  • AI Researcher (Computer Vision/Multimodal/Generative AI) About the Role We are hiring ML Researchers to develop novel approaches that...  ...and training strategies that improve realism, controllability, efficiency, and multimodal understanding — with a direct path from research... 

    SpreeAI

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to RESEARCHER, EFFICIENT INFERENCE. Be the first to apply!