RESEARCHER, EFFICIENT INFERENCE
MLSys 2020
ABOUT THE COMPANY
We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-siteABOUT THE ROLE
You'll be researching making models efficient: quantization, speculative decoding, sparse and structured attention, distillation, mixture-of-experts inference, and the training-time techniques that make those methods possible. The work spans algorithm design, careful evaluation, and pushing methods to where they actually run. This is a senior research role with a clear engineering edge. You'll spend time at the intersection of model architecture and inference performance, designing methods that move accuracy/latency/cost trade-offs in our favor (then partnering with engineers to make those wins real in production).WHAT YOU'LL DO
Research and develop quantization methods: post-training quantization, quantization-aware training, mixed-precision regimes, low-bit-width arithmetic Design and evaluate speculative decoding approaches: draft models, tree attention, parallel speculation, lookahead decoding Investigate training-time efficiency methods that compose well with inference: distillation, sparse attention, mixture-of-experts, low-rank adaptation, pruning Run controlled experiments at production scale; characterize what works on real workloads, not just toy benchmarks Co‑design methods with the inference engineering team: push results to where they actually run, not stop at the paper Read deeply across the efficient ML / efficient inference literature; translate the most useful ideas into our stack Publish when the work warrants it; share findings internally Partner with model and training researchers so efficiency choices align with model architecture and post‑training decisionsWHAT WE'RE LOOKING FOR
Strong track record of ML research on efficiency methods: quantization, speculative decoding, distillation, MoE, sparse attention, or adjacent 5+ years of hands‑on research experience Deep familiarity with both training and inference performance characteristics Fluent in PyTorch, Jax or equivalent; comfortable working at the kernel and serving‑framework level when methods require it Track record of moving efficiency research from prototype to production Strong statistical expertise: you'd notice a flawed comparison before someone else points it out Strong written communication Published research at NeurIPS, ICML, ICLR, MLSys, or comparable venuesNICE TO HAVE
PhD in ML, systems, or related field Open‑source contributions to quantization, speculative‑decoding, or efficient‑inference libraries Experience with hardware‑aware optimization and accelerator‑specific tooling Background in numerical methods, low‑precision arithmetic, or approximate computationTHIS ROLE IS PROBABLY NOT FOR YOU IF
You want to focus on pretraining large models from scratch (that's a different role) You prefer abstract algorithmic research without hands‑on implementation You want a fixed benchmark with stable targets (our targets shift with what our models actually need to do) #J-18808-Ljbffr MLSys 2020Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the RESEARCHER, EFFICIENT INFERENCE in San Francisco, CA vacancy
- MakerMaker in San Francisco seeks a senior research engineer to advance efficiency in ML models, focusing on quantization, speculative decoding, and efficient... ...techniques. The role blends model architecture with inference performance, delivering production-ready improvements....Suggested
- MLSys 2020 in San Francisco is looking for a Senior Researcher specialized in machine learning efficiency. The role involves designing and researching methods for quantization, speculative decoding, and other efficiency techniques, ensuring that research translates into...Suggested
$84.13 - $91.34 per hour
AI Researcher - Efficient AI (Contractor) Step into the innovative world of LG Electronics. As a global leader in technology, LG Electronics... ...edge areas such as model compression, quantization, efficient inference, reasoning optimization, and next-generation AI...SuggestedFull timeContract workTemporary workFor contractorsLocal areaImmediate start- LG Electronics is seeking a Contract AI Researcher focusing on Efficient AI in Santa Clara, CA, hybrid work arrangement. You will explore model compression, quantization, efficient inference, and architectures to make LLMs/VLMs faster and more deployable on devices. You...SuggestedContract work
- A leading AI research company in San Francisco is seeking an AI Researcher to drive performance and quality optimizations of AI models. As part of the research team, you will explore new model architectures and experiment with techniques like KV caching and FlashAttention...Suggested
$216.3k - $280.8k
...Foundation AI, we are leading frontier AI research across Cisco. Our mission is to advance... ...algorithms, evaluation science, inference optimization, and AI systems infrastructure... ...AI, reinforcement learning, reasoning, efficient inference, or distributed training systems...Full timeTemporary workLocal areaFlexible hours- Real-time interactivity can come from inference‑time methods applied to an existing model,... ...generation of models is designed. Department: Research Location: San Francisco What You'll Do... ...diffusion, diffusion distillation, efficient attention or state‑space models for...Visa sponsorshipRelocation package
- We are seeking an Edge AI Research Scientist to develop next-generation speech and audio AI systems that run efficiently on smartphones, wearables, and other resource-constrained... ...compression techniques, and optimize inference pipelines that enable real-time speech AI...
- ...modern AI models better, faster and more efficient. Our Algorithm Discovery Platform brings... ...biology, computational neuroscience, AI research and software engineering to develop new... ...across fidelity, robustness, latency, and inference cost in real robotic settings Own major...
- ...customers almost immediately. No speculative research track here. If you want your work to... ...large-scale training to production inference serving millions of calls a day, working... ...neural audio codecs, compressing audio efficiently without losing quality Explore LLM-Audio...Permanent employmentFull timeImmediate start
$200k - $280k
The Turbo team sits at the intersection of efficient inference (algorithms, architectures, engines) and post-training / RL systems. We build... ...implementation in the engine and/or training stack. Have a solid research foundation in your area(s) of depth: Track record of...Full time$150k - $250k
...manufacturing, consumer goods, and global social organizations.We research and deploy technologies that power AI-native operations — both... ...trade-offs between generalization and specialization, data efficiency and robustness, capability and controllability. Their work...Work at office3 days per week- Modal is building an infrastructure layer for AI at scale in San Francisco. We are seeking a research-leaning engineer to own end-to-end inference research bets for LLM serving, including speculative decoding, quantization, and memory management. You will work with the...
$218.7k - $249.6k
...at Capital One to life. Our work touches every aspect of the research life cycle, from partnering with Academia to building production... ...delivering models at scale both in terms of training data and inference volumes. Experience in delivering libraries, platform level...Full timePart timeLocal areaFlexible hours$89.99k - $143.09k
...professional obligations.Track recurring editing issues and recommend solutions to improve overall billing quality, consistency, and efficiency.Support special projects related to billing compliance, client requirements, and invoicing process enhancements.Desired...Work at officeRemote workRelocationVisa sponsorshipRelocation package- ...at Capital One to life. Our work touches every aspect of the research life cycle, from partnering with academia to building production... ...delivering models at scale both in terms of training data and inference volumes. Experience in delivering libraries, platform‑level...Flexible hours
$150k
...– i.e. rigorously, proactively, and continuously fuzz-testing them. We are looking for Research Engineers to help develop our reliability platform, with a focus on: Data-efficient alignment of evaluation models Dynamic testing of AI applications Observability and anomaly...Visa sponsorship$262.5k - $299.6k
...Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview At Capital One, we are creating trustworthy and reliable AI systems... ...delivering models at scale both in terms of training data and inference volumes. Experience in delivering libraries, platform level...Full timePart timeLocal areaFlexible hours- ...you will work with a growing team comprised of quantitative researchers, software engineers, product managers, designers, and brokerage... ...care deeply about the trade-offs between tracking error, tax efficiency, and transaction costs, continuously researching improvements...Work at officeVisa sponsorshipFlexible hours
$262.5k - $299.6k
Applied Researcher II (AI Foundations) Overview: At Capital One, we are creating trustworthy and reliable AI systems, changing banking... ...delivering models at scale both in terms of training data and inference volumes. Experience in delivering libraries, platform level...Full timePart timeLocal areaFlexible hours$295k
...leads the Automated Red Teaming (ART) effort: building scalable, research-driven systems that continuously discover failure modes in our... ...- and driving alignment on what to fix first. Care about efficiency and prioritization, and you're happy to say "no" to low-leverage...- About the job Reinforcement Learning Researcher (Humanoid) Location: San Francisco, CA (On-site at REK HQ) Reports to: CTO Company: REK... ...simulation environments (Isaac Gym, MuJoCo, PyBullet, etc.) for efficient training and domain randomization. Sim-to-Real Transfer...
- ...products behave safely across these experiences. We develop the research, training methods, and evaluations needed to make these... ...layers, and modality fusion to cross-modal reasoning, scaling, and inference tradeoffs. Have improved frontier model behavior through post-...Work at officeRelocation package
- Postdoctoral Researcher, Computer Vision (PhD), New Grad Join to apply for the Postdoctoral Researcher, Computer Vision (PhD), New Grad... ...Referrals increase your chances of interviewing at Jobright.ai by 2x Inferred from the description for this job Medical insurance Vision...Full timePart timeWork experience placementInternship
- ...invented State Space Models or SSMs, a new primitive for training efficient, large-scale foundation models. Our team combines deep... ...generative audio models. This team is where customer needs meet research, and covers the full spectrum of modeling from ideation through...Work at officeVisa sponsorshipFlexible hours
$200k - $300k
Founding ML Researcher San Francisco In office Full-time We are hiring a Founding ML Researcher in San Francisco. We are building a... ...Publications in top‑tier AI conferences Familiarity with model serving, inference optimization, or deployment at scale Compensation: $200k - $3...Full timeWork at officeVisa sponsorship$200k - $300k
Unsiloed AI — Founding ML Researcher Type: Full-time | On-site | San Francisco, CA Compensation: $200,000-$300,000 + 0.1%-1% equity Hiring... ...parsing of unstructured data, and production model serving / inference optimization. Requirements Training and deploying state-of-...Full timeH1bWork at officeVisa sponsorshipFlexible hoursWeekend work- The Token Company is seeking an ML Researcher to own a slice of open problems in applied AI, focusing on what information in an LLM context matters and how to represent it efficiently. This high-autonomy role requires running many experiments, reproducing papers, and shipping...
- ...code is provably correct. About the role Join our team as an AI Research Engineer and help us push the boundaries of what's possible in... ...Combine Reasoning algorithm and LLMs Build effective and efficient ML pipelines Collaborate with other teams to understand their...Contract work
- AI Researcher (Computer Vision/Multimodal/Generative AI) About the Role We are hiring ML Researchers to develop novel approaches that... ...and training strategies that improve realism, controllability, efficiency, and multimodal understanding — with a direct path from research...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to RESEARCHER, EFFICIENT INFERENCE. Be the first to apply!
Related searches
- senior researcher San Francisco, CA
- machine learning researcher San Francisco, CA
- researcher San Francisco, CA
- senior design researcher San Francisco, CA
- design researcher San Francisco, CA
- qualitative researcher San Francisco, CA
- data collection researcher San Francisco, CA
- product researcher San Francisco, CA
- survey researcher San Francisco, CA
- legal researcher San Francisco, CA

