Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Research Scientist, Learning & Evaluation

Studyfetch

AI Research Scientist, Learning & Evaluation Studyfetch Beverly Hills, California, United States About this position About Studyfetch StudyFetch is the #1 AI-native learning platform globally, transforming how millions of students learn through personalized AI-powered education. We’re growing fast with backing from top-tier investors and a mission that’s redefining the future of education and ethical learning. Why this role exists We're a technology company building AI-native learning products used by more than seven million students worldwide, alongside Honen, our workforce-learning platform for organizations. Both run on the Learn Engine, the intelligence that moves a learner from initial understanding to demonstrated mastery. We work with partners like NVIDIA to bring responsible, learning-first AI to the students who need it most. Nobody has settled how to measure whether an AI tutor teaches. Public benchmarks tell you a model can answer a question. They don't tell you whether a fourteen-year-old understood the explanation, whether the model handed over the answer when it should have asked a follow-up, or whether the student could still do the problem a week later. We train and tune our own models for learning outcomes rather than leaderboard scores, and that only works if someone can define what better means and prove when we've hit it. That's this role. You'll build the evaluation and measurement layer that sits under model training, product decisions, and learning science across both StudyFetch and Honen. This is a founding-team role. You'll work directly with the people making the decisions, and the standard you set is the one every model and feature gets held to. What we believe Every learner deserves the chance to succeed. StudyFetch started with one idea: high-quality, personalized learning should be within reach for anyone, at any stage of life. Honen carries that belief into the workforce. Accessible to everyone. Meet people where they are. Learning never stops. We hire people who share this conviction. The work is demanding and the hours can be long, and what sustains you through it is caring whether a real student finally understands the material. What you'll own Evaluation for the models we train. We fine-tune our own model family for tutoring, and you decide how we know whether a new checkpoint is better than the last one. That covers accuracy and reasoning, and it also covers child safety, resistance to sycophancy, and whether the model teaches Socratically instead of answering outright. You'll design the evals, run them against every candidate model, and hold the release bar. An internal benchmark for multi-turn tutoring. Single-turn Q&A benchmarks miss almost everything we care about. You'll build and maintain our benchmark for real tutoring conversations, decide what it measures, and defend those choices to researchers outside the company. Expect to publish parts of it. The link between product data and model training. The Learn Engine records what worked for past students at the personal, course, topic, and global level. You'll turn that into training signal and into evidence: which interventions moved mastery, which ones only moved engagement, and which model behaviors correlate with a student actually learning. Data quality for expert-verified content. We build assessment questions with subject-matter experts, starting in nursing licensure and expanding into medical and legal. You'll measure agreement between experts, catch where the official answer key is out of date, and design how those verifications feed back into training. Analytics across both products. Retention, activation, feature adoption, conversion, and how each of those moves when a model changes. You'll build the dashboards and reporting that product and leadership actually use, and you'll say plainly when the data can't answer the question yet. Instrumentation we don't have. You'll find the missing events, telemetry, and logging, then work with engineering to add them. Most of the interesting questions here are currently unanswerable because nobody logged the right thing. The measurement bar for the team. How we run experiments, what counts as a result, when a change ships. The patterns you set are the ones the rest of the team follows. What we're looking for You're a strong fit if either of these is true: A PhD in statistics, computer science, machine learning, economics, physics, computational social science, or another quantitative field, plus 5+ years applying it to real products or research, or Fewer credentials on paper and a track record of owning evaluation or measurement for an AI product that shipped to real users. Show us the work. Beyond that: You've evaluated LLMs in production, not just read about it. You're rigorous about causality. You write and speak clearly. You use AI every day and have informed opinions about it. The mission is why you're here. The stack you’ll work in You don't need every item below, but you should be deep in most and able to ramp quickly on the rest: Statistics: experimental design, causal inference, Bayesian and frequentist methods, hypothesis testing AI/LLM: eval frameworks, LLM-as-judge and its failure modes, RAG, embeddings, agent workflows, fine-tuning and post-training MongoDB and PostgreSQL, vector databases, warehouse and pipeline tooling Infra: GCP, and comfort reasoning about inference cost, latency, throughput, and GPU utilization Reporting: dashboards people return to, in whatever tool gets there fastest Learning science, psychometrics, or item response theory is a real plus. So is having worked with children's data and the rules that come with it. What to expect This is an in-person role at a fast pace, with periods of intense work around major launches. You'll have significant ownership and autonomy with limited oversight. The role suits researchers who do their best work with room to run. You'll be the first person in this seat. Some weeks are model evaluation, some weeks are a retention question from the founders, and you'll have to decide which one matters more that week. It's a strong fit for people who have shipped analysis that changed a decision. If your experience has been mostly reports that nobody acted on, this likely isn't the right match. 100% employer-paid Medical, Dental, and Vision; 75% dependent coverage 401(k) with employer matching Daily team dinner provided in-office A small, mission-driven team changing how the world learns

#LI-SF1

#J-18808-Ljbffr Studyfetch

Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the AI Research Scientist, Learning & Evaluation in Beverly Hills, CA vacancy
  • Studyfetch Beverly Hills, CA, is seeking an AI Research Scientist focused on learning and evaluation. You will build the evaluation layer under model training and learning science across Learn Engine and the Honen platform, driving metrics that matter for student outcomes... 
    Suggested

    Studyfetch

    Beverly Hills, CA
    16 hours ago
  • $167k - $185k

     ...and intensely curious Machine Learning Scientist with deep expertise in...  ...In this role, you'll train, evaluate, optimize, and deploy a wide...  ...staying at the forefront of AI and machine learning, especially...  ...graduate coursework, academic research, or hands-on industry... 
    Suggested
    Temporary work
    Local area
    Immediate start

    Spotter

    Culver City, CA
    1 day ago
  • $100k - $320k

     ...ecosystem, Eyeline Labs serves as the research and innovation arm of Netflix...  ...and virtual storytellers. Learn more.We are seeking a highly skilled Research Scientist with a strong background in computer...  ....Design, implement, and evaluate novel algorithms for (Dynamic)... 
    Suggested
    Worldwide

    Eyeline

    Los Angeles, CA
    4 days ago
  • $130k - $140k

     ...your career through mentoring, sponsorship, and a culture of learning. Thanks for your interest in joining our team!   Key...  ...We are currently seeking a Ph.D. level Machine Learning User Research Scientist for our Data Sciences Practice in Los Angeles, CA. In this role... 
    Suggested
    Work at office
    Local area
    Flexible hours

    Exponent

    Los Angeles, CA
    4 days ago
  • $160k - $220k

     ...Machine Learning Researcher Rainmaker is pioneering a modern cloud-seeding system to increase...  ...technologies to design, operate, and evaluate precipitation-enhancement programs....  ...attached directly to operations. Our scientists and engineers collect proprietary observations... 
    Suggested
    Work at office
    Remote work
    Relocation package
    Shift work

    Rainmaker

    El Segundo, CA
    1 day ago
  • Overview We are seeking a highly skilled Research Scientist with a strong background in computer vision, computer graphics, and machine learning. This role is ideal for candidates...  ...image synthesis. Design, implement, and evaluate novel algorithms for Dynamic Gaussian Splatting... 

    Scanline VFX LA, LLC

    Los Angeles, CA
    3 days ago
  • $99k - $225k

    Reinforcement Learning AI EngineerThe Opportunity:Are you an innovative and experienced artificial...  ...to translate reinforcement learning research into operational capability and...  ...Python and modern ML frameworks.Build and evaluate agents in simulated environments using Gym... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    El Segundo, CA
    5 days ago
  • $100k - $140k

     .... As the operator of a federally funded research and development center (FFRDC), we are broadly...  ...Materials Physicist - Non-Destructive Evaluation (NDE). In this role, you will work...  ...problem solving ~ Strong capability to learn new relevant science and engineering knowledge... 
    Full time
    Immediate start
    Remote work
    Relocation package
    Flexible hours

    The Aerospace Corporation

    El Segundo, CA
    28 days ago
  • $160k - $240k

     ...optimization at Hadrian. As our Operations Research Scientist, you'll tackle complex scheduling...  ...the U.S. Department of State. Learn more about the ITAR here.Use of AI in hiringHadrian uses AI-assisted...  ...hiring decisions. All candidate evaluations and hiring decisions are... 
    Permanent employment
    Full time
    Local area
    Flexible hours

    Hadrian

    Los Angeles, CA
    1 day ago
  • $1,100 per month

    Job Summary We are hiring an experienced AI Deep Learning Engineer to join our growing team in Los Angeles. Key Responsibilities Design and...  ...efficient deep learning models for various applications. Train and evaluate models on large datasets. Optimize model performance through... 
    Full time
    Visa sponsorship

    Charm Connect

    Los Angeles, CA
    1 day ago
  • $1,400 per month

    AI Deep Learning vacancy in Los-Angeles USA Deep Learning Researcher Join our team in Los Angeles as a Deep Learning Researcher and contribute to cutting-edge AI projects. As an Indian national, you will have the opportunity to work without language barriers and without... 
    Visa sponsorship
    Flexible hours

    Prestige Greeters

    Los Angeles, CA
    3 days ago
  • Prestige Greeters in Los Angeles is seeking a Deep Learning Researcher to join our AI team and contribute to cutting-edge projects. The role is located in Los Angeles, USA, with visa sponsorship for qualified candidates and flexible working arrangements. The position welcomes... 
    Visa sponsorship
    Flexible hours

    Prestige Greeters

    Los Angeles, CA
    3 days ago
  • A leading AI research platform is seeking an AI Trainer with expertise in Graphic and Visual Design. The role involves evaluating designs, ensuring they meet professional standards, and providing valuable insights for AI development. Candidates should have formal qualifications... 
    Remote job
    Flexible hours

    Prolific

    Los Angeles, CA
    3 days ago
  • Spotter in Culver City is seeking a Machine Learning Scientist to develop and deploy machine learning models that enhance workflows for YouTube...  ...reinforcement learning, and a strong understanding of model evaluation and product impact. You will work closely with cross-... 

    Spotter

    Culver City, CA
    1 day ago
  • $158.1k - $293.5k

     ...discovery with cutting‑edge machine learning techniques. We are seeking...  ...to develop our LLM and AI product and enable the next generation of foundational research in machine learning for scientific...  ...managing a team of engineers and scientists, and be able to foster a... 

    Genentech

    Los Angeles, CA
    1 day ago
  • $147k - $210k

     ...practical experience. Research publication experience...  ...As a Quantum Research Scientist, you will develop quantum...  ...computations. Google Quantum AI's mission is to build...  ...Responsibilities Learn more about benefits at...  ...under noise models and evaluate performance. Google is... 

    Google

    Los Angeles, CA
    2 days ago
  • AnoSys is an AI observability company building...  ...input distributions. The research challenges here are...  ...unsolved. As a Research Scientist at AnoSys, you will...  ...workflows, automated evaluation of LLM quality and safety...  ...Computer Science, Machine Learning, Statistics, Applied... 
    Immediate start

    AnoSys Technologies, Inc

    Los Angeles, CA
    3 days ago
  • $43 - $52 per hour

     ...samples and coordinating waste disposal shipments # Calculate and evaluate internal and external radiation doses of employees, visitors or...  ...or copy and paste into your browser. Privacy Notice: To learn what data we collect and how we use it, review our Privacy... 
    Hourly pay
    Local area

    Eckert & Ziegler Isotope Products, Inc.

    West Hollywood, CA
    1 day ago
  • Austin, Texas, United States Machine Learning and AI Services at Apple help hundreds of millions...  ...privacy. Description As an Applied Scientist, you will have the responsibility of...  ...in causal inference and AIML research, evaluating and integrating new frameworks where... 

    Apple Inc.

    Culver City, CA
    3 days ago
  •  ...hire a highly motivated computational scientist to fill the position of Research Bioinformatician II! The Knott...  ...first in human clinical trials. To learn more, please visit: Knott Research...  ...mining and extraction. Identifies, evaluates, and incorporates relevant algorithms... 

    CEDARS-SINAI

    West Hollywood, CA
    3 days ago
  •  ...of what’s next.The Creative Technology Research team focuses on solving high-impact technical...  ....The RoleWe are seeking a Machine Learning Researcher to define the next generation...  ...support training on large-scale clusters.Evaluation Frameworks: Conduct cross-team evaluations... 
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    Los Angeles, CA
    5 days ago
  • YO AI Labs seeks Medical Evaluation Specialists, including medical students, residents, physicians, and biomedical professionals, to contribute clinical expertise to evaluating and improving next-generation AI systems. You will create and validate challenging medical questions... 
    Remote job

    YO AI Labs

    Los Angeles, CA
    1 day ago
  • $204.44k - $324.99k

     ...our number one priority. With a wealth of learning and career development opportunities, a...  ...test Prompt Builder templates and grounded AI experiences using Salesforce data, Data...  ...RAG) patterns  Implement agent testing, evaluation, observability, and guardrails, including... 
    H1b
    Local area

    KPMG

    Los Angeles, CA
    2 days ago
  •  ...seeking a highly capable Legal Technology & AI Solutions Manager to lead the...  ...and AI-assisted development tools. Research, evaluate, test, and recommend new software, AI platforms...  ...and is committed to continuous learning and improvement. Preferred Qualifications... 

    The Lemon Pros

    Beverly Hills, CA
    9 days ago
  • $140k - $180k

     ...for highly motivated and creative Senior Scientist, Analytical Sciences who will lead...  .... Citizenship and Immigration Services. Learn more about the E-Verify program.E-Verify...  ...employees. Candidates and employees are always evaluated based on merit, qualifications, and... 
    Permanent employment
    Full time
    Immediate start
    Relocation package
    Flexible hours

    Varda Space Industries

    El Segundo, CA
    4 days ago
  • $209k - $313k

     ...themselves, live in the moment, learn about the world, and...  ...a Principal Marketing Scientist to join Snap inc.!What...  ..., including solution evaluation, adoption, tracking,...  ...measurement decisionsUtilize AI tools to more...  ...consulting firm, advertiser, or research companyPreferred... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    Los Angeles, CA
    2 days ago
  • $140k - $180k

     ...is looking for a highly motivated Senior Scientist to lead protein formulation development...  ...analytical strategies, and acceptance criteria Evaluate and implement novel formulation...  .... Citizenship and Immigration Services. Learn more about the E-Verify program.E-Verify... 
    Permanent employment
    Full time
    Immediate start
    Relocation package
    Flexible hours
    Shift work

    Varda Space Industries

    El Segundo, CA
    4 days ago
  •  ...Job Description Job Description As a Data Scientist Machine Learning, you will work within a small data science team focusing on predictive modeling...  ...abilities \nCompany Description Rule14 is an AI company and technology incubator focused on applying new approaches... 

    Rule14

    Santa Monica, CA
    9 days ago
  • $100k - $140k

     ...looking for a highly motivated early-career Scientist to join our Biologics Formulation team....  ...concentration protein formulations and evaluate formulation stability under various...  .... Citizenship and Immigration Services. Learn more about the E-Verify program.E-Verify... 
    Permanent employment
    Full time
    Immediate start
    Relocation package
    Flexible hours

    Varda Space Industries

    El Segundo, CA
    1 day ago
  • $133.2k - $172.37k

     ...day-to-day work. Join us in this mission!Scientist, Method Automation; Analytical...  ...assay workflow optimization, technology evaluations, and data-rich testing approaches that...  ...Python, multivariate analysis, machine learning, or comparable tools.Working knowledge... 
    Full time
    For contractors
    Local area

    Kite Pharma

    Santa Monica, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Research Scientist, Learning & Evaluation. Be the first to apply!