Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GenAI Evaluation Scientist - LLM Benchmarks & Failures

Scale

Scale is seeking a Machine Learning Research Scientist, Evaluations to join the GenAI Research Organization in San Francisco. You will develop rigorous evaluations, diagnose failure modes in frontier LLMs and agents, and design benchmarks for text and multimodal modalities. Collaboration with researchers and engineers will shape evaluation-driven AI development. The role emphasizes post-training techniques like SFT and RLHF, with opportunities to publish findings at top conferences and influence #J-18808-Ljbffr Scale

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the GenAI Evaluation Scientist - LLM Benchmarks & Failures in Seattle, WA vacancy
  • $165.6k - $207k

    Machine Learning Research Scientist, Evaluations Scale works with the...  ...accelerate progress in GenAI research. We are looking...  ...Engineers with expertise in LLM post-training (SFT, RLHF...  ...will focus on building benchmarks and diagnosing model failure modes in both text and multimodal... 
    Suggested
    Full time

    Scale

    Seattle, WA
    1 day ago
  •  ...MaaS platforms for large language models and contribute to end-to-end MaaS solutions across markets. You will design evaluation systems, run benchmarks, and analyze large-scale logs to drive model improvement, with opportunities to publish and pursue bold ideas within... 
    Suggested

    ByteDance

    Seattle, WA
    4 days ago
  • $142.8k - $193.2k

     ...Translations Services (TS) is seeking an Applied Scientist to be based in our Seattle office. As a...  ...* Apply your expertise in LLM models to design, develop, and implement...  ...marketplaces.* Continuously explore and evaluate state-of-the-art modeling techniques and... 
    Suggested
    Work at office
    Worldwide
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  •  ...Senior Applied Scientist The OCI AI Evaluation team builds the evidence behind model...  ...select or create appropriate benchmarks, design experiments, build...  ...data and metrics, analyze failure modes, and communicate...  ...automated evaluators, including LLM-as-a-judge methods;... 
    Suggested

    Hackajob

    Seattle, WA
    1 day ago
  • $167.1k - $226.1k

     ...for an seasoned Applied Scientist to design, build, and...  ...anomaly detection, and LLM-based reasoning — all applied...  ...(site, service, failure mode, time)- **Dose-response...  ...rationales- **LLM evaluation:** Build evaluation frameworks...  ...- Experience measuring GenAI/productivity tools'... 
    Suggested
    Flexible hours

    Amazon

    Seattle, WA
    6 days ago
  • Scale AI, Inc. is seeking Research Scientists and Research Engineers to advance LLM post-training techniques for text and multimodal data. You will develop...  ...data-driven best practices in data curation and evaluation. This role involves publishing results at top AI conferences... 

    Scale AI, Inc.

    Seattle, WA
    2 days ago
  • $167.1k - $226.1k

     ...looking for a Senior Applied Scientist to lead the development of generative...  ...model cost- Developing evaluation frameworks that catch quality...  ...problem than optimize a benchmark.What Makes This DifferentYour...  ...operates at a scale where naive LLM approaches break downYou'll work... 
    Contract work
    Flexible hours

    Amazon

    Bellevue, WA
    a month ago
  • $160k - $190k

     ...This is a hands-on IC role and our first hire dedicated to human evaluation. You'll own it end to end (what to measure, how to measure it,...  .... Calibrate automated and model-based metrics (including LLM-as-judge) against human judgment, so the team knows when to trust... 
    H1b
    Work at office
    Visa sponsorship

    Nuance Labs

    Seattle, WA
    2 days ago
  • $50 per hour

     ...and work with researchers to build better benchmarks. Responsibilities Design advanced...  ...by-step solutions with rigorous logic. Evaluate AI outputs for accuracy and quality of reasoning...  ...on cutting-edge AI projects with leading LLM companies. Pay rate: $50+/hour (... 
    Contract work
    Remote work
    Flexible hours

    Turing

    Seattle, WA
    1 day ago
  • $136k - $184k

     ...skilled, and innovative Applied Scientist with a robust background in...  ...methodologies, and automated evaluation systems to ensure the highest...  ...Nova improve performances on benchmarks. The Applied Scientist will...  ...involves developing and maintaining LLM-as-a-Judge systems, including... 
    Flexible hours

    Amazon

    Bellevue, WA
    6 days ago
  • $167.1k - $226.1k

     ...International Emerging Stores Payments as we apply LLM techniques to improve the customer...  ...to defined science problem statements2. Evaluate and identify the right ML algorithm to...  ...Models4. Partner with IML team in applying GenAI development strategies to accelerate... 
    Flexible hours

    Amazon

    Seattle, WA
    6 days ago
  • $142.8k - $193.2k

     ...Seller Assistant is our flagship GenAI-first, multi-agent system...  ...seeking a world-class Applied Scientist to help define and build the...  ...large scale data analyses, model benchmarking, model validation and model...  ...prompts and critically evaluate AI-generated outputs in a professional... 
    Worldwide
    Flexible hours

    Amazon

    Seattle, WA
    a month ago
  •  ...serving as the foundation for intelligent e-commerce agents across diverse applications.5. Agent Evaluation, Safety & Compliance:Designing evaluation metrics and benchmarks aligned with real-world business scenarios, ensuring robustness, safety, and compliance under highly... 

    TikTok

    Seattle, WA
    4 days ago
  • $142.8k - $193.2k

     ...alongside world-class scientists to build AI that transforms...  ...design, train, and evaluate systems that reason...  ...combines expertise in LLM reasoning and agentic...  ...evaluation harnesses and benchmarks to measure agent...  ...reliability, faithfulness, and failure modes in high-stakes... 
    Worldwide
    Flexible hours

    Amazon

    Seattle, WA
    12 hours ago
  • $50 per hour

     ...accelerator is seeking remote experts with a PhD in Biology, Biotechnology, or Biochemistry. You will design advanced biology questions, evaluate AI performance, and work with researchers on cutting-edge projects. This flexible role offers pay of $50+/hour and a commitment of... 
    Remote job
    Flexible hours

    Turing

    Seattle, WA
    2 days ago
  •  ...Role As a Senior Software Engineer – Evaluation, you will design and implement systems...  ...will develop evaluation methodologies, benchmarking pipelines, and monitoring tools that ensure...  ...evaluate model performance, identify failure modes, and guide improvements to data... 

    VTI Aerospace

    Seattle, WA
    25 days ago
  • State of Washington Dept. of Fish and Wildlife is seeking the Adult Salmonid Trapping and Evaluation Biologist. The role focuses on monitoring Snake River summer steelhead with emphasis at Lyons Ferry Hatchery and Southeast Washington field work. You will coordinate trapping... 
    Summer work

    State of Washington Dept. of Fish and Wildlife

    Seattle, WA
    3 days ago
  • $167.1k - $226.1k

     ...The research challenges are immense. GenAI and VLMs hold transformative promise for...  ...innovative and customer-focused applied scientist to help us make the world's best product...  ...petabytes of multimodal data with rigorous evaluation frameworks * Define research roadmaps... 
    Worldwide
    Flexible hours

    Amazon

    Seattle, WA
    15 hours ago
  • Title- Adult Salmonid Trapping and Evaluation Biologist Classification- Fish & Wildlife Biologist 3 Job Status- Full-Time / Permanent WDFW Program- Fish Program - Fish Science Division Duty Station- Dayton, Washington - Columbia County Hybrid/Telework- While this position... 
    Permanent employment
    Full time
    Contract work
    Temporary work
    Summer work
    Casual work
    Work at office
    Local area
    Remote work
    Monday to Friday
    Afternoon shift

    State of Washington Dept. of Fish and Wildlife

    Seattle, WA
    3 days ago
  •  ...WA; Atlanta, GA; New York, NY. The role focuses on engineering geology, remote sensing, GIS, mapping and geophysics to investigate failures and assess natural hazards, offering growth and leadership opportunities within a multidisciplinary team. #J-18808-Ljbffr... 
    Remote work

    Exponent Inc.

    Bellevue, WA
    5 days ago
  •  ...Machine Learning Scientist Independently partners with science, engineering, and product...  ...(e.g., including modeling approaches, evaluation techniques, data collections). Solution...  .... Generative Artificial Intelligence (GenAI) Model Building Demonstrated ability in... 
    Shift work

    Hackajob

    Seattle, WA
    3 days ago
  • $171.6k - $222.2k

     ...motivated, passionate and resourceful Applied Scientist to bring diverse perspectives, ideas,...  ...Search systems using Deep learning, GenAI, Reinforcement Learning, and optimization...  ...offline and online (A/B) experiments to evaluate proposed solutions based on in-depth data... 
    Local area
    Worldwide
    Flexible hours

    Amazon

    Seattle, WA
    1 day ago
  • $142.8k - $193.2k

     ...production.Key job responsibilitiesAs an Applied Scientist in our team, you will be responsible for...  ...massive real-world datasets.5. Develop evaluation frameworks for a system where "quality"...  ..., knowledge representation, and LLM reasoning — and the right approach hasn'... 
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $171.6k - $222.2k

     ...passionate, talented, and inventive Applied Scientist with a strong background in speech and...  ...pipelines. You’ll work across model evaluation and selection, fine-tuning, data and...  ...inference and transcription- Design datasets, benchmarks, experiments, and evaluation... 
    Local area
    Immediate start
    Flexible hours

    Amazon

    Seattle, WA
    13 days ago
  • $142.8k - $193.2k

     ...using Large Language Models (LLMs): a new LLM stack that already powers Amazon Search,...  ...across Amazon.We are hiring an Applied Scientist to push the science behind this stack:...  ...personalization.- Drive improvements on offline benchmarks as well as online experiments.About the... 
    Flexible hours

    Amazon

    Seattle, WA
    10 days ago
  • $167.1k - $226.1k

    Are you a scientist passionate about advancing Information Retrieval, NLP, and Large Language...  ....We're looking for scientists with deep LLM expertise to build our next generation...  ...architecture, training methodology, and evaluation frameworks, balancing scientific rigor... 
    Flexible hours

    Amazon

    Seattle, WA
    a month ago
  • Job Title: Research Scientist Location: Seattle, WA (Remote) Duration: 12 months contract...  ...and implementing robust processes to evaluate and improve LLM-based models. More specifically, they...  .... Proven experience developing benchmarks and metrics to evaluate AI models, preferably... 
    Contract work
    Remote work

    MindSource

    Seattle, WA
    5 days ago
  • $167.1k - $226.1k

     ...AI is looking for a Senior Applied Scientist to build Alexa+, Amazon's LLM-powered conversational assistant....  ...alignment, agentic reasoning, and evaluation — directly shaping the experience...  ...methodologies that go beyond standard benchmarks to capture real-world... 
    Worldwide
    Flexible hours

    Amazon

    Bellevue, WA
    1 day ago
  • Nuance Labs in Seattle is seeking an experienced human-evaluation scientist to turn subjective judgments into measurable signals for real-time AI avatars. You will design studies, define rubrics, and coordinate with researchers to align outcomes with product goals. You... 
    Visa sponsorship

    Nuance Labs

    Seattle, WA
    1 day ago
  • $207.48k - $368.22k

     ...events. Swiftly pinpoint root causes of failures across the entire stack, from optical transceivers...  ...platforms. - A Passion for AIOps/ML/LLM Practices: - A keen interest in the...  ...operations (e.g., RAG, tool use, safety evaluation). Preferred Qualifications: - Experience... 
    Temporary work
    Local area

    ByteDance

    Seattle, WA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GenAI Evaluation Scientist - LLM Benchmarks & Failures. Be the first to apply!