GenAI Evaluation Scientist - LLM Benchmarks & Failures
Scale
Scale is seeking a Machine Learning Research Scientist, Evaluations to join the GenAI Research Organization in San Francisco. You will develop rigorous evaluations, diagnose failure modes in frontier LLMs and agents, and design benchmarks for text and multimodal modalities. Collaboration with researchers and engineers will shape evaluation-driven AI development. The role emphasizes post-training techniques like SFT and RLHF, with opportunities to publish findings at top conferences and influence #J-18808-Ljbffr Scale
$165.6k - $207k
Machine Learning Research Scientist, Evaluations Scale works with the... ...accelerate progress in GenAI research. We are looking... ...Engineers with expertise in LLM post-training (SFT, RLHF... ...will focus on building benchmarks and diagnosing model failure modes in both text and multimodal...SuggestedFull time- ...MaaS platforms for large language models and contribute to end-to-end MaaS solutions across markets. You will design evaluation systems, run benchmarks, and analyze large-scale logs to drive model improvement, with opportunities to publish and pursue bold ideas within...Suggested
$142.8k - $193.2k
...Translations Services (TS) is seeking an Applied Scientist to be based in our Seattle office. As a... ...* Apply your expertise in LLM models to design, develop, and implement... ...marketplaces.* Continuously explore and evaluate state-of-the-art modeling techniques and...SuggestedWork at officeWorldwideFlexible hours- ...Senior Applied Scientist The OCI AI Evaluation team builds the evidence behind model... ...select or create appropriate benchmarks, design experiments, build... ...data and metrics, analyze failure modes, and communicate... ...automated evaluators, including LLM-as-a-judge methods;...Suggested
$167.1k - $226.1k
...for an seasoned Applied Scientist to design, build, and... ...anomaly detection, and LLM-based reasoning — all applied... ...(site, service, failure mode, time)- **Dose-response... ...rationales- **LLM evaluation:** Build evaluation frameworks... ...- Experience measuring GenAI/productivity tools'...SuggestedFlexible hours- Scale AI, Inc. is seeking Research Scientists and Research Engineers to advance LLM post-training techniques for text and multimodal data. You will develop... ...data-driven best practices in data curation and evaluation. This role involves publishing results at top AI conferences...
$167.1k - $226.1k
...looking for a Senior Applied Scientist to lead the development of generative... ...model cost- Developing evaluation frameworks that catch quality... ...problem than optimize a benchmark.What Makes This DifferentYour... ...operates at a scale where naive LLM approaches break downYou'll work...Contract workFlexible hours$160k - $190k
...This is a hands-on IC role and our first hire dedicated to human evaluation. You'll own it end to end (what to measure, how to measure it,... .... Calibrate automated and model-based metrics (including LLM-as-judge) against human judgment, so the team knows when to trust...H1bWork at officeVisa sponsorship$50 per hour
...and work with researchers to build better benchmarks. Responsibilities Design advanced... ...by-step solutions with rigorous logic. Evaluate AI outputs for accuracy and quality of reasoning... ...on cutting-edge AI projects with leading LLM companies. Pay rate: $50+/hour (...Contract workRemote workFlexible hours$136k - $184k
...skilled, and innovative Applied Scientist with a robust background in... ...methodologies, and automated evaluation systems to ensure the highest... ...Nova improve performances on benchmarks. The Applied Scientist will... ...involves developing and maintaining LLM-as-a-Judge systems, including...Flexible hours$167.1k - $226.1k
...International Emerging Stores Payments as we apply LLM techniques to improve the customer... ...to defined science problem statements2. Evaluate and identify the right ML algorithm to... ...Models4. Partner with IML team in applying GenAI development strategies to accelerate...Flexible hours$142.8k - $193.2k
...Seller Assistant is our flagship GenAI-first, multi-agent system... ...seeking a world-class Applied Scientist to help define and build the... ...large scale data analyses, model benchmarking, model validation and model... ...prompts and critically evaluate AI-generated outputs in a professional...WorldwideFlexible hours- ...serving as the foundation for intelligent e-commerce agents across diverse applications.5. Agent Evaluation, Safety & Compliance:Designing evaluation metrics and benchmarks aligned with real-world business scenarios, ensuring robustness, safety, and compliance under highly...
$142.8k - $193.2k
...alongside world-class scientists to build AI that transforms... ...design, train, and evaluate systems that reason... ...combines expertise in LLM reasoning and agentic... ...evaluation harnesses and benchmarks to measure agent... ...reliability, faithfulness, and failure modes in high-stakes...WorldwideFlexible hours$50 per hour
...accelerator is seeking remote experts with a PhD in Biology, Biotechnology, or Biochemistry. You will design advanced biology questions, evaluate AI performance, and work with researchers on cutting-edge projects. This flexible role offers pay of $50+/hour and a commitment of...Remote jobFlexible hours- ...Role As a Senior Software Engineer – Evaluation, you will design and implement systems... ...will develop evaluation methodologies, benchmarking pipelines, and monitoring tools that ensure... ...evaluate model performance, identify failure modes, and guide improvements to data...
- State of Washington Dept. of Fish and Wildlife is seeking the Adult Salmonid Trapping and Evaluation Biologist. The role focuses on monitoring Snake River summer steelhead with emphasis at Lyons Ferry Hatchery and Southeast Washington field work. You will coordinate trapping...Summer work
$167.1k - $226.1k
...The research challenges are immense. GenAI and VLMs hold transformative promise for... ...innovative and customer-focused applied scientist to help us make the world's best product... ...petabytes of multimodal data with rigorous evaluation frameworks * Define research roadmaps...WorldwideFlexible hours- Title- Adult Salmonid Trapping and Evaluation Biologist Classification- Fish & Wildlife Biologist 3 Job Status- Full-Time / Permanent WDFW Program- Fish Program - Fish Science Division Duty Station- Dayton, Washington - Columbia County Hybrid/Telework- While this position...Permanent employmentFull timeContract workTemporary workSummer workCasual workWork at officeLocal areaRemote workMonday to FridayAfternoon shift
- ...WA; Atlanta, GA; New York, NY. The role focuses on engineering geology, remote sensing, GIS, mapping and geophysics to investigate failures and assess natural hazards, offering growth and leadership opportunities within a multidisciplinary team. #J-18808-Ljbffr...Remote work
- ...Machine Learning Scientist Independently partners with science, engineering, and product... ...(e.g., including modeling approaches, evaluation techniques, data collections). Solution... .... Generative Artificial Intelligence (GenAI) Model Building Demonstrated ability in...Shift work
$171.6k - $222.2k
...motivated, passionate and resourceful Applied Scientist to bring diverse perspectives, ideas,... ...Search systems using Deep learning, GenAI, Reinforcement Learning, and optimization... ...offline and online (A/B) experiments to evaluate proposed solutions based on in-depth data...Local areaWorldwideFlexible hours$142.8k - $193.2k
...production.Key job responsibilitiesAs an Applied Scientist in our team, you will be responsible for... ...massive real-world datasets.5. Develop evaluation frameworks for a system where "quality"... ..., knowledge representation, and LLM reasoning — and the right approach hasn'...Flexible hours$171.6k - $222.2k
...passionate, talented, and inventive Applied Scientist with a strong background in speech and... ...pipelines. You’ll work across model evaluation and selection, fine-tuning, data and... ...inference and transcription- Design datasets, benchmarks, experiments, and evaluation...Local areaImmediate startFlexible hours$142.8k - $193.2k
...using Large Language Models (LLMs): a new LLM stack that already powers Amazon Search,... ...across Amazon.We are hiring an Applied Scientist to push the science behind this stack:... ...personalization.- Drive improvements on offline benchmarks as well as online experiments.About the...Flexible hours$167.1k - $226.1k
Are you a scientist passionate about advancing Information Retrieval, NLP, and Large Language... ....We're looking for scientists with deep LLM expertise to build our next generation... ...architecture, training methodology, and evaluation frameworks, balancing scientific rigor...Flexible hours- Job Title: Research Scientist Location: Seattle, WA (Remote) Duration: 12 months contract... ...and implementing robust processes to evaluate and improve LLM-based models. More specifically, they... .... Proven experience developing benchmarks and metrics to evaluate AI models, preferably...Contract workRemote work
$167.1k - $226.1k
...AI is looking for a Senior Applied Scientist to build Alexa+, Amazon's LLM-powered conversational assistant.... ...alignment, agentic reasoning, and evaluation — directly shaping the experience... ...methodologies that go beyond standard benchmarks to capture real-world...WorldwideFlexible hours- Nuance Labs in Seattle is seeking an experienced human-evaluation scientist to turn subjective judgments into measurable signals for real-time AI avatars. You will design studies, define rubrics, and coordinate with researchers to align outcomes with product goals. You...Visa sponsorship
$207.48k - $368.22k
...events. Swiftly pinpoint root causes of failures across the entire stack, from optical transceivers... ...platforms. - A Passion for AIOps/ML/LLM Practices: - A keen interest in the... ...operations (e.g., RAG, tool use, safety evaluation). Preferred Qualifications: - Experience...Temporary workLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GenAI Evaluation Scientist - LLM Benchmarks & Failures. Be the first to apply!
- applied scientist Seattle, WA
- remote scientist Seattle, WA
- operations research scientist Seattle, WA
- principal scientist Seattle, WA
- wetland scientist Seattle, WA
- applied sports scientist Seattle, WA
- health scientist Seattle, WA
- drug safety scientist Seattle, WA
- scientist biology Seattle, WA
- safety scientist Seattle, WA


