AI Evaluation Researcher: Post-Training Signals & Metrics
Mosaic.tech
Thinking Machines in San Francisco seeks a researcher to advance internal evaluations and signals for post-training models. You will collaborate with researchers and engineers across the research organization, shaping evaluation creation, usability, and auditing. Your work will span agentic evaluations, score systems, and scalable signals that improve data and training signals, with emphasis on robustness, truth-grounding, and generalization across environments. #J-18808-Ljbffr Mosaic.tech
$150k - $250k
About Distyl AI Distyl is an applied... ....We research and deploy technologies... ...Researchers design evaluation frameworks that... ...how metrics shape model behavior... ...and can extract signal from noisy empirical... ...rather than training or fine-tuning... ...top journals, posted amazing work on...TrainingWork at office3 days per week$200k - $280k
...architectures, engines) and post-training / RL systems. We build... ...stack. Have a solid research foundation in your... ...rollout collection and evaluation cheaper. Use these... ...as needed. Establish metrics, benchmarks, and experimentation... .... About Together AI Together AI is a...TrainingFull time$200,000 - $350,000 per day
Applied AI Researcher - Model Evaluation & Data Strategy San Francisco (in-person preferred; open to remote... ...the datasets, rubrics, reward signals, and quality controls that move performance... ...lab, foundation-model company, or post-training team; expert-data or human-eval...TrainingRemote workVisa sponsorship- ...anything. We're building the AI that finally changes... ...What, and Who Why AI Researchers are the engine of... ...legal- specific tasks. Evaluate emerging work in agentic... ..., and evals for training and measuring model performance... ...text. Define the metrics that matter, and hold...TrainingContract workWork at officeImmediate startRemote workVisa sponsorshipRelocation packageFlexible hours
$380k
...increasingly capable AI agents act... ...spans three areas:Training: Create training... ...into training signals that prevent similar... ...: Build evaluations and production metrics that identify emerging... ...frontier model research. You don’t need... ...Collaborate closely with post-training,...TrainingWork at officeRelocation package$216.3k - $280.8k
Meet the TeamAt Foundation AI, we are leading frontier AI research across Cisco. Our mission is to advance... ..., reasoning systems, scalable training algorithms, evaluation science, inference optimization,... ...expertise in foundation models, post-training and alignment, LLMs, agentic...TrainingFull timeTemporary workLocal areaFlexible hours$204k - $300k
...Advanced Technology Group (ATG) is the research division of the company. ATG’s... ...and electrical engineering, such as AI/ML, algorithms, digital signal processing, audio engineering, image... ...internal parity, and relevant education or training. Your recruiter can share more about...TrainingFull timeLocal areaWorldwideFlexible hours$196k - $230k
...We AreNotion is the collaborative AI workspace where teams and agents think... ...:We’re seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered experiences—focusing... ...scoring guidelines, and observable metrics.AI fluency and systems thinking:...Local areaShift work$262.5k - $299.6k
Applied Researcher II (AI Foundations) Overview: At Capital One, we are creating trustworthy... ...of development, from design through training, evaluation, validation, and implementation. Engage... ...willing to pay at the time of this posting. Salaries for part-time roles will be...TrainingFull timePart timeLocal areaFlexible hours$262.5k - $299.6k
Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview: At Capital One, we... ...of development, from design through training, evaluation, validation, and implementation. Engage... ...willing to pay at the time of this posting. Salaries for part-time roles will be...TrainingFull timePart timeLocal areaFlexible hours$160k - $300k
Applied AI Researcher - Video Diffusion Location: On-site, San Francisco, CA Compensation: $... ...Applied AI Researcher to help lead the training of state-of-the- art video diffusion models... ...compression, codecs, and perceptual metrics You Might Be a Fit If You're hands-on...TrainingFull time- AI Researcher (Computer Vision/Multimodal/Generative AI) About the Role We are hiring ML Researchers... ...new architectures, algorithms, and training strategies that improve realism,... ...aligned with product differentiation. Evaluate new model paradigms for scalability and...Training
$205k - $300k
...another scribe. We’re building the AI intelligence platform that... ...Experience in the KLAS Research Emerging Solutions Top 20 Report... ...novel products and features. Evaluate and refine AI models to ensure... ...high-quality, representative training and testing material. Analyze...TrainingWork at officeRemote workFlexible hours3 days per week- Verita AI works with leading AI companies to identify... ...hiring an Applied AI Researcher to work directly with clients on model evaluation and data strategy. You will... ...tasks, and evaluator-training programs. Design pilot studies... ..., AI data company, or post-training team....Training
- ...uses real neurons to improve AI models. We study how biological... ...neuroscience, AI research and software engineering to develop... ...experimentation and system‑level evaluation. You will make important... ...evaluation Identify modeling, training, and scaling risks before they...Training
$100k - $150k
...neuroscience, hardware, and AI on problems that... ...of passionate researchers and engineers... ...neural and sensor signals into reliable, real... ...ll Do Design and train state-of-the-art models... ..., training, evaluation, and deployment Develop... ..., drift) and metrics aligned with user...TrainingWork at office- We are seeking an Edge AI Research Scientist to develop next-generation speech and audio AI... ...Hardware-aware architectures Efficient training and inference techniques Contribute to... ...and deployment frameworks. Performance Evaluation Develop rigorous benchmarking methodologies...Training
$84.13 - $91.34 per hour
AI Researcher - Efficient AI (Contractor) Step into the innovative world of LG Electronics. As a global... ...multimodal models, and agentic workloads across post-training, inference, and deployment workflows. • Propose and evaluate novel compression methods (PTQ, QAT, pruning,...TrainingFull timeContract workTemporary workFor contractorsLocal areaImmediate start- ...adventure game where an AI companion is the core... ..., and top-tier AI researchers. As an early member... ...industry experience with LLM post-training You have a deep... ...understanding of LLM training and evaluation with attention to... ...by finding implicit signals in the human-AI...TrainingWork at officeVisa sponsorship
- ...app reaching 2,000-3,000 new users a day. We’re hiring an AI Researcher to build the next generation of real-time, interactive voice... ...lifecycle, from framing the question through distributed training, evaluation, and production deployment. About Kotoba Kotoba is a...Training
- ...Science San Francisco, CA, USA Posted on Jul 31, 2026 About TBC... ...uses real neurons to improve AI models. We study how... ...computational neuroscience, AI research and software engineering to... ...platform, including core modeling, training, evaluation, and deployment decisions...Training
- ...Technical Staff, Lead Researcher San Francisco, CA;... ...is building an AI Research org from... ...supply and demand signals, and longitudinal... ...compute budgets for training and inference,... ...pre-training and post-training, RL training... ..., and large-scale evaluation sweeps Full research...TrainingLocal area
- Verita AI is seeking an Applied AI Researcher to work with clients on model evaluation and data strategy. You will assess model performance, identify failure modes, and design data-driven solutions, collaborating with operations and engineering to implement scalable data...
- ...of everything we do.Schwab’s AI Strategy & Transformation team... ...company. We also build the research platform that powers AI at scale... ...technical guardrails, evaluation frameworks, and continuous monitoring... ...-in-the-loop and automated metrics for compliance at scale....Full timeWork at office
$172.5k - $260.1k
...SalesforceSalesforce is the #1 AI CRM, where humans... ...of Salesforce.Research & Insights (R&I)... ..., generative, and evaluative research using the... ..." from meaningful signals, validating AI-... ...and on-demand training with Trailhead.comExposure... ...opt out options.Posting...TrainingFull time$150k - $300k
...Join to apply for the Applied Researcher role at Variant Overview... ...This role involves designing, training, and evaluating deep neural networks and... ...end: hypotheses, datasets, metrics, results Prototype fast; turn... ...article, started with the help of AI. #J-18808-Ljbffr...TrainingFull time- ...a Google Deepmind veteran behind Project Astra, and top-tier AI researchers. As an early member of this team, you will have significant... ...You May Be A Good Fit If: You have hands-on experience in training text-to-motion or audio-to-motion models. You might have...TrainingWork at officeVisa sponsorship
$204k - $259k
...of-the-art Generative AI to create a training ground for the Waymo... ...Driver. The Simulator Evaluation team faces the... ...Engineer to build the metrics and systems that grade... ...into clear, actionable signals. The "Critic" for... ...partner closely with AI research and other simulation...TrainingFull timeRemote work- ...institutions. In 2025, we started Handshake AI and built the fastest-growing AI... ...work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the... ...with various data-intensive post-training techniques. We believe that data spend...TrainingFull timeWork at officeRemote workFlexible hours
$238k - $302k
...of-the-art Generative AI to create a training ground for the Waymo... ...Driver. The Simulator Evaluation team faces the... ...between deep technical metrics and high-level product... ...clear, reproducible signals on petabytes of data.... ...partner closely with AI research and other simulation...TrainingFull timeRemote workShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Evaluation Researcher: Post-Training Signals & Metrics. Be the first to apply!
- field researcher San Francisco, CA
- product researcher San Francisco, CA
- security researcher San Francisco, CA
- lead researcher San Francisco, CA
- data collection researcher San Francisco, CA
- machine learning researcher San Francisco, CA
- court researcher San Francisco, CA
- researcher San Francisco, CA
- postdoctoral researcher cosmetic science San Francisco, CA
- senior researcher San Francisco, CA


