Materials Scientist for AI Model Evaluation
$70 - $110 per hourSaidGig
Help advance frontier AI systems by bringing rigorous materials science and engineering judgment to the evaluation, design, and improvement of technical knowledge work. You will work closely with an AI research team to define what high quality materials reasoning looks like in practice and ensure model outputs can withstand technical scrutiny. Key Responsibilities
- Review materials science tasks and model outputs, identifying missing behaviors, weak reasoning, unsupported structure to property claims, and technically unsubstantiated conclusions.
- Write instruction specifications and golden solutions for materials problems, and create tasks that reflect real materials science and engineering work.
- Design challenging evaluation sets and benchmarks that measure progress in materials-specific reasoning.
- Partner with researchers and adjacent domain specialists to develop materials-focused skills and tools.
- Translate expert, tacit judgment into clear, teachable criteria and maintain consistent evaluation standards.
- PhD in materials science, materials engineering, or a closely related field such as chemistry, chemical engineering, applied physics, or metallurgy. A master''s degree with exceptional industrial depth may be considered.
- At least 4 years of substantive materials research or industrial R&D experience at a research university, national laboratory, or industrial research organization. Graduate coursework alone does not qualify.
- Specialization in at least one area, such as energy storage and battery materials, semiconductors and electronic materials, polymers and soft matter, structural alloys and metallurgy, characterization and microscopy, or computational materials and simulation.
- Senior-level research ownership, demonstrated through roles such as Senior Scientist, Staff Scientist, Research Lead, Principal Investigator, or a comparable senior industrial R&D position.
- Peer-reviewed publications, granted patents, or delivered materials programs are strongly preferred.
- Hands-on professional use of large language models and the ability to distinguish sound technical reasoning from plausible but incorrect answers.
- Excellent written communication and the ability to provide precise, well-structured feedback.
- Full-time W-2 employment, working 40 hours per week.
- Hybrid position based in the Bay Area, California, with on-site collaboration multiple days per week as required.
- Client-issued accounts and equipment will be provided, and work will be performed within the client’s tools alongside its research teams.
- $70 to $110 per hour.
- You must live in the Bay Area or relocate there at your own expense before the engagement begins. Relocation assistance is not available.
- Qualified applicants receive equal employment opportunity without discrimination based on legally protected characteristics. Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans during the application process.
- Submit your application through the opportunity channel. If selected, employment, onboarding, payroll, benefits, and compliance are managed by the employer of record while you work as part of the client team.
- Apple Inc. is seeking a Research Scientist/Engineer to design evaluation systems for foundation models powering Apple products. You will work hands‑on across evaluation design, experimentation, and cross‑team collaboration to drive model improvement and product quality....Suggested
$184.7k - $324.8k
Research Scientist / Engineer, Foundation Model Evaluation Cupertino, California, United States Software and Services We build frontier foundation models that... .... Minimum Qualifications 3+ years of experience in AI model evaluation, NLP, or a related area (e.g., natural...SuggestedRelocation$192.2k - $260k
...'s Delivery Foundation Model team, where you'll work alongside world-class scientists and engineers to pioneer... ...through advanced AI and foundation models.We... ...extensive training and evaluation infrastructure- Guide and... ...relationship with some of the material job duties of this...MaterialsLocal areaWorldwideFlexible hours$65 - $105 per hour
...Help improve how frontier AI models reason about real-world life sciences research. In this... ...will apply deep scientific judgment to evaluate research tasks and model outputs, define... ...senior-level progression, such as Senior Scientist, Staff Scientist, Research Scientist,...SuggestedHourly payFull timeLive inRelocationRelocation package$190k - $250k
...an artificial intelligence (AI) powered technology stack purpose... ...large-scale generative world models that learn to predict... ...We are looking for a research scientist to lead the design and development... ..., and radar outputsDesign evaluation frameworks that measure world...SuggestedTemporary workWork at officeVisa sponsorship$75 - $115 per hour
...Role Overview Help advance frontier AI models by bringing rigorous pharmaceutical research and development judgment to the evaluation, design, and improvement of domain-specific... ...level career progression, such as Senior Scientist, Principal Scientist, Director of...Hourly payFull timeContract workLive inRelocationRelocation package- ...world's first real-time speech AI platform capable of accent... ...team combines deep expertise in model innovation and systems engineering... ...We're looking for a Research Scientist who can define what "better"... ...Sanas's model families, build the evaluation infrastructure to measure it...
- Sanas in Palo Alto is seeking a Research Scientist focused on rigorous evaluation of speech AI models. You will define meaningful metrics, build scalable evaluation pipelines, and align research progress with product impact across Accent Translation, Noise Cancellation...
- AI Research Scientist, Learning & Evaluation Studyfetch Beverly Hills, California, United States About this position... .... Public benchmarks tell you a model can answer a question. They don't tell... ...real student finally understands the material. What you'll own Evaluation for the...MaterialsWork at officeWorldwide
$224k - $356.5k
...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts... ...computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems, you will play a meaningful role in crafting the...Full time$192.2k - $260k
...talented, and inventive Applied Scientist with a strong machine... ...technology in generative AI and foundational models.As part of our AI team in... ...and development of agentic evaluation frameworks and evaluation/... ...relationship with some of the material job duties of this...MaterialsLocal areaWorldwideFlexible hours$192.2k - $260k
...looking for a Senior Applied Scientist to help drive the... ...multimodal conversational AI. You will contribute... ...: advancing foundation models for speech and audio, and... ...generation - Design evaluation frameworks that capture... ...relationship with some of the material job duties of this...MaterialsLocal areaFlexible hours- Cincinnatus LLC is placing finance SMEs at a leading AI lab to critically evaluate AI model outputs for financial tasks. The role focuses on rigorous, rubric-based assessment of model performance and constructing finance-focused evaluation frameworks. Candidates should...Weekday work
$70 - $110 per hour
...Role Overview Help a leading AI research team improve how advanced AI models reason about real clinical work. In this hybrid, full-time role, you will... ...define high-quality clinical tasks, model answers, and evaluation standards alongside research and program management...Hourly payFull timeFreelanceLive inRelocationRelocation package$50 - $75 per hour
A leading tech company based in Australia is seeking an AI Model Evaluator on a contract basis. The role involves evaluating AI-generated responses, writing prompts, and providing justifications based on specific criteria. Ideal candidates will hold a Master's degree in...Hourly payContract work- Job Description - Member of Technical Staff (Language Model Evaluations) Location: San Francisco (preferred), Sydney, Melbourne, Brisbane About... ...Analysis Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises...
$300k - $320k
About the role: We are seeking a Technical Program Manager to lead our AI model evaluation initiatives across multiple workstreams. This role will be crucial in assessing the performance, capabilities, limitations, and potential risks of our AI models. Working closely with...Work at officeHome officeVisa sponsorshipRelocation package- Cincinnatus LLC is seeking a Marketing SME to join a leading GenAI team focused on evaluating AI outputs against rubrics and strengthening brand strategy in AI training data. We require 8+ years of marketing experience with top-tier brands, plus hands-on evaluation of LLM...Weekday work
$350k
...Our first goal is to democratize frontier AI R&D across scientific disciplines. We... ...AI research company and training our own models end-to-end. Our work spans areas such as... ...looking for a research engineer to build the evaluation infrastructure that tells us whether our...$60 per hour
Prolific is seeking Biology Experts and Life Science Professionals to evaluate AI-generated science and ensure compliance with scientific standards. Responsibilities include reviewing biological inquiries, validating technical claims from public databases, and critiquing...Remote jobHourly payWork from homeFlexible hours$100k - $140k
...rapidly evolving space systems. The Materials Science Department within the Space... ...of Materials Physicist - Non-Destructive Evaluation (NDE). In this role, you will work... ...Experience with experimental fixture design, CAD modeling, and prototyping using 3D printing or...MaterialsFull timeImmediate startRemote workRelocation packageFlexible hours$60 - $100 per hour
...Role Overview Help advance frontier AI models by bringing senior insurance and actuarial judgment to the evaluation of real-world insurance work. You will work directly with an AI research and program management team, translating professional standards into tasks, solutions...Hourly payFull timeLive inRelocationRelocation package- A leading AI company is seeking a legal professional for a contractor role focused on evaluating AI model outputs in legal contexts. Candidates must hold a Juris Doctor (J.D.) and have more than 3 years of experience in law. The role involves reviewing complex legal hypotheticals...For contractors10 hours per week
$192.2k - $260k
Are you a passionate scientist in the computer vision area who is aspired... ...LLMs and/or Vision Language Models. You will collaborate with... ...in the design, development, evaluation, deployment and updating of... ...relationship with some of the material job duties of this position....MaterialsLocal areaFlexible hours$100.3k - $175.9k
...rapidly evolving space systems. The Materials Science Department within the Space Materials... ...Materials Physicist - Non-Destructive Evaluation (NDE). In this program-facing role, you... ...Member of Technical Staff or Research Scientist depending on the selected candidate's...MaterialsFull timeWork experience placementImmediate startRemote workRelocation packageFlexible hours$400 per month
About the Role Mercor is partnering with a leading AI research lab to support a Frontier Code Agents project. Contributors help evaluate and improve frontier AI coding models through structured technical assessments. The work focuses on realistic data engineering workflows...- Anthropic in San Francisco seeks a Research Scientist to measure recursive-self-improvement in large models. You will design evaluations, build models of capability growth, and... ...RL, and policy teams to advance safe and reliable AI systems. #J-18808-Ljbffr Anthropic
- the company is seeking a Research Scientist to advance measurable recursive-self-improvement in large models. You will design evaluations, build models of capability growth, and interpret... ..., and a track record in evaluating AI systems. #J-18808-Ljbffr United States...
- A leading AI evaluation firm based in San Francisco seeks a Machine Learning Scientist to foster understanding of AI model performance. You'll engage in designing and analyzing comprehensive experiments while collaborating across teams. Applicants should possess a PhD...
- DeepMind seeks a Senior Research Scientist for Gemini Release Evaluations in Mountain View, CA. The role focuses... ...evaluation frameworks, and advancing models through robust release cycles. You will... ...1 year in data science, with strong AI model training and evaluation skills...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Materials Scientist for AI Model Evaluation. Be the first to apply!
- machine learning scientist California
- scientist California
- quality control scientist California
- qc scientist California
- research scientist - biology California
- applied scientist California
- support scientist California
- molecular biology scientist California
- research scientist California
- lab scientist California



