Materials Science Task Architect for AI Evaluation
Mercor
Mercor is seeking a senior scientist/engineer to design tasks in materials science and engineering that test AI models performing expert work. You write realistic prompts, assemble a data room, and define grading criteria to make evaluation objective. You will run tasks against prompts, analyze results, and iterate until the model’s reasoning can be clearly distinguished from incorrect approaches, collaborating with the project lead and domain experts. #J-18808-Ljbffr Mercor
- Mercor in San Francisco is seeking a senior evaluator to design tasks that test AI models performing real expert work. You will craft prompts, assemble... ...the model's reasoning is robustly challenged across materials science, electrical, or mechanical engineering contexts. #J-...Suggested
- Obsidian is looking for an individual to create benchmark tasks evaluating AI models on financial document understanding. You will develop multi-step requests using real-world financial files, ensuring AI can follow instructions and generate structured responses. The role...Suggested
- Mercor is partnering with a leading AI research organization to engage experienced accountants for a project that evaluates how well AI systems perform real‑world accounting work. You will design task‑specific grading criteria and define what excellent work looks like....Suggested
- Vals AI, Inc. in San Francisco is seeking exceptional researchers and research engineers to design and build the next generation of AI benchmarks that evaluate real-world LLM capabilities. You will lead development of novel benchmarks shaping how foundation models are...Suggested
- Obsidian is seeking a part-time researcher to design evaluation challenges for frontier AI in drug discovery. You will build data rooms grounded in real sources, craft grading rubrics, and validate approaches to ensure defensible conclusions. The role emphasizes rigorous...SuggestedPart timeFlexible hours
$206.4k - $379.1k
...impressive content. The AI Foundations team... ...looking for a Principal Architect to build and implement... ...analytics, and continuous evaluation frameworks.This role blends... ..., memory persistence, task decomposition, and multi... ...experience in Computer Science, Data Science, Machine...Full timeTemporary workLocal areaWorldwideFlexible hours- Mercor partners with a leading AI research organization to engage experienced accountants. You will define what excellent accounting work looks like by designing task-specific grading criteria and scoring samples with rigorous written justifications. You will apply consistent...
$142.6k - $261.5k
.... ServiceNow– ServiceNow AI Architect Manager In the digital economy... ...methods, techniques, and evaluation criteria for obtaining... ...members, ensuring successful task completion. Skills and attributes... ..., preferably in Computer Science, Information Systems Management...Summer holidayWorldwideFlexible hours- What You Will Be Doing AI/ML Architecture & Solution... ...Delivery Agentic AI Architect-Anthropic Partnership... ...structured outputs, model evaluation, safety, governance,... ...tools, execute multi-step tasks, and coordinate work... ...Machine Learning, NLP & Data Science Develop, train,...
- Mercor is seeking senior K-12 education professionals to build evaluation tasks for AI systems operating in large school district and public education contexts. The workflows are calibrated to the instructional complexity, stakeholder diversity, and student-outcome stakes...
- Mercor is seeking researchers to author AI evaluation tasks and original, executable problems for frontier models. You will source material, write prompts, and define grading criteria across subdomains with a focus on two areas in mathematics. Engagement is six weeks, part...Part timeImmediate start
- Obsidian in San Francisco builds high-fidelity simulated work environments to evaluate AI agents in marketing contexts. You will design organic growth tasks, specify required documents and dashboards, and judge AI attempts against a standard rubric. We value hands-on experience...
$180k - $240k
...About the Role We are seeking a Senior AI Agent Architect to design and deploy autonomous, multi-... ...and Autogen. You will build resilient task-delegation networks and advanced cognitive... ...enterprise databases and APIs Define evaluations, safety guardrails, and reinforcement...Full time- Senior AI Architect - Multi-Agent Systems & Platform Infrastructure Senior AI Architect - Multi... ...15% use AI for automation or compliance tasks — a gap Nivalto is built to close Your... ...components. Develop and refine test plans, evaluation pipelines, and debug tools for cross-...Full timeWork at officeRemote work
$122k - $240.5k
Position Summary Google AI Architect/AI and EngineeringJoin our AI & Engineering team... ..., security, and cost.Design, fine-tune, evaluate, and govern LLM solutions with Gemini on... ...QualificationsBachelor's degree in Computer Science, Engineering or a related technical...Local areaVisa sponsorshipFlexible hours$164.7k - $266k
...lifecycle management (CLM).What you'll doWe are seeking a Lead AI Architect to turn enterprise data, metadata, relationships, and business... ...leaders, architects, engineers, analysts, and vendors to evaluate and implement strategic AI solutionsJob DesignationHybrid: Employee...Contract workWork at officeLocal areaRemote work2 days per week- ...Next-Generation AI Research And Product Engineer Draup is a... ...development. Prototype and evaluate breakthrough AI capabilities —... ...~ BS/MS/PhD in Computer Science, AI/ML, or related field. PhD... ...a principal engineer or lead architect role. ~ Demonstrated history...Work at officeVisa sponsorship
$240k - $315k
...interpretable, and steerable AI systems. We want AI to be safe... ...role As an Applied AI Architect on the Startups team at Anthropic... ...LLM solutions, win technical evaluations, and get the most out of... ...impact AI research will be big science. At Anthropic we work as a single...Full timeWork at officeVisa sponsorshipFlexible hours- Mercor is seeking experienced accounting professionals to design and grade tasks evaluating AI performance on real-world accounting work. You will create grading criteria for reconciliations, disclosures, and close packages, and score samples with rigorous written justifications...
- OXMIQ designs GPU and AI silicon for large‑scale model inference... ...Role The Founding Principal Architect sets the architecture and technical... ...the tradeoffs among them. Evaluate inference, training, scheduling... ...BS/MS/PhD in Computer Science, Computer Engineering, or a related...
- Mercor is seeking AI experts for a benchmark dataset project evaluating models on visual document understanding and instruction-following in the Education domain. This remote role offers ~15-20 hours per week with flexible scheduling from the US or Canada. Payments are...Remote jobFlexible hours
- Conductor is seeking memory subsystem architects to own the memory subsystem for a high-bandwidth system. You will design and implement... ...controllers, and interfaces with the XPUs/compute complex. You will evaluate LLm inference workloads, memory tiering implications, and...
- ...Deployed Software Engineer to lead the software layer beneath complex partner deployments, translating bespoke environment, data, and evaluation work into repeatable infrastructure. You will own architecture for deployment services, ensure reliability and observability at...
- Mercor seeks contributors for a benchmark dataset project evaluating AI models on visual document understanding and instruction-following in the Architecture domain. This is a fully remote independent contractor role, roughly 15-20 hours per week, available to candidates...Remote jobWeekly payFor contractorsFlexible hours
- ...across campaigns and funnel data. Collaborating with Demand Gen, RevOps, Sales, Data, and Product Marketing, you’ll architect scalable processes, improve data quality, and apply AI to automate repetitive tasks while maintaining reliability and scale. #J-18808-Ljbffr Nooks
- Uber is seeking an AI Governance Lead within the Financial Risk Management (FRM) Advisory team in San Francisco. You will shape AI... ...initiatives. The role requires building scalable control architectures, evaluating AI use cases for ICFR risk, and coordinating with CAO, finance,...
$125k - $250k
...Job Description Job Description AI Partner Ecosystem Acceleration & Strategic Alliance Architect Company: HireNow Staffing (Direct Placement Partner) HireNow... ...Identify emerging AI technology partners and evaluate strategic market opportunities. Build executive...Full timeImmediate startRemote workVisa sponsorship- ...executing full-cycle recruiting for a range of roles including AI/ML research and engineering. Successful candidates will design interviews... ...relationships within the AI community, and ensure rigorous evaluation processes to maintain high hiring standards. #J-18808-Ljbffr...
$315k - $380k
...reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial... ...Role As the manager of the Solutions Architect team within Applied AI Enterprise Tech... ...C-level stakeholders in $10M+ technical evaluations and enterprise sales cycles. Multi-Segment...Visa sponsorship- ...is seeking an experienced security researcher to help mitigate AI threats and safeguard systems as AI agents become more capable.... ...lead security controls, and stress-test defenses with AI agent evaluations and penetration tests, collaborating with product and engineering...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Materials Science Task Architect for AI Evaluation. Be the first to apply!
- materials scientist San Francisco, CA
- construction materials project manager San Francisco, CA
- materials manager San Francisco, CA
- construction materials manager San Francisco, CA
- building materials sales San Francisco, CA
- outside sales building materials San Francisco, CA
- materials planning & execution specialist San Francisco, CA
- materials supervisor San Francisco, CA
- material control San Francisco, CA
- preferred materials San Francisco, CA


