Insurance AI Evaluation Specialist
Obsidian
Cincinnatus LLC is seeking an Insurance SME to join a leading AI lab's GenAI team in San Francisco. You will evaluate AI outputs against rubrics and contribute real-world underwriting judgment to training data for foundational AI models. The role requires 8+ years in insurance, strong communication, and reliable 35-hour weekday availability as part of a W-2 engagement with potential placement at a premier AI lab. #J-18808-Ljbffr Obsidian
- Cincinnatus LLC is seeking a senior Insurance SME to join a leading AI lab's GenAI team in San Francisco. You will guide underwriting-focused evaluation of AI model outputs, develop scoring rubrics, and work with cross-functional teams to ensure high-quality training data...Suggested
$20 - $26 per hour
Prolific is seeking fluent Kannada speakers to act as evaluators who compare text and voice samples to assess naturalness and authenticity. You will listen to audio clips, rate quality, and flag any mismatches in tone or pronunciation, with emphasis on cultural context...SuggestedRemote jobFlexible hours$1,750 - $2,150 per month
Obsidian is looking for experienced cybersecurity professionals to review AI systems' threat detection and vulnerability assessments. Responsibilities include evaluating AI outputs and creating realistic cybersecurity scenarios. Ideal candidates should have over 3 years...Suggested- Welo Data in San Francisco seeks a full-time AI Evaluator with professional proficiency in Portuguese (Portugal) and experience in Generative AI safety. The role involves critiquing AI outputs, identifying biases, and refining evaluation frameworks. Candidates should possess...SuggestedFull time
- A forward-thinking tech company is seeking an AI Trainer specializing in visual and graphic design to evaluate AI outputs. Responsibilities include assessing design quality and providing feedback to ensure high professional standards. Applicants should have formal education...SuggestedRemote jobFlexible hours
- Welo Data is looking for a Data Labeling Associate in San Francisco to evaluate AI systems' handling of Arabic nuances. The role requires professional-level proficiency in Arabic and 2 years of AI safety experience. Responsibilities include critiquing Arabic AI outputs...
$180k - $200k
...handle the hard stuff — payroll, health insurance, 401(k)s, and HR — so owners can focus on... ...Learn more about ourTotal Rewards philosophy.AI is a fundamental part of how work gets... ...building with AI at Gusto, including robust evaluation frameworks (evals), guardrails for code...Full timeWork at officeLocal areaRemote work2 days per week3 days per week$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering with the world’... ...largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and... ...ForAt Distyl, we build AI systems using Evaluation-Driven Development—an approach where evaluation...Work at office3 days per week$60 - $70 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...$60–$70/hour Location: Remote Role Responsibilities Evaluate AI-generated responses for safety, factual accuracy, policy compliance...Contract workSummer workRemote work$90 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...justification. Preferred Prior RLHF , preference-labeling, model-evaluation, or structured code-review work. Pixel-perfect design-to-code...Contract workSummer workLocal areaRemote work- Synthires is seeking experienced Legal Experts in the United States to contribute to advanced AI research and evaluation projects. You will evaluate AI-generated legal content, apply real-world legal reasoning, and help improve how next-generation AI systems analyze employment...Remote job
- Cincinnatus LLC is recruiting for a finance SME to join a leading AI lab's GenAI team, evaluating model outputs against rubrics and guiding financial judgment in AI training data. This is a W-2 employment placement with the option to be placed at a premier AI Lab as part...Weekday work
$240k - $280k
A leading software monitoring company is seeking a Senior Software Engineer on its AI/ML team to build evaluation infrastructure for measuring the performance of AI systems. This role involves designing datasets, creating benchmarks, and ensuring AI features behave reliably...- Obsidian is partnering with an AI research initiative to support a Frontier Code Agents project in San Francisco. You will evaluate frontier AI coding models by performing realistic data engineering tasks, reviewing ETL pipelines, data warehouses, and distributed systems...
$221k - $247k
...handle the hard stuff — payroll, health insurance, 401(k)s, and HR — so owners can focus on... ...Learn more about ourTotal Rewards philosophy.AI is a fundamental part of how work gets... ...answers for legal and business stakeholders.Evaluate, pilot, and scale AI legal-tech vendors...Full timeContract workWork at officeLocal areaShift work2 days per week3 days per week$275k - $305k
...handle the hard stuff — payroll, health insurance, 401(k)s, and HR — so owners can focus on... ...Learn more about ourTotal Rewards philosophy.AI is a fundamental part of how work gets... ...you will define how Gusto builds, deploys, evaluates, and scales AI/ML systems across the...Full timeWork at officeLocal areaRemote work2 days per week3 days per week$130k - $220k
...Artificial Analysis** Artificial Analysis is the leading independent AI benchmarking and insights company. They help engineers,... ...What This Role Actually Is** This role is best described as an AI Evaluation Engineer / Technical Generalist. It is not a traditional software...Full timeWorldwide- Mercor is hiring experienced music producers and audio engineers to evaluate generative music AI models, in partnership with a leading AI lab. The work requires evaluating AI-generated music in Hebrew and English to assess quality across genres and styles. You will rate...Immediate start
- A leading AI research organization in San Francisco is seeking a Research Engineer focused on pushing the boundaries of frontier models... ...offers the opportunity to shape innovative AI measurements and build evaluation environments that drive progress. #J-18808-Ljbffr OpenAI
- ...expanding its team with 30 Senior Management Consultants across various specialties to support frontier AI research. The role involves designing expert-level evaluation frameworks and problem formulations that today’s leading AI models struggle with, as well as...
- Synthires is seeking experienced Nurse Practitioners to contribute to AI research and evaluation of clinical reasoning, patient care, and healthcare decision-making. You will assist in assessing AI-generated clinical content and supporting safer, more effective healthcare...
- ...San Francisco is seeking an innovative Quality Engineer for their AI products. This role blends ops, strategy, and analytics to... ...in leading labs, and ensure user satisfaction through effective evaluation baselines. Competitive salary and benefits offered, with a focus...
- B Capital seeks a talented individual for an AI Evaluation role in San Francisco. This position involves conducting critical comparative analysis, refining evaluation systems, and collaborating with various teams to enhance model capabilities. The ideal candidate will have...
- Scale Labs seeks a Research Scientist focused on Frontier Risk Evaluations to design evaluation measures, harnesses and datasets for measuring risks posed by frontier AI systems. You will build harnesses to test models, collaborate with government agencies to scope evaluations...
- ...applications for a Research Scientist to design novel benchmarks and evaluate frontier language models and agents. You will lead research,... ...to external impact with opportunities to publish and present findings that influence how leading AI #J-18808-Ljbffr CerebroRelocation
- About Arena Intelligence Arena Intelligence is the open platform for evaluating how AI models perform in the real world. Created by researchers from UC Berkeley’s SkyLab, our mission is to measure and advance the frontier of AI for real-world use. Millions of people use...Permanent employmentWork at office
$216k - $270k
Scale AI, Inc. is looking for a Research Scientist specializing in Frontier Risk Evaluations to develop measures for assessing risks of advanced AI systems. In this role, you will design testing harnesses, collaborate with agencies, and publish reports to inform policymakers...- OpenAI is seeking a researcher to advance frontier evaluations and environments for safe AGI/ASI. You will help design north star model environments and steer major training runs so that research outputs translate into real-world products. Collaborate with researchers,...
- Mercor is hiring PhD and Master's scientists to author AI evaluation tasks (Sci Code) for a new benchmark in scientific computing, partnering with leading AI labs. You will author original, executable research problems that frontier models cannot solve. Engagement lasts...Part timeImmediate start
- ...tech company in New York is seeking a Research Scientist focused on Frontier Risk Evaluations. The ideal candidate will contribute to designing and creating evaluation measures for assessing AI risks and will have strong experience in machine learning and technical...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Insurance AI Evaluation Specialist. Be the first to apply!
- remote insurance agent San Francisco, CA
- medical insurance specialist San Francisco, CA
- healthcare insurance specialist San Francisco, CA
- p&c insurance agent San Francisco, CA
- medical insurance claims specialist San Francisco, CA
- insurance agents San Francisco, CA
- insurance professional San Francisco, CA
- life insurance specialist San Francisco, CA
- health insurance agents San Francisco, CA
- licensed health insurance agent San Francisco, CA


