Principal AI Researcher - LLM Evaluation & Data
Carnaby Fox
Carnaby Fox is seeking a Member of Technical Staff (AI Research) in San Francisco to help shape the research direction for frontier AI models. You will collaborate with world-class researchers to design experiments, evaluate LLMs, and improve data quality for high-stakes AI benchmarks. The role emphasizes independent ownership, pioneering benchmark design, and contributions to publications alongside top AI labs and researchers in a well-funded startup setting. #J-18808-Ljbffr Carnaby Fox
$262.5k - $299.6k
...Applied Researcher II (AI Foundations, LLM Core and Agentic AI) Overview At Capital One, we are creating trustworthy... ...with a cross‑functional team of data scientists, software engineers,... ...development, from design through training, evaluation, validation, and implementation....DataFull timePart timeLocal areaFlexible hours$197.3k - $313.7k
...SalesforceSalesforce is the #1 AI CRM, where humans with agents... ..., diverse team of researchers at Agentforce Operations. The... ...and algorithms for enterprise data, such as tabular, relational,... ...experience developing, deploying, and evaluating machine learning models in...PrincipalDataFull timeImmediate startRemote work$196k - $230k
...AreNotion is the collaborative AI workspace where teams and... ...seeking an experienced UX Researcher to define and scale how we evaluate Notion’s AI-powered... ...design, engineering, and data science can apply consistently... ...calibration of automated/LLM-judge approaches against human...DataLocal areaShift work- ...adventure game where an AI companion is the core... ...developed an in-house LLM storytelling system that... ...Astra, and top-tier AI researchers. As an early member... ...understanding of LLM training and evaluation with attention to... ...in building data processing backend, deterministic...DataWork at officeVisa sponsorship
$136.45k - $231.65k
...Principal UX Researcher Let's face it, a company whose mission is human transformation... ...and qualitative data into actionable insights that... ...scale research operations: evaluate and select research tools and... ...best practices for responsible AI usage in research processes,...PrincipalDataWork experience placementSummer holidayLocal areaFlexible hours- ...the specified location(s).As an AI Researcher within Schwab’s AI Strategy &... ...formulation through modeling, evaluation, deployment, and iteration, contributing directly to LLM and Agentic-based systems used... ...teams.Strong grounding in data analysis (Python, SQL, Pandas)...DataFull timeWork at office
- ...operationalizes responsible AI governance at scale. We're... ...We're seeking a Principal AI Security & Risk Researcher to join our founding research... ...methodologies Stay current with LLM vulnerabilities,... ...agentic systems Develop risk evaluation methodologies that adapt as...PrincipalPart timeRemote workFlexible hours
$262.5k - $299.6k
Applied Researcher II (AI Foundations) Overview: At Capital One, we are creating... ...a cross-functional team of data scientists, software... ...from design through training, evaluation, validation, and implementation... ...Engineering or related fields LLM PhD focus on NLP or...DataFull timePart timeLocal areaFlexible hours- Senior Software Engineer - LLM Evaluation Skills JavaScript JavaScript Python Overview... ..., Turing is the world's leading research accelerator for frontier AI labs and a trusted partner for global... ...frontier research with high-quality data, advanced training pipelines, plus...DataFull timeFor contractorsFlexible hours
- Distyl AI is seeking researchers to join our Post-Training team, translating foundation models into enterprise-ready systems. You will work on supervised... ..., building prototypes and conducting experiments that demonstrate impact using real-world data. #J-18808-Ljbffr DistylData
$197.3k - $313.7k
...SalesforceSalesforce is the #1 AI CRM, where humans with... ...of business with AI+ Data +CRM. Leading with our... ...— lead competitive research and industry-trend analysis... ..., vector search, and LLM workflow orchestration,... ...AI/ML observability, evaluation, and guardrails.Experience...PrincipalDataFull time- ...intelligent, auditable AI platform that spans... ...is seeking a Principal Generative AI Engineer... ...— encompassing LLM fine-tuning, retrieval... ...standards for model evaluation, system scalability... ...edge generative AI research and real-world... ...Strong background in data privacy, security,...PrincipalDataFull time
$308k - $423.5k
...using the power of tech, data, and machine learning... ...We are seeking a Principal ML / AI Engineer to be a company... ...with cutting-edge AI research and applications. This... ...of AI systems (LLM fine-tuning, RLHF, agent... ...technical standards and evaluation frameworks for safety,...PrincipalDataFull timeWork experience placementWork at officeLocal areaRemote workMonday to FridayFlexible hours3 days per week$216.3k - $280.8k
...received.Meet the TeamAt Foundation AI, we are leading frontier AI research across Cisco. Our mission is to... ...systems, scalable training algorithms, evaluation science, inference optimization,... ...At Cisco, we’re revolutionizing how data and infrastructure connect and protect...DataFull timeTemporary workLocal areaFlexible hours$150k - $250k
About Distyl AI Distyl is an applied AI technology company partnering... ...social organizations.We research and deploy technologies that power... ...measured. Researchers design evaluation frameworks that capture... ...workflowStrong Programming and Data Analysis Skills: While you might...DataWork at office3 days per week- ...We're looking for a Staff UX Researcher to join our Design Research... ...with Software Design, Product, Data Science, and Engineering,... ...research across generative, evaluative, and strategic studies to understand... ...spaces — particularly AI- and LLM-powered recommendation...DataWork at officeLocal areaRemote work
$275k - $300k
...security validation, AI-augmented adversary... ...AI security research at Postman's scale.... ...are looking for a Principal Offensive Security... ...agentic workflows, and LLM integrations.You will... ...injection, tool-use abuse, data exfiltration via... ...You can architect evaluation harnesses and...PrincipalDataWork at officeFlexible hours3 days per week- ...Job Description Principal Machine Learning... ...Artificial Intelligence (AI) Required, Work... ..., inference, evaluation, and infrastructure... ...ML systems spanning data, training, evaluation... ...in partnership with research leadership. - Own... ...- Experience with LLM inference frameworks...PrincipalDataRemote workWork from home
$270k - $340k
Principal AI Research Scientist, Research Director - AI ScalingP-1227About Databricks... ...are obsessed with enabling data teams to solve the world’s... ...of large language model (LLM) training and inference... ...state‑of‑the‑art methods and evaluating trade‑offs in quality, latency...PrincipalDataLocal areaWorldwide$206.4k - $379.1k
...impressive content. The AI Foundations team... ...'re looking for a Principal Architect to build... ...distributed systems, data architecture, and... ...analytics, and continuous evaluation frameworks.This role blends applied research and engineering... ..., data pipelines, LLM orchestration...PrincipalDataFull timeTemporary workLocal areaWorldwideFlexible hours- ...Description Accellor is an AI-native services firm... ...through advanced AI, data, and engineering capabilities... ..., and internal research workloads. This role... ...engineering, cost optimization, evaluation gates, observability,... ...for large-scale LLM and GenAI workloads....PrincipalData
$167.4k - $310.8k
...us Roche.Advances in AI, data, and computational sciences... ...development. Roche’s Research and Early Development... ...seeking a Senior or Principal Machine Learning Scientist... ...strategies, and evaluation methodologies.Model Capability... ...and project ownership.LLM Expertise: Extensive...PrincipalDataFull timeLocal areaWorldwideRelocation package$164k - $190k
Notion is seeking an experienced User Researcher to enhance AI-powered experiences through in-depth research... ...executing various research methods to evaluate how users interact with AI features and requires robust communication and data fluency. Located in San Francisco or...Data- Cohere is looking for a Member of Technical Staff in Data Analysis and Evaluation to ensure the quality and performance of large language models. The role involves designing data collection tasks, collaborating with teams, and applying statistical methods for data evaluation...Data
- ...About the Team OpenAI's research training infrastructure... ...models are trained and evaluated. The Agent Harness... ...We're looking for a Principal Software Engineer to lead... ...OpenAI OpenAI is an AI research and deployment... ...possession (including the data contained therein) upon...PrincipalDataFull time
- Our client is a fast-growing AI consulting firm helping enterprises... ..., the firm is expanding its research team. Role Overview The AI... ...role focuses on identifying, evaluating, and advancing emerging AI technologies... ...research, machine learning, data science, or applied AI Strong...DataRemote workFlexible hours
- ...Stellarus is seeking a Data Scientist, Principal to lead the development and deployment of AI-powered products across enterprise teams. You will translate cutting-edge research into real-world applications, design production-grade models, and drive rapid feature delivery...PrincipalData
- AI Researcher Location: San Francisco About Hum.ai is building planetary superintelligence. Backed... ...large foundation models (beyond just LLM fine‑tuning). Who are we? Hum is a seed‑... ...remote sensing and real world ground truth data, and are used by our customers in nature...DataRemote work
- ...Grindr, we’re at the dawn of an AI-driven evolution, and you’ll be... ...to process vast conversation data and enhance user connections.... ...ideas into tangible results. Evaluate, influence, and integrate emerging... ...Expertise in building and maintaining LLM workflows for nuanced, human‑...PrincipalDataCasual workWork at officeImmediate startFlexible hours
- ...-3,000 new users a day. We're hiring an AI Researcher to build the next generation of real-time... ...question through distributed training, evaluation, and production deployment. Kotoba is a... ...speech models that run everywhere from the data center to edge devices, and we license...Data
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal AI Researcher - LLM Evaluation & Data. Be the first to apply!
- senior researcher San Francisco, CA
- machine learning researcher San Francisco, CA
- researcher San Francisco, CA
- senior design researcher San Francisco, CA
- design researcher San Francisco, CA
- qualitative researcher San Francisco, CA
- data collection researcher San Francisco, CA
- product researcher San Francisco, CA
- survey researcher San Francisco, CA
- vulnerability researcher San Francisco, CA



