Applied AI Researcher, Benchmarking
$150k - $250kDistyl AI
Distyl AI Job Opportunity
Distyl is an applied AI technology company partnering with the world's most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations. We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys. Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s. The results reflect this approach: a 100% production deployment success rate for our customers and one of the few enterprise AI companies to run a profitable business.
What We Are Looking For
At Distyl we're pushing the envelope of AI utilization in enterprise. This requires creative researchers who don't just want to drive incremental improvements on benchmarks or optimize an existing process but instead are looking to creatively redefine how software is used.
Our researchers come from many academic backgrounds but have strong research track records, operate in an AI-native way, and would be bored staying on the rails of a traditional research org.
Key Responsibilities
The Benchmarking team defines how progress is measured. Researchers design evaluation frameworks that capture reasoning depth, interaction quality, reliability, and operational impact. They construct benchmarks that reflect real-world complexity. Their systems become the standard by which new architectures, techniques, and releases are judged.
Researchers in Benchmarking explore new paradigms for evaluating intelligent systems: adversarial robustness testing, longitudinal performance tracking, and human-in-the-loop assessment. They investigate how metrics shape model behavior and establish rigorous methodologies for quantifying emergent capability. Their insights drive both Distyl's internal research priorities and industry-wide standards.
Who You Are
Experience Designing and Running Evaluations: You've built or maintained benchmarks, test suites, or experimental frameworks to measure model or system performance
Statistical and Analytical Rigor: You design fair, reproducible experiments and can extract signal from noisy empirical results
Experience Building with Models, Not Just Building Models: We develop intelligent systems using models rather than training or fine-tuning them. Ideal candidates have expertise in compound AI systems, agentic collaboration, and associated techniques (ensembling, ReAct, graph-of-thoughts, etc.)
Proven Track Record of Research Results: Whether you've published in top journals, posted amazing work on Twitter, or somewhere else we want to see what you've done
Uses AI Every Day: Before you can revolutionize someone else's workflow, you need to revolutionize yours. You should be using tools like ChatGPT, Cursor, and Perplexity to accelerate your workflow
Strong Programming and Data Analysis Skills: While you might not consider yourself a software engineer you need to be able to build prototypes of your ideas and then perform the experiments to prove the effectiveness to a F500 Head of AI
Biases Towards Showing vs Telling: Our customers want to see the power of AI today vs discuss the most elegant idea that will take 5 years to realize
What We Offer
The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package
100% coverage of medical, dental, and vision insurance for employee and dependents
Flexible time off
Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources
Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot
Complimentary in-office lunches and snacks provided
Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems
Ownership of high-impact projects across top enterprises
A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence
Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in-office.
We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.
$262.5k - $299.6k
...we are creating trustworthy and reliable AI systems, changing banking for good. For years... ...We are committed to building world‑class applied science and engineering teams and... ...life. Our work touches every aspect of the research life cycle, from partnering with academia...SuggestedFull timePart timeLocal area- ...we are creating trustworthy and reliable AI systems, changing banking for good. We are... ...to banking. We are building world-class applied science and engineering teams and scalable... ...to life, touching every aspect of the research lifecycle, from partnering with academia...SuggestedFull timePart time
$40k - $200k
...the core team building truly AI-native helpful experiences across... ...is next in the AI space but apply it to something of true value... ...& highly ambitious AI/NLP researcher as core employees to help build... ...than working on beating AI benchmarks! :) Every day, you'll get...SuggestedVisa sponsorshipFlexible hours$192.2k - $260k
...responsible for defining and delivering AI-powered solutions that transform how... ...team is seeking an exceptional Applied Scientist to research and develop novel approaches for agent... ...skills- Establish evaluation metrics and benchmarks for agent-data interaction...SuggestedLocal areaImmediate startFlexible hoursNight shiftDay shift$218.7k - $249.6k
...Overview Applied Researcher I (AI Foundations, LLM Core and Agentic AI) Overview: At Capital One, we are creating trustworthy and reliable AI systems, changing banking for good. For years, Capital One has been leading the industry in using machine learning...SuggestedFull timePart timeLocal areaFlexible hours$262.5k - $299.6k
...Overview Applied Researcher II (AI Foundations) Overview: At Capital One, we are creating trustworthy and reliable AI systems, changing banking for good. For years, Capital One has been leading the industry in using machine learning to create real-time, intelligent...Full timePart timeLocal areaFlexible hours- JPMorganChase AI Research is a global team of research scientists, engineers and product managers... ...with internal data scientists, applied engineering teams and stakeholders across... ...design and creation of meaningful benchmarks and metrics for evaluation.Effective verbal...Work at officeShift work
$197.3k - $313.7k
...best candidate experience, please consider applying for a maximum of 3 roles within 12 months... ...SalesforceSalesforce is the #1 AI CRM, where humans with agents drive customer... ...Salesforce.Join a collaborative, diverse team of researchers at Agentforce Operations. The...Full timeImmediate startRemote work- ...team. This Data Scientist / Researcher will collaborate closely with... ...providing portfolio companies with benchmark data to help them identify... ...real impact on the world by applying data science to VC and tech... ...Capital’s areas of interest (e.g., AI, enterprise software, deep...Full timeWork experience placement
- ...technology have finally made it possible for AI to impact clinical care in a meaningful... ...the right time.Translational Scientist, Applied Machine Learning and Agentic AI, Pharma R... .... Teamwork and collaboration: Work with Research, Engineering & Data Science teams across...Full timeRemote workShift work
$228.7k - $309.4k
...Services (AWS) is looking for a Principal Applied Scientist to join the Quick Science team. Quick is AWS’s enterprise generative AI assistant that helps users answer questions... ...a key member of this team, you will lead research and development efforts in generative AI and...Local areaFlexible hours$183.8k - $248.7k
...professional users—spanning enterprises, research institutions, and government agencies—through... ...environments by deploying agentic AI solutions that bridge the gap between generative... ...future-proof their workflows.The Senior Applied Scientist will lead the development of...WorldwideFlexible hours$159k - $230k
Drive research strategy by independently identifying, prioritizing,... ...data into frameworks.Establish benchmarking and measurement programs to... ....6 years of experience in an applied research setting (e.g., product... ...operates, from accelerating AI/ML deployments to reducing...$196k - $230k
...AreNotion is the collaborative AI workspace where teams and... ...’re seeking an experienced UX Researcher to define and scale how we evaluate... ..., and data science can apply consistently.This role can be... ...help teams spot regressions, benchmark improvements, and understand when...Local areaShift work$262.5k - $299.6k
...Overview: Applied Researcher II At Capital One, we are creating trustworthy and reliable AI systems, changing banking for good. For years, Capital One has been leading the industry in using machine learning to create real-time, intelligent, automated customer experiences...Full timePart timeLocal areaFlexible hours$218.7k - $249.6k
...we are creating trustworthy and reliable AI systems, changing banking for good. For years... ...We are committed to building world‑class applied science and engineering teams and... ...life. Our work touches every aspect of the research life cycle, from partnering with academia...Full timePart timeLocal areaFlexible hours$155k - $285k
Quant Researcher - Agentic AI CTO Office Location New York Business Area Engineering and CTO Ref # 10050703 Description &... ...financial industry conducts research, trading, and risk management. Apply advanced ML techniques in an extremely rich problem space...Temporary workFor contractorsWork experience placementWork at office$147.6k - $274k
.... That’s what makes us Roche.Advances in AI, data, and computational sciences are transforming... ...drug discovery and development. Roche’s Research and Early Development organisations at... ...Learning Scientist building agents for applied small-molecule drug design. You will...Full timeLocal areaWorldwideRelocation package$200k
...States Type: Permanent ML Researcher / ML Engineer I am working with... ...and engineers, developing advanced AI models and production-scale machine learning... ...optimisation Agentic AI systems Applied machine learning research Model evaluation...Permanent employmentFull time- ...Sr. Distinguished Applied Researcher Capital One Overview At Capital One, we are creating trustworthy and reliable AI systems, changing banking for good. For years, Capital One has been leading the industry in using machine learning to create real‑time, intelligent, automated...Flexible hours
$140.39k - $166.1k
...ROLE Peloton is seeking a User Researcher to support the innovation,... ...studies, concept evaluations, benchmarking, heuristic evaluations) on Peloton... ...can easily be referenced and applied for future projects WHAT YOU... ...Comfortable incorporating AI tools into research workflows...Temporary workLocal areaRemote work$180k - $350k
Hume AI is seeking talented AI researchers interested in working with our team to build state‑of‑the‑art speech‑language models (SLMs). Our new SLM... ...support distributed model training, inference, and benchmarking as well as massive‑scale data collection, storage, preprocessing...$175k - $250k
As an AI Researcher at Vatic Labs, you will research and develop innovative AI-driven quantitative trading strategies. You will explore vast amounts of market and alternative data, inventing and applying a new generation of state-of-the-art technologies that are inspired...Work at officeNight shift- Sr. Distinguished Applied Researcher Capital One Overview At Capital One, we are creating trustworthy and reliable AI systems, changing banking for good. For years, Capital One has been leading the industry in using machine learning to create real‑time, intelligent, automated...Flexible hours
$150k - $185k
The Role: The Senior Researcher for Commerce and Growth User Experience... ...observational tests, surveys, benchmark studies, diary studies, and... ...innovative research methodologies and AI-enabled tools to streamline... ...methods and when and how to apply them Experience enabling and...Full timeWork at officeLocal areaFlexible hours$200k - $300k
Hudson River Trading (HRT) is hiring an AI Researcher to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible for developing and maintaining our most powerful models, which are used by our trading teams to drive a significant fraction of our trading. We...Work experience placementWork at officeLocal areaImmediate start$200k - $300k
Hudson River Trading (HRT) is seeking an LLM-focused AI Researcher to join the HAIL team. HAIL (HRT AI Labs) is the team at HRT responsible for developing and maintaining our most powerful AI models, which are used by our trading teams to drive a significant fraction of...Work at officeLocal areaImmediate start$168k - $264.5k
...computational engine for groundbreaking research, enabling scientists to model... ...is a unique opportunity to apply your expertise in a fast-paced... ..., and digital twins.Designing benchmarks and evaluation methods for... ...existing vacancy. NVIDIA uses AI tools in its recruiting...Full timeRemote workShift work$127.4k - $236.6k
Senior Applied Scientist, Document UnderstandingAbout the RoleThis is an applied science position... ...a PhD or Master's in Computer Science, AI, NLP, or a related field, with 5+ years... ..., lead through influence in an applied research setting, and measure success by what...Full timeLocal areaFlexible hours$141.1k - $262.1k
...what makes us Roche.Advances in AI, data, and computational... ...discovery and development. Roche’s Research and Early Development... ...evaluation criteria.Evaluation & Benchmarks: Design and implement evaluation... ...ML libraries.A passion for applying frontier AI to drug discovery...Full timeWork experience placementLocal areaWorldwideRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Applied AI Researcher, Benchmarking. Be the first to apply!
- content researcher New York, NY
- senior researcher New York, NY
- machine learning researcher New York, NY
- researcher New York, NY
- senior design researcher New York, NY
- design researcher New York, NY
- qualitative researcher New York, NY
- data collection researcher New York, NY
- product researcher New York, NY
- survey researcher New York, NY


