Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Applied Data Scientist, LLM Evaluation United States (Remote) View Role

$175k - $275k

Driverai

Austin, TX
  • Remote job

Full-Time in Austin, TX Remote (any location) - Senior - Product & Engineering - $175k - $275k Applied Data Scientist, LLM Evaluation Introduction At Driver, we’re building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily asynchronous/distributed backend server, and a frontend web application that provides a rich user experience. About Driver We’re an early-stage startup backed by Y Combinator and Google Ventures that combines first principles technical approaches and applied LLM expertise to tackle context engineering at scale. Driver builds the context layer for employees and AI agents alike to use in developing software. Working at Driver Driver is an early-stage but fast-growing startup. As such, we take advantage of that which startups can excel: delivery speed, flexibility, and enjoying working with a small close-knit team. Organizational and engineering values at Driver include first-principles thinking, correct by construction, writing things down, experimentation and iteration, pragmatism, commitment to effective communication and transparency, autonomy, and ambition. Job Overview Title : Applied Data Scientist, LLM Evaluation Location: Remote or Austin, Tx Our value is directly tied to the quality of our content at scale. The platform generates technical documentation across a complex, multi-stage pipeline — producing multiple content types at different levels of abstraction, from individual code elements up to high-level summaries. Today, changes to models, context strategies, or pipeline architecture are evaluated largely through manual review and intuition. There is no systematic way to answer: “Did this change make our output better, worse, or the same — and for which languages, repo sizes, and content types?” This is a hard problem. LLM outputs are non-deterministic — identical inputs produce different outputs across runs, and small variations at early pipeline stages compound into meaningfully different end-user content downstream. Evaluating quality requires methodology that accounts for this: statistical reasoning over multiple runs, understanding of cascade effects through the pipeline, and rubrics that balance human judgment with automated signals. This role builds the evaluation function from scratch. You’ll define what “good” means for our generated content, build the infrastructure to measure it, and create the experimental framework that lets the team ship changes with confidence. What You’ll Do You’ll own the LLM evaluation strategy at Driver — from first principles to production infrastructure. This is a foundational role: you’re not joining an existing eval team, you’re building it. As the function matures, you’ll seed and grow a team around it. Define quality metrics and build evaluation datasets. Establish what “good” looks like for each content type across the pipeline. Build and curate gold-standard evaluation datasets across languages and repo archetypes (monorepos, microservices, libraries, applications). Design rubrics that capture accuracy, completeness, usefulness, and readability. Build benchmarking and experimentation infrastructure. Create automated evaluation pipelines that score output against reference datasets. Instrument the content generation pipeline to support A/B comparisons — run the same codebase through two strategies and compare results. Build tooling for LLM-as-judge evaluation and regression detection. Integrate evaluation into CI so pipeline changes come with quality evidence. Develop automated quality signals at scale. Build quality checks that flag degraded output without requiring human review of every document. Monitor content quality trends over time. Design sampling strategies for human review that maximize signal with minimal annotation effort. Quantify tradeoffs and inform decisions. Run experiments on model selection, context strategies, and pipeline architecture changes. Quantify cost/quality/latency tradeoffs. Partner with the engineering team to turn evaluation insights into shipped improvements. Qualifications Education: Bachelor’s, Master’s, or PhD in Statistics, Machine Learning, Data Science, Computational Linguistics, or a related quantitative field. Experience: Minimum 3 — 5 years in applied science, ML engineering, or data science roles with a focus on evaluation, NLP, or generative AI. 7+ years experience preferred. Required Technical Skills Strong statistical foundations: experimental design, hypothesis testing, confidence intervals, effect sizes, power analysis. Experience designing and running evaluations for LLM or NLP systems — you’ve thought carefully about what “better” means when outputs are open-ended text. Proficient in Python and the scientific/data stack (pandas, NumPy, scipy, sklearn). Comfortable working in Jupyter notebooks for exploration and prototyping, and turning that work into automated pipelines. Experience with LLM-as-judge approaches, inter-annotator agreement, and rubric design for subjective quality assessment. Familiarity with the practical challenges of non-deterministic systems: variance decomposition, multi-run methodology, distinguishing signal from noise at scale. Strong data storytelling — you can turn experiment results into clear recommendations that drive engineering and product decisions. Preferred and Nice-to-Have Technical Skills Experience with LLM APIs and prompt engineering across multiple providers. Familiarity with evaluation frameworks (e.g., RAGAS, DeepEval, custom harnesses). Experience building data pipelines or ETL workflows (Airflow, Dagster, or similar). Comfort with SQL and working directly against production data stores. Experience with visualization tools (Matplotlib, Plotly, Streamlit) for building internal dashboards and reports. Background in code understanding, developer tools, or technical documentation. Experience building or managing annotation pipelines and human evaluation workflows. Competitive Compensation Packages - Cash & Equity Flexible Work Culture Unlimited Time Off + 12 Paid Company Holidays Life Insurance & FSA Accounts 401(k) Retirement Accounts - Traditional, Roth, or Both Quarterly Team Offsites Driver is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. #J-18808-Ljbffr Driverai

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Applied Data Scientist, LLM Evaluation United States (Remote) View Role in Austin, TX vacancy
  • $150k - $175k

     ...Applied Data Scientist, Health AI Evaluation & Datasets Remote - United States Innodata is a global data engineering company. We believe...  ...customers. Scope of the Role: Healthcare is one of the...  ...pipelines, including rubric-grounded LLM-as-judge prompts, regression... 
    Remote work
    Shift work

    Innodata Inc.

    United States
    5 days ago
  •  ...zone is seeking an AI Research Scientist to lead applied AI research projects for US-...  ...measurable experiments in LLM evaluation and RLHF data design. You will evaluate prompts...  ...safety, and usefulness. The role is remote in the United States, with compensation based on... 
    Remote job
    Hourly pay
    Flexible hours

    AIToolboard

    New York, NY
    2 days ago
  • $112k - $269k

     ...Additionally, you will apply traditional...  ...turning raw data into valuable...  ...is fully remote and does not...  ...particular state within the US...  ...LLMs, utilizing LLM APIs (OpenAI,...  ...and evaluation.A Bachelor’s...  ...range for this role to be between...  ...restricted stock units, and benefits... 
    Remote work
    Work experience placement
    Local area

    Yelp

    Austin, TX
    4 days ago
  • $166.8k

     ...The AI and Data Analytics Division...  ...a Senior Data Scientist. This role blends...  ...mission focused applied R&D projects that...  ...strategies) and evaluation (T&E, robustness...  ...such as remote sensing imagery...  ...experience with LLM/LVM/Foundation...  ...loyalty to the United States. The investigation... 
    Remote work
    For contractors
    Work experience placement
    Work at office
    Local area
    Relocation package
    Flexible hours

    Pacific Northwest National Laboratory

    Richland, WA
    5 days ago
  • Driverai is seeking an Applied Data Scientist with expertise in LLM evaluation to join its innovative team in Austin, TX. This role focuses on building the evaluation function from scratch and requires a strong background in statistics and machine learning. The successful... 
    Remote job

    Driverai

    Austin, TX
    1 day ago
  • ## Data Center Optical Technician - New AlbanyMaumee,Ohio,United StatesFind out how well you match...  ...The person in this role is responsible for...  ..., break/fix, and remote hands services, utilizing...  ...happens once you apply?** Click Here to...  ...city:** United States (US) || Ohio... 
    Remote work
    Temporary work
    Work at office
    Immediate start

    Ericsson GmbH

    New Albany, OH
    2 days ago
  • ## Sr Data Engineer6380 Rogerdale Rd, Houston, TX, United States, 77072Apply NowGet Job MatchesJob...  ..., including remote and hybrid options...  ...also play a crucial role in identifying opportunities...  ...closely with data scientists, analytics teams,...  ...any areas above, apply anyway! We love to... 
    Remote work
    Summer work
    Work at office
    Monday to Friday
    Flexible hours
    Weekend work

    Tailored Brands, Inc.

    Houston, TX
    2 hours ago
  • $174.99k - $209.98k

     ...Grafana Labs is a remote-first, open-...  .... If this role excites you,...  ...architectures, LLM integrations,...  ...third-party data platforms....  ...iteration, model evaluation, and cost...  ...on experience applying LLMs/AI to production...  ...fan-out), state management,...  ...: In the United States, the... 
    Remote job
    Full time
    Local area

    Grafana Labs

    United States
    11 hours ago
  • $154k - $200k

     ...looking for a Senior Data Scientist, Applied ML to design,...  ...reliable systems. This role is ideal for...  ...monitoring and evaluation, designing feedback...  ..., flexible and remote-friendly work options...  ...under federal, state, or local law....  ...information, political views or activity, or... 
    Remote work
    Full time
    Temporary work
    Work experience placement
    Local area
    Worldwide
    Visa sponsorship
    Flexible hours

    SpyCloud

    Austin, TX
    11 hours ago
  • $230k - $322k

     ...Engineer, Shopping Ads Remote - United States Reddit is a...  ...technical leadership role for an engineer...  ...sizing, data and label design,...  ...selection, offline evaluation, online experimentation...  ...delivery stack. Apply and adapt state-of...  ..., and political views. We invite you to... 
    Remote job
    For contractors
    Work experience placement
    Shift work

    Reddit, Inc.

    Brooklyn, NY
    2 days ago
  • $173.1k - $303k

     ...development, including data curation, training, and evaluation. Our goal is...  ...Business Units (BUs) within...  ...do in this role: Confronted...  ...creativity to apply existing...  ...applied research scientists, product managers...  ...developing LLM based...  ...personas (flexible, remote, or required... 
    Remote work
    Work experience placement
    Work at office
    Flexible hours

    Victrays

    Santa Clara, CA
    3 days ago
  •  ...Services is hiring an AI Data Strategy Engineer / Applied Scientist, LLM Data, to own data...  ...workflows, and evaluation datasets powering...  ...AI systems. The role covers acquisition...  ...data generation. Remote work options available...  .... #J-18808-Ljbffr United States Digital Space LLC
    Remote job

    United States Digital Space LLC

    New York, NY
    2 days ago
  •  ...Scottsdale, Arizona, United States Join Axon and...  ...TASER device data, build models...  ...degrees — applied mathematics and...  ...Location: This role is based out of...  ...flexibility to work remotely on Mondays,...  ...practices — the data scientists here write code...  ...a long-term view because we want... 
    Remote work

    Axon Enterprise

    Scottsdale, AZ
    1 day ago
  • $148.19k - $231.98k

     ...please consider applying for a maximum of 3 roles within 12...  ...TeamWe are a Data & AI Architects...  ...point of view, and help guide...  ...for customers evaluating, implementing...  ...not just the stated requirements....  ...education.In the United States,...  ...York; Florida - Remote; Illinois -... 
    Remote work
    Full time
    Work experience placement
    Immediate start

    Salesforce

    Irvine, CA
    3 days ago
  •  ...The Role This remote Cloud Platform engineering role enables AWS Data Lake initiatives by designing, automating...  ...modules, remote state, versioning, automated...  ...enablement. Experience applying observability, logging...  ...authorization in the United States and with MHP Americas... 
    Remote work
    Work at office

    Dr. Ing. h.c. F. Porsche AG

    Remote
    a month ago
  •  ...Welo Data is looking for English speakers to join a remote project as a Search Quality Rater. In this role, you will help improve how search engines understand...  ...search results and evaluate how helpful and relevant...  ...Applicants must be of at least 18 years of age to apply.... 
    Remote work
    Part time
    Currently hiring
    Immediate start
    Monday to Friday

    Welo Data

    United States
    more than 2 months ago
  • $160k - $174k

     ...Start your Voyage -Apply NowGet to Know...  ...the Snowflake AI Data Cloud. This role combines AI engineering...  ...can work in a remote or hybrid...  ...retrieval tuning, and evaluation.AI Engineering &...  ...Prompt Engineeringo LLM Evaluationo AI...  ...status protected by state or local law.... 
    Remote work
    Full time
    Part time
    Work experience placement
    Local area
    Flexible hours

    Benefitfocus

    New York, NY
    4 days ago
  • $175k - $225k

    Role Description Innodata is expanding its GenAI research...  ...capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-...  ...with Language Data Scientists and AI/ML Research... 
    Full time

    Innodata Inc.

    Remote
    1 day ago
  • $30 - $50 per hour

    Rex.zone is hiring mid-senior data annotators to support US-based AI/ML projects across NLP, LLM evaluation, RLHF, content safety, and computer vision. You will produce...  ...measurable model performance improvement. Fully remote role aligned with Miami job-search intent,... 
    Remote job
    Hourly pay

    Rex.zone

    Miami, FL
    1 day ago
  •  ...This is a Fully Remote Opportunity within the United States The Data Scientist will design and implement...  ...in this role will use statistics...  ..., deep learning, LLM frameworks, and domain...  ...AI solutions, applying python LLM...  ...databases, creating and evaluating statistical... 
    Remote work
    Work experience placement

    City of Hope

    United States
    3 days ago
  • $142.3k - $263.3k

    Senior Applied Scientist - AI Evaluation & Quality Systems Seattle, Washington, United States Machine Learning and AI...  ...powers the AI and LLM features behind...  ...Human‑centered AI, Data Quality...  ...services. In this role, you will develop...  ...strong point of view on when not to use... 
    Relocation
    Shift work

    Apple Inc.

    Seattle, WA
    3 days ago
  • $82.57k - $127.49k

     ...opportunity for a Data Scientist within our...  .... This role focuses on developing...  ...is open to remote or hybrid....  ...algorithms to apply to data setsDesign...  ...and evaluate models using...  ...cases, including LLM-based solutions...  ...: federal, state, and local minimum...  ...across the United States... 
    Remote work
    Minimum wage
    Full time
    Local area
    Flexible hours

    CorVel

    Irvine, CA
    3 days ago
  • $99k - $225k

    Applied Data ScientistThe Opportunity:As a data scientist, you’re excited about building intelligent...  ..., training, evaluation, and deploymentExperience...  ...during meetings.Remote: If this position...  ...the needs of the role. You may also be...  ...federal, state, local, or international... 
    Remote work
    Full time
    Contract work
    Part time
    Work at office
    Local area

    Booz Allen Hamilton

    Arlington, VA
    2 days ago
  •  ...experienced Data Scientists and Senior Data...  ....Please apply here for all...  ...recruiting team by evaluating job related...  ...applicable state laws. In...  ...Located in NYC or Remote Jobs...  ...audit may be viewed here: CoveyCompensationThe...  ...for this role includes...  ...within the United States,... 
    Remote work
    Hourly pay
    Work at office
    Local area
    Flexible hours

    Doordash

    Los Angeles, CA
    4 days ago
  •  ...us!Why this role?Evaluation is critical...  ...to measure LLM progress.As...  ...Research Scientist, Model Evaluation...  ...the state-of-the-art...  ...LLM-based data synthesis pipelines...  ...if you are remote, plus an...  ...you to apply. We strive...  ...If jobs are viewed on other sites...  ...; Seattle; United StatesEmployment... 
    Remote work
    Full time
    Work at office
    Local area
    Home office

    Cohere

    New York, NY
    1 day ago
  • $143k - $286k

     ...Location: BENTONVILLE, AR, United States; SUNNYVALE, CASalary...  ......As a Principal Data Scientist on the Applied AI team, you will...  ...code, design the evaluations, and make the architecture...  ...the real world. The role suits engineers who...  ...agentic workflows, LLM applications, and... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    1 day ago
  • $142.8k - $274.8k

     ...6-08-07Location: United States, Washington, RedmondSalary...  ...: Research, Applied, & Data...  ...through rigorous evaluations, fine-tuning, and...  ...practitioners and data scientists to join agile, cross...  ...challenges. This role offers the opportunity...  ...machine learning, and LLM techniques, and... 
    Ongoing contract
    Local area
    Worldwide
    3 days per week

    Microsoft

    Redmond, WA
    1 day ago
  • $80k - $140k

     ...OPPORTUNITY? As a Data Scientist you will...  ...applications using state of the art tools...  ...Management Applied AI team, your day...  ...offline model evaluation and validation....  ...Ability to evaluate LLM outputs....  ...plays a critical role in attracting,...  ...MinneapolisCountry:United States of... 
    Full time
    Flexible hours

    Royal Bank of Canada

    Minneapolis, MN
    1 day ago
  •  ...EngineerPlano,Texas,United StatesFind out how...  ...Grow with us!* **This role is a hybrid...  ...than a traditional data engineer. The job starts...  ...models and semantic views that ground our...  ...ve built. When you apply, share your automation...  ...and city:**United States (US) || **Hybrid:**... 
    Temporary work
    Work at office
    Immediate start
    Relocation
    3 days per week

    Ericsson GmbH

    Plano, TX
    2 days ago
  •  ...graduates. About the Role You'll work directly...  ...(agents, pipelines, evaluation tools), debug real systems...  ...and development of LLM-powered pipelines for agentic...  ...for crypto-native data sources — on-chain data...  ...adversarial robustness. Apply AI-native development... 
    Remote work
    Full time
    Internship
    Work from home

    binance

    United States
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Applied Data Scientist, LLM Evaluation United States (Remote) View Role. Be the first to apply!