Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer L5/L6 — Model Evaluations & Data Curation (MEDC)

Full-time

jobgether

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Software Engineer L5/L6 — Model Evaluations & Data Curation (MEDC) based in the United States.

The Software Engineer will build foundational infrastructure that accelerates how AI teams create, evaluate, and improve high-quality datasets.
You will transform ad-hoc data curation workflows into reusable, scalable platforms that researchers and engineers can rely on.
The role combines strong software engineering with hands-on work in LLM-driven data generation, evaluation, sampling, and quality control.
You will work closely with researchers and data scientists to understand how curation decisions influence model behavior and performance.
Your work will help make datasets more discoverable, reproducible, versioned, and production-ready across multiple AI initiatives.
The environment is highly collaborative and technically ambitious, with significant autonomy and opportunities to influence engineering practices.
At the L6 level, the role additionally calls for technical leadership and the ability to establish direction across data and evaluation infrastructure.

Accountabilities

  • Design and build reusable data curation infrastructure, including shared libraries, components, and workflows that replace fragmented notebook-based processes.

  • Develop scalable LLM-powered pipelines that transform raw catalog, metadata, and other data sources into training and evaluation datasets such as question-answer pairs and synthetic scenarios.

  • Implement large-scale batch inference workflows while balancing data quality, computational efficiency, token usage, and cost.

  • Develop sampling strategies that optimize coverage, diversity, difficulty, and representation across relevant content and member segments.

  • Create data-quality and filtering systems using techniques such as LLM-as-judge scoring, evaluation-model-based ranking, deduplication, validation, and other quality controls.

  • Partner closely with researchers to design experiments that measure how data curation choices affect downstream model behavior and performance.

  • Establish curated datasets as discoverable, reusable artifacts with clear versioning, lineage, documentation, and reproducibility.

  • Drive adoption of standardized data curation practices across engineering, research, and modeling teams.

  • At the L6 level, provide technical leadership across data and evaluation infrastructure and help define technical direction for multi-engineer initiatives.

Requirements

  • Strong software engineering expertise in Python, including experience developing reusable infrastructure, libraries, frameworks, or platforms used by other engineers and researchers.

  • Hands-on experience building LLM-driven data generation or transformation pipelines, including synthetic data generation, structured outputs, or large-scale batch inference.

  • Practical experience with data quality techniques such as sampling, filtering, deduplication, validation, and model-based quality scoring, including LLM-as-judge approaches.

  • Strong modeling intuition and an understanding of how dataset composition and curation decisions can influence model behavior and performance.

  • Experience designing experiments or evaluation approaches to measure the impact of data and modeling decisions.

  • Experience with distributed data processing technologies such as Spark, Ray, or comparable frameworks.

  • Excellent collaboration and communication skills, particularly when partnering with researchers, data scientists, and platform engineering teams.

  • For L6 roles, demonstrated experience with LLM evaluation systems is required.

  • For L6 roles, demonstrated technical leadership across data or evaluation infrastructure, including setting technical direction for multi-engineer initiatives, is required.

  • Experience with dataset versioning, lineage, artifact management, experiment tracking, or model registries is highly valued.

  • Experience optimizing large-scale LLM inference for cost, throughput, or operational efficiency is a plus.

  • Familiarity with human annotation workflows and methods for calibrating LLM judges against human ratings is beneficial.

  • Experience with pipeline orchestration frameworks such as Metaflow, Airflow, or similar tools is advantageous.

  • Background in recommendation systems, personalization, search, content catalogs, or metadata-driven applications is a plus.

Benefits

  • Annual compensation range of $600,000–$1,066,000 , with the range varying based on location and individual market factors.

  • Compensation is structured primarily around annual salary, with the flexibility to determine the desired balance between salary and stock options each year.

  • Comprehensive health insurance plans and mental health support.

  • 401(k) retirement plan with employer matching.

  • Stock option program.

  • Health Savings Accounts and Flexible Spending Accounts.

  • Family-forming benefits.

  • Life and serious injury benefits.

  • Disability programs.

  • Paid leave of absence programs.

  • Flexible paid time off for full-time salaried employees.

  • Remote work opportunity within the United States.

  • Opportunity to work on high-impact AI infrastructure spanning foundation models, evaluation, and data curation.

  • Collaborative environment with substantial technical autonomy and opportunities for senior-level technical leadership.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Software Engineer L5/L6 — Model Evaluations & Data Curation (MEDC) in United States vacancy
  •  ...what’s next. About the Team Model Evaluations and Data Curation ('MEDC') forms the flywheel of foundation...  ...the Role We are looking for a Software Engineer to build the common infrastructure...  ...evaluation systems (must-have for L6) Technical leadership across data... 
    Data
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    Remote
    21 hours ago
  • $388k

     ...GenAI workloads, capturing model inputs, features,...  ...monitor model performance, data quality, drift, latency...  ...for a hands-on senior engineer to build the frameworks...  ..., model performance, evaluation, and vendor integration...  ...will need:Experience in software, AI/ML, or platform engineering... 
    Data
    Hourly pay
    Full time
    Immediate start
    Flexible hours

    Netflix

    Los Gatos, CA
    2 days ago
  • $136.44k - $265.11k

     ...structured scientific data and AI are built into...  ...run AI agents and models directly in their workflows...  ...build the datasets, evaluations, and systems that...  ...the intersection of software engineering, biology, and frontier...  ...creating pipelines that curate, transform, and... 
    Data
    Work at office
    Local area
    Monday to Friday
    Shift work

    Benchling

    San Francisco, CA
    2 days ago
  • $272k - $431.25k

     ...generating it! Our world model team is pushing the...  ...Manager to lead world-model evaluation and benchmarking...  ...findings into better data, training recipes, model...  ..., post-training, data curation, simulation, robotics,...  ...Science, Electrical Engineering, Robotics, Machine Learning... 
    Data
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $70 per hour

     ...in simulation across 15+ U.S. states. Software Engineering builds the brains of Waymo's fully autonomous...  ...simulation logs, vehicle trajectory data, and multi-agent behavioral events into...  ...) Develop temporal graph modeling techniques to capture time-varying multi... 
    Data
    Hourly pay
    Full time
    Internship
    Summer internship
    Relocation package

    Waymo

    Mountain View, CA
    2 days ago
  • $255k - $300k

     ...ChatGPT launched. The job of the Model Capabilities team is to keep...  ...for our users and our engineers.Make inference reliable: better...  ...better" means, build or use the evaluation to test it, and are willing...  ...be proven wrong by your own data.Cost and performance instincts... 
    Data
    Local area

    Notion Labs

    New York, NY
    4 days ago
  •  ...opportunity for you to take your software engineering career to the next level. As...  ..., and work across cloud, data, and machine learning...  ...that support end-to-end ML model lifecycle — from development...  ...demonstrated ability to critically evaluate and validate AI-generated... 
    Data
    Work at office

    JP Morgan Chase

    Plano, TX
    2 days ago
  • $145k - $200k

     ...CompanyPalantir builds the world’s leading software for data-driven decisions and operations...  ....The Role We are a software engineering team with expertise in enabling ML models in production. We deploy AI...  ...and the ability to quickly evaluate and integrate new models and... 
    Data
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    2 days ago
  •  ...development of high-quality datasets and evaluation pipelines that improve and benchmark large language models for code generation and software engineering tasks. You will curate and author reference code,...  ...this by providing high-quality data, advanced training pipelines,... 
    Data
    Full time
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    SaidGig

    United States
    more than 2 months ago
  •  ...partner is looking for a Senior Software Engineer – LLM Evaluation based in United States. This part...  ...the evaluation of large language models. You will curate high-quality code, develop...  ...interest and wish you the best! Data Privacy Notice: By submitting your... 
    Data
    Full time
    Part time
    For contractors
    Remote work
    10 hours per week
    Flexible hours

    jobgether

    United States
    3 days ago
  • $125k - $150k

     ...ECS is seeking an AI Model Engineer to work in a hybrid remote...  ...a track record of evaluating, experimenting with, and...  ...robust retrieval and data fusion.Coordinate the...  ...oversight on data ingestion, curation, and storage to...  ...years of experience in software engineering and data engineeringExperience... 
    Data
    Contract work
    Work at office
    Remote work

    ECS Federal

    Fairfax, VA
    3 days ago
  • $86.8k - $198k

    Model and Simulation Software EngineerThe Opportunity: You will play a critical role...  ...DoW. Using core software engineering principles, you’ll support...  ..., and integrate diverse data sources to drive high‑fidelity...  ...generation M&S systems by evaluating new frameworks, enhancing... 
    Data
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Suffolk, VA
    3 days ago
  • $295k

     ...cutting-edge foundation AI models and end-to-end...  ...team of researchers, engineers, designers, and more,...  ...Join us!Role Overview:Evaluation is critical to making...  ...judges; refining LLM-based data synthesis pipelines; and...  ...about.You have strong software engineering skills.... 
    Data
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    1 day ago
  • $127k - $191k

    What You'll DoAs a Sr Engineer II (Team Leader, Model Operations & Enablement), you will lead the team that...  .... You will work closely with data science, engineering, and business stakeholders...  ...tools to assist in reviewing and evaluating job applications, fraud prevention, and... 
    Data
    Hourly pay
    Permanent employment
    Temporary work
    Work experience placement
    H1b
    Work at office

    Principal Financial Group

    Raleigh, NC
    4 days ago
  •  ...We are seeking an experienced Engineer, AI – AI Evaluation & Model Risk Lead to lead how AI models are evaluated...  ...AI tools, Python, cloud computing, data analysis, data modeling, model...  ...in relevant AI, machine learning, software development, or data-focused roles.... 
    Data
    Full time
    Work experience placement

    BrickRed Systems

    Washington DC
    3 days ago
  • $60 per hour

     ...simulation across 15+ U.S. states. Software Engineering builds the brains of Waymo's...  ...this role will be: The Model Eval team has a lot of...  ...Querying existing logs and data sources for aggregate...  ...fundamentals of training and evaluating models The ability to use... 
    Data
    Hourly pay
    Full time
    Internship
    Summer internship
    Relocation package

    Waymo

    Mountain View, CA
    1 day ago
  • $70 - $90 per hour

     ...Role Overview Help evaluate Neuron Kernel Interface development tasks...  ...and evaluation of advanced AI models. You will assess kernel...  ...pipeline utilization, tensor-engine throughput, and memory-bandwidth...  ...and FP32, BF16, FP8, and INT8 data types. Experience benchmarking... 
    Data
    Hourly pay
    Remote work

    SaidGig

    Remote
    a month ago
  • $130k - $260k

     ...improve through rigorous evaluation and model post-training. This...  ...an analytics-focused data science role. It is a...  ...hands-on AI systems engineering position focused on...  ...operating production software. As the Distinguished...  ...synthetic data, and curated evaluation sets—into... 
    Data
    Full time
    Contract work
    Temporary work
    Part time

    Walmart

    Bentonville, AR
    2 days ago
  •  ...the Organization The Evaluation team builds and evolves...  ...approaches that enable data-driven decisions...  ...into clear feedback for engineering and leadership, and help...  ...introspect autonomous driving software performance at...  ...prediction, and planning models. Build and maintain... 
    Data
    Full time
    Local area
    Work from home

    General Motors

    Sunnyvale, TX
    3 days ago
  • $125k - $175k

     ...NinjaTrader equips traders with award-winning software and brokerage services to navigate...  ...ll do:We are seeking a Sr. Software Engineer to join our Evaluation Services team. Our Evaluation...  ...on authentication, endpoints, data models, and architectural best practicesPartner... 
    Data
    Work at office
    Remote work
    Worldwide
    Monday to Friday
    Flexible hours
    Shift work

    NinjaTrader Group

    Chicago, IL
    4 days ago
  • $60 - $80 per hour

     ...help develop advanced large language models. In this role, you will bring...  ...campaign judgment to AI training data, partnering with research and engineering teams to improve model performance...  ...reasoning quality. Develop and improve evaluation guidelines and scoring rubrics for... 
    Data
    Hourly pay
    Weekday work

    SaidGig

    United States
    more than 2 months ago
  • $100 per hour

     ...applications by providing rigorous, real-world analysis, evaluation, and feedback. This remote, part-time contract role focuses...  ...decision-making. Assess and annotate complex financial data, reports, and model outputs using detailed rubrics and established best practices... 
    Data
    Hourly pay
    Contract work
    Part time
    For contractors
    Remote work

    SaidGig

    United States
    more than 2 months ago
  • AuraOne is seeking a remote Generalist for Data Annotation to evaluate prompts and responses against a quality rubric. You’ll help turn outputs into...  ..., and regression cases for Human Data, reviewing frontier model outputs and calibrating evaluators. Responsibilities... 
    Data
    Remote job

    AI Trainer Jobs

    New York, NY
    4 days ago
  • Surgical Planning Safety Evaluator is a remote evaluation track for reviewing surgical planning safety evaluation prompts...  ...cases, and write the kind of structured feedback the modeling team can use to retrain. AI data reviewers help turn surgical planning safety... 
    Data
    Remote job

    AI Trainer Jobs

    New York, NY
    5 days ago
  • Ignite IT Hub is seeking a senior Model Evaluation Analyst to assess the accuracy and reliability of the U.S. Census Bureau's LLM Autocoder....  ...produce evidence for Government review. Working with the contract’s Data Scientist and Government SMEs, you will evaluate enhancements... 
    Data
    Contract work
    Work at office

    Ignite IT

    Suitland, MD
    4 days ago
  • Receipt and Invoice Understanding Model Evaluator is a remote evaluation track for reviewing receipt and invoice understanding model evaluation...  ...structured feedback the modeling team can use to retrain. AI data reviewers help turn receipt and invoice understanding model... 
    Data
    Hourly pay
    For contractors
    Remote work
    10 hours per week

    AI Trainer Jobs

    New York, NY
    5 days ago
  • $15 - $20 per hour

     ...public sources and external tools. Generate high-quality human evaluation data by identifying response strengths, areas for improvement, and...  ...quality, clarity, tone, and completeness of responses. Ensure model responses align with expected conversational behavior and system... 
    Data
    Contract work
    Summer work
    Remote work

    Remote Jobs

    New York, NY
    5 days ago
  • $204k - $259k

     ...S. states. The Planner Evaluation team works on one of the key...  ...the quality of the software that drives the car. We are looking for experienced data-minded software engineers and data scientists to help...  ...analysis tools for rapid modeling and prototyping Experience... 
    Data
    Full time
    Remote work

    Waymo

    San Francisco, CA
    3 hours ago
  •  ...Description As the Manager of Model Validation & Verification (VnV)...  ...Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating...  ...will partner closely with Autonomy Software, Prediction, Planner, and ML... 
    Data
    Temporary work
    Relocation package

    Zoox

    Foster, CA
    25 days ago
  • $204k - $216k

     ...owns the quality of the models at the core of...  ..., reinforcement, and evaluation that turn open-weight...  ...work where research, data, and engineering meet: training and fine...  ...the right approach, curating the right data, running...  ...language, plus solid software engineering practice.... 
    Data

    Sapience AI Corporation

    Portland, OR
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer L5/L6 — Model Evaluations & Data Curation (MEDC). Be the first to apply!