Research
Featured story
Automated Data Quality Control With Agentic Checklists

Bad training data can leave a model worse off than no training at all. Catching it often means a person reading through the dataset manually, which stops being practical at tens of thousands of examples. It is harder still in domains like medicine, law, and business, where there is no single correct answer to check a response against.
New research introduces agentic checklist generation, a method in which an AI agent autonomously writes the quality checks itself. It builds a checklist from past runs of AutoScientist, Adaption's automated training system, then uses it to filter training data automatically.
In domains with no single correct answer, criteria have traditionally been designed and enforced by hand. Our agentic checklist generates verifiable requirements for these otherwise unverifiable domains instead.
In testing, models trained on 10,000 checklist-filtered examples had relative win rates 9.4% to 13.2% higher than models trained on 10,000 randomly sampled ones. The result held across general, medical, legal, and technology domains.
A Checklist Built by an Agent With Minimal Human Input
No single definition of "good data" works across the 44 expert domains AutoScientist supports. Each domain has its own rules and specifications for quality, so automated checks have to be domain-specific to be useful. An agentic checklist handles this by replacing the vague question "is this example good?" with specific yes-or-no questions, each tagged to the domains where it applies, so low-quality data can be found and removed before training.
Checklists for verifying datasets are not new. The novelty of our system lies in the fact that the list is written and updated by an agent with minimal human input. We use a Discovery Agent to review past AutoScientist training runs and identify critiques. From those critiques, it detects new and specific requirements for data quality across diverse domains, builds them into a checklist, and uses that checklist to filter training data automatically. A person can review the list, but does not have to.
Every AutoScientist run adds new critiques for the agent to learn from. Discovery can run continuously as experiments accumulate, and each pass produces a new version of the checklist that builds on the previous one, so coverage and specificity grow over time.
The agent tags each requirement by the domains it applies to and by difficulty, from easy to hard. Because many critiques make the same point, a compaction pass merges similar requirements when the merge adds detail, such as a specific example, and drops redundant ones.
Prompt and Completion Quality Checklist
| Stage | Quality Check | Difficulty | Domain |
|---|---|---|---|
| Prompt | The prompt does not reference missing materials (e.g., "the document above," "the table below," "the passage," "Insert HTML code here") without including them. | Easy | |
| Prompt | If the prompt provides data with known gaps or "noisy" quality, it explicitly instructs the model on how to handle these uncertainties (e.g., by stating assumptions or flagging the gaps). | Medium | data-analysis-visualization, corporate-business, science |
| Completion | When the prompt provides sufficient context to answer, the completion does not unnecessarily refuse or over-hedge (e.g., "I don't have access to...", or refusing historical facts due to "sensitivity"). | Medium | history, medical |
| Completion | The completion demonstrates a high level of technical precision and uses domain-specific terminology correctly (e.g., avoiding "Nutmeg Esophagus" when "Nutcracker Esophagus" is the correct term), and correctly interprets professional shorthand and domain-specific abbreviations without hallucinating incorrect meanings. | Hard | science, agriculture, medical, language, legal, code, technology |
Sample requirements from the checklist.
Prompt Completion Checklist Summary
| Number of Requirements | |
|---|---|
| Prompt | 16 |
| Completions | 148 |
| Difficulty | |
|---|---|
| Easy | 47 |
| Medium | 73 |
| Hard | 44 |
| Top 5 Domains | |
|---|---|
| corporate business | 48 |
| technology | 42 |
| medical | 39 |
| writing-editing-communication | 36 |
| academic-education | 28 |
The current checklist by requirement type, difficulty, and top five domains. A requirement can carry more than one domain tag.
The Checklist Removes Any Example That Misses a Requirement
The checklist runs as part of Adaptive Data, Adaption's real-time data optimization platform, which optimizes, expands, and generates AI training data. A verification agent applies the checklist in two steps. It first decides which requirements apply to an example, since a rule about medical terminology has no bearing on a business email, then checks whether the example meets each one. The agent runs in reasoning mode, which the research found necessary for it to judge relevance correctly.
Any example that misses an applicable requirement is removed. A miss is a clear error, and in supervised fine-tuning, where a model learns from examples, training on an erroneous example is worse than skipping it.
The remaining examples are scored by how many requirements they meet, the more fulfilled requirements means that the example exhibits more positive traits that the past judge critiques explicitly evaluated upon. Because the number of applicable requirements differs by domain, scores are compared within each domain, and a cutoff is tuned per domain to balance quality against how much data is kept.
Given that we need the full capability of a reasoning model to utilize the checklist effectively, efficiency is an important consideration. To increase efficiency, each of the examples evaluated upon is tagged with a primary domain and only the general requirements plus the requirements with the same tagged domain from the full checklist are considered.
Removing Low-Scoring Data Improved Win Rates
The first experiment asks what happens when the checklist removes a dataset's low-quality data and the model trains on less. It started with 50,000 examples each in medical and corporate business. We fine-tuned Gemma4-31B with LoRA twice, using a separate instance from the Discovery Agent: once on all of the data and once on what remained after the lowest-scoring 30% was removed. The training settings were held constant, so any difference comes from the data. The filtered datasets held 38,040 medical examples and 27,486 corporate business examples.
Averaged across 10 runs, the relative win rates for medical domain against the original model improved 12.1%, and for corporate business domain win rate improved 8% .

Average win rate against the original model across 10 runs, baseline versus checklist-filtered.

Relative win-rate gains of a model trained on less data after filtering through the checklist.
At a Fixed Data Size, Filtering Beat Random Selection
The second experiment asks what happens when the data size stays the same and the data is the highest quality available for that size.
Llama3.3-70B was fine-tuned on 10,000 examples for each of the general, medical, legal, and technology domains. In each domain, one set was sampled randomly from a pool of 70,000 examples, and the other came from the same pool after checklist filtering. Training settings were again fixed without optimization, so the results reflect data quality alone.
Models trained on each set were scored on win rate against the original model, using the domain-specific test sets from the AutoScientist Leaderboard, which collect the most challenging tasks on Adaption's platform for each domain. Training on checklist-filtered data outperformed random selection in every domain. The gains are relative, meaning the improvement as a share of the random-sample model's win rate: 11.4% in general, 10.1% in medical, 9.4% in legal, and 13.2% in technology.

Win rates of models trained on randomly sampled versus checklist-filtered data across 10 runs by domain.

Relative win-rate gain of models trained on checklist-filtered data over models trained on randomly sampled data.
Across both experiments, filtering with the checklist improved average win rates in every domain tested. That suggests the agent finds requirements that separate good data from bad with minimal human input.
What Becomes Possible When Quality Checks Learn
Reliable gains in AI increasingly come from curating data and improving training technique rather than adding scale. Curation depends on a standard for what good data looks like, and a standard that learns from use does not have to be rewritten as the work changes. It follows the work.
That matters most where there is no answer key. In domains like medicine, law, and business, the standard for good data can't be looked up. It has to be learned. This checklist learns from every run.
Intelligence that adapts starts with data that does too.
Date