Research
Featured Story
Invent a Dataset: Measuring Dataset Generation Abilities with Zero Seed Data

Today, Adaption is introducing new research evaluating autonomous systems that can build training datasets without any seed data. Building datasets remains one of the most manual and brittle parts of AI development. This research focuses on the most extreme yet prevalent setting real-world practitioners face: a zero-data regime.
Adaption released Invent a Dataset, which generates a high-quality AI-ready training dataset based on a user’s description and requested volume. Going from a dataset description to a dataset that captures a capability is one of the hardest agentic tasks. Typically, a core issue is diversity collapse as the volume of data requested expands. It is also hard to preserve the quality of data generations at scale.
To assess progress in the zero-data regime, we evaluate Invent API alongside APIs from Anthropic, Google, OpenAI, DeepSeek, and Zai for of a variety of dataset requests covering different tasks, domains, and languages, at sizes ranging from 200 to 20,000 samples. Some of these models do not permit training on synthetic data created with their models, so they are only included for evaluation purposes. In contrast, Invent explicitly permits commercial use for downstream training. We also include open-weight models as a feasible alternative.
Invent Produces Diverse, Quality Datasets at Any Sample Size
Invent API produces the most diverse and highest-quality datasets, with training datasets 19%-55% more diverse than every model tested. Crucially, Invent’s diversity advantage does not come at a quality cost, it is pareto-frontier, improving on both objectives at once. It also outperforms on quality, about 17% higher than the strongest baselines, Claude Opus 5 and GLM-5.3. That means, with Invent API, you can generate as much data as a project needs, without hitting a point where more data means worse data.

Diversity versus quality across five models, averaged over dataset requests from different domains.
Real-world data requests often impose explicit constraints (length, format, structure, persona), which narrow the space of valid outputs and typically reduce diversity. These constraints made the dataset request harder for most model providers, reducing diversity for every external baseline, with drops in quality of up to 22%. Under constrained requests, Invent has the highest diversity of any pipeline, about 58% above the next best (GLM-5.3).

Diversity with and without added constraints across each model.
High-Quality Datasets Lead to Higher Performing Models
Diversity and quality scores only matter if they translate into a better model. To test that, Adaption's team trained a model on 20,000 samples from each API, then had all the resulting models answer the same questions and ranked their answers head-to-head.
The model trained on Invent API's data won outright on 54% of those questions, more than double an untrained model and six times better than the dataset generated by the nearest baseline, Claude Opus 5. With Invent API, the data you generate doesn't just look better. It measurably improves the model you train on it.

Models trained on diverse, high-quality datasets see performance gains.
What This Means for Teams Building Their Own Models
Most teams start a new capability with no data. Invent API produces the highest-quality, most diverse datasets of any model tested, and the data doesn't just look better – it measurably improves the model you train on it. Unlike proprietary APIs where it is against the terms of service to use data for training, Invent is purposefully designed to enable high-quality training data for commercial purposes.
Invent a Dataset turns a single description into training data that maintains quality and diversity as it grows, so teams can build the capabilities they want without waiting for a dataset to exist.
Author
Shivalika Singh , Member of Technical Staff, Andrija Djurisic, Modelling Resident, Sara Hooker, Co-founder, Gbemileke Onilude, Modelling Resident, and Sudip Roy, Co-founderDate