AI Blog

Distilabel vs Curator vs NeMo Data Designer vs Augmentoolkit

All four frameworks orchestrate LLM calls into datasets at scale, and on that axis the differences are ergonomic. The axis that decides your outcome is whether the tool can execute a verifier inside the loop — because a judge from the generator's own family filters half your rows and adds no information. Only one of the four treats programmatic validation as a first-class stage.

By Agentic AI Wiki 12 min read

Pick a synthetic data framework on how nicely it orchestrates LLM calls and you will pick the wrong one, because all four of these do that competently and the differences are ergonomic. The stage that decides whether your dataset is worth training on is verification — and the default configuration in most tutorials is an LLM judge from the same family as the generator, which filters half your rows, raises your quality metric, and adds precisely no information. Only one of these four treats an executable validator as a first-class pipeline stage, and if you are generating tool-calling or multi-step agent data, that is the entire selection criterion.

At a glance

Four actively-used open-source frameworks, four different centres of gravity. Read the last column first.

FrameworkMaintainer / licenceShapeWhere correctness comes from
Distilabel Argilla (Hugging Face) · Apache-2.0 Typed Step / Task / Pipeline graph; generators and judges as composable nodes. AI feedback. Judges and preference scoring are the built-in primitive; executable checks are yours to add.
Bespoke Curator Bespoke Labs · Apache-2.0 A single curator.LLM class with prompt() and parse(); everything else is Python. Nothing built in. parse() is where your validation goes, and it runs in your process.
NeMo Data Designer NVIDIA · Apache-2.0 Declarative config: statistical samplers plus LLM columns, with dependency-aware generation order. Python, SQL and custom validators as declared stages, plus LLM-as-judge scoring alongside them.
Augmentoolkit Community · MIT Seven opinionated end-to-end pipelines (factual, RAG, classifier bootstrap, correction, GRPO) you configure rather than compose. Baked into each pipeline. You inherit its notion of a good row and cannot easily change it.

One maintenance fact belongs in this table and does not fit: Distilabel's README now states that the original authors have moved on and a group of community members have joined as collaborators, with the develop branch carrying the latest fixes. It is not abandoned and it is widely used, but if you are choosing infrastructure for a two-year post-training programme, that sentence is a data point of the same weight as a missing feature.

The stage everyone optimises, and the stage that decides

The five stages of a synthetic data pipeline, with the verifier highlighted A left-to-right pipeline: seed or schema, generation, verification, selection, and export to training. The generation stage is where all four tools compete and is labelled commodity; the verification stage is highlighted as the stage that decides whether the dataset is worth anything, and is the stage the tools differ on. A feedback arrow runs from verification back to generation. ONE PIPELINE, FIVE STAGES 1 · Seed documents, a schema, a sampler, or nothing yours to supply 2 · Generation prompt, call, retry, cache, batch, parse all four do this well 3 · Verification does this row hold up against something outside the generator? execute · assert · judge the whole decision 4 · Selection filter, rank, dedupe, pair for preference as good as stage 3 5 · Train SFT, DPO, RL rollouts rejection sampling: regenerate what failed Stage 2 is a commodity. Stage 3 is the product. A tool that only orchestrates generation hands you stage 3 as an exercise — which is fine if you already have a checker, and fatal if your plan was to let an LLM judge decide. For tool-calling and multi-step agent data, the only verification that discriminates is one that runs the call and compares the effect.
Generation is a commodity; four frameworks do it well. Verification is where they diverge, and it is stage three that sets the ceiling on stages four and five.

Every one of these tools was built because stage two is annoying: rate limits, retries, batch APIs, structured-output parsing, resuming a 400,000-row run that died at row 380,000. Curator's whole pitch is that it handles caching, async and fault recovery at every scale; Distilabel's is that a pipeline is a typed graph you can reason about; NeMo Data Designer's is that you declare a configuration and it handles execution, batching and token metrics for you. All three claims are true, and none of them touches the question of whether the rows are right.

The reason stage three sets the ceiling is arithmetic rather than philosophy. Rejection sampling — generate n, keep the ones that pass, retry the rest — is the technique that made synthetic post-training data work, and its yield is entirely a function of how well the filter discriminates. A filter with no discriminating power does not produce a smaller, better dataset; it produces a smaller dataset with the same error rate and a more confident owner.

Where the ground truth comes from

Three sources of ground truth in a synthetic data pipeline Three columns comparing a same-family LLM judge, a cross-family judge or human sample, and an executable verifier. The first measures style agreement and adds no information, the second is a weak independent signal, and the third is the only one that survives contact with agentic data. NO INFORMATION WEAK SIGNAL DECIDES Same-family judge The generator's sibling scores the generator's output on a rubric. Measures: agreement with its own priors. Filters 50% of rows and raises the metric, not Cross-family or human A different model family, or a small human-labelled calibration sample. Measures: something, imperfectly, with known disagreement rate. Cost scales with rows. Executable verifier Run the code. Run the query. Make the tool call and compare the effect. Measures: whether the row is true. The only kind that works for agent trajectories. the quality. Use it to calibrate the verifier, not to replace it. Needs an environment, which is the real cost of the field.
The default is column one. The thing you need for agent data is column three. Column two is for calibrating column three, not for replacing it.

Take the three columns together, because the comparison is what matters. A same-family judge scoring its sibling's output is measuring agreement with its own priors — it will reliably reject rows that are unusual and accept rows that are fluent, and both of those correlate weakly with correctness and strongly with style. A cross-family judge or a small human-labelled sample is a genuinely independent signal with a measurable disagreement rate, which makes it useful for calibration and expensive per row. An executable verifier — run the generated code, execute the SQL against a real table, make the tool call and compare the resulting state — is the only one of the three that answers whether the row is true rather than whether it looks like the rows the judge has seen.

For agentic data the gap is not a matter of degree. A synthetic tool-calling trajectory is either a sequence of calls that reaches the target state or it is not, and no amount of rubric-reading tells you which, because the plausible-looking wrong call and the correct call differ by an argument. This is the same asymmetry that governs test-time compute: verifying is cheaper and more reliable than generating, provided you have something that can verify. Which means the real cost of this field is not tokens — it is the environment the verifier runs in.

Cross-cutting comparison

Verification and scale capabilities across four synthetic data frameworks A matrix with Distilabel, Bespoke Curator, NeMo Data Designer and Augmentoolkit as rows, scored across four columns: executable validators, built-in LLM judge, scale and fault recovery, and turnkey end-to-end pipelines. NeMo Data Designer is strongest on executable validators, Curator on scale, Augmentoolkit on turnkey pipelines, and Distilabel on judges. Where each framework is strongest EXECUTABLE VALIDATORS BUILT-IN JUDGE SCALE / RECOVERY TURNKEY PIPELINES Distilabel Weak — bring your own Strong Medium Medium — recipes Bespoke Curator Weak — parse() is yours Weak Strong Weak — examples only NeMo Data Designer Strong — Python, SQL Strong Medium — via Curator Medium Augmentoolkit Medium — per pipeline Medium Weak Strong Strong Medium Weak / absent Read column one first. It is the only column whose weakness cannot be fixed by writing more of your own code — because if the framework cannot run a validator, your validator has to run outside the pipeline, and then the rejection-sampling loop stops being a loop.
Column one is the only weakness you cannot code around inside the pipeline.

Executable validation

NeMo Data Designer is alone in treating this as a declared stage: validation in Python or SQL, with custom validators, sits in the configuration next to the LLM columns and next to judge scoring, so a failed row is visible to the generation order and can be regenerated. Augmentoolkit has correctness checks, but they are internal to each of its pipelines — the factual pipeline knows what a good recall pair looks like and you inherit that definition. Distilabel and Curator both let you write any check you like, in Python, as a step or in parse(); the difference is that a check you write is a check that runs where you put it, and if it needs a container, a database or a browser, it is no longer inside the framework's retry and caching loop.

Scale and recovery

Curator is the strongest here and it is the reason people reach for it: OpenAI-native and LiteLLM backends, vLLM and Ollama for local models, batch APIs from four vendors, caching and fault recovery as design goals rather than features, and a viewer for watching data arrive. Distilabel scales adequately and its typed graph makes long pipelines legible. NeMo Data Designer leans on NeMo Curator for the scaling story, which is coherent if you are already in that stack and an additional dependency if you are not. Augmentoolkit is explicitly built to run on your own hardware without external API keys, which is a deliberate trade of throughput for independence.

Time to a first useful dataset

Reverse the order. Augmentoolkit will take a directory of documents and give you a fine-tuning set today, because it has already made every decision; Curator will take an afternoon and produce exactly what you specified; Distilabel sits between them with recipes; NeMo Data Designer asks you to express your data as a schema first, which is the slowest start and the one that pays back when the dataset needs to be regenerated with one column changed.

What happens when you need RL rather than SFT

This is where the field is moving and where the four diverge most. Reinforcement learning on verifiable rewards needs an environment, not a dataset — the verifier has to be callable during training, not just during generation. Augmentoolkit ships an experimental GRPO pipeline and Curator has added trainer integrations for going from curated data to a LoRA fine-tune, but none of the four is an environment framework, and treating a synthetic-data tool as one is the most common architectural mistake in this area.

When to pick which

SituationPickBecause
Tool-calling or multi-step agent data NeMo Data Designer Declared Python/SQL validators keep the check inside the regeneration loop, which is the only way rejection sampling stays cheap.
Millions of rows against hosted APIs Bespoke Curator Caching, batch APIs and fault recovery are the product. Write your verifier in parse() and accept that it runs in-process.
Preference pairs and AI-feedback datasets Distilabel Judges and pairwise scoring are first-class, and the Argilla path to human review is short. Read the maintenance note first.
Domain documents to a fine-tune, this week Augmentoolkit Opinionated end-to-end pipelines and local-model operation. You trade control for a working dataset today.
You have no verifier and no way to build one None of them yet Spend the week on the verifier instead. Every one of these frameworks multiplies whatever discriminating power you bring, including zero.

FAQ

Is an LLM judge ever the right filter?

Yes — for subjective axes where no executable check exists, such as tone, helpfulness or instruction adherence, and when the judge is from a different family than the generator and has been calibrated against a human-labelled sample with a reported agreement rate. What is not defensible is using a same-family judge as the primary correctness filter on data where a program could have decided.

Can I combine two of these?

Commonly, and the natural split is generation in Curator for throughput and verification outside it. The cost is that your rejection-sampling loop becomes two systems with a queue between them, so measure the yield before you commit — if your pass rate is high, the simpler single-framework path is usually cheaper overall.

Does "Apache-2.0" mean I can train on the outputs?

The framework's licence governs the framework. What you may do with the generated rows is set by the terms of whichever model produced them, which is a separate document and frequently a stricter one. Check the model's terms per run, not per framework.

How much synthetic data do I actually need?

Fewer rows than the tutorials imply, and the constraint is almost always diversity rather than volume. A pipeline that generates 100,000 rows from a single seed template produces one row a hundred thousand times; the seeding strategy in stage one deserves more of your attention than the generation throughput in stage two.

Is Distilabel safe to build on given the maintainer change?

For a project measured in months, yes — it is Apache-2.0, widely deployed and has active community collaborators. For a multi-year programme, treat it the way you would any dependency with a maintainer transition: pin it, vendor the steps you rely on, and know what replacing it would cost.

Further reading

On this wiki:

Project sources: