onrup

Preparing data

Preparing a fine-tuning dataset

How do I prepare a dataset for fine-tuning?

Collect real examples of the task done well, make them consistent, deduplicate before splitting, and check the token length distribution against your sequence length. Consistency matters more than volume: a thousand coherent examples beat ten thousand contradictory ones.

Start from work already done

The best training data usually exists before the project does. Resolved support tickets, approved summaries, merged review comments, analyst-written queries, translation memory, moderation decisions — all of these are records of the task done acceptably, labelled by outcome.

Writing examples specially for training is slower, more expensive, and produces a cleaner distribution than production will ever contain. Prefer harvesting to authoring wherever the option exists.

Filter by outcome, not by appearance

The single highest-leverage decision is what to exclude. Review comments that were ignored, first-assignment ticket labels that were later corrected, generated responses that a human edited — all of these look like valid training data and teach the model to reproduce mistakes.

Use an outcome signal wherever one exists: was it accepted, was it acted on, did it survive review. That signal is a better filter than any judgement made by reading the examples.

Consistency beats volume

Two examples giving different answers to substantively the same question teach the model that the task is ambiguous, and it will hedge. This is why a small, coherent dataset routinely outperforms a large, mixed one.

It also means the model learns the distribution you show it including the parts you did not intend. If half your examples end with a follow-up question, expect the fine-tuned model to end half its answers with a follow-up question.

The four checks before you spend anything

All four are cheap to run and all four are expensive to discover three hours into a training run.

Hold out before you do anything else

Split the held-out set first, from the same distribution as production rather than from the cleaned pool. A held-out set drawn from curated data overstates performance because production contains the messy cases that got filtered.

Then do not touch it. Every decision made by looking at it fits the model to it a little more, and the number it eventually produces is the one you will quote to other people.

Frequently asked questions

How many examples do I need?

For format and tone, a few hundred consistent ones make a visible difference. For task competence, low thousands. For preference tuning, one to three thousand genuine pairs. Coverage of the input distribution matters more than the raw count.

Should I include examples where the right answer is to refuse?

Yes, and most people do not. A model that has never seen a refusal will answer everything, confidently, including questions your context cannot support.

Terms used here

More on preparing data

Last verified 6 August 2026.

Start with the free tier

A magic link creates your account, your tenant and your first API key. No card until you ask for compute.