← All writing

FIELD NOTES

Data quality beats quantity

Your fine-tune came out mediocre, so you went looking for more data.

That’s usually the wrong fix.

In fine-tuning, 200 clean, consistent examples beat 10,000 messy ones. Almost every time.

Hand-drawn before-and-after infographic: 10,000 messy, contradictory rows versus 200 clean, consistent rows with every bad row deleted

The model copies your data — including its mistakes. Contradictory labels, sloppy formatting, off-tone answers: it learns all of it, faithfully.

What actually moves quality:

  • Consistency — same format, same voice, same answer style throughout.
  • Coverage — examples for the hard and weird cases, not only the easy ones.
  • Cleaning — ruthlessly delete the bad rows. A bad example is worse than a missing one.

Stop scaling the dataset. Start curating it.

The model is only ever as good as the examples you were willing to stand behind.