Build the Training Set: A Fine-Tuning Dataset Spec From Raw Brand Copy
Objective: Given a pile of unsorted brand copy (some on-voice, some off-voice, some duplicates), apply the lesson's data pipeline (collect, format as input/output pairs, clean and deduplicate, shuffle and split) to produce a real fine-tuning-ready dataset spec, not just a folder of raw text.
You're the content lead at Sula Vineyards, India's listed wine and hospitality company (BSE/NSE: SULA), tasked with fine-tuning a model to write on-brand tasting notes, event copy, and social captions without a stylist rewriting every output.
Sort 12 raw samples into keep/exclude, format the kept ones as input/output pairs, remove near-duplicates, then produce a train/eval split spec ready to hand to an engineer.
Which of these 12 raw copy samples actually belong in a fine-tuning dataset, and how do they need to be reshaped before a model can learn from them?
Before you start
What you'll need
- —Comfortable working in a spreadsheet or plain text file
- —Has read the lesson's data pipeline section (collect, format, clean, split)
- Input/output pair
- one fine-tuning training example: a prompt (the input) paired with the exact on-brand response the model should learn to produce (the output).
- Train/eval split
- reserving 10-20% of your examples untouched during training so you can test afterward whether the model actually learned your style, not just memorized the training set.
Free path (everything below is enough to finish)
Free, no setup, easy to hand off as a spreadsheet spec
Free tier generates plausible prompts fast; you still write and approve every output pair yourself
The process
2 steps
Step 01 of 02
The lesson's data pipeline stage says to gather your best-performing emails, social posts, and landing copy, excluding experiments, A/B tests, and anything the team didn't love, because every example teaches the model what 'on brand' means, including the bad ones.
Of 12 raw samples pulled from Sula's content archive (tasting notes, event blurbs, a bulk-mailer promo, two near-identical Instagram captions, and one AI-drafted-but-never-approved product blurb), which 8 belong in the training set and which 4 get excluded, and why?
Procedure
- List all 12 samples with their source (tasting note, email, social caption, etc.)
- Mark the bulk-mailer promo and the never-approved AI draft as excluded, neither reflects intentional brand voice
- Mark the two near-identical Instagram captions as one keep, one excluded duplicate
- Confirm the remaining 8 are each a distinct, team-approved example
Sula fine-tuning source review (12 samples) KEEP (8) s01 tasting note, Rasa Shiraz approved s02 tasting note, Dindori Reserve approved s03 event blurb, SulaFest 2026 approved s05 Instagram caption, harvest photo approved s07 landing page snippet, wine club approved s08 tasting note, Zinfandel Rose approved s09 event blurb, winery tour launch approved s11 Instagram caption, sunset vineyard shot approved EXCLUDE (4) s04 bulk-mailer discount blast reason: mass-mailed, generic, not intentional voice s06 Instagram caption, near-duplicate of s05 reason: duplicate, would bias the model toward one phrasing s10 AI-drafted product blurb, never shipped reason: never approved, not real brand voice s12 A/B test variant B, underperformed reason: explicitly excluded per lesson (experiments don't count)
Healthy
8 clean, distinct, team-approved samples move forward; every exclusion has a one-line reason.
Unhealthy
Including the bulk-mailer blast because it technically came from the brand's own email account, ignoring that it was never something the team considered 'good' copy.
What this means
A fine-tuning dataset is a curriculum, not an archive dump. Every included example actively teaches the model 'write like this'; every excluded example would teach it 'this is also fine,' which is exactly what corrupts a small dataset.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| The dataset has 12 samples but only 8 are actually on-brand | Exclude experiments, mass-mailed copy, and unapproved drafts before formatting anything | 5 min |
Step 02 of 02
The lesson specifies fine-tuning needs input/output pairs (a prompt paired with the exact on-brand response), 50-200 words per output, then a shuffle and a 10-20% eval reserve so you can test whether the model actually learned anything.
Turn the 8 kept samples into input/output pairs, then decide which samples go into the eval set instead of training.
Procedure
- Write a realistic instruction prompt for each kept sample (e.g. 'Write a tasting note for a Shiraz release')
- Paste the approved copy as the output_text for that row
- Randomize row order (avoid grouping all tasting notes together, which would bias early training)
- Assign 1 of the 8 rows (12.5%, inside the lesson's 10-20% range) to eval, the other 7 to train
Sula fine-tuning pairs (excerpt, post-shuffle) TRAIN input: "Write a short tasting note for a rose release." output: "Pale coral in the glass, this Zinfandel Rose opens with wild strawberry and a whisper of rose petal. Dry on the palate, with a clean citrus finish. Serve chilled, best alongside something grilled." input: "Write an Instagram caption for a harvest photo." output: "Grapes don't wait for a good mood. Harvest week at the vineyard, hands full, sun down by six." ...5 more train rows EVAL (held out, 1 of 8) input: "Write a landing page snippet for the wine club." output: "Six bottles, four seasons, one story each. Join the club, skip the guesswork."
Healthy
7 rows in train, 1 in eval, each row a complete input/output pair, order randomized.
Unhealthy
Skipping the eval split entirely and training on all 8 rows, leaving no way to check afterward whether the fine-tune actually generalized.
What this means
The eval row is deliberately never shown to the model during training. If the fine-tuned model can write a convincing wine-club snippet on a topic it never saw, that's evidence the fine-tune learned Sula's voice, not just memorized 7 sentences.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| No way to tell if the fine-tune actually worked after training finishes | Reserve 10-20% of pairs as an untouched eval set before training starts, never after | 5 min |
Final deliverable
An 8-row fine-tuning dataset spec (7 train, 1 eval), each row a complete input/output pair, with a one-line exclusion log for the 4 rejected samples.
See a reference example
Go Digit General Insurance, fine-tuning dataset spec (excerpt) TRAIN input: "Write a claim-approval SMS notification." output: "Good news. Your claim is approved and the payout is on its way, no paperwork needed on your end. Track it anytime in the app." EVAL (held out) input: "Write a renewal reminder email subject line." output: "Your cover doesn't renew itself. Two minutes, done." EXCLUDED (2 of 10) a bulk SMS blast, reason: mass-mailed, not intentional voice an unapproved AI draft, reason: never signed off by brand team
Success criteria
You're done when you can:
- Correctly separates approved, intentional samples from experiments/duplicates/unapproved drafts
- Every kept sample is reshaped into a complete input/output pair
- An eval split is reserved and excluded from training, matching the lesson's 10-20% range