Skip to content
Academy
Marketing Academy · Field Work●AI in Marketing
MiniBuild the Asset· 30 minutes

Build the Training Set: A Fine-Tuning Dataset Spec From Raw Brand Copy

Sula Vineyards

Objective: Given a pile of unsorted brand copy (some on-voice, some off-voice, some duplicates), apply the lesson's data pipeline (collect, format as input/output pairs, clean and deduplicate, shuffle and split) to produce a real fine-tuning-ready dataset spec, not just a folder of raw text.

You're the content lead at Sula Vineyards, India's listed wine and hospitality company (BSE/NSE: SULA), tasked with fine-tuning a model to write on-brand tasting notes, event copy, and social captions without a stylist rewriting every output.

Sort 12 raw samples into keep/exclude, format the kept ones as input/output pairs, remove near-duplicates, then produce a train/eval split spec ready to hand to an engineer.

Which of these 12 raw copy samples actually belong in a fine-tuning dataset, and how do they need to be reshaped before a model can learn from them?

LLM Fine-Tuning/Brand Voice/Dataset Curation/Prompt Engineering

Before you start

What you'll need

  • —Comfortable working in a spreadsheet or plain text file
  • —Has read the lesson's data pipeline section (collect, format, clean, split)
Input/output pair
one fine-tuning training example: a prompt (the input) paired with the exact on-brand response the model should learn to produce (the output).
Train/eval split
reserving 10-20% of your examples untouched during training so you can test afterward whether the model actually learned your style, not just memorized the training set.

Free path (everything below is enough to finish)

FreeTrack, format, and split the dataset

Free, no setup, easy to hand off as a spreadsheet spec

ChatGPT(optional)
FreemiumDraft realistic instruction prompts for each kept sample

Free tier generates plausible prompts fast; you still write and approve every output pair yourself

The process

2 steps

Step 01 of 02

Collecting only brand-approved, intentional examples

The lesson's data pipeline stage says to gather your best-performing emails, social posts, and landing copy, excluding experiments, A/B tests, and anything the team didn't love, because every example teaches the model what 'on brand' means, including the bad ones.

Of 12 raw samples pulled from Sula's content archive (tasting notes, event blurbs, a bulk-mailer promo, two near-identical Instagram captions, and one AI-drafted-but-never-approved product blurb), which 8 belong in the training set and which 4 get excluded, and why?

Google Sheets— A single sheet with columns: sample_id, text, source, approved (yes/no), reason_if_excluded.

Procedure

  1. List all 12 samples with their source (tasting note, email, social caption, etc.)
  2. Mark the bulk-mailer promo and the never-approved AI draft as excluded, neither reflects intentional brand voice
  3. Mark the two near-identical Instagram captions as one keep, one excluded duplicate
  4. Confirm the remaining 8 are each a distinct, team-approved example
Sample output
Sula fine-tuning source review (12 samples)

KEEP (8)
  s01  tasting note, Rasa Shiraz            approved
  s02  tasting note, Dindori Reserve         approved
  s03  event blurb, SulaFest 2026            approved
  s05  Instagram caption, harvest photo      approved
  s07  landing page snippet, wine club        approved
  s08  tasting note, Zinfandel Rose           approved
  s09  event blurb, winery tour launch        approved
  s11  Instagram caption, sunset vineyard shot approved

EXCLUDE (4)
  s04  bulk-mailer discount blast             reason: mass-mailed, generic, not intentional voice
  s06  Instagram caption, near-duplicate of s05  reason: duplicate, would bias the model toward one phrasing
  s10  AI-drafted product blurb, never shipped  reason: never approved, not real brand voice
  s12  A/B test variant B, underperformed       reason: explicitly excluded per lesson (experiments don't count)

Healthy

8 clean, distinct, team-approved samples move forward; every exclusion has a one-line reason.

Unhealthy

Including the bulk-mailer blast because it technically came from the brand's own email account, ignoring that it was never something the team considered 'good' copy.

What this means

A fine-tuning dataset is a curriculum, not an archive dump. Every included example actively teaches the model 'write like this'; every excluded example would teach it 'this is also fine,' which is exactly what corrupts a small dataset.

So what do I do about it?

SymptomActionEffort
The dataset has 12 samples but only 8 are actually on-brandExclude experiments, mass-mailed copy, and unapproved drafts before formatting anything5 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Formatting fine-tuning data as input/output pairs

The lesson specifies fine-tuning needs input/output pairs (a prompt paired with the exact on-brand response), 50-200 words per output, then a shuffle and a 10-20% eval reserve so you can test whether the model actually learned anything.

Turn the 8 kept samples into input/output pairs, then decide which samples go into the eval set instead of training.

Google Sheets— Same sheet, add two columns: input_prompt, output_text; then a split column (train/eval).

Procedure

  1. Write a realistic instruction prompt for each kept sample (e.g. 'Write a tasting note for a Shiraz release')
  2. Paste the approved copy as the output_text for that row
  3. Randomize row order (avoid grouping all tasting notes together, which would bias early training)
  4. Assign 1 of the 8 rows (12.5%, inside the lesson's 10-20% range) to eval, the other 7 to train
Sample output
Sula fine-tuning pairs (excerpt, post-shuffle)

TRAIN
  input: "Write a short tasting note for a rose release."
  output: "Pale coral in the glass, this Zinfandel Rose opens with wild strawberry and a whisper of rose petal. Dry on the palate, with a clean citrus finish. Serve chilled, best alongside something grilled."

  input: "Write an Instagram caption for a harvest photo."
  output: "Grapes don't wait for a good mood. Harvest week at the vineyard, hands full, sun down by six."

  ...5 more train rows

EVAL (held out, 1 of 8)
  input: "Write a landing page snippet for the wine club."
  output: "Six bottles, four seasons, one story each. Join the club, skip the guesswork."

Healthy

7 rows in train, 1 in eval, each row a complete input/output pair, order randomized.

Unhealthy

Skipping the eval split entirely and training on all 8 rows, leaving no way to check afterward whether the fine-tune actually generalized.

What this means

The eval row is deliberately never shown to the model during training. If the fine-tuned model can write a convincing wine-club snippet on a topic it never saw, that's evidence the fine-tune learned Sula's voice, not just memorized 7 sentences.

So what do I do about it?

SymptomActionEffort
No way to tell if the fine-tune actually worked after training finishesReserve 10-20% of pairs as an untouched eval set before training starts, never after5 min
YouYou can do this yourself, no engineering access required.

Final deliverable

An 8-row fine-tuning dataset spec (7 train, 1 eval), each row a complete input/output pair, with a one-line exclusion log for the 4 rejected samples.

See a reference example
Sample output
Go Digit General Insurance, fine-tuning dataset spec (excerpt)

TRAIN
  input: "Write a claim-approval SMS notification."
  output: "Good news. Your claim is approved and the payout is on its way, no paperwork needed on your end. Track it anytime in the app."

EVAL (held out)
  input: "Write a renewal reminder email subject line."
  output: "Your cover doesn't renew itself. Two minutes, done."

EXCLUDED (2 of 10)
  a bulk SMS blast, reason: mass-mailed, not intentional voice
  an unapproved AI draft, reason: never signed off by brand team

Success criteria

You're done when you can:

  • Correctly separates approved, intentional samples from experiments/duplicates/unapproved drafts
  • Every kept sample is reshaped into a complete input/output pair
  • An eval split is reserved and excluded from training, matching the lesson's 10-20% range