Diagnose the Fine-Tune: Overfitting, Underfitting, or Working as Intended?
Objective: Given three fine-tuned model outputs on unseen prompts, apply the lesson's red-flag framework (overfitting, underfitting, hallucination, niche terminology collapse) to correctly diagnose which failure mode each output shows, and recommend the matching fix.
You're evaluating a first-pass fine-tune for Bansal Wire Industries, India's largest stainless steel wire manufacturer (listed on BSE/NSE), which trained a model on 80 B2B product-page and export-catalog samples to write new distributor-facing copy at scale.
Three outputs, three unseen prompts. Diagnose each one against the lesson's four red flags, then recommend the specific fix, not a generic 'retrain it.'
Given three fine-tuned outputs that all look 'a bit off,' which specific red flag explains each one, and does the fix require more data, fewer epochs, or a different mitigation entirely?
Before you start
What you'll need
- —Has read the lesson's Red Flags to Watch and Evaluation sections
- —Understands the difference between prompting and fine-tuning (lesson's opening section)
- Overfitting
- the model repeats phrases from its training examples nearly verbatim instead of learning the underlying style, so it sounds like the training set rather than the brand on new topics.
- Niche terminology collapse
- the model drops rare, brand-specific words or phrases because too few training examples used them for the model to learn they matter.
Free path (everything below is enough to finish)
Free, structures the eval sheet the lesson's Evaluation section describes
Free tier is enough to sanity-check a single factual claim before flagging it as a hallucination
The process
Specimens to review
Diagnose the specific failure mode in this output and name the fix.
Prompt: "Write a product blurb for our new galvanized binding wire." Output: "Built for those who build. Precision-engineered. Trusted by professionals. Built for those who build, our binding wire delivers consistent tensile strength batch after batch. Built for those who build, that's our promise."
Specimen: synthetic, realistic
Diagnose the specific failure mode in this output and name the fix.
Prompt: "Write a one-line export catalog header for our stainless steel wire rope range." Output: "High-quality stainless steel wire rope products available for various industrial applications."
Specimen: synthetic, realistic
Diagnose the specific failure mode in this output and name the fix.
Prompt: "Write a compliance line for our export catalog's stainless steel wire section." Output: "All wire products are certified to ISO 20981 international export standard, ensuring guaranteed compliance in every market we serve."
Specimen: synthetic, realistic
Final deliverable
A 3-row diagnosis sheet mapping each output to its specific red flag (overfitting, underfitting, or hallucination) and a concrete, mode-specific fix for each.
See a reference example
Sula Vineyards, fine-tune diagnosis sheet (excerpt) Output: "Every bottle carries the soul of the harvest, every bottle carries the soul of the harvest." Diagnosis: overfitting, exact phrase repeats within one output Fix: reduce training epochs, add more diverse tasting-note examples Output: "This wine is made from grapes and aged before release." Diagnosis: underfitting, no distinctive voice, reads like generic base-model output Fix: collect more training examples, increase epochs
Success criteria
You're done when you can:
- Correctly diagnoses all 3 outputs against the lesson's four red-flag categories
- Distinguishes hallucination (invented fact) from overfitting (repeated phrasing) rather than treating both as 'sounds off'
- Recommends a fix that matches the diagnosed failure mode, not a generic 'retrain it'