Diagnose a Reporting Pipeline That Went Wrong
Objective: Given a written account of a marketing team's newly automated reporting pipeline that produced a bad client-facing report, apply the lesson's four common-mistakes framework to identify which specific mistake caused the failure and what the fix is.
You're auditing Stitch Fix's styling-team marketing pipeline after a client-facing partner report went out with a claim that overstated a campaign's impact, and the partner team wants to know exactly what broke before they trust the automation again.
Read the incident account, match the failure to one of the lesson's four common mistakes, and recommend the specific fix, not a generic 'add more review' answer.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free version history is enough to reconstruct the timeline
The process
1 step
Step 01 of 01
The lesson names four common mistakes: skipping the narrative step, sending unreviewed AI narratives to clients, missing a last-period comparison, and never validating one report by hand before turning on the schedule.
Incident: the team's new pipeline compared this week's referral numbers to a corrupted archive tab from three weeks prior (this week's tab had overwritten last week's before anyone noticed), producing a narrative claiming a 340% referral spike. The AI-written narrative went straight into the partner email with no human read-through. Which one or two mistakes actually caused this, and which is the root cause versus the compounding one?
Procedure
- Check whether last week's data was archived before being overwritten, per the lesson's Mistake 3
- Check whether the AI narrative was reviewed by a human before reaching the client, per Mistake 2
- Identify which failure happened first in the pipeline's timeline (root cause) versus which one let the first failure reach the client (compounding cause)
- Write the diagnosis as: root cause, then compounding cause, then the specific fix for each
DIAGNOSIS Root cause: Mistake 3, no archived 'last period' comparison. The archive tab was overwritten instead of preserved, so the pipeline compared against three-week-old data and manufactured a false 340% spike. Compounding cause: Mistake 2, unreviewed AI narrative reaching a client. Internal Slack reports can go straight from AI to team, but this was a client-facing partner email, and no human read the narrative before it sent. Fix 1 (root): Add an explicit archive-before-overwrite step, copy raw_this_week to raw_last_week before the next pull runs, never overwrite in place. Fix 2 (compounding): Route any client-facing narrative through a required human approval step before the delivery stage fires, per the lesson's client-report guidance.
Healthy
The diagnosis correctly separates root cause (bad comparison data) from compounding cause (no human gate before a client saw it) and gives a distinct fix for each.
Unhealthy
Blaming 'the AI' generally, or recommending 'add more testing' without naming which specific stage needs the archive-before-overwrite fix.
What this means
A pipeline failure usually has two layers: a data problem that created the bad output, and a process gap that let it reach someone who mattered before it was caught.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A bad number reached a client despite the pipeline running correctly on the surface | Check the archive tab's timestamp history first, then check whether a human gate existed before client delivery | 30 min |
Final deliverable
A written incident diagnosis naming the root cause, the compounding cause, and a distinct fix for each.
See a reference example
DIAGNOSIS, Peloton Partner Reporting Incident Root cause: Mistake 3, the archive tab was overwritten before the AI narrative step ran, so the comparison was against six-week-old baseline data. Compounding cause: Mistake 2, the client-facing summary skipped human review because it was treated as an internal report by mistake. Fix 1: Add a copy-before-overwrite step to stage 2 of the pipeline. Fix 2: Flag any report routed to an external distribution list as requiring the human approval gate, regardless of how it's labeled internally.
Success criteria
You're done when you can:
- Correctly identifies the missing-archive mistake as the root cause, not a symptom
- Separately identifies the missing human-review gate as the reason the bad number reached the client
- Proposes a distinct, specific fix for each cause rather than one generic recommendation