Skip to content
Academy
Marketing Academy · Field Work●Marketing Tools
CoreTeardown· 45 minutes

Spot the Broken Test: Auditing an A/B Test Report Before It Ships

Jyoti CNC Automation

Objective: Given a synthetic A/B test report for Jyoti CNC Automation's dealer-inquiry form, apply the lesson's 'Common Mistakes' checklist to find every methodology defect before the 'winning' variation gets rolled out to 100% of traffic.

You're the CRO analyst reviewing a test report a junior teammate wrote up for Jyoti CNC Automation's dealer-inquiry landing page, tested in VWO. The team wants to ship the 'winner' this week.

Read the report closely against the lesson's 5 most common test-killing mistakes and flag every one present, including any that look fine on the surface.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeTrack which of the 5 common mistakes appear in the report

Free checklist surface, no VWO/Optimizely account needed to do the review

The process

Specimens to review

Read the test report below for the Jyoti CNC dealer-inquiry form test. Find every methodology defect that would make the declared 'winner' unreliable, using the lesson's Common Mistakes checklist.

Sample output
=== TEST REPORT: Dealer Inquiry Form, Variation B ===
Test duration: Monday to Thursday (3 days).
Hypothesis: 'Let's try changing the submit button to green.'
Original goal: form submissions.
Note: form submissions were flat through day 2, so on day 3 we switched to tracking scroll depth instead, and Variation B is winning on that metric.
Results (aggregate, all traffic): Variation B is up 15% over Control. We did not check desktop vs mobile separately.
Statistical significance: not calculated, but the lift looks clearly real.
Traffic split: 50/50 between Control and Variation B, applied correctly from launch.
Recommendation: Ship Variation B to 100% of traffic this week.

Specimen: synthetic, realistic

Final deliverable

An annotated copy of the test report with all defects flagged, mapped to the specific 'Common Mistakes' category each one violates, and a recommendation on whether to ship.

See a reference example
Sample output
Concord Biotech, RFQ Form Test Report Review (excerpt)

Flagged: goal metric changed from 'form submits' to 'time on page' on day 4 (Changing the goal metric mid-test). Flagged: no significance % reported (Statistical Significance). Not flagged: 50/50 traffic split was correctly implemented.

Recommendation: Do not ship. Re-run with the original goal metric locked for a full 2-week cycle.

Success criteria

You're done when you can:

  • Identifies all 3 critical defects (short duration, goal switch, missing significance)
  • Identifies the moderate segmentation defect
  • Identifies the cosmetic hypothesis-quality defect
  • Does not flag any of the 3 distractors as defects
  • Final recommendation is 'do not ship' with a stated reason