Spot the Broken Test: Three MVMT Test Plans
Objective: Given three synthetic MVMT-style creative test plans, identify which testing discipline each one violates and why the result can't be trusted.
You've joined MVMT's growth team as a contractor for one week to review the last quarter's creative test plans before the next round launches. Three plans are queued for a repeat.
Read each test plan as written, decide whether its result is trustworthy, and flag exactly what broke it.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, no account friction, easy to sort by severity
The process
Specimens to review
Decide whether this test's conclusion ('scale Version B') is trustworthy. List every reason it isn't.
TEST PLAN: New Watch Face Ad We changed the hero image, the headline, and the CTA button color all at once. Ran both ads in the same ad set for 2 days. Version B has slightly more clicks so we're pausing A and scaling B tomorrow.
Specimen: synthetic, realistic
Decide whether this test is set up to produce a learnable result. List every reason it isn't.
TEST PLAN: Hook Test No written hypothesis, we just tried a new intro line to see what sticks. New creative is being tested directly against our 8-month-old top-performing evergreen ad. Budget and conversion target were not set before launch.
Specimen: synthetic, realistic
Decide whether this test could possibly produce a trustworthy winner. List every reason it can't.
TEST PLAN: Format Test, 'Q3 Creative Refresh' Comparing a UGC video vs. a polished studio video vs. a new CTA button text vs. a new background color, all in one experiment. Budget: $50 total, ran for 3 days, 12 conversions across all variants combined.
Specimen: synthetic, realistic
Final deliverable
A defect log across all 3 test plans, plus a rewritten one-variable, correctly-budgeted, correctly-timed test plan for the worst offender.
See a reference example
Nykaa defect log (excerpt) PLAN: Lipstick launch hook test Severity: critical — Hook, thumbnail, and price badge all changed in the same test Severity: moderate — Ran 4 days, under the 7-14 day minimum REWRITE Variable changed: hook only (first 3 seconds of a 15-second video) Budget: ₹85,000 for 2 variants, targeting 75 conversions each Duration: 10 days
Success criteria
You're done when you can:
- Correctly identifies the multi-variable violation in items 1 and 3
- Correctly identifies the missing hypothesis and unfair-comparison violations in item 2
- Does not flag any of the distractors as defects