The Sameness Check: Auditing 24 AI Variants for Real Angle Diversity
Objective: Given a 24-variant AI creative performance export tagged by angle, determine whether the test set actually covered distinct psychological angles or just cosmetic variations of one angle, and recommend which angle to scale.
You're the paid social analyst at DoorDash reviewing a Meta Advantage+ Creative test that generated 24 image variants for a new promo, and leadership wants to know which angle to scale before the next order cycle.
Group variants by underlying angle rather than by surface cosmetics, calculate performance per angle group, and flag whether the test had enough real diversity to trust the result.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, no account friction, sufficient for a 24-row export
The process
2 steps
Step 01 of 02
Testing 20 AI variants that all share the same core angle, hook, and visual structure tells you which shade of blue converts slightly better, not which underlying idea resonates, spending a testing budget while learning almost nothing transferable.
This export has 24 variants tagged with an 'angle' column, but the tags only cover 3 distinct underlying angles. Is this a valid angle test, or a sameness-fatigue trap?
Procedure
- Import the creative-performance export and pivot on the angle column
- Count how many of the 24 variants fall into each distinct angle
- Calculate average CPA per angle group, not per individual asset
- Flag any angle with fewer than 3 executions as too thin to judge confidently
DoorDash Promo Test, 24 variants by tagged angle Angle Variants Avg CPA Product Hero 18 $4.20 Social Proof 4 $3.95 Pain Point 2 $3.40
Healthy
8-12 genuinely distinct angles, each represented by a handful of executions (3-5).
Unhealthy
18 of 24 variants tagged 'Product Hero,' only 3 total underlying angles present despite 24 total assets generated.
What this means
This is sameness fatigue: 'Product Hero' looks like the strongest performer partly because it had 18 shots on goal versus 2-4 for everything else, not because the angle itself is proven best.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| One angle dominates the variant count in a creative test export | Rebrief the next test with 8-12 explicitly distinct angles and cap executions at 3-5 per angle before generating | 30 min |
Step 02 of 02
When a winning angle emerges, the lesson says to scale the underlying idea across new executions, not just the single winning asset, that is where the compounding insight lives.
Pain Point had only 2 variants but the lowest CPA per variant. Product Hero had 18 variants and the lowest CPA of any single asset. Which do you recommend leadership scale?
Procedure
- Compare average CPA per angle group, not the single best individual asset
- Flag Pain Point as promising but under-sampled (only 2 executions)
- Recommend a follow-up brief of 3-5 more Pain Point executions before crowning a winner
- Recommend scaling Product Hero cautiously, since its lead may partly reflect volume, not angle strength
Recommendation memo (excerpt) Do not scale the single best 'Product Hero' asset as the definitive winner, its angle had 9x the sample size of Pain Point. Rebrief Pain Point with 3-5 more executions next cycle before declaring a winning angle. Continue running Product Hero and Social Proof at current volume in the meantime, neither is disqualified, but neither is confirmed either.
Healthy
A scale recommendation names the angle with a note on sample-size confidence for each group.
Unhealthy
Declaring 'Product Hero' the winner and scaling the single top asset without accounting for its 9x sample advantage.
What this means
Raw asset-count wins are not the same as angle-level wins, the lesson's whole point about testing ideas, not pixels.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A recommendation names a single winning asset instead of a winning angle | Re-run the comparison at the angle level and flag any angle below 3 executions as needing more data first | 5 min |
Final deliverable
An angle-grouped performance summary with a scale recommendation and a note on which angle needs more data before a confident call.
See a reference example
ThredUp, Meta Advantage+ Creative test, angle summary (excerpt) Angle Variants Avg CPA Confidence Nostalgia/Resale 15 $2.80 Low, oversampled vs other angles UGC Unboxing 5 $2.55 Medium Sustainability 4 $2.30 Medium, promising, needs 2-3 more reps RECOMMENDATION: Do not scale Nostalgia/Resale as the confirmed winner on raw CPA alone. Rebrief Sustainability with 3 more executions before the next scaling decision.
Success criteria
You're done when you can:
- Groups variants by underlying angle rather than counting raw asset-level wins
- Flags angles with too few executions to trust before recommending a scale decision