Skip to content
Academy
Marketing Academy · Field Work●Marketing Tools
CoreBuild the Asset· 55 minutes

Building a Complete CRO Test Brief: Hypothesis, Sample Size, and a Paired Tool Stack

Utkarsh Small Finance Bank

Objective: Given Utkarsh Small Finance Bank's account-opening funnel drop-off data, build a full test brief: a hypothesis written with a metric, a sample-size calculation, and a paired experimentation-plus-behavioral tool recommendation that triangulates the quantitative result.

You're the CRO analyst at Utkarsh Small Finance Bank, the Varanasi-founded small finance bank that listed in July 2023 and closed its debut day 92% above its ₹25 IPO price. The digital account-opening funnel loses 62% of users at the KYC-document upload step, and you've been asked to write the test brief before any test is built.

Write a hypothesis with a stated metric, calculate the sample size needed, and name the paired tools (experimentation platform plus behavioral analytics tool) that will run and validate the test.

Before you start

What you'll need

Free path (everything below is enough to finish)

VWO
PaidRun the test and calculate required sample size

Free tier covers Utkarsh's account-opening funnel traffic and includes the sample-size calculator

Hotjar
FreemiumReview session recordings on the winning variant before full rollout

Free Basic plan captures a working sample of daily sessions, enough to triangulate a test result

Paid upgrades (optional, faster/deeper)

Crazy Egg(optional)
PaidScrollmap and confetti-report analysis if the team needs a second behavioral view beyond recordings

Starts at $29/month, useful once the free Hotjar session cap is regularly exceeded

The process

2 steps

Step 01 of 02

Writing a testable hypothesis with sample-size discipline

The lesson's playbook format is: 'Because we saw X, we believe changing Y for Z users will lift conversion by N%,' followed by calculating sample size up front using the platform's built-in calculator and refusing to peek early.

62% of users drop at KYC-document upload, and session recordings show repeated failed re-uploads on mobile. What's the hypothesis, and how many users does the test need before you can call a winner?

VWO— VWO's sample size calculator, fed with the current 38% completion rate and a target minimum detectable effect of 5 percentage points.

Procedure

  1. State the observed problem in one sentence citing the drop-off rate
  2. Write the hypothesis in the lesson's format: because/believe/for/by
  3. Enter baseline conversion rate and minimum detectable effect into VWO's calculator
  4. Record the required sample size per variant and the estimated days to reach it at current traffic
  5. Set a rule to not check results before the sample size is reached
Sample output
HYPOTHESIS: Because 62% of mobile users abandon at KYC-document upload after a failed re-upload attempt, we believe adding inline file-size and format guidance before the upload button will lift step completion by 8% for mobile applicants.

SAMPLE SIZE: baseline completion 38%, MDE 5pp, 95% confidence, 80% power → 3,900 users per variant (7,800 total). At current traffic of ~1,100 daily mobile applicants, that's roughly 8 days minimum runtime.

Healthy

Hypothesis names the observed behavior, the specific change, the audience, and a numeric target; sample size is calculated before the test launches.

Unhealthy

Test launches with a vague hypothesis ('improve the upload step') and no sample-size number, so the team calls a winner after 3 days on gut feel.

What this means

A hypothesis without a number is an opinion, and a test without a sample size is a coin flip dressed up as data.

So what do I do about it?

SymptomActionEffort
Team calls a winner after 3 days at 95% confidenceCalculate sample size before launch and set a minimum runtime rule the team commits to in writing30 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Triangulating a quantitative test result with behavioral evidence before shipping

The lesson's final playbook step is to 'triangulate,' pairing the quantitative result with replays and survey responses before shipping a winning variant.

The A/B test shows the inline guidance variant winning by 6%. What behavioral evidence closes the loop before you ship it to 100% of traffic?

Hotjar— Hotjar session recordings filtered to the winning variant's KYC-upload step.

Procedure

  1. Filter 20-30 session recordings to users who saw the winning variant
  2. Confirm re-upload attempts dropped compared to the control recordings
  3. Check Hotjar's post-upload survey responses for mentions of the new guidance text
  4. Only ship to 100% of traffic once both the number and the behavioral evidence agree
Sample output
BEHAVIORAL CHECK — winning variant, 25 recordings reviewed
Re-upload attempts per session: 0.4 avg (down from 1.7 in control recordings)
Survey mentions of the new file-format guidance: 9 of 14 respondents referenced it positively
DECISION: Ship to 100% of traffic. Quantitative lift and behavioral evidence agree.

Healthy

Winning variant is shipped only after recordings and survey data confirm the same story as the test's numeric result.

Unhealthy

Winning variant ships purely on the 6% lift number with no recordings reviewed, and it turns out the lift was driven by a tracking bug on one device type.

What this means

A number that agrees with what you see in a recording is trustworthy; a number that stands alone is not.

So what do I do about it?

SymptomActionEffort
A winning variant ships and the lift disappears a month laterReview 20-30 recordings of the winning variant before rolling out to 100% of traffic30 min
YouYou can do this yourself, no engineering access required.

Final deliverable

A complete test brief document: hypothesis, sample-size calculation with runtime estimate, and the paired experimentation-plus-behavioral tool stack.

See a reference example
Sample output
TEST BRIEF — RateGain Travel Technologies, hotel-partner signup form

HYPOTHESIS: Because 44% of partner-hotel signups abandon at the pricing-tier selection step, we believe replacing the three-tier comparison table with a single recommended-tier default will lift signup completion by 6% for first-time visitors.

SAMPLE SIZE: baseline completion 51%, MDE 4pp, 95% confidence, 80% power → 2,600 users per variant. At current traffic, estimated runtime is 11 days.

BEHAVIORAL PAIRING: VWO for the test, Hotjar recordings on the winning variant filtered to first-time visitors before full rollout.

Success criteria

You're done when you can:

  • Hypothesis follows the because/believe/for/by format with a stated numeric target
  • Sample size is calculated before the test launches, with a runtime estimate
  • A behavioral tool and check step are named before the winning variant ships to 100%