Build the Case: A Baseline-to-Board ROI Plan for Squarespace's AI Rollout
Objective: Build a complete internal ROI case for an AI content-tool rollout, following the lesson's 4-step framework: a pre-rollout baseline, one named metric per use case, an informal control-group design, and a miss-rate disclosure.
You're the marketing analyst at Squarespace, three weeks from a quarterly budget review where the AI content tool line item will get questioned.
Build the case leadership trusts, in order: baseline first, one named metric, a control where you can get one, and the misses reported honestly.
Before you start
What you'll need
Free path (everything below is enough to finish)
One tracker covers all three steps of the framework without extra tooling
Free, and already the source of truth for conversion data at most teams
Paid upgrades (optional, faster/deeper)
The free GA4 + Sheets combination is sufficient for one sprint's worth of manual tracking; HubSpot's dashboards save time once this becomes a recurring quarterly report.
Worth adopting only after the informal control test validates which metrics matter
The process
2 steps
Step 01 of 02
Framework step 1: capture content output, cycle time, and one quality metric for four to eight weeks before AI touches the workflow. Step 3: even an informal split, half the team on AI-assisted workflows and half not, for one sprint, turns an anecdote into a comparison leadership can trust.
Squarespace's template-launch email team has 5 writers. How do you structure a baseline and an informal control before rollout even starts?
Procedure
- Log each writer's weekly output, cycle time, and one quality score for 6 weeks before any AI tool is introduced
- At rollout, split the team: 3 writers get the AI-assisted workflow, 2 stay on the existing process for one sprint
- Track the same three numbers for both groups through the sprint
- Flag any external factor (a headcount change, a seasonal spike) that could confound the comparison
BASELINE (6 wks, all 5 writers) Avg output: 3.1 emails/writer/week | Avg cycle time: 2.8 days | Quality score: 4.1/5 SPRINT 1 (AI-assisted, n=3) vs (control, n=2) AI-assisted: 5.0 emails/writer/week, 1.6 days, 4.0/5 Control: 3.2 emails/writer/week, 2.7 days, 4.2/5
Healthy
Baseline exists before rollout, and the control group is tracked on the exact same three numbers as the AI-assisted group.
Unhealthy
The comparison starts only after rollout, with no pre-AI numbers to compare against.
What this means
A baseline turns 'it feels faster' into a number leadership can actually check.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| No baseline was captured before the tool went live | Reconstruct one from the last 6 weeks of existing reporting data if it exists, or state plainly that this rollout has no baseline and needs 6 weeks before its next review | 30 min |
Step 02 of 02
Framework step 2: pick one named metric per use case rather than five. Framework step 4: state plainly which use cases did not pay off, since a report with zero downsides is less credible than one that names its misses.
The AI tool touches 3 use cases: template-launch email copy, subject-line testing, and win-back segmentation. What's the one metric for each, and which one is the miss?
Procedure
- Assign exactly one metric per use case: email copy gets output velocity, subject lines get iteration speed, win-back gets conversion lift
- Pull the sprint-1 number for each against its baseline
- Identify the use case with no measurable lift, that's the miss to report
- Write the 3-4 sentence board summary naming the wins and the miss plainly
PER-USE-CASE RESULT Email copy velocity: 3.1 -> 5.0 emails/writer/week (win) Subject-line iteration speed: 4 -> 11 variants tested/month (win) Win-back conversion lift: 2.1% -> 2.0% (miss, no measurable lift, recommend pausing this use case)
Healthy
Each use case has exactly one metric, and at least one honest miss is named rather than omitted.
Unhealthy
All three use cases are reported as wins, or every use case shares the same generic metric.
What this means
Naming the miss before finance asks is what makes the wins credible.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A use case shows flat or negative results but isn't mentioned in the summary | Add it to the report as the named miss, with a one-line recommendation (pause, adjust, or keep watching) | 5 min |
Final deliverable
A board-ready ROI summary: a 6-week baseline, an informal control-group comparison, one named metric per use case, and an honest miss-rate line naming the use case that didn't pay off.
See a reference example
Swiggy, Q2 AI ROI board summary (excerpt) Baseline (6 wks): 2.8 assets/writer/week. Sprint 1 AI-assisted group: 4.9/writer/week vs control 2.9/writer/week. Subject-line iteration speed roughly doubled. Personalization use case showed no measurable conversion lift and is being paused pending a redesign.
Success criteria
You're done when you can:
- Includes a genuine pre-rollout baseline, not just post-rollout numbers
- Names exactly one metric per use case and reports at least one honest miss