Skip to content
Academy
Marketing Academy · Field Work●AI in Marketing
CoreBuild the Asset· 55 minutes

Build the Case: A Baseline-to-Board ROI Plan for Squarespace's AI Rollout

Squarespace

Objective: Build a complete internal ROI case for an AI content-tool rollout, following the lesson's 4-step framework: a pre-rollout baseline, one named metric per use case, an informal control-group design, and a miss-rate disclosure.

You're the marketing analyst at Squarespace, three weeks from a quarterly budget review where the AI content tool line item will get questioned.

Build the case leadership trusts, in order: baseline first, one named metric, a control where you can get one, and the misses reported honestly.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreeTrack the baseline, control split, and per-use-case metrics

One tracker covers all three steps of the framework without extra tooling

FreePull real conversion numbers for the win-back use case

Free, and already the source of truth for conversion data at most teams

Paid upgrades (optional, faster/deeper)

The free GA4 + Sheets combination is sufficient for one sprint's worth of manual tracking; HubSpot's dashboards save time once this becomes a recurring quarterly report.

HubSpot Marketing Hub(optional)
FreemiumAutomate the per-use-case metric tracking once the manual version proves the framework works

Worth adopting only after the informal control test validates which metrics matter

The process

2 steps

Step 01 of 02

Baselining output before an AI rollout

Framework step 1: capture content output, cycle time, and one quality metric for four to eight weeks before AI touches the workflow. Step 3: even an informal split, half the team on AI-assisted workflows and half not, for one sprint, turns an anecdote into a comparison leadership can trust.

Squarespace's template-launch email team has 5 writers. How do you structure a baseline and an informal control before rollout even starts?

Google Sheets— Build a 6-week tracker: 2 columns per writer, before-rollout output and a control-group flag.

Procedure

  1. Log each writer's weekly output, cycle time, and one quality score for 6 weeks before any AI tool is introduced
  2. At rollout, split the team: 3 writers get the AI-assisted workflow, 2 stay on the existing process for one sprint
  3. Track the same three numbers for both groups through the sprint
  4. Flag any external factor (a headcount change, a seasonal spike) that could confound the comparison
Sample output
BASELINE (6 wks, all 5 writers)
  Avg output: 3.1 emails/writer/week | Avg cycle time: 2.8 days | Quality score: 4.1/5

SPRINT 1 (AI-assisted, n=3) vs (control, n=2)
  AI-assisted: 5.0 emails/writer/week, 1.6 days, 4.0/5
  Control: 3.2 emails/writer/week, 2.7 days, 4.2/5

Healthy

Baseline exists before rollout, and the control group is tracked on the exact same three numbers as the AI-assisted group.

Unhealthy

The comparison starts only after rollout, with no pre-AI numbers to compare against.

What this means

A baseline turns 'it feels faster' into a number leadership can actually check.

So what do I do about it?

SymptomActionEffort
No baseline was captured before the tool went liveReconstruct one from the last 6 weeks of existing reporting data if it exists, or state plainly that this rollout has no baseline and needs 6 weeks before its next review30 min
YouYou can do this yourself, no engineering access required.

Step 02 of 02

Anchoring AI ROI claims to metrics the business already trusts

Framework step 2: pick one named metric per use case rather than five. Framework step 4: state plainly which use cases did not pay off, since a report with zero downsides is less credible than one that names its misses.

The AI tool touches 3 use cases: template-launch email copy, subject-line testing, and win-back segmentation. What's the one metric for each, and which one is the miss?

Google Analytics 4— Pull conversion and engagement numbers per use case from GA4's campaign reporting.

Procedure

  1. Assign exactly one metric per use case: email copy gets output velocity, subject lines get iteration speed, win-back gets conversion lift
  2. Pull the sprint-1 number for each against its baseline
  3. Identify the use case with no measurable lift, that's the miss to report
  4. Write the 3-4 sentence board summary naming the wins and the miss plainly
Sample output
PER-USE-CASE RESULT
  Email copy velocity: 3.1 -> 5.0 emails/writer/week (win)
  Subject-line iteration speed: 4 -> 11 variants tested/month (win)
  Win-back conversion lift: 2.1% -> 2.0% (miss, no measurable lift, recommend pausing this use case)

Healthy

Each use case has exactly one metric, and at least one honest miss is named rather than omitted.

Unhealthy

All three use cases are reported as wins, or every use case shares the same generic metric.

What this means

Naming the miss before finance asks is what makes the wins credible.

So what do I do about it?

SymptomActionEffort
A use case shows flat or negative results but isn't mentioned in the summaryAdd it to the report as the named miss, with a one-line recommendation (pause, adjust, or keep watching)5 min
YouYou can do this yourself, no engineering access required.

Final deliverable

A board-ready ROI summary: a 6-week baseline, an informal control-group comparison, one named metric per use case, and an honest miss-rate line naming the use case that didn't pay off.

See a reference example
Sample output
Swiggy, Q2 AI ROI board summary (excerpt)

Baseline (6 wks): 2.8 assets/writer/week. Sprint 1 AI-assisted group: 4.9/writer/week vs control 2.9/writer/week. Subject-line iteration speed roughly doubled. Personalization use case showed no measurable conversion lift and is being paused pending a redesign.

Success criteria

You're done when you can:

  • Includes a genuine pre-rollout baseline, not just post-rollout numbers
  • Names exactly one metric per use case and reports at least one honest miss