Find the Missing Guardrails in a Live Agentic Workflow
Objective: Audit a real-shaped CRM-triggered email workflow spec for the three failure modes the lesson names, then rewrite it with the missing guardrails added.
ThredUp's lifecycle team built an agentic workflow that watches for CRM events (signup, trial expiration, purchase anniversary) and auto-sends personalized email sequences. It's been running two weeks. You're asked to audit it before it scales to the full customer base.
Check the spec against the lesson's three failure modes, hallucinated brand voice, missing spend or reach guardrails, and approval gaps, then rewrite the spec with fixes.
Before you start
What you'll need
Free path (everything below is enough to finish)
Free, shareable, and enough structure for a repeatable audit template
Paid upgrades (optional, faster/deeper)
Visual scenario builder with per-module error handling, useful once a workflow needs enforced caps and logging
The process
2 steps
Step 01 of 02
The lesson's first failure mode: an agent writing at scale doesn't know your brand voice is 'conversational but never slangy' unless that's explicitly in its system prompt.
The current spec's agent prompt reads: 'Write a friendly email for this customer event.' No brand voice guidance is attached anywhere in the workflow. What's the risk, and what's missing?
Procedure
- Read the agent's system prompt exactly as written in the workflow spec
- Check whether brand voice guidelines are attached to the prompt or only exist in a separate brand doc nobody linked
- Flag any step where output tone could drift toward generic without a specific voice reference
- Note the fix: attach the actual brand voice doc excerpt directly into the system prompt, not just a reference to where it lives
AUDIT FINDING 1 Step: Email draft generation Gap: System prompt has no brand voice reference; output drifted toward generic marketing copy in 6 of 20 sampled sends Fix: Embed 3 brand-voice example sentences directly in the system prompt
Healthy
The system prompt contains concrete brand voice examples, not just a link or a one-word descriptor like 'friendly'.
Unhealthy
The prompt says 'write in our brand voice' with no examples anywhere in the workflow for the agent to reference.
What this means
A brand voice guideline that exists only in a separate doc might as well not exist to the agent; it can only follow what's in its own prompt.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Sampled agent output reads generic instead of on-brand | Paste concrete brand voice examples directly into the system prompt | 30 min |
Step 02 of 02
The lesson's second and third failure modes: agents connected to sending tools can publish before a human notices, and 'someone always checked it' checkpoints get quietly designed away for speed.
The spec has no maximum daily send volume and no pause-and-review step for audiences above a set size. The workflow currently has permission to send to ThredUp's full CRM list with no cap. What do you add?
Procedure
- Check the spec for a maximum daily send volume; note that none currently exists
- Check for a mandatory pause-and-review step above a defined audience size; note that none currently exists
- Add both as explicit spec lines: a hard daily send cap, and a required human sign-off before any send to an audience over 10,000
- Add a logging requirement: every agent send action gets logged with timestamp, audience size, and template used
AUDIT FINDING 2 Step: Trigger-to-send pipeline Gap: No daily send cap; no audience-size checkpoint Fix: Add 5,000/day hard cap; require marketer sign-off before any send over 10,000 recipients; log every send action
Healthy
The rewritten spec has a hard send cap and a named approval step before any large-audience send.
Unhealthy
The workflow can send to the entire CRM list overnight with no human ever reviewing volume or targeting.
What this means
A 45% executive-reported barrier to agent adoption is lack of visibility into agent decisions; a spend or reach cap plus a log is the minimum fix for that visibility gap.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| Agent workflow can send to an unbounded audience with no review | Add a hard daily cap and a mandatory approval step above a defined audience size | 30 min |
Final deliverable
A completed guardrail audit of ThredUp's CRM-triggered email workflow, plus a rewritten spec with a brand-voice-anchored prompt, a hard send cap, an audience-size approval checkpoint, and action logging.
See a reference example
Warby Parker, CRM-Trigger Workflow Audit (excerpt) FINDING 1: No brand voice examples in agent prompt — fixed by embedding 3 reference sentences FINDING 2: No send cap on trial-expiration sequence — fixed with 3,000/day cap and sign-off above 8,000 recipients FINDING 3: No action log — fixed by requiring timestamp + audience size + template on every send
Success criteria
You're done when you can:
- Identifies all three failure modes present in the flawed spec: brand voice drift, missing spend/reach cap, and missing approval gap
- Rewritten spec includes a concrete send cap number and an audience-size threshold for mandatory review