Skip to content
Academy
Marketing Academy · Field Work●AI in Marketing
CoreAudit· 45 minutes

The Retrieval Audit: Catching a Stale Knowledge Base Before a Customer Does

Slack

Objective: Given a RAG knowledge base's document list and a log of what the system actually retrieved for 10 real queries, run the lesson's monthly retrieval audit to find where stale or duplicate documents get surfaced ahead of current ones, and decide the fix for each failure.

You're the content ops lead running Slack's monthly retrieval audit on the RAG assistant that drafts sales-enablement one-pagers and support macros from Slack's own pricing, feature, and policy documents.

Four passes: trim the document set to a curated core, run the retrieval audit against 10 real queries, diagnose the versioning failures you find, and check whether document structure is causing imprecise chunk retrieval.

Before you start

What you'll need

Free path (everything below is enough to finish)

FreemiumRun the live retrieval-audit queries against the deployed assistant interface and inspect source citations

Stands in for whichever RAG-powered assistant interface a team has built, free tier covers manual monthly audits

FreeTrack document tagging, ownership, review dates, and audit results across all 4 steps

Free, sufficient for a document and audit log

The process

4 steps

Step 01 of 04

Document quality beats document quantity (Mistake 1)

The lesson's Document Types to Prioritize list ranks brand voice, current product/pricing, approved claims, top-performing content, personas, and compliance rules, and Mistake 1 warns that dumping every file you can find causes the AI to retrieve the wrong chunk, an outdated pricing sheet over the current one.

The current knowledge base has 34 uploaded documents. Cross-referencing against the 6-item priority list, only 11 clearly map to a priority category, and 6 are exact or near-duplicate versions of the same pricing page. What do you do with the other 23?

Google Sheets— Export the document list from the RAG admin panel, tag each against the 6 priority categories.

Procedure

  1. Export all 34 document titles and upload dates
  2. Tag each against the 6 priority categories from the lesson
  3. Flag exact or near-duplicate documents (same topic, different dates)
  4. Remove all documents that don't map to a priority category, keep only the newest version of any duplicate
Sample output
Slack RAG knowledge base audit, 34 documents

Maps to priority category: 11 docs (brand voice x1, current pricing x1, approved claims x3, top content x4, compliance x2)
Duplicate pricing docs: 6 (dated Q1 2025 through Q3 2026, only Q3 2026 is current)
No clear category / stale / unrelated: 17 docs

Action: keep 11 mapped docs + newest pricing doc = 12 curated documents, remove the other 22

Healthy

Cutting to a curated 12-document set and verifying retrieval quality before adding anything back.

Unhealthy

Leaving all 34 documents in place because 'more context can't hurt,' despite 6 of them contradicting each other on price.

What this means

A RAG system doesn't average conflicting documents, it retrieves whichever chunk scores closest to the query, so 5 stale pricing docs sitting next to 1 current one is a live risk, not harmless clutter.

So what do I do about it?

SymptomActionEffort
The knowledge base has grown to 30+ documents with no removal processRun this priority-category tagging pass and cut anything that doesn't map or is a superseded duplicatehalf day
YouYou can do this yourself, no engineering access required.

Step 02 of 04

Skipping retrieval audits (Mistake 3)

Mistake 3 defines a retrieval audit: take 10-15 real queries, check exactly what chunks the system retrieves, and verify the output matches current facts, done monthly.

Running this month's 10-query retrieval audit, query 4 ('what's included in the Slack Business+ plan') retrieved a chunk from a pricing doc dated January 2025, even though the curated set now only contains the Q3 2026 doc. What does that tell you the audit just caught?

ChatGPT— Run each of the 10 audit queries directly against the RAG-powered assistant interface, log the source document and date for each retrieved chunk.

Procedure

  1. Run all 10 audit queries through the assistant
  2. For each answer, open the citation/source trace and note which document and date it pulled from
  3. Flag any answer sourced from a document dated more than 90 days before today
  4. For each flag, check whether that document is still in the active knowledge base or the retrieval index simply wasn't rebuilt after removal
Sample output
Retrieval audit, 10 queries, this month

Q1: 'Slack Enterprise Grid pricing tiers'      sourced from Q3 2026 doc   CURRENT
Q4: 'Business+ plan inclusions'               sourced from Jan 2025 doc  STALE, doc removed from KB but index not rebuilt
Q7: 'Slack Connect eligibility'               sourced from Mar 2026 doc  CURRENT
...7 more rows

1 of 10 answers (10%) sourced a document that was already removed from the active set

Healthy

Catching the stale index and immediately triggering a full re-index of the vector database, then re-running query 4 to confirm it now pulls Q3 2026.

Unhealthy

Assuming removing a document from the admin panel automatically and instantly updates every retrieval, without re-testing the specific query that used to pull from it.

What this means

Deleting a source document doesn't guarantee the vector index is rebuilt immediately; the retrieval audit is the only way to catch a stale index before a customer-facing answer does.

So what do I do about it?

SymptomActionEffort
A removed document's content still shows up in an answer weeks laterTrigger a manual re-index and re-run the specific query that surfaced the stale chunk5 min
DeveloperNeeds a developer/engineer to ship the fix.

Step 03 of 04

No document versioning or review cadence (Mistake 2)

Mistake 2 requires date-stamping every document and a mandatory quarterly review, and assigning a single knowledge base owner responsible for keeping the set current.

The stale January 2025 pricing doc that showed up in query 4 had no assigned owner and no scheduled review date. Who should own fixing this, and what's the actual process fix, not just the one-time patch?

Google Sheets— Same tracking sheet, add owner and next_review_date columns to every curated document.

Procedure

  1. Assign one named owner to the 12-document curated set
  2. Add a next_review_date to every document, no more than 90 days out
  3. Set a recurring calendar reminder tied to the earliest next_review_date
  4. Document the removal-triggers-reindex step so it isn't a manual afterthought next time
Sample output
Ownership log, Slack RAG knowledge base

Document                          Owner              Next review
Pricing (Q3 2026)                 Priya (PMM)        2026-11-15
Brand voice guide v4               Priya (PMM)        2027-02-01
Compliance: data residency          Legal (Raj)        2026-10-01

Process fix logged: any document removal now triggers an automatic re-index job, not a manual ticket

Healthy

Assigning a named owner and a hard review date to every document, including a written process for what happens on removal.

Unhealthy

Fixing today's specific stale-pricing incident by hand and moving on, with no owner or review date attached to prevent the same failure next quarter.

What this means

A one-time fix addresses this month's symptom; an owner plus a review date plus a documented removal process addresses the actual cause.

So what do I do about it?

SymptomActionEffort
The same type of stale-document incident recurs every few monthsAssign an explicit owner and review cadence to every document in the knowledge base, not just the one that just failed30 min
EitherYou or a developer can handle this, depending on your access.

Step 04 of 04

Ignoring chunk size and document structure (Mistake 5)

Mistake 5 warns that unstructured walls of text produce large, imprecise retrieved chunks, while clear headers, short paragraphs, and structured formats produce more precise retrievals.

Query 4's Business+ pricing answer came back vague and generic even after the re-index fixed the staleness, because the Q3 2026 pricing doc is a single unstructured paragraph covering all four plan tiers at once. What's the actual structural fix?

Google Sheets— Open the Q3 2026 pricing document, compare its current wall-of-text format against a headed, bulleted format.

Procedure

  1. Open the current pricing doc and check whether tiers are broken into separate sections/headers or one continuous paragraph
  2. Reformat into one clearly headed section per plan tier, each with its own bullet list of inclusions
  3. Re-run query 4 after reformatting and compare the retrieved chunk's precision
  4. Apply the same headers-and-bullets structure to the next 2 highest-traffic documents
Sample output
Before: single 400-word paragraph covering Pro, Business+, and Enterprise Grid pricing together
After: 3 separate headed sections (Business+, Pro, Enterprise Grid), each 60-80 words with a bulleted inclusions list

Query 4 retrieval, before: pulled a 400-word chunk covering all 3 tiers, answer had to guess which parts applied to Business+
Query 4 retrieval, after: pulled only the 70-word Business+ section, answer became specific and accurate

Healthy

Reformatting the pricing document into per-tier sections so retrieval can pull exactly the relevant chunk instead of the whole document.

Unhealthy

Concluding the RAG system itself is broken or 'not smart enough' when the actual issue is an unstructured source document forcing an imprecise chunk.

What this means

Retrieval precision is bounded by source document structure, no amount of re-indexing fixes a chunk that's too broad because the underlying document was never split into sections.

So what do I do about it?

SymptomActionEffort
An answer is technically sourced from the current document but still reads vague or over-broadCheck the source document's structure before assuming a model or indexing problem, split it into clearly headed sections30 min
EitherYou or a developer can handle this, depending on your access.

Final deliverable

A cleaned 12-document knowledge base list with owners and review dates, plus a completed 10-query retrieval audit log with pass/fail per query.

See a reference example
Sample output
Awfis RAG retrieval audit (excerpt)

Query: 'What's included in the Awfis Fully Serviced Office plan?'
Retrieved from: Awfis Enterprise Solutions Sheet, dated July 2026   PASS

Query: 'Awfis hot desk price in Bengaluru'
Retrieved from: pricing_final_v2.docx, no date    FAIL, flagged for owner review

Audit result: 9/10 PASS, 1/10 FAIL (undated legacy document still in index)

Success criteria

You're done when you can:

  • Correctly cuts the document set to only priority-mapped, non-duplicate documents
  • Identifies the stale-index failure in the retrieval audit rather than assuming the model is wrong
  • Assigns both an owner and a review date, not just a one-time fix
  • Diagnoses the vague-answer symptom as a document-structure problem, not a retrieval-engine problem