The Retrieval Audit: Catching a Stale Knowledge Base Before a Customer Does
Objective: Given a RAG knowledge base's document list and a log of what the system actually retrieved for 10 real queries, run the lesson's monthly retrieval audit to find where stale or duplicate documents get surfaced ahead of current ones, and decide the fix for each failure.
You're the content ops lead running Slack's monthly retrieval audit on the RAG assistant that drafts sales-enablement one-pagers and support macros from Slack's own pricing, feature, and policy documents.
Four passes: trim the document set to a curated core, run the retrieval audit against 10 real queries, diagnose the versioning failures you find, and check whether document structure is causing imprecise chunk retrieval.
Before you start
What you'll need
Free path (everything below is enough to finish)
Stands in for whichever RAG-powered assistant interface a team has built, free tier covers manual monthly audits
Free, sufficient for a document and audit log
The process
4 steps
Step 01 of 04
The lesson's Document Types to Prioritize list ranks brand voice, current product/pricing, approved claims, top-performing content, personas, and compliance rules, and Mistake 1 warns that dumping every file you can find causes the AI to retrieve the wrong chunk, an outdated pricing sheet over the current one.
The current knowledge base has 34 uploaded documents. Cross-referencing against the 6-item priority list, only 11 clearly map to a priority category, and 6 are exact or near-duplicate versions of the same pricing page. What do you do with the other 23?
Procedure
- Export all 34 document titles and upload dates
- Tag each against the 6 priority categories from the lesson
- Flag exact or near-duplicate documents (same topic, different dates)
- Remove all documents that don't map to a priority category, keep only the newest version of any duplicate
Slack RAG knowledge base audit, 34 documents Maps to priority category: 11 docs (brand voice x1, current pricing x1, approved claims x3, top content x4, compliance x2) Duplicate pricing docs: 6 (dated Q1 2025 through Q3 2026, only Q3 2026 is current) No clear category / stale / unrelated: 17 docs Action: keep 11 mapped docs + newest pricing doc = 12 curated documents, remove the other 22
Healthy
Cutting to a curated 12-document set and verifying retrieval quality before adding anything back.
Unhealthy
Leaving all 34 documents in place because 'more context can't hurt,' despite 6 of them contradicting each other on price.
What this means
A RAG system doesn't average conflicting documents, it retrieves whichever chunk scores closest to the query, so 5 stale pricing docs sitting next to 1 current one is a live risk, not harmless clutter.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| The knowledge base has grown to 30+ documents with no removal process | Run this priority-category tagging pass and cut anything that doesn't map or is a superseded duplicate | half day |
Step 02 of 04
Mistake 3 defines a retrieval audit: take 10-15 real queries, check exactly what chunks the system retrieves, and verify the output matches current facts, done monthly.
Running this month's 10-query retrieval audit, query 4 ('what's included in the Slack Business+ plan') retrieved a chunk from a pricing doc dated January 2025, even though the curated set now only contains the Q3 2026 doc. What does that tell you the audit just caught?
Procedure
- Run all 10 audit queries through the assistant
- For each answer, open the citation/source trace and note which document and date it pulled from
- Flag any answer sourced from a document dated more than 90 days before today
- For each flag, check whether that document is still in the active knowledge base or the retrieval index simply wasn't rebuilt after removal
Retrieval audit, 10 queries, this month Q1: 'Slack Enterprise Grid pricing tiers' sourced from Q3 2026 doc CURRENT Q4: 'Business+ plan inclusions' sourced from Jan 2025 doc STALE, doc removed from KB but index not rebuilt Q7: 'Slack Connect eligibility' sourced from Mar 2026 doc CURRENT ...7 more rows 1 of 10 answers (10%) sourced a document that was already removed from the active set
Healthy
Catching the stale index and immediately triggering a full re-index of the vector database, then re-running query 4 to confirm it now pulls Q3 2026.
Unhealthy
Assuming removing a document from the admin panel automatically and instantly updates every retrieval, without re-testing the specific query that used to pull from it.
What this means
Deleting a source document doesn't guarantee the vector index is rebuilt immediately; the retrieval audit is the only way to catch a stale index before a customer-facing answer does.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| A removed document's content still shows up in an answer weeks later | Trigger a manual re-index and re-run the specific query that surfaced the stale chunk | 5 min |
Step 03 of 04
Mistake 2 requires date-stamping every document and a mandatory quarterly review, and assigning a single knowledge base owner responsible for keeping the set current.
The stale January 2025 pricing doc that showed up in query 4 had no assigned owner and no scheduled review date. Who should own fixing this, and what's the actual process fix, not just the one-time patch?
Procedure
- Assign one named owner to the 12-document curated set
- Add a next_review_date to every document, no more than 90 days out
- Set a recurring calendar reminder tied to the earliest next_review_date
- Document the removal-triggers-reindex step so it isn't a manual afterthought next time
Ownership log, Slack RAG knowledge base Document Owner Next review Pricing (Q3 2026) Priya (PMM) 2026-11-15 Brand voice guide v4 Priya (PMM) 2027-02-01 Compliance: data residency Legal (Raj) 2026-10-01 Process fix logged: any document removal now triggers an automatic re-index job, not a manual ticket
Healthy
Assigning a named owner and a hard review date to every document, including a written process for what happens on removal.
Unhealthy
Fixing today's specific stale-pricing incident by hand and moving on, with no owner or review date attached to prevent the same failure next quarter.
What this means
A one-time fix addresses this month's symptom; an owner plus a review date plus a documented removal process addresses the actual cause.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| The same type of stale-document incident recurs every few months | Assign an explicit owner and review cadence to every document in the knowledge base, not just the one that just failed | 30 min |
Step 04 of 04
Mistake 5 warns that unstructured walls of text produce large, imprecise retrieved chunks, while clear headers, short paragraphs, and structured formats produce more precise retrievals.
Query 4's Business+ pricing answer came back vague and generic even after the re-index fixed the staleness, because the Q3 2026 pricing doc is a single unstructured paragraph covering all four plan tiers at once. What's the actual structural fix?
Procedure
- Open the current pricing doc and check whether tiers are broken into separate sections/headers or one continuous paragraph
- Reformat into one clearly headed section per plan tier, each with its own bullet list of inclusions
- Re-run query 4 after reformatting and compare the retrieved chunk's precision
- Apply the same headers-and-bullets structure to the next 2 highest-traffic documents
Before: single 400-word paragraph covering Pro, Business+, and Enterprise Grid pricing together After: 3 separate headed sections (Business+, Pro, Enterprise Grid), each 60-80 words with a bulleted inclusions list Query 4 retrieval, before: pulled a 400-word chunk covering all 3 tiers, answer had to guess which parts applied to Business+ Query 4 retrieval, after: pulled only the 70-word Business+ section, answer became specific and accurate
Healthy
Reformatting the pricing document into per-tier sections so retrieval can pull exactly the relevant chunk instead of the whole document.
Unhealthy
Concluding the RAG system itself is broken or 'not smart enough' when the actual issue is an unstructured source document forcing an imprecise chunk.
What this means
Retrieval precision is bounded by source document structure, no amount of re-indexing fixes a chunk that's too broad because the underlying document was never split into sections.
So what do I do about it?
| Symptom | Action | Effort |
|---|---|---|
| An answer is technically sourced from the current document but still reads vague or over-broad | Check the source document's structure before assuming a model or indexing problem, split it into clearly headed sections | 30 min |
Final deliverable
A cleaned 12-document knowledge base list with owners and review dates, plus a completed 10-query retrieval audit log with pass/fail per query.
See a reference example
Awfis RAG retrieval audit (excerpt) Query: 'What's included in the Awfis Fully Serviced Office plan?' Retrieved from: Awfis Enterprise Solutions Sheet, dated July 2026 PASS Query: 'Awfis hot desk price in Bengaluru' Retrieved from: pricing_final_v2.docx, no date FAIL, flagged for owner review Audit result: 9/10 PASS, 1/10 FAIL (undated legacy document still in index)
Success criteria
You're done when you can:
- Correctly cuts the document set to only priority-mapped, non-duplicate documents
- Identifies the stale-index failure in the retrieval audit rather than assuming the model is wrong
- Assigns both an owner and a review date, not just a one-time fix
- Diagnoses the vague-answer symptom as a document-structure problem, not a retrieval-engine problem