Any platform · Updated · 9 min read · Lurto

RAG assistant answering from old documents: stale indexes, orphans and drift

Quick answer

An assistant quoting an old policy has no error state: retrieval succeeded, the answer was grounded, the citation is real, and every dashboard is green. Four faults produce it and they present identically, the document changed and nothing re-indexed it, the document was withdrawn at source and the vectors stayed, chunks from the old version survive alongside the new so retrieval returns both, or the index is fine and the right chunk simply sits below the cutoff. Diagnose them in that order, because chunks from withdrawn documents poison every retrieval experiment you run.

RAG index freshness monitor (n8n workflow)Compares your source of truth against what the index actually holds each morning and reports the three kinds of drift separately. Validated against the n8n schema before publishing.

Somebody asks the assistant about the returns window and it answers thirty days, confidently, with a citation. The policy changed to fourteen in March. The document was updated, the change was announced, and the assistant has been quoting the old number to customers for five months because nothing connected the edit to the index.

This failure has no error state. Retrieval succeeded, the answer was grounded, the citation is real, and every dashboard is green. Below are the four ways an index falls behind, how to tell which one you have, and a monitor that reports the gap each morning instead of leaving it to a complaint.

1. The document changed and nothing re-indexed it

The most common shape. Indexing ran once at launch, or runs on a schedule that only picks up new files, and edits to existing documents never trigger anything.

source:  returns-policy.md   modified 2026-03-14T09:22Z   hash a91f...
index:   doc_returns_policy  indexed  2025-11-02T14:10Z   hash 4c7e...

assistant: "You can return items within 30 days of delivery." [returns-policy.md]

The citation points at the right document, which is what makes it convincing. What it does not tell the reader is which version of that document produced the sentence. A modification timestamp comparison catches this in one query, and almost nobody runs it, because the pipeline was built to answer "is this document present" rather than "is this document current".

2. The document was withdrawn and stayed in the index

The dangerous one. A policy is deleted or superseded at the source and the vectors remain, so the assistant keeps answering from a document the company has explicitly retracted.

[retrieval] doc_8812 "pricing-2025-promo.md" (0.87)   // deleted at source in January
[assistant] "That plan is $49£40€45 a month with the launch discount applied."

Deletions rarely propagate because most indexing jobs are additive: they walk the source, upsert what they find, and have no step that asks what is in the index but no longer in the source. An orphan is worse than a stale document, since a stale one is at least a real policy that used to be right, while an orphan can be something legal asked to be taken down.

3. Chunks from the old version survive alongside the new

A subtler variant of the first two. Re-indexing ran, but chunk identifiers are derived from position or content rather than from the document, so the new version was added and the old chunks were never removed. Retrieval now sees both, and the two contradict each other.

[retrieval] chunk 4471 "returns within 14 days" (0.81)   from v3
[retrieval] chunk 2209 "returns within 30 days" (0.79)   from v1, orphaned

[assistant] "Returns are accepted within 14 to 30 days depending on the item."

That sentence exists in no document. The model was handed two contradictory passages and did what a helpful reader does, which is reconcile them. The tell is an answer that is oddly hedged or gives a range where your policy gives a number. The fix is deletion by document identifier before upsert, so a re-index replaces rather than accumulates.

4. The document is current and retrieval never reaches it

Worth separating out because the remedy is completely different. The index is fine; the right chunk sits below the cutoff, and an older document that matches the question's wording more closely wins.

query: "how long do I have to send something back"
  0.71  faq-shipping.md        "...send back within 30 days..."   [2024, superseded]
  0.68  returns-policy-v3.md   "...fourteen days from delivery..." [current]
  --- top_k = 1 ---

Six weeks of prompt engineering will not fix this, because the prompt is not the problem. Raising top_k and adding a reranking pass usually is, and so is deleting the superseded FAQ that should not have been competing in the first place. Which brings the diagnosis back to the same question as the other three: does the index contain exactly what the source contains, and nothing else?

Tell them apart before changing anything

The four faults need four different fixes and they present identically, so the ordering matters. Ask, in this sequence: does the index hold a version of this document that matches the source hash; does the index hold documents the source no longer has; does retrieval return more than one version of the same document; and only then, is the right chunk being retrieved at all. The first three are answered by comparing two lists, which takes minutes. The fourth is real retrieval work and is worth doing only once the first three are clean, because chunks from withdrawn documents will otherwise poison every experiment you run.

Check it every morning, not after a complaint

The attached workflow fetches your source listing and your index listing, compares them, and reports the three kinds of drift separately: documents never indexed, documents whose index copy is behind the source by more than a day or whose content hash differs, and documents in the index that no longer exist at source. It alerts only when something has drifted, and records a clean run when nothing has, so the execution history answers "when was the index last verified" without anyone having to remember.

It needs two things shaped by you, both marked in the workflow: a listing of source documents as { id, title, updatedAt, hash }, and the equivalent from your vector store. Pinecone, Qdrant and pgvector can all produce that metadata. The workflow validates clean against the current n8n schema, which we check before publishing anything importable.

When not to hire anyone

If your knowledge base is a handful of documents that change a few times a year, you do not need a monitor. You need a line in whatever process changes a policy that says re-index, and someone to run the comparison by hand each quarter.

If you have found one stale document and can re-index it, do that. One wrong answer traced to one document is not a project.

Where it becomes worth paying for is when documents change weekly, several people can edit them, and nobody can currently answer how many of the assistant's sources are current. At that point the fix is not a re-index, it is making the indexing an automatic consequence of the edit and having something that notices when it is not, which is a change to the pipeline rather than to the data.

What it costs to have us do it

A written diagnostic is $500£400€450, delivered in two business days, credited in full against whatever follows and refunded if it names no fixable cause. Here it means running the comparison against your live index and coming back with a count: how many documents are current, how many are behind, how many are answerable but withdrawn, and which of the four faults produced the answer that sent you looking.

Repairs land at $750£600€700 to $1,500£1,200€1,400 where the indexing job needs deletion-before-upsert and a freshness check, and $1,500£1,200€1,400 to $4,000£3,100€3,700 where the pipeline needs rebuilding so an edit at source triggers indexing on its own. Both leave the monitor running, because an index that is correct today drifts the first week nobody is watching.

Questions people ask about this

The citation points at the right document. Why is the answer wrong?

Because a citation tells you which document produced the sentence, not which version of it. Indexing that ran once at launch, or that only picks up new files, leaves edits to existing documents invisible. Comparing the source modification timestamp and content hash against the index catches it in one query, and almost nobody runs it, because the pipeline was built to answer is this document present rather than is this document current.

Why does deleting a document not remove it from the assistant?

Because most indexing jobs are additive: they walk the source, upsert what they find, and have no step that asks what is in the index but no longer at source. The orphan is worse than a stale document: a stale one is at least a real policy that used to be right, while an orphan can be something legal asked to be taken down.

The assistant gave a range where our policy gives a number.

That is the signature of two contradictory chunks in the index at once. If chunk identifiers are derived from position or content rather than from the document, a re-index adds the new version without removing the old, retrieval sees both, and the model does what a helpful reader does and reconciles them into a sentence that exists in no document. The fix is deletion by document identifier before upsert, so a re-index replaces rather than accumulates.

The document is current and the assistant still misses it.

Then it is a retrieval problem, not a freshness one, and the remedy is completely different. The right chunk sits below the cutoff while an older document that matches the question's wording more closely wins. Six weeks of prompt engineering will not fix it, because the prompt is not the problem. Raising top_k and adding a reranking pass usually is, and so is deleting the superseded FAQ that should not have been competing in the first place.

When do we need more than a re-index?

When documents change weekly, several people can edit them, and nobody can currently answer how many of the assistant's sources are current. At that point the fix is not a re-index, it is making the indexing an automatic consequence of the edit and having something that notices when it is not, a change to the pipeline rather than to the data. If your knowledge base is a handful of documents that change a few times a year, you need a line in the policy-change process that says re-index, and a manual comparison each quarter.

What each band includes

Every band below is a price agreed in writing before any invoice. The diagnostic comes off the repair in full, so if you go ahead you have paid nothing extra for the reading.

BandPriceTime
Written diagnosticWe read the workflows, logs, and prompts. Written root-cause report and a fixed repair quote. Credited in full against the fix, and refunded if it cannot name a fixable cause.from $500from £400from €4502 business days
Small fixOne clear failure: a broken integration, a bad prompt, a missing retry. When the diagnostic shows a minor break, you pay the minor price.$750-$1,500£600-£1,200€700-€1,4002-5 days
Standard rescueWhere most rescues land. Several failure points or a fragile architecture: root-cause fixes, error handling, alerts, and a trail you can audit.$1,500-$4,000£1,200-£3,100€1,400-€3,7001-2 weeks
RebuildOnly when repairing costs more than starting over. The diagnostic says so in writing, with both numbers, before you decide.$4,000-$10,000£3,100-£7,800€3,700-€9,2002-4 weeks

Full detail on the rescue page, and every other number we charge is on the pricing page.

Other symptoms we have written up

Longer reading

Send us the execution log and we will tell you what broke

Two business days and a price at the end of it. If the fix is small enough to do yourself, the report will say so.