Answers From Your Own Documents: RAG That Works

  • Home
  • <
  • Blog
  • <
  • Answers From Your Own Documents: RAG That Works
A team reviewing printed documents and reports spread across a meeting table
RETRIEVAL & KNOWLEDGE AI

Answers From Your Own Documents: RAG That Works

Neo Hives IT Solutions· 10 September 2026·13 min read

Somewhere in your business is a document that answers the question your team just spent forty minutes discussing. It is a PDF, in a folder, with a filename ending in _final_v3_updated, and the person who wrote it has left. This is the problem a knowledge assistant solves — not intelligence, retrieval. The answer already exists; nobody can find it.

This is a practical guide to building one: what retrieval-augmented generation actually is, why the hard parts are permissions and staleness rather than the model, how to handle the documents that break every naive implementation, and how to tell whether the thing is right often enough to be trusted. The general framing of where automation pays is in our guide to AI business process automation; this is the one service almost every business asks about second.

What retrieval actually is, in plain terms

A knowledge assistant does two things in sequence. It searches your documents for passages relevant to the question, then it asks a language model to answer using only those passages, with a link back to where each came from. That is the whole idea. The model supplies language and reasoning; your documents supply facts.

One clarification worth making early, because it is the first question every serious buyer asks: this does not involve training a model on your data. Your documents are searched at the moment a question is asked and the relevant extracts are passed in for that one answer. Nothing is baked into model weights, nothing leaks into anybody else's answers, and removing a document removes it from the assistant. That distinction matters commercially as well as technically — it is the difference between a project your legal team can approve and one it cannot. What does leave your network, and under whose contract, is a separate and real question, and we cover it in data privacy for AI deployments in India.

The two failure modes that actually kill these projects

Not hallucination. Not chunk size. These two:

1. Permissions. The moment an assistant can read every file in the company, it can tell an intern what the sales director earns. Permissions must be enforced at retrieval time, per user — the search itself returns only documents that person is already allowed to open. What does not work, and what gets built alarmingly often, is retrieving everything and then instructing the model not to reveal sensitive things. Instructions are not access control. If a document reaches the model, treat it as disclosed.

This has a design consequence people find inconvenient: your assistant inherits whatever mess your existing permissions are in. If half the company has access to a shared drive nobody has audited since 2021, the assistant will make that visible — quickly, and to everybody. That is not a reason to avoid building it. It is a reason to run the permissions audit as part of the project rather than discovering it in week three.

2. Staleness. An assistant that confidently quotes a policy superseded eight months ago is worse than no assistant, because a person searching manually would have noticed the 2024 date on the header and kept looking. The model will not, unless you make it. Practical defences:

  • Keep effective dates and version numbers in the indexed metadata, and have the answer state them: "according to the leave policy effective 1 April 2026…". A dated answer is one a person can sanity-check in a second.
  • Delete or explicitly mark superseded documents. The most common cause of confident wrong answers is that both the old and new versions are in the index, and retrieval has no way to prefer the newer one.
  • Prefer recency when documents conflict, and where the conflict is material, surface both and say they disagree. "Two policies address this and they differ" is a genuinely useful answer.
  • Re-index on change, not on a schedule if you can — and know your lag if you cannot.

Chunking is boring, and it decides everything

Documents get split into passages before indexing. Split them badly and the correct answer is never retrieved, no matter how good the model is. Most disappointing pilots are diagnosed as "the AI is not clever enough" when in fact the relevant sentence was cut in half.

Document typeWhat breaksWhat to do instead
Policy and contract PDFsFixed-size splits cut clauses in half, so a retrieved passage carries a condition without its exceptionSplit on structural boundaries — clause, section, heading — and keep the heading path with each passage
Tables and price listsText extraction turns a table into a column of orphaned numbers with no row or header contextExtract tables as tables, keep headers with every row, and store the surrounding caption
Scanned documentsNo text layer at all, so they index as empty and vanish silentlyOCR as a separate pipeline step, with a report of what failed rather than a silent skip
SpreadsheetsRetrieval over cells answers badly a question that is really arithmeticQuery the data, do not retrieve it — a database question deserves a database answer
Slide decksEach slide has five words and no contextMerge slides into their section, and include speaker notes
Long manualsThe answer needs three separated pages at onceOverlap passages, keep a parent-document reference, and pull neighbouring passages in when one is retrieved

The rule of thumb: a retrieved passage should be understandable on its own to a person who has not seen the rest of the document. If you would not accept it as a quotation in an email, the model should not be asked to answer from it.

Search: use both kinds

Semantic search — comparing meaning rather than words — is what makes these systems feel clever. It is also what makes them fail on the queries businesses ask most often, because it is bad at exact strings. Ask about invoice INV-2026-0847, part number HX-44B or clause 7.3, and pure semantic search will helpfully return documents about broadly similar invoices, parts and clauses.

  • Run keyword and semantic search together and combine the results. This one change fixes more real complaints than any amount of model upgrading.
  • Rerank the combined shortlist before it goes to the model. Retrieving twenty candidates and passing the best five is both more accurate and cheaper than passing twenty.
  • Filter on metadata first — department, document type, date range, and above all the user's permissions. Narrowing the haystack beats improving the needle detector.
  • Handle the vocabulary gap. Staff ask about "casual leave"; the document says "discretionary absence". A short synonym map, maintained as you watch real queries, is unglamorous and effective.

Citations are the whole trust mechanism

Every answer must link to its sources, and the link must land on the passage rather than the front page of a 200-page PDF. This matters less because people check every answer — they do not — and more because they check the first few, find the citations accurate, and calibrate their trust correctly from then on. An assistant without citations gets one wrong answer noticed and is abandoned; an assistant with citations gets the same wrong answer noticed and corrected.

The corollary is that "I don't know" has to be an available answer. A model asked a question its retrieved passages do not cover will, by default, produce something plausible. Configure it to say the documents do not contain the answer, show what it did find, and offer to route the question to a person. Users forgive an assistant that declines. They do not forgive one that invents a notice period.

When you do not need retrieval at all

Three cases where the honest answer is to build something simpler:

  • Your corpus is small. If the whole body of knowledge is thirty pages of policy, put it directly in the prompt. Retrieval adds moving parts and a new failure mode to solve a problem you do not have.
  • The question is really a database query. "How many orders shipped late last month?" is SQL. Retrieval over exported spreadsheets will produce a confident wrong number, which is the worst possible outcome.
  • The documents do not contain the answer. If the knowledge lives in one person's head, no retrieval system will find it. The project you need is writing it down — which, incidentally, is worth doing regardless.

What it costs to run

The recurring cost of a knowledge assistant is dominated by what you send the model on every question, not by storage. A few specifics worth knowing before you budget:

  • Retrieved context is the main variable. Passing ten large passages when three would answer the question multiplies your per-question cost by a factor you never see until the invoice.
  • Indexing is a one-off, until it is not. Changing the embedding model means re-indexing everything. Budget for that happening roughly once a year, and keep the pipeline reproducible so it is an afternoon rather than a project.
  • Caching earns its keep. A stable instruction block reused across every question, and repeat questions answered from cache, together take a serious bite out of the bill. In most deployments a surprising share of questions are near-duplicates.
  • Small models are often enough. When facts come from your documents, the model's job is comprehension and summarising, not knowledge. That is exactly the workload where a cheaper model performs indistinguishably.

As with any automation, the number to manage is cost per answered question at your real volume, including the questions that get retried and rephrased — not a price per million tokens.

How to tell whether it is any good

Build a fixed set of real questions with known correct answers — fifty is enough to be useful, drawn from what your team actually asks in chat and email — and score every change against it. Include the awkward ones: questions whose answer is genuinely absent, questions where two documents conflict, questions with an exact reference number, and questions a particular user should not be able to get an answer to.

  • Retrieval hit rate: was the correct passage in the retrieved set at all? Diagnose this before blaming the model, because most wrong answers are retrieval failures wearing a costume.
  • Answer accuracy against the known answer.
  • Citation accuracy: does the cited source genuinely support the claim? A right answer with a wrong citation is a future incident.
  • Abstention rate on unanswerable questions. This should be high. If it is near zero, your assistant is guessing.
  • Permission leakage: zero tolerance, tested explicitly with accounts at different access levels.

That whole discipline — and how to build the evaluation set without it becoming a research project — is the subject of how to test an AI system before customers do.

Where knowledge assistants go wrong

  • Permissions bolted on after launch, or enforced by asking the model politely.
  • Superseded documents left in the index alongside current ones.
  • Semantic search only, so every question containing a reference number fails.
  • Scanned PDFs indexed as empty files, with nobody told.
  • No citations, so the first wrong answer destroys trust in all the right ones.
  • Tables flattened into unusable text, which is how price questions get answered wrongly.
  • An assistant pointed at a document graveyard nobody curates, on the theory that more data is better.
  • Spreadsheet questions answered by retrieval instead of by querying the data.
  • No evaluation set, so quality is whatever it seemed like in the last demonstration.

How we build these at Neo Hives IT Solutions

We start with the questions, not the documents: a list of what people actually ask, where they ask it, and how long each answer currently takes to find. That list determines which documents are worth indexing, and it usually turns out to be a fraction of what was originally proposed — one well-curated corpus beats a comprehensive one.

Then permissions, before features. Then a narrow first version over a single well-understood document set, with citations on every answer and abstention switched on. Then measurement against a fixed question set, so the second version can be shown to be better rather than asserted to be. Where the assistant needs to act on what it finds rather than just answer — raise the ticket, update the record, send the reply — that crosses into agent territory, and the difference is worth being deliberate about; we set it out in chatbot, assistant or agent. The service in plainer language sits on our answers from your own documents page, and where the real problem is that the documents do not exist yet, that is closer to IT consultation — and we would rather say so.

Common questions

Do you train a model on our documents? No. Documents are searched per question and the relevant extracts are passed in for that answer only. Nothing enters model weights, and deleting a document removes it from the assistant immediately.

Will it work on scanned paper documents? With OCR as an explicit step, yes, and quality depends on the scans. The important thing is that OCR failures are reported rather than silent — the classic bug is a hundred scanned contracts indexed as empty documents and nobody noticing for a month.

How many documents is enough — or too many? Enough is however many cover the questions people actually ask; that is often a few hundred pages. Too many is a shared drive of duplicates and drafts, because retrieval will faithfully surface the draft. Curation beats volume in every deployment we have seen.

Can it run entirely inside our network? Yes, with open models on your own or a private cloud, at some cost in answer quality and a real cost in operations. It is the right choice for genuinely sensitive corpora and an expensive habit otherwise. Decide it on the sensitivity of the documents, not on general unease.

Why does it sometimes get things wrong? Usually because the right passage was never retrieved, and occasionally because two documents disagree and it picked one. Both are diagnosable, which is the point of citations — you can see what it read and fix the cause rather than argue about the model.

How long does a first version take? Weeks for one document set, with the time going into access, permissions and OCR rather than into the model. Corpora that are already tidy move considerably faster, which is the argument for tidying regardless.

Where to start

This week, collect the last twenty questions your team asked each other that were already answered somewhere in a document. Write down how long each took to resolve. That list is your evaluation set, your scope and your business case in one, and it costs an afternoon to assemble.

Our free AI readiness audit is a 45-minute conversation and a short written summary of whether your documents are in a state where this will work — including the unwelcome answer that a permissions clean-up and a decent search box should come first. Or get in touch and tell us what your team keeps asking each other.