RAG chatbot over internal documents: what it costs and when it works
A working retrieval assistant over your own documents costs $5,000–12,000 and takes 2–4 weeks on our published bands. Market quotes range far wider because the expensive part is never the model — it is document extraction, chunking, metadata and permissions. Answer quality is set by document quality, not by which model you pick.
Every company has the same pile: procedures, contracts, technical specifications, supplier terms, past quotes, meeting notes, a decade of email attachments. And the same daily cost — someone who knows where things are, being interrupted to find them.
A retrieval assistant is a genuinely good fit for that problem. It is also a category where quoted prices vary by a factor of fifty, which makes it very hard to tell whether you are being sold a small sensible thing or a large speculative one.
What RAG is, mechanically
Four steps, and it is worth understanding them because every cost line maps to one of them.
- Ingestion. Documents are collected from wherever they live and their text is extracted. PDFs, scans, spreadsheets, email attachments, shared drives.
- Chunking and indexing. Text is split into passages and converted into vectors that capture meaning rather than keywords, then stored in a searchable index alongside metadata: source, date, department, access level.
- Retrieval. A question is converted the same way, and the most relevant passages are found — filtered, critically, by what the asking user is allowed to see.
- Generation. A language model receives those passages and the question, and answers using only them, with citations back to the source.
Step four is the part everyone imagines. Steps one and two are where the work is.
Why the price range is so wide
Published industry guides put RAG builds anywhere from $15,000 for a basic single-source system to $300,000-plus for an enterprise platform, and note explicitly that the biggest cost driver is not the LLM but the data work — extraction, cleaning, chunking, metadata design, retrieval tuning and permission handling.
That is the whole explanation. Two projects described with the same three letters can differ by an order of magnitude in scope:
| Factor | Cheap end | Expensive end |
|---|---|---|
| Document format | Digital text, consistent structure | Scanned images, handwriting, mixed languages |
| Sources | One folder or one system | Six systems with different auth models |
| Permissions | Everyone sees everything | Per-document, inherited from an existing directory |
| Freshness | Re-index weekly | Near-real-time on document change |
| Output | Answer with citations | Answer plus actions taken in another system |
| Audit | None | Every query and retrieval logged and retained |
| Hosting | Managed cloud | Self-hosted model, air-gapped |
Our published band of $5,000–12,000 over 2–4 weeks sits at the focused end of that table: a working assistant over a defined document set with citations and permission filtering. It is not an enterprise knowledge platform, and it is not sold as one. The full band table is on the pricing page.
The five things that decide whether it works
Document quality beats model choice. If your procedures contradict each other, the assistant will faithfully surface contradictions. If three versions of a supplier contract are on the drive with no way to tell which is current, it will cite whichever matches the question best. Retrieval is honest about the state of your files, which some organisations find uncomfortable.
Scanned documents need OCR, and OCR is a project. A clean digital PDF is nearly free to ingest. A scanned, stamped, handwritten-annotated document is not. Ask early what proportion of the corpus is image-only, because it is the single largest swing factor in a quote.
Chunking strategy is not a detail. Splitting a contract mid-clause produces passages that retrieve well and read wrongly. Structure-aware chunking — by section, with the heading path preserved as metadata — is much of the difference between a demo and something people keep using.
Permissions must apply at retrieval, not after. If the system retrieves everything and then tries to hide what the user should not see, content leaks through summaries and paraphrases. The filter belongs in the search query. This requirement is routinely missing from proposals; ask about it explicitly.
Citations are the product. An answer without a source is a claim. An answer with a one-click link to the paragraph it came from is a tool, because the user can verify it in seconds. Never accept a build without citations.
Where it earns its cost, and where it does not
Good fits, in rough order of how often we see them work:
- Technical and product documentation. “What is the torque specification for this assembly?” Answers exist, are written down, and are hard to find.
- Supplier and customer contract terms. “What are our payment terms with this supplier and what is the notice period?” Currently answered by finding the PDF and reading it.
- Internal procedures and onboarding. “How do we handle a warranty claim on a unit out of its window?” Absorbs an enormous amount of senior people’s time.
- Historical quotes and specifications. “Have we built anything like this before, and what did we price it at?”
- Regulatory and compliance references where the source text is stable and the questions are repetitive.
Poor fits, which are worth naming because they get sold anyway:
- Questions whose answers are in a database, not a document. “What is our stock of part X?” needs a query, not retrieval. Wrapping a database in RAG makes it slower and less reliable.
- Anything requiring a calculation. Language models are not calculators. Route arithmetic to code.
- Decisions with consequences. An assistant may draft, summarise and cite. Approving a payment or releasing an order needs a person and an audit trail.
- Corpora nobody maintains. If the documents are wrong, a faster way to find them is not an improvement.
How to test it before you commit
This sequence takes about a week of your own time and eliminates most of the risk.
- Write twenty real questions, before seeing any demo. Take them from what people actually ask each other. Include five you expect to be hard.
- Write down the correct answers and their source documents. This is your test set, and building it is what makes the evaluation objective rather than impressionistic.
- Assemble a representative sample of the corpus — a few hundred documents including the ugly ones. Do not curate a clean sample; the ugly ones are the test.
- Ask for a pilot against your questions, not a generic demo. Any supplier who will only demo their own dataset is telling you something.
- Score three things per answer: is it correct, is the citation right, and would a new employee have been able to act on it? Correct-but-uncited is a fail.
- Check the permission boundary deliberately. Ask a restricted question as an unrestricted user. If anything leaks, stop.
- Get a per-query running cost and multiply by realistic monthly volume before signing anything.
If the pilot answers 15 of 20 correctly with accurate citations, you have something worth building on. If it answers 8, the problem is almost certainly your documents, and no amount of model tuning will fix it.
What it costs to run, not just to build
Build cost is a one-off. Running cost is forever, and it is the line most often left out of a business case.
| Line | Scales with | How to control it |
|---|---|---|
| Model inference per query | Number of questions × context size | Retrieve fewer, better passages; cache repeated questions |
| Embedding and re-indexing | Rate of document change | Index changed documents only, not the whole corpus |
| Vector index hosting | Corpus size | Prune obsolete documents; they were hurting quality anyway |
| Application hosting | Concurrent users | Conventional scaling |
| Document ingestion (OCR) | New scanned documents | Fix the source — digital originals cost nothing to ingest |
Industry guides put ongoing operating costs anywhere from $1,000 to $15,000 a month depending on volume and architecture, which is a range wide enough to be meaningless without your own numbers. Get a per-query cost at proposal stage, multiply by a realistic monthly question count, and multiply again by three for the growth that follows a successful launch. If that number is uncomfortable, the answer is usually to reduce the retrieved context size rather than to abandon the project.
The counterintuitive point: the largest lever on running cost is the same as the largest lever on answer quality. Retrieving five highly relevant passages is both cheaper and better than retrieving twenty mediocre ones.
Architecture decisions to make before the build
Six choices that are cheap now and expensive to reverse:
- Where the index lives. Your cloud account, in a region you choose, is the default we recommend. It keeps the data-residency conversation short.
- Which model, and whether question text leaves your infrastructure. Hosted models are better and involve an external processor. Self-hosted models avoid that and cost quality and operational effort. Decide deliberately and record the reasoning — if personal data is involved, this feeds directly into your transfer analysis, covered in GDPR and nearshore development.
- How permissions are sourced. Inherited from your existing directory or file system, or maintained separately. Inheritance is more work upfront and vastly less work forever.
- How documents get in. Push from source systems on change, or scheduled pull. Push is better and requires cooperation from the systems holding the documents.
- What happens to superseded documents. Removed from the index, or retained and marked. Retaining without marking is how an assistant confidently cites a policy that was replaced two years ago.
- Whether it takes actions. Answering questions and doing things are different products with different risk profiles. If it will eventually act, design the audit trail now.
Practical guardrails for a production rollout
- Start with one department and one corpus. Breadth is what turns a four-week project into a six-month one.
- Show the sources by default, not behind a toggle. It changes how people use the answers.
- Log every query. The query log is the best product roadmap you will ever get — it tells you exactly what people cannot find.
- Give it a way to say it does not know. An assistant that always produces an answer is producing some answers it should not.
- Set a re-indexing cadence and name an owner for document currency. Otherwise the assistant slowly becomes an archive of what used to be true.
- Keep documents and the index in your own cloud account. Access to them should be something you can revoke.
Our approach to AI in operations is set out on the AI process automation page, the wider automation context on business process automation, the delivery bands on the pricing page, and a scoped pilot can be requested through the quote form.
Frequently asked questions
What is RAG, in one paragraph?
Retrieval-augmented generation. Instead of relying on what a language model memorised during training, the system first searches your own documents for the passages most relevant to the question, then asks the model to answer using only those passages and cite them. It is the difference between an assistant that recalls and one that looks things up in your files.
Why do published RAG costs range from $5,000 to $300,000?
Because they describe different products. A single-source assistant over clean documents is a small project. A multi-source enterprise platform with per-user permissions, audit logging and compliance controls is a large one. Industry guides note the dominant cost driver is not the model but the data work — extraction, cleaning, chunking, metadata design and permission handling.
Will it hallucinate?
Less than an ungrounded chatbot, and never zero. Retrieval constrains the model to your documents and citations let a user verify any answer in one click. The remaining failure mode is a confident answer drawn from an outdated document, which is a content governance problem rather than a model problem. Design for verification, not for perfection.
Can it respect who is allowed to see what?
Yes, and it must. Permission filtering has to happen at retrieval time, so a user's query only ever searches documents they are already entitled to read. Bolting permissions on after retrieval leaks content through summaries. This is the single most commonly under-scoped requirement in RAG projects.
Do our documents leave our infrastructure?
That is an architecture choice you make deliberately. Documents and the search index can sit in your own cloud account in a region you choose. Whether question text reaches an external model provider depends on which model you use; self-hosted options avoid it entirely at some cost in quality. Decide this before the build, not after.
What does it cost to run each month?
Three lines: hosting for the index and application, model inference charged per query, and re-indexing as documents change. All three scale with usage rather than with headcount. Get a per-query cost estimate at proposal stage and multiply by realistic volume — this is the number that turns a cheap pilot into an expensive habit.
Related guides
- Resident support AI chatbot for property and estate management
- Gulf e-invoicing: what ZATCA and the UAE mandate actually require from your systems
- ERP customer portal vs dealer portal: which one your network needs
Service page: AI process automation
Let's talk about what you need.
The 30-minute discovery call is free and carries no commitment.