Enterprise AI assistant cost: the running bill, not the build price
An enterprise AI assistant carries two separate cost structures: a one-off build and a monthly bill that moves with usage. Our published band for a retrieval assistant build is $5,000 to $12,000. The running bill is rarely dominated by model calls. Integration upkeep, exception handling, re-indexing and human review usually cost more. Below a few thousand assisted interactions a month, the fixed side does not amortise and the assistant is difficult to justify on cost alone.
Most proposals for a corporate AI assistant price the build. Very few price the year that follows. That gap is the subject of this article: what the system costs to keep running once it is in daily use, which of those costs are visible at proposal stage, and which only appear on the third monthly invoice.
The distinction matters because assistants do not behave like the software your budget process was built for. A licence is a number you agree once and then forget. An assistant bills you for what your staff do with it. The finance question changes shape: not “what does it cost”, but “what does it cost per thing it handles, and how many of those things will there be”.
If you are still deciding what an assistant is and where it fits, what an enterprise AI assistant does covers that ground. This article assumes you have decided it is useful and now need to know what owning one looks like.
Two pricing models, two budget behaviours
A per-seat licence is fixed, predictable and disconnected from usage. Adoption costs nothing at the margin, so you want as many people using it as possible. The waste is invisible: you pay the same whether the licence is used daily or never, and nothing on the invoice tells you which.
Per-transaction or per-token pricing is variable and moves with success. Every document read, every question answered, every retry adds to the bill. The waste is visible, which is an advantage, but so is growth, which is uncomfortable when the growth is exactly what you asked for.
The one-line summary: a licence punishes you for buying too much, and consumption pricing punishes you for succeeding.
Three practical consequences follow.
- Consumption pricing needs a ceiling. Per-department attribution, per-user daily limits and spend alerts are not optional extras. They are what stops a retry loop turning into an invoice nobody approved.
- Approval processes designed for fixed annual spend produce a monthly variance conversation. Somebody has to own that conversation before the first quarter closes, not after.
- The return question flips. With a licence you ask whether enough people use it. With consumption pricing you ask whether each interaction is worth more than it costs. The second question is harder, and it is the correct one.
Most vendors now blend the two: a platform fee plus usage. That is a reasonable structure, but it means you have to ask which part is which, and what the platform fee actually covers. Ask separately who absorbs model provider price changes, because those flow through to you unless the contract says otherwise.
Where the cost actually accumulates
The model call is the line everybody asks about and the one least likely to decide the budget. It is also the easiest line to forecast. The surprises come from the systems on either side of it.
Integration and its upkeep. The assistant reads from your ERP, your document store, your ticketing system. Those systems change: a field is renamed, an export format shifts, a permission model is reorganised. Each change means work on the connector. This is recurring cost, not build cost, and it is the line most often quoted once and then paid for repeatedly.
Exception handling. Every assistant produces a share of output that a person must check, correct or override. That share is highest in the first months and never reaches zero. Two things get missed when this is budgeted. First, the reviewer’s time is a running cost of the assistant, not a saving from it. Second, reviewing generated output is a different task from doing the work, and it is not automatically faster. Where approval points belong, and how to keep them cheap, is covered in human-in-the-loop approval design.
Content upkeep. An assistant that answers from documents is only as correct as the documents. Someone has to own retirement of superseded procedures and correction of wrong ones. If nobody owns it, the assistant does not save time. It turns stale documents into confident answers, which is worse than no assistant.
Put crudely: the model is a metered utility, and utilities are rarely what break a budget. Maintenance of the things plugged into it is.
The volume threshold
An assistant is economic when the value it produces exceeds fixed cost plus variable cost. That is arithmetic, not opinion, and it fits on one line:
Monthly volume x saving per interaction > (monthly volume x variable cost per interaction) + monthly share of fixed cost.
The three inputs are yours to supply.
- Saving per interaction. Minutes saved, multiplied by a fully loaded hourly cost, multiplied by the share of interactions the assistant can genuinely handle. That last share is where honest models and flattering ones separate.
- Variable cost per interaction. Model call plus retrieval plus, critically, the amortised cost of reviewing the ones that get flagged.
- Fixed cost per month. The build spread over its useful life, plus hosting and index costs. Our published band for a retrieval assistant build is $5,000 to $12,000, which sits on the pricing page alongside the other delivery bands.
Run those three numbers and a break-even volume falls out. The assistant payback calculator does exactly this arithmetic with your own inputs, and it deliberately asks you for the coverage share rather than inventing an industry average, because that share depends entirely on how well your documents cover your work.
Two thresholds are worth naming explicitly.
Below a few thousand assisted interactions a month, the fixed side rarely amortises. The exact figure depends on your hourly cost and how much time each interaction saves, but the shape holds: fixed cost divided by a small volume produces a large per-interaction number, and no per-token saving rescues it. Below that line, the cheaper answer is usually written documentation and better search.
Concentration matters as much as volume. Five hundred interactions a month from three people is a different proposition from five hundred spread across eighty. The concentrated case is easier to justify, easier to measure and easier to improve, because you can watch the three people use it.
The line items nobody quotes
| Line item | Why it exists | How it behaves |
|---|---|---|
| Search index hosting | Vectors and metadata have to live somewhere queryable | Scales with corpus size and query volume; usually small, never zero |
| Re-indexing | Documents change, and a stale index answers from the old version | Batch cost tracking your document churn, not your query volume |
| Evaluation loop | A fixed test set of real questions with known-correct answers | Re-run on every model, prompt or corpus change; costs engineering time and model calls |
| Human review | Flagged outputs, sampled outputs, approvals on irreversible steps | Highest in month one, declines, never reaches zero |
| Model version changes | Providers deprecate and replace models | Irregular, and each one means re-testing against the eval set |
| Permission drift | Staff move, folders get restricted, entitlements change | Retrieval-time filtering must track the source of truth continuously |
| Logging and retention | Every question and answer is a record | Storage, plus a retention decision, plus an audit obligation in regulated work |
| Connector maintenance | Source systems change without warning you | Unpredictable timing, predictable frequency |
None of these are exotic. All of them recur. The reason they are missing from quotes is that they are invisible during a demonstration, and a demonstration is what most buying decisions are based on.
When a proposal says a line is included, ask the follow-up: included up to what volume, and what is the price past it.
Why production costs a multiple of the pilot
Pilots almost always come in cheap, and the ratio to production varies too widely to quote as a single number. What does not vary is the reason, and it is structural rather than accidental.
- The corpus changes. A pilot runs on a curated set of clean documents. Production runs on the real store, which is larger, inconsistently formatted, and contains material nobody intended to index.
- Permissions appear. A pilot typically has none. Production has to filter at retrieval time so a user only ever searches what they are already entitled to read. That is real engineering, and it is the most commonly under-scoped requirement in this category, as the RAG cost breakdown sets out in more detail.
- The question distribution widens. Pilot users ask anticipated questions. The whole company asks things nobody scripted, which raises both volume and the amount of context each answer needs.
- Operations arrive. Monitoring, alerting, a named owner, a response path for when it is wrong. None of these exist in a pilot and all of them are running cost.
- The evaluation loop becomes permanent. In a pilot you judge quality by looking at it. In production you need a repeatable test, because otherwise nobody can tell whether last week’s change helped or hurt.
- Retention and logging become obligations. Especially where the assistant touches personal data or regulated records.
The useful framing: a pilot budget is not a production estimate divided by a scale factor, it is a different set of line items. Read the pilot as an experiment in answer quality, and derive the cost estimate separately.
This is also where the category’s failure rate comes from. Gartner’s prediction that over 40% of agentic AI projects will be cancelled by the end of 2027 names escalating costs and unclear business value among the causes, and an MIT-affiliated study found that most enterprise generative AI pilots produced no measurable P&L impact. A pilot that impresses and a system that pays are different achievements.
Architecture choices that lower the running bill
Five decisions do most of the work, and all five are made at build time. Retrofitting any of them is possible and costs more than building it in.
A deterministic pre-filter. Whatever arrives in a known structure should never reach a model. Structured messages, portal submissions, standard spreadsheets: route them straight to code. This is the same principle that decides which automation layer belongs where, applied to cost. It is the highest-leverage change available, because it removes calls rather than making them cheaper.
Caching, at three levels. Repeated questions are the normal case inside a company: the same policy, the same procedure, the same price list, asked by different people in the same week. Cache the answer, cache the retrieval, and use provider-side prompt caching for the fixed part of your instructions. Cache hits cost close to nothing and return faster, which helps adoption at the same time.
Context discipline. The largest variable in a per-call bill is how much text you send. Better retrieval means fewer and more relevant passages, which lowers cost and raises accuracy together. This is the rare lever that moves both in the same direction, and it deserves more attention than model selection.
Small-model routing. Classification, routing and extraction of well-defined fields do not need your most capable model. Judgement, summarising and drafting usually do. Splitting the traffic is a genuine saving, but it is also the one lever here that can quietly cost you quality, so it is only safe with a test set that tells you when it has.
Hard ceilings. A per-user daily cap and an alert on volume anomalies. The purpose is not saving money in normal operation, it is catching a retry loop or a misconfigured job in hours rather than at month end. Runaway loops, not gradual adoption, are what produce the invoices people remember.
What to ask before signing
Seven questions, each with a reason to ask it.
- What is the cost per transaction at our expected volume, and what counts as one transaction? A “transaction” that means one user question and one that means one model call differ by a large factor.
- Which lines are fixed and which move with usage? You need the split before you can forecast anything.
- Who absorbs model provider price changes? In both directions.
- What happens to the bill if usage doubles? Look for a cliff at a tier boundary rather than a straight line.
- What triggers re-indexing, and what does it cost? This is the line most often discovered late.
- What is your evaluation method, and who pays for the runs? No method means no way to prove a change was an improvement.
- What share of interactions do you expect to need human review in the first three months, and how does that share fall? A supplier with no answer has not run one in production.
A vendor who cannot answer the first and fourth questions has not operated a system at your volume. That is worth knowing before the contract rather than after.
When the honest answer is no
When volume is low and questions rarely repeat. Write the documentation instead. It costs less and it does not carry a monthly bill.
When the source documents are out of date. Fix the content first. An assistant over stale material does not save time, it converts wrong answers into confident ones and distributes them faster.
When the process is undocumented. An assistant built over three undocumented ways of doing the same thing will confidently teach all three.
When arithmetic is the actual job. Reconciliation and calculation belong in code, which is deterministic, testable and accountable. A language model is none of those.
When the objective is a board slide. This is the most commonly stated origin of a cancelled project, and it is usually visible in the first scoping meeting.
Next step
If you want a number rather than a range, the inputs are the ones above: monthly interaction volume, minutes saved per interaction, the share your own documents can actually cover, and the systems the assistant has to read from. Send those through the quote form and we will come back with a build figure and a running cost estimate as two separate lines, including the ones most proposals leave out.
Frequently asked questions
How much does an enterprise AI assistant cost per month?
There is no single figure, because the bill is a sum of one fixed part and one variable part. Fixed: hosting, the search index, and the build amortised over its useful life. Variable: model calls, re-indexing as documents change, and human review of what the assistant produces. Ask any vendor to quote those lines separately at your expected volume. A monthly number quoted without a volume assumption behind it is not a price, it is a guess you will be billed against.
Is per-transaction pricing better than a per-seat licence?
They fail in opposite directions. A per-seat licence punishes you for buying seats nobody uses, and the waste is invisible because the invoice never changes. Per-transaction pricing punishes you for succeeding, because adoption raises the bill month by month. Neither is inherently cheaper. What decides it is whether your usage is broad and shallow, which favours consumption pricing, or narrow and heavy, which favours a fixed licence. Model both against your own volume before choosing.
Why was our pilot so much cheaper than production?
Because a pilot and a production system are not the same product at different scales. The pilot ran on a curated set of documents, had no per-user permission filtering, served friendly users asking anticipated questions, and carried no monitoring, no evaluation loop and no retention obligation. Each of those is a real engineering line in production. Read a pilot as an experiment in answer quality, not as a cost estimate you can multiply up.
Which cost lines do vendors usually leave out of a quote?
Five, consistently: re-indexing when documents change, the evaluation loop that tells you whether an update improved or degraded answers, human review of flagged and sampled outputs, connector maintenance when a source system changes, and logging plus retention of every question and answer. None are exotic and all recur. If a proposal says these are included, ask what usage volume they are included up to, and what the price is past it.
Can we lower the model bill without lowering answer quality?
Yes, and the most effective levers do both at once. Sending less but better-selected text per call reduces cost and raises accuracy, because a smaller, more relevant context is easier to answer from. Routing structured input away from the model entirely removes cost with no quality effect. Caching repeated questions is close to free. Routing simple tasks to a smaller model is the one lever that can trade quality away, so it needs a test set behind it.
Related guides
- How much does a bespoke ERP cost, really?
- Gulf e-invoicing: what ZATCA and the UAE mandate actually require from your systems
- ERP customer portal vs dealer portal: which one your network needs
Service page: Pricing bands and scope
Let's talk about what you need.
The 30-minute discovery call is free and carries no commitment.