Custom AI agent or off-the-shelf copilot: how to choose
An off-the-shelf copilot is a per-seat productivity tool: a person asks, it answers, the person still does the work. A custom AI system owns a process end to end — it reads the input, applies your rules, writes to your systems and stops at the approval point. Copilots cost per user per month forever; a bespoke build is a one-off with the source code yours.
Two years ago the question was whether AI worked. Now it is what exactly you are buying — because two very different products are sold under the same heading, and they are not substitutes for one another.
This article separates them: per-seat assistants that help individuals work faster, and bespoke systems built to own one process. It covers what each is good at, what they cost, and the situations where neither is the right answer.
What “custom AI” actually means
The phrase misleads, because the first thing it suggests is wrong: training your own model from scratch. Very few organisations should, and almost none need to. Model training is not a line item you build an operations case on.
In practice a custom AI system is three layers:
1. The model layer — rented. You use a hosted or open model. Which one you pick changes quality and cost. It does not create advantage; your competitor can rent the same thing.
2. The data layer — yours. Contracts, technical documentation, past quotes, product masters, customer correspondence. How the model reaches them — search, filtering, permission checks — is what decides whether the output is useful.
3. The decision layer — yours. Which output is written straight through, which goes to a queue, who approves it, what gets recorded. Without this layer you have a demo, not a system.
The value sits in layer three. Layer one is identical for everyone, layer two you already own, and layer three is what makes the first two worth anything.
The comparison, line by line
| Dimension | Off-the-shelf copilot | Custom AI system |
|---|---|---|
| What it does | Person asks, it answers, person does the work | Runs the process, stops at the decision point |
| Commercial model | Per user, per month, indefinitely | One-off build plus optional maintenance |
| Data access | Whatever the vendor has connected | Systems you choose, under your permission rules |
| Output | Text on a screen | A record written, a document produced, a job created |
| Customisation | Prompt level | Rules, workflow, interface and integration level |
| Source code | Vendor's | Yours |
| Cost at 50 users | Grows linearly with headcount | Independent of headcount |
| Time to value | The day licences are activated | Two to four weeks for one process |
| Exit cost | Export data, adopt a replacement | None; code and data are already yours |
A concrete anchor on the seat side: Microsoft 365 Copilot sits around $30 per user per month at the enterprise tier, and industry pricing guides note that the real all-in cost lands between $30 and $90 per user per month once the base licence is counted. Adoption is the other half of that arithmetic — the same guides observe that paid seats represent a small fraction of the eligible commercial base, which tells you an idle seat costs exactly what an active one does.
None of this makes copilots a bad purchase. It makes them a different purchase.
Why so many AI programmes get cancelled
The uncomfortable evidence belongs near the top, not buried at the end. A widely reported MIT-affiliated study found that the overwhelming majority of enterprise generative AI pilots produced no measurable P&L impact. Gartner separately predicts that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.
Read the stated reasons carefully: none of them is model capability. They are integration gaps, outputs never wired into a process, and projects without an owner. That is layer three, missing.
The practical implication is a rule you can apply this week. Programmes that start with “which model” tend to end up in that statistic. Programmes that start with “which process, which input, which output, who approves” tend not to.
A worked scenario: reading technical specifications into a quote
The process. A machinery manufacturer’s sales engineering team responds to customer specifications.
The input. A 14-page PDF specification arrives by email: 40 line items, tolerance values, delivery terms and penalty clauses.
Today. The sales engineer reads it, transcribes items into a spreadsheet, tries to recall comparable past jobs, looks up unit prices, and writes the quote. Two days.
With the system.
- Line items, quantities and tolerances are extracted into a structured table.
- Each item is matched against your product catalogue; anything it cannot match confidently is flagged, not guessed.
- Past quotes are searched for comparable specifications and surfaced with citations: which job, which customer, what tolerance, what unit price was quoted.
- Delivery terms and penalty clauses are summarised separately, with any deviation from your standard contract language highlighted.
- A quote draft is produced — without prices.
Where the human approves. In three places, all mandatory. The engineer resolves flagged items manually, enters unit prices personally — the system shows historic prices with sources but never proposes one — and presses send.
The output. Two days of work becomes a few hours, and the saved time is not taken out of accuracy, because the pricing decision never moved.
Why doesn’t the system price it? Because pricing is irreversible in effect and language models are the wrong tool for arithmetic. Calculation belongs in code; commitment belongs to a person. Every build that blurs this line loses more than it saves.
The decision test
Five questions settle it faster than any vendor demo.
Is the bottleneck individual output or a repeatable process? If people are slow at writing, a copilot helps. If a process is slow regardless of who runs it, seats will not touch it.
Does the output have to land somewhere? If the answer is “on a screen”, a copilot is probably sufficient. If it has to become a record in your ERP, a generated document, or a created work order, you need a build.
Do permissions matter? If different people must see different things, that requirement has to sit inside retrieval, and generic tools rarely expose it at the granularity you need.
How many people, for how long? Seat pricing is cheap for small teams and short horizons. Draw both curves over three years with your own headcount before deciding.
Is the capability strategic? Commodity capability — transcription, generic drafting, meeting summaries — should be bought. Process logic that reflects how your business actually works should be built.
Getting the schedule right
For a single-process build, four weeks look like this.
Week 1 — scope and data. The process is written down step by step, real input samples are collected (real ones, not cleaned ones), and access to target systems is verified. The output is a scope document and a test set. Without the test set there is nothing to measure the next three weeks against.
Week 2 — extraction and matching. The reading layer is built and the extracted fields are mapped to your data model. First accuracy measurement against the test set happens here, which is early enough for the answer to change the design.
Week 3 — workflow, approval and integration. Exception queue, approval screen and the write path into the target system. Most surprises land in this week, and the cause is nearly always undocumented behaviour in the target system.
Week 4 — parallel run. The system processes live work but nobody acts on its output; the existing process continues unchanged. The two are compared. When the gap closes, you switch over.
Skipping week four saves a week and costs you the first incident. Parallel running is not a technical step, it is how trust gets built.
Pricing
Our published bands, which sit at the focused end of the market rather than the enterprise-platform end:
- Single process: $3,000-6,000, two to four weeks. One process, one screen set, one integration.
- Multi-module system: $8,000-15,000, six to twelve weeks. Several processes, role-based permissions, reporting.
- Enterprise scale: $20,000+. Multiple sites, complex integration chains, audit requirements.
- RAG assistant over your documents: $5,000-12,000.
- Annual maintenance: 12-25%, optional.
Perpetual licence, unlimited users, source code handed over. Three factors push a quote upward: scanned rather than digital source documents, per-user permission rules, and target systems without a documented API. The full table is on the pricing page.
When not to do either
When the process is not written down. AI cannot learn a task nobody has described. Automating a process three people perform three different ways just accelerates three different errors. Write it first — that exercise often produces the improvement on its own.
When the answer lives in a database. “How many of this part do we have” is a query, not a retrieval problem. Wrapping a database in a language model makes it slower and less reliable.
When volume is low. Under a few hundred hours a year, the build cost does not come back. Multiply headcount by weekly hours by 52 before anything else.
When the underlying data is a mess. An assistant built over three versions of the same contract will confidently cite whichever matches the question. Retrieval is honest about the state of your files.
When the decision is irreversible. Payments, shipments, contractual commitments. The system drafts; a person commits.
When the goal is to be able to say you use AI. This is the single most common origin story for a cancelled project.
Measure your own situation
Five numbers, one afternoon. They will tell you which product you need better than any comparison table.
- Name one process. “General productivity” is not a process. “Extracting line items from incoming technical specifications” is.
- Headcount times weekly hours times 52. If the result is under 200 hours a year, this is probably not an automation candidate.
- What format does the input arrive in? Structured feed, digital PDF, scanned image, free-text email. If more than 30% is scanned imagery, budget accordingly — it is the largest single swing factor in a quote.
- Where does the output have to land? A screen, your ERP, an outbound email. “A screen” usually means a copilot is enough.
- Who approves, and how much time can they give it? Name the person. If an approval takes longer than about thirty seconds, nobody will do it, and the design has to change before the build starts.
Write those five answers on one page. That page will do more for the quality of the quotes you receive than any technical specification.
Next step
How we build AI into operational processes is on the AI process automation page, the kinds of organisations we work with on who we build for, and the delivery bands on the pricing page. The document-assistant side is covered in RAG chatbot for internal documents. Send the five answers above through the quote form and we will tell you which of the two products your problem actually needs.
Frequently asked questions
Does 'custom AI' mean training our own model?
Almost never, and for good reason. Training a foundation model is a capital project that produces no competitive advantage for an operations problem. Custom AI means putting a rented model behind your data, your business rules and your approval steps. The model is a commodity; the system around it is not.
We already pay for a copilot. Why would we build anything?
Because they solve different problems. A copilot helps fifty people write and summarise faster. It will not match a single invoice, post a single sales order, or route a single ticket. If your bottleneck is a repeatable process rather than individual writing speed, seats will not fix it — and most organisations end up needing both.
Where does our data go?
That is your decision and it should be written down before the build starts. Our default is that the application and database run in your own cloud account, in a region you choose. On the model side there are two options: a hosted provider API, or an open model running on your own infrastructure. The first answers better and adds an external processor; the second removes that and costs you quality and operational effort.
How long does a first system take?
Two to four weeks for a single process, six to twelve for a multi-module system. The schedule is set by data access, not by software: where the documents live, which systems expose an API, and who signs off. Answer those three and the timeline stops being a guess.
What happens when it gets something wrong?
You cannot drive the error rate to zero, so you make errors harmless instead. A well-built system drafts, cites its source and hands the decision to a person. No irreversible step — payment, shipment, external commitment — is ever left to the model alone.
Do we really get the source code?
Yes. Source code and database schema are handed over, the licence is perpetual and not tied to seat count. The practical consequence: ending the maintenance agreement stops development, not the system.
Related guides
- ERP customer portal vs dealer portal: which one your network needs
- AI automation: how European, Gulf and Turkish buyers differ
- AI workflow automation vs RPA: choosing the right layer
Service page: Package vs custom
Let's talk about what you need.
The 30-minute discovery call is free and carries no commitment.