Who we build for — Companies automating internal work

Most of what companies want AI for is not an AI problem

The requests we receive are almost always the same shape: a team is drowning in repetitive work and someone has been asked to look into AI. About half of that work should be automated with plain rules, which are cheaper and exact. The other half is genuinely language work, and that is where a model earns its cost.

KodDelta builds both halves and tells you which one you have. Rules, routing and scheduled jobs where the input is structured; retrieval-based assistants that answer from your own documents with citations where it is not. A retrieval assistant is $5,000–12,000; an automation module is $3,000–6,000 in 2–4 weeks.

The first question: rules or a model?

This decision is made before anything is designed, and getting it wrong is the most common way an AI project becomes expensive. The test is the shape of the input, not the ambition of the outcome.

Use rules when the data arrives in fields. An approval threshold, a routing decision based on department and amount, a nightly reconciliation, a reminder chain, a document generated from an order. These are deterministic, fast, cheap and auditable, and they behave identically on the thousandth run and the first. Adding a language model to them makes them slower, more expensive and less predictable, and buys nothing. That is the subject of business process automation.

Use a model when the input is language. An email that has to be understood before it can be routed. A supplier document in an unpredictable layout. A question whose answer lives in a procedure nobody can find. A first draft of a reply that a person will review. Here the variation is too wide to enumerate, and rules degenerate into a list of exceptions that nobody maintains.

Most useful systems are a mix: the model handles the unstructured edge, the rules handle everything after it, and the record of what happened lives in the operational system rather than in a chat window.

The four patterns that repay

PatternWhat the system doesWhat stays with a person
Document reading Extracts fields from incoming supplier documents, delivery notes and specifications in varied layouts, and matches them against the order or record they belong to Approving low-confidence extractions and anything where the match is ambiguous
Classification and routing Reads a free-text request or message, decides what it is about and sends it to the right queue with a summary attached Handling the item itself, and correcting misrouted cases so the categories improve
Drafted replies Produces a first draft answer grounded in your own documents and past cases, with the sources attached Reading, editing and sending; nothing leaves the company unapproved
Question answering Answers questions from the company's own procedures, contracts, specifications and operational data, with citations Deciding what to do with the answer, and flagging documents that are out of date

What all four have in common is that a wrong answer is visible and cheap. That is the real selection criterion, and it matters more than how impressive the demonstration looks.

Retrieval, in one paragraph

An assistant that answers usefully about your company does not know anything about your company. It is given the relevant passages at the moment of the question: your documents are split into sections, indexed, and searched when a question arrives, and the retrieved passages are what the model is asked to answer from. This is why the assistant can be updated by updating a document rather than by retraining anything, and why every answer can carry a citation. The mechanics, including how retrieval quality is tested and where it fails, are set out on AI integration.

Designing for wrong answers

The question is never whether the system will be wrong. It is what happens when it is, and that is a design decision rather than a model property.

  • Grounding. Answers come from retrieved company documents, not from the model's general knowledge, so there is always a source to check.
  • Citations by default. Every answer names the document and section it came from. A citation is what turns trust into verification, and it takes the user seconds.
  • An honest refusal path. When retrieval finds nothing relevant, the assistant says so and hands over. This behaviour is built and tested deliberately, because the default behaviour of any model is to produce something plausible.
  • Confidence thresholds on extraction. Below the threshold, the item goes to a person instead of into the system. The threshold is tuned against real documents, not chosen in advance.
  • Approval before anything irreversible. Nothing is sent, paid, ordered or committed without a person, regardless of how confident the system is.
  • A log of what was answered. Questions, retrieved sources and answers are recorded, which is how quality is reviewed after launch instead of guessed at.

Permissions and where things run

Two boundaries decide whether an assistant is deployable in a real company. The first is permission: the assistant must see exactly what the person asking is allowed to see, which means access rules are enforced at retrieval, per user, rather than by asking the model to be discreet. An assistant that can quote a document the user could not open is a data leak with a friendly interface.

The second is location. Three architectures are practical, and the right one depends on your governance rather than on our preference.

ArchitectureWhat it meansTrade-off
Commercial model API Questions and retrieved passages are sent to a model vendor under a contractual no-training arrangement Best answer quality and fastest to build; requires approving an external processor
Region-pinned hosting The model runs in a region you specify, commonly the EU or the UK Satisfies most residency requirements; a narrower choice of models and a higher running cost
Open-weight model on your hardware The model runs on infrastructure you control, with nothing leaving your network Maximum control; needs suitable hardware and accepts a quality gap on the hardest questions

In all three cases the document index and the operational data stay in a store you own. Where corporate data lives and who controls it is discussed further in the data location guide.

Measuring it against your own baseline

We do not publish savings percentages. A figure taken from another company's process says nothing about yours, and any number we invented would be marketing. What works instead is a before-and-after on four figures you can collect in a week.

Volume: how many items of this kind arrive per week. Handling time: how long one takes from arrival to done. Rework rate: how often it comes back or has to be corrected. Queue age: how long the oldest waiting item has been waiting. Record all four before the build, and record the same four sixty days after go-live. That comparison is defensible in front of a finance director, which is more than can be said for a vendor's percentage.

The dashboards that make those figures visible without an export are the subject of business intelligence.

How a first AI build runs

  1. Fifty real questions, or two hundred real documents

    Not a demonstration. The actual questions your team is asked, or the actual documents that arrive, including the awkward ones. This set becomes the test that decides whether the answers are good enough, and it is assembled before anything is built.

  2. Decide the authoritative source per topic

    Which document is current when three versions exist. This is the work that most often blocks a build, it is useful whether or not the assistant is built, and it belongs to the business rather than to us.

  3. Prototype against the real set

    About 2 weeks. The measure is not whether it looks impressive; it is what percentage of the real question set is answered correctly with a correct citation, and how it behaves on questions it should refuse.

  4. Live for one team, with a person in the loop

    A retrieval assistant is $5,000–12,000; a rules-based automation module is $3,000–6,000 in 2–4 weeks. It goes live for one team first, with every irreversible action requiring approval, and the answer log reviewed weekly.

  5. Extend, and connect it to the operational system

    More document sets, more processes, and the write-back that makes it operational rather than informational — creating the record, updating the order, starting the approval. A multi-department system is $8,000–15,000 over 4–8 weeks; bands are on pricing.

When we say no

  • Judgement and negotiation. Work whose value is the decision itself does not become cheaper by being automated; it becomes worse.
  • More exceptions than rules. If almost every case is special, there is no pattern to learn or encode, and the honest answer is that the process needs simplifying first.
  • Irreversible actions without review. Payments, contractual commitments and decisions about people keep a person in the loop, whatever the confidence score says.
  • Low volume. A process that happens twice a month will not repay a build, and we will say so rather than write a proposal for it.
  • Contradictory sources. Where nobody can say which document is current, an assistant will confidently quote the wrong one. Fix the source of truth, then build.

The wider automation catalogue is on automation solutions, and what a complete built system contains is on custom software.

Frequently Asked Questions

Which internal processes actually suit AI rather than plain automation?

The dividing line is whether the input is structured. If the data arrives in fields and the decision follows written rules, use rules: they are cheaper, exact and they never drift. A model earns its place where the input is unstructured language or documents, and where a confident-but-wrong answer can be caught before it does damage. Reading incoming supplier documents, routing free-text requests, drafting replies for approval and answering questions from your own documents are the four patterns that repay most reliably.

How do you stop the assistant inventing answers?

Three mechanisms together, none of them sufficient alone. The assistant answers only from retrieved company documents rather than from general knowledge, so there is a source behind every sentence. Every answer carries citations to the specific document and section, which lets a user check in seconds instead of trusting. And when nothing relevant is retrieved, the correct behaviour is to say so and hand over to a person, which has to be designed and tested deliberately because a model will otherwise produce something plausible. An assistant that cannot say it does not know is not ready to be deployed.

Does our data have to leave the company?

No, and that is a design decision made at the start rather than a consequence of the technology. Three architectures are practical: a commercial model API under a contractual no-training arrangement, a model hosted in a region you specify, or an open-weight model on hardware you control. Cost, quality and effort differ across the three, and so does what your data protection officer has to approve. In every case, retrieval runs against your own document store and access rules follow the user.

How do we know whether it worked?

By measuring the same thing before and after, using your own numbers. Before anything is built, record the baseline for the chosen process: how many items arrive per week, how long each takes, how often it is reworked, and how long the queue is. After the module is live, measure the same four figures. We publish no savings percentage, because a number from someone else's process is marketing rather than evidence.

What does a first AI build cost and how long does it take?

A retrieval-based assistant working on your own documents and data is $5,000 to $12,000. A single automation module without a model, meaning rules, routing and scheduled jobs, is $3,000 to $6,000 and runs in 4 to 6 weeks. In both cases a clickable prototype comes in about 2 weeks, and the AI work starts with a fixed set of real questions from your own team rather than with a demonstration, because that is the only way to know before spending whether the answers will be good enough.

Do we need to have our data organised before we start?

Not perfectly, but you do need to know where it is. Retrieval works on the documents and records you already have: procedures, contracts, specifications, past correspondence and the tables in your operational system. What blocks a build is not untidy data but contradictory data — three versions of the same procedure with no way to tell which is current. Deciding which source is authoritative per topic is part of the first stage.

Which work should stay with people?

Judgement, negotiation, and anything where the exceptions outnumber the rules. Also anything where being wrong is expensive and cannot be checked before it takes effect: payments, contractual commitments and decisions about people. There the useful design is preparation rather than automation — the system assembles the evidence and drafts the option, and a person decides. We decline that work rather than sell a system that gets quietly abandoned.

Send us fifty questions your team actually gets asked.

Or two hundred documents that arrive every month. We come back with what is answerable today, what needs the source of truth fixing first, and a fixed price.

Request an AI scope