/

Knowledge work automation

AI Document Processing ROI: A Costing Model That Holds Up

AI Document Processing ROI: A Costing Model That Holds Up

11 min read

Summarize

V7 Go

A smarter way to manage due diligence and underwriting

Eighty percent of people who use AI at work say it has made them more productive. Thirty-seven percent of the firms they work for can attribute any earnings impact to it at all. Six percent can attribute five percent or more of EBIT. Those three numbers come from the same survey of 1,719 executives across 97 countries. The distance between the first and the last is where most AI ROI cases stall, somewhere between a successful pilot and a signed contract.

If you are the person inside a private markets firm carrying that case, you already know the shape of the problem. The pilot worked. Somebody on the deal team says it saves them two days a week and means it. Finance asks what the return is, and the honest answer is that nobody has costed it properly. The vendor quoted a licence. The licence is not the cost. And the saving is measured in hours that never appeared in a budget line to begin with.

This article is the costing model for AI document processing. It covers what the work actually costs across four layers rather than one, and which denominator survives a finance review. It shows why the architecture you choose sets your cost curve more than the model price does. And it sets out what to ask a vendor for in writing before you commit. It also covers the cases where the arithmetic does not work, because those exist and pretending otherwise is how a champion loses credibility.

Everything here works against any vendor, including this one.

Private Markets

Turn complex deal documents into faster investment decisions.

Private Markets

Turn complex deal documents into faster investment decisions.

Why AI saves everyone time and almost no one money

Time returned to a person is not money returned to a firm. That is the whole gap, and it is not a vendor conspiracy or a measurement failure. It is arithmetic.

Chart contrasting 95% of PE funds saying AI meets expectations with only 7% of portfolio companies at enterprise-scale deployment.

Conviction runs far ahead of deployment. Believing the technology works is not the same as having it in production at scale.

An associate who gets four hours back on Thursday does not hand the firm four hours of cost. They read one more data room, or they leave at seven instead of nine. Both are worth having. Neither shows up in the P&L unless the firm does something structural with the recovered capacity. It takes on deal volume it was previously turning away. It retires a spend line it was previously paying for. Or it delays a hire it was previously planning. Until one of those three things happens, the AI cost savings are real for the individual and invisible to finance.

McKinsey's State of AI survey for 2026 puts hard numbers on the split. Eighty percent of respondents report improved individual productivity and half report better decisions. Thirty-seven percent report any EBIT impact, a figure that has not moved year on year. Six percent qualify as high performers, meaning they attribute at least five percent of EBIT to AI and call the impact significant. Adoption is close to universal. Attribution is rare.

The same firm's work on measuring AI value names the failure mode directly. It calls it the pilot trap: an organisation releases pilot after pilot without ever scaling one, because teams cannot agree on what to measure or how to attribute the improvement. Sixty percent of organisations using generative AI in at least one function have seen no enterprise-wide EBIT impact. The technology is not the bottleneck in those firms. The accounting is.

For a private markets team the structural change that closes the gap is usually coverage rather than headcount. A mid-market fund that reviews eighty to a hundred targets in detail out of six hundred sourced is not short of analysts. It is short of analyst hours per target, which is why the screening decision gets made on a teaser and a phone call. Move the constraint and the return is not a salary saved. It is the deals that were previously screened out on capacity grounds and are now screened out on merit, or in.

Flow diagram showing AI compressing six M&A deal stages, from deal sourcing to legal close, with manual vs AI-assisted timelines.

Where the hours actually sit across a deal. The stages with the largest manual load are the ones where a costing model has something to work with.

That is a harder number to defend than a headcount reduction, and it is also the true one. Build the case on it and be explicit that you are doing so. A finance function that has seen three vendors promise a headcount saving that never materialised will trust a coverage argument more than a fourth promise of the same kind.

What AI document processing actually costs

The licence is one of four cost layers, and it is rarely the largest. Costing only the layer the vendor invoices you for is the most common error in this category. It is why so many first-year budgets are overrun by something the buyer never thought to count.

Cost layer

What it covers

How it behaves

Who usually forgets it

Inference

Model calls per document, per field, per retry. Priced per token by the model provider

Scales with volume and with how much of each document the system reads

Nobody. This is the layer everyone budgets

Platform

The workflow layer, storage, integrations, user seats, audit logging

Mostly fixed, steps up at usage tiers

Nobody. This is the invoice

Build and configuration

Scoping the use case, defining the schema, configuring steps, connecting sources, testing against real documents

Front-loaded, then small. Recurs when a document format or a template changes

Most buyers, when implementation is bundled and therefore invisible

Review labour

Human verification of the output before anyone acts on it

Scales with volume and inversely with accuracy. Often the largest single line

Almost everyone, because no vendor volunteers it

Review labour is the layer that decides whether the case works. Say a system returns a hundred fields per document and an analyst has to check every one of them. The workflow has not removed the manual step. It has moved it from extraction to verification, and verification is only cheaper if the reviewer can confirm a value faster than they could have found it. That is a property of the software, not of the model. A value that opens its own source document at the right line is a two-second check. A value that does not is a search.

The comparison that matters is not against zero. It is against what the work costs today. Most teams have never priced the current process either, which makes the baseline the first piece of real work in any business case.

A useful discipline is to cost the current process before looking at any vendor. Count the documents that pass through the workflow in a normal quarter. Count the fields or the decisions each one produces. Time the work honestly, including the waiting, the chasing and the second pass that happens when a number looks wrong. Multiply by whatever your firm uses as a fully loaded hourly cost. Do not use a benchmark for that figure. Your finance team has the real one and will not accept anyone else's.

The number most business cases get wrong

Cost per page and cost per seat are both the wrong ROI denominator. Cost per verified field is the one that survives a finance review, because it is the only one of the three that prices the review step into the unit.

Cost per page fails because pages are not the unit of work. A fund report and a capital account statement can run to the same length and still differ by an order of magnitude. What separates them is how much has to be pulled out and how hard each value is to find. Cost per seat fails for the opposite reason. It prices access rather than output, so it looks stable while the workload doubles underneath it, and it gives you nothing to compare against the manual baseline.

A bar chart comparing invoice processing costs for a team before and after implementing AI, showing a significant cost reduction.

Usage recorded per session with a running cost. A denominator you can audit starts with spend you can actually see.

Cost per verified field is built from four inputs, all of which a team can measure in a fortnight:

  • Inference and platform spend for the period, taken from the usage record rather than from a forecast

  • The number of fields the workflow produced in that period

  • The share of those fields a human had to open, check or correct

  • The average time that check took, at your firm's fully loaded hourly cost

Divide the first by the second and you have a machine cost per field. Multiply the third by the fourth and add it, and you have the number to compare against the manual baseline. It is the only figure in this article that a CFO can trace back to an invoice and a timesheet, which is precisely why it is the one to build on.

The share of fields needing review is where the arithmetic gets unforgiving, and it compounds faster than most people expect. At 97% accuracy per field, a document with twenty-five fields has something wrong on it more often than not. If your review policy is to check everything whenever anything might be wrong, per-field accuracy has to be very high indeed before the review line falls at all. Systems that let a reviewer confirm a value in two seconds change that calculation more than a two-point accuracy improvement does.

Two decimal places are not the point. An order of magnitude is. If the manual process costs roughly four dollars a field and the assisted one costs roughly forty cents, the case does not depend on getting either figure exactly right. If it costs three dollars sixty against three dollars ten, it does, and you should be suspicious of a case that close, because the error bars on both sides are wider than the gap.

Architecture, not model price, sets your cost curve

Two document processing systems can use the identical model on the identical documents and differ by more than two and a half times in cost per question. The gap widens as the document estate grows. That is not a pricing difference. It is a difference in how much work the system re-does every time somebody asks it something.

We measured it. Forty questions over synthetic private-equity disclosure documents, three configurations, two independently generated corpora, between four and eight repeats each, run at 10, 100 and 1,000 documents. One arm used a context layer built at ingestion. One used retrieval alone, on the production stack: documents chunked and indexed in pgvector with metadata filtering, plus keyword search and the ability to read whole documents. All arms used the same model, the same prompt and the same question set.

Measure at 1,000 documents

Context layer

Retrieval only

Cost per question

$0.128

$0.338

Tool calls per question

6.0

16.7

Median response time

18.8s

52.6s

Slow-case response time (95th percentile)

38.3s

224.1s

Accuracy

93.8%

72.1%

Cost behaviour from 10 to 1,000 documents

Flat

Rising

The flat line is the finding that matters for a budget. Structuring the content at ingestion means the reading has already happened, so a question costs six tool calls whether the corpus holds ten documents or a thousand. Retrieval does its reading at query time, so a larger index means more searching, more reformulation and more retries. Sixteen point seven tool calls per question is not a system that is failing. It is a system working very hard to compensate, and every one of those calls is billed.

Two caveats belong with those numbers, and we would rather publish them than have somebody find them. The first is that our 100-document measurement dips below both the 10 and the 1,000 result. Nearly all of it is two questions about which holdings a fund owns exclusively, where the system occasionally re-derives the split by hand instead of trusting the structure it already has. Both recover at 1,000 documents, which a problem caused by scale would not do. It is a reasoning bug worth about one and a half questions in forty. The second is that the test understates the combined configuration: every question was answerable from the structured layer alone, so search fired on 1.4% of them. The case where an analyst asks for a covenant detail or a footnote nobody modelled is exactly where the combination should pull furthest ahead, and this benchmark does not cover it.

The practical consequence for a budget is simple. Ask a vendor to quote cost per question or per document at your target volume rather than at pilot volume, then ask what happens to that figure when the corpus is ten times larger. We wrote up the underlying mechanism separately in knowledge graphs versus vector databases, and the product built on it is the Context Graph. The distinction is not academic. It is the difference between a cost line that is flat and one that grows with your own success.

Model choice still matters, just less than people assume, and mostly through step design rather than through picking a cheaper model overall. Published rates from Anthropic and OpenAI vary by more than an order of magnitude across a single provider's range. A workflow that sends every step to the largest available model is paying frontier prices to reformat a date. Give each step the tool the work requires instead. Code where the answer is arithmetic, a named model only where the judgment is real. That spends a fraction as much and gives up nothing on the steps that matter.

AI Implementation

Start with one workflow, then roll it out across the firm.

AI Implementation

Start with one workflow, then roll it out across the firm.

Building an ROI case a CFO will sign

A defensible case has three components and most drafts are missing the second. You need a baseline, a counterfactual, and a single ledger that carries costs and benefits together.

Bar chart of insurance AI time savings, led by a 96% cut in underwriting turnaround and 82% in FNOL intake.

Return expressed as a curve rather than a headline figure. The shape of the line is what a finance function will interrogate.

The baseline is what the work costs today, measured rather than estimated, using the method in the previous section. The counterfactual is what would have happened without the purchase, and it is the component that separates a case from a wish. If the team was going to hire two analysts and now hires one, the saving is one salary. If the team was never going to hire and now processes forty percent more documents with the same people, the benefit is coverage. Coverage has to be expressed as something the firm values: more targets screened to the same depth, LP queries answered in a day instead of a week, or a quarterly reporting cycle that closes on time without weekend work.

McKinsey's measurement framework runs five layers, from technical performance up through adoption and operational KPIs to financial impact. It carries one explicit instruction: total cost of ownership, including cloud and token spend, sits in the same ledger as the benefits. That last point is the one to enforce internally. A document processing case that reports hours saved in one paper and infrastructure spend in another is not a case. It is two papers that never have to meet.

Attribution is the part that gets argued about, so agree the rule before you start collecting rather than after. Pick one workflow, measure it for a full cycle before deployment, measure the same workflow for a full cycle after, and accept that the comparison is imperfect. A quarter with an unusual deal is not a controlled experiment. Say so in the paper, give the range, and let the range be honest. A case with error bars is more persuasive to a finance function than a point estimate that turns out to be wrong, because the second kind costs the champion their credibility on the next request.

From pilot to production without falling into the trap

The pilot trap is not caused by pilots failing. It is caused by pilots succeeding and never scaling, because success was never defined in terms that would trigger the next decision. Define it in advance. Name the volume the pilot must handle, the accuracy threshold it must clear on your own documents rather than on a demo set, and the review time per document it must achieve. Then name the date on which the decision gets made either way.

Start with one workflow that has a clear owner, a repeating cycle and a measurable output. LP reporting and portfolio monitoring both qualify, because they run on a fixed calendar and produce the same artefacts each quarter, which makes before-and-after comparison straightforward. Turning a CIM into an investment memo qualifies for a different reason: the volume is high enough that a small per-document saving compounds quickly. Screening a data room qualifies least, at least at first, because deal timing makes the cycles hard to compare.

Keep the review gate in the design. A workflow that removes the human check to improve its numbers is optimising the wrong variable, and in a regulated firm it is also the version that cannot be deployed. Security and access review belong in the same conversation for the same reason. Roughly a quarter of the buying conversations we have raise sub-processors, data residency and retention before anyone has looked at a price. A case that reaches procurement without those answers attached loses a month waiting for them.

What to ask for in writing

Ask for these as text, not as answers on a call. A vendor who will put a number in an email has thought about it. A vendor who will only say it out loud has not, or would rather you could not quote them later.

Table showing three enterprise use cases of LLMs: complex invoice processing, regulatory filing automation, and healthcare documentation. Each column includes a real-world scenario and specific ways LLMs help, such as extracting tax data from invoices, adapting to changing regulations, or normalizing medical terminology.

Ask which of these lines is fixed, which is metered and what triggers a move between tiers.

  • What is the total first-year cost at our volume, itemised? Licence, inference, implementation and support, as separate lines. If inference is metered, ask for the estimate and the assumptions behind it.

  • What is the cost per document or per question at ten times our current volume? This is the cost-curve question. A vendor whose architecture holds will answer it flatly. One whose architecture does not will change the subject to per-seat pricing.

  • What share of outputs do comparable customers review, and how long does a review take? You are asking them to quantify your largest cost line. An honest answer includes a range and the conditions that move it.

  • What happens to the workflow when a document format changes? Reconfiguration is a recurring cost. Ask whether it is billed, who does it, and how long it takes.

  • Who are your sub-processors, where does our data sit, and what is the retention policy? Ask before pricing, not after. It determines whether the deal can close at all.

A vendor that answers all five in writing is not necessarily the right one, but a vendor that will answer none of them has told you something useful for free.

Where this case is weakest

There are four situations where the arithmetic does not work, and recognising them early is worth more than an ROI case that gets rejected in month three.

Funnel chart of the mid-market PE pipeline, narrowing from 600 targets and 80-100 detailed reviews down to a single closed deal.

Volume is the first test. A funnel this shape has the document throughput to justify the exercise; a firm doing a handful of deals a year does not.

Low volume. Below a few hundred documents a quarter through a single repeating workflow, the build and configuration layer dominates and never amortises. A four-person family office reviewing a dozen managers a year should use an assistant well rather than buy infrastructure, and any vendor telling them otherwise is selling rather than advising.

No clean baseline. If the work is currently done by three different people in three different ways, and nobody has ever timed it, you cannot prove a saving because you have nothing to prove it against. Fix the process description first. This often turns out to be the more valuable project.

A workflow that changes every quarter. Configuration cost recurs each time the schema or the source format moves. A stable workflow amortises that cost over thousands of documents. An unstable one pays it again and again, and the payback period stretches past the point where anyone is still tracking it.

A saving nobody will bank. If the freed hours will not be converted into coverage, a deferred hire or a retired spend line, the return is genuine for the people doing the work and invisible to the firm. That is a legitimate reason to buy, and it is worth being clear that it is the reason. The earlier V7 guide to securing return on an AI investment takes the wider enterprise view of the same question; this piece is the private markets version of it.

Where the case does work, it usually works decisively rather than marginally, which is the other thing the percentage-saving articles get wrong. A team that goes from reading a fifth of a data room to reading all of it has not improved a metric by fifteen percent. It has changed what the team is able to know before it commits capital, which is the argument that AI-assisted diligence has always rested on. Price it honestly, and it does not need to be oversold.

For reference on where the lines sit, V7 Go's own pricing is published. V7 Go is not a CRM, a data room, a portfolio monitoring platform or a fund accounting system, and a costing exercise that compares it to one of those is comparing the wrong things. It is the workflow layer that sits between the documents and whichever of those systems your firm already runs.

What is the 30% rule for AI?

There is no formal 30% rule. The phrase gets used for at least three different claims: that AI should cut a process cost by about 30%, that roughly 30% of a knowledge role is automatable, or that 30% of an AI budget should go to change management rather than technology. None is a standard, and none should appear in a business case. The useful version of the idea is narrower. A saving under roughly a third is hard to detect through normal variance in how long work takes, so a case built on a 10% improvement will be difficult to prove after the fact even if it is real. Build the case on a change large enough to see, in a workflow measured for a full cycle before and after, and use your own figures rather than any rule of thumb.

+

How much does AI document processing actually cost per document?

It depends far more on how many values you need from each document and how many of them a person has to check than on the document itself. The machine portion is usually small. In our own benchmark on private-equity disclosure documents, a question against a structured context layer cost about 13 cents and stayed flat as the corpus grew from 10 documents to 1,000. The retrieval-only configuration cost about 34 cents and kept climbing. The human review portion is typically the larger line. Work out your own figure as inference and platform spend divided by fields produced, then add the share of fields a person reviews multiplied by review time at your fully loaded hourly cost.

+

Is AI going to get cheaper or more expensive?

Per unit of model capability it has got steadily cheaper, and the published rate cards from the major providers span more than an order of magnitude between their smallest and largest models. Total spend often rises anyway, because falling unit prices invite more usage and more ambitious workflows. The more useful question for a budget is not where model prices go but whether your own cost per question stays flat as your document estate grows. Systems that do their reading at ingestion hold that line. Systems that search the whole corpus at query time do more work as the index grows, so their cost per question rises even while the price per token falls.

+

What is a realistic payback period?

For a single repeating document workflow at reasonable volume, most teams should expect the operating cost covered within one to two quarterly cycles after go-live. The build and configuration layer amortises over the first year. That assumes the workflow is stable, the baseline was measured rather than estimated, and somebody is tracking it. If a projection shows payback in weeks, check what has been left out of the cost side. If it shows payback beyond eighteen months, the volume is probably too low or the workflow too unstable, and it is better to say so before signing than to discover it in month nine.

+

Does the human review cost ever go away?

Separate the fixed layer from the metered one and cap the metered one. Platform and configuration costs are largely fixed and can be committed annually. Inference scales with volume. Ask the vendor for spend controls at workspace level, set a period budget, then review actual usage monthly for the first two quarters until you know the shape of your own demand. Model the budget on your busiest quarter rather than your average, because in private markets the busy quarter is the one where the tool matters most and the worst time to discover a cap. Ask what happens when a budget is reached: whether the workflow queues, degrades or stops.

+

How do we budget when document volume is unpredictable?

Go is more accurate and robust than calling a model provider directly. By breaking down complex tasks into reasoning steps with Index Knowledge, Go enables LLMs to query your data more accurately than an out of the box API call. Combining this with conditional logic, which can route high sensitivity data to a human review, Go builds robustness into your AI powered workflows.

+

Casimir is a seasoned tech journalist and content creator specializing in AI implementation and new technologies. His expertise lies in LLM orchestration, chatbots, generative AI applications, and computer vision.

Precision AI for Institutional Workflows

Build once.
Deploy across teams.
Improve over time.

Precision AI for Institutional Workflows

Build once.
Deploy across teams.
Improve over time.

Precision AI for Institutional Workflows

Build once.
Deploy across teams.
Improve over time.