/

AI implementation

The AI Data Orchestration Layer Private Markets Firms Are Missing

The AI Data Orchestration Layer Private Markets Firms Are Missing

13 min read

Summarize

V7 Go

A smarter way to manage due diligence and underwriting

A mid-market fund has a CRM, a virtual data room, a portfolio monitoring platform and a fund accounting system. Four systems, four renewals, four implementations that took a quarter each. Every one of them is a place to put structured data. Not one of them produces any.

The thing that turns two hundred PDFs into fields somebody can sort is an associate. That associate is the AI data orchestration layer, and for most of the history of private markets there was no other candidate for the job. Documents arrive unstructured. Systems expect structure. A person closes the gap by reading and retyping. The layer has always existed. It was just staffed rather than built.

Worth being precise about the term, because it comes from enterprise IT and arrives in private markets carrying some baggage. An orchestration layer is the part of a stack that turns raw inputs into structured, typed, repeatable data and delivers it where it belongs. Enterprise IT named this years ago. Private markets has had the problem for longer, never named it, and consequently keeps solving it by buying another destination.

This article covers where the gap actually sits, why connecting more systems together does not close it, and the four properties a layer needs before a deal team should let its output anywhere near a system of record.

In this article:

  • Why every system in the private markets stack is a destination and none is a source.

  • The difference between connectivity and data, and why protocols do not close it.

  • The four properties that make an orchestration layer trustworthy enough to feed a ledger.

  • Five questions that separate a real layer from another place to put things.

Private Markets

Turn complex deal documents into faster investment decisions.

Private Markets

Turn complex deal documents into faster investment decisions.

Every system in your stack assumes the data has already arrived

The private markets stack is well supplied with destinations and has no source. That is the whole problem, and it is easiest to see by following one document.

A portfolio company sends its quarterly report. Somewhere in it is an EBITDA figure, a headcount number, a net debt position and a revenue line that may or may not be defined the way last quarter's was. Those values need to reach the portfolio monitoring platform. Ask how they get there and the answer is always the same: an analyst opens the PDF, finds them, and types them in. The monitoring platform did not read the report. It received keystrokes.

The same thing happens at the other end of the funnel. A teaser arrives, somebody skims it, and the target details reach the CRM because a person decided which fields mattered and filled them in. We have written before about moving deal documents into a CRM, and the pattern holds across every one of these handoffs. The system on the receiving end is passive. It waits. Something else has to turn a document into a record, and in most firms that something is a person with a second monitor and a lot of patience.

Multiply that by the document set an active fund actually handles. Teasers and pitch decks at the top of the funnel. Data room contents during diligence. Management accounts, audited financials, quality of earnings reports. GP and fund reports, capital account statements, DDQs, limited partnership agreements and side letters on the investor side. Different desks call them different things, and the format varies with whoever produced them, which is precisely why no system upstream can be relied on to hand over anything tidy.

Funnel chart of the mid-market PE pipeline, narrowing from 600 targets and 80-100 detailed reviews down to a single closed deal.

The volume compounds it. Sutton Place Strategies put the mid-market pipeline at roughly 600 targets reviewed for every deal that closes, with 80 to 100 of those going to detailed review. Every one of those detailed reviews is a document package somebody has to read. Teams handle it by reading the ten or twenty percent they have time for and calling it coverage.

So the fix is a better data warehouse? Wrong again. A warehouse is another destination, and a very good one, but it has the same entry requirement as everything else: the data has to already be structured when it shows up. Buying storage for data you cannot produce is how firms end up with four systems, an integration budget, and an analyst still typing.

The numbers on AI adoption in private equity make more sense once you see the stack this way. FTI Consulting's 2026 Private Equity AI Radar found 95% of funds reporting that AI is meeting expectations, while only 7% of portfolio companies have reached enterprise-scale deployment. Conviction is nearly universal. Production is not.

Chart contrasting 95% of PE funds saying AI meets expectations with only 7% of portfolio companies at enterprise-scale deployment.

That gap is usually read as a change management story, or a talent story, or a governance story. It is mostly a plumbing story. Pilots work because a human curates the inputs. Production does not, because nobody is curating the inputs at production volume, and the layer that would do it does not exist. Allvue's 2026 GP Outlook Survey puts a hard number on the same thing from the other direction: 64% of firms still run critical workflows in Excel despite having bought purpose-built systems designed to replace it. The systems arrived. The data did not.

There is a second failure underneath the first, and it survives even when somebody does the typing. Private markets does not enforce standardisation the way public markets do. Fund structures differ firm to firm and sometimes fund to fund. Portfolio company financials arrive in whatever the company's own accounting system produces, with a different chart of accounts and a different fiscal calendar. EBITDA, NAV and distributions can each mean something subtly different depending on which fund, which company and which system produced the number. The same survey found 92% of firms describing their data as moderately organised at best, and 47% naming systems integration among their largest technology problems, which is what happens when you try to connect systems whose definitions were never reconciled in the first place.

Retyping a value into a system does not decide what the value means. That decision gets made somewhere, by someone, usually in a hurry, and it never gets written down. The tools a deal team runs day to day assume that work is already finished.

Connectivity is not the same as data

The industry's answer to fragmentation has been to connect things, and connecting things is genuinely useful. It is also a different problem from the one described above.

First came integrations. Then the Model Context Protocol, which gives AI systems a consistent way to reach enterprise data instead of demanding a bespoke connector for every tool. Both are worth having. V7 Go runs as an MCP server for exactly this reason, so a team can drive a workflow from Claude if that is where they already work.

Allvue framed the limit of this better than anyone else in the category, in a piece on why private equity's data problem has to be solved first: MCP will hand you the wire, but it will not hand you a data strategy. The point is worth extending rather than restating. A protocol reaches a system. It does not reach a PDF sitting in a data room folder, and it has nothing to say about which of two systems is right when they disagree about what NAV means for the same fund.

Connectivity moves data that is already structured. Everything in the paragraph above about quarterly reports is still unstructured on the other side of the wire.

This is also where the assistant comparison usually gets made badly. A general assistant pointed at a data room is a real improvement over reading it manually, and teams that already work this way are the ones best positioned to go further. The limit is not capability. It is that an assistant re-solves the problem every time you ask, so the twentieth document takes a different route through the reasoning than the first, and the output arrives in a different shape. Fine for exploration. Disqualifying for anything that feeds a ledger.

What has to be true before you push a number into your system of record

Four properties. Without them the orchestration layer is a liability, because it produces values that look authoritative and carry no evidence.

Everyone in this category agrees the data foundation matters. Allvue says it, Arcesium's Bryan Dougherty says it in a piece on why private markets AI ambitions stall, Acuity's five-capability framework says it. What none of them specifies is what ready means as a technical property. Here is the specification.

Architecture diagram of a V7 Go agent turning document inputs into structured outputs through a five-step workflow.

The path is fixed, not planned

The workflow is a defined sequence of steps. Each step carries its own instruction, its own tool and its own output type. Nothing gets re-planned at runtime, which means the five hundredth run takes the same route as the first. A step told to check a figure against the firm's own criteria checks it every time, not on the runs where it happens to occur to the model.

Each step uses the tool the work requires

A named model where the answer is judgment, such as reading a covenant and deciding whether it is unusual. Python where the answer is arithmetic, a date format or a cross-foot that either ties or does not. A specific integration where the value lives in another system and the correct move is to go and get it rather than infer it. Deciding this field by field is the difference between a workflow and a very long prompt.

Outputs are typed

Fields, tables, selects, nested structures, declared before the run rather than discovered after it. This is the mechanical reason output can land in the same template every time. Prose that has to be re-read and re-keyed is not a layer. It is a draft.

Every value opens its source

A citation that lands on the exact sentence in the PDF, or the exact cell in the spreadsheet. The point is not reassurance. The point is that a disputed number can be resolved in four seconds by someone who was not on the deal, which is what makes the output usable by an investment committee rather than just by the analyst who ran it. Review gates sit in the same place, deliberately: the layer compresses the reading and the retyping, and leaves the judgment where it was.

Taken separately, none of these four is remarkable. Together they change what the output is for. A value that came out of a defined step, in a declared type, with a citation to the sentence it was read from and a reviewer's name against it, is not a suggestion. It is a record with provenance, and it can go into a ledger, a model or an investment committee memo without anyone re-verifying it by hand. A value produced by a system missing any one of the four cannot, which is why so many pilots produce impressive demonstrations and nothing a firm will actually rely on.

V7 Go is not a CRM, not a data room, not a portfolio monitoring platform and not a fund accounting system. It is the layer that feeds them, which is a duller claim and a more useful one.

The layer should get better at your firm over time

A pipe moves things. A context layer remembers them, and the difference shows up about six months in.

When a deal team turns over, the firm loses the reasoning behind decisions it already made. Not the files, which sit in the data room forever. The reasoning: why the sector was passed on, which covenant pattern triggered the concern, what the comparable looked like eighteen months ago. New analysts rebuild that from scratch, badly, and the firm pays for the same thinking twice.

A layer that runs against a persistent context layer rather than against one document at a time does not have that problem. Work arrives with the firm's own history attached. When a new information pack comes in, the screening workflow already knows the firm passed on three similar businesses last year and why. We have written separately about how a context graph gets built in private equity and venture capital, so the mechanics are not worth repeating here.

Split illustration on a black background with a white three-dimensional grid on one side and a node-and-edge graph with an orange central node on the other, contrasting a knowledge graph with a vector database.

The distinction the orchestration layer turns on. An index stores passages and finds the ones that look like your question. A graph stores objects and the relationships between them, so a question that spans four documents is a traversal rather than a search.

The ontology is the part you configure

Most of what gets called a knowledge layer is entity extraction with no opinion: run a model over the corpus, collect every proper noun it finds, and hope the result is navigable. It usually is not. What you get is a web of near-duplicate nodes where the same fund appears four times under three spellings, and the only way back out is to search it, which is the problem you started with.

A Context Graph is declared before it is built. You choose the object types the firm actually asks about, which for a private markets firm is some combination of funds, managers, portfolio companies, securities, commitments, entities and people. You choose the relationships that hold between them: this company is held by that fund, which is managed by that GP, in which this entity holds a commitment. And you choose the metrics that get tracked against each object, so net asset value, EBITDA or headcount are fields with a history rather than numbers in a paragraph.

V7 Go onboarding screen asking a new user to select their industry (Private Markets or Venture Capital) to build a context graph.

The ontology choice, made at setup. Private markets and venture capital ask different questions of the same document types, so they get different object models rather than a shared one with unused fields.

Three properties follow from declaring the model first, and each one answers a question a spreadsheet cannot.

Entity resolution happens once, at ingestion. Nine entities whose names begin with the same two words get separated when the whole document is in view and there is time to be careful, rather than at query time under a deadline. By the time an analyst asks the question, the ambiguity is already gone.

Relationships are traversable, so a set can be enumerated. Ask how many portfolio companies a fund holds exclusively and the answer comes from following edges, which either returns the complete set or visibly does not. Ask the same question of a search index and you get a number that is usually close, occasionally wrong, and carries no signal telling the two cases apart.

Change is a first-class object rather than a diff. Because the same metric is attached to the same resolved entity every quarter, a movement in net asset value, a covenant heading toward its limit, or a manager quietly drifting from the strategy that was underwritten are all queryable rather than noticeable. That is the difference between monitoring and remembering to check.

Mind map showing connections between SharePoint, Partners, Funds, Drive, Assets, Files

The same layer viewed from the source side. Documents arrive from the places they already arrive from; what changes is that a file resolves to the fund, the manager and the commitment it concerns rather than to a folder.

This is also why the layer is a layer rather than a product category. The object model differs by firm, the document set differs by strategy, and the questions differ by desk. The mechanism underneath does not.

What matters for the orchestration argument is narrower. A layer without memory converts documents at a constant rate forever. A layer with memory converts them against an accumulating asset, which means the same workflow produces a better answer in year two than it did in month one, without anyone rewriting it. The Everest Group architecture that Acuity summarises in its five capabilities for AI-ready private markets firms makes the same structural point: each layer depends on the integrity of the one beneath it, and the data and integration layer is at the bottom.

Which is the part everyone skips.

AI Implementation

Start with one workflow, then roll it out across the firm.

AI Implementation

Start with one workflow, then roll it out across the firm.

How to tell whether you are buying a layer or another destination

Five questions, all of them answerable in a demo, none of them answerable with a slide.

  1. Does the same input produce the same output shape on the hundredth run, and can you watch that happen rather than take it on trust?

  2. Can you see which step produced a given value, and open the source document at the exact place the value came from?

  3. Is the output typed, or is it prose that somebody on your team will re-key into the system that needed it?

  4. Where is the review gate, who owns it, and what happens to the run when a reviewer rejects a value?

  5. Who builds the first workflow, and are they still involved in month six?

The fifth one decides more implementations than the first four. Configuration against a firm's own documents is where these projects succeed or quietly stall, which is why V7 Go treats implementation as part of what you buy rather than a services line item: scope the use case, build a proof of concept before anyone commits, configure against your documents, connect the sources and the outputs, stay involved after go-live.

Two more worth raising before anyone asks, because a quarter of buyers ask and waiting reads as evasive. SOC 2 Type II, and where your data sits and who can reach it. Any vendor that needs a follow-up call to answer either one has told you something.

One caution on sequencing. None of this argues for pausing while a data programme runs for eighteen months. The layer is what produces governed data in the first place, so waiting for clean data before building it inverts the dependency. Start with one workflow where the input volume is high and the output destination is unambiguous, such as quarterly report ingestion into monitoring, or DDQ completion. Get that one running end to end, with the citations and the review gate in place, and use what it teaches you about your own definitional conflicts before extending to the next.

The firms that get value out of AI in private markets over the next few years will not be the ones with the largest model budget. FTI's 95% number says conviction is already universal and is not the constraint. They will be the firms that stopped treating document-to-data as a staffing problem, named the layer, and bought one. Everything else in the stack has been waiting on it.

What is an AI data orchestration layer in private markets?

It is the part of the technology stack that turns unstructured documents into structured, typed data and delivers that data into the systems that need it. In private markets the inputs are things like data room contents, quarterly portfolio company reports, capital account statements, DDQs and side letters. The outputs are the fields a CRM, portfolio monitoring platform or fund accounting system expects to receive. The term comes from enterprise IT, where orchestration layers coordinate models, data pipelines and applications. Private markets firms have had the same problem for longer without a name for it, because the layer has traditionally been staffed rather than built. An analyst reading a PDF and typing values into a system is performing orchestration. Naming it as a layer makes it possible to ask what a built version would have to do differently, and to evaluate vendors against that rather than against a feature list.

+

Is a data orchestration layer the same thing as a data warehouse or a CRM?

No, and the distinction matters when you are deciding what to buy next. A data warehouse, a CRM, a virtual data room and a portfolio monitoring platform are all destinations. They store, organise and present data that is already structured when it arrives. None of them produces structure from a document. An orchestration layer sits upstream of all of them and does the production step: reading the source material, extracting values, typing them, attaching provenance and routing them onward. A firm that buys another destination when its actual constraint is production ends up with more systems, a larger integration budget and the same analyst doing the same retyping. The test is simple. Ask what the system requires as input. If the answer is structured data, it is a destination, not a layer.

+

Does the Model Context Protocol already solve this problem?

MCP solves a real problem, but a different one. It gives AI systems a consistent, governed way to reach enterprise data sources instead of requiring a bespoke integration for every tool. That is connectivity, and it is worth having. What it does not do is reach data that is not in a system yet. A large share of private markets data lives in PDFs, spreadsheets and email attachments that no protocol can query, because there is nothing on the other end to answer. MCP also has nothing to say about definitional conflicts. If two systems disagree about what NAV means for the same fund, a connector faithfully delivers both answers. Extraction, normalisation and reconciliation still have to happen somewhere, and that somewhere is the orchestration layer. The two are complementary: connectivity moves structured data, and the layer produces it.

+

How is this different from using a general AI assistant on a data room?

An assistant pointed at a data room is a genuine improvement over reading it manually, and teams already working that way tend to be the best prepared to go further. The limit is consistency rather than capability. An assistant re-solves the problem each time it is asked, so the twentieth document can take a different route through the reasoning than the first and the answer can arrive in a different shape. That is fine when you are exploring and disqualifying when the output feeds a ledger, a memo template or an investor report. An orchestration layer runs a defined sequence of steps in a fixed order, with declared output types, so the shape is identical every run and the variance sits only in the source documents. The comparison is assistant versus infrastructure, not one vendor versus another.

+

Where does human review sit in an orchestration layer?

Four properties, and they are testable in a demo. First, the path is fixed rather than planned at runtime, so the hundredth run follows the same sequence of steps as the first. Second, each step uses the tool the work requires: a model where the answer is judgment, code where the answer is arithmetic or a date format, an integration where the value lives in another system. Third, outputs are typed, with fields, tables and nested structures declared before the run, so results land in the same template every time instead of arriving as prose somebody re-keys. Fourth, every value opens its source, landing on the exact sentence or cell it came from, so a disputed figure can be resolved by someone who was not on the deal. Without all four, the layer produces values that look authoritative and carry no evidence.

+

What has to be true before a firm pushes AI output into a system of record?

Go is more accurate and robust than calling a model provider directly. By breaking down complex tasks into reasoning steps with Index Knowledge, Go enables LLMs to query your data more accurately than an out of the box API call. Combining this with conditional logic, which can route high sensitivity data to a human review, Go builds robustness into your AI powered workflows.

+

Casimir is a seasoned tech journalist and content creator specializing in AI implementation and new technologies. His expertise lies in LLM orchestration, chatbots, generative AI applications, and computer vision.

Precision AI for Institutional Workflows

Build once.
Deploy across teams.
Improve over time.

Precision AI for Institutional Workflows

Build once.
Deploy across teams.
Improve over time.

Precision AI for Institutional Workflows

Build once.
Deploy across teams.
Improve over time.