FinTech · AI-assisted private equity · one case, one stealable principle

For expert skeptics, provenance is the product.

Private-equity analysts get paid to doubt confident numbers — a score they can’t cross-examine is a smooth pitch with no references, and experts walk past smooth pitches. So this product refused to launch as a naked score: from day one, every claim carried its source, and thin evidence drew an honest “I’m not sure.” Screening ran 60% faster — because the machine came with references.

60% fasterdeal screening · pre/post rollout
3sources sat beside every score, not behind it
Leadanalysts now open the memo with it

FinTech · AI-Assisted Private Equity Investing · 2025

Role
Lead Product Designer
Team
Cross-functional — engineers, data scientists, PM
Surface
LLM over deal docs for PE investing
Screening
60% faster measured pre- vs post-rollout
Status
Shipped · under NDA
Owned
Product definition, the interaction design, and the abstention + citation UX — built before the first release, so every claim traced to its source on day one.
Shared
Eval design and threshold tuning, with the data-science team.
Not mine
The model’s accuracy — the ML team’s result to defend — and the retrieval backend.
The principle this case proves

A score is a claim. Experts don’t buy claims — they buy claims with references.

The house belief: get the score accurate enough and the analysts will use it. The score was accurate. The analysts audited it by hand anyway — because a number with no sources isn’t evidence, it’s homework.

One principle · proven in the room where doubt is the job

The stakes: an accurate score that added work instead of removing it.

Doubt is the job here, and these analysts are very good at it. A score they couldn’t cross-examine wasn’t a tool — it was a liability with a UI, and they treated it like one. It didn’t help that the number came out of a large language model reading messy deal documents: the same technology that, pushed too far, will invent a confident answer it can’t back up.

The LLM read board packs, statements and interview notes like a junior analyst, and scored deal risk accurately. But on screen, a ninety-percent-confident score looked identical to a fifty-percent guess, and nothing said whether a claim came from a real document or was quietly made up. So analysts built parallel scorecards by hand and demoted the AI to sanity check.

The accuracy is the ML team’s result to defend. Mine is everything that made an accurate score acted on: the evidence beside the number, the abstention state, the override logged instead of lost. The job was to make a correct score survive cross-examination by the most skeptical reader in the building — and to hold the launch until it could.

The test: if the principle holds, belief should move without touching accuracy.

An analyst won’t stake a recommendation on a number they can’t defend to the investment committee — so an unexplainable score, however accurate, becomes one more thing to audit by hand. The model added work instead of removing it.

The mechanism: references beside the number, and an honest “I’m not sure.”

WHAT SHIPPED BEFORE RELEASE Two conditions, and the failure each one bought off. CONDITION 01 Every generated number carries its source 3 sources behind a score, inline — not in a tooltip prevents: a confident figure nobody can check on deal day CONDITION 02 Thin evidence draws an abstention a visible “I’m not sure about this one” prevents: a fluent hallucination entering a capital decision Ship — only once both are true. Until then the model was accurate and unusable. cost: weeks of delay The trade only pays when the reader is an expert paid to doubt the answer. For a low-stakes call nobody audits, holding the launch would have been wrong.
The two conditions I would not ship without — and the failure each one bought off

The product served one analysis in two interpretation modes — because analysts read differently. Some want the full written argument to sit with and annotate; others want to interrogate. So the system generated a descriptive mode — a detailed written analysis of the deal, the scores explained in prose — and a conversational mode for questioning that same analysis. Neither was a free-form chatbot: underneath both runs an agentic architecture with strict operational boundaries — specialised components playing as an orchestra, not a single improvising soloist — so whichever mode an analyst chose, the answers stayed structured, auditable, evidence-based.

ONE ANALYSIS, TWO READERS The design refused to pick a temperament for them. One analysis scored, cited, abstaining where thin A written report the scores argued in prose — for the reader who wants the case made once, in order A conversation interrogate any line — for the reader who arrives with one question and no patience Pipeline dashboard every live deal in one view — the surface both readers return to Same evidence underneath both. Citations travel with the claim, whichever way it is read.
One analysis, two readings — over the dashboard both readers return to

Above the single deal sat the scoring dashboard. The whole pipeline scored side-by-side — the multi-deal view where screening actually starts — and any deal drills down into the per-signal breakdown, the written analysis, and the conversation. One analysis, three surfaces, no forced workflow.

So I built an “explain this score” surface that grounds the model in its evidence. An analyst could pull any rating apart into the signals behind it, challenge the weighting, and watch the score answer back — and every signal came stapled to the exact document it was retrieved from, so no claim floated free of a source. Making the model survive that, before release rather than after, is the argument the whole project hung on — and the reason it took the weeks it did.

Three things were non-negotiable: a cited source behind every generated number; a visible “I’m not sure about this one” where the model abstains instead of bluffing — an LLM’s worst failure is a fluent wrong claim; and a logged override when the analyst disagreed. The override drew the longest fight: logging disagreement felt exposing. It was actually the moat. Every correction became structured data the system learns from, so the platform comes to reflect how this firm thinks. A rented model any competitor can rent too; a record of the firm’s own judgment, nobody can.

The abstention threshold was a design argument, not just a statistical one. Every claim carried its own confidence; run low, and it rendered as “I’m not sure about this one” — a signal promoted from the eval dashboard to a state the analyst can act on. The pitch is the whole difficulty: too eager and analysts learn to skim past the crying wolf; too shy and one fluent wrong claim burns the trust bank. The threshold fails toward humility — the cheaper mistake.

Trust also had to be cheap to enter. Deal flow lives in email — decks, spreadsheets, forwarded intros — and no analyst re-keys email into a system. So the platform meets them inside Outlook: tag a message, drag an attachment, get “Received, thanks” while agents stitch the pile into a deal profile — and the prequalification score surfaces back in the inbox, where prioritisation actually happens. The fuzziest judgment got the same discipline: management quality — pure gut feel, traditionally — became a scorecard of leadership tenure, employee sentiment and filed financials on one timeline, every input traceable.

Plate 01 · Deal screening — reconstruction · all names and figures synthetic Redrawn from memory at production fidelity — the client’s pixels stay theirs. This is the surface the case describes: a verb instead of a naked score, reasons with the evidence beside them, and the honest “I’m not sure.”

Pull a score apart and challenge it below — the signals, their 3 cited sources, their weights, and the score answering back:

Reconstruction — anonymised, rebuilt from memory for illustration · client under NDA
Deal Atlas-7 Risk 62 Medium confidence
Q3 board pack · p.12

Top customer is 38% of recurring revenue — concentration pushes the score up.

Weight 3
Audited statements · FY24

Margins hold steady across three years; working capital reads clean.

Weight 2
Founder interview · note 07

Churn figure couldn’t be matched against the raw export.

“I’m not sure about this one” — low confidence, kept visible

Weight 1

Disagreement logged — training signal

Keyboard-friendly · nothing you click here leaves the page
DEAL SCREENING — EVALUATION PATH A document in. A defensible score out — or nothing. Inbox tag a mail, drop the deck Retrieve claims matched to sources Score risk read from the docs GATE is every claim grounded? Abstain — “not sure about this one” thin evidence never becomes a confident number Score, with its sources attached 3 sources behind every number, inline Analyst reads it two modes: written report · conversation Override, logged structured data, not a comment box the correction trains the next read The 60% was not the model getting faster — it was analysts no longer re-verifying by hand.
The path a document takes — and the branch where the system declines to answer
FinTech deal risk: explain before the verdict A wide left-to-right diagram on cream paper. Three source documents on the left — Financials, Market data and Diligence notes — each traced by a thin line into a central circular deal-risk score reading 62. From the score a line leads to a panel on the right titled “Explain this score”, listing three weighted factors with bars: revenue concentration, market volatility and management tenure. Explain before the verdict FINTECH · DEAL RISK Financials Market data Diligence notes DEAL 62 RISK SCORE Explain this score Revenue concentration +34% Market volatility +19% Management tenure -9% Sources traced Verdict scored Reasons shown
The argument arrives before the verdict — every score opens onto its three sources

What holding the principle cost — and the objection it had to survive.

The strongest objection, kept in

“Ship the score now, add citations later — provenance is polish.” Reasonable, budget-shaped, and wrong for this room: with expert skeptics you get exactly one first impression, and a bluffing tool never earns a second audit. The product-owner call was to perfect it before the release — provenance went in as a launch condition, not a fast-follow.

Building it first wasn’t free. The accurate LLM sat finished while I built the surface that could defend it — the retrieval that tied each claim to a source, the abstention state, the override — weeks the team could feel, against a model that already worked. Waiting was the product owner’s call as much as mine; what I own is the argument for it, because shipping a correct score nobody acted on would have been the more expensive mistake.

Accuracy gets you a correct number. It doesn’t get you a decision. The gap between the two is the whole job — and it’s made of trust, not math.

Falsifiable evidence — 60% faster, and the skeptics now open with the score.

Interrogation turned into reliance · every number with its window

Screening ran 60% faster, and analysts went from ignoring the model to leading their deal memos with it. The LLM didn’t get more accurate — they could finally trace every claim to its cited source and see where its confidence came from. What the NDA keeps off this page — the sample, the eval design, the baseline, who ran the count — I walk through on a call, artifacts open.

Screening speed
60% faster deal screening · pre/post rollout
Adoption
From ignoring the model to leading the deal memo with it · lead analysts now open with the score
Explainability
3 cited sources sat beside every score, not behind it · every claim grounded, low-confidence claims flagged as abstention
Override loop
Logged disagreement became a training signal the model learned from

Where this wouldn’t transfer

Those weeks bought explainability, and that trade only pays when the user is an expert with a reputation riding on the answer. For a low-stakes, high-volume decision nobody audits, spending them would have been the wrong call.

Where the principle breaks.

Provenance-first assumes the reader will actually open the references — true of analysts whose names go on the memo, untrue of users grazing a feed. It also taxes latency and retrieval cost on every claim; in a product where a wrong answer is cheap, that tax buys nothing. And citations cannot save a model that’s guessing: a source beside a wrong number just makes the error easier to find — which is the point, and which some rooms don’t want. If your users don’t doubt for a living, this case argues against copying it.

Want the full version of this work?

Details are confidential; I’ll walk through the artifacts — the service blueprint included — and numbers on a call under mutual NDA.

Send me the role

Design Patterns Used in This Case

This project produced two patterns I now reach for whenever AI meets domain experts:

  • ML Explainability Patterns: How to ground each generated claim in the cited source it was retrieved from — and surface abstention when the model isn’t sure — so a non-technical expert can read the why, not just the verdict.
  • Human-in-Loop Patterns: The override-and-explain loop, the visible weighting on each risk factor, and disagreement handled with respect — the moves that keep the expert in command.

I write about this in my newsletter ↗ (opens in a new tab)

That was trust built for experts betting other people’s money. Next: take the model away entirely — the same trust problem, between two hundred people.