- Role
- Product & Design Lead
- Team
- Alongside the engineers and data scientists who built the model
- Surface
- Technical-DD platform · VC + PE
- Cycle
- 3 wks → 4 days
- Status
- Shipped · under NDA
Remove the wrong friction and you get speed. Add the right friction and you get a verdict someone will sign.
A partner has four days to decide whether a startup’s technology is real, and millions riding on the answer. An LLM could read the whole codebase in an afternoon — chunked and retrieved, since no context window holds a repo whole. The diligence still took three weeks. The bottleneck was never the reading.
A diagnosis · symptom → wrong suspects → the real culprit
The stakes: a fast model, and a verdict nobody would sign.
VC and PE partners stake millions on a startup’s technology on claims they will never personally check. They can’t read the codebase. They can’t reproduce the benchmark. So the deal stalls in a three-week diligence cycle while analysts assemble something a partner is willing to put their name to.
An LLM handled the hard technical work in an afternoon: read the unstructured repos, read the design docs, and analyse each finding across sixteen dimensions. The extraction was never the problem.
The problem was that a partner would read a clean machine verdict and still pick up the phone to re-check it by hand. The LLM produced an answer. It did not produce something anyone would stand on. That is what kept the cycle at three weeks.
The test: if conviction is the bottleneck, more evidence-work should beat more accuracy.
The obvious fixes all pointed at the engine. Make the model read more. Train it on more deals. Surface a single, confident risk score per company so the partner has one number to act on. Every one of those sharpens a tool the partner had already decided not to use.
Because a partner staking millions does not distrust a number for being imprecise. They distrust it for being unaccountable. A higher-confidence score the partner still can’t trace back to evidence is just a more persuasive thing to be wrong about — and the first time it’s wrong on a real deal, the whole platform goes back in a drawer.
The mechanism: no source, no verdict — sign-off stays locked until the evidence is opened.
The design was tested where it would live: usability sessions with the investment analysts who’d stake their names on its output. The verdict was right — which is exactly why it was dangerous. On deal day, a correct verdict the partner can’t audit is indistinguishable from a confident hallucination. The missing piece was never accuracy. It was the path from verdict down to source: the repo scan, the infra audit, the interview note — every claim traceable to the document it came from.
Every signal carries its own confidence and its own cited source, and every score maps to a move a partner can make — proceed, attach a condition, send an analyst back. Citations are table stakes in 2026; the work was making them cheap enough to check that partners actually checked. One decision organised everything: no verdict is reviewable until its evidence is. Sign-off stays locked until every signal has been opened.
Try it — open each signal to see the evidence, then sign:
Dossier · Company K
Technical verdict: proceed with conditions
Core modules read clean; test coverage thins around the billing paths.
source: repo scan · static analysis pass 02
A single job queue becomes the chokepoint well before the growth plan needs it to hold.
source: design-doc review · infra audit 01
Merge cadence is steady, but two contributors carry most of the commit history.
source: commit history · release cadence
Claims in the deck match the demo; one benchmark could not be reproduced.
source: founder interviews · reference calls
Analysed across 16 dimensions · 12 further findings stayed below the confidence bar and never reached this report
review every signal first
That sign-off gate looks like a small interaction. It was the whole argument. It turns the LLM from an oracle you either believe or ignore into an evidence file a partner walks through, each verdict grounded in its cited source — which is the difference between a tool that sits in a drawer and one that compresses three weeks into four days. The analyst’s overrides feed back to the model, so the trust runs both ways.
What it cost — and the objection the principle had to survive.
“Partners are busy — every extra click loses a user.” True in consumer software, and exactly wrong here: a partner who breezes to a verdict can’t defend it to their investment committee, so the breezy verdict gets re-done by hand. The friction wasn’t a tax on the decision. It was the decision, distributed into openable pieces.
The honest line: the gate adds friction. A partner who just wants the number has to open every signal before they can sign — I chose to slow the one moment everyone wanted to speed up. The first reaction was exactly what you’d expect: why am I clicking through evidence when the machine already decided?
The friction is the product. A verdict a partner clicks through is a verdict they’ll defend in the investment committee; a one-tap score is one they’ll quietly re-check by hand, which is how you end up back at three weeks. I priced the extra clicks in on purpose. The thing that actually compressed the cycle wasn’t a faster model — it was a partner who, for the first time, didn’t feel the need to verify it twice.
Falsifiable evidence — three weeks to four days, measured across engagements.
What changed, and for whom
And the 3 wks → 4 days belongs to the whole platform — model, retrieval, pipeline, surface; the part I claim is the design that made partners willing to stand on it for live deals.
Where this wouldn’t transfer
Deliberate friction survives only when the person clicking is personally accountable for the verdict. Give the same evidence gate to a user with no downside and they’ll route around it — the friction has to buy them something they actually need.
Where the principle breaks.
Deliberate friction earns its keep only when the signer bears the downside — a partner with capital at risk, a clinician, an underwriter. Put the same locks in front of a user with nothing at stake and you’ve built a nag screen; they’ll satisfice, click through, and resent you. It also assumes the evidence is worth opening: gating sign-off on weak sources teaches people to stop reading. Friction is a precision instrument — this case argues for placing it, not sprinkling it.
The framework is the part worth walking through.
Details are confidential; I’ll walk through the decisions, artifacts — the service blueprint included — and numbers on a call under mutual NDA.
Send me the roleDesign Patterns Demonstrated
- Confidence Score Patterns: Per-signal and per-document confidence, each one mapping to a next step a partner can take.
- ML Explainability Patterns: Provenance on every claim — a clean drill from the summary down to the source signal and code-quality measure behind it.
- Human-in-Loop Patterns: Analyst overrides fed back to the model; partner sign-off as real workflow, not a rubber stamp.
I write about this in my newsletter ↗ (opens in a new tab)
The gate priced the doubt. Next: back to where the whole argument starts paying — a business model rewritten from inside the design seat.