Skip to the book

The Trust Layer

Fifteen years designing the half-second where a person decides to believe a machine.

Founding / Staff / Director · Available · 4 weeks’ notice

Arpit Maheshwari · First edition · 2026 · The Trust Layer · Vol. I

You’re reading the no-script edition — the same book, without the page-turns. The full text follows. The interactive version, and the same content as a scrolling page, live at arpitmaheshwari.com.

How I Lead

Hire me and week one looks like this: I'm reading eval results before opening a design file, sitting silent on customer calls, and writing the diagnosis nobody assigned. By week two we're arguing productively. The best call in an AI product is rarely the interface — it's what the system is willing to admit it doesn't know.

What gets measured is a design decision.

By Friday of week one I've read your evals and sat in your customer calls. I shape what gets measured, then I ship the front-end — the CSS I own goes out under my name in the PR.

I design the wrong-answer screen first.

No AI feature ships until I've watched someone fail to use it. If I can't draw how the system fails, the happy path doesn't matter. Trust is built in the error states.

The org chart is the hardest wireframe.

Most UX problems are misaligned teams wearing UX clothes, so I design the organization before the interface. Then I write the system down — the next designer should inherit more than my taste.

Override is a feature, not a failure.

A user correcting the model is the training data the next version needs. I design the override as a first-class move — logged, visible, fed back into the next eval, so people see their fingerprints on next week's calls.

I read the data before I open Figma.

SQL, raw support tickets, model evals — I want the signal before the summary. Decisions grounded in the data survive the review; the ones I made on instinct don't.

About

Fifteen years, five industries, the same half-second: the model surfaces something true, and the person at the screen pauses. Not because the model is wrong. Because they don’t know how to bet on it yet. This book is everything I’ve worked out about that pause.

Selected Work

Every brief opened with “improve the UX.” Every diagnosis ended somewhere else. Fifteen years of this work mostly lives behind NDAs; the six below are the shape of all of it — two told in full, four as decision walkthroughs.

  • PTC University — Learning Connector — EdTech · Non-NDA · $1M/yr
    Five learning platforms, one survivor, eleven languages. Drawing the screens was easy; convincing a company to retire four products and rewire its revenue model was the work that mattered.
  • Telefónica MyO2 & Priority Moments — Telecom · Non-NDA · 4M+
    Two O2 UK products at national scale. Every screen drawn by me, then coded by me — mobile web for four million pockets.
  • AI-Assisted Private Equity Investing — FinTech · NDA · 60% faster
    Analysts get paid to doubt confident numbers. I held the launch until the score could argue its own case; then they leaned on it.
  • Programmatic Advertising Platform — AdTech · NDA · 2 wks → 3 hrs
    Traders watched an algorithm beat them and still played hunches. The fix wasn't a better model — a score tied to one action, reasons named out loud, an override that taught next week.
  • OrgOS · Transparent Org Tooling — Org Design · NDA · 200
    Eight modules doing the coordination work a management layer usually does. Designed for 200 people with no managers — 250 run on it today.
  • Technical Due Diligence Platform — VC/PE · NDA · 3 wks → 4 days
    Partners stake millions on claims they'll never personally verify. Extracting the signals was the model's job; getting partners to stand on the extraction was mine.

The Method — how the work gets made

Every product is a series of bets someone else has to accept. The process is a machine for making each bet smaller, better-evidenced, and easier to say yes to.

Act I — the wager worth making. Desirable, feasible, viable — the overlap is the bet. Everything outside it dies in review.

Act II — the spiral. Listen → Structure → Prove → Land, in loops; every loop ends in front of a user. Prototypes are AI-assisted working HTML, tested with users, shipped as part of the codebase. Research is a rhythm, not a phase.

Act III — the loop that never closes. MVP, then version n. Shipping is the first honest data — what users do returns as the next brief.

Confidence is earned in loops, not declared in launches.

01 · PTC University — Learning Connector

EdTech · Non-NDA

Five platforms, one survivor. The redesign took a quarter — the case for killing four products took a year. That was the design work.

Role
Lead Product Designer
Span
2014–2019 · two squads
Surface
Web LMS · 11 languages
Result
Shipped · in production

PTC sold its software on perpetual licenses: pay once, own forever. Around it sat five learning platforms — Learning Connector, LearningExchange, Precision LMS, Digital Guides and IoTU — five logins, five lines on an invoice. Three weeks in the customer-success recordings told me nothing was wrong with the navigation. The brief said “redesign the UX.” I argued the contract was the broken interface.

Perpetual licenses meant no recurring revenue, which meant stale content, which meant engineers learned on YouTube instead. The CRO had 60% of revenue on perpetual and said so loudly. Consolidation meant telling four executives their product was now a tab — a case made in the language of risk and P&L, not pixels.

Three decisions did the load-bearing work:

One data model before one UI

Rebuilt the content model first — one skill graph every platform mapped onto — so merging was a data migration, not a turf war.

Localisation as an architecture decision

Built knowing German and Russian run 30% longer: short labels, shallow hierarchy, no text in images. Nine languages, built to scale to eleven, so no region could fork off.

A switch-off ladder

Sequenced the four shutdowns so each VP watched their users land softly before the portal went dark. New customers on subscription from Q3 2017; existing ones protected for 24 months.

  • $1M Saved per year — print + shipping
  • 5→1 Platforms consolidated
  • 9→11 Languages, one pipeline
  • 550k+ Registered · 350k+ active

The miss, written down: my first accessibility pass buried screen-reader users in verbose ARIA labels — testing showed they skim, not listen. A week navigating with the monitor off, then I recoded the front end.

The redesign took a quarter. The case for deleting four products took a year — and that was the actual design work.— PTC University, project note

02 · Telefónica MyO2 & Priority Moments

Telecom · Non-NDA

Drawn by me, then coded by me — every screen of two O2 UK products on mobile web, at a scale where rounding errors have populations.

Role
Designer + Front-end
Client
O2 UK (Telefónica) · via Equal Experts
Status
Shipped · public

The one move: own both sides of the handoff. I drew the screens, then wrote the front-end that shipped them — nothing lost in translation between a design file and an engineer who never saw the intent. MyO2 and Priority Moments, mobile web, national scale.

MyO2 — the whole account, alone

O2 UK's self-service app: data and usage, the bill, a tariff change, an upgrade — the whole account without dialing anyone. The math is blunt: every self-service task that lands is a contact-centre call that never happens. It went on to serve more than four million users.

Priority Moments — a reason to open it

O2's loyalty programme: rewards from Odeon, M&S, Caffè Nero, matched by interest, behaviour and location. Launched July 2011; 2.6M registrations in year one, 2.5M+ active. The launch figures are O2's record — I joined in 2013 and owned the reward and offer screens.

Same designer, same stack, opposite job

MyO2 is a utility; Priority is a habit. Both on mobile web under a top UK brand, where small things stop being small — a tap target, a spinner, an exact billing figure lands on a stadium at once. The outcome figures are public, reported by O2 and Equal Experts. The claim is the work: every screen designed and built by me.

  • 4M+ MyO2 users served
  • 2.6M Priority sign-ups · yr 1
  • 5★ Priority App Store rating

03 · AI-Assisted Private Equity Investing

FinTech · NDA

I held the release until the model could defend its own scores. Screening ran 60% faster once it could.

Role
Lead Product Designer
Surface
AI for private-equity investing
Status
Shipped · under NDA

The model worked — that was never the problem. PE analysts are paid to doubt confident numbers, and a score they couldn't cross-examine wasn't a tool; it was a liability they treated like one. So I held the release until it could defend its own scores — without that trust the tool was dead on arrival, however good the model.

Explain before the verdict

An “explain this score” surface let an analyst pull any rating apart into its signals, challenge the weighting, and watch the score answer back. The three sources sat next to the number, not buried behind it.

Design the decline

A visible “I'm not sure about this one” state for the low-confidence cases, so the model could refuse to bluff. Analysts trusted the confident answers more once they'd watched it decline.

Disagreement on record

A logged override when the analyst disagreed. It drew the longest argument — logging dissent felt exposing — but it built the training signal that sharpened the model over time.

  • 60% faster Time per diligence pass
  • 3 Sources behind every score
  • Lead Analysts now open with it

04 · Programmatic Advertising Platform

AdTech · NDA

The scoreboard said the algorithm beat the traders. They read the scoreboard and played their hunches anyway. That’s a missing interface, not a modeling failure.

Role
Lead Product Designer
Surface
DSP recommendation UI
Status
Shipped · under NDA

The engine outperformed the buyers, visibly, and adoption sat near zero. It handed down verdicts without arguments. A bare number asks for faith; traders deal in collateral, so they ignored it.

A score tied to one action

Each recommendation resolved to one of three verbs — act, review, or ignore — and never reached the screen as a naked 87%.

Reasons named out loud

A reasoning panel named the exact signals behind each call — not a tooltip, the actual argument — so a buyer could check the model's work the way they'd check an analyst's.

An override that teaches

The override logged the correction and fed next week's model. The first time a buyer watched their pushback sharpen the next week's calls, fighting flipped to coaching. Planning fell from about two weeks to three hours; the model never changed.

  • 2 wks → 3 hrs Campaign planning time
  • 3 Actions a score maps to
  • Why Reasoning on every call

05 · OrgOS · Transparent Org Tooling

Org Design · NDA

Two hundred people. No managers. Eight modules doing the job of an org chart — coordination that doesn't smuggle a boss back in through the side door.

Role
Design Lead
Surface
Internal operating system
Status
Shipped · under NDA

Transparency does the coordinating — salaries, finances, assignments, reviews, open to everyone. That holds at forty; at two hundred the hallway stops scaling. Leading four engineering streams and a PM, I made the flat org legible without imposing a hierarchy. Every obvious feature — assignment, approval, escalation — was a manager wearing a different name. The job was saying no to each one.

Read access is the feature

Who's on what, who's blocked, who decides — visible to everyone, always. Pull, not push. Coordination came from information, not instruction.

Commitments, not assignments

People pull work and publish commitments in the open. The system tracks promises kept; it never hands out tasks.

Eight modules, one grammar

Staffing, comp, OKRs, onboarding all spoke one object model — so the org could rebuild its own process with nobody in the room to arbitrate. The user base and the org structure were the same two hundred humans, so every design call was an organizational one.

  • 250 On it today · designed for 200
  • 0 Managers in the loop
  • 8 Modules, one grammar

06 · Technical Due Diligence Platform

VC/PE · NDA

Partners bet millions on claims they'll never personally check. The model extracted the evidence; the design made them willing to stand on it. Three-week cycles ran to four days.

Role
Design Lead
Surface
Technical-DD platform · VC + PE
Status
Shipped · under NDA

The hard part isn't finding the signals — models do that. It's getting a partner to attach their reputation to an extraction they didn't perform. So the design budget went to provenance, confidence that maps to a next step, and a clean drill from summary to source.

Score at the signal level

Confidence on each technical signal — code quality, architecture risk, team velocity, founder credibility — not one opaque verdict, each one mapping to a next step a partner can take.

Provenance on every claim

Each score named the signals that drove it, with a clean drill from summary to source, so a partner could inspect the reasoning before committing capital.

Dissent on record

Analyst overrides fed back into the model, and partner sign-off was real workflow, not a rubber stamp. Disagreement became training signal, not noise.

  • 3 wks → 4 days Diligence cycle time
  • VC + PE Both fund types served
  • 4 Signal classes scored

A Field Guide to Trust

Everything here ran in production, failed somewhere specific, and came back stronger. The tradeoffs are written down because I paid for them once already — so you don’t have to. Each pattern follows the same arc: what it does, where it failed me, what I changed.

01 · Confidence Score Patterns

A score with no action attached is decoration.

How much certainty to show, in what form, and the threshold at which a number earns the right to drive a decision instead of decorating a dashboard. Five ways to put a number on certainty, and when each one earns or burns trust.

Do

  • Anchor the score to an action — act, review, or ignore — not just a bare number.
  • Show the score's own track record so people can calibrate their trust.
  • Round to the precision you'd be willing to defend out loud.

Don’t

  • Render 87.3% when what you actually mean is “probably.”
  • Let a high score auto-execute with no visible way to override.
  • Reuse one confidence scale across decisions of wildly different stakes.

AdTech · Programmatic: Buyers ignored the recommendation until the score sat beside what it had gotten right last month. Confidence with a memory got acted on; confidence alone never did.

02 · Failure States

If you can't design the wrong answer, the feature isn't ready.

What the screen says when the model can't deliver. Saying it honestly — and making recovery from a miss cheaper than the mistake itself — is what keeps users from leaving.

Do

  • Design the wrong-answer screen before you design the happy path.
  • Make recovery from a miss cheaper than the mistake itself.
  • Say what the system doesn't know, plainly and early.

Don’t

  • Hide uncertainty behind a confident-looking default.
  • File “what if it's wrong” as an edge case to handle later.
  • Apologise for an error without offering the next step.

FinTech · Due Diligence: I shipped the “I'm not sure about this one” state first. Analysts trusted the confident answers more once they'd watched the model decline to bluff.

03 · Explainability

Turn a black-box output into an argument a person can audit.

Showing a non-technical person why the machine decided, at a depth they can use without a stats degree — the difference between obeying a score and owning the decision.

Do

  • Show the two or three inputs that actually moved the result.
  • Let the user trace from the output back to the evidence.
  • Make “I disagree” a first-class, recorded action.

Don’t

  • Dump every feature weight on screen and call it transparency.
  • Explain after the decision instead of before it.
  • Mistake a tooltip for an account of the reasoning.

FinTech · Due Diligence: Analysts went from ignoring the risk score to leading their memos with it — once “explain this score” surfaced the three documents behind the number.

04 · Human-in-the-Loop

Keep the person in command when the stakes outgrow the model.

Where and how the person corrects the system — turning corrections into the training signal the next version needs, so the workflow scales without growing overhead.

Do

  • Make the human's edit visibly improve the next result.
  • Put the control where the decision happens, never buried in settings.
  • Default to the human's last call when the stakes are high.

Don’t

  • Ask for approval on everything until approval means nothing.
  • Treat corrections as exceptions instead of as training signal.
  • Make overriding feel like a fight with the product.

AdTech · Programmatic: When a buyer's override visibly retrained the next week's recommendation, correcting the model stopped feeling like rework and started feeling like teaching.

05 · Provenance & Citations

An answer you can trace is an answer you'll defend.

Explainability says why the model decided; provenance says where the evidence came from — the exact source behind every claim, one click from the number to the document that produced it.

Do

  • Put the source next to the claim, not behind a 'details' link.
  • Let a person open the original document the model read, unedited.
  • Say how many sources back a number — and flag the one that disagreed.

Don’t

  • Cite a source the user can't actually open and verify.
  • Summarise the evidence so heavily that the trail goes cold.
  • Reveal provenance only after the answer is challenged.

VC · Technical Diligence: Partners signed off faster once every score carried a clean drill from summary to the source document — they'll stand on an extraction they can open, never one they can't.

06 · The Capability Contract

Trust starts with an honest “no.”

The model's promise, stated up front: what this system is for, where it taps out, and what it hands back to a human — set before the first use, not after the first complaint.

Do

  • State the model's limits in the interface, not just the docs.
  • Hand off to a human the moment a request leaves the model's competence.
  • Make the boundary specific — 'I can't price illiquid assets,' not 'results may vary.'

Don’t

  • Imply the product can do things it can't, then degrade silently.
  • Bury scope in a terms page nobody reads.
  • Treat 'out of scope' as an error instead of an honest answer.

AdTech · Programmatic: Each call resolved to act, review, or ignore — and “ignore” was the model admitting it had nothing worth saying. The honest no is what made buyers believe the act.

07 · Calibration & Track Record

Show whether the confidence has been right before.

A confidence number is a claim; its track record is the evidence. Show whether “80% sure” has actually been right about 80% of the time — so a person learns how hard to lean, and watches that judgment improve as the history grows.

Do

  • Show the model's hit rate beside its current confidence.
  • Break the record down by the kind of case, not one global average.
  • Let the history update in the open, so trust is earned, not assumed.

Don’t

  • Show a confidence number with no past to back it.
  • Average away the cases where the model is reliably wrong.
  • Reset the track record silently every time the model changes.

FinTech · Due Diligence: Ninety days and forty-two deals in, the score had been right often enough that analysts stopped re-checking the confident calls. The history earned the trust the number alone couldn't.

08 · Reversibility

People act on the model when it's easy to walk back.

Adoption stalls when acting feels risky, not when the model is wrong. Make the action cheap to undo — one click to reverse, a clear path back, no permanent damage — and people will try the recommendation they'd otherwise ignore.

Do

  • Make acting on a recommendation one click to reverse.
  • Show the way back before the person commits.
  • Stage risky changes so they can be halted, not just rolled back.

Don’t

  • Hide undo, or make reversing cost more than the original action.
  • Make a wrong call feel permanent.
  • Force an irreversible commit to get any value from the model.

EdTech · PTC University: Retiring four products, I sequenced the shutdowns so every team watched their users land softly before the lights went out — a migration you could halt beats a leap you can't take back.

Notes & Writing

One idea per issue, on getting humans to act on machines. The fastest way to know how I think before you hire me.

  • The Agentic MVP: Why Your Next Launch Will Be Lovable — 2026 · Essay
    How the rise of agentic systems is rewriting what 'minimum viable' means — and why lovability is now the bar.
  • The AI Fight Club — 2026 · Field note
    Weaponizing Claude and Gemini for Bulletproof Products. A practical method for stress-testing AI interfaces.
  • Designing for the eval, not the mock — 2025 · Talk
    What changes when designers own what gets measured.
  • Transparency as coordination — 2024 · Essay
    What a zero-manager org taught me about interface design.
  • Reading the support tickets myself — 2024 · Field note
    The cheapest research method nobody on the design team wants to do.

Curriculum Vitæ

The facts, in order — for the founder who checks the work before the call. Every number arrives holding its baseline.

Focus

Model-Layer Design · Data-Intensive UI · Design Leadership · Organizational Design · Research & Evals · System Architecture

Experience

  • 2019— — Solution Consultant
    Sahaj AI · AI products for AdTech, HRTech & Private Equity
  • 2014–19 — Lead Product Designer
    PTC Inc. · PTC University — Learning Connector · 550k+ registered, 350k+ active · NASA, Apple, Boeing, Airbus & more
  • 2012–14 — Front End Specialist
    Equal Experts · O2 UK consumer apps · 4M+ users
  • 2010–12 — Systems Analyst
    Tata Consultancy Services · mobile for Fortune 500
  • 2009 — Intern
    Nokia Networks · Indore

Education

  • 2017–18 — Executive MBA, Business Analytics
    Institute of Management Technology · Ghaziabad
  • 2006–10 — B.E., Electronics & Communication
    Shri Vaishnav Institute of Technology & Science · Indore

Contact

One seat. Full-time. Yours to offer.

Your model is right. Your users still won't bet on it. That half-second of doubt is the only thing I design. I'm choosing one role: founding or staff designer at an AI product company, or a director-level seat where the trust layer is the job description. Available — 4 weeks' notice.

  • Available · 4 weeks' notice
  • Founding / Staff / Director
  • Remote · GMT+5:30

Send me the role · LinkedIn · The Trust Layer on Substack

I design the trust layer of AI products — the surface where a person decides to act on the model.

No copyright · Design is for all