Skip to the book

Human in the Loop

Sixteen years designing the half-second where a person decides to bet on a machine.

Staff / Principal or Founding · Available · 4 weeks’ notice

Arpit Maheshwari · First edition · 2026 · Human in the Loop · Vol. I

You’re reading the no-script edition — the same book, without the page-turns. The full text follows. The interactive version, and the same content as a scrolling page, live at arpitmaheshwari.com.

How I Lead

Hire me and week one looks like this: I’m reading eval results before opening a design file, sitting silent on customer calls, and writing the diagnosis nobody assigned. By week two we’re arguing productively. The best call in an AI product is rarely the interface — it’s what the system is willing to admit it doesn’t know.

What gets measured is a design decision.

By Friday of week one I’ve read your evals and sat in your customer calls. I shape what gets measured, then I ship the front-end — the CSS I own goes out under my name in the PR.

I design the wrong-answer screen first.

No AI feature ships until I’ve watched someone fail to use it. If I can’t draw how the system fails, the happy path doesn’t matter. Trust is built in the error states.

The org chart is the hardest wireframe.

Most UX problems are misaligned teams wearing UX clothes, so I design the organisation before the interface. Then I write the system down — the next designer should inherit more than my taste.

Override is a feature, not a failure.

A user correcting the model is the training data the next version needs. I design the override as a first-class move — logged, visible, fed back into the next eval, so people see their fingerprints on next week’s calls.

I read the data before I open Figma.

SQL, raw support tickets, model evals — I want the signal before the summary. Decisions grounded in the data survive the review; the ones I made on instinct don’t.

About

Sixteen years, five industries, the same half-second: the model surfaces something true, and the person at the screen pauses. Not because the model is wrong. Because they don’t know how to bet on it yet. This book is everything I’ve worked out about that pause.

Selected Work

Every brief opened with “improve the UX.” Every diagnosis ended somewhere else. Sixteen years of this work mostly lives behind NDAs; the seven below are the shape of all of it — one told in full, six as decision walkthroughs.

  • An Ad Agency Became the Market’s Aggregator — AdTech · Named client · 2 wks → 3 hrs
    Traders watched an algorithm beat them and still played hunches. The fix wasn’t a better model — recommendations as full campaign plans, KPIs attached, one click to customise, every edit teaching next week.
  • Deal-Screening AI That Cites Its Sources — FinTech · NDA · 60% faster
    It did not ship until every generated number could name its source and thin evidence drew an honest abstention — analysts stopped ignoring the model and started leading their deal memos with it.
  • Due Diligence You Can’t Rubber-Stamp — VC/PE · NDA · 3 wks → 4 days
    Partners stake millions on claims they’ll never personally verify. So sign-off stays locked until the evidence is read — the friction was the product.
  • The Software That Replaced the Org Chart — Org Design · NDA · 200
    Eight modules doing the coordination work a management layer usually does. Designed for 200 people with no managers — 250 run on it today.
  • The Redesign That Asked PTC to Kill Four Products — EdTech · Non-NDA · $1M/yr
    Five learning platforms, one survivor, eleven languages. Drawing the screens was easy; convincing a company to retire four products and rewire its revenue model was the work that mattered.
  • Two O2 UK Products, Four Million People — Telecom · Non-NDA · 4M+
    Two O2 UK products at national scale. Every screen co-designed, every screen coded by me — mobile web for four million pockets.
  • The App People Broke on Purpose — Consumer · PlanIt · 0 complaints
    The same client’s data, pointed at a commuter. A London transit app that answered “will I be able to breathe when I get on?” — and whose users triggered the error screen on purpose, just to watch the little train be sorry. The failure state became the reason they came back.

The Method — how the work gets made

Every product is a series of bets someone else has to accept. The process is a machine for making each bet smaller, better-evidenced, and easier to say yes to.

Act I — the wager worth making. Desirable, feasible, viable — the overlap is the bet. Everything outside it dies in review.

Act II — the spiral. Listen → Structure → Prove → Land, in loops; every loop ends in front of a user. Listen is user research and requirement gathering; Structure is journey maps, service blueprints, information architecture. Prototypes are AI-assisted working HTML, usability-tested, shipped as part of the codebase. Research is a rhythm, not a phase.

Act III — the loop that never closes. MVP, then version n. Shipping is the first honest data — what users do returns as the next brief.

Confidence is earned in loops, not declared in launches.

01 · The Redesign That Asked PTC to Kill Four Products

EdTech · Non-NDA

Five platforms, one survivor. The redesign took a quarter — the case for killing four products took a year. That was the design work.

Role
Product & Design Lead
Span
2014–2019
Surface
Web LMS · 11 languages
Result
Shipped · in production

PTC sold its software on perpetual licenses: pay once, own forever. Around it sat five learning platforms — Learning Connector, LearningExchange, Precision LMS, Digital Guides and IoTU — five logins, five lines on an invoice. Three weeks in the customer-success recordings told me nothing was wrong with the navigation. The brief said “redesign the UX.” I argued the contract was the broken interface. Research had already named the survivor: usability testing kept showing customers didn’t use our product names — they called everything “PTC University.”

Perpetual licenses meant no recurring revenue, which meant stale content, which meant engineers learned on YouTube instead. The CRO had 100% of revenue on perpetual and said so loudly. Consolidation meant telling four executives their product was now a tab — a case made in the language of risk and P&L, not pixels.

Three decisions did the load-bearing work:

One data model before one UI

Rebuilt the content model first — one skill graph every platform mapped onto — so merging was a data migration, not a turf war.

Diagram: 100% perpetual licences at the start; free tutorials and trainings let learners experience the product; that experience converted them to premium subscriptions, moving subscription from 0% to 64% of new bookings in five quarters.
The funnel was the product pitch — 0% to 64% of new bookings in five quarters.

Localisation as an architecture decision

Built knowing German runs ~30% longer: short labels, shallow hierarchy, no text in images. Nine languages, built to scale to eleven, so no region could fork off.

A switch-off ladder

Sequenced the four shutdowns so each VP watched their users land softly before the portal went dark. New customers on subscription from Q3 2017; existing ones protected for 24 months.

Diagram: five platforms consolidate to one over 24 months — one skill graph first so merging was a data migration, then four sunsets each landing their users softly, ending at one research-named front door with a licence-keyed homepage.
Data model first, four soft landings, one researched name.
  • $1M Saved per year — print + shipping
  • 5→1 Platforms consolidated
  • 9→11 Languages, one pipeline
  • 550k+ Registered · 350k+ active

The miss, written down: my first accessibility pass buried screen-reader users in verbose ARIA labels — a usability study showed they skim, not listen. A week navigating with the monitor off, then I recoded the front end.

The redesign took a quarter. The case for deleting four products took a year — and that was the actual design work.— PTC University, project note

02 · Two O2 UK Products, Four Million People

Telecom · Non-NDA

Designed with one co-designer, coded by me alone — every screen of two O2 UK products on mobile web, at a scale where rounding errors have populations.

Role
Designer + Front-end
Client
O2 UK (Telefónica) · via Equal Experts
Status
Shipped · public

The one move: own both sides of the handoff. I designed the screens with a co-designer, then wrote the front-end that shipped them — nothing lost in translation between a design file and an engineer who never saw the intent. MyO2 and Priority Moments, mobile web, national scale.

MyO2 — the whole account, alone

O2 UK’s self-service app: data and usage, the bill, a tariff change, an upgrade — the whole account without dialing anyone. A replatforming, not a fresh start — the legacy system made responsive across iOS, Android, Windows Phone and web, moved cautiously with millions of subscribers on it, in step with Telefónica’s brand and copy teams. It went on to serve more than four million users.

Priority Moments — a reason to open it

O2’s loyalty programme: a geolocated list of rewards near you — Odeon, M&S, Caffè Nero — every redemption a reason to stay. Launched July 2011; 2.6M registrations in year one, 2.5M+ active. The launch figures are O2’s record — I joined in 2013 and owned the reward and offer screens.

Same designer, same stack, opposite job

MyO2 is a utility; Priority is a habit. Both on mobile web under a top UK brand, where small things stop being small — a tap target, a spinner, an exact billing figure lands on a stadium at once. The outcome figures are public, reported by O2 and Equal Experts. The claim is exact: every screen co-designed, and every line of the front-end that shipped them written by me.

Diagram: the legacy MyO2 system rebuilt as one responsive build serving iOS, Android, Windows Phone and desktop web from the same code — moved cautiously with millions of subscribers aboard, in step with Telefónica brand and copy teams.
A replatforming, not a fresh start — one build, four surfaces.
  • 4M+ MyO2 users served
  • 2.6M Priority sign-ups · yr 1
  • 5★ Priority App Store rating

03 · Deal-Screening AI That Cites Its Sources

FinTech · NDA

The model shipped already able to defend its own scores. Screening ran 60% faster once it could.

Role
Lead Product Designer
Surface
AI for private-equity investing
Status
Shipped · client named, on the public record

The model was already good enough for the workflow; adoption was the bottleneck. PE analysts are paid to doubt confident numbers, and a score they couldn’t cross-examine wasn’t a tool; it was a liability they treated like one. So it shipped already able to defend its own scores — without that trust the tool was dead on arrival, however good the model.

One analysis, two reading modes: a written descriptive report — the scores argued in prose — beside the conversation that interrogates it, over a pipeline-wide scoring dashboard. Analysts read differently; the design refused to pick a temperament for them.

Diagram: two launch conditions — every generated number carries its source, and thin evidence draws an abstention — each with the failure it prevents, gating the release.
The two conditions I would not ship without — and the failure each one bought off.
Diagram: a document enters from the inbox, claims are matched to sources, risk is scored, then a gate decides — thin evidence abstains, grounded evidence emits a score with its sources; the analyst reads it and any override is logged back to train the next read.
The path a document takes — and the branch where the system declines to answer.
Diagram: one analysis rendered two ways — a written report and a conversation — both over a pipeline-wide scoring dashboard, sharing the same evidence.
One analysis, two readings — over the dashboard both readers return to.

Explain before the verdict

An “explain this score” surface let an analyst pull any rating apart into its signals, challenge the weighting, and watch the score answer back. The three sources sat next to the number, not buried behind it.

Design the decline

A visible “I’m not sure about this one” state for the low-confidence cases, so the model could refuse to bluff. Analysts trusted the confident answers more once they’d watched it decline.

Disagreement on record

A logged override when the analyst disagreed. It drew the longest argument — logging dissent felt exposing — but it built the training signal that sharpened the model over time.

  • 60% faster Time per diligence pass
  • 3 Sources behind every score
  • Lead Analysts now open with it

04 · An Ad Agency Became the Market’s Aggregator

AdTech · Named client — Talon Outdoor

The algorithm beat the traders, and they played their hunches anyway — until they could argue back, one click to reshape the plan, every edit teaching next week’s calls.

Role
Lead Product Designer
Surface
End-to-end aggregator · 6 systems
Status
Shipped · under NDA

The principle this case proves: users don’t adopt the most accurate system — they adopt the one that lets them argue back.

The engine outperformed the buyers, visibly, and adoption sat near zero. Its raw output was a list of billboards, handed down without arguments. A bare list asks for faith; traders deal in collateral, so they ignored it.

A plan, not a pick

The screen turned the model’s raw list into a full campaign plan a trader could defend — never a bare list, never a bare number.

Diagram: the engine produced a ranked list of billboards; the screen turned it into a full campaign plan with KPIs and reasoning attached, customisable in one click, with every edit logged into the next round.
The model didn’t change — what it handed the trader did.
Diagram: the old brief named a place; once the plans were believed the brief could name an audience and a moment, and the system chose the sites.
What adoption unlocked — a brief that names an audience, not an address.
Diagram: five connected systems forming a two-sided marketplace, with the audience-intelligence platform between demand and supply.
Six systems, one connected platform — demand on one side, supply on the other.

KPIs on the plan itself

Reach, estimated ROI, price trend — the case for the plan sat on the plan, checkable before anyone committed budget.

Customisation that teaches

Reshaping the plan took one click; every edit was logged and fed next week’s recommendations. Watching their pushback land flipped fighting into coaching. Six systems carried the principle across the platform — the trading floor (Plato), the audience intelligence (Ada), the creative management, the play reporting, and a free inventory SaaS for media owners, and a white-label booking product franchisors brand as their own, unified under one design system — and the structural result was public: a media agency became the market’s aggregator. The client’s leadership later put the arc on the public record: platform bookings grew from under 5% to 100% of UK bookings. Planning fell from about two weeks to three hours; the model never changed.

  • 2 wks → 3 hrs Campaign planning time
  • £69k Media-value gain per client
  • Why Reasoning on every call

05 · The Software That Replaced the Org Chart

Org Design · NDA

Two hundred people. No managers. Eight modules doing the job of an org chart — coordination that doesn’t smuggle a boss back in through the side door.

Role
Product & Design Lead
Surface
Internal operating system
Status
Shipped · under NDA

Transparency does the coordinating — salaries, finances, assignments, reviews, open to everyone. That holds at forty; at two hundred the hallway stops scaling. Leading four engineering streams and a PM, I made the flat org legible without imposing a hierarchy. Every obvious feature — assignment, approval, escalation — was a manager wearing a different name. The job was saying no to each one.

Read access is the feature

Who’s on what, who’s blocked, who decides — visible to everyone, always. Pull, not push. Coordination came from information, not instruction.

Commitments, not assignments

People pull work and publish commitments in the open. The system tracks promises kept; it never hands out tasks.

Diagram: three requested features — assign, approve, escalate — each struck through as a manager wearing a different name, replaced by one mechanism: visibility, with commitments published in the open. Eight modules speak one object model.
The features I said no to — and the one mechanism that replaced them.

Eight modules, one grammar

Staffing, comp, OKRs, onboarding all spoke one object model — so the org could rebuild its own process with nobody in the room to arbitrate. The user base and the org structure were the same two hundred humans, so every design call was an organisational one.

  • 250 On it today · designed for 200
  • 0 Managers in the loop
  • 8 Modules, one grammar

06 · Due Diligence You Can’t Rubber-Stamp

VC/PE · NDA

Partners bet millions on claims they’ll never personally check. So the design made them do more work — no sign-off until the evidence is read. The friction cut three weeks to four days.

Role
Product & Design Lead
Surface
Technical-DD platform · VC + PE
Status
Shipped · under NDA

The hard part isn’t finding the signals — models do that. It’s getting a partner to attach their reputation to an extraction they didn’t perform. So the design budget went to provenance, confidence that maps to a next step, and a clean drill from summary to source.

Score at the signal level

Confidence carried per signal — a finding crossed into a signal only when the analysis cleared the confidence bar, each one holding the evidence behind it — never one opaque verdict, each one mapping to a next step a partner can take.

Provenance on every claim

Each score named the signals that drove it, with a clean drill from summary to source, so a partner could inspect the reasoning before committing capital.

Diagram: a partner verdict over three signals, two with their evidence trails opened and one not yet opened, with the sign-off control locked until every trail has been read.
Sign-off stays locked until the evidence has been read.

Dissent on record

Analyst overrides fed back into the model, and partner sign-off was real workflow, not a rubber stamp. Disagreement became training signal, not noise.

  • 3 wks → 4 days Diligence cycle time
  • VC + PE Both fund types served
  • 16 Dimensions analysed per finding

A Field Guide to Trust

Everything here ran in production, failed somewhere specific, and came back stronger. The tradeoffs are written down because I paid for them once already — so you don’t have to. Each pattern follows the same arc: what it does, where it failed me, what I changed.

01 · Confidence Score Patterns

An unexplained 87% is a shrug with decimals. A score earns its pixels only when it resolves to a verb.

How much certainty to show, in what form, and the threshold at which a number earns the right to drive a decision instead of decorating a dashboard. Five ways to put a number on certainty, and when each one earns or burns trust.

Do

  • Anchor the score to an action — act, review, or ignore — not just a bare number.
  • Show the score’s own track record so people can calibrate their trust.
  • Round to the precision you’d be willing to defend out loud.

Don’t

  • Render 87.3% when what you actually mean is “probably.”
  • Let a high score auto-execute with no visible way to override.
  • Reuse one confidence scale across decisions of wildly different stakes.

AdTech · Programmatic: Buyers ignored the recommendation until it arrived as a plan with the KPIs to check it. Evidence they could audit got acted on; a bare recommendation never did.

02 · Failure States

Users forgive a model for being wrong. They never forgive it for bluffing.

What the screen says when the model can’t deliver. Saying it honestly — and making recovery from a miss cheaper than the mistake itself — is what keeps users from leaving.

Do

  • Design the wrong-answer screen before you design the happy path.
  • Make recovery from a miss cheaper than the mistake itself.
  • Say what the system doesn’t know, plainly and early.

Don’t

  • Hide uncertainty behind a confident-looking default.
  • File “what if it’s wrong” as an edge case to handle later.
  • Apologise for an error without offering the next step.

FinTech · Due Diligence: I shipped the “I’m not sure about this one” state first. Analysts trusted the confident answers more once they’d watched the model decline to bluff.

03 · Explainability

A recommendation you can trace, you’ll defend. One you can’t, you’ll quietly rebuild around.

Showing a non-technical person why the machine decided, at a depth they can use without a stats degree — the difference between obeying a score and owning the decision.

Do

  • Show the two or three inputs that actually moved the result.
  • Let the user trace from the output back to the evidence.
  • Make “I disagree” a first-class, recorded action.

Don’t

  • Dump every feature weight on screen and call it transparency.
  • Explain after the decision instead of before it.
  • Mistake a tooltip for an account of the reasoning.

FinTech · Due Diligence: Analysts went from ignoring the risk score to leading their memos with it — once “explain this score” surfaced the three documents behind the number.

04 · Human-in-the-Loop

An override is not defiance. It is the training data the next version needs.

Where and how the person corrects the system — turning corrections into the training signal the next version needs, so the workflow scales without growing overhead.

Do

  • Make the human’s edit visibly improve the next result.
  • Put the control where the decision happens, never buried in settings.
  • Default to the human’s last call when the stakes are high.

Don’t

  • Ask for approval on everything until approval means nothing.
  • Treat corrections as exceptions instead of as training signal.
  • Make overriding feel like a fight with the product.

AdTech · Programmatic: When a buyer’s override visibly retrained the next week’s recommendation, correcting the model stopped feeling like rework and started feeling like teaching.

05 · Provenance & Citations

A claim with its source beside it is evidence. The same claim without one is prose.

Explainability says why the model decided; provenance says where the evidence came from — the exact source behind every claim, one click from the number to the document that produced it.

Do

  • Put the source next to the claim, not behind a ’details’ link.
  • Let a person open the original document the model read, unedited.
  • Say how many sources back a number — and flag the one that disagreed.

Don’t

  • Cite a source the user can’t actually open and verify.
  • Summarise the evidence so heavily that the trail goes cold.
  • Reveal provenance only after the answer is challenged.

VC · Technical Diligence: Partners signed off faster once every score carried a clean drill from summary to the source document — they’ll stand on an extraction they can open, never one they can’t.

06 · The Capability Contract

The honest “no” is what makes the confident answer believable.

The model’s promise, stated up front: what this system is for, where it taps out, and what it hands back to a human — set before the first use, not after the first complaint.

Do

  • State the model’s limits in the interface, not just the docs.
  • Hand off to a human the moment a request leaves the model’s competence.
  • Make the boundary specific — ’I can’t price illiquid assets,’ not ’results may vary.’

Don’t

  • Imply the product can do things it can’t, then degrade silently.
  • Bury scope in a terms page nobody reads.
  • Treat ’out of scope’ as an error instead of an honest answer.

FinTech · PE screening: Every claim cited its source, or the model said “I’m not sure” out loud. The honest no is what made analysts believe the yes.

07 · Calibration & Track Record

“80% sure” is a promise. Show whether it has been kept.

A confidence number is a claim; its track record is the evidence. Show whether “80% sure” has actually been right about 80% of the time — so a person learns how hard to lean, and watches that judgment improve as the history grows.

Do

  • Show the model’s hit rate beside its current confidence.
  • Break the record down by the kind of case, not one global average.
  • Let the history update in the open, so trust is earned, not assumed.

Don’t

  • Show a confidence number with no past to back it.
  • Average away the cases where the model is reliably wrong.
  • Reset the track record silently every time the model changes.

FinTech · Due Diligence: Once the track record was long enough to read, the score had been right often enough that analysts stopped re-checking the confident calls. The history earned the trust the number alone couldn’t.

08 · Reversibility

People don’t act when the model is right. They act when being wrong is cheap.

Adoption stalls when acting feels risky, not when the model is wrong. Make the action cheap to undo — one click to reverse, a clear path back, no permanent damage — and people will try the recommendation they’d otherwise ignore.

Do

  • Make acting on a recommendation one click to reverse.
  • Show the way back before the person commits.
  • Stage risky changes so they can be halted, not just rolled back.

Don’t

  • Hide undo, or make reversing cost more than the original action.
  • Make a wrong call feel permanent.
  • Force an irreversible commit to get any value from the model.

EdTech · PTC University: Retiring four products, I sequenced the shutdowns so every team watched their users land softly before the lights went out — a migration you could halt beats a leap you can’t take back.

Notes & Writing

One idea per issue, on getting humans to act on machines. The fastest way to know how I think before you hire me.

A loose thread runs through all five: what happens when a system that behaves like a colleague still needs to be governed like software — the confidence to let it act, and the discipline to check it did.

  • The Agentic MVP: Why Your Next Launch Will Be Lovable — 2026 · Essay
    How the rise of agentic systems is rewriting what ’minimum viable’ means — and why lovability is now the bar.
  • The AI Fight Club — 2026 · Field note
    Weaponizing Claude and Gemini for Bulletproof Products. A practical method for stress-testing AI interfaces.
  • The New Renaissance — 2026 · Essay
    How agentic AI is shifting knowledge work from executing tasks to exercising judgment.
  • AI as Your Startup Cofounder — 2026 · Essay
    Treating AI as a 24/7 strategic partner — if you govern its decision-making first.
  • How I Forbid AI from Hallucinating — 2026 · Essay
    Forcing AI to ask before it assumes — clarifying questions before any implementation work begins.

Curriculum Vitæ

The facts, in order — for the founder who checks the work before the call. Every number arrives holding its baseline.

Focus

Model-Layer Design · Product Definition & Roadmaps · Data-Intensive UI · Design Leadership · Organisational Design · Research & Evals · System Architecture

Experience

  • 2019— — Product & Design Lead — AI Products
    Sahaj AI · AdTech, HRTech & Private Equity
  • 2014–19 — Product & Design Lead
    PTC Inc. · PTC University — Learning Connector · 550k+ registered, 350k+ active · NASA, Boeing, Toyota, Airbus & Apple
  • 2012–14 — Front End Specialist
    Equal Experts · O2 UK consumer apps · 4M+ users
  • 2010–12 — Systems Analyst
    Tata Consultancy Services · mobile for Fortune 500
  • 2009 — Intern
    Nokia Networks · Indore

Education

  • 2017–18 — Executive MBA, Business Analytics
    Institute of Management Technology · Ghaziabad
  • 2006–10 — B.E., Electronics & Communication
    Shri Vaishnav Institute of Technology & Science · Indore

Contact

One seat. Full-time. Yours to offer.

Your model is right. Your users still won’t bet on it. That half-second of doubt is the only thing I design. I’m choosing one role: staff / principal designer or founding product-design lead at an AI product company — open to a hands-on director seat. Available — 4 weeks’ notice.

  • Available · 4 weeks’ notice
  • Staff / Principal or Founding lead
  • Remote · GMT+5:30

Send me the role · LinkedIn · Human in the Loop on Substack

I do human-in-the-loop design for AI products — the surface where a person decides to act on the model.

No copyright · Design is for all