How I Design Confidence Scores for AI Products
A recommendation engine that beat human buyers — and got overruled half the time anyway. What I learned about the half-second where a person decides whether to trust a number.
Design the Failure State First
A model that demoed at 94% confidence and was wrong 40% of the time in production. Why the happy path is the smaller half of the job, and the failure state is where trust is won or lost.
The Explainability Layer: Making AI Legible
A compliance officer asked one question the system couldn't answer — and a trade went through unsupervised. Why "why this, not that?" is a design problem, not a data-science one.
Human-in-the-Loop: When to Ask, When to Override
A 94%-accurate recommender that students skipped 60% of the time — out of click fatigue, not disagreement. When asking for confirmation stops respecting the user and starts taxing them.