In 2019 a recommendation engine I was working on beat the traders it was built for — visibly, on their own numbers. Adoption sat near zero. I spent the next two weeks making it smarter. That was the mistake.
I got the diagnosis wrong first
The client was the UK’s largest out-of-home advertising company, and the engine did something the desk genuinely could not: it read demand, supply and price across the whole market and proposed where a campaign should run. It outperformed the traders’ own picks. They ignored it anyway.
My first instinct was the engineer’s instinct, and it was wrong. If a correct thing is not being used, make it more correct. Two weeks went into sharpening a model that was already sharp. Nothing moved, because accuracy was never the problem. I keep that miss on the case page too, because it is the most useful part of the story: the wrong suspect is expensive, and it looks exactly like diligence while you are chasing it.
The traders were not being stubborn. They were being rational. A bare recommendation asks for faith, and traders deal in collateral — their name goes on the booking, and “the system said so” is not a defence you can carry into a client review. The model was not being out-argued. It was being ignored, which is a different failure and has a different fix.
The model did not need to be more right. It needed to hand over something a professional could put their name on.
What I actually changed
Nothing in the model. What changed was what it handed over.
The engine’s raw output was a ranked list of billboards. A list is an instruction: take it or leave it, and there is no third move. So the screen stopped serving a list and started serving a full campaign plan — the sites, the channel split, the creative rotation, and the KPIs the plan was committing to, printed on the plan itself. Every component customisable in one click.
That last part is the whole thing. The trader could argue back — pull a cluster, shift budget between channels, reject a panel on pricing — and every one of those edits was logged and fed into the next week’s recommendations. Disagreement stopped being friction and became fuel. A trader who overrode Friday’s plan watched the following week’s arrive already reflecting the override, which is a very different feeling from being audited by a machine. Overrides fell, not because anyone was told to stop overriding, but because there was less left to argue with.
What it bought, and where it stops
Campaign planning went from two weeks to three hours. Against the traders’ own baselines the plans carried £69,000 more media value per client on average. The platform went from under 5% of the client’s UK bookings to all of them — that part is on the public record, not my claim.
The honest limit: argue-back only works when there is an accountable expert on the other side. Give the same interface to someone with no standing to disagree and you have built a slower list. And it cannot rescue a weak model — a plan built on bad reasoning is just a bad plan with more surface area to defend.
What I believe now
Adoption is not a function of accuracy. It is a function of whether a professional can put their name on the output. Most of the AI-adoption problems I have been handed since have arrived dressed as model problems and turned out to be handover problems — the system was right, and it gave the person nothing to hold. The fix is rarely inside the model. It is in what the model hands over, and whether the person can push back on it and be heard. Right isn’t enough. The handover has to earn the tap.
Want the playbook, not the story? The human-in-the-loop patterns behind this — when to ask, when to act, and how to make an override earn its keep.
See the pattern reference: The Override Is Training Data →Continue exploring