← ReferralCandy

Blog

We Didn't Trust AI With Our Attribution Numbers. So We Made It Earn That Trust.

Author
Revanth Tadikonda
Date
2026-08-23
We Didn't Trust AI With Our Attribution Numbers. So We Made It Earn That Trust.

A merchant clicks one of our ads on a Tuesday. They think about it for three days. On Friday they go looking in the Shopify App Store, find us there, and install.

Every conventional attribution tool credits the App Store.

The hop through the storefront severs the journey that actually created the demand: the ad, the blog post, the search. Only the doorway gets recorded. For a long time our numbers told us the App Store was our strongest acquisition channel, and our numbers were confidently wrong.

This is a structural hole rather than a tagging mistake. The storefront sits between your marketing and your customer by design, and it collects the credit on the way past. If you sell through any app store or marketplace, your numbers have the same hole.

Half a picture, and no way to know which half

Our best view before all this came from analytics on the App Store side. It could see roughly how merchants arrived at our listing, nothing upstream of it, and only for the installs it could match at all. Half a picture, with no way to know which half.

We had also tried pointing AI at the problem, more than once. It produced plausible answers quickly, which was exactly the problem. An interactive AI analysis cannot be reproduced. It is brilliant while you are inside it, gone when you close it, and offers no way to tell which of its answers were right. You cannot run a company metric on "the AI said so." So AI stayed a toy for us rather than a source of record.

Finding marketing insights with AI turns out to be the easy part. The hard part, and the part almost nobody does, is making the answers trustworthy enough to become the numbers your company reports.

We designed a test instead of a demo

The second time around we changed what we were asking for. Rather than asking the AI for an answer, we asked it to survive a serious attempt to prove its answer wrong. Three ideas did the work. All three are copyable, and none of them require anything exotic.

Blind testing. We gave Claude, Anthropic's Fable 5 model, our raw data and one goal: attribute every new merchant to the channel that acquired them. (New merchants, for us, means merchants who signed up and actually started doing business, and we exclude anyone who bounced within minutes.) We deliberately withheld our existing method. If it rediscovered our techniques on its own, our method was sound. If it found something better, we would learn something. It did both. It rebuilt our core technique from scratch, which is the strongest validation that method has ever received, and then it found things we had missed. It also identified structural flaws in the logic of the earlier method, which a second, independent AI later confirmed.

Neutral adjudication. Where the new answers disagreed with the old method's answers, we let neither side grade itself. We then ran a second, independent AI analysis, which pulled the raw evidence for every disputed merchant and ruled on it. It was told nothing about which method was which, or which one was ours. The verdict came back 16 to 1 in favor of the new approach across every case it could check. The one it lost was useful too: it taught us a rule we now apply everywhere, which is that no claim survives without reproducible evidence behind it.

An answer key. Before any of this could touch fresh data, we froze the validated results as an answer key and required the production system to reproduce them. It matched on more than 99% of cases: all but a handful, each one individually explained. The rule is simple. A system has to score close to perfectly on the past before it is allowed to grade the present.

And the process caught us. Twice, we proposed rules that sounded obviously right to us, and twice the evidence falsified them. Both falsifications are written down so nobody proposes them again. The point of the process was never that the AI is always right. It is that nobody gets to be right without evidence, including us.

What it found

The last click records the doorway, not the reason. The App Store surface always fires at the moment of install, because that is where installs happen. It is the door everyone walks through. Nothing was broken in the tracking. The logic was crediting the last hop, and the last hop is always the same place. When a merchant had a real upstream touch inside a sensible window, whether an ad click, a search, or an AI recommendation, that touch is what earned the credit. Under the old last-hop logic we had been systematically over-crediting the platform and under-crediting the marketing that created the demand in the first place. This is the bug we named at the start: the doorway collects the credit.

A third of our "paid" wasn't paid. Organic category browsing had been lumped in with paid category ads, inflating paid App Store performance by roughly a third. One mapping error, quietly flattering a channel for months.

ChatGPT is a real acquisition channel now. Our old reporting had no category for it at all, so AI referrals were being filed under social media. Once the category existed, we watched it move from essentially zero to a visible and growing source of new merchants inside a single quarter. If your attribution vocabulary predates AI assistants, you cannot see this happening to you even while it happens.

Fresh numbers are floors. The identity data that links anonymous visits to real accounts keeps accruing for weeks after a merchant signs up. That means the newest months always start low and rise for around two months afterwards. We now label the two freshest months as provisional, which has spared us the Monday-morning ritual of panicking at a wall of red week-over-week deltas that only ever meant "it is early."

Where it landed

Coverage went from 58% to 68% of new merchants attributed. Just as usefully, the remaining gap is decomposed and explained, so we know what is recoverable and what is genuinely dark, instead of shrugging at it.

The method runs as a system now. Twice a week, one command, a few minutes, refreshing a live dashboard the team uses. The AI analysis that used to evaporate the moment we closed it became the company's number of record, because it earned the position.

Three rules for putting AI near your metrics

Test blind. Withhold your existing method and see whether the AI converges on it or surpasses it. Convergence validates you. Divergence teaches you. Either outcome is worth more than a demo.

Never let anyone grade their own work. Not the AI, and not you. Use a neutral judge with access to the raw evidence, blind to whose answer is whose.

Freeze an answer key. No AI output becomes a number of record until a reproducible system can match the validated past.

What stood between us and useful AI analysis was accountability, not capability. Build the accountability, and the capability is finally usable.