Your “Direct / Unknown” bucket is the most honest number on the page
A big unattributed bucket is not a broken report. It is a measurement — and there are only four things it can be. Here is how to read it, how to shrink it, and why any tool that hides it is lying to you politely.
The first reaction to an attribution report showing 38% Direct / Unknown is almost always that the report is broken. Occasionally that is true. Far more often, that bucket is the most honest number on the page, and the useful question is not how to make it smaller but what it is made of.
The unattributed bucket is not the tool failing to answer. It is the tool declining to make something up.
This matters more every year, because the amount of activity that leaves no trace is growing. Search itself has become a place where things happen without a click.
Somebody can read your answer in an AI summary, remember your name, and call you three weeks later from a number your ad platform has never seen. That sale is real, that marketing worked, and no click connects them. Multiply it across a year and you have a large, growing, genuinely unattributable slice of revenue — which is a measurement problem, not a marketing failure.
What Direct / Unknown actually means
In revenue attribution, Direct / Unknown means the sale did not match any source record you supplied. It is a statement about your data coverage, not about the customer's behaviour — and it is not the same thing as “Direct” in web analytics, which only means no referrer header was present.
That distinction sounds pedantic and is the source of a great deal of wasted effort. Teams see a large Direct figure, assume it means “people typing our URL”, conclude their brand is stronger than they thought, and spend a quarter congratulating themselves. The truthful reading is usually duller: a system was not exported.
It can only be four things
Every unattributed sale falls into one of four categories: a source you are not exporting, a person your data cannot match, a sale with genuinely no marketing source, or a source that predates your data. Only two of them are fixable, and knowing which is which is the whole job.
| Cause | What it looks like | Fixable? | First move |
|---|---|---|---|
| A source you are not exporting | Referrals, walk-ins, repeat customers, calls to a main line that bypasses the tracker | Yes — and it is usually the cheapest fix | List every system that touches a lead. Export the ones that were missed |
| A person your data cannot match | Personal email one side, work email the other; a different phone; a company name where a person was expected | Yes | Normalise identifiers on both sides before blaming the method |
| Genuinely no marketing source | Word of mouth, reputation, the customer's brother-in-law | No — and it should not be | Measure it deliberately: add a source field at intake |
| A source older than your data | An eighteen-month-old lead closing this quarter | Partly | Extend the source export further back than the sales window |
The first row is the commonest by a wide margin, and it is the one people look at last. In most businesses there is at least one system nobody thought of as a marketing system — a booking tool, a main switchboard log, a partner portal, the inbox that quotes go out from — quietly producing a meaningful share of revenue with no channel attached to it.
Why redistributing it is worse than showing it
Proportional redistribution spreads unmatched revenue across the channels that did match. It inflates every channel by the same factor, cannot be seen by the reader, and makes it impossible to tell which part of a number was measured and which was assumed.
Consider a report with £1,000,000 of revenue, £620,000 of it matched. Redistribution takes the £380,000 nobody could attribute and shares it out in proportion. Google Ads goes from £250,000 to £403,000. Everything looks better. Nothing got better.
| Channel | Measured | After redistribution | Overstated by |
|---|---|---|---|
| Google Ads | £250,000 | £403,000 | £153,000 |
| Meta | £180,000 | £290,000 | £110,000 |
| £120,000 | £194,000 | £74,000 | |
| Referral | £70,000 | £113,000 | £43,000 |
| Unattributed | £380,000 | £0 — hidden | — |
Now imagine defending the second column to a CFO who asks how Google Ads produced £403,000 when the CRM shows £250,000 of matched deals. There is no good answer, because the honest one is “we assumed the unmatched revenue behaved like the matched revenue” — an assumption with no evidence behind it, and one that is wrong in a specific, predictable direction. Unmatched revenue skews heavily toward referral and repeat business, which are precisely the channels you did *not* pay for.
Redistribution does not just inflate the numbers. It systematically transfers credit from the channels you get for free to the channels you are paying for.
Any attribution model that cannot show you its unattributed bucket has one. It just isn't telling you where it went.
How to shrink it, in order of return
Widening the source export moves the number more than every other fix combined. Then normalise identifiers, then extend the date range, then capture a source at the point of sale.
1. Widen the source export
Write down every system through which a customer could conceivably first make contact. Not every marketing system — every system. The booking widget, the main phone line, the WhatsApp number on the van, the partner referral form, the inbox where quotes go out. Then check which of those you are actually exporting.
In practice this single step often halves the unattributed bucket, because the missing system turns out to be the one producing the largest deals slowly and quietly.
2. Normalise identifiers on both sides
- Phone numbers to E.164, using a real library rather than a regular expression.
- Emails lowercased, plus-addressing stripped, dot-handling decided deliberately per provider.
- Shared inboxes (info@, admin@, accounts@) excluded rather than matched — one address against forty sales is a bug wearing a match's clothing.
This is usually worth several percentage points on its own, and it is entirely free.
3. Extend the source date range
If your sales export covers Q2 and your leads export covers Q2, every long-cycle deal is unattributable by construction. Pull the source side back at least one full sales cycle further than the revenue side. If you do not know your sales cycle, that is worth finding out before anything else on this list.
4. Capture a source at the point of sale
One dropdown on an intake form — “How did you hear about us?” — permanently fixes the referral problem, which is otherwise unfixable by any amount of data engineering. It is low-tech, slightly unfashionable, and more effective than most attribution software.
The bucket is growing, and it is not your fault
Unattributed revenue is structurally increasing because more of the buying journey now happens where no click is recorded — AI answers, private messages, podcasts, communities and search results that never send a visit. Attribution did not get worse; the observable surface got smaller.
Ten years ago a considered purchase produced a searchable trail: a search, a click, a few pages, a form. Today the same purchase can involve an AI summary that quotes you without linking, a screenshot forwarded in a group chat, a recommendation in a Slack community, and a phone call from a number that has never touched your site. Every one of those is real marketing doing real work, and not one leaves a referrer.
This is what people mean by dark social, and it is why a rising Direct figure is often a sign that your marketing is working in places you cannot see rather than a sign that it has stopped working. The trap is responding by defunding the channels you *can* measure, which is the one action guaranteed to make the picture worse.
A growing unattributed bucket is usually evidence of reach you cannot see, not evidence of waste. Defunding what you can measure to chase what you cannot is how good marketing gets cut.
How to present this to a CFO without losing the room
Lead with the matched revenue and the method, state the unattributed share as a coverage figure with a plan attached, and never let the first mention of the gap come from the person reviewing your budget.
The failure mode is predictable. A marketing lead presents a confident attribution deck, someone in finance notices the numbers do not reconcile with the CRM, and the entire report — including the parts that were correct — loses credibility in about forty seconds. It is very hard to get it back in the same meeting.
The alternative costs one slide and buys a great deal. Something like: *of £1.0m closed last quarter, we traced £620k to a source. Google Ads produced £250k of that, at £58k spend. The £380k we could not trace is mostly the booking system, which we are not exporting yet — that is next month's fix, and I expect the traced share to go from 62% to about 80%.*
- You have given a number, a method, a known limitation and a dated plan.
- You have pre-empted the only hard question, which removes its force entirely.
- You have made the next quarter's improvement measurable, which is how a marketing budget stops being an argument and starts being a forecast.
The counter-intuitive part is that admitting the gap makes the rest of the report *more* persuasive, not less. Everyone in that room knows attribution is imperfect. The person who says so first is the one being trusted with the estimate.
What a healthy number looks like
There is no universal target. A healthy unattributed bucket is one that shrinks as coverage improves and whose remainder you can explain. The direction and the explanation matter more than the percentage.
| What you see | Most likely meaning | What to do |
|---|---|---|
| Under 10% | Either exceptional data discipline — or redistribution happening somewhere | Ask the tool where unmatched revenue goes. If it cannot say, assume the worst |
| 20–35% | Normal for a business with real offline and referral revenue | Widen exports; expect steady improvement rather than a cliff |
| 40–60% | Almost always a coverage problem, not a matching problem | Find the system that is not being exported before touching anything else |
| Over 60% | The source export is probably one system rather than all of them | Stop optimising the match logic. Go and find the data |
A business that knows 22% of its revenue arrives through word of mouth has something genuinely useful: a reason to invest in the referral programme, and a defensible explanation for why paid channels look smaller than the ad platforms claim. That is a better position than a tidy 4% that nobody can account for.
Three mistakes that make the bucket look worse than it is
Most inflated unattributed figures come from three avoidable errors: matching on the wrong window, treating shared inboxes as people, and counting revenue that was never attributable in the first place.
Matching a quarter against a quarter
If the sales export covers April to June and the leads export covers April to June, every deal with a cycle longer than a few weeks is unattributable by construction. The lead that produced June's largest invoice was created in February and is simply not in the file. This one error alone routinely accounts for ten to twenty points of apparent unattributed revenue, and it is fixed by changing a date on an export.
Letting one address represent forty customers
Shared inboxes — info@, accounts@, admin@, the address on the side of the van — appear against many unrelated sales. Matching on them produces confident nonsense: one lead credited with a year of revenue, and every other sale in that group left unattributed because the wrong record already claimed it. Exclude them explicitly, on both sides, and treat the affected sales as unmatchable rather than as one enormous customer.
Counting revenue that never had a marketing source
Renewals, contract expansions, upsells to existing accounts and inter-company transfers frequently sit in the same sales export as new business. None of them was produced by a campaign, and including them in the denominator makes the attributed share look far worse than it is. Separate new business from existing-customer revenue before calculating anything — a 40% unattributed figure across all revenue is often 18% across new business, which is a completely different conversation.
| Fix applied | Unattributed share | Change |
|---|---|---|
| Starting point | 41% | — |
| Extend source export back one sales cycle | 29% | −12 points |
| Exclude shared inboxes from matching | 26% | −3 points |
| Report new business separately from renewals | 17% | −9 points |
None of those three required new software, a new vendor, or a single line of tracking code. They required looking at what was in the file.
The opinion, since you have read this far
The attribution industry has spent a decade optimising for reports that look complete rather than reports that are true, and buyers have rewarded it for doing so.
It is not really the vendors' fault, or not only. A seamless dashboard wins the demo. A dashboard with a 38% hole in it loses to the one that shows 4%, even when the 4% is manufactured — because in a forty-minute evaluation nobody asks where the rest went. The incentive runs the wrong way, and it runs that way all the way down.
So the useful skill is not picking the tool with the best-looking numbers. It is asking one question in the demo — *show me the unattributed bucket* — and watching what happens next. The answer tells you more about the product than the rest of the call combined.
Ask any attribution vendor to show you their unattributed bucket. The pause before the answer is the review.
CloseRev reports unattributed revenue as its own line and never redistributes it. Every matched sale is auditable — which lead, which channel, and on what basis — so any channel figure can be traced back to the individual sales behind it.
Questions people actually ask
- What does Direct / Unknown mean in an attribution report?
- It is revenue that could not be tied to any marketing source. In a person-based attribution report it means the sale did not match any lead, call or campaign record you supplied — not that the customer typed your URL directly.
- Is a high Direct / Unknown percentage bad?
- Not necessarily. It is a measurement of coverage, not of performance. A bucket that shrinks as you widen your source exports is healthy; a suspiciously small one usually means the tool redistributed the unmatched revenue across other channels instead of reporting it.
- What is a normal amount of unattributed revenue?
- Commonly between 20% and 50%, depending on how completely the lead side is exported and how consistently phone numbers and emails are recorded. There is no universal target, and anyone quoting one without seeing your data is guessing.
- How do I reduce unattributed revenue?
- In order of return: widen the source export to include systems that were left out, normalise phone numbers and emails on both sides, extend the source date range further back than the sales range, and capture a source field at the point of sale for anything arriving without one.
- Should unattributed revenue be redistributed across channels?
- No. Proportional redistribution inflates every channel by the same factor, is invisible to the reader, and always flatters the tool doing it. It converts a measurement into an assumption without saying so.
- Is Direct traffic the same as unattributed revenue?
- No, and conflating them causes real mistakes. In web analytics “Direct” means no referrer was present. In revenue attribution, unattributed means no source record matched the customer. A sale can be genuinely direct, genuinely referred but unrecorded, or simply unmatchable — and those need different responses.