How long should you keep customer data for attribution?
Attribution needs history and privacy law needs forgetting, and those two facts are in direct tension. A practical retention policy for marketing data: what to keep, what to aggregate, what to delete, and how to prove you did.
Attribution wants a long memory. A sale closing today might trace back to an enquiry from two years ago, and the only way to see that is to still hold the enquiry. Privacy law wants the opposite: personal data should exist for as long as it is needed and then stop existing. Both positions are correct, and any honest retention policy is a negotiated settlement between them rather than a victory for either.
The tension between attribution and privacy is real and it is resolvable, because the thing you need long-term is the answer, not the personal data that produced it.
That sentence is the whole article, and the rest is how to apply it without breaking your reporting or your compliance position. It is worth doing properly: retention is one of the few areas where a couple of hours of thought produces a policy that holds for years.
Why there is no number in the law
Data protection regimes ask you to keep personal data no longer than necessary for a specified purpose, and to be able to justify the period. They deliberately do not give a figure, because the necessary period genuinely differs between a takeaway order and a mortgage application.
This frustrates people who want a compliance checklist, and the frustration is understandable. But a fixed statutory period would be worse in both directions: too short for businesses with multi-year sales cycles, and far too long for businesses handling casual transactions. The principle-based approach puts the reasoning where the facts are.
What it means in practice is that your defence is the reasoning, not the number. A regulator asking why you hold three years of phone numbers will accept a clear answer about sales cycle length and reporting needs. They will not accept that three years felt about right, and they will be actively unimpressed by data being kept because nobody thought about it.
So the deliverable is a short written policy: what categories you hold, for what purpose, for how long, and what happens at the end. One page. The act of writing it usually reveals two or three things you are holding for no reason at all, which is the first benefit before anybody external ever asks.
The three tiers of attribution data
Attribution data separates cleanly into raw source files, matched working records, and aggregated results. They carry very different risk and very different value over time, and the whole policy follows from treating them differently.
| Tier | Contains | Value over time | Suggested retention |
|---|---|---|---|
| Raw uploaded files | Everything in the export, including unused columns | Falls to near zero once processed and reconciled | Days to weeks |
| Matched records | Identifiers, amounts, dates, source, match evidence | High while cycles are open; falls after | Two to three years, or one cycle plus a year |
| Aggregated results | Channel and period totals, no personal data | Permanent — this is your trend line | Indefinitely |
The raw file is the item most often mishandled, because it feels like the safe original that ought to be preserved. It is in fact the riskiest thing in the system: it contains every column the exporter happened to include, frequently including free-text notes and fields nobody reviewed, and it usually ends up in more places than the processed data does.
The aggregated tier is the one that makes long retention unnecessary. Once you have recorded that paid search produced eight hundred and ten thousand pounds in the third quarter, you never need the underlying rows again to state that fact. The trend line survives the deletion of everything that built it.
Setting the identifier retention period
Base the retention period for identifiers on your sales cycle plus a reporting margin, not on a round number. A business with a four-month cycle needs far less history than one with a two-year cycle, and choosing by cycle makes the period defensible.
The arithmetic is straightforward. Take the length within which the large majority of your deals close — the ninetieth percentile, not the average, because the average is dragged down by quick wins — add the longest period over which you report and compare, and add a modest buffer for late-arriving corrections. That figure is your period, and you can explain it in a sentence.
For most mid-market businesses this lands between eighteen months and three years. Enterprises with genuinely long procurement cycles land higher and should say why. Businesses with same-week purchasing should land far lower, and frequently hold years of data purely because deletion was never configured.
Write the period down alongside the reasoning and revisit it when the business changes. A company that moves upmarket lengthens its cycle, and a retention period set for the old business will start quietly destroying the evidence for the new one.
Deletion has to be a mechanism, not an intention
A retention policy that depends on somebody remembering to delete things will not be kept. The period has to be enforced by a scheduled process, and the process has to leave a record that it ran.
This is where most policies fail, and the failure is unglamorous: the policy exists, it is well written, and no code implements it. Two years later the data is still there, which is worse than having no policy at all, because you have documented an obligation and then demonstrably not met it.
The record matters as much as the deletion. Being able to say that a deletion job ran on a date, covered a defined set, and completed, converts a claim into evidence. Without it, the only proof that data was deleted is that it is not there, which is indistinguishable from never having collected it properly.
CloseRev records deletions in a ledger — what was deleted, when, and why — so a retention claim can be shown rather than asserted. The same ledger covers customer-requested deletion and scheduled expiry, because a regulator is likely to ask about both.
Erasure requests and the aggregate question
When an individual asks to be erased, their records go. Already-published aggregate totals that included them do not have to be recomputed, because those totals no longer constitute personal data once no individual is identifiable within them.
This is a point that causes real anxiety and should not. If it were otherwise, every historical report in every business would have to be restated whenever a single person exercised a right, which is neither required nor sensible. The obligation attaches to personal data, and a channel total is not personal data.
Where care is needed is with aggregates thin enough that an individual could be identified from them — a channel with two customers in a quarter, for instance. That is an edge case rather than the normal one, and it is handled by not publishing aggregates below a sensible threshold.
The practical requirement is that erasure actually reaches everywhere the data went, including raw files, exports somebody downloaded, and any backups within their own retention window. Mapping where the data goes is the harder half of honouring the request.
The processor and controller split
If you use an attribution tool, you are the controller and the tool is your processor. You decide the retention period; the tool is obliged to implement it, to delete on your instruction, and to return or destroy data when the relationship ends.
This matters when choosing a vendor, and it is worth asking directly rather than assuming. Can the retention period be configured, or is it fixed by the vendor. What happens at the end of a contract, and how long do they hold data after cancellation. Is deletion actually deletion, or is it a flag on a row that stays in the database forever.
The last question is more pointed than it sounds. Soft deletion is a perfectly reasonable engineering pattern and a poor answer to a data subject request, and plenty of systems built on it describe themselves as deleting data. It is a fair thing to ask about and the answer tells you a good deal about how seriously the vendor has thought about this.
Also confirm where the data is stored, whether it is isolated between customers, and whether it is encrypted at rest. None of those are retention questions strictly, and all of them will be asked in the same procurement conversation.
Minimise at the door, not in the warehouse
The cheapest retention decision is the one made before collection. Data you never took requires no policy, no deletion job, no breach notification and no answer to a subject access request.
Attribution needs remarkably little: an identifier, an amount, a date, and a source. Most exports carry far more than that because the person building the report ticked the columns that were already on a template, and every additional column becomes a permanent obligation attached to a field nobody will ever query.
The habit worth building is to review the field list once, at the start, and to treat additions as requiring a justification. It takes ten minutes and it removes entire categories of risk permanently. Free-text notes in particular should never enter a marketing system, because they contain whatever anybody typed, which over a few years is genuinely anything.
There is a secondary benefit that matters operationally. Narrow exports are faster to produce, faster to process, easier to reconcile and less likely to break, so the discipline that protects you legally also makes the monthly process less painful. That alignment is rare enough to be worth exploiting.
Backups are the part everybody forgets
Deleting a record from a live system does not delete it from backups, and backups typically have their own retention schedule measured in months. A deletion policy that does not account for them is incomplete in a way that is easy to overlook and awkward to explain.
The accepted approach is not to surgically edit backups, which is impractical and risks corrupting them. It is to document the backup retention period, ensure it is bounded, and confirm that a restored backup would be reprocessed against the deletion log so that erased records do not silently return to the live system.
That last clause is the one that gets missed. A restore after an incident can resurrect data somebody asked to have deleted, and unless the restore procedure includes reapplying deletions, nobody will notice. Writing the step into the runbook costs nothing and closes a genuine gap.
It is also worth knowing your backup window as a number, because you will be asked. Saying that erased data persists in backups for no more than thirty-five days and is then gone is a complete and satisfactory answer. Not knowing is not.
When retention and reporting genuinely conflict
Occasionally the honest answer is that a report cannot be produced because the underlying data was correctly deleted. That is an acceptable outcome, and pretending otherwise is what leads businesses to keep everything forever.
The scenario is familiar: somebody wants to re-run last year's analysis with a new segmentation, and the identifiers needed to do it are gone. The instinct is to conclude that the retention period was too short. Usually the correct conclusion is that the aggregate you should have kept was not specific enough, and the fix is forward-looking rather than a reason to stop deleting.
This is worth planning for deliberately. When designing what to retain in aggregate, think about the cuts you are likely to want later — by channel, by region, by product line, by new versus returning — and keep those breakdowns rather than a single total. Aggregates are cheap and contain no personal data, so being generous there costs almost nothing and removes most future regret.
What you should resist is the argument that any conceivable future analysis justifies indefinite retention of identifiers. That reasoning has no stopping point, which is exactly why the law does not accept it.
What good looks like on one page
A workable policy states the categories held, the purpose of each, the period, the deletion mechanism, and who owns the review. Five headings, one page, revisited annually.
- Raw source files: held for reconciliation only, deleted within a stated short window after processing.
- Matched records with identifiers: held for the sales cycle plus reporting margin, with the figure and its reasoning stated.
- Aggregate results: retained indefinitely, containing no personal data.
- Deletion: scheduled, automatic, logged, and covering exports and backups as well as the primary store.
- Review: an owner and a date, so the policy is revisited when the business changes rather than when a regulator asks.
The value of writing this down is not primarily defensive. It is that it forces a decision about what you actually need, and the answer is almost always less than what you are currently holding. Every category you can strike is risk you no longer carry and a question you no longer have to answer.
Keep the answers forever and the identifiers as briefly as the answers allow. Almost every retention problem in marketing data comes from confusing the two.
Questions people actually ask
- How long can I keep customer data for marketing attribution?
- For as long as you have a documented, specific purpose that requires it, and no longer. There is no fixed period in law. In practice, keeping raw identifiers for two to three years and aggregated results indefinitely satisfies almost every attribution need while limiting exposure.
- Does GDPR specify a retention period for marketing data?
- No. The GDPR requires that personal data be kept no longer than necessary for the purposes it was collected for, and that you can justify the period you chose. The absence of a number is deliberate: the regulation asks you to reason about it and record the reasoning.
- Can I keep attribution results after deleting the underlying data?
- Yes, and this is the key move that resolves most of the tension. Once a period is reported, the channel-level totals contain no personal data, so they can be retained indefinitely while the identifiers that produced them are deleted on schedule.
- What happens to attribution when a customer asks to be deleted?
- Their identifiers and records must go. Their contribution to already-published aggregate totals does not have to be recomputed, because those totals are no longer personal data. Delete the row, keep the history, and be able to show that this is what happened.
- How long should I keep raw uploaded CSV files?
- Much less time than you think. The raw file is the highest-risk artefact you hold, because it contains everything including the columns you did not need. Keeping it for a short reconciliation window and then deleting it is usually the right trade.
- Do I need to delete data if I am only a processor?
- You need to do what your controller instructs and what your contract commits you to, including deletion on request and at the end of the relationship. Being a processor changes who decides the retention period; it does not remove the obligation to have one and honour it.