Cookie preferences

Choose what we may use. Strictly necessary cookies keep you signed in and can't be turned off; everything else is your call and changes nothing about how the site works. You can change your mind any time from the footer.

A stack of blank white cards beside a single matching card, on a two-tone paper background.

One customer, three leads: deduplicate before you credit anyone

Photo by Kaboompics.com (opens in a new tab) on Pexels

The same person fills in a form from a Facebook ad, calls from a Google ad a week later, and opens the chat on your site the day before they buy. Your lead report counts three leads, every channel's cost per lead looks better than it is, and if you are not careful three channels claim one sale. Here is how to deduplicate on the person rather than the record, and why the rule you pick for the winner changes the answer.

Contents
  1. How one person becomes three leads
  2. Why your systems do not catch them
  3. Deduplicate on the person, not the record
  4. What duplicates do to cost per lead
  5. First touch or last touch: which lead gets the sale
  6. What deduplication cannot fix
  7. A short routine for every reporting period

One customer often arrives as three leads: a form, a call and a chat, sometimes from two campaigns. Count people, not records, by matching on normalized phone and email; credit each sale once, to one lead chosen by a rule you set in advance; and report how many records were duplicates, because that number changes every channel's cost per lead.

Every business that runs more than one channel has duplicate leads, and most of them do not know how many. The duplicates are not errors in the usual sense. Each record describes a real contact: somebody really did fill in that form, really did call that number, really did open that chat. The error is in what happens next, when the records are counted as if each one were a different person.

That counting error spreads into everything built on top of it. Lead volume looks higher than it is. Cost per lead looks lower. Conversion rates look worse, because the same number of customers is being divided by more leads. And when sales are credited back to sources without a rule, a single customer can be claimed by every channel they touched.

How one person becomes three leads

A duplicate lead is a second or third record for somebody who is already a lead. It appears whenever one customer contacts you through more than one route, or through one route more than once, and each contact creates its own record in whatever system received it.

The common patterns are all ordinary customer behavior:

  • Form, then call. Somebody submits a quote request from a Facebook ad, gets impatient, and calls the number on your site two days later. The form is in your CRM; the call is in your call tracker.
  • Two campaigns. The same person clicks a search ad in June and a retargeting ad in July and fills in the form both times.
  • Chat, then form. The chat widget captures an email address, the visitor leaves, and returns the next week to book online.
  • A lead marketplace and you. A homeowner requests quotes through a lead marketplace, then finds you directly and calls.
  • The same channel, twice. Somebody submits a form, hears nothing for an hour, and submits it again.

None of these is rare. The more channels you run and the slower your response, the more of them you get, which means the businesses with the most duplicates are often the ones spending the most on leads.

Why your systems do not catch them

Each system deduplicates on its own key, inside its own walls. A CRM folds a new contact into an existing one with the same email address; an ad platform counts one conversion per click; a call tracker counts callers. None of them sees the other systems' records, so the same person passes through every check as somebody new.

That rule is sensible for an email marketing tool and it leaves a gap exactly where local and sales-led businesses need it closed. A form fill carries an email. A phone call carries a phone number and nothing else. The CRM cannot know that the caller and the form-filler are the same person, because the only key it deduplicates on is missing from one of the two records. If you use HubSpot or any similar CRM, assume that calls and forms from the same person are sitting in different places.

"One per click" is deduplication within one platform and one click. The same person clicking two different ads, or clicking a Google ad and later a Facebook ad, is counted again by each. That is not a flaw in either platform; it is the boundary of what a platform can see. It is also why the sum of every platform's reported leads is always larger than the number of people who contacted you.

Every system deduplicates inside its own walls, on its own key. The only place all of a customer's contacts can be seen together is a file you assemble yourself, so that is where deduplication has to happen.

Deduplicate on the person, not the record

To deduplicate across sources, put every lead from every system into one file with a source column, normalize the phone numbers and email addresses, and treat any two records that share either identifier as the same person. Keep every original row; the deduplication is a decision about which row counts, not a deletion.

  1. Combine the sources. Export forms, calls, chats and marketplace leads into one file, one row per contact, with the date and a source or channel column on every row.
  2. Normalize the identifiers. Convert every phone number to one format, ideally E.164 with the country code, and lowercase and trim every email. Without this, "(555) 010-2030" and "+15550102030" look like two people. Phone normalization is the single step most likely to go wrong.
  3. Link on phone or email. Two rows that share a normalized phone or a normalized email belong to the same person. Chains count: if the form shares an email with the chat and the chat shares a phone with the call, all three are one person.
  4. Do not merge on names. "J. Smith" and "John Smith" might be one person or two. A name can support a match a person reviews; it should never create one on its own.
  5. Pick the counting row by a written rule. Usually the earliest contact. Mark it, keep the others, and record the rule in the report.

Do this outside your CRM, or at least without merging inside it. A CRM export should arrive as the system holds it, because a merge made to tidy up attribution usually discards the losing record's source field, and with it the evidence of which channel the person came through first.

What duplicates do to cost per lead

Cost per lead divides spend by records. Once duplicates are removed, it divides spend by new people, and the channels whose leads are mostly people who already knew you get noticeably more expensive. The ranking can change even though not a single dollar of spend did.

Take a month with 400 raw lead records across three sources. Deduplicated on phone and email, they turn out to be 310 people. Each person is assigned to the source of their earliest contact.

Raw and deduplicated cost per lead for one month, earliest contact wins
SourceSpendRaw lead recordsNew people (first contact)Cost per raw leadCost per new person
Meta Ads$7,000160140$43.75$50.00
Google Ads$9,000150105$60.00$85.71
Website chat and organic$09065——
Total$16,000400310$40.00$51.61

The arithmetic: $7,000 ÷ 160 = $43.75 and $7,000 ÷ 140 = $50.00; $9,000 ÷ 150 = $60.00 and $9,000 ÷ 105 = $85.71; across the whole month, $16,000 ÷ 400 = $40.00 against $16,000 ÷ 310 = $51.61. Ninety of the 400 records, 22.5%, were people who had already contacted you.

Meta's cost per lead rises by 14%. Google's rises by 43%, because 45 of its 150 records were people who had first arrived through Meta or the website. That is a common shape: search ads catch people looking for a business they have already heard of, and so collect a larger share of second contacts. It does not mean search is wasteful; it means its raw cost per lead was describing a different thing from Meta's.

A channel full of second contacts looks cheap per record and expensive per new person. Report both figures, and the share of records that were duplicates, so nobody cuts the channel that finds people to fund the one that collects them.

First touch or last touch: which lead gets the sale

Once one person has several leads, a sale needs a rule to pick one of them. First touch credits the earliest contact, the source that brought the person to you. Last touch credits the latest contact before the sale. Both are defensible; they answer different questions, and they can move a large share of revenue between channels.

Continue the example. Of the 310 people, 40 bought, at an average of $3,000. Credit each sale to the person's earliest lead and then to their latest lead before the sale, and the channels trade places:

The same 40 sales credited by first touch and by last touch
SourceSales, first touchRevenue, first touchSales, last touchRevenue, last touch
Meta Ads18$54,00011$33,000
Google Ads14$42,00017$51,000
Website chat and organic8$24,00012$36,000
Total40$120,00040$120,000

Total revenue is $120,000 under both rules, and that is the property that matters most: each sale is counted once. What moves is $21,000 of credit — seven sales — from Meta to Google and the website. Under first touch, Meta returns $54,000 on $7,000, or 7.7 to one. Under last touch it returns $33,000, or 4.7 to one. Neither figure is wrong. One describes who found the customer; the other describes who was last in the room.

We think first touch is the better default for a business deciding where to spend, because the expensive problem is finding new customers, and the last contact before a sale is disproportionately the main phone line, the chat window and branded search — routes people use once they have already decided. We set out the wider argument in our post on attribution models. What is not defensible is choosing after seeing which rule flatters the channel you wanted to keep.

The rule you must never use is no rule at all: crediting the sale to every lead that shares the customer's phone or email. In this example that hands out credit for more than 40 sales, and every channel's return on spend is overstated at once. It is the same arithmetic that makes the platforms' reported totals add up to more than your revenue.

First touch and last touch move credit between channels; they never change the total. A report that credits one sale to three leads has changed the total, and that is the one mistake that makes every other number in it wrong.

What deduplication cannot fix

Matching on phone and email finds the same person when they used the same identifiers. It cannot link a customer who called from a work phone and bought with a personal one, or a household where one partner made the inquiry and the other signed. Those stay separate, and the honest report says so.

This limit is worth stating plainly because the temptation is to close it with guesses: merge on surname and ZIP code, link records within the same street, treat similar names as one person. Each guess makes the duplicate count look more complete and makes the attribution less true. A sale that cannot be linked to a lead on a real identifier is better left unattributed than credited to a lead that might have been somebody else.

  • Ask for the phone number on every form. It is the identifier most likely to appear in both the lead and the sale.
  • Capture the caller's number on every call. A call log without numbers cannot be deduplicated or matched.
  • Record which identifier made each link. An email match and a phone match are both strong; anything weaker belongs in a review queue, not in the totals.
  • Watch the [match rate](/blog/match-rates). If it falls after a change to a form or a phone system, an identifier has stopped arriving.

CloseRev takes one lead file per import, so forms, calls and chats go into a single file with a source column. For each sale it looks for leads with the same email first and, failing that, the same phone number in E.164, then credits the earliest of those leads; that is first touch, and it is the only deduplication it does. It does not merge on names or households, and sales it cannot match stay in Direct / Unknown.

A short routine for every reporting period

Deduplication is not a one-off cleanup. Done every period, it takes minutes and keeps three numbers honest: how many new people each channel brought in, what each of them cost, and which single lead each sale is credited to.

  1. Export every lead source for the period, plus at least one sales cycle before it, into one file with a source column.
  2. Normalize phones and emails, and link records that share either one.
  3. Count new people per channel by earliest contact, and report the share of records that were duplicates.
  4. Credit each closed sale to one lead by the rule you chose in advance, and check that credited revenue does not exceed total revenue.
  5. Leave sales with no linked lead in their own line rather than spreading them across channels.

The output is smaller than the platforms' figures, and it is supposed to be. Every lead in it is a person, every sale is counted once, and every number can be traced back to the rows that produced it.

Questions people actually ask

What is a duplicate lead?
A lead record for a person who is already a lead. It happens when one customer reaches you more than once — a form, then a call, then a chat — or through two campaigns, and each contact creates its own record. The records are real; the extra people are not. A lead report that counts records rather than people overstates demand and flatters every channel's cost per lead.
How do you deduplicate leads?
On the person, using identifiers that belong to them: a normalized phone number and a lowercased, trimmed email address. Two records that share either one are the same person. Names are too unreliable to merge on automatically. Keep every original record, mark which one counts, and write down the rule you used to choose it.
Should duplicate leads count toward cost per lead?
No. Cost per lead should be spend divided by the new people a channel brought in, not by the records it produced. A channel whose leads are mostly people who already contacted you through another source is cheaper per record and more expensive per new customer, and only the second figure is useful for budgeting.
Does first touch or last touch handle duplicates better?
Neither handles them better; they answer different questions. First touch credits the source that brought the person to you, which is right for judging which channels create customers. Last touch credits the final contact before the sale, which tends to favor branded search, chat and the main phone line. Pick one in advance and say which.
Should I merge duplicate contacts in my CRM before exporting?
Not to improve attribution. A merge inside the CRM usually keeps one record's source and discards the other's, and it cannot be undone cleanly. Export the records as they are, deduplicate where the rule is visible, and keep the originals so anybody can check which lead a sale was credited to and why.

See it on your own numbers.

Two exports and a few minutes. 14 days free, no card, nothing to install.