You bought the list. The reps say it is bad. The vendor says the match rate was 92%. Both of those can be true, and the way to find out which one matters is not another call with the account manager — it is four tests you can run yourself, on a sample you chose, in an afternoon.
This page is the procedure. It is written for the buyer who already has a bad feeling and needs to turn it into a number, a credit claim, or a cancellation. Run it on the first delivery from any new vendor, and quarterly on the ones you keep.
Disclosure: Scout Data is our product and we sell homeowner data, which means this procedure gets run on us. We would rather be tested this way than trusted on a match rate, and nothing below is a test we would fail differently from anyone else.
Before the tests: draw your own sample
Every vendor will offer a sample. Decline it, or take it and set it aside. The rows a vendor chooses to show you are the rows they are proudest of, and an audit of those tells you how good the vendor’s best work is, which was never the question.
Instead, pull a random few hundred rows from the file that was actually delivered, after your own standardisation and before anything else touches it. Random means a method you can describe — every n-th row after shuffling, a random-number column sorted, anything that is not “the first three hundred” or “the ones for the county we care about.” Record the draw date and keep the row IDs. Everything that follows is a claim about this sample, and the sample has to be defensible.
Test 1 — Match-back
Join the sample against everything you already hold, on a property key rather than a name or a phone: your CRM, every prior delivery from every vendor, your suppression file, your customer list. Names vary in spelling and phones get reassigned; a normalised property identifier does neither.
Three shares fall out, and each one means something different:
- Already in your CRM or a prior delivery. Rows you paid for twice. A legitimate vendor dedupes against a suppression file you supply; a vendor who was not offered one has an excuse, once.
- On your suppression file. Rows you should never have been sent — prior opt-outs, complaints, customers. If you supplied the file and these came through anyway, the vendor’s suppression step is not connected to their delivery step, which is a compliance finding as much as a billing one. Building a suppression list covers the negative-control test that proves yours is wired in.
- Already a customer. A second rep is about to pitch someone who already bought. Small share, disproportionate damage.
Match-back is the fastest test and the one vendors least expect, because it uses data they have never seen. It also answers a question the vendor cannot: what fraction of this file was new to you.
Test 2 — Duplicate check
Now look inside the file. Duplicates are the same household appearing more than once, and row-level dedupe catches almost none of them, because real duplicates are 124 N Maple St and 124 North Maple Street, or the same parcel listed once under an owner’s name and once under a trust.
Dedupe the whole delivery — not just the sample — at household level on the normalised property key, and count the collapse. Then check the phones: the same number attached to several addresses is a different kind of duplicate, and usually a sign the number was matched to an area rather than a person. A file where one mobile appears against four properties is a file where that mobile belongs to none of them. List hygiene for call centers has the standardise-then-dedupe order this depends on.
Test 3 — Disconnect rate
Dial the sample, or run it through a line-status check, and count the numbers that are disconnected, out of service, or ring to a fax or a business. Do it within days of delivery, because phone data decays and a disconnect found in month three is arguable in a way a disconnect found in week one is not.
Two things to record beside the rate. First, line type: whether the file told you each number was a mobile or a landline, and whether it was right. Line type is a compliance input for automated dialling, and a file with partial or wrong line-type coverage has a problem larger than its disconnect rate. Second, the vendor’s replacement policy: whether a dead number is credited, replaced with a new one, or simply yours now. Confirm which before the first order; the audit only turns into money if the policy exists.
Test 4 — Wrong-party rate
The test that separates a good vendor from a good match rate. When the dialled sample connects, log whether the person who answered is the owner the row named — as its own disposition, distinct from “not interested” and from “bad number.” A working number that reaches a prior resident, a tenant, or a relative counts as a contact on every report and converts at nothing.
A high wrong-party share on a file with a low disconnect rate is the signature of address-anchored matching: the vendor attached whichever number their data associates with the house, rather than resolving who owns it and matching a number to that person. The two methods produce similar match rates and very different floors. Skip tracing accuracy covers why the match rate on the invoice does not tell you which one you bought.
The scorecard
| Test | Record | What a bad reading means |
|---|---|---|
| Match-back | % already owned · % on suppression · % customers | Paid twice; suppression not wired in; a compliance exposure |
| Duplicates | % collapsed at household level · phones on multiple rows | Billed rows that were the same household; area-matched phones |
| Disconnect | % dead or non-residential · line-type accuracy | Stale phone data; a compliance problem if line type is wrong |
| Wrong-party | % of connects not the named owner | Address-anchored matching; the floor will convert at nothing |
| Net | Rows that passed all four ÷ rows billed | The share of the invoice that could ever have produced a sale |
The last row is the one to carry into the conversation with the vendor, because it converts the invoice into a price per usable row — the same unit the skip tracing cost calculator uses to compare quotes before purchase. We are not going to tell you what a good reading is on any row, because it depends on what you supplied, what the contract promised and what the file cost; the point is that you now have a number where you had a feeling.
What to do with the result
- Claim the credits the contract names. Send the row IDs, the draw date and the test each row failed. A claim with rows attached gets paid; a claim with adjectives gets a call.
- Fix what was yours. If the match-back failed because you never supplied a suppression file, that is on you, and the next order should carry one. If duplicates were across your own prior deliveries, the vendor could not have known.
- Re-run on the next delivery. One bad file is a question; two is a pattern. A vendor whose second delivery scores the same after the first was disputed has told you what the third will look like.
- Put the net figure in the ledger. Rows that passed all four tests are the denominator for everything the file goes on to produce. How to measure lead source ROI is where that number ends up.
Run it before the purchase, too
Everything above works as a pre-purchase test with one change: ask for a paid trial delivery of a size you choose, from a geography you choose, rather than a free sample the vendor chooses. Then draw your own rows from it and run the four tests before the volume order. Bulk skip tracing covers how to structure that week-long trial, and how to buy solar leads has the questions to get answered in writing beforehand — exclusivity, source, replacement policy — so the audit has something to hold the vendor to.
Frequently asked questions
How do I know if a lead vendor is selling bad data?
Run four tests on a sample you drew yourself: match the rows back to what you already own, check for duplicates inside the file and across prior deliveries, dial or line-check the numbers for disconnects, and log wrong-party contact as its own outcome. Each produces a percentage you can put beside the invoice. A vendor is selling bad data when those percentages, not their match rate, say so.
How big a sample do I need to audit a lead list?
Enough to dial in an afternoon and still see a pattern — a few hundred rows drawn at random from the delivery is usually the practical size. The sample must be yours: pulled by you from the file that arrived, not the sample the vendor sent to win the account. A vendor’s sample is their best rows; yours is their average rows.
What is a match-back?
Joining the delivered rows against records you already hold — your CRM, prior deliveries from any vendor, your suppression file, your customers — on a property key rather than a name. The share that matches is the share you paid for twice, plus the share you should never have been sent at all. It is the fastest of the four tests and the one vendors least expect.
Can I get a refund for bad leads?
Usually only for the categories the contract names — disconnected numbers, duplicates, wrong-party — and only if you can show them. That is what the audit is for. A credit claim that says “these were bad” gets argued; one that says “212 of 300 sampled rows, drawn on this date, failed these tests, and here are the rows” gets paid. Confirm the credit terms before the first order, not after.