Most free lead samples prove nothing, and it is not usually the vendor’s fault. The vendor sends five hundred rows. The floor manager hands them to the best rep, who dials them once, on a Tuesday afternoon, with a script tuned for a different file, and reports back that they “felt good” or “felt bad.” Nobody wrote down what the current list did that same afternoon. Two weeks later the sample is a memory and the decision gets made on price.
A sample can be a real test. It has to be designed like one — a single variable, everything else held constant, an outcome defined before the first dial — and it has to be small enough in ambition to measure the one thing a sample can measure. This page is that method, written as a neutral procedure that works on any vendor, including us. Our own free test is at the bottom as the worked example.
Disclosure: Scout Data is our product. We offer a free 1,000-record test to call centers, and this page is the method we would like to be tested with. It is also the method that would catch us if the file were bad.
Get a sample that can fail
A sample the vendor chose is a sample of the vendor’s best work, and an audit of that tells you how good the vendor is on a good day. Ask instead for rows produced by the vendor’s normal process against the spec you would actually buy: your geography, your segment, your filters, your suppression file applied. If a vendor cannot build to a spec for a sample, they cannot build to it for an order either, and you have learned the first thing.
Size it to the test, not to generosity. The sample has to be dialed three times inside two weeks by the reps you have, so a floor of five can work a thousand records comfortably and a floor of thirty should ask for more — a sample nobody finishes measures the first pass only. And take it in the same format you would take an order: one number per row, line type, scrub date, the pitch fields. A sample that arrives as a spreadsheet of ten columns you will never see again proves nothing about the file that would ship.
Hold everything else constant
The only thing that should differ between the sample and the list you dial today is the file. Everything the floor controls stays fixed:
- Same reps. Not the best rep on the sample and the floor on the incumbent. Rotate both files through the same people; if the floor is large, assign every rep a share of each.
- Same hours. Interleave the two lists inside each shift — alternate hour blocks, or split each block — rather than sample in the morning and incumbent in the afternoon. Hour of day moves contact rates on its own — the one contact rate this site can source, at the door, swings by a third across the day — and a test that confounds it with the file has measured the clock.
- Same script, including the opener. If the sample carries a fact the incumbent does not — a permit date, a footprint — the honest design is to let the script use it, because that is the product; but note that the test is then measuring file plus opener, and say so in the write-up.
- Same dialer settings and caller IDs. Pacing, attempt limits, the number pool, the calling window. A new campaign on a fresh number pool will outperform a tired one regardless of the file.
- Same disposition codes, and specifically these four as separate outcomes: right party reached, wrong party reached, bad or non-residential number, no answer. A code set without a distinct wrong-party outcome cannot run this test at all.
Three passes over one to two weeks
One pass measures who happened to be near their phone on one day. The second and third recover the households the first missed and let the hour-of-day pattern average out; they also let a rep who had a bad Tuesday have a normal Thursday. Two weeks is long enough for three passes at a sensible attempt cadence and short enough that phone data does not decay inside the window, so the difference you measure is about matching and not about age. Stop at three: a fourth pass on a sample starts measuring the floor’s persistence, which is a fine thing to know and not what the sample was for.
Measure two rates, log four outcomes
Two numbers come out, and both are comparisons against the incumbent list dialed the same way in the same fortnight.
- Connect rate — connects as a share of dials. Define “connect” once, in one sentence, before the test, and apply it to both files; the four definitions floors use interchangeably are in what is a good contact rate.
- Wrong-number rate — wrong party plus bad number, as a share of connects or of dials. Pick one denominator. As a share of connects it isolates how the file was matched; as a share of dials it folds in how old the numbers are. Both are useful; neither is comparable to the other.
The four dispositions are the whole instrument. A floor whose codes collapse “wrong person” into “not interested” has no way to tell a file that reaches tenants from a file that reaches owners who decline, and those are opposite results.
Ignore deals
The strongest temptation in the test is to count appointments, and the arithmetic is why not. Take a thousand records, three passes: 3,000 dials. Suppose 15% connect — 450 connects, enough to pin the connect rate to within a few points either way. Now suppose 3% of connects become a set: 13 or 14 appointments. At 1%, four or five. The gap between those two outcomes is nine appointments, and nine is well inside what one strong rep, one bad week or one promotion moves on its own. Held appointments are a fraction of sets; closed deals a fraction of those, decided by the closer and the month. The sample cannot see any of it.
The rates in that paragraph are illustrative placeholders chosen to make the arithmetic legible, not benchmarks. Substitute your floor’s own numbers; the conclusion about sample size does not change.
So decide on the two rates alone. A vendor who asks to be judged on deals from a thousand rows is asking to be judged on noise, and a floor manager who accepts is agreeing to buy on a feeling with a spreadsheet attached.
Reading the result
| Against the incumbent | What it usually means | Decision |
|---|---|---|
| Connect up, wrong numbers down | Better matched and fresher; the file did what it claimed | Buy, then audit the first paid delivery |
| Connect up, wrong numbers up | More numbers ring, more of them ring the wrong person — often several numbers per row, address-matched | Ask how the phone was tied to the owner; usually no |
| Connect down, wrong numbers down | Cleaner file, fewer lines answering — check line type and the calling window before blaming the data | Re-run one pass with the timezone fixed, then decide |
| Both about the same | The files are equivalent on the phones | Price, terms and the pitch fields decide it |
Write the result up as two rates, two files, one fortnight, the reps and hours named, and keep it. A vendor who wins the test in September can be held to the same rates in March.
The worked example, in full
The free 1,000-record test is this method with a vendor on the other end. A fifteen-minute call establishes the market, the dialer and who the floor sells to. We build a thousand owner-occupied records to that pitch — one active mobile per home, matched by name to the owner on title, scrubbed against the do-not-call registry before delivery — and ship the file within 24 hours. The floor dials it three times over a week or two on its own dialer, next to the list it uses today, with the check-in call booked before the batch ships so the result has a date on it. On that call we put connect rate and wrong-number rate side by side. Deals are not part of the comparison, for the reason above. If the batch loses, the floor keeps the records.
That is one vendor’s version of the test. The method belongs to the buyer, and it works on any file.
After the sample
A sample that wins earns a first paid delivery, not a contract. Draw your own rows from that delivery and run the four post-delivery tests — match-back, duplicates, disconnects, wrong-party — in how to audit a lead vendor, because the delivered file is the vendor’s average work where the sample was, at best, their normal. Then stamp every row with its source and keep the two rates per source, per month, in the ledger described in how to measure lead source ROI. The buying sequence around all of this — what to get in writing before and after the test — is in how to buy call center leads.
Frequently asked questions
How do I test a free lead sample properly?
Run it as an experiment with one variable. Get a sample built to your real spec and geography, dial it with the same reps, in the same hours, on the same script and dialer settings as the list you use today, with that list running alongside, three passes over one to two weeks. Log right party, wrong party and bad number as separate outcomes. Then compare connect rate and wrong-number rate between the two files. Nothing else in the test is a result.
How many records should a free lead sample have?
Enough to dial three times in two weeks with the reps you have, and enough that the connect rate has a few hundred connects behind it. A thousand records is a workable size for a small floor; a larger floor can ask for more, but the constraint is finishing all three passes, not the size of the file. A sample nobody finishes dialing measures the first pass only, and the first pass overstates the difference between any two lists.
Why should I ignore deals when testing a lead sample?
Because a sample is large enough to measure connects and far too small to measure sales. A thousand records dialed three times might produce a few hundred connects, which pins a connect rate to within a few points, and a handful of appointments, which pins nothing — the difference between four and nine sets is noise, and a closed deal depends on the closer and the month more than on the file. Judge the file on what it controls: whether the phone rings the right person.
What is a wrong-number rate?
The share of dials, or of connects, that reached a line the named homeowner does not answer — a disconnected number, a business, a tenant, a previous owner, a relative. Pick one denominator and keep it: as a share of connects it measures how the file was matched; as a share of dials it also folds in how old the phone data is. Either is fine as long as both files are measured the same way.