Most data vendor evaluations end up measuring noise. A prospect types ten names into a search box, gets results for eight of them, and calls that an 80% match rate. Then they sign a contract, run their first real batch, and discover the actual coverage on their specific use case is 45%.
The vendors who benefit from poor testing are the ones who can't survive good testing.
We can. And here's the methodology.
The problem with how most people test
The typical evaluation goes like this: a sales rep gives you a trial, you search for a few famous executives or companies you already know, the results come back looking great, and you move forward. You’re facing three failure modes with this approach:
Famous names are not representative. Bill Gates, Satya Nadella, and the CEOs of Fortune 500 companies are among the most data-rich profiles on the internet. Every enrichment vendor has them. Testing famous names tells you nothing about coverage on your actual prospect list.
Random names aren't representative either. If you pick random records from your head, you'll unconsciously pick memorable ones like people you've encountered recently, companies you recognize, industries you know. Your real list is more diverse and more uneven.
One vendor at a time is not a comparison. If you test Vendor A this week and Vendor B next week on different records, you're actually just comparing test conditions.
Good testing fixes all three of these, and it's not complicated.
Step 1: Build a test set from data you already trust
The foundation of a good evaluation is a test set with known ground truth. These are records where you already know the correct answer, so you can evaluate not just match rate but accuracy.
The best source is your own CRM. Pull 200–300 contacts where you have high-confidence data: current job title, company, email, LinkedIn URL. These are people your team has actually spoken to, verified manually, or confirmed recently.
Why this works:
You're testing on records that look like your real use case
You can check accuracy, not just match rate
You control the quality of the benchmark
Diversify the set intentionally. Don't just pull your most recent 200 contacts, that's just a convenience sample. Try:
Region (don't make it all US/UK if your real list isn't)
Company size (mix startups, SMBs, enterprise)
Seniority level (IC, manager, VP, C-suite)
Industry (include at least 5–6 verticals)
Data age (include records last touched 6–18 months ago to test freshness)
Aim for 200+ records. Fewer than 100 gives you too little signal to distinguish real performance from luck.
Step 2: Decide which input identifier you'll actually use
Different identifier types produce very different match rates. The delta between them is greater than most vendors advertise.
From highest match rate to lowest:
LinkedIn URL: the most reliable identifier for people data. If you have the LinkedIn URL, use it. Match rates are meaningfully higher than any other input.
Professional email: strong for company lookup, variable for person lookup depending on whether the vendor can reverse-resolve email to LinkedIn identity.
Name + current company: works reasonably well for senior executives at larger companies, degrades quickly for common names, smaller companies, or anyone with a recent job change.
Name + domain: similar to name + company, with added ambiguity when a company has multiple domains.
The critical rule is to test with the identifier type you'll actually use in production. If your CRM has LinkedIn URLs for 30% of records and name+company for the rest, run two separate evaluations and weight them accordingly. Vendors who perform well on LinkedIn URL lookups sometimes perform poorly on name-based resolution, and vice versa.
Step 3: Run the same set through multiple vendors simultaneously
This is the step most evaluations skip, and it's the most important one.
Run your test set through every vendor you're seriously considering in the same week, using the same input format. Use live API calls, not cached results — cached responses reflect historical quality, not current indexing. Most vendors offer trial access that allows this.
For each record, track:
Match: Did the vendor return any result?
Fill rate: For the records that matched, what percentage of your required fields were populated? (Don't count fields you don't need.)
Accuracy: For matched records, compare the returned data against your known ground truth. Count exact matches and meaningful mismatches separately.
Freshness: Where available, check the data's last-updated timestamp. Stale data that looks correct today may be wrong by the time you use it at scale.
Build a simple spreadsheet. One row per test record, one column per vendor for each metric. The aggregate view will tell you what no demo ever will.
Step 4: Score what actually matters for your use case
Not all metrics matter equally for every use case. Before you score, decide which ones are decision-critical for yours:
SDR outbound → match rate and email fill rate are what move the needle
CRM enrichment → accuracy and freshness matter more than raw match rate
Lead scoring models → field coverage consistency across your entire record base
Market intelligence → company-level coverage, employee count accuracy, industry classification
AI pipelines → structured response consistency and freshness of organizational data
Weight your scoring accordingly. A vendor with a 70% match rate and 95% accuracy may be worth more than one with a 90% match rate and 65% accuracy. This depends entirely on what you're building.
What good results actually look like
Realistic benchmarks for well-structured tests (LinkedIn URL input on a diverse B2B record set):
Match rate: 75–90% is strong. Below 60% should prompt questions. Above 95% on a diverse set warrants scrutiny.
Field fill rate on returned records: 80%+ on core fields (name, title, company) is table stakes. Email fill rate varies dramatically by vendor and matters most for outbound use cases.
Accuracy on verified ground truth: 85%+ on job title and company is achievable. Lower on title for records older than 12 months (job changes are the hardest thing for any enrichment vendor to keep current).
If a vendor's numbers look suspiciously good across every dimension, test their edge cases: smaller companies (under 50 employees), non-English-speaking markets, people who changed jobs in the last 6 months. That's where dataset quality separates.
The vendors who discourage rigorous testing tell you something
Every enrichment vendor will show you cherry-picked accuracy examples. Not every vendor will tell you their real-world match rates on a representative test set. The ones who do, or better yet the ones who give you the framework to find out yourself, are the ones who have invested in data quality rather than just data volume.
We publish this methodology because we want you to run it. Not because we'll win on every dimension for every use case, but because the buyers who test properly make better decisions, and the ones who test us properly tend to become customers.
Test us. Test our competitors. Use this framework. Then decide.
→ Start a free trial and bring your own test set:enrichlayer.com