Part 1 mapped the six layers of a data intelligence system — data acquisition, schema normalization, entity resolution, attribute enrichment, signal modeling, and decision interfaces — and the failure modes that live in each.
That post covers the mechanics of each layer and the specific ways they fail, like bad merges, drifted schemas, and stale attributes. But it leaves an open question: Why does one system earn a team's trust while another, built the same way, never does?
Size matters not here. A bigger dataset is not automatically a more trusted one. Part 1’s whole point was that a false merge three layers down corrupts every metric sitting above it, and more records only widen the potential for damage.
What separates a system from being reliable or unreliable is whether the people using it can tell what the data represents, how it changes, and where it breaks. We call this interpretability, which is how far you can open a system up and reason about it.
Buyers rarely get to inspect the layers directly. They evaluate stand-ins like record counts, schema breadth, update frequency, benchmarks, dashboards, and case studies. Those look like clarity, but they leave questions about reliability unanswered:
- How stable is the schema when a new source gets added?
- Which value wins when two sources disagree?
- Which attributes are observed, and which are inferred?
- What happens to a record when two companies get merged into one by mistake?
- When a value updates, which downstream signals move with it?
These are the same failure modes Part 1 catalogued, asked from the outside. A team can't run on a dataset until it can answer these questions.
From collecting data to interpreting it
Plenty of systems still sell themselves as collection infrastructure claiming coverage, freshness, and volume. Part 1's acquisition layer is exactly that, and its job is to determine coverage, not quality. Coverage matters, but it never tells you whether the data survives contact with a real workflow.
As a system takes in more kinds of input, the problem moves from getting the data to reading it correctly. This means someone other than the person who built it can pick up the dataset and understand what a field means, how often it changes, and where its limitations are.
Ten million company records with undocumented merge logic can be harder to use than one million with clear field definitions and a predictable refresh schedule. With the big set, you can't trust its numbers, because you don't know how the records were merged. With the small set, you can see exactly where each number comes from.
How buyers judge before they compare
Formal comparison happens late in the process, once a team is already running trials, filling in scorecards, and lining up vendors in a spreadsheet. But which vendors even make it onto that spreadsheet gets decided earlier, and mostly on gut feel.
As buyers gain experience, they start to expect a good system to work properly, and they learn what to check. Once a buyer has watched one merge mistake — for example where two different companies treated as one throws off a market-size number (the same mistake Part 1 warned about) — a claim like "we have millions of records" stops being impressive. A big pile of records tells you nothing about whether the records inside are right.
Once a buyer can tell which facts in the data are actually observed and which are only guessed, saying "500 data points per company" doesn’t mean much because many of those points may be guesses instead of checked facts. And once they realize that "refreshed daily" might mean only one part of the data updates while the rest stays old, they may start asking what actually changes each day.
By this point, the buyer no longer takes a vendor's claims at face value. They've built a mental checklist for testing how a data system actually behaves, and they run every vendor through it.
Where vocabulary stops being enough
A lot of writing about data systems remains theoretical. It defines terms, lists benefits, and walks through use cases. That’s helpful for a newcomer, but it’s not everything when it comes to evaluation.
Knowing what enrichment means doesn't tell you how it behaves when two sources describe the same company with different schemas. Knowing what entity resolution is doesn't tell you how a bad merge distorts a territory map or a market-sizing estimate. "The data updates regularly" doesn't say which layer updates, on what cadence, or what moves downstream when it does.
Writing that reconciles these gaps gets bookmarked, quoted in internal docs, and passed around a team, because it answers legitimate questions about data systems in the real world.
What to inspect before you trust a dataset
Before you trust a dataset in production, put the vendor through these questions:
- Which fields are observed, and which are inferred? Observed fields come from a real source; inferred fields are model guesses. A vendor who can't separate the two is selling you both as if they're equal.
- What's the fill rate on each field? A field present on 15% of records is a bonus, not a foundation. Ask for coverage field by field, not one headline total.
- Does each record carry a last-updated date, and what counts as stale? "Refreshed daily" with no per-record timestamp is unverifiable.
- How does it decide when two records are the same company? Ask for the rules. For example, does an exact domain match beat a subdomain, and does it treat a parent company and its subsidiaries as one or keep them separate? Get this wrong and two different companies merge into one. No error fires, and every number built on them is off (covered in Part 1's entity resolution layer).
- When two sources disagree, which one wins? If there's no rule, conflicts get resolved at random.
- Is there a field dictionary and a schema map? If a vendor can hand you a definition for every field and how it's derived, the data is legible. If not, you'll reverse-engineer it in production.
What this looks like in a real system
Every Enrich Layer API response carries a meta.last_updated timestamp, so freshness is a value you read off each record; a profile that hasn't been refreshed in 29 days comes back marked stale rather than served as current.
When a record is sparse, a meta.thin_profile flag says so, instead of a thin profile passing for a complete one. The resolution logic is documented down to the tie-breakers. The company endpoint matches exact domains before subdomains and parents before subsidiaries, and person lookups return similarity scores on a 0–1 scale, so you can tell how confident a given match is.
On a landing page these read as footnotes. You notice them when you have to defend a number three layers downstream, which is exactly when they matter.
Interpretability is what you're really assessing
Every data system gets more complex as it grows. The one a team keeps trusting is the one it can still inspect, the one that can say which fields are observed, how fresh each record is, and how it resolves conflicts. Part 1 gave you a diagnostic for your own stack, and the checklist above turns the same lens on the vendors you're evaluating. A dataset is only as useful as it is inspectable.
