Field notes
CRM enrichment without the hallucinations
AI enrichment is half verified fact and half confident guessing, and the method that tells them apart is the entire product.
Trustworthy AI enrichment isn’t magic. Every value is checked against a named source, the source is logged on the record, and anything the system isn’t confident about goes to a person instead of into your CRM as fact.
On this page
Half of what gets sold as 'AI enrichment' is a model making things up with confidence.
You ask for a company's headcount, its industry, the CFO's name. It returns all three, instantly, formatted, certain. Two are right. One is a guess dressed as a fact, and it just went into your CRM wearing the same font as the truth.
That's the problem with enrichment as most people buy it: it doesn't tell you which parts it knows and which parts it invented. And the share it invents on the hard-to-find fields is not a rounding error - [STAT: measured error rate of unverified AI enrichment on private-company firmographics].
The fix isn't a cleverer model. It's method: provenance, confidence thresholds, and a person on the fields that matter. Boring, checkable, and the only version that survives contact with an auditor.
Why AI enrichment hallucinates in the first place
A language model's job is to produce a plausible answer, not a true one. Those are different tasks that happen to overlap most of the time.
Ask for a fact it has seen a thousand times and it gets it right. Ask for one it has never seen - a private company's exact headcount, a role that changed last month - and it won't stop to say so. It generates the most likely-looking answer. Likely-looking and correct are not the same thing, and enrichment is exactly where they part company.
So the goal isn't a model that never guesses. It's a system that knows when it is guessing.
Provenance: every field points back to a source
The first rule: no value enters the CRM without a source attached.
If the system writes '250 employees', it also writes where it read that - the filing, the company page, the profile, with a link and a date. Enrichment without provenance is a rumour with good formatting. Enrichment with provenance is a claim you can check in one click.
If a field can't cite where it came from, it doesn't get written. That one rule removes most of the danger, because a fabricated value has no source to point to, so it never makes it in. This is the spine of every enrichment build I ship.
Confidence thresholds: let the system say it doesn't know
The second rule: the system scores its own confidence, and low confidence is not treated as fact.
A match on a clear, verified source is high confidence and can be written automatically. A weak match. A stale source - and sources go stale fast, with HubSpot's own data putting B2B contact decay at about 22.5% a year. Two sources that disagree. All low confidence, and none of it written as truth. It gets held, flagged, and sent to a person.
Most enrichment tools have one setting: fill the field. An honest one has three - fill it, hold it, or ask a human - and it chooses based on how sure it actually is.
The threshold is a dial, not a law. On a field where a wrong answer is cheap, you can let more through. On a field you will report to a regulator, you set the bar high and send more to review. You decide where guessing is allowed.
Human-in-the-loop on the fields that matter
People assume human review is the slow, expensive part. Done right, it's tiny.
The machine handles the fields it can verify - the overwhelming majority - on its own. Only the low-confidence handful reaches a person, already sorted, the conflicting sources shown side by side. The human isn't enriching. They're adjudicating, a few records at a time, on exactly the fields where a wrong answer would cost something.
This matters more the more regulated you are. I spent more than a decade in regulated financial services, inside an eleven-country group. A wrong field there wasn't an inconvenience. It was a compliance problem you'd be defending later. Keeping a human on the fields that carry risk isn't caution for its own sake. It's the job.
What to ask before you buy enrichment
Three questions separate the honest tools from the confident ones.
Where did this value come from - can it show the source for a given field. What happens when it isn't sure - does it write a guess or hold the field. And who checks the low-confidence ones - is there a human step, or does everything land as fact. A vendor who answers all three is selling method. A vendor who changes the subject to model size is selling a guess with good formatting.
The method is the product
The selling point of good enrichment isn't the model. It's the discipline around the model.
Provenance, so every value can be traced. Confidence, so the system knows its own limits. A human on the fields worth a human. None of it is exciting to demo, because the exciting demo is the one that fills two hundred fields in four seconds and never mentions that a quarter of them are guesses. It's the same principle I hold to when training an agent on real data: ground it in something checkable, keep a person on what matters.
A model that fills every field is easy. A model that admits which fields it can't fill is the one you can trust.
Ask the vendor which one they're selling.
Common questions
Does AI data enrichment hallucinate?
Unguarded enrichment can, because a language model produces the most plausible-looking answer even when it hasn't actually found the fact. The fix is method: verify each value against a named source, log where it came from, and route low-confidence fields to a human instead of writing them as fact.
What is provenance in CRM enrichment?
Provenance means every enriched value carries the source it came from - the filing, page, or profile, with a link and a date. If a field can't cite a source, it doesn't get written, which is what stops fabricated values from entering the CRM.
How do you keep AI enrichment accurate?
Score confidence on every field, write only the high-confidence matches automatically, and send weak or conflicting ones to a person to adjudicate. Accuracy comes from the discipline around the model, not from the model alone.
This is the kind of AI we build into a revenue engine - verified, logged, and owned by you.
Get the next one
New field notes twice a month. We don’t do newsletters, follow RevOps XL on LinkedIn instead.
Follow on LinkedIn ↗