The first enrichment run I ever paid for came back 58% complete. I stared at that file for a long time. We’d bought a well-reviewed provider, sent them 2,000 contacts, and got back a list where nearly half the email column was still blank.
Nobody had told me the quiet truth of B2B data: no single provider covers everyone. Not one.
Waterfall enrichment is the fix for exactly that problem. So let’s break down what it is, how the cascade works, how to order your providers, and the hidden costs the vendor pages skip.
What is waterfall enrichment?
Waterfall enrichment queries multiple data providers in sequence, so records one provider misses get caught by the next. Provider one gets first try. Whatever it can’t fill “falls” to provider two, then provider three, until the field is found or the list runs out.
The name comes from that falling motion. And the payoff is coverage: the percentage of your records that end up with the field you wanted. It’s a technique inside the broader B2B data enrichment process, not a replacement for it.
📌 TL;DR: Waterfall enrichment chains several data providers so each one only handles what the previous one missed. You get higher coverage than any single provider offers, you only pay each provider for its own layer, and the two things to plan for are provider order and the hidden costs of running several vendors at once.
How does waterfall enrichment work?
A waterfall works by cascading each unmatched record to the next provider in line until someone fills it. Here’s the flow for a single field, say a work email:
- Your record goes to provider 1. If it returns the email, the record exits the waterfall. Done.
- If provider 1 misses, the record cascades to provider 2. Same rule.
- Misses cascade again to provider 3, and so on down your list of sources.
- A verification step checks whatever came back before it touches your CRM.
→ Record → Provider 1 (hit? stop) → Provider 2 (hit? stop) → Provider 3 → Verify → CRM.
Two terms you’ll see constantly. The hit rate is how often a given provider fills the requests it receives. And coverage is the share of your whole list that ends up filled after all levels run. Waterfalls exist because stacking imperfect hit rates raises total coverage, which is the entire point of data enrichment in the first place.
What can you waterfall?
Any field with multiple possible sources can run through a waterfall. In practice, B2B teams waterfall four things:
- Work emails: the most common waterfall, because email coverage varies wildly between providers
- Phone numbers: especially direct dials and mobiles, the scarcest fields in B2B data
- Firmographics: industry, headcount, and revenue, where providers disagree more than you’d hope
- Company domains: resolving a company name to its website, the match key everything else depends on
Emails and phones benefit most. Company-level data has smaller gaps between providers, so a two-level waterfall is usually plenty there.
How do you order the providers?
Order providers by accuracy for your specific market first, then by cost per match, with verification always last. That one sentence is most of the strategy. Let me unpack it.
Accuracy comes first because the first provider to answer wins. If a sloppy source sits at level one, its wrong answers exit the waterfall looking like hits, and your better providers never get asked. Test each candidate on a sample of YOUR list (not a generic benchmark) and put the one that’s most accurate for your region and industry up top.
Between providers of similar accuracy, put the cheaper one earlier. It’ll absorb the bulk of the volume, and the expensive specialist only sees the hard leftovers. That’s the leftover queue it’s actually worth paying premium rates for. Our data enrichment API roundup is a good starting shortlist for choosing those candidates.
💡 Ordering tip: re-test your order quarterly. Provider hit rates drift as their databases refresh, and the level order that was right in January can quietly become the expensive order by summer. Your per-level hit-rate report will tell you.
Sequential or parallel: which waterfall style wins?
Sequential waterfalls save money; parallel ones save time. In sequential mode, providers run one after another and stop at the first hit. In parallel mode, you query several providers at once and keep the best answer.
| Sequential | Parallel | |
|---|---|---|
| Cost | Lower (stop at first hit) | Higher (you pay every provider) |
| Speed | Slower (levels stack) | Fast (one round trip) |
| Best for | Batch list projects | Real-time flows like inbound forms |
| Answer quality | First hit wins | You can compare and pick |
The honest default? Sequential. Most teams enrich in batches where a few extra minutes cost nothing and duplicate provider fees cost plenty. Go parallel only where a human is waiting on the result.
The hidden costs nobody mentions
Waterfalls raise coverage, but they also multiply everything else. Four costs to plan for before you build:
You pay for misses at every level. Some providers charge per query, not per match. A record that falls through three paid-per-query levels costs you three times before anyone fills it. Check each provider’s pricing shape before assigning it a level.
Latency stacks. Each level adds its own response time. Fine for overnight batches. Painful in a real-time flow, which is why parallel exists.
Providers disagree. Level one says 200 employees, level three says 950. You need a rule for which value wins (usually the higher-accuracy source, with a timestamp), or your CRM becomes an argument.
Compliance multiplies. Every provider in the chain is another data processor touching your contacts.
🧠 Compliance note: a three-provider waterfall means three vendor agreements, three sourcing stories, and three places an EU regulator could ask about. Under GDPR you stay the data controller for all of it, so collect each provider's sourcing documentation up front, not after a question arrives.
A worked example: the three-level email waterfall
Numbers make this concrete, so here’s an illustrative run. The percentages are round examples to show the mechanics, not measurements. Say you start with 1,000 contacts needing work emails:
- Level 1 (your most accurate provider) fills 600 of 1,000. → 400 fall through.
- Level 2 fills half of those 400. → 200 fall through.
- Level 3 (the specialist) catches 80 of the hard 200. → 120 remain unmatched.
- Verification then screens all 880 found emails and flags the risky ones before anything gets imported.
Total coverage: 880 of 1,000. No single level came close to that alone, and that’s the whole argument for waterfalls in one list. The verification pass at the end is not optional, by the way. It’s what keeps appended emails from torching your sender reputation, and an email verification API does it automatically as the final level.
How do you set up a waterfall step by step?
You set up a waterfall by defining fields, choosing ordered providers, setting stop conditions, and monitoring per-level results. Concretely:
- Define the fields you’re waterfalling. Emails only? Emails and phones? Each field gets its own cascade.
- Shortlist 3-4 providers and test each on the same 200-row sample of your real list.
- Set the order: accuracy first, then cost, verification last.
- Set stop conditions: a stop condition is the rule that ends the cascade for a record, usually “field found” or “all levels tried.” Add a confidence threshold if your providers return one.
- Add the verification level so nothing unverified reaches the CRM.
- Monitor per-level hit rates monthly. A level that stops earning its position gets reordered or dropped.
Should you build or buy your waterfall?
Buy the orchestration if enrichment isn’t your product; build it only if you have engineers to spare and unusual needs. Platforms like Clay and Apollo sell waterfall features off the shelf, and orchestration (the software that routes records between providers and applies your rules) is exactly the undifferentiated plumbing that’s rarely worth building. But the decision has real math on both sides, and the build vs buy comparison walks through it line by line.
My first waterfall, honestly
Remember that 58%-complete file from my Hamburg days? Here’s what happened next. We added a second provider and pushed the misses through it, by hand at first, with a spreadsheet and stubbornness. Coverage on that 2,000-contact list climbed to a little over 80%, and the second provider only billed us for the smaller leftover pile.
The surprise wasn’t the coverage. It was WHICH records each provider caught. Our first provider was great for US tech and weak for German Mittelstand companies, and the second was the mirror image. Nobody’s marketing told me that. The leftover file did.
So test on your own list. Always.
How we know this (and what to double-check)
This guide draws on hands-on enrichment projects with multi-provider setups plus the public documentation of waterfall platforms as of 2026. Two limits worth naming. Hit rates vary a lot by industry, region, and list quality, so the example numbers above show mechanics, not promises. And provider capabilities change fast, so re-test your own waterfall order on a fresh sample before trusting last quarter’s setup.
Frequently asked questions
What is waterfall enrichment in simple terms?
Waterfall enrichment tries multiple data providers in order, so each provider only handles what the previous one missed. The result is higher total coverage than any single provider can deliver alone.
How does the waterfall method work?
A record queries provider one; a hit exits the cascade, a miss falls to provider two, and so on. A verification step screens the results before they reach your CRM.
What are the best waterfall enrichment tools?
Look for orchestration platforms that let you plug in your own choice of providers, set the order, and see per-level hit rates. Clay and Apollo are the best-known examples, and several data vendors now bundle waterfall features. Judge them on provider flexibility and reporting, not the label.
What is waterfall enrichment in Apollo and Clay?
Both platforms offer built-in waterfalls that query several integrated data sources in sequence for you. Same cascade concept as this guide, packaged as a feature, with each platform’s own provider lineup and credit rules.
Can you do waterfall enrichment for free?
Partly. You can chain the free tiers of several providers by hand: run your list through one, export the misses, run those through the next. It’s slow and manual, but the cascade logic is identical and it costs nothing except time.
Is sequential or parallel enrichment better?
Sequential is better for cost, parallel for speed. Batch projects should default to sequential, since stopping at the first hit avoids duplicate charges. Save parallel calls for real-time flows where someone is waiting.
How many providers should a waterfall have?
Three to four data levels plus verification covers most teams. Each added level catches fewer records than the one before it, so past four the extra coverage rarely justifies the extra vendor management.
It’s time to stop accepting 60% coverage
You don’t need a platform migration to start. Take your last enrichment export, count the blanks, and push just those misses through one more provider this week. That’s a two-level waterfall, and you’ll see your real numbers by Friday.
Coverage is a choice now. Choose it. You’ve got this.
Running a waterfall already? Tell me in the comments which level surprised you. Mine was level two.