Your average retention looks stable. Comforting, and possibly a lie. Because averages blend everyone together, improving new customers can mask collapsing old ones. And vice versa.
Cohort analysis is how you un-blend the picture.
📌 TL;DR: Cohort analysis groups customers by when (or how) they started and follows each group separately over time. The classic output is a retention grid: rows are signup months, columns are months since signup. It answers the question averages can't: is what we changed actually working?
What Is Cohort Analysis?
Cohort analysis groups customers by a shared starting point (usually signup month or quarter) and tracks each group’s behavior over time. Retention, spending, engagement. Instead of one blended average, you get a family of curves. One per vintage, each aging on its own timeline.
The classic visualization is the cohort grid. Rows are start periods, columns are periods since start, cells show the metric. Read down a column to compare vintages at the same relationship age. That’s the view that isolates whether things are improving.
What do cohorts reveal that averages hide?
- Whether changes worked: did customers acquired after the onboarding revamp retain better at month three than those before it? The grid answers directly
- Where the cliff is: most products lose customers at a consistent relationship age. Cohorts expose the month it happens
- Quality shifts in acquisition: a channel change shows up as newer vintages behaving differently from older ones
- The truth under growth: fast acquisition can mask terrible retention in blended numbers. Cohorts don’t allow the disguise
And cohorts don’t have to be time-based. Grouping by acquisition channel, plan, or segment answers “which KIND of customer works?” That feeds directly into lifetime value math and churn diagnosis. The dependency, as ever: reliable start dates and consistent customer records. Otherwise the grid is beautifully formatted noise.
Reading a retention grid like a practitioner
The grid rewards knowing where to look. Start down the first column: month-one retention by vintage. That’s your onboarding report card over time. Improving means recent changes land. Degrading means new customers hit new friction.
Then read along the rows: each vintage aging. This shows the retention curve’s shape: where it cliffs, where it flattens. And the flattening matters enormously. The plateau height is your long-term retained core, the asset everything else is built on.
The diagonal carries calendar events. An outage or pricing change hits every cohort in the same calendar month, so it appears as a diagonal stripe across the grid. That’s the signature separating “something happened to everyone in March” from “month-three customers always wobble.”
The triangle’s edge, the newest and shortest rows, tempts overreading. Two data points on a young cohort are weather, not climate. So resist judging vintages before they’ve aged past your known cliff.
Cohorts beyond retention
- Revenue cohorts: cumulative revenue per vintage, the empirical CLV curve drawn from actual money instead of assumptions
- Behavioral cohorts: grouped by first action (imported data vs started empty; attended onboarding vs skipped). This is where activation insights hide
- Channel and campaign cohorts: acquisition source quality measured by how vintages BEHAVE, not what they cost. The antidote to cheap-lead celebrations
- Plan and price cohorts: how packaging changes ripple through retention, visible only when vintages stay separate
Running cohort analysis credibly
The mechanics are unforgiving about data discipline. You need stable start-date definitions: signup, activation, or first payment. Pick one and freeze it. You need consistent activity definitions across eras. And you need clean identity, because duplicate accounts smear one customer across false cohorts, and merged accounts teleport history.
Where definitions changed mid-history, break the grid honestly instead of comparing across the seam. A cohort grid built on shifting definitions is the most authoritative-looking wrong artifact in analytics. Which is exactly why the underlying record quality deserves an audit before the grid deserves belief.
A worked read: one grid, three findings
Picture a monthly retention grid for a B2B tool, and three real patterns it might surface.
Finding one: every cohort drops hardest between months one and two, then flattens. The product has an activation cliff. So onboarding investment should aim exactly there, not at the flat middle.
Finding two: cohorts from March onward hold five points better at month three than earlier vintages. The Q1 onboarding revamp is working. And you can only see it because the vintages stayed separate.
Finding three: a horizontal stripe of weakness in one calendar month, across ALL cohorts. That’s not a vintage story. It’s an event story. Check the incident log and the churn reasons for that month.
Three findings, three owners, three different actions, all from one honest grid. That’s the method’s entire pitch. Averages would have shown a single wobbling line and invited a single wrong debate.
Common Mistakes
Cohort grids look authoritative even when the reading is wrong. So check yourself against the traps practitioners hit most.
- Judging young cohorts early: two months of data on a fresh vintage is noise. Wait until a cohort ages past your known cliff before comparing it to anything
- Comparing across a definition change: if “active” meant one thing in January and another in June, the retention curve breaks at the seam. Split the grid there and compare within eras
- Celebrating month one while the plateau sinks: better early numbers can coexist with a shrinking long-term core. Watch both ends of the retention curve, not just the start
- Blending channels into one grid: a strong signup month from a weak channel drags the vintage down and hides the real story. Split by source whenever the acquisition mix shifts
- Survivorship reads: asking month-six customers why they stayed tells you nothing about the ones who left in month two. Pair the grid with churned-customer follow-up
But the deepest mistake is treating the grid as a report instead of a question machine. Every odd cell deserves a why. That’s where the vintage comparison earns its keep.
Frequently Asked Questions
What is cohort analysis in simple terms?
Grouping customers by when they started and tracking each group separately over time, instead of blending everyone into one average. It’s how you see whether newer customers behave better than older ones did.
What is an example of cohort analysis?
Comparing month-3 retention of customers who signed up each month this year, and seeing whether the vintages after a product change hold on better. Down-the-column comparison is the whole trick.
Why use cohort analysis instead of averages?
Averages mix customers at different relationship ages and hide diverging trends; cohorts keep vintages separate so cause and effect stay visible. Most “stable” averages hide at least one moving story.
What does a good retention curve look like?
An early decline that flattens into a stable plateau, which is your durable customer core. Curves that never flatten describe a leaky bucket; raising the plateau beats slowing the early slide in long-term value.
What time grain should cohorts use?
Match the product’s natural rhythm: weekly cohorts for high-frequency products, monthly for most B2B, quarterly for long sales cycles. Too fine a grain produces noisy triangles; too coarse hides the cliffs you’re looking for.
What tools do you need for cohort analysis?
Any environment that can group by start period and pivot by elapsed time: SQL and a spreadsheet suffice; product-analytics tools add convenience. The constraint is never tooling. It’s clean start dates and stable definitions in the underlying records.