Most companies that say they are scaling are not.
The difference shows only when it is too late to matter. So we built the study that shows it early: the first matched, forensic study of scaling success and failure in business-to-business technology companies.
182
companies studied
76
verified scaling failures, reconstructed from primary records
106
companies that scaled and held
$27bn
of real invested capital that never came back
$117bn
of peak paper value erased. Counted separately, never added to the cash figure, because paper is not money.
WHY THIS STUDY EXISTS
Four studies dominate every boardroom conversation about growth. Bain’s Founder’s Mentality. Startup Genome’s Scaleup Report. Jim Collins’s Good to Great. Lee and Kim’s 2024 paper in the Strategic Management Journal.
Here is what none of them will tell you: not one examined the mid-market business-to-business technology company. Bain studied giants. Startup Genome studied seedlings. Collins studied the large public companies of a vanished era. Lee and Kim studied job adverts.
The company that defines the mid-market technology portfolio had never been the subject of a matched study of why scaling holds or fails.
This is that study.
THE COMPARISON
| This research | Bain, The Founder’s Mentality | Startup Genome, Scaleup Report | Collins, Good to Great | Lee and Kim, 2024 | |
|---|---|---|---|---|---|
| Who it studied | 182 business-to-business technology companies, across the growth range | Database of roughly 8,000 global public companies, all sectors and sizes; about 100 executive interviews | Roughly 7,000 young startups, pre-scale | Large established public companies of the 1965 to 1995 era | Over 38,000 United States startups, mostly early stage, via 6.3 million job postings |
| The mid-market technology operator’s companies? | Yes, and across every ownership structure theirs might have | No. Mostly far larger. | No. Mostly far smaller. | No. Fortune-500 scale. | No. Pre-scale. |
| What counted as success | Durable operating scale: value that held, defined before analysis began | Sustained growth without stall-out | Reaching a US$50m valuation mark | Fifteen years of stock-market outperformance | Timing of first managerial and sales hires |
| Studied failure directly | Yes: 76 verified scaling failures, forensically reconstructed | Indirectly: crises inside survivors | Only as a line on a chart | No failures: comparison firms all survived | Only as a statistical outcome |
| Counted the money lost | Yes: two measures, reported separately, never summed | No | No | Stock returns only | No |
| Covers 2020 to 2025 | Yes, and isolates the cheap-capital vintage as its own pattern | No: fieldwork ended around 2015 | Partly, through a valuation lens | No: data ends in the 1990s | Partly |
| Re-tests its own winners | Yes: a dated durability test that demoted two of its own successes | No | No | No: several exemplars later collapsed in print | Not applicable |
| Ends in a decision | Yes: Scale Up, First Optimise, or Scale Back | A mindset | Benchmarks | A philosophy | A finding |
| Evidence inspectable | Yes: every source graded, every case block openable, rejected explanations published | Proprietary | Partly; founder self-report | Not reproducible | Yes, in principle |
FIVE CLAIMS NO OTHER STUDY CAN MAKE
1. It studied the right population.
Findings from Fortune-500 conglomerates or seed-stage startups reach a mid-market portfolio only by analogy, and analogy is where growth advice goes to die. This research needs no leap. Venture-backed, sponsor-backed, listed and founder-controlled companies appear on both sides of the line, so no reader gets to think their ownership structure exempts them.
2. It measured the outcome that pays.
Startup Genome’s finishing line is a US$50 million valuation mark – precisely the species of paper the 2021 vintage taught every investor to distrust. Collins’s finishing line was share-price momentum, and the market later delivered its own verdict on several of his exemplars. This study’s finishing line is the only one that pays: value that held.
3. It counted the money, honestly.
Two measures, reported separately, never added together, because one counts paper that was never cash and the other counts cash that never came back. Most failure research sums everything in sight to manufacture the largest possible headline. This study refuses. The refusal is the credential.
4. It covers the era every current portfolio is living through.
The cheap-capital surge, the megaround cluster, and the correction that followed. And it treats that era with suspicion rather than excitement: no pattern was retained unless it also held outside the zero-rate vintage. What survives that filter is an operating signal, not a cycle artefact.
5. It ends in a decision, not a philosophy.
Bain ends in a mindset. Collins ends in virtues. Startup Genome ends in benchmarks. This research ends in one of three verdicts a board can act on Monday morning: Scale Up, First Optimise, or Scale Back.
THE TEST NO OTHER STUDY DARED TO RUN
In 2001, Good to Great crowned eleven companies as enduringly great. Circuit City went bankrupt in 2009. Fannie Mae passed into government conservatorship in 2008. The book kept selling. The verdicts were never withdrawn.
This study runs a dated durability test on its own success corpus. When that test found two success cases whose value and operating trajectory had genuinely reset, it did not quietly keep them. It reclassified them to the failure corpus, on the record, with the date on the entry.
No famous study of company growth has ever demoted its own success stories. This one did. That is the difference between research that flatters its conclusions and research built to survive them.
THE TWO QUESTIONS SCEPTICS ASK
“Only 182 companies?”
Look at where the depth actually sits. Bain’s celebrated 8,000 were screened by database; the mechanism work rests on about one hundred interviews. Collins’s five-year, twenty-one-researcher programme analysed eleven companies in depth. This study coded 182, every one against the same instrument, from graded primary sources, with the full evidence block for every case openable and checkable. On the measure that matters – depth per company – this is not the small study in the field. It is the large one.
“Where is the famous name on the cover?”
The famous studies ask for trust. This one hands you the means to check. Every source is graded, and any claim without primary verification is downgraded, not promoted. Every undisclosed private figure is recorded as a ceiling, never guessed. Every rejected explanation is published alongside the retained ones, with the evidence that killed it. A blind second coder re-scored a random sample of cases against definitions alone. Name one famous growth study that will show you its rejects. There is not one.
WHAT THIS STUDY DOES NOT CLAIM
Honesty about limits is part of the method, so here they are, unprompted. The study supports mechanism claims, not frequency claims: it shows how failure unfolds where a pattern is present, against matched cases where its absence coincides with the opposite outcome. It never says what percentage of failures a pattern causes. The design is observational, not experimental. The claims hold for business-to-business technology companies on the organic-growth question, and nowhere else. Every durability judgement carries a date, and every judgement is re-tested on a fixed date each year, with the aggregate results published whichever way they fall.
THE FULL RECORD
The complete methodology – membership rules, screening funnel, source-grade hierarchy, coding architecture, the blind second-coder result, the rejected explanations, and the technical appendices – is published in full.
The full findings are the subject of a forthcoming book. The instrument the study produced, the Scaling Diagnostic, is in use now.