The Truthscaling Study
Most companies that say they are scaling are not. The difference shows only when it is too late to matter. So this study reads the companies that failed and the companies that held against the same instrument, and the instrument was locked before the first company was read.
WHAT IT IS
The Truthscaling Study is a study of 182 business-to-business technology companies: 78 that failed to scale and 104 that scaled and held. All 182 were scored on one instrument, locked before any company was read. The register holds 199 companies. Twelve moved after they had been coded, ten off the failure side and two off the success side, and none was deleted. Every ruling is recorded and reversible, and the register is published.
It was built over eighteen months from primary records rather than from interviews about what people remember.
182
companies studied
78
that failed to scale
104
that scaled and held
$27bn
of real invested capital that never came back
$117bn
of peak paper value erased. Counted separately, never added to the cash figure, because paper is not money.
WHY IT WAS BUILT THIS WAY
Most work on scaling reads the winners. Reading only the winners cannot tell you which of their habits mattered, because the companies that had the same habits and failed are not in the room.
So this study reads both. The failures and the successes are scored against the same frame, and the frame was closed before the first company was read. That order matters more than the size of the study. It means no company could shape the thing it was measured with.
WHAT THE INSTRUMENT TESTS
Four conditions. A company is scaling only while all four hold at the same time. The weakest one decides, and strength on three can never buy back a failure on the fourth.
Durability. It keeps. The money you already earn keeps coming. Customers won last year are still paying this year, at the same price or better, without being rescued, heavily discounted, or propped up by extra people.
Repeatability. It repeats. New customers are won by a method the company owns, not by luck or by one person’s relationships. It is proved by transfer: the same win, produced by a different seller, at a different kind of customer, at the normal price.
Compounding. It multiplies. Growth arrives that nobody paid for. Customers already on board buy more each year they stay, work delivered for one customer helps win the next, and each new win costs less than the one before it.
Absorbable pace. It absorbs. The company takes on customers, staff and markets no faster than it can carry them. It is read from what breaks first: complaints rising against last year, customers waiting longer to get started, and new joiners taking longer to become productive than they did a year ago.
WHAT IT FOUND
The failures were not short of a market or a product. They were short of an accurate account of themselves.
Of the 76 failures scored so far, 67 mistook external endorsement for evidence that the business worked. A funding round, a logo, an analyst’s note or a competitor’s valuation was treated as proof, when none of them is. That is the most common pattern in the failure set.
Next most common, 63 of the 76, was installing a playbook that had worked somewhere else without testing whether it fitted here. Then 43 whose leadership could not revise a theory because their standing rested on having been right before. Then 29 who saw the parts of the business but never the system connecting them. Then 19 who had been told the bad news for so long that they had stopped hearing it.
Most companies carried more than one.
WHY THIS STUDY EXISTS
Four studies dominate boardroom conversation about growth: Bain’s Founder’s Mentality, Startup Genome’s Scaleup Report, Jim Collins’s Good to Great, and Lee and Kim’s 2024 paper in the Strategic Management Journal.
None of them examined the mid-market business-to-business technology company. Bain’s sample sat far above it and Startup Genome’s far below. Collins studied large United States public companies between 1965 and 1995. Lee and Kim worked from job postings at mostly early-stage firms.
So the company that makes up most of a mid-market technology portfolio had never been the subject of a study that read scaling failures and scaling successes together.
This is that study.
THE COMPARISON
| This research | Bain, The Founder’s Mentality | Startup Genome, Scaleup Report | Collins, Good to Great | Lee and Kim, 2024 | |
|---|---|---|---|---|---|
| Who it studied | 182 business-to-business technology companies, across the growth range | Database of roughly 8,000 global public companies, all sectors and sizes; about 100 executive interviews | Roughly 7,000 young startups, pre-scale | Large established public companies of the 1965 to 1995 era | Over 38,000 United States startups, mostly early stage, via 6.3 million job postings |
| The mid-market technology operator’s companies? | Yes, and across every ownership structure theirs might have | No. Mostly far larger. | No. Mostly far smaller. | No. Fortune-500 scale. | No. Pre-scale. |
| What counted as success | Durable operating scale: value that held, defined before analysis began | Sustained growth without stall-out | Reaching a US$50m valuation mark | Fifteen years of stock-market outperformance | Timing of first managerial and sales hires |
| Studied failure directly | Yes: 78 companies that failed to scale, reconstructed from primary records | Indirectly: crises inside survivors | Only as a line on a chart | No failures: comparison firms all survived | Only as a statistical outcome |
| Counted the money lost | Yes: two measures, reported separately, never summed | No | No | Stock returns only | No |
| Covers 2020 to 2025 | Yes, and isolates the cheap-capital vintage as its own pattern | No: fieldwork ended around 2015 | Partly, through a valuation lens | No: data ends in the 1990s | Partly |
| Re-tests its own winners | Yes: a dated durability test that demoted two of its own successes | No | No | No: several exemplars later collapsed in print | Not applicable |
| Ends in a decision | Yes: Scale Up, First Optimise, or Scale Back | A mindset | Benchmarks | A philosophy | A finding |
| Evidence inspectable | Yes: every source graded, every case block openable, every company ruling published including the twelve that were reversed | Proprietary | Partly; founder self-report | Not reproducible | Yes, in principle |
FIVE THINGS THE FOUR STUDIES ABOVE DO NOT DO
1. It studied the right population.
Findings drawn from Fortune-500 conglomerates or seed-stage startups reach a mid-market portfolio only by analogy, and an analogy is not evidence about your company. This study needs no such step across. Venture-backed, sponsor-backed, listed and founder-controlled companies all appear on both sides of the line, so no ownership structure is exempt from the findings.
2. It measured the outcome that pays.
Startup Genome’s finishing line is a US$50 million valuation. That is paper, and the correction of 2022 showed how far paper can move without any cash changing hands. Collins’s finishing line was fifteen years of share-price outperformance, and several of his eleven companies later failed. This study’s finishing line is value that held.
3. It counted the money, honestly.
Two measures, reported separately and never added together. One counts paper that was never cash. The other counts cash that never came back. Adding them would produce a larger number and a false one, so the study does not add them.
4. It covers the era every current portfolio is living through.
The period is 2020 to 2025: the cheap-capital surge, the megaround cluster, and the correction that followed. The cheap-capital years were tested as a rival explanation rather than assumed away, and the vintage of every case is on the register.
5. It ends in a decision, not a philosophy.
Bain ends in a mindset, Collins in a set of virtues, Startup Genome in benchmarks. This study ends in one of three verdicts a board can act on: Scale Up, First Optimise, or Scale Back.
THE STUDY RE-TESTS ITS OWN WINNERS
In 2001, Good to Great named eleven companies as enduringly great. Circuit City went bankrupt in 2009 and Fannie Mae passed into government conservatorship in 2008. Neither verdict was ever withdrawn.
This study re-tests its own success cases on a fixed date each year. When that test found two cases whose value and operating trajectory had reset, it moved them to the failure set, on the record, with the date on the entry.
We know of no other study of company growth that has demoted its own success cases. Re-testing is what makes the success set checkable, rather than fixed at the date it was published.
THE TWO QUESTIONS SCEPTICS ASK
“Only 182 companies?”
The number to compare is depth per company, not sample size. Bain’s 8,000 companies were screened by database, and the mechanism work rests on about one hundred interviews. Collins’s five-year programme, with twenty-one researchers, analysed eleven companies in depth. This study coded 182 companies, each against the same instrument, from graded primary sources, with the evidence block for every case openable and checkable. Measured on depth per company, this is the larger study and not the smaller one.
“Where is the famous name on the cover?”
A famous name asks you to trust the result. This study gives you the means to check it. Every source is graded, and a claim without primary verification is downgraded rather than promoted. Undisclosed private figures are recorded as ceilings and never guessed. Every company the Study ruled out is published with the rule that ruled it out, and so are the twelve rulings reversed after coding. Nothing was deleted to tidy the count. Twenty-eight cases were coded a second time, blind, against the definitions alone, and both reliability figures are published below. We are not aware of another growth study that publishes the rulings it reversed.
THE RELIABILITY FIGURES
Twenty-eight cases were coded a second time, from scratch, against the definitions alone, with the original scores withheld. Two standard measures of agreement were calculated on that re-code. Both are published.
Gwet’s AC1: 0.86. Cohen’s kappa: 0.34.
The two figures disagree, and the reason is known. Cohen’s kappa falls sharply when most cases score the same way on most items, which is what a study of this design produces. Gwet’s AC1 was built to correct for that effect. Common practice is to publish the higher figure and leave the lower one out. Both are published here.
WHAT THIS STUDY DOES NOT CLAIM
These limits are published with the Study rather than kept for anyone who asks.
It has no denominator. It does not know how many companies it could have examined, so it makes no claim to be complete, representative or the largest of anything.
It kept no screening record. The candidates considered before the rulings were never counted, and that record is not reconstructed now.
Two of the 78 failures are not yet fully coded, so every figure about what the failures exhibited is stated on 76.
For 64 of the 104 successes the financials are undisclosed. For those, “held” means no recorded reset of a quarter or more against peak and no distress event, as at 18 June 2026. It does not mean disclosed figures were checked. Thirty-two are affirmatively still scaling and eight are too recent to judge.
Fact verification is partial. The success workbook records most of its atomic claims as not yet publication-verified and treats them as not-in-source until that verification pass closes. The workbook says so on its own front sheet.
103 of the 182 had a named private equity operating partner. The rest were venture-backed, listed, or held by a corporate parent. Every case on the register records how far that company sits from the ten million to two hundred and fifty million pound band this work is calibrated to.
THE REGISTER
Every ruling in the Study is recorded and reversible, including the twelve that were changed after coding. All 199 companies and their rulings are published in full.
HOW TO CITE IT
Ward, Mark C. The Truthscaling Study, 2026 edition. Revenue Arc, 2026. revenuearc.co.uk/research
THE FULL RECORD
The account of what the Study found is a book called Truthscaling, out in 2027. The instrument the Study produced, the Scaling Diagnostic, is in use now.