Wes Ellis./ a personal notebook
Technology. Stories. Side projects.
A few things worth writing down.
← Back to Writing Journal

Writing Journal

Measuring the Flourish: What a 50-Case Benchmark Held Up and What It Didn't

A hand filling in a form on a clipboard with a pen.

Part 2 of the thread The Art of Wander

THE SHORT VERSION4 points
  • The Flourishing Systems Model scores seven variables per era. Fifty historical cases were frozen before any serious scoring.
  • Scoring is period-locked and firewalled: each interval is judged using only evidence from its own time, by a scorer that never sees the rest of the case.
  • The reliability checks were humbling. With scorers doing their own research, the composite score came in at alpha 0.59, which isn't publishable.
  • The fix was separating research from scoring. With a shared evidence packet, agreement jumped to 0.94, but only on eight intervals so far.

The book I'm working on rests on one claim: things reach a flourish, and then people optimize them to death. I wrote about the idea itself in The Art of Wander. This note is about the harder part, trying to measure it, and what happened when I made the measurement try to prove me wrong. Some of it held up. A lot of my early confidence didn't.

The model and the fifty cases

The Flourishing Systems Model (FSM) scores a system on seven things, each 0 to 100, relative to other systems of its kind in the same era: quality, diversity, accessibility, resilience, participant agency, complexity and extraction. The last two count against it. Version 1.0 combines them into one flourish score, F.

The first dozen cases (Disney, McDonald's, Kodak, Apple and so on) were scored in the original conversation by a model that already knew every ending. All twelve passed. That's a warning sign, not a result, so they're now a quarantined development set that never counts toward accuracy.

The benchmark is fifty cases, numbered and frozen before scoring, with no swapping out inconvenient ones. They're scored in a fixed random order so the famous collapses don't go first. Ten are sealed holdouts: score up to a cutoff year, seal a prediction, then look at what happened.

How the scoring works

The rules live in a frozen protocol, with every change logged as a numbered amendment. The parts that matter most:

  • Period lock. Score each interval as if you're standing at its end. A 2012 retrospective on Nokia can't inform a 2006 score. There's even a banned-phrase list, "golden age," "in hindsight," "never recovered," because hindsight sneaks in through vocabulary.
  • Three layers, never mixed. Evidence, raw variable scores, and derived numbers are kept apart. A scoring record with an F value in it is void. That's what lets me change the formula later and rerun all fifty with a script instead of a year of research.
  • A firewall. Each interval goes to a fresh scorer with a fixed prompt. It never sees other intervals, curves, or what I expect.
  • Retrieve before you cite. No citing a document the scorer didn't open, and numbers need a page or quote.
  • Null is a valid answer. An invented number costs more than ten honest blanks.

What didn't hold up

The formula had an arithmetic bug. The weights sum to 1.15, not 1.00. Participant agency got added mid-conversation and nothing was rebalanced. So F was never a 0-100 scale, and the fixed flourish threshold of 85 I'd proposed is unreachable. It also means the conversation's scores, Apple at 94 and so on, were narrative numbers wearing an equation. The formula stays frozen anyway, because fixing it mid-run would be fitting the model to the results.

The composite score wasn't reliable. Twenty-nine intervals were scored twice, independently. Using Krippendorff's alpha, the usual agreement statistic, extraction came out at 0.817. Everything else was shakier, and the composite F landed at 0.593. For reference, the usual bar for drawing even tentative conclusions is 0.667.

Some cases can't be scored this way. Three of Coney Island's five intervals rested entirely on after-the-fact sources, because the 1926 and 1950 record isn't digitized. Cases like that are now marked not scorable rather than scored to a lower standard, and it could be a sixth of the cohort.

Heads up

The protocol requires dumb baselines, like "peak = midpoint of the window," and they aren't in the logs yet. Until FSM beats them there's no accuracy claim, and the conversation's 75-80% target is off the table.

What held up

The research step was the noise, not the model. For eight intervals, one researcher built a shared evidence packet and three scorers judged only from it. Mean agreement across the seven variables went from 0.619 to 0.892, and the composite went from 0.584 to 0.942. I'd been measuring the variability of literature searches and blaming the instrument. The honest caveat: agreement isn't accuracy, and a one-sided packet would produce three scorers confidently agreeing on the wrong answer.

The direction of extraction tells two kinds of decline apart. When a system is being harvested, extraction rises while it declines. When it's being displaced by something better, extraction falls. That sign was right on 4 of 4 cases where the outcome is known. Video rental is the clean example: a plateau, then 45.3 and 41.0, with extraction dropping. Pagers are the case I'm most sure of historically, and their shift sits inside the error band, so they're labeled unconfirmed.

Fabricated citations went away. The first outside scorer produced 4 impossible citations in 143. After the retrieve-before-you-cite rule, the next batches had 0 in 14, then 0 in 177. Meanwhile my own checking script falsely accused good records three times. Check the checker.

Dropping the weakest variable barely moves anything. Without participant agency, five of six peak years stay put and the weights sum to exactly 1.00. It's a 2.0 candidate, not adopted.

The most interesting single case is the Las Vegas Strip: record gaming revenue of $8.9 billion in 2023, and the instrument reads falling flourish with extraction climbing to 75. If it measures anything real, that's the thesis caught in the act.

Keeping an archive this traceable was its own project. Every claim cites the conversation turn it came from, using the Conversation Project Kit I built for exactly that.