The Shrinking Reign of the Flagship
2026-08-27draft — publish-gated on ethics clearance
Draft 1 — 2026-08-27. Third in the pricing-ledger series. Publish-gated on
ethics clearance. Reign dates are the ledger-validated lineage from analyses
#1–2 (dataset/flagship_reigns.csv); rerun via
scripts/analysis_flagship_reign.py.
Here’s a number that should be in every AI procurement deck and isn’t: the median reign of an OpenAI flagship model — launch until its successor launches — was 9.4 months before 2025. Since January 2025 it’s been 2.8 months.
davinci sat at the top for roughly 27 months. GPT-4 got about eight. GPT-4o got eleven. And then the wheel started spinning: GPT-4.1 held the spot for under four months, GPT-5 for four, 5.2 for three, 5.4 for seven weeks. The current flagship, GPT-5.6, is seven weeks in as this is written — which, at the 2026 cadence, makes it middle-aged.
Why this is a procurement problem, not trivia
Three clocks run against a team adopting a model:
- The eval clock. A serious evaluation-plus-integration cycle at a regulated company runs 6–12 weeks. At a 2.8-month flagship cadence, the model you certified is one generation behind by the time you ship — the certification treadmill never catches the market.
- The pricing clock. As analysis #1 showed, successive flagships now arrive at higher prices ($1.25 → $1.75 → $2.50 → $5.00). Staying current and staying on budget became different strategies in 2025.
- The contract clock. An annual commitment now spans ~4 flagship generations. Whatever you negotiated, you negotiated for a catalog that won’t exist by renewal.
The strategic implication cuts against instinct: chasing the frontier is now a treadmill strategy, and deliberately running N-1 — yesterday’s flagship at a discount — is the emerging rational default. The reign data is why the N-1 tier exists at all: providers keep prior generations listed (GPT-4 Turbo’s pricing stayed listed for 34+ months) precisely because a growing share of buyers can’t re-certify every quarter.
What this piece is not claiming
Reign ends when the next flagship launches — not when a model shuts down. Old models typically remain available long after they leave the top spot; primacy and availability are different lifetimes. The companion question — how fast providers actually shut down models, the number that breaks production systems — requires provider deprecation pages, which we’re assembling separately; the community-ledger removal events we checked are dominated by catalog housekeeping (44% land on three cleanup dates) and would be dishonest to present as deprecations. When that dataset lands, this piece gets its sequel.
One more honesty note: “flagship” here is OpenAI’s top general-purpose API model, the same lineage as analyses #1–2. Generations are the unit — same-generation repricings and point-releases (5.1, 5.3) don’t restart the clock. Change the definitions and the exact medians move; the collapse from double-digit months to single digits survives any reasonable definition.
The forecastable question
This sets up the first number we intend to put a probability on: does the 2026 cadence hold? Ten flagship generations in twenty months is either the new physics of the frontier — continuous-delivery model development — or a competitive spasm that consolidates. The reign series is now short enough to update monthly and long enough to score a forecast against. Watch this space.
Method: reign = launch-to-successor-launch over the ledger-validated
flagship lineage. Current flagship’s bar is censored (marked “ongoing”).
Median split at 2025-01-01; davinci enters the pre-2025 median at its
observed window only (understating the old regime, i.e., biasing against
our own conclusion). Audit table: dataset/flagship_reigns.csv.