In March 2023, the model at the top of the AI leaderboard billed $37.50 per million blended tokens. Twenty-nine months later the throne cost $3.44 — a soft drink instead of a steak dinner. Then, for the first time in the industry’s short history, the price of being best started going back up.

Everyone tracks what today’s models cost. Almost nobody keeps the ledger of what leaderboard position cost at the moment it was held — the price tag attached to a seat on the podium, quarter by quarter, as the podium itself moved. We built that ledger: 59 models, 65 dated list-price observations, and 35 LMArena Elo snapshots spanning March 2023 to August 2026, each tagged with the rating methodology in force when it was measured. The dataset lives in our own database and every chart below is drawn from it — you can explore it interactively or read on for the three charts that tell the story.

At a glance

  • The blended list price of the Arena #1 model fell about 11× — GPT-4’s $37.50 to GPT-5’s $3.44 — between March 2023 and August 2025, then rose ~6× to Claude Fable 5’s $20 by July 2026.
  • GPT-4o-class capability (the ~1300-Elo band) went from $7.50 to $0.48 per blended million in seven months — 16× — and the 2026 floor tier runs below $0.20.
  • LMArena’s rating system was rebuilt three times along the way; ratings are re-baselined at each break, so no chart here draws a line across one.
  • Per-token list price flatters reasoning models, which burn far more tokens per task. This is a list-price ledger, not a cost-per-task benchmark — the one caveat that frames everything below.

The chart nobody keeps

The deflation of machine intelligence is well documented in the aggregate. Andreessen Horowitz’s “LLMflation” essay pegs it at “the cost is decreasing by 10x every year” for equivalent performance. Epoch AI, measuring the price to hit fixed benchmark scores, found declines anywhere from 9× to 900× per year depending on the milestone. And Artificial Analysis now publishes the modern metric — dollars to complete a standardized battery of tasks — precisely because per-token price stopped being a fair capability measure once reasoning models started thinking in tens of millions of tokens.

Those are all cost-at-fixed-capability studies, and they answer their question well. Ours is different and more like a market question: the leaderboard is a ranking of seats, the seats have list prices, and both the ranking and the prices carry dates. What did the view from the top cost in 2023? What does it cost now? Who briefly sold frontier capability at a discount, and who charged a premium for it? Per-token list price — for all its flaws as a capability metric — is exactly the right unit for that ledger, because it is the number the market actually posted at the time.

Three eras, three rulers

One honesty requirement before any trend line: Arena Elo is not one continuous scale. The rating pipeline was rebuilt three times — online Elo gave way to a Bradley-Terry model in January 2024, style control became the default (with an offset re-scaling) in May 2025, and a frequency re-weighting followed in July 2025. Each rebuild re-baselines the numbers. A model rated 1250 in 2024 and a model rated 1250 in 2026 do not share a yardstick, and any chart that regresses across those breaks is manufacturing a trend out of a methodology change.

So the scatter below keeps the eras apart. Within each panel, the teal line traces the cheapest price at any given rating — the value frontier of its moment.

Three-panel scatter of blended price versus Arena Elo, one panel per rating era, with a teal frontier line tracing the cheapest model at each rating level

The frontier marches down and to the right inside every era. GPT-4.5, alone at the top of the middle panel at $93.75, is the most expensive off-frontier outlier in the dataset — withdrawn from the API within months.

The recurring pattern inside each panel: a provider plants a flag at the top right, and within months, cheaper entrants pull the frontier line down underneath it. In the middle panel, Gemini 2.0 Flash and DeepSeek’s V3 and R1 drag the cheapest path to ~1400 Elo below a dollar. In the right panel the same shape is forming again — DeepSeek V4 Pro holds the cheap corner at $0.65 while the $10–$20 cluster crowds the top.

The price of the throne

Chart the #1 seat itself and the story gets a plot twist.

Step chart of the blended list price of the Arena number-one model from 2023 to 2026, falling from $37.50 to $3.44 and then rising to $20

The throne got 11× cheaper in 29 months, then 6× pricier in twelve. Dashed verticals mark the two Arena methodology breaks.

The long slide is the familiar part: GPT-4 at $37.50, Claude 3 Opus briefly at $30, GPT-4o at $7.50 and then $4.38 after its October 2024 cut, Gemini 2.5 Pro and then GPT-5 at $3.44. GPT-5’s August 2025 launch price — $1.25 in, $10 out — was the correction that reset the whole market’s expectations of what the frontier should cost.

The rebound is the new part. Gemini 3 Pro took the crown at $4.50. Claude Opus 4.6 took it back at $10. Claude Fable 5 holds it today at $20 — and OpenAI’s own price ladder moved the same direction, with GPT-5.5 and 5.6 Sol at double GPT-5’s per-token rate. The economics underneath are no mystery: reasoning models shifted the product from tokens to work, and the labs began pricing the premium tier accordingly. Kimi K3 — the first Chinese open-weight flagship to raise prices, at three to five times its predecessor — suggests the rebound is structural, not a coincidence of two Western labs.

Two things can be true at once, and the research behind this dataset is emphatic about the second one: per-token prices at the frontier rose through 2026, and the cost to complete a fixed batch of hard tasks kept falling, because newer models spend their tokens far more efficiently. A ledger of list prices records the first fact; Artificial Analysis’s per-task index records the second. Read together, they say the frontier stopped competing on token price and started competing on delivered work.

Yesterday’s frontier, discounted

The throne is the headline, but the deeper economic engine is what happens to last year’s frontier capability.

Step chart showing the cheapest blended price for a model rated near 1300 Elo falling from $7.50 to $0.48 in seven months, with a dashed attributed continuation below $0.20

GPT-4o-class capability: $7.50 in May 2024, $0.48 by December 2024. The dashed continuation marks the research-attributed 2026 floor tier, which no longer carries a current Arena listing at this band.

In May 2024, roughly-1300-Elo capability meant buying OpenAI’s flagship at $7.50 blended. Seven months later DeepSeek V3 posted a ~1315 rating at $0.48 — sixteen times cheaper, and the moment the open-weight discount stopped being a rounding error and became the market’s pricing floor. By 2026 that class of capability had fallen off the leaderboard’s radar entirely and into the utility tier: DeepSeek V4 Flash lists at $0.18 blended, and Google’s Flash-Lite and OpenAI’s nano tiers bracket it. Yesterday’s miracle is today’s commodity; the ledger just puts numbers and dates on how fast the conveyor belt runs.

Explore the data

The full dataset — models, dated price points including mid-life cuts and raises, era-tagged Elo snapshots, Artificial Analysis cross-references, and a simplified Arena-#1 timeline — is browsable in the interactive dashboard, with filters by provider, era, and date. It is served from the same database this site runs on.

Method notes, briefly. Blended price is (3 × input + 1 × output) ÷ 4, the conventional chat-workload weighting; it understates output-heavy reasoning workloads by design, which is one more reason this is a list-price ledger rather than a spend estimate. Prices are base list rates at or below 200K context — no cache, batch, or off-peak rates. Elo figures marked approximate in the dataset came from ranges or third-party snapshot trackers (many 2026 figures vary ±5–15 points between trackers); a handful of prices the underlying research could not confirm are flagged unverified rather than silently dropped. The Arena-#1 timeline is deliberately simplified from eighteen documented crown changes; refresh-level churn is folded into the reigning model family.


Sourcing: compiled from a commissioned deep-research dataset (August 2026) reconciling provider pricing pages and launch announcements, LMArena leaderboard data and BenchLM snapshot history, and Artificial Analysis model pages; deflation prior art from Epoch AI (“LLM inference price trends”) and Andreessen Horowitz (“LLMflation”, November 2024), both quoted verbatim from source. Dataset, loader, and chart code are in the aistatus repository; the interactive dashboard runs on Deltabase.