Every notable model since ChatGPT, scored on one capability scale. Thirteen labs. One line that refuses to flatten. Press play and watch the race run.
Every line is a lab's best model to date. Every dot is a release. The ringed dots are the moments that changed the game. Drag across the chart to zoom, all the way down to single days: more labels appear as you go. Below the axis runs the meanwhile lane: grey diamonds are the business (crashes, buyouts, kill switches), terracotta dots are new ways of working (Claude code, codex, cowork).
| date | model | lab | index | weights | why it mattered |
|---|
The frontier is American, the capex is monstrous (some $725b of hyperscaler spend this year), and the state is now inside the release pipeline: gpt-5.6 shipped on Washington's schedule, fable 5 went dark on Washington's order. The worry is the red cluster: five Chinese labs, six points back, at a sixth of the price.
Cut off from the best chips, China went open and efficient. The open crown has changed hands five times in eighteen months, and stanford's AI index puts the model performance gap at 2.7% for roughly 23x less spend. Open weights are not charity. They are distribution, and they are working.
One European line, thirty-six points off the lead, and a June reminder that rented intelligence can be unplugged in 90 minutes. In July 2024 Mistral sat under three points off the frontier; the reasoning era turned the race into a capital game. The counterplay is compute and openness, and we mapped the compute half: the compute gap.
The honest read: Europe did not lose the frontier race. It stopped being able to enter it around the time the game changed from clever pretraining to industrial-scale reinforcement learning. Mistral saw this early and pivoted: ocr 4 for documents, leanstral for formal verification, robostral for robots, a roughly €20b valuation built on being the model factory for European industry rather than chasing the staircase.
Meanwhile the new European bets refuse to play the same game. Yann lecun's ami raised $1b to build world models. David silver's ineffable intelligence raised $1.1b to make models that learn without human data. The europa consortium is training a 400b-parameter open model across all 24 EU languages on EuroHPC compute, and openeurollm ships its first fully open weights this month. And after Washington switched off fable 5 for 19 days in June, sovereignty stopped being an abstraction in Brussels.
The chart does not say Europe is out. It says Europe is betting the next chart looks different. Closing the gap.
When Washington issued its export directive on June 12, Anthropic took fable 5, the highest point on this entire chart, dark across the planet in about an hour and a half. 19 days later it came back. Every government in Europe took the same note: a nation that rents its intelligence can be unplugged overnight.
The y-axis is the ctg capability index: the artificial analysis intelligence index v4.1 (June 2026) where a model is listed, and our estimate on the same scale where it is not. Estimates lean on contemporaneous index versions, headline benchmarks (gpqa diamond, swe-bench verified, aime, humanity's last exam, arc-agi-2) and launch-window comparisons, and every estimate is marked in its tooltip. Cursor's composer models are coding specialists scored partly from the separate coding agent index, so treat their dots as directional. The five Chinese open-weights labs share a family of reds on purpose: read them as one bloc with five names.
Dates are first public availability or announcement. Each line is a lab's best score to date, which is why lines only step up: dots below a line are releases that did not move that lab's frontier. Sources: lab announcements, artificial analysis, lmarena, and the ctg. daily briefings, April to September 2026. Compiled September 17, 2026. Spotted an error? Tell us at ctg.show.