Live desk, refreshed daily
AI model rankings: who's #1 right now
The leaderboard people actually cite, the price of each model, and this month's launches, on one page that updates itself.
As of September 13, 2026, Claude Fable 5.1 (max) from Anthropic leads LMArena's text leaderboard, with Claude Opus 5 (max) and Claude Opus 4.6 (high) close behind.
Among the name-brand models, GLM 5.3 Flash is the cheapest to run at $0.25 per million output tokens and Claude Fable 5.1 the most expensive at $50.00.
The top of the table is usually a statistical tie: when two models' margins overlap, the votes cannot separate them.
The leaderboard
LMArena's text leaderboard: people chat with two unnamed models side by side and vote for the better answer. One row per model; reasoning tiers collapsed to the best one.
| # | Model | Maker | Score | Votes |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 (max) | Anthropic | 1,508 ±9 | 5,783 |
| 2 | Claude Opus 5 (max) | Anthropic | 1,505 ±5 | 20,706 |
| 3 | Claude Opus 4.6 (high) | Anthropic | 1,503 ±4 | 71,993 |
| 4 | Gemini 3.8 Flash (high) | 1,495 ±9 | 5,076 | |
| 5 | Claude Fable 5 | Anthropic | 1,493 ±5 | 30,057 |
| 6 | Gemini 3.7 Flash (high) | 1,490 ±9 | 5,640 | |
| 7 | Claude Opus 4.7 (high) | Anthropic | 1,490 ±4 | 60,002 |
| 8 | Muse Spark 1.3 (max) | Meta | 1,490 ±9 | 4,723 |
| 9 | Muse Spark 1.2 (xhigh) | Meta | 1,489 ±11 | 3,227 |
| 10 | Gemini 3.5 Flash (high) | 1,482 ±4 | 38,257 | |
| 11 | Qwen3.8 (max) | Alibaba | 1,481 ±6 | 16,670 |
| 12 | Muse Spark 1.1 | Meta | 1,480 ±5 | 27,615 |
| 13 | Gemini 3.1 Pro Preview | 1,480 ±3 | 106,951 | |
| 14 | Gemini 3 Pro | 1,479 ±4 | 40,654 | |
| 15 | Gemini 3.6 Flash (high) | 1,476 ±5 | 26,445 | |
| 16 | GLM 5.3 (max)Open weights | Z.ai | 1,475 ±6 | 10,960 |
| 17 | Qwen3.7 Max Preview | Alibaba | 1,474 ±10 | 3,705 |
| 18 | Muse Spark | Meta | 1,473 ±6 | 13,565 |
| 19 | Kimi K3 (max)Open weights | Moonshot | 1,472 ±5 | 20,987 |
| 20 | GLM 5.3 FlashOpen weights | Z.ai | 1,472 ±7 | 10,038 |
| 21 | Qwen3.5 Max Preview | Alibaba | 1,471 ±5 | 21,476 |
| 22 | GPT-5.5 (high) | OpenAI | 1,471 ±4 | 64,924 |
| 23 | GPT-5.4 (high) | OpenAI | 1,470 ±4 | 60,537 |
| 24 | ERNIE 5.1 | Baidu | 1,468 ±5 | 37,058 |
| 25 | GLM 5.2 (max)Open weights | Z.ai | 1,467 ±5 | 36,798 |
LMArena text leaderboard, published September 13, 2026. Score is the arena rating with its 95% margin; rows whose margins overlap are a statistical tie.
Best open-weight models
The same LMArena votes, only models whose weights you can download and run yourself. The grey number is where each one sits on the overall board.
| # | Model | Maker | Score | Votes |
|---|---|---|---|---|
| 1 | GLM 5.3 (max)#20 overall | Z.ai | 1,475 ±6 | 10,960 |
| 2 | Kimi K3 (max)#23 overall | Moonshot | 1,472 ±5 | 20,987 |
| 3 | GLM 5.3 Flash#24 overall | Z.ai | 1,472 ±7 | 10,038 |
| 4 | GLM 5.2 (max)#29 overall | Z.ai | 1,467 ±5 | 36,798 |
| 5 | Mimo V2.5 Pro#32 overall | Xiaomi | 1,465 ±4 | 60,919 |
| 6 | GLM 5.1#33 overall | Z.ai | 1,462 ±4 | 48,901 |
| 7 | Kimi K2.6#38 overall | Moonshot | 1,455 ±5 | 37,502 |
| 8 | DeepSeek V4 Pro#43 overall | DeepSeek | 1,451 ±4 | 54,130 |
| 9 | GLM 5#51 overall | Z.ai | 1,446 ±5 | 27,605 |
| 10 | Kimi K2.5 (thinking)#53 overall | Moonshot | 1,446 ±4 | 70,513 |
| 11 | DeepSeek V4 Pro High Preview#54 overall | DeepSeek | 1,445 ±4 | 51,485 |
| 12 | Nvidia Nemotron 3 Ultra 550b A55B Nvfp4#55 overall | Nvidia | 1,445 ±7 | 10,693 |
| 13 | DeepSeek V4 Pro High 20260813#58 overall | DeepSeek | 1,444 ±7 | 9,008 |
| 14 | Gemma 4 31b#64 overall | 1,442 ±8 | 5,894 | |
| 15 | Hy3#65 overall | Tencent | 1,441 ±8 | 8,047 |
LMArena text leaderboard, published September 13, 2026. Score is the arena rating with its 95% margin; rows whose margins overlap are a statistical tie.
What each model costs
List price per million tokens for the current model in each family, cheapest first. A million tokens is roughly 750,000 words of text in, or out.
| Model | Maker | Input | Output | $1 buys | Context |
|---|---|---|---|---|---|
| GLM 5.3 Flash | Z.ai | $0.075 | $0.25 | ≈3M words | 1.3M |
| DeepSeek V4.1 Flash | DeepSeek | $0.15 | $0.60 | ≈1.3M words | 1M |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | ≈625K words | 1.1M |
| DeepSeek V4 Pro 0813 | DeepSeek | $0.66 | $1.98 | ≈378.8K words | 1M |
| Gemini 3.8 Flash | $0.75 | $3.75 | ≈200K words | 1M | |
| GPT-5.4 Mini | OpenAI | $0.75 | $4.50 | ≈166.7K words | 400K |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | ≈150K words | 200K |
| GLM 5.3 | Z.ai | $1.40 | $4.40 | ≈170.5K words | 1.3M |
| Grok 4.6 | xAI | $2.00 | $6.00 | ≈125K words | 500K |
| GPT-5.6 Sol | OpenAI | $2.00 | $10.00 | ≈75K words | 1.1M |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | ≈75K words | 1M |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | ≈62.5K words | 1.1M |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 | ≈62.5K words | 1M | |
| Kimi K3 | Moonshot | $2.648 | $13.283 | ≈56.5K words | 1M |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | ≈30K words | 1M |
| GPT-6 Astra | OpenAI | $10.00 | $50.00 | ≈15K words | 1.1M |
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | ≈15K words | 1M |
US dollars per million tokens, input and output. OpenRouter's public model list, checked September 15, 2026. These are list prices; check the maker's own pricing page before you budget, and note that thinking modes bill their reasoning as output.
New this month
Models from makers you have heard of, listed in the last 30 days, newest first.
| Listed | Model | Maker | $/M in / out |
|---|---|---|---|
| 5 days agoSep 10, 2026 | DeepSeek V4.1 Flash | DeepSeek | $0.15 / $0.60 |
| 11 days agoSep 4, 2026 | GPT-6 Astra | OpenAI | $10.00 / $50.00 |
| 11 days agoSep 4, 2026 | GPT-6 Astra Pro | OpenAI | $10.00 / $50.00 |
| 12 days agoSep 3, 2026 | Qwen3.8 Max (0902) | Alibaba | $2.00 / $6.00 |
| 13 days agoSep 2, 2026 | Muse Spark 1.3 Contributor | Meta | $0.10 / $0.20 |
| 13 days agoSep 2, 2026 | Muse Spark 1.3 | Meta | $1.25 / $4.25 |
| 13 days agoSep 2, 2026 | Gemini 3.8 Flash | $0.75 / $3.75 | |
| 14 days agoSep 1, 2026 | Claude Fable 5.1 | Anthropic | $10.00 / $50.00 |
| 15 days agoAug 31, 2026 | Granite 4.2 8B | IBM | $0.06 / $0.25 |
| 18 days agoAug 28, 2026 | Hy4 preview | Tencent | $0.834 / $2.501 |
| 20 days agoAug 26, 2026 | Qwen3.8 Flash | Alibaba | $0.15 / $0.47 |
| 20 days agoAug 26, 2026 | GLM 5.3 Flash | Z.ai | $0.075 / $0.25 |
| 25 days agoAug 21, 2026 | Muse Spark 1.2 Contributor | Meta | $0.10 / $0.20 |
| 25 days agoAug 21, 2026 | DeepSeek V4 Flash Vision Exp | DeepSeek | $0.22 / $0.66 |
| 26 days agoAug 20, 2026 | Hy-MT2-1.8B | Tencent | $0.044 / $0.177 |
| 26 days agoAug 20, 2026 | Hy-MT2-30B-A3B | Tencent | $0.074 / $0.295 |
| 27 days agoAug 19, 2026 | Hy-MT2-7B | Tencent | $0.074 / $0.295 |
| 28 days agoAug 18, 2026 | GLM 5.3 | Z.ai | $1.40 / $4.40 |
No new models from major makers in the last 30 days.
Listings from the last 30 days by major makers, checked September 15, 2026. Source: OpenRouter's model list.
How the ranking works
LMArena (the former Chatbot Arena from UC Berkeley's LMSYS group) shows a visitor two unnamed models answering the same prompt and asks which answer was better. Millions of those votes feed a Bradley-Terry rating, the same family of math as chess Elo, so a model's score means how often it wins against the field. The 95% margin next to each score is the uncertainty: a model with 5,000 votes has a wider margin than one with 70,000. We show LMArena's default overall text board, one row per model. LMArena lists the same model several times at different reasoning settings ("high", "max"); we keep the best-ranked one and note the setting in brackets.
How to read the prices
AI companies bill by the token, and a token is about three-quarters of an English word, so a million tokens is roughly 750,000 words. Input is what you send (your question, the document you paste); output is what the model writes back. Output is the expensive side, and "thinking" or reasoning modes bill their hidden reasoning as output too, so a hard question can cost several times what the answer length suggests. The prices here are the list prices on OpenRouter's public catalogue, which mirror the makers' own rates; OpenAI, Anthropic and Google also sell batch tiers at half price for jobs that can wait. Always confirm on the maker's pricing page before you budget.
Which one should you actually use?
If you use AI through a chat app rather than an API, the ranking matters less than it looks. The top ten sit within a few points of each other, the free tiers of ChatGPT, Claude and Gemini all run models from that group, and the differences that matter day to day are the app's features: memory, file handling, voice, and whether your workplace allows it. Pick by the app you like and the price you pay. If you are paying per token, the cost table is the page to read: the gap between the cheapest and priciest name-brand model is more than a hundredfold, and for routine work the cheap end is usually good enough.
What this page does not tell you
A head-to-head vote measures which answer people preferred in a chat window. It does not measure coding accuracy, long-document work, image understanding, speed, or how a model behaves inside an agent. LMArena runs separate boards for several of those, and benchmark suites such as Artificial Analysis publish their own indexes. Treat this table as the popular vote, not the whole election.
Questions people ask
Which AI model is the best right now?
By LMArena's blind head-to-head votes, the model at the top of the table on this page. Check the margin column: when the top few overlap, they are tied for practical purposes, and the order can swap from one week to the next.
How often does this page update?
Once a day. LMArena publishes new leaderboard snapshots every few days and OpenRouter's price list changes whenever a maker changes a price or lists a model; the desk picks both up on its next daily run and stamps the date it last checked under each table.
What does 'open weights' mean?
The maker has published the model's weights under a licence that lets you download and run it yourself, on your own hardware or a cloud you rent, instead of only through the maker's API. Proprietary models can be used only through the company that made them.
Why is a model listed twice on LMArena but once here?
LMArena scores the same model at different reasoning settings as separate rows. We keep the best-ranked setting for each model so the table reads as a list of models, and show the setting in brackets after the name.
Is the cheapest model good enough?
For most everyday writing, summarising and question-answering, yes: the cheap tier from a major maker is usually a smaller version of that maker's flagship. Pay for the expensive tier when you need long, careful reasoning, complex code, or the best possible answer on the first try.
Keep reading
- GPT-6 Astra is out: what changed
- Claude Fable 5.1: the model that hunts bugs
- Why all AIs sound the same
- Why Opus 4.6 ended the copilot era
- Reasoning models explained: AI that thinks before it speaks
Rankings data: LMArena. Price and listing data: OpenRouter. Both are credited on each table; this site is not affiliated with either.