Live desk, refreshed daily

AI model rankings: who's #1 right now

The leaderboard people actually cite, the price of each model, and this month's launches, on one page that updates itself.

As of September 13, 2026, Claude Fable 5.1 (max) from Anthropic leads LMArena's text leaderboard, with Claude Opus 5 (max) and Claude Opus 4.6 (high) close behind.

Among the name-brand models, GLM 5.3 Flash is the cheapest to run at $0.25 per million output tokens and Claude Fable 5.1 the most expensive at $50.00.

The top of the table is usually a statistical tie: when two models' margins overlap, the votes cannot separate them.

The leaderboard

LMArena's text leaderboard: people chat with two unnamed models side by side and vote for the better answer. One row per model; reasoning tiers collapsed to the best one.

#ModelMakerScoreVotes
1 Claude Fable 5.1 (max) Anthropic 1,508 ±9 5,783
2 Claude Opus 5 (max) Anthropic 1,505 ±5 20,706
3 Claude Opus 4.6 (high) Anthropic 1,503 ±4 71,993
4 Gemini 3.8 Flash (high) Google 1,495 ±9 5,076
5 Claude Fable 5 Anthropic 1,493 ±5 30,057
6 Gemini 3.7 Flash (high) Google 1,490 ±9 5,640
7 Claude Opus 4.7 (high) Anthropic 1,490 ±4 60,002
8 Muse Spark 1.3 (max) Meta 1,490 ±9 4,723
9 Muse Spark 1.2 (xhigh) Meta 1,489 ±11 3,227
10 Gemini 3.5 Flash (high) Google 1,482 ±4 38,257
11 Qwen3.8 (max) Alibaba 1,481 ±6 16,670
12 Muse Spark 1.1 Meta 1,480 ±5 27,615
13 Gemini 3.1 Pro Preview Google 1,480 ±3 106,951
14 Gemini 3 Pro Google 1,479 ±4 40,654
15 Gemini 3.6 Flash (high) Google 1,476 ±5 26,445
16 GLM 5.3 (max)Open weights Z.ai 1,475 ±6 10,960
17 Qwen3.7 Max Preview Alibaba 1,474 ±10 3,705
18 Muse Spark Meta 1,473 ±6 13,565
19 Kimi K3 (max)Open weights Moonshot 1,472 ±5 20,987
20 GLM 5.3 FlashOpen weights Z.ai 1,472 ±7 10,038
21 Qwen3.5 Max Preview Alibaba 1,471 ±5 21,476
22 GPT-5.5 (high) OpenAI 1,471 ±4 64,924
23 GPT-5.4 (high) OpenAI 1,470 ±4 60,537
24 ERNIE 5.1 Baidu 1,468 ±5 37,058
25 GLM 5.2 (max)Open weights Z.ai 1,467 ±5 36,798

LMArena text leaderboard, published September 13, 2026. Score is the arena rating with its 95% margin; rows whose margins overlap are a statistical tie.

Best open-weight models

The same LMArena votes, only models whose weights you can download and run yourself. The grey number is where each one sits on the overall board.

#ModelMakerScoreVotes
1 GLM 5.3 (max)#20 overall Z.ai 1,475 ±6 10,960
2 Kimi K3 (max)#23 overall Moonshot 1,472 ±5 20,987
3 GLM 5.3 Flash#24 overall Z.ai 1,472 ±7 10,038
4 GLM 5.2 (max)#29 overall Z.ai 1,467 ±5 36,798
5 Mimo V2.5 Pro#32 overall Xiaomi 1,465 ±4 60,919
6 GLM 5.1#33 overall Z.ai 1,462 ±4 48,901
7 Kimi K2.6#38 overall Moonshot 1,455 ±5 37,502
8 DeepSeek V4 Pro#43 overall DeepSeek 1,451 ±4 54,130
9 GLM 5#51 overall Z.ai 1,446 ±5 27,605
10 Kimi K2.5 (thinking)#53 overall Moonshot 1,446 ±4 70,513
11 DeepSeek V4 Pro High Preview#54 overall DeepSeek 1,445 ±4 51,485
12 Nvidia Nemotron 3 Ultra 550b A55B Nvfp4#55 overall Nvidia 1,445 ±7 10,693
13 DeepSeek V4 Pro High 20260813#58 overall DeepSeek 1,444 ±7 9,008
14 Gemma 4 31b#64 overall Google 1,442 ±8 5,894
15 Hy3#65 overall Tencent 1,441 ±8 8,047

LMArena text leaderboard, published September 13, 2026. Score is the arena rating with its 95% margin; rows whose margins overlap are a statistical tie.

Advertisement

What each model costs

List price per million tokens for the current model in each family, cheapest first. A million tokens is roughly 750,000 words of text in, or out.

ModelMakerInputOutput$1 buysContext
GLM 5.3 Flash Z.ai $0.075 $0.25 ≈3M words 1.3M
DeepSeek V4.1 Flash DeepSeek $0.15 $0.60 ≈1.3M words 1M
GPT-5.6 Luna OpenAI $0.20 $1.20 ≈625K words 1.1M
DeepSeek V4 Pro 0813 DeepSeek $0.66 $1.98 ≈378.8K words 1M
Gemini 3.8 Flash Google $0.75 $3.75 ≈200K words 1M
GPT-5.4 Mini OpenAI $0.75 $4.50 ≈166.7K words 400K
Claude Haiku 4.5 Anthropic $1.00 $5.00 ≈150K words 200K
GLM 5.3 Z.ai $1.40 $4.40 ≈170.5K words 1.3M
Grok 4.6 xAI $2.00 $6.00 ≈125K words 500K
GPT-5.6 Sol OpenAI $2.00 $10.00 ≈75K words 1.1M
Claude Sonnet 5 Anthropic $2.00 $10.00 ≈75K words 1M
GPT-5.6 Terra OpenAI $2.00 $12.00 ≈62.5K words 1.1M
Gemini 3.1 Pro Preview Google $2.00 $12.00 ≈62.5K words 1M
Kimi K3 Moonshot $2.648 $13.283 ≈56.5K words 1M
Claude Opus 5 Anthropic $5.00 $25.00 ≈30K words 1M
GPT-6 Astra OpenAI $10.00 $50.00 ≈15K words 1.1M
Claude Fable 5.1 Anthropic $10.00 $50.00 ≈15K words 1M

US dollars per million tokens, input and output. OpenRouter's public model list, checked September 15, 2026. These are list prices; check the maker's own pricing page before you budget, and note that thinking modes bill their reasoning as output.

New this month

Models from makers you have heard of, listed in the last 30 days, newest first.

ListedModelMaker$/M in / out
5 days agoSep 10, 2026 DeepSeek V4.1 Flash DeepSeek $0.15 / $0.60
11 days agoSep 4, 2026 GPT-6 Astra OpenAI $10.00 / $50.00
11 days agoSep 4, 2026 GPT-6 Astra Pro OpenAI $10.00 / $50.00
12 days agoSep 3, 2026 Qwen3.8 Max (0902) Alibaba $2.00 / $6.00
13 days agoSep 2, 2026 Muse Spark 1.3 Contributor Meta $0.10 / $0.20
13 days agoSep 2, 2026 Muse Spark 1.3 Meta $1.25 / $4.25
13 days agoSep 2, 2026 Gemini 3.8 Flash Google $0.75 / $3.75
14 days agoSep 1, 2026 Claude Fable 5.1 Anthropic $10.00 / $50.00
15 days agoAug 31, 2026 Granite 4.2 8B IBM $0.06 / $0.25
18 days agoAug 28, 2026 Hy4 preview Tencent $0.834 / $2.501
20 days agoAug 26, 2026 Qwen3.8 Flash Alibaba $0.15 / $0.47
20 days agoAug 26, 2026 GLM 5.3 Flash Z.ai $0.075 / $0.25
25 days agoAug 21, 2026 Muse Spark 1.2 Contributor Meta $0.10 / $0.20
25 days agoAug 21, 2026 DeepSeek V4 Flash Vision Exp DeepSeek $0.22 / $0.66
26 days agoAug 20, 2026 Hy-MT2-1.8B Tencent $0.044 / $0.177
26 days agoAug 20, 2026 Hy-MT2-30B-A3B Tencent $0.074 / $0.295
27 days agoAug 19, 2026 Hy-MT2-7B Tencent $0.074 / $0.295
28 days agoAug 18, 2026 GLM 5.3 Z.ai $1.40 / $4.40

Listings from the last 30 days by major makers, checked September 15, 2026. Source: OpenRouter's model list.

How the ranking works

LMArena (the former Chatbot Arena from UC Berkeley's LMSYS group) shows a visitor two unnamed models answering the same prompt and asks which answer was better. Millions of those votes feed a Bradley-Terry rating, the same family of math as chess Elo, so a model's score means how often it wins against the field. The 95% margin next to each score is the uncertainty: a model with 5,000 votes has a wider margin than one with 70,000. We show LMArena's default overall text board, one row per model. LMArena lists the same model several times at different reasoning settings ("high", "max"); we keep the best-ranked one and note the setting in brackets.

How to read the prices

AI companies bill by the token, and a token is about three-quarters of an English word, so a million tokens is roughly 750,000 words. Input is what you send (your question, the document you paste); output is what the model writes back. Output is the expensive side, and "thinking" or reasoning modes bill their hidden reasoning as output too, so a hard question can cost several times what the answer length suggests. The prices here are the list prices on OpenRouter's public catalogue, which mirror the makers' own rates; OpenAI, Anthropic and Google also sell batch tiers at half price for jobs that can wait. Always confirm on the maker's pricing page before you budget.

Advertisement

Which one should you actually use?

If you use AI through a chat app rather than an API, the ranking matters less than it looks. The top ten sit within a few points of each other, the free tiers of ChatGPT, Claude and Gemini all run models from that group, and the differences that matter day to day are the app's features: memory, file handling, voice, and whether your workplace allows it. Pick by the app you like and the price you pay. If you are paying per token, the cost table is the page to read: the gap between the cheapest and priciest name-brand model is more than a hundredfold, and for routine work the cheap end is usually good enough.

What this page does not tell you

A head-to-head vote measures which answer people preferred in a chat window. It does not measure coding accuracy, long-document work, image understanding, speed, or how a model behaves inside an agent. LMArena runs separate boards for several of those, and benchmark suites such as Artificial Analysis publish their own indexes. Treat this table as the popular vote, not the whole election.

Questions people ask

Which AI model is the best right now?

By LMArena's blind head-to-head votes, the model at the top of the table on this page. Check the margin column: when the top few overlap, they are tied for practical purposes, and the order can swap from one week to the next.

How often does this page update?

Once a day. LMArena publishes new leaderboard snapshots every few days and OpenRouter's price list changes whenever a maker changes a price or lists a model; the desk picks both up on its next daily run and stamps the date it last checked under each table.

What does 'open weights' mean?

The maker has published the model's weights under a licence that lets you download and run it yourself, on your own hardware or a cloud you rent, instead of only through the maker's API. Proprietary models can be used only through the company that made them.

Why is a model listed twice on LMArena but once here?

LMArena scores the same model at different reasoning settings as separate rows. We keep the best-ranked setting for each model so the table reads as a list of models, and show the setting in brackets after the name.

Is the cheapest model good enough?

For most everyday writing, summarising and question-answering, yes: the cheap tier from a major maker is usually a smaller version of that maker's flagship. Pay for the expensive tier when you need long, careful reasoning, complex code, or the best possible answer on the first try.

Keep reading

Rankings data: LMArena. Price and listing data: OpenRouter. Both are credited on each table; this site is not affiliated with either.

Advertisement