Open-weight LLM landscape

Comparing benchmark scores, pricing, and energy use across open-weight language models and their providers.

Built with GLM-5.2 on Neuralwatt · 1.60 kWh · 193.7 g CO₂ · Data: Artificial Analysis · models.dev · Neuralwatt · OpenRouter · generated 2026-09-02 07:17 UTC

§ Model comparison — 44 open-weight models

Zhipu AIAlibabaMoonshot AIDeepSeekMiniMax (minimax.io)Thinking MachinesTencentXiaomiNvidiaMetaStepFun (China)MistralGoogleOpenAICohereMistralArcee AI
X axis
Y axis
Inference provider availability
KV cache
Provider type

Blended cost = (7·cache + 2·input + 1·output)/10. When a provider omits cache-hit pricing, the input price is used as an upper bound (cache hits are never more expensive than a regular input token). 44 models shown: open-weight (via models.dev), with an Agentic Index score AND input + output pricing. Models without input or output pricing are excluded. Neuralwatt energy values in tooltips are measured when an AA model matches a NW model; otherwise the tooltip shows an estimated energy derived from the NW cost ↔ energy regression (see the Neuralwatt scatter below).

§ Neuralwatt — Energy use vs. cost (12 models, 16k–64k band)

X: NW blended cost per 1M tokens = (7·cache + 2·input + 1·output)/10 (USD, from Neuralwatt pricing). Y: energy per request at the 16k–64k prompt-size band (mWh, scraped from the Neuralwatt portal). Dashed line: linear regression on base models only (excluding -fast / -short variants). Hover any dot for details.

§ Open-weight models — 44 models by Agentic Index

#1
59.10
agentic
coding74.8
intel.59.5
blended $/1M$0.90
tokens/s77
#2
58.20
agentic
coding71.5
intel.57.5
blended $/1M$0.10
tokens/s42
#4

Qwen3.8-Flash-Next

Alibaba · weights
released 1w ago · 2026-08-26 slug: qwen3-8-flash-next 262k RTteximavid
56.40
agentic
coding73.1
intel.55.8
blended $/1M$0.09
tokens/s89
#5
54.30
agentic
coding76.2
intel.59.7
blended $/1M$2.31
tokens/s38
#7
49.60
agentic
coding68.8
intel.53.2
blended $/1M$0.69
tokens/s54
#8
48.40
agentic
coding69.1
intel.51.8
blended $/1M$0.23
tokens/s108
#9
45.70
agentic
coding68.8
intel.52.6
blended $/1M$0.90
tokens/s68
#10
36.10
agentic
coding58.6
intel.45.4
blended $/1M$0.22
tokens/s131
#14
31.20
agentic
coding61.8
intel.45.1
blended $/1M$0.70
tokens/s33
#15
30.60
agentic
coding55.8
intel.41.0
blended $/1M$0.84
tokens/s56
#16
30.30
agentic
coding60.8
intel.43.0
blended $/1M$0.72
tokens/s45
#17
29.50
agentic
coding60.2
intel.42.9
blended $/1M$0.47
tokens/s37
#19

Nemotron 3 Ultra 550B A55B (Reasoning)

Nvidia
released 3mo ago · 2026-06-04 slug: nvidia-nemotron-3-ultra-550b-a55b 1000k RTtex
27.50
agentic
coding49.3
intel.38.3
blended $/1M$0.58
tokens/s69
#20
26.20
agentic
coding45.3
intel.34.5
blended $/1M$0.76
tokens/s95
#21
25.90
agentic
coding52.6
intel.38.9
blended $/1M$0.22
tokens/s64
#22

Muse Glimmer (high)

Meta · weights
released 3w ago · 2026-08-10 slug: muse-glimmer 131k RTtexima
22.90
agentic
coding49.0
intel.35.1
blended $/1M$0.23
tokens/s103
#24
21.70
agentic
coding46.8
intel.36.0
blended $/1M$0.49
tokens/s65
#27
19.80
agentic
coding48.2
intel.34.3
blended $/1M$0.90
tokens/s87
#29

Gemma 4 31B (Reasoning)

Google · weights
released 5mo ago · 2026-04-02 slug: gemma-4-31b 262k RTtexima
14.40
agentic
coding43.4
intel.29.7
blended $/1M$0.00
tokens/s35
#30

Nemotron 3.5 Lightning

Nvidia
released 3w ago · 2026-08-11 slug: nemotron-3-5-lightning 262k RTtex
13.80
agentic
coding26.8
intel.23.6
blended $/1M$0.07
tokens/s308
#31
13.40
agentic
coding30.4
intel.24.1
blended $/1M$0.20
tokens/s156
#32

Gemma 4 26B A4B (Reasoning)

Google · weights
released 5mo ago · 2026-04-02 slug: gemma-4-26b-a4b 262k RTtexima
1 providers:EmpirioLabs AI
11.00
agentic
coding39.3
intel.26.1
blended $/1M$0.12
tokens/s
#33

Command A+

Cohere
released 3mo ago · 2026-05-20 slug: command-a-plus 128k RTtexima
9.20
agentic
coding27.8
intel.22.8
blended $/1M$0.00
tokens/s241
#35

Nemotron 3 Super 120B A12B (Reasoning)

Nvidia
released 5mo ago · 2026-03-11 slug: nvidia-nemotron-3-super-120b-a12b 262k RTtex
1 providers:Requesty
8.80
agentic
coding37.7
intel.25.7
blended $/1M$0.26
tokens/s140
#37

Mistral Small 3.1

Mistral · weights
released 1y ago · 2025-03-17 slug: mistral-small-3-1 128k Ttexima
5.30
agentic
coding26.3
intel.14.9
blended $/1M$0.12
tokens/s139
#39
3.10
agentic
coding20.7
intel.15.2
blended $/1M$0.07
tokens/s108
#40

North Mini Code

Cohere
released 2mo ago · 2026-06-09 slug: north-mini-code 256k RTtex
1 providers:NanoGPT
3.10
agentic
coding36.5
intel.20.2
blended $/1M$0.00
tokens/s115

§ Providers (155)

Providers offering at least one of the 44 open-weight models above, with count of those models. Click for docs.

HQ location 🇺🇸 US 🇨🇳 China 🌍 Other ❓ Unknown KV cache 💰 Priced ❌ Not priced Provider type Model maker Proxy Other
NanoGPT34Kilo Gateway33OpenRouter33Vercel AI Gateway31Hugging Face30Deep Infra🇺🇸 US29Cortecs24Merge Gateway24DevPass (LLM Gateway)21LLM Gateway21EmpirioLabs AI20OrcaRouter19Eden AI18Ofox18Requesty18NovitaAI17Pioneer17Together AI17Kenari16Charm Hyper15CrossModel15Abacus14Baseten🇺🇸 US14CrofAI14OpenCode Go14SiliconFlow🇸🇬 SG🇺🇸 US14Nvidia🇺🇸 US13TensorX🇮🇪 IE🇫🇮 FI13ZenMux13DigitalOcean12Ollama Cloud12Venice AI🇺🇸 US12Weights & Biases12Jalapeno Cloud11OpenCode Zen11SCNet Token Plan11Alibaba (China)10AIHubMix9ClinePass9Impossibl9Nebius Token Factory🇳🇱 NL9TokenGo9ai&9Berget.AI🇸🇪 SE8GreenPT8Neuralwatt🇺🇸 US8SiliconFlow (China)8Alibaba Token Plan7Alibaba Token Plan (China)7Qiniu7Synthetic7Auriko6FastRouter6FrogBot6HPC-AI6Neon6Vancine6Vivgrid6Vultr6routing.run6Alibaba🇸🇬 SG🇨🇳 CN5Ambient5Arcee5Azure🇺🇸 US5Crusoe🇺🇸 US5GMI Cloud🇺🇸 US5Helicone5OVHcloud AI Endpoints5QVAC5Regolo AI5Scaleway🇫🇷 FR5UnoRouter5Volcengine Ark Coding Plan5Z.AI5Zhipu AI5Zhipu AI Coding Plan5above.dev5evroc5DInference4EBCloud4Groq🇺🇸 US4Jiekou.AI4LLMTR4Moonshot AI🇸🇬 SG4Moonshot AI (China)4Poe4Tinfoil4Wafer4Z.AI Coding Plan4302.AI3Alibaba Coding Plan3Alibaba Coding Plan (China)3Cloudflare AI Gateway3Friendli🇺🇸 US3Inceptron🇸🇪 SE🇫🇮 FI3Lilac3Meganova3Mixlayer3Modal🇺🇸 US3SCX.ai3STACKIT3AKI.IO2AMD2AnyAPI2Azure Cognitive Services2Cerebras🇺🇸 US2CoralBricks2DeepSeek🇨🇳 CN2GitHub Copilot2IO.NET🇺🇸 US2MiniMax (minimax.io)🇸🇬 SG🇺🇸 US2MiniMax (minimaxi.com)2MiniMax Token Plan (minimax.io)2MiniMax Token Plan (minimaxi.com)2Model Oracle AI2Modelis2NEAR AI Cloud2OpenReason2Perplexity Agent2Privatemode AI2RunInfra2SenseNova (China)2submodel2ArgyllDev🇬🇧 GB1Blue Claw1CloudFerro Sherlock1D.Run (China)1DaoXE1Hetzner1InferX1Infomaniak1IteraCompute1KUAE Cloud Coding Plan1Kosmik Compute1LMStudio1Moark1Opper1Pendra1SaladCloud AI Gateway1StepFun (China)1StepFun (Global)1StepFun Step Plan (China)1StepFun Step Plan (Global)1Subconscious1Tencent Coding Plan (China)1Tencent Token Plan1Tencent TokenHub1Thinking Machines1Xiaomi🇨🇳 CN🇸🇬 SG🇳🇱 NL1Xiaomi Token Plan (China)1Xiaomi Token Plan (Europe)1Xiaomi Token Plan (Singapore)1Zenifra1iFlow1watsonx.ai1