Comparing benchmark scores, pricing, and energy use across open-weight language models and their providers.
Built with GLM-5.2 on Neuralwatt · 1.60 kWh · 193.7 g CO₂ · Data: Artificial Analysis · models.dev · Neuralwatt · OpenRouter · generated 2026-09-02 07:17 UTC
Blended cost = (7·cache + 2·input + 1·output)/10. When a provider omits cache-hit pricing, the input price is used as an upper bound (cache hits are never more expensive than a regular input token). 44 models shown: open-weight (via models.dev), with an Agentic Index score AND input + output pricing. Models without input or output pricing are excluded. Neuralwatt energy values in tooltips are measured when an AA model matches a NW model; otherwise the tooltip shows an estimated energy derived from the NW cost ↔ energy regression (see the Neuralwatt scatter below).
X: NW blended cost per 1M tokens = (7·cache + 2·input + 1·output)/10 (USD, from Neuralwatt pricing).
Y: energy per request at the 16k–64k prompt-size band (mWh, scraped from the
Neuralwatt portal).
Dashed line: linear regression on base models only (excluding -fast / -short variants).
Hover any dot for details.
Providers offering at least one of the 44 open-weight models above, with count of those models. Click for docs.