MODEL LEDGER
The Wang Report · AI Desk · Hong Kong

The table a procurement or platform team needs and nobody keeps: 36 frontier and leading open models with what each costs per million tokens, how much context it holds, what it accepts, whether its weights are open, the Artificial Analysis capability indices where they are published, and the routes a firm can run it on. Prices and context from OpenRouter's public catalogue of 229 models, refreshed daily (last Sep 9). Hosting routes are verified at the vendor level only; region availability in Hong Kong and Singapore and contractual terms are not asserted here.

Changed This Week
Sep 9
Price change: Moonshot (Kimi) Kimi K2.7 Code
from $0.66 / $3.4 to $0.71 / $3.5 per million tokens
Sep 9
Price change: DeepSeek DeepSeek V4 Flash 0731
from $0.14 / $0.28 to $0.07 / $0.18 per million tokens
Sep 9
Price change: DeepSeek DeepSeek V4 Pro 0813
from $0.66 / $1.98 to $0.58 / $1.74 per million tokens
Sep 9
New: Anthropic Claude Opus 4.8
released May 27
Sep 8
Price change: Tencent Hunyuan Hy3
from $0.13 / $0.53 to $0.08 / $0.33 per million tokens
Sep 8
Price change: DeepSeek DeepSeek V4 Pro 0813
from $1.05 / $3.15 to $0.66 / $1.98 per million tokens
Sep 8
Price change: DeepSeek DeepSeek V4 Flash Vision Exp
from $0.44 / $1.32 to $0.22 / $0.66 per million tokens
Sep 7
New: Tencent Hunyuan Hy3 preview
released Apr 22
Sep 7
New: Tencent Hunyuan Hy3
released Jul 6
OpenAI
Sep 4
GPT-6 Astra
released Sep 4 · 1.1M ctx · 128k max output · file, image, text in · tool use
$10 in / $50 out per million tokens · $1 cached input
hosted only · Artificial Analysis: intelligence 52.8, coding 76.9, agentic 51.5
Runs on: Direct API, Microsoft Azure · hosted only (open-weight gpt-oss models excepted)
Sep 4
GPT-6 Astra Pro
released Sep 4 · 1.1M ctx · 128k max output · file, image, text in · tool use
$10 in / $50 out per million tokens · $1 cached input
hosted only · no published index
Runs on: Direct API, Microsoft Azure · hosted only (open-weight gpt-oss models excepted)
Jul 9
GPT-5.6 Luna Pro
released Jul 9 · 1.1M ctx · 128k max output · file, image, text in · tool use
$0.2 in / $1.2 out per million tokens · $0.02 cached input
hosted only · no published index
Runs on: Direct API, Microsoft Azure · hosted only (open-weight gpt-oss models excepted)
Jul 9
GPT-5.6 Sol
released Jul 9 · 1.1M ctx · 128k max output · file, image, text in · tool use
$2 in / $10 out per million tokens · $0.2 cached input
hosted only · Artificial Analysis: intelligence 47.1, coding 77.4, agentic 50.5
Runs on: Direct API, Microsoft Azure · hosted only (open-weight gpt-oss models excepted)
Jul 9
GPT-5.6 Terra
released Jul 9 · 1.1M ctx · 128k max output · file, image, text in · tool use
$2 in / $12 out per million tokens · $0.2 cached input
hosted only · Artificial Analysis: intelligence 42.3, coding 76.7, agentic 43.7
Runs on: Direct API, Microsoft Azure · hosted only (open-weight gpt-oss models excepted)
Alibaba Qwen
Sep 3
Qwen3.8 Max (0902)
released Sep 3 · 1M ctx · 131k max output · text, image, video in · tool use
$2 in / $6 out per million tokens · $0.25 cached input
hosted only · Artificial Analysis: intelligence 40.3, coding 71.8, agentic 49.6
Runs on: Alibaba Cloud Model Studio, self-host (open weights, most sizes) · mostly open weights
Aug 26
Qwen3.8 Flash
released Aug 26 · 1M ctx · 131k max output · text, image, video in · tool use
$0.15 in / $0.47 out per million tokens · $0.02 cached input
open weights (Qwen/Qwen3.8-Flash-Next) · no published index
Runs on: Alibaba Cloud Model Studio, self-host (open weights, most sizes) · mostly open weights
Aug 14
Qwen3.8 27B
released Aug 14 · 1M ctx · 131k max output · text, image, video in · tool use
$0.42 in / $3 out per million tokens · $0.09 cached input
open weights (Qwen/Qwen3.8-27B) · Artificial Analysis: intelligence 33.9, coding 68.1, agentic 46.5
Runs on: Alibaba Cloud Model Studio, self-host (open weights, most sizes) · mostly open weights
Google
Sep 2
Gemini 3.8 Flash
released Sep 2 · 1M ctx · 66k max output · text, image, video, file, audio in · tool use
$0.75 in / $3.75 out per million tokens · $0.08 cached input
hosted only · Artificial Analysis: intelligence 41.2, coding 76.3, agentic 41.1
Runs on: Gemini API, Google Vertex AI · hosted only (Gemma open weights excepted)
Aug 13
Gemini 3.7 Flash
released Aug 13 · 1M ctx · 66k max output · text, image, video, file, audio in · tool use
$0.75 in / $3.75 out per million tokens · $0.08 cached input
hosted only · Artificial Analysis: intelligence 39.4, coding 76.1, agentic 36.4
Runs on: Gemini API, Google Vertex AI · hosted only (Gemma open weights excepted)
Jul 21
Gemini 3.6 Flash
released Jul 21 · 1M ctx · 66k max output · text, image, video, file, audio in · tool use
$0.75 in / $3.75 out per million tokens · $0.08 cached input
hosted only · Artificial Analysis: intelligence 34.3, coding 69.2, agentic 30.2
Runs on: Gemini API, Google Vertex AI · hosted only (Gemma open weights excepted)
Anthropic
Sep 1
Claude Fable 5.1
released Sep 1 · 1M ctx · 128k max output · text, image, file in · tool use
$10 in / $50 out per million tokens · $0.25 cached input
hosted only · Artificial Analysis: intelligence 53.4, coding 81.6, agentic 58
Runs on: Direct API, Amazon Bedrock, Google Vertex AI · hosted only
Jul 24
Claude Opus 5
released Jul 24 · 1M ctx · 128k max output · text, image, file in · tool use
$5 in / $25 out per million tokens · $0.5 cached input
hosted only · Artificial Analysis: intelligence 50.7, coding 78, agentic 56.2
Runs on: Direct API, Amazon Bedrock, Google Vertex AI · hosted only
Jun 30
Claude Sonnet 5
released Jun 30 · 1M ctx · 128k max output · text, image, file in · tool use
$2 in / $10 out per million tokens · $0.2 cached input
hosted only · Artificial Analysis: intelligence 38.4, coding 71.5, agentic 44.3
Runs on: Direct API, Amazon Bedrock, Google Vertex AI · hosted only
Jun 9
Claude Fable 5
released Jun 9 · 1M ctx · 128k max output · text, image, file in · tool use
$10 in / $50 out per million tokens · $1 cached input
hosted only · Artificial Analysis: intelligence 49.7, coding 76.5, agentic 51
Runs on: Direct API, Amazon Bedrock, Google Vertex AI · hosted only
May 27
Claude Opus 4.8
released May 27 · 1M ctx · 128k max output · text, image, file in · tool use
$5 in / $25 out per million tokens · $0.5 cached input
hosted only · Artificial Analysis: intelligence 42, coding 74.3, agentic 42.6
Runs on: Direct API, Amazon Bedrock, Google Vertex AI · hosted only
Tencent Hunyuan
Aug 28
Hy4 preview
released Aug 28 · 1M ctx · 64k max output · text in · tool use
$0.83 in / $2.5 out per million tokens · $0.04 cached input
open weights (tencent/Hy4-preview) · no published index
Runs on: Tencent Cloud, self-host (open-weight models) · mixed
Jul 6
Hy3
released Jul 6 · 262k ctx · 128k max output · text in · tool use
$0.08 in / $0.33 out per million tokens · $0.02 cached input
open weights (tencent/Hy3) · no published index
Runs on: Tencent Cloud, self-host (open-weight models) · mixed
Apr 22
Hy3 preview
released Apr 22 · 262k ctx · 236k max output · text in · tool use
$0.18 in / $0.6 out per million tokens · $0.06 cached input
open weights (tencent/Hy3-preview) · Artificial Analysis: intelligence 25.8, coding 58.8, agentic 25.6
Runs on: Tencent Cloud, self-host (open-weight models) · mixed
Zhipu (GLM)
Aug 26
GLM 5.3 Flash
released Aug 26 · 1.3M ctx · 131k max output · text, image, video in · tool use
$0.08 in / $0.25 out per million tokens · $0.02 cached input
open weights (zai-org/GLM-5.3-Flash) · Artificial Analysis: intelligence 41.9, coding 71.5, agentic 51.2
Runs on: Direct API, self-host (open weights) · open weights
Aug 18
GLM 5.3
released Aug 18 · 1.3M ctx · 944k max output · text in · tool use
$1.4 in / $4.4 out per million tokens · $0.26 cached input
open weights (zai-org/GLM-5.3) · Artificial Analysis: intelligence 44.9, coding 74.8, agentic 53.4
Runs on: Direct API, self-host (open weights) · open weights
Jun 16
GLM 5.2
released Jun 16 · 1M ctx · 131k max output · text in · tool use
$0.97 in / $3.04 out per million tokens · $0.19 cached input
open weights (zai-org/GLM-5.2) · Artificial Analysis: intelligence n/a, coding 68.8, agentic 39.4
Runs on: Direct API, self-host (open weights) · open weights
DeepSeek
Aug 21
DeepSeek V4 Flash Vision Exp
released Aug 21 · 1M ctx · 384k max output · text, image in · tool use
$0.22 in / $0.66 out per million tokens · $0.01 cached input
open weights (deepseek-ai/DeepSeek-V4-Flash-Vision-Exp) · no published index
Runs on: Direct API, self-host (open weights), cloud marketplaces · open weights
Aug 12
DeepSeek V4 Pro 0813
released Aug 12 · 1M ctx · 393k max output · text in · tool use
$0.58 in / $1.74 out per million tokens · $0.06 cached input
open weights (deepseek-ai/DeepSeek-V4-Pro-0813) · Artificial Analysis: intelligence 36.3, coding 68.8, agentic 42.3
Runs on: Direct API, self-host (open weights), cloud marketplaces · open weights
Jul 31
DeepSeek V4 Flash 0731
released Jul 31 · 1.3M ctx · 944k max output · text in · tool use
$0.07 in / $0.18 out per million tokens · $0.02 cached input
open weights (deepseek-ai/DeepSeek-V4-Flash-0731) · Artificial Analysis: intelligence 34.5, coding 69.1, agentic 41.7
Runs on: Direct API, self-host (open weights), cloud marketplaces · open weights
ByteDance Seed
Aug 12
Seed 2.1 Turbo
released Aug 12 · 262k ctx · 236k max output · text, image, video in · tool use
$0.5 in / $2.5 out per million tokens
hosted only · no published index
Runs on: Volcano Engine · mixed
Aug 12
Seed-2.0-Code
released Aug 12 · 262k ctx · 131k max output · text, image, video in · tool use
$0.5 in / $3 out per million tokens
hosted only · no published index
Runs on: Volcano Engine · mixed
xAI
Aug 12
Grok 4.6
released Aug 12 · 500k ctx · 450k max output · text, image, file in · tool use
$2 in / $6 out per million tokens · $0.5 cached input
hosted only · Artificial Analysis: intelligence 44.4, coding 76.8, agentic 53.4
Runs on: Direct API · hosted only
Jul 8
Grok 4.5
released Jul 8 · 500k ctx · 450k max output · text, image, file in · tool use
$2 in / $6 out per million tokens · $0.3 cached input
hosted only · Artificial Analysis: intelligence 39.1, coding 72.4, agentic 42.1
Runs on: Direct API · hosted only
May 20
Grok Build 0.1
released May 20 · 256k ctx · 230k max output · text, image, file in · tool use
$1 in / $2 out per million tokens · $0.2 cached input
hosted only · Artificial Analysis: intelligence n/a, coding 51.5, agentic n/a
Runs on: Direct API · hosted only
Moonshot (Kimi)
Jul 16
Kimi K3
released Jul 16 · 1M ctx · 944k max output · text, image, video in · tool use
$3 in / $15 out per million tokens · $0.3 cached input
open weights (moonshotai/Kimi-K3) · Artificial Analysis: intelligence 43.8, coding 76.2, agentic 50.6
Runs on: Direct API, self-host (open weights) · open weights
Jun 12
Kimi K2.7 Code
released Jun 12 · 262k ctx · 236k max output · text, image in · tool use
$0.71 in / $3.5 out per million tokens · $0.15 cached input
open weights (moonshotai/Kimi-K2.7-Code) · Artificial Analysis: intelligence 26.3, coding 60.8, agentic 22.5
Runs on: Direct API, self-host (open weights) · open weights
Apr 20
Kimi K2.6
released Apr 20 · 262k ctx · 236k max output · text, image in · tool use
$0.95 in / $4 out per million tokens · $0.16 cached input
open weights (moonshotai/Kimi-K2.6) · Artificial Analysis: intelligence n/a, coding 61.8, agentic 22.1
Runs on: Direct API, self-host (open weights) · open weights
MiniMax
May 31
MiniMax M3
released May 31 · 1M ctx · 512k max output · text, image, video in · tool use
$0.3 in / $1.2 out per million tokens · $0.06 cached input
open weights (MiniMaxAI/Minimax-M3) · Artificial Analysis: intelligence 29.6, coding 58.6, agentic 30.8
Runs on: Direct API, self-host (open weights) · open weights
Mar 18
MiniMax M2.7
released Mar 18 · 205k ctx · 131k max output · text in · tool use
$0.3 in / $1.2 out per million tokens · $0.06 cached input
open weights (MiniMaxAI/MiniMax-M2.7) · Artificial Analysis: intelligence 23.2, coding 52.6, agentic 16.8
Runs on: Direct API, self-host (open weights) · open weights
Mistral
Apr 30
Mistral Medium 3.5
released Apr 30 · 262k ctx · 210k max output · text, image, file in · tool use
$1.5 in / $7.5 out per million tokens
hosted only · Artificial Analysis: intelligence 14.9, coding 46.9, agentic 9.4
Runs on: Direct API, Amazon Bedrock, Microsoft Azure, self-host (open-weight models) · mixed

The rows on this page are drawn from the sources named in the footer and may be reused with attribution to The Wang Report and to that source; the curation, the series and the annotations are The Wang Report's own. Data on this site. Named parties have a right of reply.

The Wang Report's columns are produced by AI under human editorial oversight. See our Editorial Standards.