Model prices

Current enabled models and USD prices

Pricing is informational. An individual price may be configured for each individual account.
Input-token caching and routed requestsBecause Model Gate can route requests across available upstream capacity, provider-side input-token or prompt caching may work, but it is not guaranteed for every request. When deterministic cache affinity is important, use a direct provider API account or a direct API key.

All prices are in US dollars per one million tokens. This table is generated from enabled models in the administration panel.

ModelInputOutputCache readCache writeReasoning
claude-opus-5$4$20$0.4$5$20
claude-sonnet-5$1.6$8$0.16$2$1.6
deepseek-v4-flash$0.112$0.224$0.0022$0.112$0.224
deepseek-v4-pro$0.348$0.696$0.0029$0.348$0.696
gpt-5.6-luna$0.16$0.32$0.16$0.2$0.16
gpt-5.6-sol$4$24$0.4$5$24
gpt-5.6-terra$1.6$9.6$0.16$2$9.6
gpt-6-astra$8$40$0.8$9.6$40
kimi-k2.7-code$1.52$6.4$0.304$1.52$6.4
kimi-k3$2.4$12$0.24$2.4$12
mimo-v2.5$0.12$0.24$0.032$0.12$0.24
mimo-v2.5-pro$0.36$0.72$0.032$0.36$0.72
minimax-m3$0.36$1.44$0.072$0.36$1.44
qwen3.7-max$2$6$2$2$6
qwen3.7-plus$0.96$3.84$0.96$0.96$3.84
claude-opus-4-6$4$20$0.4$5$20
claude-opus-4-7$4$20$0.4$5$20
claude-opus-4-8$4$20$0.4$5$20
claude-sonnet-4-6$2.4$12$0.24$3$12
deepseek-v4.1-flash$0.12$0.48$0.024$0.12$0.48
gpt-5.4$2$12$0.2$3$12
gpt-5.5$4$24$0.4$4.8$24
qwen3.8-flash$0.12$0.376$0.0128$0.16$0.376
qwen3.8-max$1.6$4.8$0.2$2$4.8

The normal table shows synchronous and native async=true pricing. If your pricing template has a Batch request price coefficient other than 1, the page displays that coefficient above the table. It applies only to Claude/OpenAI-compatible batch items executed by Model Gate.

Your account catalog

The authenticated Models page shows compact rows with Input, Output and Cache read prices in USD per one million provider tokens. The displayed account rate includes the pricing configuration applicable to your account; the official provider reference is informational and is not the amount debited. An official price is struck through only when the account rate is lower. Missing references are displayed as unknown rather than as a zero price.

Search by name, model ID or provider, narrow by provider/capability, or sort by output price. Details opens the model's context limits, advertised capabilities, modalities and the complete five-category rate table, including Cache write and Reasoning. Missing limits remain “Not published.” Prices in the detail table retain their stored decimal precision. Individual accounts additionally see the configured Tokens display coefficients; Business accounts retain USD pricing.