Skip to content

Plans & Limits

CheapestInference sells unlimited time-block subscriptions on a model pool. Two pools are active: the Frontier Pool bundles Kimi K2.7 + Kimi K2.6 + GLM 5.2 + MiniMax M3, and the Core Pool bundles DeepSeek V4 Flash + MiMo v2.5. A subscription covers every model in its pool.

How pricing works — reserve daily time blocks

Section titled “How pricing works — reserve daily time blocks”

The day is split into three 8-hour time blocks, aligned to global peak hours and shown in your local timezone. Reserve one, two, or all three blocks — all three gives full 24/7 coverage.

Core Pool — DeepSeek V4 Flash, MiMo v2.5

Section titled “Core Pool — DeepSeek V4 Flash, MiMo v2.5”
BlockWindow (UTC)Price
Asia-Pacific00:00 – 08:00$19.99/mo
Europe08:00 – 16:00$17.99/mo
Americas16:00 – 24:00$14.99/mo

From $12.74/mo with annual billing (~15% off), available per block.

Frontier Pool — Kimi K2.7, Kimi K2.6, GLM 5.2, MiniMax M3

Section titled “Frontier Pool — Kimi K2.7, Kimi K2.6, GLM 5.2, MiniMax M3”
BlockWindow (UTC)Price
Asia-Pacific00:00 – 08:00$59/mo
Europe08:00 – 16:00$61/mo
Americas16:00 – 24:00$57/mo

From $48.45/mo with annual billing (~15% off), available per block.

During your reserved hours, usage is truly unlimited — no tokens to count, no overage charges. Each key handles a limited number of simultaneous requests — create more keys (one per seat) to run requests in parallel. Outside your reserved blocks, the key does not serve traffic.

The price you subscribe at is the price you keep, for as long as your subscription stays active.

  • Renewals never cost more. As long as you don’t cancel, your subscription renews at the price you signed up at — regardless of what new subscribers pay.
  • Price drops reach you automatically. Pricing is dynamic, and when it moves down it moves down for everyone: if the current price of your subscription ever falls below what you’re paying, your renewal is charged at the lower price. You always pay the minimum of your contracted price and the current price — reductions are never reserved for new customers.
  • No silent changes, ever. If circumstances beyond our control ever made a price change unavoidable — something we don’t expect — we would email you in advance, before anything changed. You will never be charged more without being told first.

Every model in your pool is available throughout your reserved blocks. The lineup of a pool can change over time — GET /v1/models is always the authoritative live list.

Your subscription doesn’t count tokens — there are no credits to top up and no per-token charges on your bill. The rates below are not what you pay: they are the models’ standard per-token API prices, shown as a reference so you can size the value of a flat fee against your own usage.

Frontier Pool — from $48.45/mo with annual billing:

ModelModel IDPer-token elsewhere (in)Per-token elsewhere (out)
Kimi K2.7kimi-k2.7$0.950/M$4.000/M
Kimi K2.6kimi-k2.6$0.950/M$4.000/M
GLM 5.2glm-5.2$1.400/M$4.400/M
MiniMax M3minimax-m3$0.300/M$1.200/M

Core Pool — blocks from $14.99/mo, from $12.74/mo with annual billing:

ModelModel IDPer-token elsewhere (in)Per-token elsewhere (out)
DeepSeek V4 Flashdeepseek-v4-flash$0.140/M$0.280/M
MiMo v2.5mimo-v2.5$0.140/M$0.280/M

Per-model details: Kimi K2.7, Kimi K2.6, GLM 5.2, MiniMax M3, DeepSeek V4 Flash, MiMo v2.5.

Additional pools open through community pledges. Pledge for a model at cheapestinference.com/pools; when enough subscribers commit to a specific model, a new pool activates. See the Unlimited Subscriptions API for the full subscribe flow and response shapes.

MethodFor
Card (Stripe)Time-block subscriptions (monthly + annual)
USDC on BaseSubscriptions
x402Agent subscriptions (no human setup needed)

Card subscriptions (Stripe) renew automatically each cycle (monthly or yearly) until you cancel — cancel anytime and access continues to the end of the paid period. USDC and x402 subscriptions are one-time (30 days, no auto-renewal) — you renew manually.

Each key handles a limited number of simultaneous requests during your reserved block. To run more requests in parallel — or to isolate clients — reserve additional seats and use one key each. Keys are independent, so one busy key never starves the others.