Skip to content

Blog

Combined keys: stack your subscriptions into one credential

On an Unlimited subscription, the thing you’re actually shaping isn’t tokens — it’s capacity over time. Each subscription gives you a daily coverage window (the 8-hour blocks you reserve) and, within it, some number of requests you can run in parallel. Your real capacity is those two dimensions multiplied: parallel slots × hours.

Most people grow along both axes. You buy a second block to cover more of the day, or a second subscription to run more requests at once during your busy hours. The awkward part used to be the bookkeeping: every subscription minted its own key, so a serious setup meant three or four keys, each live at different times, each with its own capacity — and you juggling which one to paste where.

Combined keys remove the juggling. A combined key (it looks like sk-ci-meta-…) folds two or more of your subscriptions into a single credential, and it behaves like the union of everything underneath it.

Coverage adds up. A combined key is live whenever any of its subscriptions has an open window. Reserve the Europe block on one subscription and the Americas block on another, combine them, and the one key covers both stretches of the day.

Overlap stacks parallel capacity. Where two subscriptions cover the same hour, their parallel slots add together. That’s the lever for concurrency: if you need to run more requests side by side during your peak hours, buy a second subscription over those hours and combine it in.

A concrete example. Say you hold one full-day (24h) subscription and add a second subscription on just the Europe block. The combined key gives you:

  • Double the parallel capacity during the Europe hours — the two subscriptions overlap there, so their slots stack.
  • Baseline capacity the rest of the day — only the 24h subscription is covering those hours.

Coverage is 24/7 (from the full-day subscription); parallel capacity is shaped to peak exactly when you work.

Combining is about capacity, not billing. Each subscription keeps its own monthly allowance and its own renewal date — nothing is pooled or co-mingled. The practical upside is resilience: if one subscription lapses or you let it cancel, it simply drops out of the union. The combined key keeps working with whatever subscriptions remain — no dead key, no scramble to re-issue credentials.

Requests route by the model you ask for. So a combined key can span subscriptions on different pools — a Core Pool subscription and a Frontier Pool subscription under one key — and each request lands wherever its model lives. Ask for deepseek-v4-flash and it’s served from your Core subscription; ask for kimi-k2.7 and it’s served from your Frontier one. GET /v1/models on a combined key lists every model you can reach across all of them.

Combining a subscription into a key removes that subscription’s standalone key — each subscription has exactly one credential at a time. That’s deliberate: it means there’s never ambiguity about what capacity a given key carries. A key’s coverage and parallel capacity are always the exact sum of the subscriptions it holds, nothing more, nothing hidden. The dashboard walks you through it when you combine.

In the Keys page, choose Create API Key and multi-select the subscriptions you want to combine. Before you commit, a preview shows the resulting daily coverage and peak parallel capacity, so you can see the shape you’re buying into. Prefer the API? The Management API does the same thing programmatically.

Full walkthrough in the combined keys guide. If you’re still deciding which blocks to reserve, the live menu and per-block prices are on the pools page.


CheapestInference serves frontier open-weights models — Kimi K2.7, Kimi K2.6, GLM 5.2, MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash, MiMo v2.5 (Core Pool) — through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

Qwen coding plans in 2026: what you actually get

Qwen3-Coder — Alibaba’s open-weights coder family — is one of the most capable coding models you can run over an API, and one of the most searched-for. If you want to run it as your daily coding model, you have a few realistic routes. As with Kimi and GLM, they differ less in headline price than in cost shape: what happens to your bill and your workflow when a heavy week hits.

Evaluating Qwen coding plans? Here’s how flat-rate unlimited access to comparable open-weights models stacks up. We don’t serve Qwen — but if what you’re really after is a fixed monthly bill for a capable open-weights coder, the same cost-shape question applies to models like Kimi K2.7, GLM 5.2, DeepSeek V4 Flash, MiMo v2.5, and MiniMax M3. Compare the flat-rate pools →

Alibaba Cloud sells a subscription Coding Plan on Model Studio aimed specifically at coding-tool usage of Qwen (plus a few third-party models). It’s first-party access, integrated with Qwen Code and compatible with Claude Code, Cline, and Cursor, and the Qwen coder models — qwen3-coder-plus, qwen3-coder-next, and the newer qwen3.x-plus line — arrive there first.

The trade-off is that the plan is quota-based: each tier grants a request allowance that resets on a schedule — with per-few-hours, weekly, and monthly request caps — and burning through it mid-refactor means waiting for the reset or moving up a tier. The tier lineup itself has already shifted once in 2026 (the entry-level Lite tier stopped accepting new orders), so any number printed here would go stale — check Alibaba’s Coding Plan page for the current tiers and quotas.

Good fit: you want first-party access to the newest Qwen coder models and your volume fits inside a tier’s request quota.

Qwen3-Coder is available per-token from Alibaba’s Model Studio / DashScope API and from several aggregators. No tiers, no resets — you pay for exactly the tokens you burn, which is ideal while you’re evaluating the model or your usage is light. New Model Studio accounts also get a time-limited free-token trial (region-restricted, expiring after a fixed window), useful for a first look — see Alibaba’s pricing page for the current allowance.

The catch is structural, not Qwen-specific: coding agents re-send their whole context on every tool call, so token volume compounds with every iteration. A capable coder will happily churn through long agent sessions — great for output, open-ended for the invoice. Per-token Qwen is cheap per request and unpredictable per month.

Good fit: a few million tokens a month, spiky schedules, or benchmarking before committing.

Route 3: unlimited time blocks on comparable open-weights models

Section titled “Route 3: unlimited time blocks on comparable open-weights models”

The third shape is the one we sell, so apply the usual discount for self-interest — and note the honest caveat up front: we don’t serve Qwen. What we offer is the same cost shape for a lineup of comparable open-weights coders. You reserve one or more daily 8-hour time blocks and get unlimited usage during them: no token allowances, no request quotas, no resets — a monthly number that’s fixed the day you subscribe. Capacity is shaped by per-key concurrency instead of token or request budgets, so an agent that loops all afternoon changes nothing on the bill.

Two things matter if you’re weighing this against a Qwen plan:

  1. Comparable models, one fee. A Frontier Pool block covers Kimi K2.7, Kimi K2.6, GLM 5.2, and MiniMax M3 (1M context); the Core Pool covers DeepSeek V4 Flash and MiMo v2.5 — all open-weights, switchable per request. If you were choosing Qwen for capability-per-dollar, these are in the same class.
  2. It runs in the same tools. The API speaks both the Anthropic and OpenAI formats, so it drops into Claude Code, Cline, Roo Code, or Qwen Code — no wrapper, just a base-URL change.

Current block pricing is on the pools page.

Good fit: you want a fixed monthly bill for a capable open-weights coder and predictable working hours — and you’re not locked to the Qwen name specifically.

RouteCost shapeLimitsModelsBest for
Official Qwen Coding PlanFixed monthly subscriptionRequest quotas that reset (per-few-hours / weekly / monthly)First-party Qwen coder models (plus some third-party)First-party Qwen, volume inside quota
Per-token APIPay per token usedNone — spend scales with usageAny Qwen model on Model Studio / aggregatorsLight, spiky, or exploratory use
Flat-rate time blocks (us)Fixed monthly, per blockConcurrency-shaped; no token or request capsComparable open-weights (Kimi, GLM, DeepSeek, MiMo, MiniMax) — not QwenPredictable hours, fixed bill, model-agnostic

Official Qwen Coding Plan — first-party access, day-one Qwen coder updates, your volume fits the request quota. Per-token — light, spiky, or exploratory usage; pay only for what you burn. Flat-rate time blocks — heavy daily coding in predictable hours where you want a constant bill and you’re open to a comparable open-weights model rather than Qwen specifically.

All three answer the same underlying question. It isn’t “which Qwen tier is cheapest” — it’s which cost shape matches how you work, and whether you need the Qwen name or just a capable open-weights coder at a fixed price.


CheapestInference serves Kimi K2.7, Kimi K2.6, GLM 5.2, and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. We do not serve Qwen. See the pools or get started.

Kimi coding plans in 2026: K2.7, K2.6, and what you actually get

Kimi K2.7 — Moonshot AI’s open-weights flagship — is currently one of the most capable coding models you can run over an API, and its predecessor K2.6 remains a strong, cheaper-to-serve option. If you want Kimi as your daily coding model, you have three realistic routes. As with GLM, they differ less in headline price than in cost shape: what happens to your bill and your workflow when a heavy week hits.

Moonshot sells subscription plans aimed specifically at coding-tool usage of Kimi. It’s first-party access, tightly integrated with their own tooling, and the models arrive there first. The trade-off is that the plans are quota-based: each tier grants a usage allowance that resets on a schedule, and burning through it mid-refactor means waiting for the reset or moving up a tier. Tiers and allowances change often enough that any number printed here would go stale — check Moonshot’s pricing page for the current shape.

Good fit: you want first-party access and your coding volume fits comfortably inside a tier’s allowance.

Kimi K2.7 and K2.6 are available per-token from Moonshot’s open platform and from several aggregators. No tiers, no resets — you pay for exactly the tokens you burn, which is ideal while you’re evaluating the model or your usage is light.

The catch is structural, not Kimi-specific: coding agents re-send their whole context on every tool call, so token volume compounds with every iteration. A model as eager to work as K2.7 will happily churn through long agent sessions — great for output, open-ended for the invoice. Per-token Kimi is cheap per request and unpredictable per month.

Good fit: a few million tokens a month, spiky schedules, or benchmarking before committing.

The third shape is the one we sell, so apply the usual discount for self-interest — but the mechanics are easy to verify. You reserve one or more daily 8-hour time blocks and get unlimited Kimi usage during them: no token allowances, no resets, a monthly number that’s fixed the day you subscribe. Capacity is shaped by per-key concurrency instead of token budgets, so an agent that loops all afternoon changes nothing on the bill.

Three properties matter for coding specifically:

  1. Both Kimis under one fee. A Frontier Pool block covers Kimi K2.7 and K2.6 (model ids kimi-k2.7, kimi-k2.6), switchable per request — use K2.7 for the hard problems and K2.6 where it’s already enough.
  2. It runs inside Claude Code natively. The API speaks both the Anthropic and OpenAI formats, so Kimi drops into Claude Code, Cline, Roo Code, or whatever tool you already use — no wrapper, just a base-URL change.
  3. The subscription isn’t Kimi-only. The same block also covers GLM 5.2 and MiniMax M3 (1M context). If Kimi is your main model but not your only one, that’s four coding plans for the price of one.

Current block pricing is on the pools page.

Good fit: Kimi is your daily driver, your working hours are roughly predictable, and you want the bill to be a constant instead of a variable.

Official Moonshot plan — first-party access, day-one model updates, your volume fits the quota. Per-token — light, spiky, or exploratory usage; pay only for what you burn. Time-block unlimited — heavy daily coding or agent work in predictable hours; fixed cost, both K2 generations plus two more frontier models under one fee.

All three routes serve the same open-weights models. The question isn’t “which Kimi is better” — it’s which cost shape matches how you work.


CheapestInference serves Kimi K2.7, Kimi K2.6, GLM 5.2, and MiniMax M3 (Frontier Pool) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

GLM 5.2 coding plans in 2026: what you actually get

GLM 5.2 — Zhipu AI’s (Z.ai) frontier coding and reasoning model — has become one of the most searched-for open-weights models for coding work. If you’re trying to run it as your daily coding model, you have three realistic routes, and they differ less in price than in cost shape: what happens to your bill and your workflow when usage spikes.

Z.ai sells subscription tiers aimed at coding-tool usage of GLM. You get first-party access and tight integration with their own tooling. The trade-off is that the plans are quota-based: each tier grants an amount of usage that resets on a schedule, and hitting the ceiling mid-task means waiting or upgrading. For usage details and current tiers, check Z.ai’s pricing page — quotas and tiers change often enough that any number printed here would go stale.

Good fit: you want first-party access and your usage fits comfortably inside a tier’s quota.

GLM 5.2 is available per-token from several inference providers and aggregators. No quotas, no tiers — you pay for exactly what you use, which is ideal for light or unpredictable usage.

The catch is the same one that applies to every agent workload: coding agents re-send their whole context on every tool call, so token volume scales with iterations. Per-token GLM is cheap per request and open-ended per month — the bill is a dependent variable of how hard your agent worked.

Good fit: a few million tokens a month, spiky schedules, or evaluation before committing to anything.

The third shape is the one we sell, so discount accordingly — but the mechanics are simple to verify. You reserve one or more daily 8-hour time blocks and get unlimited GLM 5.2 usage during them: no token caps, no quota resets, a fixed monthly number decided at subscription time. Capacity is shaped by per-key concurrency rather than token budgets, so a runaway agent loop changes nothing on the invoice.

Two properties matter for coding specifically:

  1. It works as a drop-in coding plan. The API is Anthropic- and OpenAI-compatible, so GLM 5.2 runs inside Claude Code, Cline, Roo Code, or any tool you already use — model id glm-5.2.
  2. The subscription isn’t GLM-only. A Frontier Pool block covers Kimi K2.7, Kimi K2.6, and MiniMax M3 (1M context) too, switchable per request. If GLM is your main model but not your only one, that’s four coding plans for the price of one.

Current block pricing is on the pools page.

Good fit: GLM is your daily-driver coding model, your hours are roughly predictable, and you want the bill to be a constant.

Official Z.ai plan — first-party access, quota fits your volume, you use their tooling. Per-token — light, spiky, or exploratory usage; pay only for what you burn. Time-block unlimited — heavy daily coding or agent work in predictable hours; fixed cost, no quota anxiety, multiple frontier models under one fee.

All three serve the same open-weights model. The question isn’t “which GLM is better” — it’s which cost shape matches how you work.


CheapestInference serves GLM 5.2, Kimi K2.7, Kimi K2.6, and MiniMax M3 (Frontier Pool, from $48.45/mo billed annually) and DeepSeek V4 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.

Unlimited DeepSeek: what a flat monthly subscription changes

DeepSeek has a well-earned reputation as the budget option among frontier-quality models. Per-token rates for DeepSeek V4 Flash run around $0.14 per million input tokens and $0.28 per million output — an order of magnitude below closed-source flagships.

So why would anyone pay a flat monthly fee for it?

Because per-token pricing has a property that doesn’t care how low the rate is: cost scales with tokens, and agent tokens scale with iterations, not value. Cheap per token is not the same as cheap per month.


The math nobody runs until the invoice arrives

Section titled “The math nobody runs until the invoice arrives”

A coding agent re-sends its growing context on every tool call. A typical task burns 300–500K tokens; an active developer runs dozens of tasks a day. Being conservative:

Tokens/dayTokens/monthPer-token cost (V4 Flash rates)
Light use2M60M~$11/mo
Daily driver15M450M~$80/mo
Heavy agent loops50M1.5B~$270/mo

The rate is tiny. The bill is not — and it’s unpredictable, because next month’s iteration count is unknowable in advance.

A time-block subscription inverts this: you reserve a daily 8-hour window and usage inside it is unlimited. The number on your invoice is decided when you subscribe, not by how many times your agent loops. DeepSeek V4 Flash is served in the Core Pool — current pricing is on the pools page.

At “daily driver” volume, the flat block is cheaper than even DeepSeek’s per-token rates — and the gap only widens from there.

  • DeepSeek V4 Flash with a 1M-token context window — whole codebases, long documents, extended agent runs in a single request.
  • No token caps during your blocks. The plan is unlimited in tokens; capacity is shaped by per-key concurrency instead, so one busy key never affects another.
  • MiMo v2.5 included. A Core Pool subscription covers every model in the pool — Xiaomi’s MiMo v2.5 shares the same 1M-context class.
  • Drop-in API. OpenAI-compatible (/v1/chat/completions) and Anthropic-compatible (/anthropic/v1/messages) — point your SDK, Cline, or Claude Code at it with model id deepseek-v4-flash.
from openai import OpenAI
client = OpenAI(
base_url="https://api.cheapestinference.com/v1",
api_key="sk-...", # subscriber key
)
r = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Review this repo for race conditions: ..."}],
)

When per-token DeepSeek is still the right call

Section titled “When per-token DeepSeek is still the right call”

Honesty clause: if your usage is light or spiky — a few million tokens a month, unpredictable hours — per-token is cheaper and you should use it. The flat block wins when usage is heavy and concentrated in predictable hours: agent development, batch processing, a working day of assisted coding. That’s the break-even logic in one sentence; the full break-even analysis is here.


CheapestInference serves DeepSeek V4 Flash and MiMo v2.5 (Core Pool, from $12.74/mo billed annually) and Kimi K2.7, Kimi K2.6, GLM 5.2, and MiniMax M3 (Frontier Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.