Billing questions
The questions that actually reach support, answered without a sales call. If yours is not here, the API reference covers the request side.
Why do I see several billing rows after sending one message#
A single operation in an agent tool can trigger several real model calls behind the scenes:
- The tool asks the model to understand the task.
- It calls another model to read files and plan edits.
- It calls the model again after tool execution.
- Some clients also retry, or split long work into several requests.
Why did I choose one model and see a different one in billing#
Some clients, agents and routing policies call a smaller model for auxiliary work: classification, planning, summarisation, short file summaries. That is not your main model selection failing — part of the workflow simply used a different model for a supporting call. The catalogue shows the price of each so you can tell a supporting call from a substitution.
What are cache reads and cache writes, and why is the price different#
Many providers prompt-cache aggressively. A cache write is the provider storing reusable context — long system prompts, project files, previous conversation. A cache read is the provider re-serving context it already has.
Writes usually cost more than reads, because the provider has to process and store the context. Reads are usually cheaper. This is the single biggest reason two agents on the same model can produce very different bills: the one with warm cache reads is being served mostly cached context.
Why do some requests have few input tokens but still cost more#
Usually one of four things:
- The request triggered cache writes.
- The model used a higher unit price.
- The client made multiple hidden calls around the visible action.
- Output or reasoning tokens were higher than expected.
Read the model id, input tokens, output tokens, cache read tokens and cache write tokens together before drawing a conclusion from the input count alone.
Why do expected charge and actual deduction differ#
Expected charge is derived from the model usage table. Actual deduction can differ because of rounding, discounts, cached billing rules, account-level settlement rules, or provider-side adjustments. A small gap is expected; if it looks abnormal, keep the request id and contact support.
Why does Fast mode use more balance#
In Fast mode the model list price is 2× the standard-mode price, while your existing account discount is unchanged. So the effective cost roughly doubles per token — but so does the speed, up to 2.5× on GPT-5.6 Sol. Fast mode in the API reference covers how to turn it on and when the response still reports the legacy priority value.
A Fast hit in billing means the request was processed with Fast mode enabled. If you did not expect that, check whether a client library is setting service_tier on your behalf.
Why is my balance consumed faster over time#
Agent tools send larger context as a project grows. Long conversations, more files and repeated tool calls all increase token usage — and each new turn re-establishes context.
To reduce cost:
- Start a new conversation when the old context is no longer needed.
- Avoid asking the agent to read the whole project unless it genuinely has to.
- Prefer smaller models for simple tasks, and pin them — see pinning models in Claude Code.
- Use streaming for perceived latency. It does not reduce token usage.
- Watch the cache write rows. Frequent rewrites raise cost.
- Set a budget ceiling in the dashboard so a runaway loop cannot drain the balance quietly.
Do credits expire#
No. Buy credits when you want and use them whenever. There is no subscription and no expiry — see the FAQ on the home page for the rest of the billing questions that are not request-specific.
When should I contact support#
Reach out if you see:
- Requests you did not make.
- A model id that is clearly not the one you configured.
- An actual deduction far from the usage details.
- Requests that failed but still appear as charged.
- You need help reading cache or token records.
Please include the request time, model id, request id, and a screenshot of the billing record. With the request id the answer is usually immediate.
Bring the request id to Telegram — it is the fastest way to get a specific billing row explained.