Models & routing
Every model id, the protocol it needs, and what you actually pay. The list column is the provider’s published rate, so the discount is something you can check rather than take on trust.
Protocol matrix#
The router does not translate protocols silently. Pick the endpoint that matches the model family — the reasoning for why is in the API reference.
| Model family | Endpoint | Base URL suffix | Fast mode |
|---|---|---|---|
| Claude (Anthropic) | /v1/messages | none | no |
| GPT (OpenAI) | /v1/chat/completions | /v1 | yes |
| GPT Responses (OpenAI) | /v1/responses | /v1 | yes |
| Gemini (Google) | /v1beta/models/{model}:generateContent | none | no |
| TypeSafe (jev) | /v1/messages | none | no |
| GPT Image 2 | /v1/images/generations | /v1 | no |
Catalogue and rates#
Prices are per million tokens in USD. List is the provider’s published rate; VipAI is what you are charged.
| Model | Id | Context | Cache read | List in | VipAI in | VipAI out | Discount |
|---|---|---|---|---|---|---|---|
| Claude Fable 5 | claude-fable-5 | 1M | $0.30 | $10.00 | $3.00 | $14.60 | 70% |
| Claude Opus 4.8 | claude-opus-4-8 | 1M | $0.16 | $5.00 | $1.60 | $8.00 | 68% |
| Claude Sonnet 5 | claude-sonnet-5 | 1M | $0.06 | $3.00 | $0.60 | $3.10 | 80% |
| Gemini 3.1 Pro (Preview) | gemini-3.1-pro-preview | 2M | $0.035 | $2.00 | $0.35 | $2.10 | 83% |
| GPT-5.5 | gpt-5.5 | 1M | $0.07 | $5.00 | $0.70 | $4.00 | 86% |
| GPT-5.6 Sol | gpt-5.6-sol | 1M | $0.065 | $5.00 | $0.65 | $3.90 | 87% |
| Claude Fable 5.1 | claude-fable-5-1 | 1M | — | $10.00 | $2.40 | $12.00 | 76% |
| Claude Haiku 4.5 | claude-haiku-4-5-20251001 | 200K | $0.10 | $1.00 | $1.00 | $5.00 | — |
| Claude Haiku 4.5 | claude-haiku-4-5 | 200K | — | $1.00 | $0.34 | $1.70 | 66% |
| Claude Opus 4.6 | claude-opus-4-6 | 1M | $0.50 | $5.00 | $5.00 | $25.00 | — |
| Claude Opus 4.7 | claude-opus-4-7 | 1M | $0.50 | $5.00 | $5.00 | $25.00 | — |
| Claude Opus 5 | claude-opus-5 | 1M | — | $5.00 | $1.05 | $5.25 | 79% |
| Claude Opus 5.5 | claude-opus-5-5 | 1M | — | $4.00 | $1.04 | $5.20 | 74% |
| Claude Sonnet 4.6 | claude-sonnet-4-6 | 1M | — | $3.00 | $0.60 | $3.00 | 80% |
| DS DeepSeek Flash | deepseek-flash | 1M | $0.003 | $0.15 | $0.15 | $0.60 | — |
| DS DeepSeek Flash (free) | deepseek-flash-free | 1M | $0 | $0.15 | $0.0015 | $0.006 | 99% |
| DS DeepSeek V4 Flash | deepseek-v4-flash | 1M | $0.0028 | $0.14 | $0.14 | $0.28 | — |
| DS DeepSeek V4 Flash (free) | deepseek-v4-flash-free | 1M | $0 | $0.14 | $0.0014 | $0.0028 | 99% |
| DS DeepSeek V4 Flash Vision | deepseek-v4-flash-vision-exp | 1M | $0.0044 | $0.22 | $0.22 | $0.66 | — |
| DS DeepSeek V4 Pro | deepseek-v4-pro | 1M | $0.0046 | $0.43 | $0.55 | $1.10 | — |
| DS DeepSeek V4 Pro (0425) | deepseek-v4-pro-260425 | 1M | $0.0036 | $0.43 | $0.43 | $0.87 | — |
| GLM GLM-5.2 | glm-5.2 | 1M | $0.32 | $1.40 | $1.70 | $5.30 | — |
| GLM-5.3 | glm-5.3 | 1M | $0.26 | $1.40 | $1.40 | $4.40 | — |
| GLM-5.3 Flash | glm-5.3-flash | 1M | $0.13 | $0.70 | $0.70 | $2.20 | — |
| GLM-5.3 Flash (free) | glm-5.3-flash-free | 1M | $0.0013 | $0.70 | $0.007 | $0.022 | 99% |
| Gemini 3.1 Flash Image | gemini-3.1-flash-image | 1M | — | $1.50 | $0.30 | $1.80 | 80% |
| Gemini 3.5 Flash | gemini-3.5-flash | 1M | — | $1.50 | $0.30 | $1.80 | 80% |
| Gemini 3.5 Flash-Lite | gemini-3.5-flash-lite | 1M | — | $0.30 | $0.069 | $0.57 | 77% |
| Gemini 3.6 Flash | gemini-3.6-flash | 1M | — | $1.50 | $0.32 | $1.57 | 79% |
| Gemini 3.7 Flash | gemini-3.7-flash | 1M | — | $1.50 | $0.28 | $1.42 | 81% |
| Gemini 3.8 Flash | gemini-3.8-flash | 1M | — | $1.50 | $0.27 | $1.35 | 82% |
| Kimi K3 | kimi-k3 | 1M | — | $3.00 | $0.75 | $3.75 | 75% |
| Kimi K3 (1M) | kimi-k3[1M] | 1M | — | $3.00 | $0.75 | $3.75 | 75% |
| GPT-5.4 | gpt-5.4 | 1M | $0.035 | $2.50 | $0.35 | $2.10 | 86% |
| GPT-5.4 mini | gpt-5.4-mini | 1M | $0.011 | $0.75 | $0.11 | $0.66 | 85% |
| GPT-5.6 Luna | gpt-5.6-luna | 1M | $0.014 | $1.00 | $0.14 | $0.80 | 86% |
| GPT-5.6 Terra | gpt-5.6-terra | 1M | $0.035 | $2.50 | $0.35 | $2.00 | 86% |
| GPT-6 Luna | gpt-6-luna | 1M | — | $0.10 | $0.011 | $0.055 | 89% |
| GPT-6 Sol | gpt-6-sol | 1M | — | $2.00 | $0.22 | $1.10 | 89% |
| Grok 4.6 | grok-4.6 | 500K | $0.10 | $2.00 | $0.40 | $1.20 | 80% |
Complimentary models#
Model ids ending in -free, plus jev, run at a 99% promotional discount as a limited daily offer. They are subject to both a personal allowance and a shared platform capacity — rate limits has the numbers.
| Complimentary id | Lane |
|---|---|
deepseek-flash-free | DeepSeek lane |
deepseek-v4-flash-free | DeepSeek lane |
glm-5.3-flash-free | GLM lane |
jev | TypeSafe lane |
Listing everything from code#
Rather than hard-coding the list above, ask for it — the catalogue rotates as providers add and retire models.
from openai import OpenAI
client = OpenAI(
base_url="https://api.vipai.site/v1",
api_key="YOUR_API_KEY",
)
for m in client.models.list():
print(m.id)If a lane goes offline#
The router fails loudly rather than substituting a cheaper model. If a request returns a lane error, the model you selected is genuinely unavailable at that moment — retry, or pin a different id. Nothing is silently swapped behind your back, which is the behaviour to rely on when you are validating output quality rather than chasing a green checkmark.
Create a key, send one request, and watch it land in the usage view before you point an agent tool at it.