Free test tokens — try any model

Models & routing

5 min readAPI v1Report a problem

Every model id, the protocol it needs, and what you actually pay. The list column is the provider’s published rate, so the discount is something you can check rather than take on trust.

Protocol matrix#

The router does not translate protocols silently. Pick the endpoint that matches the model family — the reasoning for why is in the API reference.

Model familyEndpointBase URL suffixFast mode
Claude (Anthropic)/v1/messagesnoneno
GPT (OpenAI)/v1/chat/completions/v1yes
GPT Responses (OpenAI)/v1/responses/v1yes
Gemini (Google)/v1beta/models/{model}:generateContentnoneno
TypeSafe (jev)/v1/messagesnoneno
GPT Image 2/v1/images/generations/v1no

Catalogue and rates#

Prices are per million tokens in USD. List is the provider’s published rate; VipAI is what you are charged.

ModelIdContextCache readList inVipAI inVipAI outDiscount
Claude Fable 5claude-fable-51M$0.30$10.00$3.00$14.6070%
Claude Opus 4.8claude-opus-4-81M$0.16$5.00$1.60$8.0068%
Claude Sonnet 5claude-sonnet-51M$0.06$3.00$0.60$3.1080%
Gemini 3.1 Pro (Preview)gemini-3.1-pro-preview2M$0.035$2.00$0.35$2.1083%
GPT-5.5gpt-5.51M$0.07$5.00$0.70$4.0086%
GPT-5.6 Solgpt-5.6-sol1M$0.065$5.00$0.65$3.9087%
Claude Fable 5.1claude-fable-5-11M—$10.00$2.40$12.0076%
Claude Haiku 4.5claude-haiku-4-5-20251001200K$0.10$1.00$1.00$5.00—
Claude Haiku 4.5claude-haiku-4-5200K—$1.00$0.34$1.7066%
Claude Opus 4.6claude-opus-4-61M$0.50$5.00$5.00$25.00—
Claude Opus 4.7claude-opus-4-71M$0.50$5.00$5.00$25.00—
Claude Opus 5claude-opus-51M—$5.00$1.05$5.2579%
Claude Opus 5.5claude-opus-5-51M—$4.00$1.04$5.2074%
Claude Sonnet 4.6claude-sonnet-4-61M—$3.00$0.60$3.0080%
DS DeepSeek Flashdeepseek-flash1M$0.003$0.15$0.15$0.60—
DS DeepSeek Flash (free)deepseek-flash-free1M$0$0.15$0.0015$0.00699%
DS DeepSeek V4 Flashdeepseek-v4-flash1M$0.0028$0.14$0.14$0.28—
DS DeepSeek V4 Flash (free)deepseek-v4-flash-free1M$0$0.14$0.0014$0.002899%
DS DeepSeek V4 Flash Visiondeepseek-v4-flash-vision-exp1M$0.0044$0.22$0.22$0.66—
DS DeepSeek V4 Prodeepseek-v4-pro1M$0.0046$0.43$0.55$1.10—
DS DeepSeek V4 Pro (0425)deepseek-v4-pro-2604251M$0.0036$0.43$0.43$0.87—
GLM GLM-5.2glm-5.21M$0.32$1.40$1.70$5.30—
GLM-5.3glm-5.31M$0.26$1.40$1.40$4.40—
GLM-5.3 Flashglm-5.3-flash1M$0.13$0.70$0.70$2.20—
GLM-5.3 Flash (free)glm-5.3-flash-free1M$0.0013$0.70$0.007$0.02299%
Gemini 3.1 Flash Imagegemini-3.1-flash-image1M—$1.50$0.30$1.8080%
Gemini 3.5 Flashgemini-3.5-flash1M—$1.50$0.30$1.8080%
Gemini 3.5 Flash-Litegemini-3.5-flash-lite1M—$0.30$0.069$0.5777%
Gemini 3.6 Flashgemini-3.6-flash1M—$1.50$0.32$1.5779%
Gemini 3.7 Flashgemini-3.7-flash1M—$1.50$0.28$1.4281%
Gemini 3.8 Flashgemini-3.8-flash1M—$1.50$0.27$1.3582%
Kimi K3kimi-k31M—$3.00$0.75$3.7575%
Kimi K3 (1M)kimi-k3[1M]1M—$3.00$0.75$3.7575%
GPT-5.4gpt-5.41M$0.035$2.50$0.35$2.1086%
GPT-5.4 minigpt-5.4-mini1M$0.011$0.75$0.11$0.6685%
GPT-5.6 Lunagpt-5.6-luna1M$0.014$1.00$0.14$0.8086%
GPT-5.6 Terragpt-5.6-terra1M$0.035$2.50$0.35$2.0086%
GPT-6 Lunagpt-6-luna1M—$0.10$0.011$0.05589%
GPT-6 Solgpt-6-sol1M—$2.00$0.22$1.1089%
Grok 4.6grok-4.6500K$0.10$2.00$0.40$1.2080%

Complimentary models#

Model ids ending in -free, plus jev, run at a 99% promotional discount as a limited daily offer. They are subject to both a personal allowance and a shared platform capacity — rate limits has the numbers.

Complimentary idLane
deepseek-flash-freeDeepSeek lane
deepseek-v4-flash-freeDeepSeek lane
glm-5.3-flash-freeGLM lane
jevTypeSafe lane

Listing everything from code#

Rather than hard-coding the list above, ask for it — the catalogue rotates as providers add and retire models.

models.py
from openai import OpenAI

client = OpenAI(
    base_url="https://api.vipai.site/v1",
    api_key="YOUR_API_KEY",
)

for m in client.models.list():
    print(m.id)

If a lane goes offline#

The router fails loudly rather than substituting a cheaper model. If a request returns a lane error, the model you selected is genuinely unavailable at that moment — retry, or pin a different id. Nothing is silently swapped behind your back, which is the behaviour to rely on when you are validating output quality rather than chasing a green checkmark.

Ready to route real traffic?

Create a key, send one request, and watch it land in the usage view before you point an agent tool at it.