API integration
VipAI accepts four wire formats. Point the official SDK at the base URL and your application code does not change — the requests, the streaming and the SDK behaviour are the provider’s own.
Authentication#
Each protocol expects a different header. Sending the wrong one is the single most common cause of a 401, and it is worth getting right before anything else.
| Protocol | Header | Base URL suffix |
|---|---|---|
| Anthropic | x-api-key: YOUR_API_KEY plus anthropic-version: 2023-06-01 | none |
| OpenAI | Authorization: Bearer YOUR_API_KEY, suffix /v1 | as given |
| Gemini | Authorization: Bearer YOUR_API_KEY, or x-goog-api-key | none |
Available models#
Fetch the live list rather than hard-coding it — the catalogue rotates as providers add and retire models. Models & routing lists the current set with prices.
https://api.vipai.site/v1/modelsBearercurl https://api.vipai.site/v1/models \
-H "Authorization: Bearer $VIPAI_API_KEY"Anthropic native format#
This is the correct protocol for Claude, and the only one that keeps prompt caching and extended thinking intact. Agent tools must be configured with it.
https://api.vipai.site/v1/messagesx-api-keyBasic request#
curl https://api.vipai.site/v1/messages \
-H "x-api-key: YOUR_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-fable-5-1",
"max_tokens": 1024,
"messages": [{ "role": "user", "content": "Introduce yourself in one sentence" }]
}'Streaming with SSE#
Add "stream": true. The response becomes text/event-stream and the event order is fixed: message_start → content_block_start → repeated content_block_delta → content_block_stop → message_delta → message_stop.
curl https://api.vipai.site/v1/messages \
-H "x-api-key: YOUR_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-fable-5-1",
"max_tokens": 1024,
"stream": true,
"messages": [{ "role": "user", "content": "Write a short poem" }]
}'OpenAI-compatible format#
https://api.vipai.site/v1/chat/completionsBearerChat Completions#
curl https://api.vipai.site/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"max_tokens": 1024,
"messages": [{ "role": "user", "content": "Hello" }]
}'Add "stream": true and VipAI returns standard OpenAI SSE chunks in the data: {…} form, terminated by data: [DONE].
Responses API#
If your application already speaks the Responses API, call it directly instead of translating down to Chat Completions.
https://api.vipai.site/v1/responsesGPT models onlycurl https://api.vipai.site/v1/responses \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"input": "Introduce VipAI in one sentence"
}'Fast mode#
Fast mode — formerly Priority processing — is available on the OpenAI-compatible endpoints. Add the service tier to a Chat Completions or Responses request:
"service_tier": "fast"fast is the current value; the legacy priority still works and behaves identically. Fast mode bills at 2× the standard rate and runs GPT-5.6 Sol up to 2.5× faster. On GPT-5.6 and earlier the response object may still report "priority" — that is expected, not a bug. Why Fast mode uses more balance walks through the arithmetic.
Gemini native format#
https://api.vipai.site/v1beta/models/{model}:generateContentBearerBasic request#
curl https://api.vipai.site/v1beta/models/gemini-3.5-flash:generateContent \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{
"contents": [
{ "role": "user", "parts": [{ "text": "Introduce yourself in one sentence" }] }
]
}'Streaming with SSE#
Switch the method to streamGenerateContent and add alt=sse:
curl "https://api.vipai.site/v1beta/models/gemini-3.5-flash:streamGenerateContent?alt=sse" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{
"contents": [
{ "role": "user", "parts": [{ "text": "Write a short poem" }] }
]
}'Environment variables#
Gemini SDKs disagree about how to spell a custom endpoint — the field may be base_url, baseURL, apiEndpoint, or an environment variable. The rule underneath is the same: point the base URL at VipAI and use your key.
export GOOGLE_GEMINI_BASE_URL="https://api.vipai.site"
export GEMINI_API_KEY="YOUR_API_KEY"
export GEMINI_API_KEY_AUTH_MECHANISM="bearer"Image generation#
The image model gpt-image-2 lives behind its own endpoints and authenticates with Authorization: Bearer. Raise the client timeout to 300 seconds — image calls routinely outlast a default 30-second budget and will fail against it.
https://api.vipai.site/v1/images/generationsBearercurl https://api.vipai.site/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{
"model": "gpt-image-2",
"prompt": "An orange cat typing on a keyboard, illustration style"
}'https://api.vipai.site/v1/images/editsmultipart/form-datacurl https://api.vipai.site/v1/images/edits \
-H "Authorization: Bearer YOUR_API_KEY" \
-F model="gpt-image-2" \
-F image="@photo.png" \
-F prompt="Replace the background with a starry sky"Connect agent tools#
Claude Code uses the Anthropic protocol, Codex uses the OpenAI protocol, Gemini CLI uses the Gemini protocol. In every case you set a base URL and a key — the tool itself is untouched. The Claude Code walkthrough goes deeper on pinning the models so the CLI’s choices are explicit and priced.
| Tool | Protocol | Variables |
|---|---|---|
| Claude Code, Claude Desktop | Anthropic | ANTHROPIC_BASE_URL, ANTHROPIC_API_KEY |
| Codex CLI, Codex Desktop, Cline, Cherry Studio | OpenAI | OPENAI_BASE_URL, OPENAI_API_KEY |
| Gemini CLI | Gemini | GOOGLE_GEMINI_BASE_URL, GEMINI_API_KEY, GEMINI_API_KEY_AUTH_MECHANISM |
| OpenCode, OpenClaw, Trae, WorkBuddy | OpenAI | OPENAI_BASE_URL, OPENAI_API_KEY |
export ANTHROPIC_BASE_URL="https://api.vipai.site"
export ANTHROPIC_API_KEY="YOUR_API_KEY"FAQ#
Getting 401 / authentication failed#
Check the protocol-to-header mapping first: Anthropic uses x-api-key, OpenAI and Gemini use Authorization: Bearer. Then confirm the key is complete (48 characters), has no stray whitespace, and still exists in the dashboard. Finally check the base URL: OpenAI SDKs need the /v1 suffix and Anthropic SDKs must not have it.
Calling Claude through the OpenAI format errors, costs more, or performs worse#
Use the Anthropic native protocol for Claude whenever you can. Claude Code and other agent tools must be configured with it. The OpenAI-compatible shape can lose prompt caching and thinking, which raises cost and lowers answer quality; it is only appropriate for simple chat. The /v1/responses endpoint does not support Claude or Gemini at all and returns 400.
Model unavailable#
Fetch the live list with GET /v1/models first. Check the id spelling, keep it lowercase, and mind - versus .. Also confirm the endpoint matches the model family — a Gemini id on the OpenAI endpoint will not resolve.
Timeout or slow first token#
Large reasoning models such as Opus and Fable can take several seconds to reach the first token. That is not a failure. In production set "stream": true and raise the client read timeout; the server allows responses up to 600 seconds.
Key security#
Store keys in environment variables or a secret manager. Do not hard-code them, commit them, or bundle them into a browser client. If a key leaks, delete it in the dashboard and issue a replacement immediately.
The Claude Code guide is the fastest route from zero to a working agent tool, and rate limits covers what the complimentary models allow.