Free test tokens — try any model

API integration

9 min readAPI v1Report a problem

VipAI accepts four wire formats. Point the official SDK at the base URL and your application code does not change — the requests, the streaming and the SDK behaviour are the provider’s own.

Authentication#

Each protocol expects a different header. Sending the wrong one is the single most common cause of a 401, and it is worth getting right before anything else.

ProtocolHeaderBase URL suffix
Anthropicx-api-key: YOUR_API_KEY plus anthropic-version: 2023-06-01none
OpenAIAuthorization: Bearer YOUR_API_KEY, suffix /v1as given
GeminiAuthorization: Bearer YOUR_API_KEY, or x-goog-api-keynone

Available models#

Fetch the live list rather than hard-coding it — the catalogue rotates as providers add and retire models. Models & routing lists the current set with prices.

GEThttps://api.vipai.site/v1/modelsBearer
request.sh
curl https://api.vipai.site/v1/models \
  -H "Authorization: Bearer $VIPAI_API_KEY"

Anthropic native format#

This is the correct protocol for Claude, and the only one that keeps prompt caching and extended thinking intact. Agent tools must be configured with it.

POSThttps://api.vipai.site/v1/messagesx-api-key

Basic request#

request.sh
curl https://api.vipai.site/v1/messages \
  -H "x-api-key: YOUR_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-fable-5-1",
    "max_tokens": 1024,
    "messages": [{ "role": "user", "content": "Introduce yourself in one sentence" }]
  }'

Streaming with SSE#

Add "stream": true. The response becomes text/event-stream and the event order is fixed: message_start → content_block_start → repeated content_block_delta → content_block_stop → message_delta → message_stop.

request.sh
curl https://api.vipai.site/v1/messages \
  -H "x-api-key: YOUR_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-fable-5-1",
    "max_tokens": 1024,
    "stream": true,
    "messages": [{ "role": "user", "content": "Write a short poem" }]
  }'

OpenAI-compatible format#

POSThttps://api.vipai.site/v1/chat/completionsBearer

Chat Completions#

request.sh
curl https://api.vipai.site/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "max_tokens": 1024,
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Add "stream": true and VipAI returns standard OpenAI SSE chunks in the data: {…} form, terminated by data: [DONE].

Responses API#

If your application already speaks the Responses API, call it directly instead of translating down to Chat Completions.

POSThttps://api.vipai.site/v1/responsesGPT models only
request.sh
curl https://api.vipai.site/v1/responses \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "input": "Introduce VipAI in one sentence"
  }'

Fast mode#

Fast mode — formerly Priority processing — is available on the OpenAI-compatible endpoints. Add the service tier to a Chat Completions or Responses request:

request.json
"service_tier": "fast"

fast is the current value; the legacy priority still works and behaves identically. Fast mode bills at 2× the standard rate and runs GPT-5.6 Sol up to 2.5× faster. On GPT-5.6 and earlier the response object may still report "priority" — that is expected, not a bug. Why Fast mode uses more balance walks through the arithmetic.

Gemini native format#

POSThttps://api.vipai.site/v1beta/models/{model}:generateContentBearer

Basic request#

request.sh
curl https://api.vipai.site/v1beta/models/gemini-3.5-flash:generateContent \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "contents": [
      { "role": "user", "parts": [{ "text": "Introduce yourself in one sentence" }] }
    ]
  }'

Streaming with SSE#

Switch the method to streamGenerateContent and add alt=sse:

request.sh
curl "https://api.vipai.site/v1beta/models/gemini-3.5-flash:streamGenerateContent?alt=sse" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "contents": [
      { "role": "user", "parts": [{ "text": "Write a short poem" }] }
    ]
  }'

Environment variables#

Gemini SDKs disagree about how to spell a custom endpoint — the field may be base_url, baseURL, apiEndpoint, or an environment variable. The rule underneath is the same: point the base URL at VipAI and use your key.

~/.zshrc
export GOOGLE_GEMINI_BASE_URL="https://api.vipai.site"
export GEMINI_API_KEY="YOUR_API_KEY"
export GEMINI_API_KEY_AUTH_MECHANISM="bearer"

Image generation#

The image model gpt-image-2 lives behind its own endpoints and authenticates with Authorization: Bearer. Raise the client timeout to 300 seconds — image calls routinely outlast a default 30-second budget and will fail against it.

POSThttps://api.vipai.site/v1/images/generationsBearer
request.sh
curl https://api.vipai.site/v1/images/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "gpt-image-2",
    "prompt": "An orange cat typing on a keyboard, illustration style"
  }'
POSThttps://api.vipai.site/v1/images/editsmultipart/form-data
request.sh
curl https://api.vipai.site/v1/images/edits \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F model="gpt-image-2" \
  -F image="@photo.png" \
  -F prompt="Replace the background with a starry sky"

Connect agent tools#

Claude Code uses the Anthropic protocol, Codex uses the OpenAI protocol, Gemini CLI uses the Gemini protocol. In every case you set a base URL and a key — the tool itself is untouched. The Claude Code walkthrough goes deeper on pinning the models so the CLI’s choices are explicit and priced.

ToolProtocolVariables
Claude Code, Claude DesktopAnthropicANTHROPIC_BASE_URL, ANTHROPIC_API_KEY
Codex CLI, Codex Desktop, Cline, Cherry StudioOpenAIOPENAI_BASE_URL, OPENAI_API_KEY
Gemini CLIGeminiGOOGLE_GEMINI_BASE_URL, GEMINI_API_KEY, GEMINI_API_KEY_AUTH_MECHANISM
OpenCode, OpenClaw, Trae, WorkBuddyOpenAIOPENAI_BASE_URL, OPENAI_API_KEY
~/.zshrc
export ANTHROPIC_BASE_URL="https://api.vipai.site"
export ANTHROPIC_API_KEY="YOUR_API_KEY"

FAQ#

Getting 401 / authentication failed#

Check the protocol-to-header mapping first: Anthropic uses x-api-key, OpenAI and Gemini use Authorization: Bearer. Then confirm the key is complete (48 characters), has no stray whitespace, and still exists in the dashboard. Finally check the base URL: OpenAI SDKs need the /v1 suffix and Anthropic SDKs must not have it.

Calling Claude through the OpenAI format errors, costs more, or performs worse#

Use the Anthropic native protocol for Claude whenever you can. Claude Code and other agent tools must be configured with it. The OpenAI-compatible shape can lose prompt caching and thinking, which raises cost and lowers answer quality; it is only appropriate for simple chat. The /v1/responses endpoint does not support Claude or Gemini at all and returns 400.

Model unavailable#

Fetch the live list with GET /v1/models first. Check the id spelling, keep it lowercase, and mind - versus .. Also confirm the endpoint matches the model family — a Gemini id on the OpenAI endpoint will not resolve.

Timeout or slow first token#

Large reasoning models such as Opus and Fable can take several seconds to reach the first token. That is not a failure. In production set "stream": true and raise the client read timeout; the server allows responses up to 600 seconds.

Key security#

Store keys in environment variables or a secret manager. Do not hard-code them, commit them, or bundle them into a browser client. If a key leaks, delete it in the dashboard and issue a replacement immediately.

Need a hand getting set up?

The Claude Code guide is the fastest route from zero to a working agent tool, and rate limits covers what the complimentary models allow.