API reference
Tokun supports three API formats on one metered pipeline: the OpenAI Chat Completions and Responses APIs, and the Anthropic Messages API. Use the one your client already supports. Model IDs, billing, and routing work the same way on all three.
Base URL and authentication
Authenticate every request with a Tokun API key (starts with sk-) from the API keys page. The base URL depends on the API format you call:
| API format | Base URL | Auth header |
|---|---|---|
OpenAI (Chat Completions, Responses) | https://api.tokun.sh/v1 | Authorization: Bearer sk-… |
Anthropic (Messages) | https://api.tokun.sh | x-api-key: sk-… (or Authorization: Bearer) |
/v1 suffix because the Anthropic SDK adds the path itself. Use https://api.tokun.sh for Messages and https://api.tokun.sh/v1 for the OpenAI endpoints.If your client can't set an Authorization header, you can put the key in the URL path instead: POST /v1/{sk-key}/chat/completions.
POST /v1/chat/completions
The OpenAI Chat Completions API. Send a standard OpenAI request body with a Tokun model ID. Set "stream": true to stream tokens over SSE. Responses are byte-compatible with the official OpenAI format.
curl https://api.tokun.sh/v1/chat/completions \
-H "Authorization: Bearer $TOKUN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-opus-4-8",
"messages": [{"role": "user", "content": "Say hello in one word."}]
}'POST /v1/responses
The OpenAI Responses API, with the same authentication and model IDs as Chat Completions. Use it if your SDK targets the Responses API. It is stateless: store is always false, previous_response_id is rejected, and GET /v1/responses/{id} returns 404. Use the Anthropic Messages API if you need signed thinking blocks to round-trip across turns.
POST /v1/messages
The Anthropic Messages API, used by the Anthropic SDKs and Claude Code. Buffered or streaming with the official Anthropic event protocol; signed thinking blocks round-trip in multi-turn tool use. Authenticate with x-api-key: sk-… or Authorization: Bearer.
curl https://api.tokun.sh/v1/messages \
-H "x-api-key: $TOKUN_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-opus-4-8",
"max_tokens": 64,
"messages": [{"role": "user", "content": "Say hello in one word."}]
}'POST /v1/messages/count_tokens
Returns a conservative token estimate for a Messages request. It is free: no charge and no balance hold.
GET /v1/models
Official model pools
Official model pools appear in the model catalog and use a pool ID such as flash-pool. Choose an available pool from the live catalog; pool names and members are configured by Tokun administrators.
When the pool enables member selection, use pool-id/member-name, replacing the model's lab prefix with the pool ID. The selected member is the only model that can execute; the pool's pricing mode still applies. When selection is disabled, this form returns 400.
# Set TOKUN_API_KEY and TOKUN_POOL_ID from your account and the live catalog.
curl "${TOKUN_BASE_URL:-https://api.tokun.sh/v1}/chat/completions" \
-H "Authorization: Bearer $TOKUN_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\":\"$TOKUN_POOL_ID\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}],\"max_tokens\":64}"# The pool must enable member selection. TOKUN_MEMBER is the member name without its lab prefix.
curl "${TOKUN_BASE_URL:-https://api.tokun.sh/v1}/chat/completions" \
-H "Authorization: Bearer $TOKUN_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\":\"$TOKUN_POOL_ID/$TOKUN_MEMBER\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}],\"max_tokens\":64}"Every pool response echoes the model ID you sent: the pool ID for an automatic request, and pool-id/member-name when you selected a member. The canonical ID of the member that served the request travels in the tokun-resolved-model header, and in executed_public_model on the usage record. Your bill follows the configured pool pricing mode.
A pool stays listed when every member is unavailable. Calling it returns 503; requests never leave the pool to find another model.
Pool api_surfaces is the union of its members' supported APIs. Each request uses a member that supports its API format. Context and output limits are the minimum curated values across members; unknown limits are omitted. Pool count_tokens uses the first supported Anthropic member in routing order and returns its tokenizer count without a charge.
Lists the available model IDs and each model's context window. The response format follows the request: a caller that sends x-api-key or anthropic-version (as Anthropic clients do) gets Anthropic's list format; every other caller gets OpenAI's list format. Listing is free (no charge, no hold).
Errors
| Status | Meaning | Fix |
|---|---|---|
401 | Missing or invalid API key | Send a valid Tokun sk- key. |
402 | Your account balance cannot cover the request | Top up on the Billing page. |
403 | The account has funds, but this key is over its own budget cap | Send the request with a key that has no budget cap, or create a key with a higher cap. |
404 | /v1 added to the Anthropic base URL | Use https://api.tokun.sh (no /v1) for Messages. |
400 — unknown model | Model ID not recognized | Use a model or pool ID from the live catalog. An unknown ID does not select another model. |
429 / 5xx | Rate limit or provider error | Passed through from the provider that served the request. Retry with backoff. |
insufficient balance. 403 is the per-key budget cap: the message is credential budget exceeded. Tokun checks the balance first, so an empty account returns 402 even when the key is also over its cap. On Messages the same two statuses arrive in the Anthropic envelope, typed billing_error and permission_error.