Docs
From zero to first routed token in about five minutes
The gateway speaks the OpenAI request shape. If your stack already calls a compatible endpoint, switching is a base URL, a key, and a deploy.
Quickstart
Three steps, one deploy
Sign up, point your client at the gateway, and pass a model id from the catalog. Everything after that is your existing code.
01
Create a key
Sign up, copy your key. It starts with lw-sk-. One key reaches every provider in the catalog.
02
Point at the gateway
Set the base URL to https://api.liquidwhale.com/v1 and swap the key. The request shape is OpenAI-compatible, so nothing else moves.
03
Call any model
Pass the model id from the catalog. If a provider wobbles, the gateway reroutes and reports it in response headers.
Environment
Two exports, any shell. Your SDK picks them up untouched.
export OPENAI_BASE_URL="https://api.liquidwhale.com/v1"
export OPENAI_API_KEY="lw-sk-your-key"First call
The same request in all three flavors, copy button included.
curl https://api.liquidwhale.com/v1/chat/completions \
-H "Authorization: Bearer $LIQUIDWHALE_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "helios-5",
"messages": [
{ "role": "user", "content": "Summarize this changelog in one line." }
]
}'Compatibility
The shape is open. The exit is one env var
Requests and responses keep the standard chat completion fields, including tool calls, streaming deltas, and structured output. Nothing is proprietary, so pointing the same code somewhere else later is a two line diff. We think that is the only honest way to build infrastructure.
SDKs: python, node, go, and raw curl all work.
# before: a provider endpoint
OPENAI_BASE_URL=https://api.provider.example/v1
OPENAI_API_KEY=sk-...
# after: the Liquid Whale gateway, same shape
OPENAI_BASE_URL=https://api.liquidwhale.com/v1
OPENAI_API_KEY=lw-sk-your-keyReference / endpoints
The full surface
Eight routes on one base URL. Modality tells you which body shape to send; catalog and usage routes are read-only.
| Method | Path | Serves | Notes | Free |
|---|---|---|---|---|
| POST | /v1/chat/completions | text | Chat and instruction models with streaming, tool calls, and structured output. | yes |
| POST | /v1/completions | text | Legacy text completion shape, kept identical for older stacks. | paid |
| POST | /v1/images/generations | image | Text to image across frontier diffusion models, billed per image. | paid |
| POST | /v1/videos/generations | video | Text to video with shot length and resolution controls, billed per second. | paid |
| POST | /v1/audio/transcriptions | speech-to-text | Audio in, timestamped text out, with automatic language detection. | paid |
| POST | /v1/audio/speech | text-to-speech | Text in, natural audio out, in dozens of voices and languages. | paid |
| GET | /v1/models | catalog | Every routed model with context window, modality, and live price. | yes |
| GET | /v1/usage | analytics | Spend, tokens, and latency per key, model, and day. | yes |
Routing headers
Debugging is reading a header
Every response reports the region, the provider that served it, how many attempts the router made, and a request id that ties into log search. When something looks off, the trail is already in your logs.
# every response carries the routing trail
x-lw-route: fra1
x-lw-provider: solstice
x-lw-attempts: 1
x-lw-request-id: req_9f2c41d8Reference / errors and limits
Error codes, rate limits, and the failover ladder
What the gateway does on each failure class, which codes retry automatically, and the limits on every tier.
Provider returns 429 rate limit
Retry on the next provider serving the same model class, with jittered backoff.
x-lw-attempts: 2Provider returns 5xx or times out
Circuit opens, traffic reroutes instantly, provider is retested in the background.
x-lw-failover: provider-2Provider degrades above latency budget
Weighted routing shifts share toward faster healthy providers.
x-lw-route: eu-fastModel unavailable in a region
Request lands on the nearest healthy replica that passes eval parity.
x-lw-region: fra1
Error codes
| Code | Name | Meaning | Retry |
|---|---|---|---|
| 400 | invalid_request_error | Malformed body, unknown parameter, or unsupported value. | No |
| 401 | authentication_error | Missing or revoked API key, or wrong key prefix lw-sk-. | No |
| 402 | insufficient_credits | Prepaid balance exhausted. Top up or enable auto recharge. | After top-up |
| 404 | model_not_found | Unknown model id for your key. List /v1/models to verify. | No |
| 429 | rate_limit_error | Key or model concurrency limit hit. The gateway retries upstream for you. | Automatic |
| 500 | routing_error | No healthy provider available for this model right now. | Yes, with backoff |
Rate limits
| Scope | Free tier | Usage tier |
|---|---|---|
| Requests per minute | 50 | 1,000, scaling on request |
| Concurrent requests | 4 | 200, scaling on request |
| Free-tier eligible models | All flagged rows in the catalog | Same, plus everything paid |
| Monthly free allowance | Every eligible model, unlimited days | Not applicable |
| Key payload size | 10 MB | 50 MB |
Docs done, time to call
Your key is one signup away
Create the key, export two variables, and rerun your eval suite against any model on the leaderboard. The free tier is waiting.