Docs

From zero to first routed token in about five minutes

The gateway speaks the OpenAI request shape. If your stack already calls a compatible endpoint, switching is a base URL, a key, and a deploy.

Quickstart

Three steps, one deploy

Sign up, point your client at the gateway, and pass a model id from the catalog. Everything after that is your existing code.

  1. 01

    Create a key

    Sign up, copy your key. It starts with lw-sk-. One key reaches every provider in the catalog.

  2. 02

    Point at the gateway

    Set the base URL to https://api.liquidwhale.com/v1 and swap the key. The request shape is OpenAI-compatible, so nothing else moves.

  3. 03

    Call any model

    Pass the model id from the catalog. If a provider wobbles, the gateway reroutes and reports it in response headers.

Environment

Two exports, any shell. Your SDK picks them up untouched.

env.sh
export OPENAI_BASE_URL="https://api.liquidwhale.com/v1"
export OPENAI_API_KEY="lw-sk-your-key"

First call

The same request in all three flavors, copy button included.

quickstart.curl
curl https://api.liquidwhale.com/v1/chat/completions \
  -H "Authorization: Bearer $LIQUIDWHALE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "helios-5",
    "messages": [
      { "role": "user", "content": "Summarize this changelog in one line." }
    ]
  }'

Compatibility

The shape is open. The exit is one env var

Requests and responses keep the standard chat completion fields, including tool calls, streaming deltas, and structured output. Nothing is proprietary, so pointing the same code somewhere else later is a two line diff. We think that is the only honest way to build infrastructure.

SDKs: python, node, go, and raw curl all work.

migrate.env
# before: a provider endpoint
OPENAI_BASE_URL=https://api.provider.example/v1
OPENAI_API_KEY=sk-...

# after: the Liquid Whale gateway, same shape
OPENAI_BASE_URL=https://api.liquidwhale.com/v1
OPENAI_API_KEY=lw-sk-your-key

Reference / endpoints

The full surface

Eight routes on one base URL. Modality tells you which body shape to send; catalog and usage routes are read-only.

Gateway API endpoints
MethodPathServesNotesFree
POST/v1/chat/completionstextChat and instruction models with streaming, tool calls, and structured output.yes
POST/v1/completionstextLegacy text completion shape, kept identical for older stacks.paid
POST/v1/images/generationsimageText to image across frontier diffusion models, billed per image.paid
POST/v1/videos/generationsvideoText to video with shot length and resolution controls, billed per second.paid
POST/v1/audio/transcriptionsspeech-to-textAudio in, timestamped text out, with automatic language detection.paid
POST/v1/audio/speechtext-to-speechText in, natural audio out, in dozens of voices and languages.paid
GET/v1/modelscatalogEvery routed model with context window, modality, and live price.yes
GET/v1/usageanalyticsSpend, tokens, and latency per key, model, and day.yes

Routing headers

Debugging is reading a header

Every response reports the region, the provider that served it, how many attempts the router made, and a request id that ties into log search. When something looks off, the trail is already in your logs.

response.headers
# every response carries the routing trail
x-lw-route: fra1
x-lw-provider: solstice
x-lw-attempts: 1
x-lw-request-id: req_9f2c41d8

Reference / errors and limits

Error codes, rate limits, and the failover ladder

What the gateway does on each failure class, which codes retry automatically, and the limits on every tier.

  • Provider returns 429 rate limit

    Retry on the next provider serving the same model class, with jittered backoff.

    x-lw-attempts: 2
  • Provider returns 5xx or times out

    Circuit opens, traffic reroutes instantly, provider is retested in the background.

    x-lw-failover: provider-2
  • Provider degrades above latency budget

    Weighted routing shifts share toward faster healthy providers.

    x-lw-route: eu-fast
  • Model unavailable in a region

    Request lands on the nearest healthy replica that passes eval parity.

    x-lw-region: fra1

Error codes

CodeNameMeaningRetry
400invalid_request_errorMalformed body, unknown parameter, or unsupported value.No
401authentication_errorMissing or revoked API key, or wrong key prefix lw-sk-.No
402insufficient_creditsPrepaid balance exhausted. Top up or enable auto recharge.After top-up
404model_not_foundUnknown model id for your key. List /v1/models to verify.No
429rate_limit_errorKey or model concurrency limit hit. The gateway retries upstream for you.Automatic
500routing_errorNo healthy provider available for this model right now.Yes, with backoff

Rate limits

ScopeFree tierUsage tier
Requests per minute501,000, scaling on request
Concurrent requests4200, scaling on request
Free-tier eligible modelsAll flagged rows in the catalogSame, plus everything paid
Monthly free allowanceEvery eligible model, unlimited daysNot applicable
Key payload size10 MB50 MB

Docs done, time to call

Your key is one signup away

Create the key, export two variables, and rerun your eval suite against any model on the leaderboard. The free tier is waiting.