AI model gateway

One key. Every model. One bill.

Liquid Whale fronts hundreds of models from 40+ providers behind one OpenAI-compatible API. Smart routing with automatic fallbacks, cost analytics on every request, and per-token prices passed through with zero markup.

Your first request

hello-gateway.sh
curl https://api.liquidwhale.com/v1/chat/completions \
  -H "Authorization: Bearer $LIQUIDWHALE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "helios-5",
    "messages": [
      { "role": "user", "content": "Hello, gateway." }
    ]
  }'

Swap the base URL and key. The request shape stays the same.

Routing uptime

99.98%

Models routed

300+

Providers

40+

Markup

0%

  • Helios 5
  • Meridian Max 4.1
  • Tycho 2.5 Deep
  • Foghorn 4 Coastal
  • Northbeam V3.1
  • Vela 3 Max
  • Aurora Large 3
  • Kestrel 4
  • Beacon A
  • Argosy K2
  • Tidal 1.1 Pro
  • Maelstrom 3
  • Echo large-v3
  • Siren Multilingual v2
  • 0 +

    models routed

  • 0 +

    providers on one bill

  • 0.00 %

    routing uptime, 90 days

  • 0 %

    markup on provider prices

01 / The gateway

Everything between your code and every provider, handled

Six things the gateway does so your team does not have to. Each one ships enabled by default, on every key, at every price.

  • 01

    One key, one bill, one SDK

    A single API key fronts hundreds of models from 40+ providers. One invoice at the end of the month, one dashboard, one client library to maintain.

    1 key, 300+ models

  • 02

    Drop-in compatible

    The request shape is OpenAI-compatible. Change the base URL and the key, keep your code. Teams migrate in an afternoon, not a quarter.

    base url + key

  • 03

    Smart routing

    Every request is scored against live latency, price, and capacity, then sent to the best healthy provider for that model class.

    scored per request

  • 04

    Automatic fallbacks

    Provider outage or rate limit? The gateway retries the next healthy provider and reports the attempt chain in response headers.

    x-lw-failover

  • 05

    Cost tracking and analytics

    Per key, per model, per day: tokens, spend, and latency. Set budgets, get alerts, and export everything through the usage endpoint.

    GET /v1/usage

  • 06

    Zero markup on tokens

    You pay the provider price, per token, with no platform fee folded in. What the catalog lists is what the invoice shows.

    0% platform fee

02 / Surface

Eight endpoints, five modalities, one base URL

Text, image, video, speech to text, and text to speech, plus catalog and usage. Everything hangs off the same authenticated prefix.

Gateway API endpoints
MethodPathServesNotesFree
POST/v1/chat/completionstextChat and instruction models with streaming, tool calls, and structured output.yes
POST/v1/completionstextLegacy text completion shape, kept identical for older stacks.paid
POST/v1/images/generationsimageText to image across frontier diffusion models, billed per image.paid
POST/v1/videos/generationsvideoText to video with shot length and resolution controls, billed per second.paid
POST/v1/audio/transcriptionsspeech-to-textAudio in, timestamped text out, with automatic language detection.paid
POST/v1/audio/speechtext-to-speechText in, natural audio out, in dozens of voices and languages.paid
GET/v1/modelscatalogEvery routed model with context window, modality, and live price.yes
GET/v1/usageanalyticsSpend, tokens, and latency per key, model, and day.yes

Cookbook

Copy, paste, and you are routing

Four recipes covering the shapes most teams ship first. Tabs switch between curl, python, and node, and every block copies with one click.

Chat completion

The classic call, identical shape to the SDK you already use.

chat.curl
curl https://api.liquidwhale.com/v1/chat/completions \
  -H "Authorization: Bearer $LIQUIDWHALE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "helios-5",
    "messages": [
      { "role": "user", "content": "Summarize this changelog in one line." }
    ]
  }'

Streaming

Server-sent events, token by token, with the same delta shape.

stream.curl
curl https://api.liquidwhale.com/v1/chat/completions \
  -H "Authorization: Bearer $LIQUIDWHALE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meridian-pro-4.5",
    "stream": true,
    "messages": [
      { "role": "user", "content": "Draft a one-paragraph release note." }
    ]
  }'

Image and speech to text

Diffusion models and transcription under the same key.

media.curl
# text to image
curl https://api.liquidwhale.com/v1/images/generations \
  -H "Authorization: Bearer $LIQUIDWHALE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tidal-1.1-pro",
    "prompt": "Aerial view of deep blue waves, white foam",
    "size": "1536x1024"
  }'

# speech to text
curl https://api.liquidwhale.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $LIQUIDWHALE_KEY" \
  -F model="echo-large-v3" \
  -F file="@meeting.mp3" \
  -F response_format="verbose_json"

Speech and video

Voices out, video out, billed per character and per second.

voice.curl
# text to speech
curl https://api.liquidwhale.com/v1/audio/speech \
  -H "Authorization: Bearer $LIQUIDWHALE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "siren-v2",
    "voice": "aurora",
    "input": "One key. One bill. The whole ocean."
  }'

# text to video
curl https://api.liquidwhale.com/v1/videos/generations \
  -H "Authorization: Bearer $LIQUIDWHALE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "maelstrom-3",
    "prompt": "Slow aerial over ocean swell at dawn",
    "seconds": 8
  }'

Why a whale

Whales cross the whole ocean and never touch the bottom. Your traffic should work the same way.

A whale is enormous, calm, and everywhere at once. That is the property we wanted for model traffic: deep enough to carry anything, stable enough to route around storms, and wide enough that every provider sits under the same surface. The gateway is the ocean, the models are what moves through it.

Read the full story
The tail of a whale rising out of calm misty water
fig. 01 Surface first. Depth on demand.

04 / Resilience

Outages happen upstream. Downtime does not have to be yours

The routing layer watches every provider continuously. When one wobbles, your traffic moves and the response headers tell the story.

  • Provider returns 429 rate limit

    Retry on the next provider serving the same model class, with jittered backoff.

    x-lw-attempts: 2
  • Provider returns 5xx or times out

    Circuit opens, traffic reroutes instantly, provider is retested in the background.

    x-lw-failover: provider-2
  • Provider degrades above latency budget

    Weighted routing shifts share toward faster healthy providers.

    x-lw-route: eu-fast
  • Model unavailable in a region

    Request lands on the nearest healthy replica that passes eval parity.

    x-lw-region: fra1

Error codes

CodeNameMeaningRetry
400invalid_request_errorMalformed body, unknown parameter, or unsupported value.No
401authentication_errorMissing or revoked API key, or wrong key prefix lw-sk-.No
402insufficient_creditsPrepaid balance exhausted. Top up or enable auto recharge.After top-up
404model_not_foundUnknown model id for your key. List /v1/models to verify.No
429rate_limit_errorKey or model concurrency limit hit. The gateway retries upstream for you.Automatic
500routing_errorNo healthy provider available for this model right now.Yes, with backoff

Rate limits

ScopeFree tierUsage tier
Requests per minute501,000, scaling on request
Concurrent requests4200, scaling on request
Free-tier eligible modelsAll flagged rows in the catalogSame, plus everything paid
Monthly free allowanceEvery eligible model, unlimited daysNot applicable
Key payload size10 MB50 MB

The ocean underneath

Two hundred data centers, forty providers, one route table. The swell moves, the surface stays level.

fig. 02, sea foam, open water

03 / Leaderboard

What the world actually routes, refreshed weekly

Ranked by routed tokens over the last 30 days across the gateway. Snapshot: Week 39, 2026. Every figure comes from real traffic, not vendor claims.

Text model leaderboard by routed tokens, 30 days
#ModelProviderTokens, 30dRequests, 30dShareP50 latencyContextPrice in/outFree
01Helios 5 Solstice Labs412.8B148.2M320 ms400k$1.25 / $10.00-
02Meridian Pro 4.5 Meridian AI356.2B121.7M280 ms200k$3.00 / $15.00-
03Tycho 2.5 Swift Tycho Systems289.6B174.3M190 ms1M$0.30 / $2.50-
04Foghorn 4 Coastal freeFoghorn AI201.4B96.8M240 ms1M$0.22 / $0.85yes
05Northbeam V3.1 freeNorthbeam168.9B71.2M410 ms128k$0.56 / $1.68yes
06Helios 5 Mini Solstice Labs122.7B139.5M210 ms400k$0.25 / $2.00-

Full leaderboard, all 14 text models

05 / Pricing

Provider prices, passed straight through

No platform fee on tokens. The rate in the catalog is the rate on your invoice, whether you route one request or ten billion.

Token prices per one million tokens
ModelProviderInput / 1MOutput / 1MContextFree tier
Helios 5 Solstice Labs$1.25$10.00400k-
Helios 5 Mini Solstice Labs$0.25$2.00400k-
Meridian Pro 4.5 Meridian AI$3.00$15.00200k-
Meridian Max 4.1 Meridian AI$15.00$75.00200k-
Tycho 2.5 Deep Tycho Systems$1.25$10.001M-
Tycho 2.5 Swift Tycho Systems$0.30$2.501M-

All prices and the free tier

Get started

Point your stack at the ocean

Create a key, change two lines, and route your first request in minutes. The free tier covers every model flagged free, forever, with the same routing as paid traffic.