AI Gateways
One door in front of many models: route by model, survive a dead backend with fallbacks, rate-limit per key, answer repeats from cache, and know what every token costs.
An interactive AI Infrastructure lesson: 20 steps, about 32 minutes, on a live simulation in your browser.
The shop's support bot runs on Llama 3.1 8B, cheap and fast. The code assistant needs Llama 3.1 70B, slower and smarter. Each is already a service: chat on one GPU, code tensor-parallel across four.
Clients should not know any of that. An AI gateway is one HTTP door in front of many models: it reads which model each request asks for, routes it to the right service, and answers for everything around serving — retries, limits, caching and the bill. This lesson builds edge, the shop's gateway, one job at a time.
What you will learn
One door, many models
- Two models behind one door: A gateway separates what clients ask for (a model) from what serves it (replicas on GPUs). Routes map the first to the second.
- Follow one request through: Every gateway request is route, then forward, then account. The serving numbers belong to the service; the routing numbers belong to the gateway.
Routing by model
- The other model goes elsewhere: Routes match on the model name the client sends. The name is a contract between the client and the gateway.
- A model with no route: An unmatched request is refused, not guessed. Defaults are routes too: written, visible, and billed to the right model.
Every token has a price
- Same question, two prices: Cost per request ≈ tokens × price. The gateway multiplies what it already counted (tokens) by what only it knows (the price list).
- Where the easy questions go: Model routing is cost control. The cheapest model that answers well is the right default; the price list makes the default stick.
- Drill: price one request: Spend = (prompt + output tokens) ÷ 1M × price. Four decimal places of nothing per request; millions of requests per line item.
- Prices are policy, not physics: Services report physics; the gateway reports money. Keep the price list next to the routes, because both are policy.
When a backend dies
- Chat goes dark at 2 a.m.: A fallback trades money and latency for availability. A slow expensive answer beats the correct error page.
- The fallback's bill: Every fallback has a price gap and the incident multiplies it. Route back as soon as the primary is healthy.
- Both backends dark: 502 means the gateway kept its promise and the backends broke theirs. Retries belong above this layer, not inside it.
- Bring chat back: Routes are static; health is dynamic. Recovery means the same requests match the same routes and land somewhere healthy again.
Rate limits per key
- One key, one bucket: A rate limit is a per-key bucket checked before the expensive work. Fitting traffic never notices it.
- The scraper at midnight: Reject at the cheapest layer that can decide. A gateway 429 costs nothing; a service 429 costs the wait that discovered it.
- Drill: size the bucket: Size a limit between the peak you must serve and the abuse you must stop, with headroom on both sides.
Answering from cache
- The same question twice: An exact-prompt hit skips prefill and decode both. Gateway cache is for repeats; prefix cache is for shared beginnings.
- One word changes: Exact cache never answers a question it was not asked. If near-duplicates dominate, that is a different cache, not a looser exact one.
- Drill: count the service calls: Cache math: one miss warms it, every repeat inside the TTL is free.
Recap & playground
- Cheat sheet
- Playground