Skip to content
Ishaan Reddy

Cookbook · Inference

Deployment

2 mindeploymentinferencecost-modelingstreaming

Serving is not deployment

Serving covers running a model efficiently once it's up. Deployment is the layer above that: deciding where the model runs, how requests reach it, what happens when it's overloaded, and what it costs to keep running. A well-optimized serving stack behind a badly-designed deployment can still be slow, expensive, or unreliable in practice.

Where the model runs

Managed API (calling a hosted model over an API, your own or a third party's): no infrastructure to manage, but you're paying per-token and subject to someone else's rate limits, latency, and uptime.

Self-hosted server: full control over hardware, batching, and cost structure, but you own the operational burden (scaling, monitoring, failover) that a managed API absorbs for you.

On-device / edge: the model runs on the user's own hardware (a laptop, phone, or embedded device), which removes network latency and per-request cost entirely, but constrains model size hard to whatever that device can hold, usually pushing toward small models and aggressive quantization.

Application-layer concerns serving alone doesn't solve

Streaming: sending generated tokens to the client as they're produced rather than waiting for the full response, which matters for perceived latency even when total generation time is unchanged. Requires the application layer to handle partial, incrementally-arriving output, not just a single final response.

Rate limiting and backpressure: deciding what happens when demand exceeds capacity, queue requests, reject them, or degrade gracefully (a shorter response, a cheaper model), rather than letting the serving layer fall over under load.

Cost modeling: the real unit economics of a deployment are usually dollars per output token (or per request), which depends on model size, quantization, hardware utilization, and request patterns together, not any single one of those in isolation. A smaller, well-utilized deployment can beat a larger, underutilized one on cost even at lower raw throughput.

Where to look further

  • Serving: the efficiency techniques (batching, KV cache management, speculative decoding) that sit underneath any deployment.
  • Quantization: the main lever for making a model small enough to deploy on constrained hardware.