Inference & Serving

Sampling, the KV cache, quantization, batching — where latency and cost are made