AI
How the models actually run: what the serving layer does with your request, what it costs, and which behaviour you can predict rather than measure.
- LLMs
Inference as a system with a cost model, not an oracle — caching, context, tokens and latency.
1 post