How model serving actually behaves once it meets production traffic.
A community resource for inference engineering. Measurement you can reproduce, benchmarks built around real workloads rather than leaderboard rank, and the reading that explains your latency and your bill.
Sections
inference.academy/
- feed/
Papers, posts and release notes worth the read, with a note on why.
- explainers/
Simulations you can run for the concepts that serving actually turns on.
- at home/
Running models on your own Mac or GPU: what fits, which engine, and how to experiment without buying hardware.
- api bill/
What drives an LLM API bill, whether to route through OpenRouter, and where the real discounts are.
- glossary/
Sixty terms, each with one number that pins it to a real machine.