How model serving actually behaves once it meets production traffic.
A community resource for inference engineering. Measurement you can reproduce, benchmarks built around real workloads rather than leaderboard rank, and the reading that explains your latency and your bill.
Sections
inference.academy/
- feed/
Papers, posts and release notes worth the read, with a note on why.
- benchmarks/ (not built yet)
What a workload costs and how it feels, per model, per host.
- explainers/
Simulations you can run for the concepts that serving actually turns on.
- events/ (not built yet)
Talks and meetups, listed when they are real and dated.