Now reading
Inference Engineering
Baseten · Oct 2026
Digging into how model serving actually works under the hood — batching, KV cache, and why latency is a systems problem more than a model problem.
Essays, posts, and docs I'm working through at the moment, plus a short archive of what I've finished. One line on why each one is worth the time.
Now reading
Baseten · Oct 2026
Digging into how model serving actually works under the hood — batching, KV cache, and why latency is a systems problem more than a model problem.
1 in progress · 0 finished