Writing
RSS·
Inference Lab
Field reports on inference performance, from HTTP buffering and Metal traces to memory limits, benchmark lineage, and vLLM compilation.
Aug 30 - Sep 20, 2026 9 reports 120 min
Published research
Understanding GPU-Level Bottlenecks in Large Language Model InferencePeer-reviewed conference paper · Springer · 2026
- 01 Your LLM's time to first token might be measuring your HTTP client August 30, 2026 8 min
- 02 A Metal trace is not your workload until you attribute it by process August 31, 2026 8 min
- 03 The weights fit. The inference workload didn't. September 1, 2026 14 min
- 04 Parsing 7,585 XML references to count 400 Metal dispatches September 2, 2026 11 min
- 05 The fastest passing system was not the cheapest one September 3, 2026 12 min
- 06 A benchmark result without lineage is just a screenshot September 4, 2026 25 min
- 07 The compiled run had a lower request-latency sum. I still could not claim break-even. September 5, 2026 18 min
- 08 A cache hit is not proof that you skipped the work September 13, 2026 12 min
- 09 Context should be a build artifact, not a prompt assembled at runtime. September 20, 2026 12 min
Pinned
More writing