Writing
RSS·
Inference Lab
Seven field reports on inference performance, from HTTP buffering and Metal traces to memory limits, benchmark lineage, and vLLM compilation.
Aug 30 - Sep 5, 2026 7 reports 96 min
- 01 Your LLM's time to first token might be measuring your HTTP client August 30, 2026 8 min
- 02 A Metal trace is not your workload until you attribute it by process August 31, 2026 8 min
- 03 The weights fit. The inference workload didn't. September 1, 2026 14 min
- 04 Parsing 7,585 XML references to count 400 Metal dispatches September 2, 2026 11 min
- 05 The fastest passing system was not the cheapest one September 3, 2026 12 min
- 06 A benchmark result without lineage is just a screenshot September 4, 2026 25 min
- 07 The compiled run had a lower request-latency sum. I still could not claim break-even. September 5, 2026 18 min
Pinned
More writing