Peer-reviewed conference paper · Springer, Cham ·
Understanding GPU-Level Bottlenecks in Large Language Model Inference
Shubhanshu Kushwaha · Mamata Samal · Siddhant Khare
ICDEC 2025 proceedings · Lecture Notes in Networks and Systems, vol. 2003 · pp. 211-224
My first public research paper, based on LLMTraceFX, the open-source profiler I built. It examines memory bandwidth, GPU interconnects, kernel overhead, and prefill/decoding.
- My contribution
- Software · Conceptualization · Investigation · Methodology · Writing - original draft · Validation