Scaling Laws for Efficient Inference in Large Language Models
A. Mitchell, R. Patel, S. Kim, J. Hernandez, L. Chen
Scaling-law analysis for inference efficiency in transformer language models. We map model size, batch settings, and latency, and show configurations that raise throughput about 3.2x without measured quality loss.
Read on arXiv