FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
Achieving low latency and high throughput under volatile demand requires deep understanding of real-world serving workloads, yet existing studies often rely on proxy traces or coarse-grained characterizations that fail to capture the heterogeneity of modern multi-model LLM platforms. We present FineServe, an in-the-wild, multi-model LLM serving workload dataset collected from a global commercial marketplace, enabling fine-grained characterization of real-world serving dynamics across heterogeneous models and tasks. Leveraging FineServe, we conduct a comprehensive analysis of arrival dynamics and token behavior, revealing fundamentally different fluctuation regimes across model architectures, scales and task intents.
- ▪Achieving low latency and high throughput under volatile demand requires deep understanding of real-world serving workloads, yet existing studies often rely on proxy traces or coarse-grained characterizations that fail to capture the hetero
- ▪We present FineServe, an in-the-wild, multi-model LLM serving workload dataset collected from a global commercial marketplace, enabling fine-grained characterization of real-world serving dynamics across heterogeneous models and tasks.
- ▪Leveraging FineServe, we conduct a comprehensive analysis of arrival dynamics and token behavior, revealing fundamentally different fluctuation regimes across model architectures, scales and task intents.
Opening excerpt (first ~120 words) tap to expand
Computer Science > Artificial Intelligence arXiv:2607.19349 (cs) [Submitted on 17 Apr 2026] Title:FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads Authors:Tiancheng Zhang, Shaoyuan Huang, Mingyuan Wang, Yunfeng Zhao, Xiaofei Wang, Wenyu Wang View a PDF of the paper titled FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads, by Tiancheng Zhang and 5 other authors View PDF HTML (experimental) Abstract:Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at arXiv.org.