Abstract
Large Language Models (LLMs) are transforming recommendation into a generative task, yet their industrial deployment is hindered by the prohibitive latency of processing long, personalized prompts where standard prefix caching fails due to non-contiguous reuse patterns. To address this, we propose RcLLM, a distributed inference system that accelerates generative recommendation via Beyond-Prefix KV Caching. RcLLM decomposes prompts into reusable blocks and manages the massive scale of item catalogs through a stratified distributed storage architecture: compact user histories are replicated for zero-latency retrieval, while massive item caches are sharded using a similarity-aware placement algorithm. RcLLM effectively eliminates redundant quadratic computation with two strategic techniques: 1) an affinity-based global scheduler to maximize data locality, and 2) a selective attention mechanism to correct approximation errors. Evaluations on real-world datasets demonstrate that RcLLM reduces Time-To-First-Token (TTFT) by 1.31×-9.51× compared to state-of-the-art prefix caching systems, enabling real-time serving compliance with negligible impact on recommendation accuracy.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - 2026 IEEE 46th International Conference on Distributed Computing Systems, ICDCS 2026 |
| Editors | Cristina Ceballos |
| Publisher | IEEE |
| Pages | 642-652 |
| Number of pages | 11 |
| ISBN (Electronic) | 9798319529794 |
| ISBN (Print) | 9798319529800 |
| DOIs | |
| Publication status | Published - 22 Jun 2026 |
| Event | 2026 IEEE 46th International Conference on Distributed Computing Systems, ICDCS 2026 - Seoul, Korea, Republic of Duration: 22 Jun 2026 → 25 Jun 2026 |
Publication series
| Name | Proceedings - International Conference on Distributed Computing Systems |
|---|---|
| ISSN (Print) | 1063-6927 |
| ISSN (Electronic) | 2575-8411 |
Conference
| Conference | 2026 IEEE 46th International Conference on Distributed Computing Systems, ICDCS 2026 |
|---|---|
| Country/Territory | Korea, Republic of |
| City | Seoul |
| Period | 22/06/26 → 25/06/26 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 9 Industry, Innovation, and Infrastructure
User-Defined Keywords
- inference optimization
- kv cache reuse
- large language models
- recommendation systems
Fingerprint
Dive into the research topics of 'RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver