TY - JOUR
T1 - Cooperative Edge Caching for Profit Maximization Based on Multi-agent Deep Reinforcement Learning
AU - Liang, Yuzhu
AU - Huang, Haiyang
AU - Xu, Changfu
AU - Zou, Haodong
AU - Fan, Xinggang
AU - Li, Yupeng
AU - Wang, Tian
N1 - The above work was supported in part by the Joint Funds of the National Natural Science Foundation of China under Grant U25A20436, the National Natural Science Foundation of China (NSFC) (62372047), Guangxi Key Research & Development Program (FN2504240036,2025FN96441087), Guangdong S&T Programme (No. 2025B0101120006), the Natural Science Foundation of Guangdong Province (2024A1515011323), the Supplemental Funds for Major Scientific Research Projects of Beijing Normal University, Zhuhai (ZHPT2023002), the Fundamental Research Funds for the Central Universities, Guangdong Province Educational Science Planning Research Project under 2025JKZG069, Higher Education Research Topics of Guangdong Association of Higher Education in the 14th Five-Year Plan under 24GYB207.
Publisher Copyright:
© 2002-2012 IEEE.
PY - 2026/7/6
Y1 - 2026/7/6
N2 - The rapid growth of multimedia services and latency-sensitive applications such as online gaming has imposed stringent requirements on content delivery, rendering traditional cloud-centric architectures increasingly inadequate. Edge caching offers a promising solution, yet its efficiency is fundamentally constrained by limited storage, dynamic demand, and strategic interactions among distributed Edge nodes (ENs). This paper studies cooperative edge caching through the lens of joint optimization of caching, pricing, and request scheduling (JOCPRS) and develops an EC-MADRL framework, where multi-agent deep reinforcement learning is leveraged to support scalable and adaptive decision-making in dynamic environments. We formulate the JOCPRS problem as a hierarchical cooperative game, in which autonomous ENs jointly consider pricing, caching configurations, and request scheduling to maximize network wide profit while minimizing user access costs. The JOCPRS problem is shown to be NP-hard. To address the challenges of dynamic strategy adaptation, inter-agent coupling, and conflict resolution, we design a centralized-training–distributed-execution MADRL architecture that learns stable cooperative pricing policies under given caching configurations and drives the system toward stable hierarchical outcomes. The proposed framework effectively balances cooperation and competition among ENs, enabling efficient resource utilization under time-varying demand and network conditions. Edge testbed experiments indicate that EC-MADRL consistently outperforms state-of-the-art baselines, including pipeline-optimized LRU (P4LRU), single agent DDQN-based caching (DDQN-ECMP), and conventional MADRL approaches. Specifically, EC-MADRL achieves an average of 25.58% higher network profit, 43.32% lower user latency, and 17.30% reduction in user access cost, validating its effectiveness and robustness for large-scale cooperative systems.
AB - The rapid growth of multimedia services and latency-sensitive applications such as online gaming has imposed stringent requirements on content delivery, rendering traditional cloud-centric architectures increasingly inadequate. Edge caching offers a promising solution, yet its efficiency is fundamentally constrained by limited storage, dynamic demand, and strategic interactions among distributed Edge nodes (ENs). This paper studies cooperative edge caching through the lens of joint optimization of caching, pricing, and request scheduling (JOCPRS) and develops an EC-MADRL framework, where multi-agent deep reinforcement learning is leveraged to support scalable and adaptive decision-making in dynamic environments. We formulate the JOCPRS problem as a hierarchical cooperative game, in which autonomous ENs jointly consider pricing, caching configurations, and request scheduling to maximize network wide profit while minimizing user access costs. The JOCPRS problem is shown to be NP-hard. To address the challenges of dynamic strategy adaptation, inter-agent coupling, and conflict resolution, we design a centralized-training–distributed-execution MADRL architecture that learns stable cooperative pricing policies under given caching configurations and drives the system toward stable hierarchical outcomes. The proposed framework effectively balances cooperation and competition among ENs, enabling efficient resource utilization under time-varying demand and network conditions. Edge testbed experiments indicate that EC-MADRL consistently outperforms state-of-the-art baselines, including pipeline-optimized LRU (P4LRU), single agent DDQN-based caching (DDQN-ECMP), and conventional MADRL approaches. Specifically, EC-MADRL achieves an average of 25.58% higher network profit, 43.32% lower user latency, and 17.30% reduction in user access cost, validating its effectiveness and robustness for large-scale cooperative systems.
KW - Edge caching
KW - Edge computing
KW - multi-agent deep reinforcement learning
KW - profit maximization
UR - https://www.scopus.com/pages/publications/105044354973
U2 - 10.1109/TMC.2026.3710512
DO - 10.1109/TMC.2026.3710512
M3 - Journal article
AN - SCOPUS:105044354973
SN - 1536-1233
JO - IEEE Transactions on Mobile Computing
JF - IEEE Transactions on Mobile Computing
ER -