Skip to main navigation Skip to search Skip to main content

Cooperative Edge Caching for Profit Maximization Based on Multi-agent Deep Reinforcement Learning

  • Yuzhu Liang
  • , Haiyang Huang
  • , Changfu Xu
  • , Haodong Zou
  • , Xinggang Fan
  • , Yupeng Li
  • , Tian Wang*
  • *Corresponding author for this work

Research output: Contribution to journalJournal articlepeer-review

Abstract

The rapid growth of multimedia services and latency-sensitive applications such as online gaming has imposed stringent requirements on content delivery, rendering traditional cloud-centric architectures increasingly inadequate. Edge caching offers a promising solution, yet its efficiency is fundamentally constrained by limited storage, dynamic demand, and strategic interactions among distributed Edge nodes (ENs). This paper studies cooperative edge caching through the lens of joint optimization of caching, pricing, and request scheduling (JOCPRS) and develops an EC-MADRL framework, where multi-agent deep reinforcement learning is leveraged to support scalable and adaptive decision-making in dynamic environments. We formulate the JOCPRS problem as a hierarchical cooperative game, in which autonomous ENs jointly consider pricing, caching configurations, and request scheduling to maximize network wide profit while minimizing user access costs. The JOCPRS problem is shown to be NP-hard. To address the challenges of dynamic strategy adaptation, inter-agent coupling, and conflict resolution, we design a centralized-training–distributed-execution MADRL architecture that learns stable cooperative pricing policies under given caching configurations and drives the system toward stable hierarchical outcomes. The proposed framework effectively balances cooperation and competition among ENs, enabling efficient resource utilization under time-varying demand and network conditions. Edge testbed experiments indicate that EC-MADRL consistently outperforms state-of-the-art baselines, including pipeline-optimized LRU (P4LRU), single agent DDQN-based caching (DDQN-ECMP), and conventional MADRL approaches. Specifically, EC-MADRL achieves an average of 25.58% higher network profit, 43.32% lower user latency, and 17.30% reduction in user access cost, validating its effectiveness and robustness for large-scale cooperative systems.

Original languageEnglish
Number of pages14
JournalIEEE Transactions on Mobile Computing
DOIs
Publication statusE-pub ahead of print - 6 Jul 2026

User-Defined Keywords

  • Edge caching
  • Edge computing
  • multi-agent deep reinforcement learning
  • profit maximization

Fingerprint

Dive into the research topics of 'Cooperative Edge Caching for Profit Maximization Based on Multi-agent Deep Reinforcement Learning'. Together they form a unique fingerprint.

Cite this