AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.9 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Meta reinforcement learning based resource pre-caching for emergency rescue edge cloud services

Peizu SHAO1Churan ZHOU1Siwei JIANG1Shangjing LIN1( )Weifeng SHAN2Weimin WU3Jilong LI4
School of Electronic Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China
Unit 96901 of the Chinese People's Liberation Army, Beijing 100080, China
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China
Institute of Internet Audiovisual Technology, Academy of Broadcasting Science, National Radio and Television Administration, Beijing 100866, China
Show Author Information

Abstract

Objective

In disaster emergencies, edge cloud service resource pre-caching faces two major challenges. First, post-disaster service demands exhibit strong burstiness and temporal dynamics. Insufficient pre-cached resources may cause request queuing and prolonged response latency, whereas excessive pre-caching wastes computing resources and incurs additional operating costs. Second, because disaster events are rare, historical data are insufficient for newly affected regions, new service types, and newly deployed edge nodes. Conventional prediction and reinforcement learning methods therefore struggle to obtain effective pre-caching policies rapidly under few-sample conditions. To address these challenges, this paper proposes an edge cloud service resource pre-caching method based on meta-reinforcement learning (Meta-RL).

Methods

Because user behavior data from real disaster scenarios are scarce, this paper analyzes service traffic in routine operational settings and trains the agent through interactions with daily service environments. The goal is to learn transferable workload evolution patterns and adapt quickly to emergency service modes. User request traffic is analyzed across Internet data centers (IDCs) and service types. The results show that traffic variations among IDCs exhibit correlations at monthly and intraday scales, indicating transferable common structures across regional workloads. Meanwhile, differences in request baselines, peak intensities, and fluctuation ranges reveal environmental heterogeneity. Across service types, different services share temporal trends but differ in request scale, peak duration, dispersion degree, and tail-load distribution. These observations indicate that daily service scenarios provide transferable experience, whereas data distribution differences hinder independent reinforcement learning training. Therefore, a Meta-RL framework is constructed to learn transferable policy initialization parameters through multitask training.

Results

In the proposed framework, the edge cloud service resource pre-caching process is formulated as a Markov decision process. The agent determines the number of resources to pre-cache in the next time slot based on the current resource queue length and request waiting queue length. The reward function jointly constrains redundant cached resources and waiting requests, guiding the agent to balance service responsiveness and resource efficiency. To evaluate this trade-off, a comprehensive pooling indicator is designed from request hit capability and resource waste rate. The hit rate measures the ability of pre-cached resources to immediately satisfy user requests, whereas the waste rate characterizes the proportion of cached resources that remain unused. For policy optimization, proximal policy optimization (PPO) is adopted as the basic framework, and a meta proximal policy optimization (METAPPO) algorithm is developed. METAPPO performs task-specific adaptation through inner-loop updates and optimizes shared initial policy parameters via outer-loop updates, improving adaptability to new IDCs, service types, and emergency tasks.

Conclusions

Simulation experiments compare METAPPO with PPO and other baseline methods across edge cloud service environments. The results show that METAPPO accelerates convergence across multiple scenarios. After a scenario change, METAPPO still exhibits favorable convergence performance, demonstrating that the meta-learning mechanism improves policy adaptation efficiency. Although some conventional methods may achieve competitive steady-state performance after sufficient training, they usually require more historical samples and longer training processes. Conversely, METAPPO leverages cross-task experience from daily service scenarios and rapidly generates effective resource pre-caching policies under few-sample conditions, supporting fast resource allocation adaptation in multiregion, multiservice, and newly emerging emergency edge cloud scenarios.

CLC number: TP393.1 Document code: A Article ID: 1000-0054(2026)09-1890-12

References

【1】
【1】
 
 
Journal of Tsinghua University (Science and Technology)
Pages 1890-1901

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
SHAO P, ZHOU C, JIANG S, et al. Meta reinforcement learning based resource pre-caching for emergency rescue edge cloud services. Journal of Tsinghua University (Science and Technology), 2026, 66(9): 1890-1901. https://doi.org/10.16511/j.cnki.qhdxxb.2026.27.052

7

Views

0

Downloads

0

Crossref

0

Scopus

0

CSCD

Received: 13 April 2026
Published: 14 September 2026
© Journal of Tsinghua University (Science and Technology). All rights reserved.