As application migration to the cloud becomes the mainstream approach for deployment, application runtime management requires large-scale workload prediction to ensure resource efficiency and system stability. However, existing forecasting models primarily focus on improving accuracy, often overlooking the impact of storage, training time, and inference time, leading to excessive computational overhead. Furthermore, cloud workloads exhibit high heterogeneity, with diverse patterns across containers, making it difficult for conventional models to generalize effectively. To address these challenges, this paper proposes FELT, a large-scale cloud workload prediction model through adaptive feature-enhanced and similarity-aware Transformer. FELT characterizes workload dynamics from both waveform and value perspectives, reducing model complexity while capturing macroscopic and detailed variations. We introduce a feature-enhanced workload similarity-aware algorithm that adaptively groups containers with similar workload patterns in real time by analyzing both historical and recent similarity, improving robustness in heterogeneous environments. Additionally, we leverage Transformer with customized position encoding and attention masks based on real-time workload similarity and employ multi-head self-attention for parallelized training, achieving a balance between accuracy and efficiency. Extensive experiments on public datasets demonstrate FELT’s superiority in both prediction accuracy and overhead, with ablation studies further validating the effectiveness of its components.
Publications
- Article type
- Year
- Co-author
Year
Open Access
Online First
Tsinghua Science and Technology
Published: 26 September 2025
Downloads:140
Total 1
京公网安备11010802044758号