AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.1 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Feature-Wise Linear Modulation for Heterogeneous-Frequency Multimodal Fusion in Temporal Sequence Encoders

Maurice Kyla OctavianoJin-Taek Seong( )
Graduate School of Data Science, Chonnam National University, Gwangju, Republic of Korea
Show Author Information

Abstract

Integrating high-frequency sequential signals with low-frequency contextual descriptors into a unified deep encoder is a recurring challenge in computational modelling, exemplified by cross-sectional stock ranking where price dynamics must be jointly modelled with quarterly accounting fundamentals. Existing approaches use late concatenation, where the contextual signal influences only the final prediction head and cannot shape upstream feature extraction. We propose Feature-wise Linear Modulation (FiLM) as an intermediate conditioning mechanism: fundamentals generate per-channel scaling (gamma) and shifting (beta) parameters that affinely transform the encoder’s intermediate representations before aggregation. The same price sequence thus yields different temporal features depending on the firm’s fundamental profile, which we hypothesise reduces signal variability across heterogeneous market regimes by allowing the encoder to amplify or suppress patterns based on contextual quality. We instantiate FiLM across recurrent (LSTM), convolutional (TCN), and attention-based (iTransformer) encoders, evaluated on China A-share equities (2010–2024). Across the convolutional and recurrent encoder families, the primary benefit of FiLM conditioning is improved signal stability—formally measured as the standard deviation of RankIC across rebalancing dates—and risk-adjusted performance, rather than mean predictive accuracy. The gain depends critically on where modulation is applied: pre-aggregation conditioning on temporally-rich representations produces the largest variance reduction. FiLM-TCN, which modulates the convolutional feature map before pooling, achieves RankIC of 0.1415, annualised Sharpe of 1.633, and IC hit rate of 80.4% net of transaction costs. The insight that intermediate conditioning improves signal stability rather than raw accuracy may inform analogous fusion problems in other sequential modelling domains.

References

【1】
【1】
 
 
Computers, Materials & Continua
Article number: 80

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Octaviano MK, Seong J-T. Feature-Wise Linear Modulation for Heterogeneous-Frequency Multimodal Fusion in Temporal Sequence Encoders. Computers, Materials & Continua, 2026, 88(3): 80. https://doi.org/10.32604/cmc.2026.082842

7

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 24 March 2026
Accepted: 27 May 2026
Published: 23 July 2026
© The Author 2026.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.