AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Home iFuture Article
PDF (259.4 KB)
Collect
AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Original Article | Open Access | Just Accepted

Decoupled visual processing: Efficient multimodal adaptation via modality-specific transformer substitution

Mingkuan Feng1Zhengqi Wen2( )Jianhua Tao1,2( )

1 Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China

2 Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China

Show Author Information

Abstract

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for visual instruction tuning is computationally expensive and often unnecessary, as the representation requirements for visual and textual tokens diverge significantly in the deeper layers of the network. In this paper, we propose Decoupled Visual Processing (DVP), an efficient training framework that replaces the upper decoder layers of a pretrained LLM with a lightweight, independently trainable single transformer block dedicated exclusively to visual token processing. Specifically, after shared processing through the first half of the decoder layers, visual and textual tokens are split: visual tokens are routed through a newly initialized single transformer block while textual to-kens continue through the original frozen de-coder layers. The two streams are then concatenated before the language modeling head. During training, only the single transformer block is updated, dramatically reducing the number of trainable parameters. Experiments on the LLaVA-1.5 framework demonstrate that DVP achieves competitive performance on MME, POPE, and ChartQA benchmarks while training only a fraction of the total parameters, suggesting that visual representations in MLLMs can be effectively learned through a decoupled, parameter-efficient pathway.

References

【1】
【1】
 
 
iFuture

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Feng M, Wen Z, Tao J. Decoupled visual processing: Efficient multimodal adaptation via modality-specific transformer substitution. iFuture, 2026, https://doi.org/10.26599/IF.2026.9710005

271

Views

6

Downloads

0

Crossref

Received: 26 April 2026
Revised: 04 July 2026
Accepted: 16 July 2026
Available online: 16 July 2026

© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).