In the past years, multimodal large language models (MLLMs) have demonstrated remarkable performance in tasks such as visual question answering and visual understanding and reasoning. However, the extensive model size and high training and inference costs have hindered the widespread application of MLLMs in academia and industry. Thus, studying efficient and lightweight MLLMs has enormous potential, especially in edge computing scenarios. In this survey, we provide a comprehensive and systematic review of the current state of efficient MLLMs. Specifically, this survey summarizes the timeline of representative efficient MLLMs, the current state of research in structures and strategies, and the applications. Finally, the limitations of current efficient MLLM research and promising future directions are discussed.
- Article type
- Year
- Co-author
Open Access
Review
Issue
Open Access
Research Article
Issue
A significant performance boost has been achieved in point cloud semantic segmentation by utilization of the encoder-decoder architecture and novel convolution operations for point clouds. However, co-occurrence relationships within a local region which can directly influence segmentation results are usually ignored by current works. In this paper, we propose a neighborhood co-occurrence matrix (NCM) to model local co-occurrence relationships in a point cloud. Wegenerate target NCM and prediction NCM fromsemantic labels and a prediction map respectively. Then,Kullback-Leibler (KL) divergence is used to maximize the similarity between the target and prediction NCMs to learn the co-occurrence relationship. Moreover, for large scenes where the NCMs for a sampled point cloud and the whole scene differ greatly, we introduce a reverse form of KL divergence which can better handle the difference to supervise the prediction NCMs. We integrate our method into an existing backbone and conduct comprehensive experiments on three datasets: Semantic3D for outdoor space segmentation, and S3DIS and ScanNet v2 for indoor scene segmentation. Results indicate that our method can significantly improve upon the backbone and outperform many leading competitors.
Open Access
Research Article
Issue
The technique of facial attribute manipulation has found increasing application, but it remains challenging to restrict editing of attributes so that a face’s unique details are preserved. In this paper, we introduce our method, which we call amask-adversarialautoencoder (M-AAE). It combines a variational autoencoder (VAE) and a generative adversarial network (GAN) for photorealistic image generation. We use partial dilated layers to modify a few pixels in the feature maps of an encoder, changing the attribute strength continuously without hindering global information. Our training objectives for the VAE and GAN are reinforced by supervision of face recognition loss and cycle consistency loss, to faithfully preserve facial details. Moreover, we generate facial masks to enforce background consistency, which allows our training to focus on the foreground face rather than the background. Experimental results demonstrate that our method can generate high-quality images with varying attributes, and outperforms existing methods in detail preservation.
Open Access
Research Article
Issue
Recent years have witnessed the emergence of image decomposition techniques which effectively separate an image into a piecewise smooth base layer and several residual detail layers. However, the intricacy of detail patterns in some cases may result in side-effects including remnant textures, wrongly-smoothed edges, and distorted appearance. We introduce a new way to construct an edge-preserving image decomposition with properties of detail smoothing, edge retention, and shape fitting. Our method has three main steps: suppressing high-contrast details via a windowed variation similarity measure, detecting salient edges to produce an edge-guided image, and fitting the original shape using a weighted least squares framework. Experimental results indicate that the proposed approach can appropriately smooth non-edge regions even when textures and structures are similar in scale. The effectiveness of our approach is demonstrated in the contexts of detail manipulation, HDR tone mapping, and image abstraction.
京公网安备11010802044758号