Sort:
Open Access Review Article Issue
Towards depth foundation models: Recent trends in vision-based depth estimation
Computational Visual Media 2026, 12(2): 243-271
Published: 20 March 2026
Abstract PDF (23.2 MB) Collect
Downloads:251

Depth estimation is a fundamental task in 3D computer vision, crucial for applications such as 3D reconstruction, free-viewpoint rendering, robotics, autonomous driving, and AR/VR technologies. Traditional methods relying on hardware sensors like LiDAR are often limited by their high costs, low resolution, and sensitivity to the environment, limiting their applicability to real-world scenarios. Recent advances in vision-based methods offer a promising alternative, yet they face challenges in generalization and stability due to either the low capacity of model architectures or reliance on domain-specific and small-scale datasets. The emergence of scaling laws and foundation models in other domains has inspired the development of “depth foundation models”: deep neural networks trained on large datasets with strong zero-shot generalization capabilities. This paper surveys the evolution of deep learning architectures and paradigms for depth estimation across monocular, stereo, multi-view, and monocular video settings. We explore the potential of these models to address existing challenges and we also provide a comprehensive overview of large-scale datasets that can facilitate their development. By identifying key architectures and training strategies, we aim to highlight the path towards robust depth foundation models, offering insights for future research and applications.

Open Access Research Article Issue
A biophysical-based skin model for heterogeneous volume rendering
Computational Visual Media 2025, 11(2): 289-303
Published: 08 May 2025
Abstract PDF (13.8 MB) Collect
Downloads:273

Realistic human skin rendering has been a long-standing challenge in computer graphics. Recently, biophysical-based skin rendering has received increasing attention, as it provides a more realistic skin-rendering and a more intuitive way to adjust the skin style. In this work, we present a novel heterogeneous biophysical-based volume rendering method for human skin that improves the realism of skin appearance while easily simulating various types of skin effects, including skin diseases, by modifying biological coefficient textures. Specifically, we introduce a two-layer skin representation by mesh deformation that explicitly models the epidermis and dermis with heterogeneous volumetric medium layers containing the corresponding spatially varying melanin and hemoglobin, respectively. Furthermore, to better facilitate skin acquisition, we introduced a learning-based framework that automatically estimates spatially varying biological coefficients from an albedo texture, enabling biophysical-based and intuitive editing, such as tanning, pathological vitiligo, and freckles. We illustrated the effects of multiple skin-editing applications and demonstrated superior quality to the commonly used random walk skin-rendering method, with more convincing skin details regarding subsurface scattering.

Open Access Research Article Issue
Hybrid mesh-neural representation for 3D transparent object reconstruction
Computational Visual Media 2025, 11(1): 123-140
Published: 28 February 2025
Abstract PDF (16.3 MB) Collect
Downloads:195

In this study, we propose a novel method to reconstruct the 3D shapes of transparent objects using images captured by handheld cameras under natural lighting conditions. It combines the advantages of an explicit mesh and multi-layer perceptron (MLP) network as a hybrid representation to simplify the capture settings used in recent studies. After obtaining an initial shape through multi-view silhouettes, we introduced surface-based local MLPs to encode the vertex displacement field (VDF) for reconstructing surface details. The design of local MLPs allowed representation of the VDF in a piecewise manner using two-layer MLP networks to support the optimization algorithm. Defining local MLPs on the surface instead of on the volume also reduced the search space. Such a hybrid representation enabled us to relax the ray–pixel correspondences that represent the light path constraint to our designed ray–cell correspondences, which significantly simplified the implementation of a single-image-based environment-matting algorithm. We evaluated our representation and reconstruction algorithm on several transparent objects based on ground truth models. The experimental results show that our method produces high-quality reconstructions that are superior to those of state-of-the-art methods using a simplified data-acquisition setup.

Open Access Research Article Issue
A causal convolutional neural network for multi-subject motion modeling and generation
Computational Visual Media 2024, 10(1): 45-59
Published: 30 November 2023
Abstract PDF (4.7 MB) Collect
Downloads:82

Inspired by the success of WaveNet in multi-subject speech synthesis, we propose a novel neural network based on causal convolutions for multi-subject motion modeling and generation. The network can capture the intrinsic characteristics of the motion of different subjects, such as the influence of skeleton scale variation on motion style. Moreover, after fine-tuning the network using a small motion dataset for a novel skeleton that is not included in the training dataset, it is able to synthesize high-quality motions with a personalized style for the novel skeleton. The experimental results demonstrate that our network can model the intrinsic characteristics of motions well and can be applied to various motion modeling and synthesis tasks.

Total 4