Publications
Article type
Sort:
Open Access Article Issue
NestLipGNN: A Hierarchical Graph Neural Network Framework with Nested Multi-Granularity Learning for Robust Visual Speech Recognition
Computers, Materials & Continua 2026, 88(1)
Published: 08 May 2026
Abstract PDF (2.1 MB) Collect
Downloads:0

Visual speech recognition (VSR) aims to infer spoken content from visual observations of articulatory movements. Despite significant progress, it remains a challenging task in computer vision and speech processing. Its difficulty arises from pronounced speaker-to-speaker variability, the presence of homophenes (phonemes that are visually indistinguishable), changes in illumination, and the intrinsically high-dimensional nature of spatiotemporal lip dynamics. In this work, we propose NestLipGNN, a graph-based framework that integrates Graph Neural Networks (GNNs) with a nested multi-granularity learning strategy for visual speech recognition. We construct dynamic lip graphs from facial landmarks to model both spatial relationships between lip regions and their temporal motion during speech articulation. The proposed nested learning architecture supports hierarchical feature extraction across several levels of linguistic abstraction, spanning phoneme-level articulatory units, viseme-level visual speech categories, and word-level semantic representations. We further introduce a Temporal Graph Attention mechanism (T-GAT) that adaptively reweights the importance of distinct lip regions over time. We also introduce a graph-based contrastive learning objective to improve the discrimination of visually similar speech patterns, directly confronting the challenge of homophene resolution. Experiments on the LRW, LRS2, LRS3, and GRID datasets show that NestLipGNN improves recognition accuracy compared with existing methods, obtaining 92.3% word-level accuracy on LRW and delivering a 2.1% absolute performance gain over prior methods. Comprehensive ablation analyses confirm the contribution of each architectural component.

Open Access Article Issue
Deep Learning-Based Faulty Wood Detection with Area Attention
Computers, Materials & Continua 2025, 85(1): 1495-1514
Published: 29 August 2025
Abstract PDF (1.4 MB) Collect
Downloads:2

Improving consumer satisfaction with the appearance and surface quality of wood-based products requires inspection methods that are both accurate and efficient. The adoption of artificial intelligence (AI) for surface evaluation has emerged as a promising solution. Since the visual appeal of wooden products directly impacts their market value and overall business success, effective quality control is crucial. However, conventional inspection techniques often fail to meet performance requirements due to limited accuracy and slow processing times. To address these shortcomings, the authors propose a real-time deep learning-based system for evaluating surface appearance quality. The method integrates object detection and classification within an area attention framework and leverages R-ELAN for advanced fine-tuning. This architecture supports precise identification and classification of multiple objects, even under ambiguous or visually complex conditions. Furthermore, the model is computationally efficient and well-suited to moderate or domain-specific datasets commonly found in industrial inspection tasks. Experimental validation on the Zenodo dataset shows that the model achieves an average precision (AP) of 60.6%, outperforming the current state-of-the-art YOLOv12 model (55.3%), with a fast inference time of approximately 70 milliseconds. These results underscore the potential of AI-powered methods to enhance surface quality inspection in the wood manufacturing sector.

Total 2