Publications
Sort:
Open Access Research Article Issue
Robust multi-person vital-sign sensing in indoor environments using FMCW MIMO radar
Electronic Research Archive 2026, 34(3): 1477-1505
Published: 12 February 2026
Abstract PDF (1.7 MB) Collect
Downloads:25

Contactless monitoring of respiration and heart rate in shared indoor spaces—such as hospital wards, eldercare facilities, and sleep labs—demands solutions that are unobtrusive, privacy-preserving, and robust to static reflections, multipath, and inter-subject interference. This paper presents a complete millimetre-wave frequency-modulated continuous-wave (FMCW) technology - multiple-input multiple-output (MIMO) radar framework for multi-subject vital-sign sensing with three key contributions: (i) clutter-robust target discovery to stabilize detections in realistic indoor scenes; (ii) high-resolution spatial separation to suppress inter-subject leakage; and (iii) hybrid time–frequency decomposition to improve heartbeat isolation by mitigating respiration harmonics and noise. Experiments on single-, two-, and three-subject datasets achieve Mean Absolute Error/ Root Mean Squared Error (MAE/RMSE) of 2.81/4.13 Beats Per Minute (BPM) for respiration and 2.43/3.61 BPM for heart rate. Compared with filtering-, Ensemble Empirical Mode Decomposition (EEMD), Compressive Sensing - Orthogonal Matching Pursuit (CS-OMP), and mmRH-Millimetre-wave radar heart-rate (mmRH) based baselines, the proposed framework yields substantially lower heart-rate errors while maintaining reliable multi-subject operation in shared indoor environments.

Open Access Research Article Issue
Lightweight high-performance pose recognition network: HR-LiteNet
Electronic Research Archive 2024, 32(2): 1145-1159
Published: 29 January 2024
Abstract PDF (1.1 MB) Collect
Downloads:10

To address the limited resources of mobile devices and embedded platforms, we propose a lightweight pose recognition network named HR-LiteNet. Built upon a high-resolution architecture, the network incorporates depthwise separable convolutions, Ghost modules, and the Convolutional Block Attention Module to construct L_block and L_basic modules, aiming to reduce network parameters and computational complexity while maintaining high accuracy. Experimental results demonstrate that on the MPII validation dataset, HR-LiteNet achieves an accuracy of 83.643% while reducing the parameter count by approximately 26.58 M and lowering computational complexity by 8.04 GFLOPs compared to the HRNet network. Moreover, HR-LiteNet outperforms other lightweight models in terms of parameter count and computational requirements while maintaining high accuracy. This design provides a novel solution for pose recognition in resource-constrained environments, striking a balance between accuracy and lightweight demands.

Open Access Research Article Issue
Audio2DiffuGesture: Generating a diverse co-speech gesture based on a diffusion model
Electronic Research Archive 2024, 32(9): 5392-5408
Published: 26 September 2024
Abstract PDF (26.6 MB) Collect
Downloads:93

People use a combination of language and gestures to convey intentions, making the generation of natural co-speech gestures a challenging task. In audio-driven gesture generation, relying solely on features extracted from raw audio waveforms limits the model's ability to fully learn the joint distribution between audio and gestures. To address this limitation, we integrated key features from both raw audio waveforms and Mel-spectrograms. Specifically, we employed cascaded 1D convolutions to extract features from the audio waveform and a two-stage attention mechanism to capture features from the Mel-spectrogram. The fused features were then input into a Transformer with cross-dimension attention for sequence modeling, which mitigated accumulated non-autoregressive errors and reduced redundant information. We developed a diffusion model-based Audio to Diffusion Gesture (A2DG) generation pipeline capable of producing high-quality and diverse gestures. Our method demonstrated superior performance in extensive experiments compared to established baselines. Regarding the TED Gesture and TED Expressive datasets, the Fréchet Gesture Distance (FGD) performance improved by 16.8 and 56%, respectively. Additionally, a user study validated that the co-speech gestures generated by our method are more vivid and realistic.

Total 3