In the world of Six-Generation 6G network, real-time medical streaming plays an important part in providing fast and accurate services. A new method known as cross-modal semantic communication helps in sending meaningful information to different types of data. This paper presents a deep neural network based system that uses 6G cross-modal semantic communication systems to handle medical streaming in real time. The proposed system includes a semantic encoder, a semantic decoder, and a method capable of measuring similarity meaning across various data types. The semantic encoder extracts important features like text, sound, and pictures from different types of medical data, and combines them to create an integrated information. After this, the semantic decoder redesigns the data as per the required format. Using Siamese and pseudo-Siamese networks, the cross-modal semantic similarity evaluation technique compares the meaning of the original and redesigned data across different types of data, resulting in improved encoding and decoding processes. Experimental results show that the proposed framework excels in semantic similarity and real-time performance compared to traditional communication systems. This deep neural networks based encoding and decoding framework enables efficient and effective real-time medical streaming in 6G cross-modal semantic communication systems.
- Article type
- Year
- Co-author
Open Access
Research Article
Issue
Open Access
Research Article
Just Accepted
Facial expression recognition (FER) is a crit-ical component in many fields, such as human-computer interaction and affective computing. However, existing FER methods face several key challenges such as label ambiguity, mixed emotional expressions, and class imbalance. To ad-dress these issues, we propose a novel framework based on Label Distribution Learning (LDL) that captures the complex and compound nature of real-world emotions. Our approach introduces MixFeature, a feature-level augmentation strategy that synthesizes new samples with label distribution by mix-ing single-label ones. Such process is guided by the Robert Plutchik’s Emotion Wheel to ensure semantic consistency in the generated label distribution. By modeling emotions as label distribution, our method provides a more nuanced repre-sentation of blended emotional states. Extensive experiments on widely used datasets demonstrate that our framework sig-nificantly outperforms existing single-label and LDL methods in recognition accuracy and robustness, particularly in han-dling ambiguous and mixed emotions and addressing class imbalance.
Open Access
Research Article
Issue
Automatic depression recognition is essential to depression diagnosis. In this paper, we investigate the problem of depression recognition from facial images, each of which is labeled with one Beck Depression Inventory (BDI-II) score. Because of the ambiguity between one facial image and the depression score, the annotators may not present the accurate score but tend to give those around the ground-truth one. To solve the problem, this paper adopts label distribution to annotate each image, in which each (score) label has a relevance degree. First, we apply the Gaussian distribution to generate the depression score distributions, in which the ground-truth score attains the highest degree, while the neighborhood scores also have degrees to some extent. Thus, each image can contribute to not only its ground-truth score but also neighborhood scores. Second, we generate the depression severity level distributions from the score distributions according to the mapping relationship between BDI-II score and severity level. Finally, we propose a novel method to learn joinT depression scoRE And level distribuTion, termed as TREAT. In the experiments, we compare TREAT with several state-of-the-art methods on three publicly released datasets AVEC 2013, AVEC 2014, and AVEC 2019, and the experimental results justified that TREAT achieves the best performance.
京公网安备11010802044758号