Face recognition (FR) technology has numerous applications in artificial intelligence including biometrics, security, authentication, law enforcement, and surveillance. Deep learning (DL) models, notably convolutional neural networks (CNNs), have shown promising results in the field of FR. However CNNs are easily fooled since they do not encode position and orientation correlations between features. Hinton et al. envisioned Capsule Networks as a more robust design capable of retaining pose information and spatial correlations to recognize objects more like the brain does. Lower-level capsules hold 8-dimensional vectors of attributes like position, hue, texture, and so on, which are routed to higher-level capsules via a new routing by agreement algorithm. This provides capsule networks with viewpoint invariance, which has previously evaded CNNs. This research presents a FR model based on capsule networks that was tested using the LFW dataset, COMSATS face dataset, and own acquired photos using cameras measuring 128 × 128 pixels, 40 × 40 pixels, and 30 × 30 pixels. The trained model outperforms state-of-the-art algorithms, achieving 95.82% test accuracy and performing well on unseen faces that have been blurred or rotated. Additionally, the suggested model outperformed the recently released approaches on the COMSATS face dataset, achieving a high accuracy of 92.47%. Based on the results of this research as well as previous results, capsule networks perform better than deeper CNNs on unobserved altered data because of their special equivariance properties.
- Article type
- Year
Open Access
Article
Issue
Open Access
Research Article
Issue
Speech disorders have a significant impact on quality of life, as they decrease the ability to define one's character, exercise autonomy, and frequently affect relationships and self-esteem, particularly in young children. Dysarthria is a neurological illness that affects motor speech pronunciation. Young children who experience this disorder have no issue with their understanding, but they have a problem expressing their words. They might struggle to communicate precisely and smoothly with their friends and family members due to this illness. A dysarthric child has significant trouble with communication, as this disorder causes poorly pronounced phonemes and poor speech articulation. To address this condition, numerous speech assistive technologies have been developed for consumers with dysarthria, tailored to the level of severity. Currently, deep learning (DL) systems offer potential for objective evaluation, thereby improving diagnostic accuracy. Its goal is to systematically analyze present approaches for detecting dysarthria based on severity levels. In this manuscript, a novel pediatric dysarthria disorder detection framework using residual recurrent neural network and transformer (PD3F-RRNNT) technique is proposed. The PD3F-RRNNT technique aims to develop a real-time recognition method for accurately detecting dysarthria speech disorders in children, supporting early diagnosis and intervention. Initially, the audio processing phase involved various steps, including voice activity detection (VAD), noise removal, pre-emphasis, framing, windowing, and normalization, to transform and extract significant data from audio signals. Furthermore, the PD3F-RRNNT method utilizes the transformer-attention-based U-Net (TransAttUnet) technique for feature extraction. Finally, the residual bidirectional gated recurrent unit (RBG) method is employed to detect and classify speech disorders accurately. The experimental validation of the PD3F-RRNNT model is performed under the dysarthria and non-dysarthria speech dataset. The comparison analysis of the PD3F-RRNNT model revealed a superior accuracy value of 99.50% compared to existing techniques.
Open Access
Article
Issue
ESystems based on EHRs (Electronic health records) have been in use for many years and their amplified realizations have been felt recently. They still have been pioneering collections of massive volumes of health data. Duplicate detections involve discovering records referring to the same practical components, indicating tasks, which are generally dependent on several input parameters that experts yield. Record linkage specifies the issue of finding identical records across various data sources. The similarity existing between two records is characterized based on domain-based similarity functions over different features. De-duplication of one dataset or the linkage of multiple data sets has become a highly significant operation in the data processing stages of different data mining programmes. The objective is to match all the records associated with the same entity. Various measures have been in use for representing the quality and complexity about data linkage algorithms, and many other novel metrics have been introduced. An outline of the problem existing in the measurement of data linkage and de-duplication quality and complexity is presented. This article focuses on the reprocessing of health data that is horizontally divided among data custodians, with the purpose of custodians giving similar features to sets of patients. The first step in this technique is about an automatic selection of training examples with superior quality from the compared record pairs and the second step involves training the reciprocal neuro-fuzzy inference system (RANFIS) classifier. Using the Optimal Threshold classifier, it is presumed that there is information about the original match status for all compared record pairs (i.e., Ant Lion Optimization), and therefore an optimal threshold can be computed based on the respective RANFIS. Febrl, Clinical Decision (CD), and Cork Open Research Archive (CORA) data repository help analyze the proposed method with evaluated benchmarks with current techniques.
京公网安备11010802044758号