Ontology alignment is a crucial task in the field of knowledge fusion. It provides an essential basis for constructing large-scale high-quality knowledge graphs. However, the previous ontology alignment works face the three problems: lack of implicit semantic within reference mapping (i.e., labeled set), falsely high similarity between class embeddings, and imbalanced training data. To solve these problems, this paper proposes an active learning ontology alignment approach with attribute self-adaptation mechanism and heterogeneous feature fusion (ASHF). The active learning framework and attribute self-adaptation mechanism aim to avoid false positives aligned classes by reconstructing the reference mapping, and to obtain stable performance for sparse ontologies. The heterogeneous feature fusion strategy calculates the similarity between classes by selecting the more distinguished semantic features of classes. Experimental results on eight public datasets show that the proposed model outperforms the state-of-art methods, demonstrating the effectiveness and superiority of the proposed approach in this paper.
- Article type
- Year
- Co-author
Open Access
Research Article
Issue
Open Access
Issue
Multimodal Sentiment Classification (MSC) uses multimodal data, such as images and texts, to identify the users’ sentiment polarities from the information posted by users on the Internet. MSC has attracted considerable attention because of its wide applications in social computing and opinion mining. However, improper correlation strategies can cause erroneous fusion as the texts and the images that are unrelated to each other may integrate. Moreover, simply concatenating them modal by modal, even with true correlation, cannot fully capture the features within and between modals. To solve these problems, this paper proposes a Cross-Modal Complementary Network (CMCN) with hierarchical fusion for MSC. The CMCN is designed as a hierarchical structure with three key modules, namely, the feature extraction module to extract features from texts and images, the feature attention module to learn both text and image attention features generated by an image-text correlation generator, and the cross-modal hierarchical fusion module to fuse features within and between modals. Such a CMCN provides a hierarchical fusion framework that can fully integrate different modal features and helps reduce the risk of integrating unrelated modal features. Extensive experimental results on three public datasets show that the proposed approach significantly outperforms the state-of-the-art methods.
Open Access
Issue
Density-based approaches in content extraction, whose task is to extract contents from Web pages, are commonly used to obtain page contents that are critical to many Web mining applications. However, traditional density-based approaches cannot effectively manage pages that contain short contents and long noises. To overcome this problem, in this paper, we propose a content extraction approach for obtaining content from news pages that combines a segmentation-like approach and a density-based approach. A tool called BlockExtractor was developed based on this approach. BlockExtractor identifies contents in three steps. First, it looks for all Block-Level Elements (BLE) & Inline Elements (IE) blocks, which are designed to roughly segment pages into blocks. Second, it computes the densities of each BLE&IE block and its element to eliminate noises. Third, it removes all redundant BLE&IE blocks that have emerged in other pages from the same site. Compared with three other density-based approaches, our approach shows significant advantages in both precision and recall.
京公网安备11010802044758号