AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

Caching Document Identifiers to Speedup Query Processing in Search Servers

Department of Computer Science, National University of Luján, Luján 6700, Argentina
Center for Research and Development in Information and Communications Technologies (CIDETIC), National University of Luján, Luján 6700, Argentina
Show Author Information

Abstract

Modern search systems have become a fundamental tool for accessing the massive amount of information stored in different repositories. These systems use sophisticated techniques to efficiently process a high volume of queries (thus optimizing energy consumption). One of these techniques is caching, which is implemented at different levels of a search architecture. In this work, we propose a novel caching strategy that speeds up dynamic pruning techniques (such as Maxscore) by exploiting the information of the lowest (Min) and highest (Max) document identifiers that appear as the result of a previously submitted query. We name this technique as Min/Max caching. The idea is to use Min/Max information for pruning the terms’ posting lists in the query before executing the ranking algorithm in a document-at-a-time (DAAT) approach. The proposed technique uses low memory resources, returns safe results, and complements other levels of caching (if present). We also combine the approach with different access policies. Extensive experimentation on real-world data shows that the proposed method increases query processing speedup up to 1.23x and can also reduce high-percentile tail latency (up to 2.0x speedup), an essential requirement for operational scenarios. We evaluate different access and eviction cache policies based on different decision criteria. Our findings confirm that considering the cost of the cached items (cost-aware policies) allows more computation savings.

Electronic Supplementary Material

Download File(s)
JCST-2212-13040-Highlights.pdf (132.5 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 870-886

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Tolosa GH, Lavallén P, Ríssola EA. Caching Document Identifiers to Speedup Query Processing in Search Servers. Journal of Computer Science and Technology, 2025, 40(3): 870-886. https://doi.org/10.1007/s11390-024-3040-9

895

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 20 December 2022
Accepted: 30 April 2024
Published: 30 April 2025
© Institute of Computing Technology, Chinese Academy of Sciences 2025