This paper addresses the challenges of low efficiency in manual inventory counting and the weak generalization of existing visual methods in retail inventory management by proposing the cooler shelf inventory recognition (CSIR) framework. The framework takes multimodal inputs, encodes multi-angle images using a Vision Transformer, and aligns the resulting features to the latent space of the LLaMA (Large Language Model Meta AI) decoder through linear projection. A decoder-only language model is then constructed to enable end-to-end inventory information generation. This paper combines a domain-specific tokenizer that serializes shelf positions, product types, and inventory levels into discrete tokens to support autoregressive generation, and constructs a real-scenario dataset containing 17,000 samples covering complex conditions such as multi-view, reflective surfaces, and dense arrangements, with multidimensional evaluation metrics. Experimental results show that the proposed method achieves a tolerance-free overall accuracy of 70.17%, representing an approximately 10% improvement over the detection baseline, along with a 5.5-fold increase in inference efficiency. These results effectively reduce labor costs and inventory discrepancies, providing a scalable and reproducible reference solution for automated inventory management.
- Article type
- Year
- Co-author
Open Access
Issue
Open Access
Issue
Aiming at the shortcomings that the existing image-text matching algorithms based on common-sense learning cannot effectively match the intractable negative samples in image-text sample pairs, and the generalization ability of the models is weak and ineffective on large-scale datasets, a novel image-text matching model called Active Mining Sample Pair Semantics image-text matching model is proposed. Firstly, the proposed Adaptive Hierarchical Reinforcement Loss has diversified learning modes, and on top of the traditional triple loss, predictive candidate instances (pairs of intractable sample pairs) are added to aid in training. Its active learning mode enables model to more focus on the intractable negative samples through a penalizing mechanism to enhance the discriminative ability. In addition, the proposed model can also adaptively mine more hidden relevant semantic representations from uncommented items, which greatly improves the performance and generalization ability of model. Finally, experimental results on Flickr30K and MSCOCO datasets show that this proposed method is superior to the existing advanced comparison methods.
京公网安备11010802044758号