Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Robust and accurate multimodal information extraction from tobacco box labels using MLDP

作者:Songling Huang, Shuai Zhang, Minghua Han, Jie Yang, Yu Bai · 发表于:Scientific Reports · 年份:2026 · DOI:10.1038/s41598-026-60482-1 · 研究领域:Text and Document Classification Technologies、Handwritten Text Recognition Techniques、Image Retrieval and Classification Techniques

With the advancement of smart re-baking factories and the accelerating digital transformation of agriculture, tobacco re-baking is facing an urgent need to shift from manual inspection to intelligent automated systems for quality traceability and information management. However, traditional cigarette box label information extraction methods still suffer from limited efficiency, insufficient accuracy, and poor robustness under real industrial conditions. To address these challenges, we propose MLDP (Multimodal Learning for Document Processing), a multimodal framework tailored for robust information extraction from tobacco box labels.MLDP integrates textual and visual information in a unified framework. For the text modality, multiple DeBERTa-based variants, including BiLSTM-DeBERTa, Multi-Sample Dropout-DeBERTa, Distil-DeBERTa, Two-Stage Dropout-DeBERTa, and LongMax-Dropout-DeBERTa, are combined through an Optuna-driven weighted ensemble strategy to balance model diversity and computational efficiency. For the image modality, PaddleOCR is adopted for text detection and recognition, and a Levenshtein distance-based correction mechanism is introduced to reduce OCR errors, especially in numeric fields. To further enhance cross-modal complementarity, MLDP employs a dynamic weighted fusion strategy together with a decision matrix for modality-level error correction, improving extraction robustness under noisy and complex industrial conditions.Experiments on a real-world multimodal ...