Scholay

学术搜索 · AI 审稿 · LaTeX 协作

A Tibetan Ancient Uchen Text Line Dataset for OCR

作者:Zhuome Gongqu, Peng Luo, Dongzhou Jiayang, Cairang Jia, Jiacuo Cizhen, Dongzhu Renqing · 年份:2024 · DOI:10.1109/icicml63543.2024.10958135 · 研究领域:Natural Language Processing Techniques、Handwritten Text Recognition Techniques、Mathematics, Computing, and Information Processing

Characters written by hand that represent distinct writing styles are referred to as handwritten text; these are especially common in ancient texts. In the field of digitizing ancient literature, Tibetan ancient document image recognition is a focus of both domestic and international research. However, technological advancement in this field has been hampered by the absence of a specific recognition dataset for image recognition systems of Tibetan ancient documents. This paper addresses this by proposing a dataset for Uchen style text line recognition in Tibetan ancient documents. This dataset includes basic synthetic images of multi-style Uchen text lines as well as real-world images of multi-format Tibetan Uchen text lines that have been manually annotated. Cross-Validation and several rounds of annotation by various people guarantee the quality of the data. Additionally, three popular text line recognition models are used to assess the dataset and validate its usability.