Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Creation of custom recognition profiles for historical documents

作者:Adam Dudczak, Aleksandra Nowak, Tomasz Parkoła · 年份:2014 · DOI:10.1145/2595188.2595209 · 被引用次数:3 · 研究领域:Handwritten Text Recognition Techniques、Mathematics, Computing, and Information Processing、Natural Language Processing Techniques

This paper describes creation of custom OCR profiles for improved recognition of text from historical documents. Presented workflow is based on tools developed by Poznań Supercomputing and Networking Center in the framework of SYNAT project. OCR customization consists of three steps. The first two include preparation of training material in a dedicated web application called Cutouts and processing of the training material using command-line tool called a page-generator. The results of the latter step are passed into Tesseract OCR engine. In the final step Tesseract creates a recognition profile which might be used for an OCR in Virtual Transcription Laboratory (VTL). The described tools allow to crowdsource the most tedious parts of the mentioned process: training material preparation and post OCR correction.