Scholay

学术搜索 · AI 审稿 · LaTeX 协作

Initial learning of document structure

作者:Andreas R. Dengel · 年份:2002 · DOI:10.1109/icdar.1993.395776 · 被引用次数:39 · 研究领域:Rough Sets and Fuzzy Logic、Text and Document Classification Technologies、Data Mining Algorithms and Applications

Proposes an approach for automatically generating a decision tree which is applied as a model for the logical labeling of business letters. Instead of top-down determination of the discriminating attributes, the system inspects a finite set of document instances that are presented to a learner in a bottom-up position. The learner itself figures out local similarities, rates them with respect to the overall structure, and determines the best structural match of two instances (neighborhood). The entire decision tree is grown step by step deducing subtrees by forming generalizations from a neighborhood. Consequently, heuristics are learned for structurally discriminating documents during subsequent classification.>