Analyzing the effectiveness and applicability of co-training
作者:Kamal Nigam, Rayid Ghani · 年份:2000 · DOI:10.1145/354756.354805 · 被引用次数:1025 · 研究领域:Machine Learning and Algorithms、Optimization and Search Problems、Algorithms and Data Compression
Recently there has been signi cant i n terest in supervised learning algorithms that combine labeled and unlabeled data for text learning tasks.The co-training setting [1] applies to datasets that have a natural separation of their features into two disjoint sets.We demonstrate that when learning from labeled and unlabeled data, algorithms explicitly leveraging a natural independent split of the features outperform algorithms that do not.When a natural split does not exist, co-training algorithms that manufacture a feature split may out-perform algorithms not using a split.These results help explain why co-training algorithms are both discriminative in nature and robust to the assumptions of their embedded classi ers.