Qualitative data cleaning
作者:Xu Chu, Ihab F. Ilyas · 发表于:Proceedings of the VLDB Endowment · 年份:2016 · DOI:10.14778/3007263.3007320 · 被引用次数:40 · 研究领域:Data Quality and Management、Privacy-Preserving Technologies in Data、Data Mining Algorithms and Applications
Data quality is one of the most important problems in data management, since dirty data often leads to inaccurate data analytics results and wrong business decisions. Data cleaning exercise often consist of two phases: error detection and error repairing. Error detection techniques can either be quantitative or qualitative; and error repairing is performed by applying data transformation scripts or by involving human experts, and sometimes both. In this tutorial, we discuss the main facets and directions in designing qualitative data cleaning techniques. We present a taxonomy of current qualitative error detection techniques, as well as a taxonomy of current data repairing techniques. We will also discuss proposals for tackling the challenges for cleaning "big data" in terms of scale and distribution.