The influence of training sample size on the accuracy of deep learning models for the prediction of soil properties with near-infrared spectroscopy data
作者:Wartini Ng, Budiman Minasny, Wanderson de Sousa Mendes, José Alexandre Melo Demattê · 发表于:SOIL · 年份:2020 · DOI:10.5194/soil-6-565-2020 · 被引用次数:216 · 研究领域:Soil Geostatistics and Mapping、Mineral Processing and Grinding、Spectroscopy and Chemometric Analyses
The number of samples used in the calibration data set affects the quality of the generated predictive models using visible, near and shortwave infrared (VIS–NIR–SWIR) spectroscopy for soil attributes. Recently, the convolutional neural network (CNN) has been regarded as a highly accurate model for predicting soil properties on a large database. However, it has not yet been ascertained how large the sample size should be for CNN model to be effective. This paper investigates the effect of the training sample size on the accuracy of deep learning and machine learning models. It aims at providing an estimate of how many calibration samples are needed to improve the model performance of soil properties predictions with CNN as compared to conventional machine learning models. In addition, this paper also looks at a way to interpret the CNN models, which are commonly labelled as a black box. It is hypothesised that the performance of machine learning models will increase with an increasing number of training samples, but it will plateau when it reaches a certain number, while the performance of CNN will keep improving. The performances of two machine learning models (partial least squares regression – PLSR; Cubist) are compared against the CNN model. A VIS–NIR–SWIR spectra library from Brazil, containing 4251 unique sites with averages of two to three samples per depth (a total of 12 044 samples), was divided into calibration (3188 sites) and validation (1063 sites) sets. A subset...