A Study of text mining for Chinese coal mine safety based on BILSTM-CRR-LDA
作者:Yinuo Xu, Yuntao Liang, Baoshan Jia, Si Xiong, Qiyao Xiao · 发表于:Research Square · 年份:2022 · DOI:10.21203/rs.3.rs-2179141/v1 · 被引用次数:2 · 研究领域:Evaluation and Optimization Models、Risk and Safety Analysis
Abstract This is because text mining can reduce the time invested in pre-contextual research on coal mine safety in China. Therefore, this paper uses a combined BiLSTM-CRF model to handle the recognition of data. The optimal number of topics for LDA is determined by combining the perplexity and the topic variance, and this is done by manually annotating the coal mine safety corpus with the BIO annotation method. The BiLSTM-CRF combined model was then used to initially split the data into words and manually correct them. The combined BiLSTM-CRF model was repeatedly trained until the combined model met the word separation requirements. Finally, the LDA model was used to mine the data for topics. The results show that the combined BiLSTM-CRF model is able to recognise long proper nouns and achieve recognition in both English and Chinese. The combination of confusion and topic variance to determine the optimal number of topics for LDA can provide a reference for manual determination of the optimal number of topics for LDA.