Construction and application of materials database under data-driven research paradigm
作者:Junjie Feng, Kun Zhou, Muchen Li, Xinjiang Wang, Lijun Zhang · 发表于:Chinese Science Bulletin (Chinese Version) · 年份:2024 · DOI:10.1360/tb-2024-0946 · 被引用次数:3 · 研究领域:Machine Learning in Materials Science、Nuclear Materials and Properties
In recent years, with the rapid development of big data and artificial intelligence technologies, the data-driven paradigm for materials research has increasingly demonstrated significant advantages in areas such as mining structure-property relationships and screening and designing novel materials. It relies on vast amounts of data, employing high-throughput screening or artificial intelligence aided methods to shorten the materials research and development cycle while reducing scientific research costs. In this paradigm, data serves not only as the foundation but also plays a crucial role in expanding the search space for high-throughput screening and enhancing the performance of artificial intelligence models. Consequently, establishing high-quality materials databases has become an essential component of data-driven materials research. Since the launch of the Materials Genome Initiative, both theoretical and experimental databases have seen continuous growth in number and scale, characterized by automatic data generation, diversified data access, and standardized data representation. These databases strive to meet the demands for high concurrency and scalability in the context of big data while improving data reliability and usability. This review focuses on materials database construction under data-driven research paradigm, outlining the key of database construction from four aspects: Data generation, data preprocessing, data storage, and data access. In the data genera...