DOI: 10.1021/acscatal.6c03456 ISSN: 2155-5435

Data as the Backbone of Artificial Intelligence: Insights from Heterogeneous Catalysis

Dongjae Shin, Ruchika Mahajan, Zan Lian, Jake Heinlein, Matteo Cargnello, Christopher J. Tassone, Kirsten T. Winther

Abstract

Machine learning (ML) and artificial intelligence (AI) are changing how catalysis science is performed, with the prospect of accelerating the discovery of catalysis knowledge. Despite the substantial efforts and resources invested in the development and application of AI/ML approaches, their performance still exhibits clear limitations due to the lack of usable data. Although decades of work have produced an abundance of catalysis data, much of it was generated without sufficient understanding of AI/ML or community-wide consensus on terminology and reporting standards. As a result, not all of this data is suitable for training models. It is therefore essential to define what constitutes AI-ready data and understand the efforts required to obtain it, in order to prepare for the era of AI-driven catalysis. In this perspective, we define the essential properties of AI-ready data that enable the construction of robust catalysis AI/ML models. Then, we discuss current practices in computational and experimental catalysis through the lens of data readiness for AI/ML applications. Lastly, we discuss community-level efforts needed to establish large-scale datasets ready for AI-driven catalysis research.

More from our Archive