DOI: 10.1002/csc2.70367 ISSN: 0011-183X

Can artificial intelligence effectively mimic a maize breeder's selection decisions?

Isaías Ariza‐Hernández, Rex Bernardo

Abstract

Genetic improvement depends on multi‐trait selection decisions whereby breeders advance a subset of candidates from a larger population. Our objective was to assess whether artificial intelligence via machine learning can effectively mimic selection decisions in maize ( Zea mays L.). Seven experienced maize breeders independently made selection decisions in seven biparental populations from within‐ and across‐environment phenotypic data. Selection decisions were mimicked via two linear models and four nonlinear models. In across‐population prediction, the probability that a random breeder‐selected candidate was ranked ahead of a random non‐selected candidate was about 0.90 for the best prediction models. This high probability, which did not depend on the proportion selected by the breeder, was supported by about 0.70–0.80 of the breeder's decisions being exactly replicated by machine learning. The models performed better in identifying discarded candidates than selected candidates. Logistic regression and gradient boosting performed the best, whereas k ‐nearest neighbors performed the poorest. Hyperparameter tuning did not improve model performance. Including check hybrids in the prediction models was not helpful. As expected, yield was the most important trait, with selection being weaker but generally consistent for lower moisture, higher test weight, shorter plants, and lower ears. Overall, our findings indicated that breeder selection decisions in maize have a learnable structure that is consistent enough to be modeled effectively in a decision‐support system. We speculate that the effectiveness of machine learning to mimic selection decisions in maize will only increase as more data become available to train the models.