DOI: 10.1002/agj2.70511 ISSN: 0002-1962

Select or discard? A machine learning approach in common bean breeding

Paulo Henrique Cerutti, Luan Tiago dos Santos Carbinari, Jefferson Luís Meirelles Coimbra, Mauro Bitencourt de Souza, Carlos Zacarias Joaquim

Abstract

The selection of superior genotypes in common bean ( Phaseolus vulgaris L.) breeding is challenging due to the broad genetic variability in segregating populations. In this context, machine learning (ML) has emerged as a promising tool to support decision‐making in plant breeding. This study aimed to evaluate the potential of ML to classify common bean genotypes according to their selection value. A total of 5136 plants, including inbred lines, cultivars, and segregating generations (F 2 to F 10 ), were field‐evaluated. Traits analyzed included plant height, stem diameter (SD), number of pods (NP), first pod insertion height, and grain weight per plant (GWPP). The Random Forest classifier algorithm was applied to the dataset, split into training (70%) and testing (30%) sets. Genotypes were categorized based on GWPP quartiles as follows: poor (<4.99 g), fair (5.00–8.50 g), good (8.51–13.93 g), and excellent (>13.93 g). The model achieved an accuracy of 0.68, with high precision for “excellent” (0.78) and “poor” (0.77) classes. Strong correlations were observed between GWPP and NP ( r  = 0.93) and SD ( r  = 0.64), supporting indirect selection strategies. Top‐ranked individuals, such as the BAF07 accession and BRS Embaixador cultivar, were identified as potential parents. These findings reinforce the usefulness of supervised ML models in common bean breeding, particularly for selecting genotypes with extreme performance, and highlight their potential to enhance efficiency, precision, and speed in breeding programs.

More from our Archive