Predicting second-language learning performance through deep linguistic embeddings and behavioral indicators: A RoBERTa–IFPA–decision tree framework
Yingnan Wang, Xinyu Liang, Zhen Liu, Zhongxing ZhaoThis study proposes a hybrid data-driven framework that integrates deep linguistic, behavioral, and psychological features to predict second-language learning performance with high accuracy and interpretability. Using the Education First Cambridge Open Language Dataset (EFCAMDAT), this study employs a Robustly Optimized BERT Approach (RoBERTa), a transformer-based language model, to extract context-sensitive lexical, syntactic, and semantic representations from learner essays. Behavioral and psychological indicators, including engagement, motivation, and self-control, are combined with deep linguistic embeddings to construct a multimodal dataset. The improved flower pollination algorithm (IFPA) is used to identify the most informative features, while a decision tree performs the final predictive modeling because of its transparent structure and ability to handle heterogeneous feature types. Fivefold cross-validation on the EFCAMDAT corpus showed that the proposed RoBERTa + IFPA + decision tree framework outperformed six classical machine-learning baselines and three neural architectures, achieving an overall accuracy of 92.36% and an F1-score of 0.91. Statistical significance testing (p < 0.05) supported the observed performance gains, while the interpretability analysis identified lexical complexity, engagement level, and motivational stability as key predictors. The analysis of errors showed that misclassifications were mostly found at intermediate proficiency levels and reflected greater overlap among learner profiles rather than model failure. Overall, this framework offers an evidence-based connection between artificial intelligence and second-language acquisition, with implications for adaptive language assessment and individualized learning support.