Prediction of software bug severity using transformers and sentiment signals
Rehab Duwairi, Shatha Al-Mallak, Husam SuleimanAbstract
Software bugs adversely affect software quality, security, and release schedules. Accurately predicting bug severity can improve triage and resource allocation. This study investigates transformer-based bug-severity prediction using approximately 20,000 Bugzilla reports from the Eclipse project. To strengthen the signal in short bug reports, the text was augmented with sentiment labels generated using Senti4SD The sentiment-augmented representation was adopted based on exploratory pilot observations suggesting that sentiment labels provide useful cues for short bug summaries. We also evaluate both full and partial fine-tuning of BERT for binary (Severe vs. Non-severe) and multiclass (Blocker, Critical, Major, Minor, Trivial) classification. We compare the proposed models against five classical baselines, namely Random Forest, Naïve Bayes, Multinomial Naïve Bayes, Extra Trees, and k-Nearest Neighbors. For multiclass classification, fully fine-tuned BERT achieved 79.0 % accuracy and a 78.8 % F1-score. By contrast, partially fine-tuned BERT achieved 78.7 % accuracy and a 78.0 % F1-score. Both models outperformed the strongest baseline, Extra Trees, which achieved a 77.9 % F1-score. In binary classification, partial fine-tuning slightly outperformed full fine-tuning, reaching an 88 % F1-score versus 87 % for full fine-tuning. The results also show that sentiment augmentation supports prediction for short bug summaries and that partial fine-tuning provides performance close to full fine-tuning with lower training burden. These findings support the use of transformer-based models as practical decision-support tools for software bug triage.