A Multi-Stage Post-Training Framework for Domain-Specific Language Models in Fault Diagnosis
Wei Zhang, Hui Fang, Tongle Wu, Chaoqun Wang, Libo Xu, Jiajun Bu, Yueyao Yu, Qiming ZhongThe rapid advancement of large language models (LLMs) has created new opportunities for intelligent fault diagnosis, particularly in complex industrial systems, such as heating, ventilation, and air conditioning (HVAC) in urban rail transit. Although LLMs have shown strong general reasoning capabilities, adapting them to domain-specific fault diagnosis tasks remains challenging. This is particularly true for textual maintenance records, where sparse and brief entries can only provide limited context for effective knowledge adaptation. To address this challenge, we propose a multi-stage post-training framework based on LLMs. The framework consists of three components: (1) data augmentation via retrieval-augmented generation (RAG) to enrich brief maintenance records with domain knowledge and reasoning traces; (2) supervised fine-tuning (SFT) for domain-specific adaptation; and (3) reinforcement learning with group relative policy optimization (GRPO), using a task-specific reward that separately evaluates root-cause identification and maintenance action recommendation. The framework is applied to a real-world textual HVAC fault dataset derived from Ningbo Rail Transit, covering 33 equipment categories and over 100 fault types. With Qwen3-0.6B as the base model, the proposed method significantly improves diagnostic accuracy and reasoning quality, achieving a 97% increase in model-based diagnostic accuracy (from 0.323 to 0.635) and a 48% improvement in human expert evaluation scores (from 0.509 to 0.754). Moreover, conventional machine learning baselines, such as Support Vector Machine (SVM), Random Forest (RF), and Multi-Layer Perceptron (MLP), achieve accuracies below 0.4 in this task, further highlighting the superiority of the proposed LLM-based framework. These results indicate that the proposed multi-stage post-training framework effectively improves LLM performance on real-world text-based fault diagnosis. It provides a practical and extensible solution for intelligent maintenance decision support in complex electromechanical systems.