A Study on Locally Runnable Large Language Models for Bearing Fault Diagnosis
Mehadi Hasan Shawon, Prashant KumarLarge language model (LLM) agents can perform prognostics and health management (PHM) tasks such as bearing fault diagnosis, but most published systems rely on large, paid, cloud-hosted models that a small or medium enterprise (SME) cannot self-host. The diagnostic accuracy that can be achieved on free, offline, commodity hardware is a practical question. Using an execution-based evaluation (Pass@1, Pass all 3, macro F1), this paper benchmarks eight small, publicly accessible, locally runnable LLMs (1B–9B parameters, via Ollama) on vibration-derived features for three-class bearing fault diagnosis (Healthy, Outer race, and Inner race). The proposed work is evaluated on three independent datasets, Paderborn, CWRU, and HUST, across a 0–5-shot ablation and at two decoding temperatures to separate accuracy from reliability. The findings replicate across all three datasets, namely, a model-capability gate that only models near 7B parameters and above clears the majority-class floor; few-shot prompting is non-monotonic; accuracy and reliability are distinct axes; and a single free model, gemma2:9b (5.4 GB), is the most accurate and among the most reliable, with no task-specific training. We further propose ensemble agreement gating, a training-free reliability rule that withholds predictions when several free local models disagree, raising accuracy on the answered subset (e.g., 0.77 → 0.87 on Paderborn). In a head-to-head on identical prompts and data, the free local models match a current frontier model in the settings tested at zero cost and fully offline, and we characterize where such a training-free local approach is and is not appropriate for resource-constrained operators of rotating machinery. On the same features, however, a simple supervised baseline such as logistic regression, and even an untrained physics rule, outperform all eight LLMs, so we present this as a cautionary benchmark: the contribution of the free local approach is training-free deployment and a reliability gating rule, not classification accuracy.