An Energy-Efficient Hierarchical Federated Learning Protocol with Downward Feature Transfer: A Simulation-Based Feasibility Study for Low-Power Edge Nodes
Luciano Radrigan, Anibal S. Morales, Pedro Toledo, Sebastian E. Godoy, Ernesto Guerra-VallejosElectric motors consume over 45% of global electricity and are a primary source of unplanned industrial downtime. Real-time fault detection at scale faces severe constraints, including distributed topologies, intermittent connectivity, and strict energy budgets on battery-powered edge nodes. Existing hierarchical federated learning approaches address resource disparities across tiers but lack downward feature transfer. This prevents resource-constrained edge sensors from utilizing cloud-learned representations when local fault data is sparse. This paper proposes a hierarchical federated cyber-physical architecture featuring three online cross-layer transfer mechanisms: warm-start Convolutional Neural Network (CNN) weight extraction, Long Short-Term Memory (LSTM) embedding alignment, and adaptive teacher–student distillation. This work is best characterized as a hierarchical federated-learning protocol and edge-hardware feasibility study: it validates the communication protocol, cross-layer transfer mechanisms, and sensor-tier hardware budget end-to-end, using the Gym-Electric-Motor (GEM) simulator as a controlled, reproducible, and openly available substitute for physically instrumented motor faults, rather than as a validated physical motor-fault diagnosis system. The framework is evaluated on a ten-node low-power System-on-Chip (SoC) microcontroller, low-power single-board computer, and cloud computing platform test bench, using GEM-simulated operating trajectories as a controlled, reproducible proxy for non-IID motor fault conditions in five-class motor fault detection. Under this test bench, the framework achieves an over 3-fold convergence speedup and reduces the sensor–cloud accuracy gap by nearly 74% (from 10.1% down to 2.7%) at a 200× lower compute budget. It improves minority-class diagnostic reliability, with F1 scores improving by 45% on average, while achieving over 95% sensor accuracy on this test bench, and the proposed mechanism also limits Macro-F1 degradation under injected sensor noise (SNR = 10 dB) to 8.0%, versus up to 22.5% for independent per-tier training. From an embedded electronics implementation perspective, the system operates under a 3.3 ms latency and consumes only 2.1 mJ per inference on the low-power SoC hardware—outperforming prior edge PdM deployments that report 4–6 mJ per inference—indicating a feasible architecture for energy-constrained edge intelligence, pending validation on physically measured fault data.