Simplifying AI-Based AHU Forecasting for Sustainable Building Operation: Do Seasonal and Engineered Features Improve Prediction Accuracy?
Dalia Mohammed Talat Ebrahim Ali, Violeta Motuzienė, Rasa Džiugaitė-TumėnienėFeature engineering has become a common step in AI-based HVAC forecasting, often involving variables calculated from raw building management system (BMS) measurements, such as temperature differences, setpoint tracking deviations, airflow balance indicators, rolling statistics, and temporal or seasonal descriptors. Accurate short-term forecasting can provide a baseline of expected operation for anomaly and fault detection and can support control optimization and operator decision making. However, real-world deployment is complicated due to differences in BMS sensor availability and data quality, as well as the preprocessing and maintenance burden associated with complex feature sets. The actual contribution of these features to the performance of AI forecasting remains underexplored, particularly for short-term prediction of air handling unit (AHU) operation. This study evaluates the impact of features on short-term AHU forecasting using three deep learning (DL) architectures: Temporal Convolutional Networks (TCNs), Long Short-Term Memory (LSTM) networks, and a hybrid CNN–LSTM model. An actual operational AHU dataset from a BMS was used to predict key operational variables, including supply and extract air temperatures, supply and extract fan operating signals, and supply air temperature setpoint-tracking error. Fan signal balance was additionally evaluated as a derived indicator calculated from the two predicted fan signals. Four input configurations were evaluated: (i) full (74 inputs), containing raw BMS measurements, short-cycle temporal variables, engineered and dynamic features, and annual-calendar information; (ii) no annual calendar (68 inputs), identical to full but excluding annual-calendar variables; (iii) raw + short-cycle temporal (20 inputs); and (iv) raw-only (12 inputs). The models used a 60-min input history to forecast the following 30-min at one-minute resolution. Persistence and Ridge models were included as reference baselines. All models were trained and tested on identical data splits and forecasting horizons to ensure a fair comparison. Each DL experiment was repeated across five independent runs, and performance was evaluated using MAE, RMSE, and R2. The TCN showed the strongest overall DL performance. Raw-only achieved the highest mean R2 in 11 of 15 architecture–target comparisons using just 12 inputs. The best mean DL R2 ranged from 0.916 for the fan signals to 0.993 for extract air temperature. Annual-calendar features improved the TCN results but provided no consistent benefit for the LSTM or CNN–LSTM. Ridge slightly outperformed the best DL configurations for temperature-related targets, reflecting the strong short-term continuity of these signals. These findings show that recent raw BMS measurements contain most of the information needed for accurate 30-min AHU forecasting, while explicit seasonal and engineered features provide limited additional value. The resulting simpler models may in the future be used as forecasting components in predictive control and fault detection systems. However, their control and energy-saving benefits must be tested separately.