DOI: 10.3390/fi18100511 ISSN: 1999-5903

Hybrid Content- and Session-Based Anomaly Detection for Web Server Logs

Abdul Rehman, Mouhammad Nouman, Muhsin Hassanu

Rule-based inspection of web server logs cannot detect attacks it has not already been told to look for, and most anomaly-detection studies validate their methods on a single dataset, which risks overstating how well a method generalises to new traffic. This paper presents a hybrid framework that fuses a content-based autoencoder, which scores individual HTTP requests, with a session-based variational LSTM autoencoder, which scores per-client request sequences. The framework is evaluated across eight datasets spanning three log formats, including the public CSIC 2010 HTTP dataset, under both within-dataset and cross-dataset threshold-transfer protocols. Correcting a session-construction fault that had mixed different clients’ requests into the same window raises the session branch’s AUC on CSIC 2010 from 0.63 to 0.98, evaluated on a held-out split disjoint from the requests the corrected model was trained on; the branch depends on genuine multi-request sessions and produces no score where the traffic does not contain them. Fusion under raw score averaging is not uniformly beneficial: on CSIC 2010, weighting fully toward the session branch outperforms every blended raw-average weight tested, though normalising each branch’s score against its own training-score distribution before averaging largely closes this gap. A threshold calibrated once on normal CSIC 2010 traffic and transferred unchanged to three attack-heavy datasets raises fused F1 six- to seven-fold over a threshold tuned separately on each dataset, because those datasets contain too much attack traffic for self-calibration to represent normal behaviour. The framework also achieves a lower false positive rate than a labelled Random Forest baseline, at comparable or better AUC than unsupervised baselines, and runs fast enough for near-real-time use. Taken together, these results show that session structure and threshold calibration can matter as much as model choice.