Graph-Aware Reinforcement Learning for Adaptive APT Threat Hunting in Dynamic Enterprise Networks
Bahar Memarpour, Kimia Memarpour, Kimia Shirini, Sina Samadi Gharehveran, Siamak Pedrammehr, Hussain Mohammed Dipu KabirAdvanced Persistent Threats (APTs) are particularly challenging for enterprise intrusion detection because they are time-evolving, distributed, and difficult to detect under changing network conditions. Conventional machine-learning-based intrusion detection systems (IDSs) often rely on node-level features and may be vulnerable to concept drift, topological changes, and severe class imbalance, potentially achieving high overall accuracy while overlooking rare but critical malicious activities. This study proposes a Graph-Aware Deep Reinforcement Learning (GADRL) framework that integrates Graph Neural Networks (GNNs), Double Deep Q-Networks (DDQNs), and Prioritized Experience Replay (PER) for adaptive APT detection. Network traffic is modeled as a temporal graph to learn spatial–temporal relational representations of communication patterns and neighborhood interactions. These representations are then provided to a DDQN-based threat-hunting agent, while PER prioritizes high-severity and rare attack experiences during training. The framework is evaluated using five temporally ordered snapshots of enterprise network traffic to examine its performance under evolving network conditions. On an unseen future snapshot, GADRL achieves 100.00% recall, outperforming both a Random Forest baseline and a graph-agnostic reinforcement-learning ablation. This improvement in recall is accompanied by increased false-positive alerts, indicating a deliberate trade-off that prioritizes minimizing missed intrusions over reducing false alarms. Overall, the results demonstrate that combining topology-aware graph representation learning with priority-aware reinforcement learning can improve adaptive detection of evolving APTs under temporally changing network conditions.