An adaptive firewall framework using deep reinforcement learning for threat-aware cyber threat detection and mitigation policy learning

dc.contributor.authorAlam, Mohammad Zahangir
dc.contributor.authorSultana, Sharmin
dc.contributor.organizationfi=tietotekniikan laitos|en=Department of Computing|
dc.contributor.organization-code1.2.246.10.2458963.20.85312822902
dc.converis.publication-id526817221
dc.converis.urlhttps://research.utu.fi/converis/portal/Publication/526817221
dc.date.accessioned2026-07-31T20:11:13Z
dc.description.abstract<p>Rapid digital infrastructure growth has enabled sophisticated multi-vector cyberattacks that bypass traditional signature-based intrusion detection and static firewall configurations. Current defenses rely on fixed rule-based engines, signature-dependent classifiers, or threshold-based anomaly detectors, all of which struggle with con current threat patterns, zero-day attacks, and high-dimensional decision spaces. Machine and deep learning approaches for intrusion detection also face limitations, including slow adaptation to novel threats, high false positive rates, and limited support for dynamic firewall policy adjustment. To address these limitations, we propose an Adaptive Firewall Framework using Deep Reinforcement Learning (DRL) for threat-aware cyber threat detection and mitigation policy learning. The framework incorporates a policy-optimization DRL agent with a novel Threat-Aware Prioritized Experience Replay (TA-PER) mechanism, which prioritizes threat expe riences according to operational severity, CVSS criticality, and buffer rarity. Through state-space modeling, threat-sensitive reward shaping, and incremental threat scoring, the DRL-TA-PER agent learns firewall response policies, including traffic blocking, isolation, adaptive rerouting, throttling, and benign traffic allowance. The framework is evaluated in an offline sequential setting using public benchmark datasets, where network-flow records are processed as ordered decision steps. Evaluation on CIC-IDS-2018, UNSW-NB15, CIC-DDoS2019, and CIC-IoT-2023 shows that the proposed framework outperforms K-Nearest Neighbor, Random Forest, Lo gistic Regression, Support Vector Machine, Convolutional Neural Network, Long Short-Term Memory, standard DRL with uniform experience replay, Snort, and Suricata. The model achieved 96.8% accuracy, 97.2% recall, and an AUC-ROC of 0.981 across five trials, indicating that threat-aware replay prioritization can improve DRL-based firewall policy learning under offline benchmark evaluation conditions.<br></p>
dc.identifier.eissn2590-0056
dc.identifier.urihttps://www.utupub.fi/handle/11111/62812
dc.identifier.urlhttps://doi.org/10.1016/j.array.2026.101081
dc.identifier.urnURN:NBN:fi-fe20260730113586
dc.language.isoen
dc.okm.affiliatedauthorAlam, Mohammad
dc.okm.discipline113 Computer and information sciencesen_GB
dc.okm.discipline113 Tietojenkäsittely ja informaatiotieteetfi_FI
dc.okm.internationalcopublicationinternational co-publication
dc.okm.internationalityInternational publication
dc.okm.typeA1 ScientificArticle
dc.publisherElsevier
dc.publisher.countryUnited Statesen_GB
dc.publisher.countryYhdysvallat (USA)fi_FI
dc.publisher.country-codeUS
dc.relation.articlenumber101081
dc.relation.doi10.1016/j.array.2026.101081
dc.relation.ispartofjournalArray
dc.relation.volume31
dc.titleAn adaptive firewall framework using deep reinforcement learning for threat-aware cyber threat detection and mitigation policy learning
dc.year.issued2026

Tiedostot

Näytetään 1 - 1 / 1
Ladataan...
Name:
1-s2.0-S2590005626004042-main.pdf
Size:
5.05 MB
Format:
Adobe Portable Document Format