{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,9]],"date-time":"2025-09-09T20:54:37Z","timestamp":1757451277212,"version":"3.41.0"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2020,4,3]],"date-time":"2020-04-03T00:00:00Z","timestamp":1585872000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Army Research Office under MURI","award":["W911NF-13-1-0421"],"award-info":[{"award-number":["W911NF-13-1-0421"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2020,6,30]]},"abstract":"<jats:p>\n            Cyber-security is an important societal concern. Cyber-attacks have increased in numbers as well as in the extent of damage caused in every attack. Large organizations operate a Cyber Security Operation Center (CSOC), which forms the first line of cyber-defense. The inspection of cyber-alerts is a critical part of CSOC operations (defender or blue team). Recent work proposed a reinforcement learning (RL) based approach for the defender\u2019s decision-making to prevent the cyber-alert queue length from growing large and overwhelming the defender. In this article, we perform a red team (adversarial) evaluation of this approach. With the recent attacks on learning-based decision-making systems, it is even more important to test the limits of the defender\u2019s RL approach. Toward that end, we learn several adversarial alert generation policies and the\n            <jats:italic>best response<\/jats:italic>\n            against them for various defender\u2019s inspection policy. Surprisingly, we find the defender\u2019s policies to be quite robust to the best response of the attacker. In order to explain this observation, we extend the earlier defender\u2019s RL model to a game model with adversarial RL, and show that there exist defender policies that can be robust against any adversarial policy. We also derive a competitive baseline from the game theory model and compare it to the defender\u2019s RL approach. However, when we go further to exploit the assumptions made in the Markov Decision Process (MDP) in the defender\u2019s RL model, we discover an attacker policy that overwhelms the defender. We use a\n            <jats:italic>double oracle<\/jats:italic>\n            like approach to retrain the defender with episodes from this discovered attacker policy. This made the defender robust to the discovered attacker policy and no further harmful attacker policies were discovered. Overall, the adversarial RL and double oracle approach in RL are general techniques that are applicable to other RL usage in adversarial environments.\n          <\/jats:p>","DOI":"10.1145\/3377554","type":"journal-article","created":{"date-parts":[[2020,4,3]],"date-time":"2020-04-03T14:20:03Z","timestamp":1585923603000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Two Can Play That Game"],"prefix":"10.1145","volume":"11","author":[{"given":"Ankit","family":"Shah","sequence":"first","affiliation":[{"name":"University of South Florida, Tampa, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arunesh","family":"Sinha","sequence":"additional","affiliation":[{"name":"Singapore Management University, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rajesh","family":"Ganesan","sequence":"additional","affiliation":[{"name":"George Mason University, Fairfax, VA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sushil","family":"Jajodia","sequence":"additional","affiliation":[{"name":"George Mason University, Fairfax, VA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hasan","family":"Cam","sequence":"additional","affiliation":[{"name":"Army Research Lab, Adelphi, MD"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,4,3]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Retrieved","author":"Ackerman Bob","year":"2017","unstructured":"Bob Ackerman . 2017 . The healthcare industry is in a world of cybersecurity hurt . Retrieved August 10, 2018 from https:\/\/tcrn.ch\/2OqmOX9.techcrunch.com. Bob Ackerman. 2017. The healthcare industry is in a world of cybersecurity hurt. Retrieved August 10, 2018 from https:\/\/tcrn.ch\/2OqmOX9.techcrunch.com."},{"volume-title":"Applications of Dynamic Games in Queues","author":"Altman Eitan","key":"e_1_2_1_2_1","unstructured":"Eitan Altman . 2005. Applications of Dynamic Games in Queues . Birkh\u00e4user Boston , 309--342. Eitan Altman. 2005. Applications of Dynamic Games in Queues. Birkh\u00e4user Boston, 309--342."},{"key":"e_1_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Itai Ashlagi Brendan Lucier and Moshe Tennenholtz. 2013. Equilibria of online scheduling algorithms. In AAAI.  Itai Ashlagi Brendan Lucier and Moshe Tennenholtz. 2013. Equilibria of online scheduling algorithms. In AAAI.","DOI":"10.1609\/aaai.v27i1.8631"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-62416-7_19"},{"volume-title":"The Tao of Network Security Monitoring: Beyond Intrusion Detection","author":"Bejtlich Richard","key":"e_1_2_1_5_1","unstructured":"Richard Bejtlich . 2005. The Tao of Network Security Monitoring: Beyond Intrusion Detection . Pearson Education Inc . Richard Bejtlich. 2005. The Tao of Network Security Monitoring: Beyond Intrusion Detection. Pearson Education Inc."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2014.103"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of IJCAI.","author":"Blocki Jeremiah","year":"2013","unstructured":"Jeremiah Blocki , Nicolas Christin , Anupam Datta , Ariel D. Procaccia , and Arunesh Sinha . 2013 . Audit games . In Proceedings of IJCAI. Jeremiah Blocki, Nicolas Christin, Anupam Datta, Ariel D. Procaccia, and Arunesh Sinha. 2013. Audit games. In Proceedings of IJCAI."},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of AAAI.","author":"Brown Matthew","year":"2016","unstructured":"Matthew Brown , Arunesh Sinha , Aaron Schlenker , and Milind Tambe . 2016 . One size does not fit all: A game-theoretic approach for dynamically and effectively screening for threats . In Proceedings of AAAI. Matthew Brown, Arunesh Sinha, Aaron Schlenker, and Milind Tambe. 2016. One size does not fit all: A game-theoretic approach for dynamically and effectively screening for threats. In Proceedings of AAAI."},{"key":"e_1_2_1_9_1","volume-title":"Intrusion detection evasion: How Attackers get past the burglar alarm. Retrieved","author":"Carlo Corbin","year":"2019","unstructured":"Corbin Carlo . 2003. Intrusion detection evasion: How Attackers get past the burglar alarm. Retrieved February 10, 2019 from https:\/\/www.sans.org\/reading-room\/whitepapers\/detection\/intrusion-detection-evasion-attackers-burglar-alarm-1284. SANS Institute . Corbin Carlo. 2003. Intrusion detection evasion: How Attackers get past the burglar alarm. Retrieved February 10, 2019 from https:\/\/www.sans.org\/reading-room\/whitepapers\/detection\/intrusion-detection-evasion-attackers-burglar-alarm-1284. SANS Institute."},{"volume-title":"On Regenerative Processes in Queueing Theory","author":"Cohen Jacob W.","key":"e_1_2_1_10_1","unstructured":"Jacob W. Cohen . 2012. On Regenerative Processes in Queueing Theory . Vol. 121 . Springer Science 8 Business Media. Jacob W. Cohen. 2012. On Regenerative Processes in Queueing Theory. Vol. 121. Springer Science 8 Business Media."},{"volume-title":"The Single Server Queue","author":"Cohen Jacob Willem","key":"e_1_2_1_11_1","unstructured":"Jacob Willem Cohen and Anthony Browne . 1982. The Single Server Queue . Vol. 8 . North-Holland Amsterdam . Jacob Willem Cohen and Anthony Browne. 1982. The Single Server Queue. Vol. 8. North-Holland Amsterdam."},{"volume-title":"Implementing Intrusion Detection Systems","author":"Crothers Tim","key":"e_1_2_1_12_1","unstructured":"Tim Crothers . 2002. Implementing Intrusion Detection Systems . Wiley Publishing Inc . Tim Crothers. 2002. Implementing Intrusion Detection Systems. Wiley Publishing Inc."},{"volume-title":"VizSEC 2007: Proceedings of the Workshop on Visualization for Computer Security","author":"D\u2019Amico Anita","key":"e_1_2_1_13_1","unstructured":"Anita D\u2019Amico and Kirsten Whitley . 2008. VizSEC 2007: Proceedings of the Workshop on Visualization for Computer Security . Springer Berlin , 19--37. DOI:https:\/\/doi.org\/10.1007\/978-3-540-78243-8_2 10.1007\/978-3-540-78243-8_2 Anita D\u2019Amico and Kirsten Whitley. 2008. VizSEC 2007: Proceedings of the Workshop on Visualization for Computer Security. Springer Berlin, 19--37. DOI:https:\/\/doi.org\/10.1007\/978-3-540-78243-8_2"},{"key":"e_1_2_1_14_1","volume-title":"International Conference on Artificial Intelligence. 526--532","author":"Durkota Karel","year":"2015","unstructured":"Karel Durkota , Viliam Lisy , Branislav Bo\u0161ansky , and Christopher Kiekintveld . 2015 . Optimal network security hardening using attack graph games . In International Conference on Artificial Intelligence. 526--532 . Karel Durkota, Viliam Lisy, Branislav Bo\u0161ansky, and Christopher Kiekintveld. 2015. Optimal network security hardening using attack graph games. In International Conference on Artificial Intelligence. 526--532."},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of AAAI.","author":"Elmalech Avshalom","year":"2015","unstructured":"Avshalom Elmalech , David Sarne , Avi Rosenfeld , and Eden Shalom Erez . 2015 . When suboptimal rules . In Proceedings of AAAI. Avshalom Elmalech, David Sarne, Avi Rosenfeld, and Eden Shalom Erez. 2015. When suboptimal rules. In Proceedings of AAAI."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2914795"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1214\/14-AAP1095"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 565--573","author":"Grze\u015b Marek","year":"2017","unstructured":"Marek Grze\u015b . 2017 . Reward shaping in episodic reinforcement learning . In Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 565--573 . Marek Grze\u015b. 2017. Reward shaping in episodic reinforcement learning. In Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 565--573."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/1597148.1597262"},{"key":"e_1_2_1_20_1","volume-title":"Adversarial attacks on neural network policies. arXiv preprint arXiv:1702.02284","author":"Huang Sandy","year":"2017","unstructured":"Sandy Huang , Nicolas Papernot , Ian Goodfellow , Yan Duan , and Pieter Abbeel . 2017. Adversarial attacks on neural network policies. arXiv preprint arXiv:1702.02284 ( 2017 ). Sandy Huang, Nicolas Papernot, Ian Goodfellow, Yan Duan, and Pieter Abbeel. 2017. Adversarial attacks on neural network policies. arXiv preprint arXiv:1702.02284 (2017)."},{"key":"e_1_2_1_21_1","unstructured":"Marc Lanctot Vinicius Zambaldi Audrunas Gruslys Angeliki Lazaridou Karl Tuyls Julien P\u00e9rolat David Silver and Thore Graepel. 2017. A unified game-theoretic approach to multiagent reinforcement learning. In Advances in Neural Information Processing Systems. 4190--4203.  Marc Lanctot Vinicius Zambaldi Audrunas Gruslys Angeliki Lazaridou Karl Tuyls Julien P\u00e9rolat David Silver and Thore Graepel. 2017. A unified game-theoretic approach to multiagent reinforcement learning. In Advances in Neural Information Processing Systems. 4190--4203."},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Aron Laszka Jian Lou and Yevgeniy Vorobeychik. 2016. Multi-defender strategic filtering against spear-phishing attacks. In AAAI.  Aron Laszka Jian Lou and Yevgeniy Vorobeychik. 2016. Multi-defender strategic filtering against spear-phishing attacks. In AAAI.","DOI":"10.1609\/aaai.v30i1.10020"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2017\/523"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2017\/525"},{"key":"e_1_2_1_25_1","volume-title":"Retrieved","author":"Manion Carl","year":"2016","unstructured":"Carl Manion . 2016 . How to Avoid Wasting Time on False Positives . Retrieved August 10, 2018 from https:\/\/www.rsaconference.com\/blogs\/how-to-avoid-wasting-time-on-false-positives. Raytheon Foreground Security. Carl Manion. 2016. How to Avoid Wasting Time on False Positives. Retrieved August 10, 2018 from https:\/\/www.rsaconference.com\/blogs\/how-to-avoid-wasting-time-on-false-positives. Raytheon Foreground Security."},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the 20th International Conference on Machine Learning (ICML-03)","author":"McMahan H. Brendan","year":"2003","unstructured":"H. Brendan McMahan , Geoffrey J. Gordon , and Avrim Blum . 2003 . Planning in the presence of cost functions controlled by an adversary . In Proceedings of the 20th International Conference on Machine Learning (ICML-03) . 536--543. H. Brendan McMahan, Geoffrey J. Gordon, and Avrim Blum. 2003. Planning in the presence of cost functions controlled by an adversary. In Proceedings of the 20th International Conference on Machine Learning (ICML-03). 536--543."},{"key":"e_1_2_1_27_1","doi-asserted-by":"crossref","unstructured":"Jean-Fran\u00c3\u011fois Mertens Sylvain Sorin and Shmuel Zamir. 2015. Repeated Games. Cambridge University Press.  Jean-Fran\u00c3\u011fois Mertens Sylvain Sorin and Shmuel Zamir. 2015. Repeated Games. Cambridge University Press.","DOI":"10.1017\/CBO9781139343275"},{"key":"e_1_2_1_28_1","first-page":"278","article-title":"Policy invariance under reward transformations: Theory and application to reward shaping","volume":"99","author":"Ng Andrew Y.","year":"1999","unstructured":"Andrew Y. Ng , Daishi Harada , and Stuart Russell . 1999 . Policy invariance under reward transformations: Theory and application to reward shaping . In ICML , Vol. 99. 278 -- 287 . Andrew Y. Ng, Daishi Harada, and Stuart Russell. 1999. Policy invariance under reward transformations: Theory and application to reward shaping. In ICML, Vol. 99. 278--287.","journal-title":"ICML"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/NCS.2018.00008"},{"key":"e_1_2_1_30_1","volume-title":"International Conference on Machine Learning. 2817--2826","author":"Pinto Lerrel","year":"2017","unstructured":"Lerrel Pinto , James Davidson , Rahul Sukthankar , and Abhinav Gupta . 2017 . Robust adversarial reinforcement learning . In International Conference on Machine Learning. 2817--2826 . Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta. 2017. Robust adversarial reinforcement learning. In International Conference on Machine Learning. 2817--2826."},{"key":"e_1_2_1_31_1","unstructured":"David Pollard. 2015. A few good inequalities. Retrieved from http:\/\/www.stat.yale.edu\/ pollard\/Books\/Mini\/Basic.pdf.  David Pollard. 2015. A few good inequalities. Retrieved from http:\/\/www.stat.yale.edu\/ pollard\/Books\/Mini\/Basic.pdf."},{"key":"e_1_2_1_32_1","volume-title":"Abbas Ghaemi Bafghi, and Mohsen Kahani","author":"Rasoulifard Amin","year":"2008","unstructured":"Amin Rasoulifard , Abbas Ghaemi Bafghi, and Mohsen Kahani . 2008 . Incremental hybrid intrusion detection using ensemble of weak classifiers. In Advances in Computer Science and Engineering. Springer , 577--584. Amin Rasoulifard, Abbas Ghaemi Bafghi, and Mohsen Kahani. 2008. Incremental hybrid intrusion detection using ensemble of weak classifiers. In Advances in Computer Science and Engineering. Springer, 577--584."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939778"},{"key":"e_1_2_1_34_1","volume-title":"International Conference on Autonomous Agents and MultiAgent Systems.","author":"Schlenker Aaron","year":"2018","unstructured":"Aaron Schlenker , Omkar Thakoor , Haifeng Xu , Fei Fang , Milind Tambe , Long Tran-Thanh , Phebe Vayanos , and Yevgeniy Vorobeychik . 2018 . Deceiving cyber adversaries: A game theoretic approach . In International Conference on Autonomous Agents and MultiAgent Systems. Aaron Schlenker, Omkar Thakoor, Haifeng Xu, Fei Fang, Milind Tambe, Long Tran-Thanh, Phebe Vayanos, and Yevgeniy Vorobeychik. 2018. Deceiving cyber adversaries: A game theoretic approach. In International Conference on Autonomous Agents and MultiAgent Systems."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2017\/54"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173457"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10207-017-0365-1"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2018.2871744"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/3176788.3176793"},{"key":"e_1_2_1_40_1","volume-title":"Retrieved","author":"Underwood Kimberly","year":"2017","unstructured":"Kimberly Underwood . 2017 . Cyber Attacks on Government Agencies on the Rise . Retrieved August 10, 2018 from https:\/\/www.afcea.org\/content\/cyber-attacks-government-agencies-rise. Kimberly Underwood. 2017. Cyber Attacks on Government Agencies on the Rise. Retrieved August 10, 2018 from https:\/\/www.afcea.org\/content\/cyber-attacks-government-agencies-rise."},{"key":"e_1_2_1_41_1","volume-title":"Lantao Yu, Yi Wu, Rohit Singh, Lucas Joppa, and Fei Fang.","author":"Wang Yufei","year":"2018","unstructured":"Yufei Wang , Zheyuan Ryan Shi , Lantao Yu, Yi Wu, Rohit Singh, Lucas Joppa, and Fei Fang. 2018 . Deep reinforcement learning for green security games with real-time information. 33, 1 (2018), 1401\u20131408. Yufei Wang, Zheyuan Ryan Shi, Lantao Yu, Yi Wu, Rohit Singh, Lucas Joppa, and Fei Fang. 2018. Deep reinforcement learning for green security games with real-time information. 33, 1 (2018), 1401\u20131408."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2018.00136"},{"key":"e_1_2_1_43_1","doi-asserted-by":"crossref","unstructured":"Mengchen Zhao Bo An and Christopher Kiekintveld. 2016. Optimizing personalized email filtering thresholds to mitigate sequential spear phishing attacks. In AAAI.  Mengchen Zhao Bo An and Christopher Kiekintveld. 2016. Optimizing personalized email filtering thresholds to mitigate sequential spear phishing attacks. In AAAI.","DOI":"10.1609\/aaai.v30i1.10030"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3377554","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3377554","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:33:18Z","timestamp":1750199598000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3377554"}},"subtitle":["An Adversarial Evaluation of a Cyber-Alert Inspection System"],"short-title":[],"issued":{"date-parts":[[2020,4,3]]},"references-count":43,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2020,6,30]]}},"alternative-id":["10.1145\/3377554"],"URL":"https:\/\/doi.org\/10.1145\/3377554","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"type":"print","value":"2157-6904"},{"type":"electronic","value":"2157-6912"}],"subject":[],"published":{"date-parts":[[2020,4,3]]},"assertion":[{"value":"2019-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-04-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}