{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T16:38:37Z","timestamp":1753893517370,"version":"3.41.2"},"reference-count":46,"publisher":"American Association for the Advancement of Science (AAAS)","content-domain":{"domain":["spj.science.org"],"crossmark-restriction":true},"short-container-title":["Intell Comput"],"published-print":{"date-parts":[[2023,1]]},"abstract":"<jats:p>Reinforcement learning (RL) is indispensable for building intelligent decision-making agents. However, current RL algorithms suffer from statistical and computational inefficiencies that render them useless in most real-world applications. We argue that high-value information in the real world is essential for intelligent decision-making; however, it is not addressed by most RL formalisms. Through a closer investigation of high-value information, it becomes evident that, to exploit high-value information, there is a need to formalize intelligent decision-making as bounded-optimal lifelong RL. Thus, the challenge of achieving intelligent decision-making is summarized as effectively surfing information, specifically regarding handling the non-IID (independent and identically distributed) information stream while operating with limited resources. This study discusses the design of an intelligent decision-making agent and examines its primary challenges, which are (a) online learning for non-IID data streams, (b) efficient reasoning with limited resources, and (c) the exploration\u2013exploitation dilemma. We review relevant problems and research in the field of RL literature and conclude that current RL methods are insufficient to address these challenges. We propose that an agent capable of overcoming these challenges could effectively surf the information overload in the real world and achieve sample- and compute-efficient intelligent decision-making.<\/jats:p>","DOI":"10.34133\/icomputing.0041","type":"journal-article","created":{"date-parts":[[2023,6,15]],"date-time":"2023-06-15T13:56:42Z","timestamp":1686837402000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark_01","source":"Crossref","is-referenced-by-count":0,"title":["Surfing Information: The Challenge of Intelligent Decision-Making"],"prefix":"10.34133","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0920-7895","authenticated-orcid":false,"given":"Chenyang","family":"Wu","sequence":"first","affiliation":[{"name":"National Key Lab for Novel Software Technology, Nanjing University, Nanjing 210023, China."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9238-4747","authenticated-orcid":true,"given":"Zongzhang","family":"Zhang","sequence":"additional","affiliation":[{"name":"National Key Lab for Novel Software Technology, Nanjing University, Nanjing 210023, China."}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"221","published-online":{"date-parts":[[2023,7,6]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Sutton RS Barto AG. Reinforcement learning: An introduction . Cambridge (MA): MIT Press; 2018."},{"key":"e_1_3_3_3_2","doi-asserted-by":"crossref","DOI":"10.1016\/j.artint.2021.103535","article-title":"Reward is enough","volume":"299","author":"Silver D","year":"2021","unstructured":"Silver D, Singh S, Precup D, Sutton RS. Reward is enough. Artif Intell. 2021;299:Article 103535.","journal-title":"Artif Intell"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-03051-4"},{"key":"e_1_3_3_5_2","unstructured":"Magureanu S Combes R Proutiere A. Lipschitz bandits: Regret lower bound and optimal algorithms. Paper presented at: Proceedings of the 27th Conference on Learning Theory; 2014 May 19; Barcelona Spain."},{"key":"e_1_3_3_6_2","unstructured":"Dong K Yang J Ma T. Provable model-based nonlinear bandit and reinforcement learning: Shelve optimism embrace virtual curvature. Paper presented at: Advances in Neural Information Processing Systems. 2021; 34 :26168\u201326182."},{"key":"e_1_3_3_7_2","doi-asserted-by":"crossref","first-page":"253","DOI":"10.1613\/jair.3912","article-title":"The arcade learning environment: An evaluation platform for general agents","volume":"47","author":"Bellemare MG","year":"2013","unstructured":"Bellemare MG, Naddaf Y, Veness J, Bowling M. The arcade learning environment: An evaluation platform for general agents. J Artif Intell Res. 2013;47:253\u2013279.","journal-title":"J Artif Intell Res"},{"issue":"3","key":"e_1_3_3_8_2","doi-asserted-by":"crossref","first-page":"441","DOI":"10.1287\/moor.12.3.441","article-title":"The complexity of Markov decision processes","volume":"12","author":"Papadimitriou CH","year":"1987","unstructured":"Papadimitriou CH, Tsitsiklis JN. The complexity of Markov decision processes. Math Oper Res. 1987;12(3):441\u2013450.","journal-title":"Math Oper Res"},{"key":"e_1_3_3_9_2","unstructured":"LeCun Y. A path towards autonomous machine intelligence version 0.9.2. Preprint posted on openreview; 2022 Jun 27. https:\/\/openreview.net\/pdf?id=BZ5a1r-kVsf."},{"key":"e_1_3_3_10_2","unstructured":"Russell S. Fundamental issues of artificial intelligence . Cham: Springer; 2016."},{"key":"e_1_3_3_11_2","unstructured":"Lu X Roy BV Dwaracherla V Ibrahimi M Osband I Wen Z. Reinforcement learning bit by bit. ArXiv 2021. https:\/\/doi.org\/10.48550\/arXiv.2103.04047"},{"issue":"1","key":"e_1_3_3_12_2","first-page":"575","article-title":"Provably bounded-optimal agents","volume":"2","author":"Russell SJ","year":"1994","unstructured":"Russell SJ, Subramanian D. Provably bounded-optimal agents. J Artif Intell Res. 1994;2(1):575\u2013609.","journal-title":"J Artif Intell Res"},{"key":"e_1_3_3_13_2","doi-asserted-by":"crossref","first-page":"1401","DOI":"10.1613\/jair.1.13673","article-title":"Towards continual reinforcement learning: A review and perspectives","volume":"75","author":"Khetarpal K","year":"2022","unstructured":"Khetarpal K, Riemer M, Rish I, Precup D. Towards continual reinforcement learning: A review and perspectives. J Artif Intell Res. 2022;75:1401\u20131476.","journal-title":"J Artif Intell Res"},{"key":"e_1_3_3_14_2","doi-asserted-by":"crossref","unstructured":"Brachman R Levesque H. Knowledge representation and reasoning . San Francisco (CA): Morgan Kaufmann; 2004.","DOI":"10.1016\/B978-155860932-7\/50099-6"},{"key":"e_1_3_3_15_2","unstructured":"Russell S Norvig P. Artificial intelligence: A modern approach . Hoboken (NJ): Pearson; 2009."},{"issue":"226","key":"e_1_3_3_16_2","doi-asserted-by":"crossref","DOI":"10.1098\/rspa.2021.0068","article-title":"Inductive biases for deep learning of higher-level cognition","volume":"478","author":"Goyal A","year":"2022","unstructured":"Goyal A, Bengio Y. Inductive biases for deep learning of higher-level cognition. Proc Society A: Math Phys Eng Sci. 2022;478(226):Article 20210068.","journal-title":"Proc Society A: Math Phys Eng Sci"},{"issue":"7","key":"e_1_3_3_17_2","doi-asserted-by":"crossref","first-page":"1341","DOI":"10.1162\/neco.1996.8.7.1341","article-title":"The lack of a priori distinctions between learning algorithms","volume":"8","author":"Wolpert DH","year":"1996","unstructured":"Wolpert DH. The lack of a priori distinctions between learning algorithms. Neural Comput. 1996;8(7):1341\u20131390.","journal-title":"Neural Comput"},{"issue":"1","key":"e_1_3_3_18_2","doi-asserted-by":"crossref","first-page":"361","DOI":"10.1016\/0004-3702(91)90015-C","article-title":"Principles of metareasoning","volume":"49","author":"Russell S","year":"1991","unstructured":"Russell S, Wefald E. Principles of metareasoning. Artif Intell. 1991;49(1):361\u2013395.","journal-title":"Artif Intell"},{"key":"e_1_3_3_19_2","unstructured":"Thrun S. Efficient exploration in reinforcement learning . Tech. Rep. CMU-CS-92-102Pittsburgh; PA: Carnegie Mellon University; 1992."},{"key":"e_1_3_3_20_2","unstructured":"Gama J. Knowledge discovery from data streams . Boca Raton (FL): Chapman & Hall\/CRC; 2010."},{"issue":"3","key":"e_1_3_3_21_2","doi-asserted-by":"crossref","DOI":"10.1002\/widm.1405","article-title":"Data stream analysis: Foundations, major tasks and tools","volume":"11","author":"Bahri M","year":"2021","unstructured":"Bahri M, Bifet A, Gama J, Gomes HM, Maniu S. Data stream analysis: Foundations, major tasks and tools. Data Min Knowl Disc. 2021;11(3):Article e1405.","journal-title":"Data Min Knowl Disc"},{"key":"e_1_3_3_22_2","unstructured":"Zhou Z-H. Stream efficient learning. ArXiv 2023. https:\/\/doi.org\/10.48550\/arXiv.2305.02217"},{"key":"e_1_3_3_23_2","unstructured":"Sutton RS. Temporal credit assignment in reinforcement learning [thesis]. [Amherst (MA)]: University of Massachusetts Amherst; 1984."},{"key":"e_1_3_3_24_2","unstructured":"Cohen T Welling M. Paper presented at: Proceedings of the 33rd International Conference on Machine Learning; 2016 Jun 19\u201324; New York USA."},{"issue":"1","key":"e_1_3_3_25_2","first-page":"857","article-title":"Self-supervised learning: Generative or contrastive","volume":"35","author":"Liu X","year":"2021","unstructured":"Liu X, Zhang F, Hou Z, Mian L, Wang Z, Zhang J, Tang J. Self-supervised learning: Generative or contrastive. IEEE Trans Knowl Data Eng. 2021;35(1):857\u2013876.","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"e_1_3_3_26_2","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1017\/S0962492921000027","article-title":"Deep learning: A statistical viewpoint","volume":"30","author":"Bartlett PL","year":"2021","unstructured":"Bartlett PL, Montanari A, Rakhlin A. Deep learning: A statistical viewpoint. Acta Numerica. 2021;30:87\u2013201.","journal-title":"Acta Numerica"},{"key":"e_1_3_3_27_2","unstructured":"Smith SL Dherin B Barrett D De S. On the origin of implicit regularization in stochastic gradient descent. Paper presented at: International Conference on Learning Representations; 2022; Virtual."},{"key":"e_1_3_3_28_2","doi-asserted-by":"crossref","unstructured":"Xu Z-QJ Zhang Y Xiao Y. Training behavior of deep neural network in frequency domain. Neural Inform Process. 2019;264\u2013274.","DOI":"10.1007\/978-3-030-36708-4_22"},{"issue":"9","key":"e_1_3_3_29_2","first-page":"5149","article-title":"Meta-learning in neural networks: A survey","volume":"44","author":"Hospedales T","year":"2022","unstructured":"Hospedales T, Antoniou A, Micaelli P, Storkey A. Meta-learning in neural networks: A survey. IEEE Trans Pattern Anal Mach Intell. 2022;44(9):5149\u20135169.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"1","key":"e_1_3_3_30_2","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1109\/JPROC.2020.3004555","article-title":"A comprehensive survey on transfer learning","volume":"109","author":"Zhuang F","year":"2021","unstructured":"Zhuang F, Qi Z, Duan K, Xi D, Zhu Y, Zhu H, Xiong H,He Q. A comprehensive survey on transfer learning. Proc IEEE. 2021;109(1):43\u201376.","journal-title":"Proc IEEE"},{"key":"e_1_3_3_31_2","unstructured":"Silver DL Yang Q Li L. Lifelong machine learning systems: Beyond learning algorithms. In Lifelong machine learning ; California USA: AAAI; 2013."},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2019.01.012"},{"issue":"8","key":"e_1_3_3_33_2","doi-asserted-by":"crossref","DOI":"10.1093\/nsr\/nwac123","article-title":"Open-environment machine learning","volume":"9","author":"Zhou Z-H","year":"2022","unstructured":"Zhou Z-H. Open-environment machine learning. Natl Sci Rev. 2022;9(8):Article nwac123.","journal-title":"Natl Sci Rev"},{"issue":"3","key":"e_1_3_3_34_2","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1109\/2.33","article-title":"The ART of adaptive pattern recognition by a self-organizing neural network","volume":"21","author":"Carpenter G","year":"1988","unstructured":"Carpenter G, Grossberg S. The ART of adaptive pattern recognition by a self-organizing neural network. Computer. 1988;21(3):77\u201388.","journal-title":"Computer"},{"key":"e_1_3_3_35_2","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1016\/S0079-7421(08)60536-8","article-title":"Catastrophic interference in connectionist networks: The sequential learning problem","volume":"24","author":"McCloskey M","year":"1989","unstructured":"McCloskey M, Cohen NJ. Catastrophic interference in connectionist networks: The sequential learning problem. Psychol Learn Motiv. 1989;24:109\u2013165.","journal-title":"Psychol Learn Motiv"},{"issue":"1","key":"e_1_3_3_36_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/0010-0277(88)90031-5","article-title":"Connectionism and cognitive architecture: A critical analysis","volume":"28","author":"Fodor JA","year":"1988","unstructured":"Fodor JA, Pylyshyn ZW. Connectionism and cognitive architecture: A critical analysis. Cognition. 1988;28(1):3\u201371.","journal-title":"Cognition"},{"key":"e_1_3_3_37_2","first-page":"603","volume":"107","author":"Russin J","year":"2020","unstructured":"Russin J, O\u2019Reilly RC, Bengio Y. Deep learning needs a prefrontal cortex. Paper presented at: ICLR Bridging AI and Cognitive Science (BAICS) Workshop. 2020;107:603\u2013616.","journal-title":"ICLR Bridging AI and Cognitive Science (BAICS) Workshop"},{"key":"e_1_3_3_38_2","unstructured":"Mnih V Kavukcuoglu K Silver D Graves A Antonoglou I Wierstra D Riedmiller M. Playing Atari with deep reinforcement learning. ArXiv 2013. https:\/\/doi.org\/10.48550\/arXiv.1312.5602"},{"issue":"1","key":"e_1_3_3_39_2","doi-asserted-by":"crossref","first-page":"3302","DOI":"10.1609\/aaai.v32i1.11595","article-title":"Selective experience replay for lifelong learning","volume":"32","author":"Isele D","year":"2018","unstructured":"Isele D, Cosgun A. Selective experience replay for lifelong learning. AAAI. 2018;32(1):3302\u20133309.","journal-title":"AAAI"},{"key":"e_1_3_3_40_2","unstructured":"Rolnick D Ahuja A Schwarz J Lillicrap T Wayne G. Experience replay for continual learning. Paper presented at: Advances in Neural Information Processing Systems; 2019; 32 ."},{"key":"e_1_3_3_41_2","unstructured":"Li X Shang J Das S Ryoo MS. Does self-supervised learning really improve reinforcement learning from pixels?. Paper presented at: Advances in Neural Information Processing Systems; 2022; 35 :30865\u201330881."},{"key":"e_1_3_3_42_2","unstructured":"Li W et\u00a0al. A survey on transformers in reinforcement learning. ArXiv 2023. https:\/\/doi.org\/10.48550\/arXiv.2301.03044"},{"key":"e_1_3_3_43_2","unstructured":"Mialon G Dessi R Lomeli M Nalmpantis C Pasunuru R Raileanu R Roziere B Schick T Dwivedi-Yu J Celikyilmaz A et\u00a0al. Augmented language models: A survey. ArXiv 2023. https:\/\/doi.org\/10.48550\/arXiv.2302.07842"},{"key":"e_1_3_3_44_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.inffus.2022.03.003","article-title":"Exploration in deep reinforcement learning: A survey","volume":"85","author":"Ladosz P","year":"2022","unstructured":"Ladosz P, Weng L, Kim M, Oh H. Exploration in deep reinforcement learning: A survey. Inf Fusion. 2022;85:1\u201322.","journal-title":"Inf Fusion"},{"key":"e_1_3_3_45_2","unstructured":"Azar MG Osband I Munos R. Minimax Regret Bounds for Reinforcement Learning. Paper presented at: Proceedings of the 34th International Conference on Machine Learning; 2017; 70 :263\u2013272."},{"key":"e_1_3_3_46_2","unstructured":"Jin C Yang Z Wang Z Jordan MI. Provably e\ufb00icient reinforcement learning with linear function approximation.Paper presented at: Proceedings of 33rd Conference on Learning Theory; 2020; 125 :2137\u20132143."},{"key":"e_1_3_3_47_2","unstructured":"Wu C Li T Zhang Z Yu Y. Bayesian optimistic optimization: Optimistic exploration for model-based reinforcement learning. Paper presented at: Advances in Neural Information Processing Systems; 2022;35:14210\u201314223."}],"container-title":["Intelligent Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/spj.science.org\/doi\/pdf\/10.34133\/icomputing.0041","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,9]],"date-time":"2024-01-09T14:10:23Z","timestamp":1704809423000},"score":1,"resource":{"primary":{"URL":"https:\/\/spj.science.org\/doi\/10.34133\/icomputing.0041"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1]]},"references-count":46,"alternative-id":["10.34133\/icomputing.0041"],"URL":"https:\/\/doi.org\/10.34133\/icomputing.0041","relation":{},"ISSN":["2771-5892"],"issn-type":[{"type":"electronic","value":"2771-5892"}],"subject":[],"published":{"date-parts":[[2023,1]]},"assertion":[{"value":"2023-03-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-29","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"0041"}}