{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T17:10:55Z","timestamp":1784913055635,"version":"3.55.0"},"reference-count":52,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2025,5,12]],"date-time":"2025-05-12T00:00:00Z","timestamp":1747008000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,5,12]],"date-time":"2025-05-12T00:00:00Z","timestamp":1747008000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100000145","name":"Division of Information and Intelligent Systems","doi-asserted-by":"publisher","award":["997427"],"award-info":[{"award-number":["997427"]}],"id":[{"id":"10.13039\/100000145","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000145","name":"Division of Information and Intelligent Systems","doi-asserted-by":"publisher","award":["997356"],"award-info":[{"award-number":["997356"]}],"id":[{"id":"10.13039\/100000145","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007297","name":"Office of Naval Research Global","doi-asserted-by":"publisher","award":["997041"],"award-info":[{"award-number":["997041"]}],"id":[{"id":"10.13039\/100007297","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"publisher","award":["997694"],"award-info":[{"award-number":["997694"]}],"id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["The VLDB Journal"],"published-print":{"date-parts":[[2025,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>The ability to reuse trained models in Reinforcement Learning (RL) holds substantial practical value in particular for complex tasks. While model reusability is widely studied for supervised models in data management, to the best of our knowledge, this is the first ever principled study that is proposed for RL. To capture trained policies, we develop a framework based on an expressive and lossless graph data model that accommodates Temporal Difference Learning and Deep-RL based RL algorithms. Our framework is able to capture arbitrary reward functions that can be composed at inference time. The framework comes with theoretical guarantees and shows that it yields the same result as policies trained from scratch. We design a parameterized algorithm that strikes a balance between efficiency and quality w.r.t cumulative reward. Our experiments with two common RL tasks (query refinement and robot movement) corroborate our theory and show the effectiveness and efficiency of our algorithms.\n<\/jats:p>","DOI":"10.1007\/s00778-025-00920-0","type":"journal-article","created":{"date-parts":[[2025,5,12]],"date-time":"2025-05-12T14:29:38Z","timestamp":1747060178000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Model reusability in Reinforcement Learning"],"prefix":"10.1007","volume":"34","author":[{"given":"Sepideh","family":"Nikookar","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sohrab","family":"Namazi Nia","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3475-8138","authenticated-orcid":false,"given":"Senjuti","family":"Basu Roy","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sihem","family":"Amer-Yahia","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Behrooz","family":"Omidvar-Tehrani","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,5,12]]},"reference":[{"key":"920_CR1","doi-asserted-by":"crossref","unstructured":"Abbeel, P., Coates, A., Quigley, M., Ng, A.: An application of reinforcement learning to aerobatic helicopter flight. In: Advances in Neural Information Processing Systems 19 (2006)","DOI":"10.7551\/mitpress\/7503.003.0006"},{"key":"920_CR2","unstructured":"Anonymous Technical report. https:\/\/www.dropbox.com\/scl\/fo\/ikd6w88mv2jsa11bdlj7v\/h?rlkey=lrpsox76rhyvvmosu4m7ckruf&dl=0 (2024)"},{"issue":"6","key":"920_CR3","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1109\/MSP.2017.2743240","volume":"34","author":"K Arulkumaran","year":"2017","unstructured":"Arulkumaran, K., Deisenroth, M.P., Brundage, M., Bharath, A.A.: Deep reinforcement learning: a brief survey. IEEE Signal Process. Mag. 34(6), 26\u201338 (2017)","journal-title":"IEEE Signal Process. Mag."},{"key":"920_CR4","doi-asserted-by":"crossref","unstructured":"Bacon, P.L., Harb, J., Precup, D.: The option-critic architecture. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol\u00a031 (2017)","DOI":"10.1609\/aaai.v31i1.10916"},{"key":"920_CR5","doi-asserted-by":"crossref","unstructured":"Bar\u00a0El, O., Milo, T., Somech, A.: Automatically generating data exploration sessions using deep reinforcement learning. In: Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data (2020)","DOI":"10.1145\/3318464.3389779"},{"key":"920_CR6","unstructured":"Barrett, S., Taylor, M.E., Stone, P.: Transfer learning for reinforcement learning on a physical robot. In: Ninth International Conference on Autonomous Agents and Multiagent Systems-Adaptive Learning Agents Workshop (AAMAS-ALA), vol.\u00a01 (2010)"},{"key":"920_CR7","doi-asserted-by":"crossref","unstructured":"Cai, H., Ren, K., Zhang, W., Malialis, K., Wang, J., Yu, Y., Guo, D.: Real-time bidding by reinforcement learning in display advertising. WSDM \u201917, pp. 661\u2013670 (2017)","DOI":"10.1145\/3018661.3018702"},{"issue":"1","key":"920_CR8","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1017\/S0963548306007887","volume":"16","author":"P Charbit","year":"2007","unstructured":"Charbit, P., Thomass\u00e9, S., Yeo, A.: The minimum feedback arc set problem is np-hard for tournaments. Comb. Probab. Comput. 16(1), 1\u20134 (2007)","journal-title":"Comb. Probab. Comput."},{"key":"920_CR9","unstructured":"Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)"},{"issue":"6","key":"920_CR10","doi-asserted-by":"publisher","first-page":"319","DOI":"10.1016\/0020-0190(93)90079-O","volume":"47","author":"P Eades","year":"1993","unstructured":"Eades, P., Lin, X., Smyth, W.F.: A fast and effective heuristic for the feedback arc set problem. Inf. Process. Lett. 47(6), 319\u2013323 (1993)","journal-title":"Inf. Process. Lett."},{"key":"920_CR11","doi-asserted-by":"crossref","unstructured":"Fomin, F., Lokshtanov, D., Raman, V., Saurabh, S.: Fast local search algorithm for weighted feedback arc set in tournaments. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol.\u00a024, pp. 65\u201370 (2010)","DOI":"10.1609\/aaai.v24i1.7557"},{"issue":"1","key":"920_CR12","first-page":"1437","volume":"16","author":"J Garc\u0131a","year":"2015","unstructured":"Garc\u0131a, J., Fern\u00e1ndez, F.: A comprehensive survey on safe reinforcement learning. J. Mach. Learn. Res. 16(1), 1437\u20131480 (2015)","journal-title":"J. Mach. Learn. Res."},{"key":"920_CR13","unstructured":"Garey, M.R., Johnson, D.S.: Computers and intractability. A Guide to the (1979)"},{"key":"920_CR14","unstructured":"Gupta, A., Mendonca, R., Liu, Y., Abbeel, P., Levine, S.: Meta-reinforcement learning of structured exploration strategies. In: Advances in Neural Information Processing Systems 31 (2018)"},{"key":"920_CR15","doi-asserted-by":"publisher","DOI":"10.1016\/j.autcon.2023.104817","volume":"150","author":"C Jiang","year":"2023","unstructured":"Jiang, C., Li, X., Lin, J., Liu, M., Ma, Z.: Adaptive control of resource flow to optimize construction work and cash flow via online deep reinforcement learning. Autom. Constr. 150, 104817 (2023)","journal-title":"Autom. Constr."},{"key":"920_CR16","doi-asserted-by":"publisher","first-page":"237","DOI":"10.1613\/jair.301","volume":"4","author":"LP Kaelbling","year":"1996","unstructured":"Kaelbling, L.P., Littman, M.L., Moore, A.W.: Reinforcement learning: a survey. J. Artif. Intell. Res. 4, 237\u2013285 (1996)","journal-title":"J. Artif. Intell. Res."},{"key":"920_CR17","doi-asserted-by":"crossref","unstructured":"Kraska, T., Beutel, A., Chi, EH., Dean, J., Polyzotis, N.: The case for learned index structures. In: Proceedings of the 2018 International Conference on Management of Data, pp. 489\u2013504 (2018)","DOI":"10.1145\/3183713.3196909"},{"key":"920_CR18","unstructured":"Lezzar, F., Zidani, A., Atef, C.: A collaborative web-based application for health care tasks planning. In: Proceedings of the 4th International Conference on Web and Information Technologies, Citeseer, pp. 30\u201339 (2012)"},{"key":"920_CR19","unstructured":"Li, Y.: Deep reinforcement learning: an overview. arXiv preprint arXiv:1701.07274 (2017)"},{"issue":"3","key":"920_CR20","doi-asserted-by":"publisher","first-page":"414","DOI":"10.14778\/3494124.3494127","volume":"15","author":"Y Lu","year":"2021","unstructured":"Lu, Y., Kandula, S., K\u00f6nig, A.C., Chaudhuri, S.: Pre-training summarization models of structured datasets for cardinality estimation. Proc. VLDB Endow. 15(3), 414\u2013426 (2021)","journal-title":"Proc. VLDB Endow."},{"key":"920_CR21","unstructured":"Mahadevan, S., Theocharous, G.: Optimizing production manufacturing using reinforcement learning. In: Proceedings of the Eleventh International Florida Artificial Intelligence Research Society Conference, Citeseer, vol. 372, p. 377 (1998)"},{"key":"920_CR22","doi-asserted-by":"crossref","unstructured":"Mao, H., Alizadeh, M., Menache, I., Kandula, S.: Resource management with deep reinforcement learning. In: Proceedings of the 15th ACM Workshop on Hot Topics in Networks, pp. 50\u201356 (2016)","DOI":"10.1145\/3005745.3005750"},{"key":"920_CR23","unstructured":"McMahan, B., Moore, E., Ramage, D., Hampson, S., y\u00a0Arcas, BA.: Communication-efficient learning of deep networks from decentralized data. In: Artificial intelligence and statistics, PMLR, pp 1273\u20131282 (2017)"},{"key":"920_CR24","doi-asserted-by":"crossref","unstructured":"Nakandala, S., Kumar, A.: Nautilus: An optimized system for deep transfer learning over evolving training datasets. In: Proceedings of the 2022 International Conference on Management of Data, pp. 506\u2013520 (2022)","DOI":"10.1145\/3514221.3517846"},{"key":"920_CR25","doi-asserted-by":"crossref","unstructured":"Nikookar, S., Sakharkar, P., Smagh, B., Amer-Yahia, S., Roy, S.B.: Guided task planning under complex constraints. In: 2022 IEEE 38th International Conference on Data Engineering (ICDE), pp. 833\u2013845 , IEEE (2022)","DOI":"10.1109\/ICDE53745.2022.00067"},{"key":"920_CR26","doi-asserted-by":"crossref","unstructured":"Nikookar, S., Sakharkar, P., Somasunder, S., Basu\u00a0Roy, S., Bienkowski, A., Macesker, M, Pattipati, KR., Sidoti, D.: Cooperative route planning framework for multiple distributed assets in maritime applications. In: Proceedings of the 2022 International Conference on Management of Data, pp. 1518\u20131527 (2022b)","DOI":"10.1145\/3514221.3526131"},{"key":"920_CR27","unstructured":"Obando-Ceron, J.S., Castro, P.S.: Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research. CoRR arXiv:2011.14826 (2020)"},{"key":"920_CR28","doi-asserted-by":"crossref","unstructured":"Omidvar-Tehrani, B., Personnaz, A., Amer-Yahia, S.: Guided text-based item exploration. In: Hasan M.A., Xiong, L. (eds.) Proceedings of the 31st ACM International Conference on Information & Knowledge Management, Atlanta, GA, USA, October 17\u201321, 2022, pp. 3410\u20133420. ACM (2022)","DOI":"10.1145\/3511808.3557141"},{"key":"920_CR29","doi-asserted-by":"crossref","unstructured":"Pednault, E., Abe, N., Zadrozny, B.: Sequential cost-sensitive decision making with reinforcement learning. In: Proceedings of the Eighth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 259\u2013268 (2002)","DOI":"10.1145\/775047.775086"},{"key":"920_CR30","unstructured":"Pertsch, K., Lee, Y., Lim, J.: Accelerating reinforcement learning with learned skill priors. In: Conference on Robot Learning, PMLR, pp. 188\u2013204 (2021)"},{"key":"920_CR31","unstructured":"Pineau, J.: Reproducible, reusable, and robust reinforcement learning. In: Advances in Neural Information Processing Systems (2018)"},{"key":"920_CR32","doi-asserted-by":"crossref","unstructured":"Rosset, C., Jose, D., Ghosh, G., Mitra, B., Tiwary, S.: Optimizing query evaluations using reinforcement learning for web search. In: The 41st International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 1193\u20131196 (2018)","DOI":"10.1145\/3209978.3210127"},{"key":"920_CR33","doi-asserted-by":"crossref","unstructured":"Seleznova, M., Omidvar-Tehrani, B., Amer-Yahia, S., Simon, E.: Guided exploration of user groups. Proc. VLDB Endow. (PVLDB) 13(9), 1469\u20131482 (2020)","DOI":"10.14778\/3397230.3397242"},{"key":"920_CR34","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1007\/s10994-010-5229-0","volume":"84","author":"SM Shortreed","year":"2011","unstructured":"Shortreed, S.M., Laber, E., Lizotte, D.J., Stroup, T.S., Pineau, J., Murphy, S.A.: Informing sequential clinical decision-making through reinforcement learning: an empirical study. Mach. Learn. 84, 109\u2013136 (2011)","journal-title":"Mach. Learn."},{"key":"920_CR35","doi-asserted-by":"publisher","first-page":"323","DOI":"10.1023\/A:1022680823223","volume":"8","author":"SP Singh","year":"1992","unstructured":"Singh, S.P.: Transfer of learning by composing solutions of elemental sequential tasks. Mach. Learn. 8, 323\u2013339 (1992)","journal-title":"Mach. Learn."},{"key":"920_CR36","doi-asserted-by":"crossref","unstructured":"Singh, S.P., Sutton, R.S.: Reinforcement learning with replacing eligibility traces. Mach. Learn. 22, 123\u2013158 (1996)","DOI":"10.1007\/BF00114726"},{"issue":"4","key":"920_CR37","doi-asserted-by":"publisher","first-page":"42","DOI":"10.1145\/1107499.1107504","volume":"34","author":"M Stonebraker","year":"2005","unstructured":"Stonebraker, M., \u00c7etintemel, U., Zdonik, S.: The 8 requirements of real-time stream processing. ACM SIGMOD Rec. 34(4), 42\u201347 (2005)","journal-title":"ACM SIGMOD Rec."},{"key":"920_CR38","volume-title":"Reinforcement Learning: An Introduction","author":"RS Sutton","year":"2018","unstructured":"Sutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction. MIT Press, Cambridge (2018)"},{"key":"920_CR39","unstructured":"Taylor, M.E., Stone, P.: Transfer learning for reinforcement learning domains: a survey. J. Mach. Learn. Res. 10(7), 1633\u20131685 (2009)"},{"issue":"1","key":"920_CR40","doi-asserted-by":"publisher","first-page":"1","DOI":"10.31449\/inf.v45i1.3104","volume":"45","author":"A Tlili","year":"2021","unstructured":"Tlili, A., Chikhi, S.: Risks analyzing and management in software project management using fuzzy cognitive maps with reinforcement learning. Informatica 45(1), 1\u201324 (2021)","journal-title":"Informatica"},{"key":"920_CR41","doi-asserted-by":"crossref","unstructured":"Torrey, L., Shavlik, J.: Transfer learning. In: Handbook of Research on Machine Learning Applications and Trends: Algorithms, Methods, and Techniques, IGI global, pp. 242\u2013264 (2010)","DOI":"10.4018\/978-1-60566-766-9.ch011"},{"issue":"2","key":"920_CR42","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1145\/2641190.2641198","volume":"15","author":"J Vanschoren","year":"2014","unstructured":"Vanschoren, J., Van Rijn, J.N., Bischl, B., Torgo, L.: Openml: networked science in machine learning. ACM SIGKDD Explor. Newsl. 15(2), 49\u201360 (2014)","journal-title":"ACM SIGKDD Explor. Newsl."},{"issue":"9","key":"920_CR43","doi-asserted-by":"publisher","first-page":"1363","DOI":"10.3390\/electronics9091363","volume":"9","author":"Varghese N Vithayathil","year":"2020","unstructured":"Vithayathil, Varghese N., Mahmoud, Q.H.: A survey of multi-task deep reinforcement learning. Electronics 9(9), 1363 (2020)","journal-title":"Electronics"},{"key":"920_CR44","unstructured":"Von\u00a0Hessling, A., Goel, A.K.: Abstracting reusable cases from reinforcement learning. In: ICCBR Workshops, pp. 227\u2013236 (2005)"},{"key":"920_CR45","unstructured":"Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., Freitas, N.: Dueling network architectures for deep reinforcement learning. In: International Conference on Machine Learning, PMLR, pp. 1995\u20132003 (2016)"},{"key":"920_CR46","doi-asserted-by":"crossref","unstructured":"Watkins, C.J., Dayan, P.: Q-learning. Mach. Learn. 8, 279\u2013292 (1992)","DOI":"10.1007\/BF00992698"},{"key":"920_CR47","doi-asserted-by":"crossref","unstructured":"Wei, H., et\u00a0al.: Intellilight: a reinforcement learning approach for intelligent traffic light control. In: SIGKDD (2018)","DOI":"10.1145\/3219819.3220096"},{"key":"920_CR48","doi-asserted-by":"crossref","unstructured":"Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et\u00a0al.: Huggingface\u2019s transformers: state-of-the-art natural language processing. arXiv preprint arXiv:1910.03771 (2019)","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"issue":"9","key":"920_CR49","doi-asserted-by":"publisher","first-page":"1798","DOI":"10.14778\/3538598.3538603","volume":"15","author":"B Youngmann","year":"2022","unstructured":"Youngmann, B., Amer-Yahia, S., Personnaz, A.: Guided exploration of data summaries. Proc. VLDB Endow. 15(9), 1798\u20131807 (2022)","journal-title":"Proc. VLDB Endow."},{"key":"920_CR50","doi-asserted-by":"crossref","unstructured":"Yu, Y.: Towards sample efficient reinforcement learning. In: IJCAI, pp. 5739\u20135743 (2018)","DOI":"10.24963\/ijcai.2018\/820"},{"issue":"6","key":"920_CR51","doi-asserted-by":"publisher","first-page":"2204","DOI":"10.1109\/TNNLS.2018.2803729","volume":"29","author":"Y Yu","year":"2018","unstructured":"Yu, Y., Chen, S.Y., Da, Q., Zhou, Z.H.: Reusable reinforcement learning via shallow trails. IEEE Trans. Neural Netw. Learn. Syst. 29(6), 2204\u20132215 (2018)","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"920_CR52","doi-asserted-by":"crossref","unstructured":"Zhao, C., He, Y.: Auto-em: End-to-end fuzzy entity-matching using pre-trained deep models and transfer learning. In: The World Wide Web Conference, pp. 2413\u20132424 (2019)","DOI":"10.1145\/3308558.3313578"}],"container-title":["The VLDB Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00778-025-00920-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00778-025-00920-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00778-025-00920-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,6]],"date-time":"2025-09-06T14:14:00Z","timestamp":1757168040000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00778-025-00920-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,12]]},"references-count":52,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,7]]}},"alternative-id":["920"],"URL":"https:\/\/doi.org\/10.1007\/s00778-025-00920-0","relation":{},"ISSN":["1066-8888","0949-877X"],"issn-type":[{"value":"1066-8888","type":"print"},{"value":"0949-877X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,12]]},"assertion":[{"value":"5 May 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 January 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 April 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 May 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"41"}}