{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T16:44:38Z","timestamp":1781196278907,"version":"3.54.1"},"reference-count":48,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T00:00:00Z","timestamp":1701043200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T00:00:00Z","timestamp":1701043200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004238","name":"Universit\u00e4t Potsdam","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004238","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Bus Inf Syst Eng"],"published-print":{"date-parts":[[2024,8]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Nowadays, customers as well as retailers look for increased sustainability. Recommerce markets \u2013 which offer the opportunity to trade-in and resell used products \u2013 are constantly growing and help to use resources more efficiently. To manage the additional prices for the trade-in and the resale of used product versions challenges retailers as substitution and cannibalization effects have to be taken into account. An unknown customer behavior as well as competition with other merchants regarding both sales and buying back resources further increases the problem\u2019s complexity. Reinforcement learning (RL) algorithms offer the potential to deal with such tasks. However, before being applied in practice, self-learning algorithms need to be tested synthetically to examine whether they and which work in different market scenarios. In the paper, the authors evaluate and compare different state-of-the-art RL algorithms within a recommerce market simulation framework. They find that RL agents outperform rule-based benchmark strategies in duopoly and oligopoly scenarios. Further, the authors investigate the competition between RL agents via self-play and study how performance results are affected if more or less information is observable (cf. state components). Using an ablation study, they test the influence of various model parameters and infer managerial insights. Finally, to be able to apply self-learning agents in practice, the authors show how to calibrate synthetic test environments from observable data to be used for effective pre-training.<\/jats:p>","DOI":"10.1007\/s12599-023-00841-8","type":"journal-article","created":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T09:03:09Z","timestamp":1701075789000},"page":"441-463","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Self-learning Agents for Recommerce Markets"],"prefix":"10.1007","volume":"66","author":[{"given":"Jan","family":"Groeneveld","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Judith","family":"Herrmann","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nikkel","family":"Mollenhauer","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Leonard","family":"Dree\u00dfen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nick","family":"Bessin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Johann Schulze","family":"Tast","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexander","family":"Kastius","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Johannes","family":"Huegle","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rainer","family":"Schlosser","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,11,27]]},"reference":[{"key":"841_CR1","volume-title":"Reinforcement learning and optimal control","author":"DP Bertsekas","year":"2019","unstructured":"Bertsekas DP (2019) Reinforcement learning and optimal control. Athena Scientific, Nashua"},{"issue":"5","key":"841_CR2","first-page":"308","volume":"33","author":"NM Bocken","year":"2016","unstructured":"Bocken NM, de Pauw I, Bakker C, van der Grinten B (2016) Product design and business model strategies for a circular economy. J Indust Prod Eng 33(5):308\u2013320","journal-title":"J Indust Prod Eng"},{"key":"841_CR4","unstructured":"Brockman G, Cheung V, Pettersson L, Schneider J, Schulman J, Tang J, Zaremba W (2016) Openai gym. arXiv preprint arXiv:1606.01540"},{"issue":"107","key":"841_CR5","first-page":"957","volume":"165","author":"F Chen","year":"2022","unstructured":"Chen F, Lu A, Wu H, Dou R, Wang X (2022) Optimal strategies on pricing and resource allocation for cloud services with service guarantees. Comput Ind Eng 165(107):957","journal-title":"Comput Ind Eng"},{"key":"841_CR6","doi-asserted-by":"publisher","first-page":"704","DOI":"10.1111\/poms.12295","volume":"24","author":"M Chen","year":"2015","unstructured":"Chen M, Chen ZL (2015) Recent developments in dynamic pricing research: multiple products, competition, and limited demand information. Prod Oper Manag 24:704\u2013731","journal-title":"Prod Oper Manag"},{"key":"841_CR7","unstructured":"Colony GF (2005) As I.T. goes, so goes Forrester? New York Times https:\/\/www.nytimes.com\/2005\/02\/18\/business\/yourmoney\/as-it-goes-so-goes-forrester.html. Accessed 21 June 2022"},{"key":"841_CR8","first-page":"343","volume":"3","author":"B Commoner","year":"1972","unstructured":"Commoner B (1972) The environmental cost of economic growth. Popul Resour Environ 3:343\u201363","journal-title":"Popul Resour Environ"},{"key":"841_CR3","first-page":"1","volume":"20","author":"AV den Boer","year":"2015","unstructured":"den Boer AV (2015) Dynamic pricing and learning: historical origins, current research, and new directions. Surv Oper Res Manag Sci 20:1\u201318","journal-title":"Surv Oper Res Manag Sci"},{"issue":"3\u20134","key":"841_CR9","doi-asserted-by":"publisher","first-page":"245","DOI":"10.1023\/A:1023427023289","volume":"3","author":"JM DiMicco","year":"2003","unstructured":"DiMicco JM, Maes P, Greenwald A (2003) Learning curve: a simulation-based approach to dynamic pricing. Electron Commer Res 3(3\u20134):245\u2013276","journal-title":"Electron Commer Res"},{"key":"841_CR10","unstructured":"Fujimoto S, van Hoof H, Meger D (2018) Addressing function approximation error in actor-critic methods. CoRR abs\/1802.09477, arXiv:1802.09477"},{"key":"841_CR12","doi-asserted-by":"publisher","first-page":"596","DOI":"10.1057\/s41272-022-00390-x","volume":"21","author":"T Gerpott","year":"2022","unstructured":"Gerpott T, Berends J (2022) Competitive pricing on online markets: a literature review. J Reven Pricing Manag 21:596\u2013622","journal-title":"J Reven Pricing Manag"},{"key":"841_CR13","first-page":"715","volume":"84","author":"J G\u00f6nsch","year":"2014","unstructured":"G\u00f6nsch J (2014) Buying used products for remanufacturing: negotiating or posted pricing. J Bus Econ 84:715\u2013747","journal-title":"J Bus Econ"},{"key":"841_CR14","unstructured":"Haarnoja T, Zhou A, Abbeel P, Levine S (2018) Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. In: ICML 2018, Proceedings of machine learning research, vol 80, pp 1856\u20131865"},{"key":"841_CR15","doi-asserted-by":"publisher","first-page":"207","DOI":"10.1007\/s12599-020-00671-y","volume":"63","author":"F Hawlitschek","year":"2021","unstructured":"Hawlitschek F (2021) The future of waste management. Bus Inf Syst Eng 63:207\u2013211","journal-title":"Bus Inf Syst Eng"},{"key":"841_CR16","unstructured":"Hill A et al (2018) Stable baselines. https:\/\/github.com\/hill-a\/stable-baselines, accessed 21 June 2022"},{"key":"841_CR17","doi-asserted-by":"publisher","first-page":"50","DOI":"10.1057\/s41272-021-00285-3","volume":"21","author":"A Kastius","year":"2022","unstructured":"Kastius A, Schlosser R (2022) Dynamic pricing under competition using reinforcement learning. J Reven Pricing Manag 21:50\u201363","journal-title":"J Reven Pricing Manag"},{"issue":"6","key":"841_CR18","doi-asserted-by":"publisher","first-page":"731","DOI":"10.1016\/S1389-1286(00)00026-8","volume":"32","author":"JO Kephart","year":"2000","unstructured":"Kephart JO, Hanson JE, Greenwald A (2000) Dynamic pricing by software agents. Comput Netw 32(6):731\u2013752","journal-title":"Comput Netw"},{"key":"841_CR19","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1016\/j.resconrec.2017.09.005","volume":"127","author":"J Kirchherr","year":"2017","unstructured":"Kirchherr J, Reike D, Hekkert M (2017) Conceptualizing the circular economy: an analysis of 114 definitions. Res Conserv Recycl 127:221\u2013232","journal-title":"Res Conserv Recycl"},{"key":"841_CR20","doi-asserted-by":"publisher","first-page":"397","DOI":"10.1016\/j.ejor.2019.06.034","volume":"284","author":"R Klein","year":"2020","unstructured":"Klein R, Koch S, Steinhardt C, Strauss A (2020) A review of revenue management: recent generalizations and advances in industry applications. Europ J Oper Res 284:397\u2013412","journal-title":"Europ J Oper Res"},{"key":"841_CR21","doi-asserted-by":"crossref","unstructured":"Maestre R, Duque JR, Rubio A, Ar\u00e9valo J (2018) Reinforcement learning for fair dynamic pricing. In: Arai K, Kapoor S, Bhatia R (eds) Intelligent Systems and Applications - Proceedings of the 2018 Intelligent Systems Conference, IntelliSys 2018, Advances in Intelligent Systems and Computing. Springer, Heidelberg, vol 868, pp 120\u2013135","DOI":"10.1007\/978-3-030-01054-6_8"},{"issue":"7540","key":"841_CR22","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih V et al (2015) Human-level control through deep reinforcement learning. Nature 518(7540):529\u2013533","journal-title":"Nature"},{"key":"841_CR23","unstructured":"Mnih V et al (2016) Asynchronous methods for deep reinforcement learning. In: International conference on machine learning, PMLR, pp 1928\u20131937"},{"key":"841_CR24","unstructured":"Paszke A et al (2019) PyTorch: An imperative style, high-performance deep learning library. In: Wallach H et al (eds) Advances in neural information processing systems 32, pp 8024\u20138035"},{"key":"841_CR25","unstructured":"Rabe L (2020) Reuse und Secondhand in Deutschland. Wuppertal Institut. https:\/\/de.statista.com\/statistik\/daten\/studie\/1248873\/umfrage\/bevorzugter-kanal-fuer-den-verkauf-von-secondhand-produkten-in-deutschland. Accessed 21 June 2022"},{"issue":"3","key":"841_CR26","doi-asserted-by":"publisher","first-page":"1181","DOI":"10.1016\/j.ijforecast.2019.07.001","volume":"36","author":"D Salinas","year":"2020","unstructured":"Salinas D, Flunkert V, Gasthaus J, Januschowski T (2020) DeepAR: probabilistic forecasting with autoregressive recurrent networks. Int J Forecast 36(3):1181\u20131191","journal-title":"Int J Forecast"},{"issue":"2","key":"841_CR27","doi-asserted-by":"publisher","first-page":"239","DOI":"10.1287\/mnsc.1030.0186","volume":"50","author":"RC Savaskan","year":"2004","unstructured":"Savaskan RC, Bhattacharya S, Van Wassenhove LN (2004) Closed-loop supply chain models with product remanufacturing. Manag Sci 50(2):239\u2013252","journal-title":"Manag Sci"},{"key":"841_CR28","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1016\/j.cor.2018.07.011","volume":"100","author":"R Schlosser","year":"2018","unstructured":"Schlosser R, Boissier M (2018) Dealing with the dimensionality curse in dynamic pricing competition: using frequent repricing to compensate imperfect market anticipations. Comput Oper Res 100:26\u201342","journal-title":"Comput Oper Res"},{"key":"841_CR29","doi-asserted-by":"crossref","unstructured":"Schlosser R, Boissier M (2018b) Dynamic pricing under competition on online marketplaces: a data-driven approach. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining pp 705\u2013714","DOI":"10.1145\/3219819.3219833"},{"key":"841_CR30","doi-asserted-by":"publisher","first-page":"451","DOI":"10.1057\/s41272-019-00206-5","volume":"18","author":"R Schlosser","year":"2019","unstructured":"Schlosser R, Richly K (2019) Dynamic pricing under competition with data-driven price anticipations and endogenous reference price effects. J Reven Pricing Manag 18:451\u2013464","journal-title":"J Reven Pricing Manag"},{"key":"841_CR31","doi-asserted-by":"publisher","first-page":"108117","DOI":"10.1016\/j.ijpe.2021.108117","volume":"236","author":"R Schlosser","year":"2021","unstructured":"Schlosser R, Chenavaz R, Dimitrov S (2021) Circular economy: joint dynamic pricing and recycling investments. Intl J Prod Econ 236:108117","journal-title":"Intl J Prod Econ"},{"key":"841_CR32","unstructured":"Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O (2017) Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347"},{"key":"841_CR33","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1057\/s41272-022-00370-1","volume":"21","author":"SAM Shihab","year":"2022","unstructured":"Shihab SAM, Wei P (2022) A deep reinforcement learning approach to seat inventory control for airline revenue management. J Reven Pricing Manag 21:1\u201317","journal-title":"J Reven Pricing Manag"},{"key":"841_CR34","unstructured":"Silver D, Lever G, Heess N, Degris T, Wierstra D, Riedmiller M (2014) Deterministic policy gradient algorithms. In: ICML\u201914, Vol. I, p 387-395"},{"issue":"7676","key":"841_CR35","doi-asserted-by":"publisher","first-page":"354","DOI":"10.1038\/nature24270","volume":"550","author":"D Silver","year":"2017","unstructured":"Silver D et al (2017) Mastering the game of go without human knowledge. Nature 550(7676):354\u2013359","journal-title":"Nature"},{"issue":"7595","key":"841_CR36","doi-asserted-by":"publisher","first-page":"435","DOI":"10.1038\/531435a","volume":"531","author":"WR Stahel","year":"2016","unstructured":"Stahel WR (2016) The circular economy. Nature 531(7595):435\u2013438","journal-title":"Nature"},{"key":"841_CR37","unstructured":"Statista (2020) Wie \u00e4u\u00dfert sich bei ihnen der fokus auf nachhaltige mode beim shopping? Statista Research Department https:\/\/de.statista.com\/statistik\/daten\/studie\/1179997\/umfrage\/umfrage-unter-verbrauchern-zu-nachhaltigemmodekauf-in-deutschland\/. Accessed 21 June 2022"},{"key":"841_CR38","doi-asserted-by":"publisher","first-page":"375","DOI":"10.1016\/j.ejor.2018.01.011","volume":"271","author":"AK Strauss","year":"2018","unstructured":"Strauss AK, Klein R, Steinhardt C (2018) A review of choice-based revenue management: theory and methods. Europ J Oper Res 271:375\u2013387","journal-title":"Europ J Oper Res"},{"key":"841_CR39","unstructured":"Sutton RS, Barto AG (2018) Reinforcement learning - an introduction. In: Adaptive computation and machine learning, 2nd edn. MIT Press, Cambridge"},{"key":"841_CR40","volume-title":"The theory and practice of revenue management","author":"KT Talluri","year":"2006","unstructured":"Talluri KT, Van Ryzin GJ (2006) The theory and practice of revenue management. Springer, Heidelberg"},{"key":"841_CR41","unstructured":"Teh YW et al (2017) Distral: robust multitask reinforcement learning. In: Advances in neural information processing systems 30: Annual conference on neural information processing systems 2017, pp 4496\u20134506"},{"key":"841_CR42","doi-asserted-by":"publisher","first-page":"385","DOI":"10.1007\/s12599-020-00657-w","volume":"62","author":"O Thomas","year":"2020","unstructured":"Thomas O et al (2020) Global crises and the role of BISE. Bus Inf Syst Eng 62:385\u2013396","journal-title":"Bus Inf Syst Eng"},{"issue":"108","key":"841_CR43","first-page":"567","volume":"172","author":"Y Tsao","year":"2022","unstructured":"Tsao Y, Beyene TD, Thanh V, Gebeyehu SG (2022) Power distribution network design considering dynamic and differential pricing, buy-back, and carbon trading. Comput Ind Eng 172(108):567","journal-title":"Comput Ind Eng"},{"issue":"102","key":"841_CR44","first-page":"829","volume":"121","author":"B Turan","year":"2020","unstructured":"Turan B, Pedarsani R, Alizadeh M (2020) Dynamic pricing and fleet management for electric autonomous mobility on demand systems. Transp Res Part C: Emerg Technol 121(102):829","journal-title":"Transp Res Part C: Emerg Technol"},{"key":"841_CR11","doi-asserted-by":"publisher","first-page":"185","DOI":"10.1057\/s41272-018-00164-4","volume":"18","author":"R van de Geer","year":"2019","unstructured":"van de Geer R, den Boer A, Bayliss C et al (2019) Dynamic pricing and learning with competition: insights from the dynamic pricing challenge at the 2017 INFORMS RM & pricing conference. J Reven Pricing Manag 18:185\u2013203","journal-title":"J Reven Pricing Manag"},{"key":"841_CR45","doi-asserted-by":"publisher","first-page":"325","DOI":"10.1007\/s12599-021-00705-z","volume":"63","author":"C Weinhardt","year":"2021","unstructured":"Weinhardt C et al (2021) Welcome to economies in IS! Bus Inf Syst Eng 63:325\u2013328","journal-title":"Bus Inf Syst Eng"},{"issue":"108","key":"841_CR46","first-page":"290","volume":"169","author":"D Wen","year":"2022","unstructured":"Wen D, Xiao T, Dastani M (2022) Pricing strategy and collection rate for a supply chain considering environmental responsibility behaviors and rationality degree. Comput Ind Eng 169(108):290","journal-title":"Comput Ind Eng"},{"issue":"108","key":"841_CR47","first-page":"440","volume":"171","author":"Y Yang","year":"2022","unstructured":"Yang Y, Chu W, Wu C (2022) Learning customer preferences and dynamic pricing for perishable products. Comput Ind Eng 171(108):440","journal-title":"Comput Ind Eng"},{"key":"841_CR48","unstructured":"Zhu Z, Lin K, Zhou J (2020) Transfer learning in deep reinforcement learning: a survey. CoRR abs\/2009.07888, arXiv:2009.07888"}],"container-title":["Business &amp; Information Systems Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12599-023-00841-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s12599-023-00841-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12599-023-00841-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,31]],"date-time":"2024-08-31T09:21:39Z","timestamp":1725096099000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s12599-023-00841-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,27]]},"references-count":48,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,8]]}},"alternative-id":["841"],"URL":"https:\/\/doi.org\/10.1007\/s12599-023-00841-8","relation":{},"ISSN":["2363-7005","1867-0202"],"issn-type":[{"value":"2363-7005","type":"print"},{"value":"1867-0202","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,27]]},"assertion":[{"value":"9 November 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 September 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 November 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}