{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T08:51:19Z","timestamp":1785919879043,"version":"3.56.0"},"reference-count":85,"publisher":"Institute for Operations Research and the Management Sciences (INFORMS)","issue":"8","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Management Science"],"published-print":{"date-parts":[[2026,8]]},"abstract":"<jats:p>We formulate a collaborative learning and decision-making problem involving contextual information. In current business practices, pricing and recommendation decisions often are made jointly by multiple teams in sequence. The decision-making processes for different teams can be controlled by either a centralized or decentralized planner. We propose a simple collaboration framework that integrates the learning about decision making in an unknown environment. The main challenge in a decentralized framework is that the decision-making process in other teams is unknown, but the subsequent decisions are mutually dependent. From a practical concern about high exploration costs and implementation complexity, we propose a simple greedy algorithm for centralized planners and a \u201cgreedy\u201d + \u201cweighted sampling\u201d (GWS) algorithm for both centralized and decentralized planners to balance the learning and earning. We show that the exploration-free greedy algorithm can achieve the optimal rate when context diversity holds. The GWS algorithm works effectively for either centralized or decentralized planners under a much weaker condition, which we call context variation. Furthermore, we extend our framework to the multiproduct pricing and ranking problem and study the model misspecification issue. We validate our results using simulations on synthetic and real data. Numerical studies show the superior performance of the two proposed frameworks for different types of planners.<\/jats:p>\n                  <jats:p>This paper was accepted by J. George Shanthikumar, data science.<\/jats:p>\n                  <jats:p>Supplemental Material: The online appendices and data files are available at https:\/\/doi.org\/10.1287\/mnsc.2023.00320 .<\/jats:p>","DOI":"10.1287\/mnsc.2023.00320","type":"journal-article","created":{"date-parts":[[2025,11,11]],"date-time":"2025-11-11T14:36:51Z","timestamp":1762871811000},"page":"6472-6492","source":"Crossref","is-referenced-by-count":1,"title":["Collaborative Learning and Decision Making on Pricing and Recommendation: A Simple Framework for Planning"],"prefix":"10.1287","volume":"72","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9235-1411","authenticated-orcid":false,"given":"Junyu","family":"Cao","sequence":"first","affiliation":[{"name":"McCombs School of Business, The University of Texas at Austin, Austin, Texas 78712"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"109","reference":[{"key":"B1","first-page":"236","volume":"92","author":"Abbasi-Yadkori Y","year":"2009","journal-title":"Proc. COLT Workshop On-Line Learn. Limited Feedback"},{"key":"B2","first-page":"2312","author":"Abbasi-Yadkori Y","year":"2011","journal-title":"Adv. Neural Inform. Processing Systems"},{"key":"B3","first-page":"19","volume-title":"Conf. Artificial Intelligence Statist.","author":"Agarwal A","year":"2012"},{"key":"B4","unstructured":"Agarwal A, Hsu D, Kale S, Langford J, Li L, Schapire R (2014) Taming the monster: A fast and simple algorithm for contextual bandits.\n                      Proc. Internat. Conf. Machine Learn.\n                      (PMLR, New York), 1638\u20131646."},{"key":"B5","unstructured":"Agrawal S, Goyal N (2013) Thompson sampling for contextual bandits with linear payoffs.\n                      Proc. Internat. Conf. Machine Learn.\n                      (PMLR, New York), 127\u2013135."},{"key":"B6","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2015.1459"},{"key":"B7","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2020.3680"},{"key":"B8","unstructured":"Banerjee S, Sinclair SR, Tambe M, Xu L, Yu CL (2022) Artificial replay: A meta-algorithm for harnessing historical data in bandits. Preprint, submitted September 30, https:\/\/arxiv.org\/abs\/2210.00025."},{"key":"B9","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2020.3605"},{"key":"B10","first-page":"1713","volume":"33","author":"Bayati M","year":"2020","journal-title":"Adv. Neural Inform. Processing Systems"},{"key":"B11","doi-asserted-by":"publisher","DOI":"10.1287\/opre.1080.0640"},{"issue":"1","key":"B12","first-page":"5928","volume":"22","author":"Bietti A","year":"2021","journal-title":"J. Machine Learn. Res."},{"key":"B13","author":"Bird S","year":"2016","journal-title":"Proc. Workshop Fairness Accountability Transparency Machine Learn."},{"key":"B14","doi-asserted-by":"publisher","DOI":"10.1287\/opre.1120.1057"},{"key":"B15","unstructured":"Cao J, Gao R (2021) Contextual decision-making under parametric uncertainty and data-driven optimistic optimization. Preprint, submitted October 14, https:\/\/optimization-online.org\/2021\/10\/8634\/."},{"key":"B16","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2022.01130"},{"key":"B17","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2023.4940"},{"key":"B18","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2020.2016"},{"key":"B19","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2019.1857"},{"key":"B20","first-page":"784","volume":"1","author":"Chen X","year":"2012","journal-title":"The Oxford Handbook of Pricing Management"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2022.2347"},{"key":"B22","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2023.4859"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2020.3772"},{"key":"B24","unstructured":"Chu W, Li L, Reyzin L, Schapire R (2011) Contextual bandits with linear payoff functions.\n                      Proc. 14th Internat. Conf. Artificial Intelligence Statist\n                      . 208\u2013214."},{"key":"B25","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2022.4317"},{"key":"B26","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2019.3485"},{"key":"B27","unstructured":"Dani V, Hayes TP, Kakade SM (2008) Stochastic linear optimization under bandit feedback.\n                      Proc. Conf. Learn. Theory."},{"key":"B28","author":"den Boer AV","year":"2022","journal-title":"Management Sci. 68(10):7112\u20137130."},{"key":"B29","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2013.1788"},{"key":"B30","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2015.1397"},{"key":"B31","doi-asserted-by":"crossref","unstructured":"Dimakopoulou M, Zhou Z, Athey S, Imbens G (2019) Balanced linear contextual bandits.\n                      Proc. AAAI Conf. Artificial Intelligence\n                      , vol. 33, 3445\u20133453.","DOI":"10.1609\/aaai.v33i01.33013445"},{"key":"B32","first-page":"6003","volume":"33","author":"Dubey A","year":"2020","journal-title":"Adv. Neural Inform. Processing Systems"},{"key":"B33","author":"El Housni O","year":"2022","journal-title":"Oper. Res. 71(4):1197\u20131215."},{"key":"B34","doi-asserted-by":"publisher","DOI":"10.1287\/msom.2018.0756"},{"key":"B35","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2018.1755"},{"key":"B36","unstructured":"Filippi S, Cappe O, Garivier A, Szepesv\u00e1ri C (2010) Parametric bandits: The generalized linear case. Lafferty JD, Williams CKI, Shawe-Taylor J, Zemel RS, Culotta A, eds.\n                      Adv. Neural Inform. Processing Systems\n                      , vol. 23 (Curran Associates, Inc., Red Hook, NY), 586\u2013594."},{"key":"B37","unstructured":"Foster D, Rakhlin A (2020) Beyond UCB: Optimal and efficient contextual bandits with regression oracles.\n                      Proc. Internat. Conf. Machine Learn.\n                      (PMLR, New York), 3199\u20133210."},{"key":"B38","unstructured":"Foster D, Agarwal A, Dudik M, Luo H, Schapire R (2018) Practical contextual bandits with regression oracles.\n                      Proc. Internat. Conf. Machine Learn\n                      . (PMLR, New York), 1539\u20131548."},{"key":"B39","doi-asserted-by":"publisher","DOI":"10.1111\/j.2517-6161.1979.tb01068.x"},{"key":"B40","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2015.1439"},{"key":"B41","first-page":"27057","volume":"34","author":"Huang R","year":"2021","journal-title":"Adv. Neural Inform. Processing Systems"},{"key":"B42","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2016.2491"},{"issue":"1","key":"B43","first-page":"315","volume":"20","author":"Javanmard A","year":"2019","journal-title":"J. Machine Learn. Res."},{"key":"B44","unstructured":"Jun KS, Willett R, Wright S, Nowak R (2019) Bilinear bandits with low-rank structure.\n                      Proc. Internat. Conf. Machine Learn\n                      . (PMLR, New York), 3163\u20133172."},{"key":"B45","doi-asserted-by":"publisher","DOI":"10.1561\/2200000083"},{"key":"B46","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2019.1948"},{"key":"B47","unstructured":"Kannan S, Morgenstern JH, Roth A, Waggoner B, Wu ZS (2018) A smoothed analysis of the greedy algorithm for the linear contextual bandit problem.\n                      Adv. Neural Inform. Processing Systems\n                      , vol. 31 (Curran Associates Inc., Red Hook, NY), 2227\u20132236."},{"key":"B48","unstructured":"Kao H, Wei CY, Subramanian V (2022) Decentralized cooperative reinforcement learning with hierarchical information structure.\n                      Proc. Internat. Conf. Algorithmic Learn. Theory\n                      (PMLR, New York), 573\u2013605."},{"key":"B49","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2014.1294"},{"key":"B50","doi-asserted-by":"publisher","DOI":"10.1287\/moor.2016.0807"},{"key":"B51","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2017.1713"},{"key":"B52","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2022.0112"},{"key":"B53","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.1110.1402"},{"issue":"1","key":"B54","first-page":"5402","volume":"21","author":"Krishnamurthy A","year":"2020","journal-title":"J. Machine Learn. Res."},{"key":"B55","unstructured":"Kveton B, Zaheer M, Szepesvari C, Li L, Ghavamzadeh M, Boutilier C (2020) Randomized exploration in generalized linear bandits.\n                      Proc. Internat. Conf. Artificial Intelligence Statist\n                      . 2066\u20132076."},{"key":"B56","doi-asserted-by":"publisher","DOI":"10.1017\/9781108571401"},{"key":"B57","unstructured":"Li L, Lu Y, Zhou D (2017) Provably optimal algorithms for generalized linear contextual bandits.\n                      Proc. 34th Internat. Conf. Machine Learn.\n                      , vol. 70, 2071\u20132080."},{"key":"B58","doi-asserted-by":"crossref","unstructured":"Li L, Chu W, Langford J, Schapire RE (2010) A contextual-bandit approach to personalized news article recommendation.\n                      Proc. 19th Internat. Conf. World Wide Web\n                      , 61\u2013670.","DOI":"10.1145\/1772690.1772758"},{"key":"B59","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2020.2975749"},{"key":"B60","doi-asserted-by":"publisher","DOI":"10.1287\/mksc.2022.1403"},{"key":"B61","unstructured":"Lu Y, Meisami A, Tewari A (2021) Low-rank generalized linear bandit problems.\n                      Proc. Internat. Conf. Artificial Intelligence Statist\n                      . (PMLR, New York), 460\u2013468."},{"key":"B62","first-page":"1273","volume-title":"Artificial Intelligence and Statistics","author":"McMahan B","year":"2017"},{"key":"B63","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2009.2031725"},{"issue":"2","key":"B64","first-page":"525","volume":"23","author":"Miao S","year":"2021","journal-title":"Manufacturing Service Oper. Management"},{"key":"B65","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2018.3194"},{"key":"B66","doi-asserted-by":"publisher","DOI":"10.52202\/068431-0250"},{"key":"B67","doi-asserted-by":"crossref","unstructured":"Qiang S, Bayati M (2016) Dynamic pricing with demand covariates. Preprint, submitted April 25, https:\/\/arxiv.org\/abs\/1604.07463.","DOI":"10.2139\/ssrn.2765257"},{"key":"B68","doi-asserted-by":"publisher","DOI":"10.1287\/moor.1100.0446"},{"key":"B69","doi-asserted-by":"publisher","DOI":"10.1287\/moor.2014.0650"},{"issue":"1","key":"B70","first-page":"2442","volume":"17","author":"Russo D","year":"2016","journal-title":"J. Machine Learn. Res."},{"key":"B71","doi-asserted-by":"publisher","DOI":"10.1287\/msom.2020.0900"},{"key":"B72","doi-asserted-by":"crossref","unstructured":"Shi C, Shen C (2021) Federated multi-armed bandits.\n                      Proc. AAAI Conf. Artificial Intelligence\n                      , vol. 35, 9603\u20139611.","DOI":"10.1609\/aaai.v35i11.17156"},{"key":"B73","author":"Shin D","year":"2022","journal-title":"Management Sci. 69(2):824\u2013845."},{"key":"B74","author":"Simchi-Levi D","year":"2021","journal-title":"Math. Oper. Res. 47(3):1904\u20131931."},{"key":"B75","doi-asserted-by":"publisher","DOI":"10.1561\/2200000068"},{"key":"B76","doi-asserted-by":"crossref","unstructured":"Sun H, Li X, Teo C-P (2025) Partition and prosper: Design and pricing of single bundle.\n                      Oper. Res.\n                      73(4):1983\u20132001.","DOI":"10.1287\/opre.2022.0465"},{"key":"B77","unstructured":"Valko M, Korda N, Munos R, Flaounas I, Cristianini N (2013) Finite-time analysis of kernelised contextual bandits. Preprint, submitted September 26, https:\/\/arxiv.org\/abs\/1309.6869."},{"key":"B78","doi-asserted-by":"publisher","DOI":"10.1016\/j.orl.2012.08.003"},{"key":"B79","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2005.844079"},{"key":"B80","doi-asserted-by":"publisher","DOI":"10.1287\/opre.2021.0802"},{"key":"B81","doi-asserted-by":"publisher","DOI":"10.1287\/mnsc.2022.01167"},{"key":"B82","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1979.10481033"},{"key":"B83","unstructured":"Xu Y, Zeevi A (2020) Upper counterfactual confidence bounds: A new optimism principle for contextual bandits. Preprint, submitted July 15, https:\/\/arxiv.org\/abs\/2007.07876."},{"key":"B84","doi-asserted-by":"publisher","DOI":"10.1007\/s11590-014-0795-x"},{"key":"B85","unstructured":"Zhou D, Li L, Gu Q (2020) Neural contextual bandits with UCB-based exploration.\n                      Internat. Conf. Machine Learn\n                      . (PMLR, New York), 11492\u201311502."}],"container-title":["Management Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/pubsonline.informs.org\/doi\/pdf\/10.1287\/mnsc.2023.00320","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T08:32:42Z","timestamp":1785918762000},"score":1,"resource":{"primary":{"URL":"https:\/\/pubsonline.informs.org\/doi\/10.1287\/mnsc.2023.00320"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8]]},"references-count":85,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2026,8]]}},"alternative-id":["10.1287\/mnsc.2023.00320"],"URL":"https:\/\/doi.org\/10.1287\/mnsc.2023.00320","relation":{},"ISSN":["0025-1909","1526-5501"],"issn-type":[{"value":"0025-1909","type":"print"},{"value":"1526-5501","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,8]]}}}