{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T21:37:38Z","timestamp":1784842658733,"version":"3.55.0"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2020,3,18]],"date-time":"2020-03-18T00:00:00Z","timestamp":1584489600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2020,3,18]]},"abstract":"<jats:p>With the recent proliferation of mobile health technologies, health scientists are increasingly interested in developing just-in-time adaptive interventions (JITAIs), typically delivered via notifications on mobile devices and designed to help users prevent negative health outcomes and to promote the adoption and maintenance of healthy behaviors. A JITAI involves a sequence of decision rules (i.e., treatment policies) that take the user's current context as input and specify whether and what type of intervention should be provided at the moment. In this work, we describe a reinforcement learning (RL) algorithm that continuously learns and improves the treatment policy embedded in the JITAI as data is being collected from the user. This work is motivated by our collaboration on designing an RL algorithm for HeartSteps V2 based on data collected HeartSteps V1. HeartSteps is a physical activity mobile health application. The RL algorithm developed in this work is being used in HeartSteps V2 to decide, five times per day, whether to deliver a context-tailored activity suggestion.<\/jats:p>","DOI":"10.1145\/3381007","type":"journal-article","created":{"date-parts":[[2020,3,18]],"date-time":"2020-03-18T18:54:31Z","timestamp":1584557671000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":121,"title":["Personalized HeartSteps"],"prefix":"10.1145","volume":"4","author":[{"given":"Peng","family":"Liao","sequence":"first","affiliation":[{"name":"University of Michigan, Ann Arbor, MI"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kristjan","family":"Greenewald","sequence":"additional","affiliation":[{"name":"IBM Research, Cambridge, MA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Predrag","family":"Klasnja","sequence":"additional","affiliation":[{"name":"University of Michigan, Ann Arbor, MI"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Susan","family":"Murphy","sequence":"additional","affiliation":[{"name":"Harvard University, Cambridge, MA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,3,18]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"International Conference on Machine Learning. 127--135","author":"Agrawal Shipra","year":"2013"},{"key":"e_1_2_1_2_1","volume-title":"Lucas Lehnert, and Michael L Littman.","author":"Arumugam Dilip","year":"2018"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCST.2016.2580661"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.2017.1305274"},{"key":"e_1_2_1_5_1","unstructured":"Olivier Chapelle and Lihong Li. 2011. An empirical evaluation of thompson sampling. In Advances in Neural Information Processing Systems. 2249--2257.  Olivier Chapelle and Lihong Li. 2011. An empirical evaluation of thompson sampling. In Advances in Neural Information Processing Systems. 2249--2257."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1740-9713.2015.00863.x"},{"key":"e_1_2_1_7_1","volume-title":"Estimation considerations in contextual bandits. arXiv preprint arXiv:1711.07077","author":"Dimakopoulou Maria","year":"2017"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1136\/jnnp.35.2.234"},{"key":"e_1_2_1_9_1","volume-title":"Neural Information Processing Systems 2013 Workshop on Bayesian Optimization","author":"Fonteneau Rapha\u00ebl","year":"2013"},{"key":"e_1_2_1_10_1","volume-title":"Can the artificial intelligence technique of reinforcement learning use continuously-monitored digital data to optimize treatment for weight loss? Journal of behavioral medicine","author":"Forman Evan M","year":"2018"},{"key":"e_1_2_1_11_1","volume-title":"Neural Information Processing Systems 2015 Workshop on Deep Reinforcement Learning","author":"Fran\u00e7ois-Lavet Vincent","year":"2015"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/3298023.3298115"},{"key":"e_1_2_1_13_1","unstructured":"Kristjan Greenewald Ambuj Tewari Susan Murphy and Predag Klasnja. 2017. Action centered contextual bandits. In Advances in Neural Information Processing Systems. 5977--5985.  Kristjan Greenewald Ambuj Tewari Susan Murphy and Predag Klasnja. 2017. Action centered contextual bandits. In Advances in Neural Information Processing Systems. 5977--5985."},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 1181--1189","author":"Jiang Nan","year":"2015"},{"key":"e_1_2_1_15_1","volume-title":"Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. In International Conference on Machine Learning. 652--661","author":"Jiang Nan","year":"2016"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-34106-9_18"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1409635.1409656"},{"key":"e_1_2_1_18_1","doi-asserted-by":"crossref","unstructured":"P. Klasnja E.B. Hekler S. Shiffman A. Boruvka D. Almirall A. Tewari and S.A. Murphy. 2015. Micro-randomized trials: An experimental design for developing just-in-time adaptive interventions. Health Psychology 34 S (2015) 1220.  P. Klasnja E.B. Hekler S. Shiffman A. Boruvka D. Almirall A. Tewari and S.A. Murphy. 2015. Micro-randomized trials: An experimental design for developing just-in-time adaptive interventions. Health Psychology 34 S (2015) 1220.","DOI":"10.1037\/hea0000305"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1093\/abm\/kay067"},{"key":"e_1_2_1_20_1","volume-title":"Zhiwei Steven Wu, and Vasilis Syrgkanis","author":"Krishnamurthy Akshay","year":"2018"},{"key":"e_1_2_1_21_1","volume-title":"Thirty-Second AAAI Conference on Artificial Intelligence.","author":"Lehnert Lucas","year":"2018"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772758"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/73.1.13"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287057"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1002\/sim.6847"},{"key":"e_1_2_1_26_1","volume-title":"Development of a control-oriented model of social cognitive theory for optimized mHealth behavioral interventions","author":"Martin Cesar A","year":"2018"},{"key":"e_1_2_1_27_1","volume-title":"Non-stationary bandits with habituation and recovery dynamics. Operations Research","author":"Mintz Yonatan","year":"2019"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12160-016-9830-8"},{"key":"e_1_2_1_29_1","unstructured":"Ian Osband Daniel Russo and Benjamin Van Roy. 2013. (More) efficient reinforcement learning via posterior sampling. In Advances in Neural Information Processing Systems. 3003--3011.  Ian Osband Daniel Russo and Benjamin Van Roy. 2013. (More) efficient reinforcement learning via posterior sampling. In Advances in Neural Information Processing Systems. 3003--3011."},{"key":"e_1_2_1_30_1","volume-title":"On optimistic versus randomized exploration in reinforcement learning. arXiv preprint arXiv:1706.04241","author":"Osband Ian","year":"2017"},{"key":"e_1_2_1_31_1","volume-title":"International Conference on Machine Learning. 2701--2710","author":"Osband Ian","year":"2017"},{"key":"e_1_2_1_32_1","unstructured":"Yi Ouyang Mukul Gagrani Ashutosh Nayyar and Rahul Jain. 2017. Learning unknown markov decision processes: A thompson sampling approach. In Advances in Neural Information Processing Systems. 1333--1342.  Yi Ouyang Mukul Gagrani Ashutosh Nayyar and Rahul Jain. 2017. Learning unknown markov decision processes: A thompson sampling approach. In Advances in Neural Information Processing Systems. 1333--1342."},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the 8th International Conference on Pervasive Computing Technologies for Healthcare","author":"Paredes Pablo"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2750858.2805840"},{"key":"e_1_2_1_35_1","volume-title":"Optimization of Behavioral, Biobehavioral, and Biomedical Interventions","author":"Rivera Daniel E"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1287\/moor.2014.0650"},{"key":"e_1_2_1_37_1","volume-title":"Abbas Kazerouni, Ian Osband, Zheng Wen, et al.","author":"Russo Daniel J","year":"2018"},{"key":"e_1_2_1_38_1","volume-title":"Reinforcement learning: An introduction","author":"Sutton Richard S"},{"key":"e_1_2_1_39_1","volume-title":"Mobile Health","author":"Tewari Ambuj"},{"key":"e_1_2_1_40_1","volume-title":"International Conference on Machine Learning. 2139--2148","author":"Thomas Philip","year":"2016"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.2196\/jmir.7994"},{"key":"e_1_2_1_42_1","volume-title":"Companion Proceedings of the 23rd International on Intelligent User Interfaces: 2nd Workshop on Theory-Informed User Modeling for Tailoring and Personalizing Interfaces (HUMANIZE).","author":"Zhou Mo","year":"2018"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3381007","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3381007","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:44:58Z","timestamp":1750203898000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3381007"}},"subtitle":["A Reinforcement Learning Algorithm for Optimizing Physical Activity"],"short-title":[],"issued":{"date-parts":[[2020,3,18]]},"references-count":42,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,3,18]]}},"alternative-id":["10.1145\/3381007"],"URL":"https:\/\/doi.org\/10.1145\/3381007","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,3,18]]},"assertion":[{"value":"2020-03-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}