{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T13:36:06Z","timestamp":1780407366612,"version":"3.54.1"},"reference-count":67,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2021,8,11]],"date-time":"2021-08-11T00:00:00Z","timestamp":1628640000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Australian Research Council Discovery Early Career Researcher","award":["DE180100315"],"award-info":[{"award-number":["DE180100315"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Comput.-Hum. Interact."],"published-print":{"date-parts":[[2021,8,31]]},"abstract":"<jats:p>Annotation is an effective reading strategy people often undertake while interacting with digital text. It involves highlighting pieces of text and making notes about them. Annotating while reading in a desktop environment is considered trivial but, in a mobile setting where people read while hand-holding devices, the task of highlighting and typing notes on a mobile display is challenging. In this article, we introduce GAVIN, a gaze-assisted voice note-taking application, which enables readers to seamlessly take voice notes on digital documents by implicitly anchoring them to text passages. We first conducted a contextual enquiry focusing on participants\u2019 note-taking practices on digital documents. Using these findings, we propose a method which leverages eye-tracking and machine learning techniques to annotate voice notes with reference text passages. To evaluate our approach, we recruited 32 participants performing voice note-taking. Following, we trained a classifier on the data collected to predict text passage where participants made voice notes. Lastly, we employed the classifier to built GAVIN and conducted a user study to demonstrate the feasibility of the system. This research demonstrates the feasibility of using gaze as a resource for implicit anchoring of voice notes, enabling the design of systems that allow users to record voice notes with minimal effort and high accuracy.<\/jats:p>","DOI":"10.1145\/3453988","type":"journal-article","created":{"date-parts":[[2021,8,12]],"date-time":"2021-08-12T00:39:36Z","timestamp":1628728776000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["GAVIN"],"prefix":"10.1145","volume":"28","author":[{"given":"Anam Ahmad","family":"Khan","sequence":"first","affiliation":[{"name":"The University of Melbourne, Melbourne, Victoria, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joshua","family":"Newn","sequence":"additional","affiliation":[{"name":"The University of Melbourne, Melbourne, Victoria, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ryan M.","family":"Kelly","sequence":"additional","affiliation":[{"name":"The University of Melbourne, Melbourne, Victoria, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Namrata","family":"Srivastava","sequence":"additional","affiliation":[{"name":"The University of Melbourne, Melbourne, Victoria, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"James","family":"Bailey","sequence":"additional","affiliation":[{"name":"The University of Melbourne, Melbourne, Victoria, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eduardo","family":"Velloso","sequence":"additional","affiliation":[{"name":"The University of Melbourne, Melbourne, Victoria, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,8,11]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3351227"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2820783.2820885"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the 8th IEEE International Conference on Cloud Computing (CLOUD\u201915)","author":"Aransay Ignacio","year":"2015","unstructured":"Ignacio Aransay , M. Z. Sancho , P. A. Garc\u0131a , and J. M. M. Fern\u00e1ndez . 2015 . Self-Organizing maps for detecting abnormal thermal behavior in data centers . In Proceedings of the 8th IEEE International Conference on Cloud Computing (CLOUD\u201915) . 138\u2013145. Ignacio Aransay, M. Z. Sancho, P. A. Garc\u0131a, and J. M. M. Fern\u00e1ndez. 2015. Self-Organizing maps for detecting abnormal thermal behavior in data centers. In Proceedings of the 8th IEEE International Conference on Cloud Computing (CLOUD\u201915). 138\u2013145."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2016.7899814"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the 9th Chais Conference for the Study of Innovation and Learning Technologies. 28\u201335","author":"Ben-Yehudah Gal","year":"2014","unstructured":"Gal Ben-Yehudah and Yoram Eshet-Alkalai . 2014 . The influence of text annotation tools on print and digital reading comprehension . In Proceedings of the 9th Chais Conference for the Study of Innovation and Learning Technologies. 28\u201335 . Gal Ben-Yehudah and Yoram Eshet-Alkalai. 2014. The influence of text annotation tools on print and digital reading comprehension. In Proceedings of the 9th Chais Conference for the Study of Innovation and Learning Technologies. 28\u201335."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1186\/s12913-019-4185-z"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1080\/14640748308402115"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/291224.291229"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2168556.2168593"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1753326.1753417"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1358628.1358805"},{"key":"e_1_2_1_12_1","first-page":"122","article-title":"Audio feedback versus written feedback: Instructors\u2019 and students\u2019 perspectives","volume":"10","author":"Cavanaugh Andrew J.","year":"2014","unstructured":"Andrew J. Cavanaugh and Liyan Song . 2014 . Audio feedback versus written feedback: Instructors\u2019 and students\u2019 perspectives . Journal of Online Learning and Teaching 10 , 1 (2014), 122 . Andrew J. Cavanaugh and Liyan Song. 2014. Audio feedback versus written feedback: Instructors\u2019 and students\u2019 perspectives. Journal of Online Learning and Teaching 10, 1 (2014), 122.","journal-title":"Journal of Online Learning and Teaching"},{"key":"e_1_2_1_13_1","volume-title":"Raghuwanshi","author":"Chaudhari P.","year":"2014","unstructured":"P. Chaudhari , Dipti P. Rana , Rupa G. Mehta , Narendra J. Mistry , and Mukesh M . Raghuwanshi . 2014 . Discretization of temporal data: a survey. arXiv preprint arXiv:1402.4283 (2014). P. Chaudhari, Dipti P. Rana, Rupa G. Mehta, Narendra J. Mistry, and Mukesh M. Raghuwanshi. 2014. Discretization of temporal data: a survey. arXiv preprint arXiv:1402.4283 (2014)."},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the Workshop on Visualization for the Digital Humanities.","author":"Cheema Muhammad Faisal","year":"2016","unstructured":"Muhammad Faisal Cheema , Stefan J\u00e4nicke , and Gerik Scheuermann . 2016 . AnnotateVis : Combining traditional close reading with visual text analysis . In Proceedings of the Workshop on Visualization for the Digital Humanities. Muhammad Faisal Cheema, Stefan J\u00e4nicke, and Gerik Scheuermann. 2016. AnnotateVis : Combining traditional close reading with visual text analysis. In Proceedings of the Workshop on Visualization for the Digital Humanities."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2702123.2702271"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1177\/1469787417735614"},{"key":"e_1_2_1_17_1","first-page":"152","article-title":"Can you hear me now? Providing feedback using audio commenting technology","volume":"29","author":"Dagen Allison Swan","year":"2008","unstructured":"Allison Swan Dagen , C. Matter , Steven Rinehart , and Philip Ice . 2008 . Can you hear me now? Providing feedback using audio commenting technology . College Reading Association Yearbook 29 (2008), 152 \u2013 166 . Allison Swan Dagen, C. Matter, Steven Rinehart, and Philip Ice. 2008. Can you hear me now? Providing feedback using audio commenting technology. College Reading Association Yearbook 29 (2008), 152\u2013166.","journal-title":"College Reading Association Yearbook"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2168556.2168575"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4419-1428-6_849"},{"key":"e_1_2_1_20_1","unstructured":"M. Filetti H. R. Tavakoli N. Ravaja and G. Jacucci. 2019. PeyeDF: An eye-tracking application for reading and self-indexing research. arXiv e-prints (April 2019). arXiv:cs.HC\/1904.12152.  M. Filetti H. R. Tavakoli N. Ravaja and G. Jacucci. 2019. PeyeDF: An eye-tracking application for reading and self-indexing research. arXiv e-prints (April 2019). arXiv:cs.HC\/1904.12152."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1080\/09658211.2017.1383434"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1030397.1030414"},{"key":"e_1_2_1_23_1","unstructured":"J. Grabowski. 2005. Speaking writing and memory span performance: replicating the Bourdin and Fayol results on cognitive load in German children and adults. Studies in Writing: PREPUBLICATIONS ARCHIVES. (2005). DOI:https:\/\/sigwriting. publication-archive.com\/publication\/1\/163.  J. Grabowski. 2005. Speaking writing and memory span performance: replicating the Bourdin and Fayol results on cognitive load in German children and adults. Studies in Writing: PREPUBLICATIONS ARCHIVES. (2005). DOI:https:\/\/sigwriting. publication-archive.com\/publication\/1\/163."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-04962-0_53"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2663204.2663277"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the CSUN 1997 Conference. 1\u20137.","author":"Hatfield Franz","unstructured":"Franz Hatfield and Eric A. Jenkins . 1997. An interface integrating eye gaze and voice recognition for hands-free computer access . In Proceedings of the CSUN 1997 Conference. 1\u20137. Franz Hatfield and Eric A. Jenkins. 1997. An interface integrating eye gaze and voice recognition for hands-free computer access. In Proceedings of the CSUN 1997 Conference. 1\u20137."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1753326.1753413"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.24059\/olj.v11i2.1724"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.learninstruc.2020.101344"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/97243.97246"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.5944\/openpraxis.10.4.909"},{"key":"e_1_2_1_32_1","volume-title":"Using voice note-taking to promote learners","author":"Khan Anam Ahmad","year":"2012","unstructured":"Anam Ahmad Khan , Sadia Nawaz , Joshua Newn , Jason M. Lodge , James Bailey , and Eduardo Velloso . 2020. Using voice note-taking to promote learners \u2019 conceptual understanding. arXiv: 2012 .02927. Anam Ahmad Khan, Sadia Nawaz, Joshua Newn, Jason M. Lodge, James Bailey, and Eduardo Velloso. 2020. Using voice note-taking to promote learners\u2019 conceptual understanding. arXiv:2012.02927."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2013.339"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISM.2017.116"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.3758\/s13423-011-0168-8"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/1542084.1542097"},{"key":"e_1_2_1_38_1","volume-title":"Methods of hierarchical clustering. CoRR abs\/1105.0121","author":"Murtagh Fionn","year":"2011","unstructured":"Fionn Murtagh and Pedro Contreras . 2011. Methods of hierarchical clustering. CoRR abs\/1105.0121 ( 2011 ). arXiv:1105.0121. http:\/\/arxiv.org\/abs\/1105.0121. Fionn Murtagh and Pedro Contreras. 2011. Methods of hierarchical clustering. CoRR abs\/1105.0121 (2011). arXiv:1105.0121. http:\/\/arxiv.org\/abs\/1105.0121."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2638728.2638783"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.3233\/978-1-61499-512-8-531"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1029632.1029663"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267242.3267288"},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the Eurographics Ireland Workshop.","author":"O\u2019Donovan J.","unstructured":"J. O\u2019Donovan , J. Ward , S. Hodgins , and V. Sundstedt . 2009. Rabbit run: Gaze and voice based game interaction . In Proceedings of the Eurographics Ireland Workshop. J. O\u2019Donovan, J. Ward, S. Hodgins, and V. Sundstedt. 2009. Rabbit run: Gaze and voice based game interaction. In Proceedings of the Eurographics Ireland Workshop."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1080\/07294360.2013.841651"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.2307\/4128941"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/2745555.2746644"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0364-0213(02)00078-2"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/355017.355028"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3129340"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/332040.332445"},{"key":"e_1_2_1_51_1","first-page":"22","article-title":"Ideas in practice: Developmental writers\u2019 attitudes toward audio and written feedback","volume":"30","author":"Sipple Susan","year":"2007","unstructured":"Susan Sipple . 2007 . Ideas in practice: Developmental writers\u2019 attitudes toward audio and written feedback . Journal of Developmental Education 30 , 3 (2007), 22 . Susan Sipple. 2007. Ideas in practice: Developmental writers\u2019 attitudes toward audio and written feedback. Journal of Developmental Education 30, 3 (2007), 22.","journal-title":"Journal of Developmental Education"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSMC.2012.6378301"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287067"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/169059.169150"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1177\/1098214005283748"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.3906\/elk-1710-128"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/1983302.1983311"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2009.5202770"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-17534-3_16"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.11120\/beej.2014.00022"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1080\/07421222.1989.11517849"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/258549.258700"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/1400885.1400972"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/2639189.2639251"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.13189\/ujer.2017.050309"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2015.7333762"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290607.3312918"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2011.08.008"}],"container-title":["ACM Transactions on Computer-Human Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3453988","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3453988","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:47:51Z","timestamp":1750193271000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3453988"}},"subtitle":["Gaze-Assisted Voice-Based Implicit Note-taking"],"short-title":[],"issued":{"date-parts":[[2021,8,11]]},"references-count":67,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,8,31]]}},"alternative-id":["10.1145\/3453988"],"URL":"https:\/\/doi.org\/10.1145\/3453988","relation":{},"ISSN":["1073-0516","1557-7325"],"issn-type":[{"value":"1073-0516","type":"print"},{"value":"1557-7325","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,8,11]]},"assertion":[{"value":"2020-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-03-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-08-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}