{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,30]],"date-time":"2025-12-30T17:53:32Z","timestamp":1767117212706,"version":"3.41.0"},"reference-count":104,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2025,2,10]],"date-time":"2025-02-10T00:00:00Z","timestamp":1739145600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Interact. Intell. Syst."],"published-print":{"date-parts":[[2025,3,31]]},"abstract":"<jats:p>\n            Multimodal scene search of conversations is essential for unlocking valuable insights into social dynamics and enhancing our communication. While experts in conversational analysis have their own knowledge and skills to find key scenes, a lack of comprehensive, user-friendly tools that streamline the processing of diverse multimodal queries impedes efficiency and objectivity. To address this gap, we developed\n            <jats:italic>ConverSearch<\/jats:italic>\n            , a visual-programming-based tool based on insights for effective interface and implementation design derived from a formative study with experts. The tool allows experts to integrate various machine learning algorithms to capture human behavioral cues without the need for coding. Our user study, employing the System Usability Scale (SUS) and satisfaction metrics, demonstrated high user preference, reflecting the tool\u2019s ease of use and effectiveness in supporting scene search tasks. Additionally, through a deployment trial within industrial organizations, we confirmed the tool\u2019s objectivity, reusability, and potential to enhance expert workflows. This suggests the advantages of expert-AI collaboration in domains requiring human contextual understanding and demonstrates how customizable, transparent tools yielding reusable artifacts can support expert-driven tasks in complex, multimodal environments.\n          <\/jats:p>","DOI":"10.1145\/3709012","type":"journal-article","created":{"date-parts":[[2024,12,23]],"date-time":"2024-12-23T13:26:43Z","timestamp":1734960403000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["ConverSearch: Supporting Experts in Human Behavior Analysis of Conversational Videos with a Multimodal Scene Search Tool"],"prefix":"10.1145","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7868-4754","authenticated-orcid":false,"given":"Riku","family":"Arakawa","sequence":"first","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, Pennsylvania, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3270-1974","authenticated-orcid":false,"given":"Kiyosu","family":"Maeda","sequence":"additional","affiliation":[{"name":"Princeton University, Princeton, New Jersey, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2558-735X","authenticated-orcid":false,"given":"Hiromu","family":"Yakura","sequence":"additional","affiliation":[{"name":"Max-Planck Institute for Human Development, Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,2,10]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/1922649.1922653"},{"key":"e_1_3_2_3_2","first-page":"71:1","volume-title":"Proceedings of the ACM on Interactive, Mobile, Wearable, and Ubiquitous Technologies","volume":"3","author":"Ahuja Karan","year":"2019","unstructured":"Karan Ahuja, Dohyun Kim, Franceska Xhakaj, Virag Varga, Anne Xie, Stanley Zhang, Jay Eric Townsend, Chris Harrison, Amy Ogan, and Yuvraj Agarwal. 2019. EduSense: Practical classroom sensing at scale. Proceedings of the ACM on Interactive, Mobile, Wearable, and Ubiquitous Technologies 3, 3 (2019), 71:1\u201371:26. DOI: 10.1145\/3351229"},{"key":"e_1_3_2_4_2","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1145\/1026653.1026656","volume-title":"Proceedings of the 1st ACM Workshop on Continuous Archival and Retrieval of Personal Experiences","author":"Aizawa Kiyoharu","year":"2004","unstructured":"Kiyoharu Aizawa, Datchakorn Tancharoen, Shinya Kawasaki, and Toshihiko Yamasaki. 2004. Efficient retrieval of life log based on context and content. In Proceedings of the 1st ACM Workshop on Continuous Archival and Retrieval of Personal Experiences. ACM, New York, NY, 22\u201331. DOI: 10.1145\/1026653.1026656"},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1145\/3463948.3469069","volume-title":"Proceedings of the 4th Annual on Lifelog Search Challenge","author":"Alam Naushad","year":"2021","unstructured":"Naushad Alam, Yvette Graham, and Cathal Gurrin. 2021. Memento: A prototype lifelog search engine for LSC\u201921. In Proceedings of the 4th Annual on Lifelog Search Challenge. ACM, New York, NY, 53\u201358. DOI: 10.1145\/3463948.3469069"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2496269"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1011395131992"},{"key":"e_1_3_2_8_2","first-page":"3:1","volume-title":"Proceedings of the 2019 ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Amershi Saleema","year":"2019","unstructured":"Saleema Amershi, Daniel S. Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi T. Iqbal, Paul N. Bennett, Kori Inkpen, et\u00a0al. 2019. Guidelines for human-AI interaction. In Proceedings of the 2019 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 3:1\u20133:13. DOI: 10.1145\/3290605.3300233"},{"key":"e_1_3_2_9_2","first-page":"6948","volume-title":"Proceedings of the 2021 IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Ando Shintaro","year":"2021","unstructured":"Shintaro Ando and Hiromasa Fujihara. 2021. Construction of a large-scale Japanese ASR corpus on TV recordings. In Proceedings of the 2021 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, New York, NY, 6948\u20136952. DOI: 10.1109\/ICASSP39728.2021.9413425"},{"key":"e_1_3_2_10_2","first-page":"46:1","volume-title":"Proceedings of the ACM on Interactive, Mobile, Wearable, and Ubiquitous Technologies","volume":"7","author":"Arakawa Riku","year":"2023","unstructured":"Riku Arakawa, Karan Ahuja, Kristie Mak, Gwendolyn Thompson, Sam Shaaban, Oliver Lindhiem, and Mayank Goel. 2023. LemurDx: Using unconstrained passive sensing for an objective measurement of hyperactivity in children with no parent input. Proceedings of the ACM on Interactive, Mobile, Wearable, and Ubiquitous Technologies 7, 2 (2023), 46:1\u201346:23. DOI: 10.1145\/3596244"},{"key":"e_1_3_2_11_2","first-page":"572","volume-title":"Proceedings of the 2019 ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Arakawa Riku","year":"2019","unstructured":"Riku Arakawa and Hiromu Yakura. 2019. REsCUE: A framework for REal-time feedback on behavioral CUEs using multimodal anomaly detection. In Proceedings of the 2019 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 572. DOI: 10.1145\/3290605.3300802"},{"key":"e_1_3_2_12_2","first-page":"574:1","volume-title":"Proceedings of the 2020 ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Arakawa Riku","year":"2020","unstructured":"Riku Arakawa and Hiromu Yakura. 2020. INWARD: A computer-supported tool for video-reflection improves efficiency and effectiveness in executive coaching. In Proceedings of the 2020 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 574:1\u2013574:13. DOI: 10.1145\/3313831.3376703"},{"key":"e_1_3_2_13_2","first-page":"378:1","volume-title":"Proceedings of the Extended Abstracts of the 2023 ACM SIGCHI Conference on Human Factors in Computing Systems.","author":"Arakawa Riku","year":"2023","unstructured":"Riku Arakawa and Hiromu Yakura. 2023. AI for Human assessment: What do professional assessors need? In Proceedings of the Extended Abstracts of the 2023 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 378:1\u2013378:7. DOI: 10.1145\/3544549.3573849"},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","first-page":"70","DOI":"10.1145\/3503161.3548363","volume-title":"Proceedings of the 30th ACM International Conference on Multimedia","author":"Balazia Michal","year":"2022","unstructured":"Michal Balazia, Philipp M\u00fcller, \u00c1kos Levente T\u00e1nczos, August von Liechtenstein, and Fran\u00e7ois Br\u00e9mond. 2022. Bodily behaviors in social interaction: Novel annotations and state-of-the-art evaluation. In Proceedings of the 30th ACM International Conference on Multimedia. ACM, New York, NY, 70\u201379. DOI: 10.1145\/3503161.3548363"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1080\/10447310802205776"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","first-page":"577","DOI":"10.1007\/978-3-030-37734-2_47","volume-title":"Proceedings of the 6th International Conference on Multimedia Modeling","author":"Baur Tobias","year":"2020","unstructured":"Tobias Baur, Sina Clausen, Alexander Heimerl, Florian Lingenfelser, Wolfgang Lutz, and Elisabeth Andr\u00e9. 2020. NOVA: A tool for explanatory multimodal behavior analysis and its application to psychotherapy. In Proceedings of the 6th International Conference on Multimedia Modeling. Springer, Cham, Swizterland, 577\u2013588. DOI: 10.1007\/978-3-030-37734-2_47"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","first-page":"160","DOI":"10.1007\/978-3-319-02714-2_14","volume-title":"Proceedings of the 4th International Workshop on Human Behavior Understanding","author":"Baur Tobias","year":"2013","unstructured":"Tobias Baur, Ionut Damian, Florian Lingenfelser, Johannes Wagner, and Elisabeth Andr\u00e9. 2013. NovA: Automated analysis of nonverbal signals in social interactions. In Proceedings of the 4th International Workshop on Human Behavior Understanding. Springer, Cham, Swizterland, 160\u2013171. DOI: 10.1007\/978-3-319-02714-2_14"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/S13218-020-00632-3"},{"key":"e_1_3_2_19_2","first-page":"448:1","volume-title":"Proceedings of the 2024 ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Bedmutha Manas Satish","year":"2024","unstructured":"Manas Satish Bedmutha, Anuujin Tsedenbal, Kelly Tobar, Sarah Borsotto, Kimberly R. Sladek, Deepansha Singh, Reggie Casanova-Perez, Emily Bascom, Brian Wood, Janice Sabin, et\u00a0al. 2024. ConverSense: An automated approach to assess patient-provider interactions using social signals. In Proceedings of the 2024 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 448:1\u2013448:22. DOI: 10.1145\/3613904.3641998"},{"key":"e_1_3_2_20_2","first-page":"1","volume-title":"Proceedings of the ACM on Human-Computer Interaction","volume":"6","author":"Benke Ivo","year":"2022","unstructured":"Ivo Benke, Maren Schneider, Xuanhui Liu, and Alexander Maedche. 2022. TeamSpiritous - A retrospective emotional competence development system for video-meetings. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2 (2022), 1\u201328. DOI: 10.1145\/3555117"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1177\/0265532212455394"},{"key":"e_1_3_2_22_2","first-page":"7124","volume-title":"Proceedings of the 2020 IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Bredin Herv\u00e9","year":"2020","unstructured":"Herv\u00e9 Bredin, Ruiqing Yin, Juan Manuel Coria, Gregory Gelly, Pavel Korshunov, Marvin Lavechin, Diego Fustes, Hadrien Titeux, Wassim Bouaziz, and Marie-Philippe Gill. 2020. Pyannote.audio: Neural building blocks for speaker diarization. In Proceedings of the 2020 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, New York, NY, 7124\u20137128. DOI: 10.1109\/ICASSP40776.2020.9052974"},{"key":"e_1_3_2_23_2","first-page":"207","volume-title":"Usability Evaluation in Industry","author":"Brooke John","year":"1996","unstructured":"John Brooke. 1996. SUS: A \u201dquick and dirty\u201d usability scale. In Usability Evaluation in Industry. Patrick W. Jordan, B. Thomas, Ian Lyall McClelland, and Bernard Weerdmeester (Eds.), CRC Press, London, UK, 207\u2013212."},{"key":"e_1_3_2_24_2","first-page":"2065","volume-title":"Proceedings of the 4th International Conference on Language Resources and Evaluation","author":"Brugman Hennie","year":"2004","unstructured":"Hennie Brugman and Albert Russel. 2004. Annotating multi-media\/multi-modal resources with ELAN. In Proceedings of the 4th International Conference on Language Resources and Evaluation. ELRA, Luxembourg, 2065\u20132068."},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1017\/9781316676202"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1002\/047134608X.W1707"},{"key":"e_1_3_2_27_2","first-page":"1","volume-title":"Proceedings of the Extended Abstracts of the 2020 ACM SIGCHI Conference on Human Factors in Computing Systems.","author":"Carney Michelle","year":"2020","unstructured":"Michelle Carney, Barron Webster, Irene Alvarado, Kyle Phillips, Noura Howell, Jordan Griffith, Jonas Jongejan, Amit Pitaru, and Alexander Chen. 2020. Teachable machine: Approachable web-based tool for exploring machine learning classification. In Proceedings of the Extended Abstracts of the 2020 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 1\u20138. DOI: 10.1145\/3334480.3382839"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2024.3388521"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/b0-08-044854-2\/00796-3"},{"key":"e_1_3_2_30_2","first-page":"565","volume-title":"Proceedings of the 33rd ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Damian Ionut","year":"2015","unstructured":"Ionut Damian, Chiew Seng Sean Tan, Tobias Baur, Johannes Sch\u00f6ning, Kris Luyten, and Elisabeth Andr\u00e9. 2015. Augmenting social interactions: Realtime behavioural feedback using social signal processing techniques. In Proceedings of the 33rd ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 565\u2013574. DOI: 10.1145\/2702123.2702314"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","DOI":"10.1037\/10001-000","volume-title":"The Expression of the Emotions in Man and Animals","author":"Darwin Charles","year":"1872","unstructured":"Charles Darwin. 1872. The Expression of the Emotions in Man and Animals. John Murray, London, UK."},{"key":"e_1_3_2_32_2","first-page":"67","article-title":"Should \u201duh\u201d and \u201dum\u201d be categorized as markers of disfluency? The use of fillers in a challenging conversational context","volume":"4","author":"Degand Liesbeth","year":"2019","unstructured":"Liesbeth Degand, Ga\u00ebtanelle Gilquin, Laurence Meurant, and Anne Catherine Simon. 2019. Should \u201duh\u201d and \u201dum\u201d be categorized as markers of disfluency? The use of fillers in a challenging conversational context. Fluency and Disfluency across Languages and Language Varieties 4 (2019), 67.","journal-title":"Fluency and Disfluency across Languages and Language Varieties"},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","first-page":"1577","DOI":"10.1145\/3404835.3462806","volume-title":"Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Deldjoo Yashar","year":"2021","unstructured":"Yashar Deldjoo, Johanne R. Trippas, and Hamed Zamani. 2021. Towards multi-modal conversational information seeking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, New York, NY, 1577\u20131587. DOI: 10.1145\/3404835.3462806"},{"key":"e_1_3_2_34_2","first-page":"5202","volume-title":"Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Deng Jiankang","year":"2020","unstructured":"Jiankang Deng, Jia Guo, Evangelos Ververas, Irene Kotsia, and Stefanos Zafeiriou. 2020. RetinaFace: Single-shot multi-level face localisation in the wild. In Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. CVF\/IEEE, Washington, DC, 5202\u20135211. DOI: 10.1109\/CVPR42600.2020.00525"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1037\/0033-2909.111.2.203"},{"key":"e_1_3_2_36_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. ACL, Stroudsburg, PA, 4171\u20134186. DOI: 10.18653\/v1\/n19-1423"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","first-page":"271","DOI":"10.1145\/1322192.1322239","volume-title":"Proceedings of the 9th International Conference on Multimodal Interfaces","author":"Dong Wen","year":"2007","unstructured":"Wen Dong, Bruno Lepri, Alessandro Cappelletti, Alex Pentland, Fabio Pianesi, and Massimo Zancanaro. 2007. Using the influence model to recognize functional roles in meetings. In Proceedings of the 9th International Conference on Multimodal Interfaces. ACM, New York, NY, 271\u2013278. DOI: 10.1145\/1322192.1322239"},{"key":"e_1_3_2_38_2","first-page":"23 pages","volume-title":"Proceedings of the 2023 ACM SIGCHI Conference on Human Factors in Computing Systems.","author":"Du Ruofei","year":"2023","unstructured":"Ruofei Du, Na Li, Jing Jin, Michelle Carney, Scott Miles, Maria Kleiner, Xiuxiu Yuan, Yinda Zhang, Anuva Kulkarni, Xingyu \u201cBruce\u201d Liu, et\u00a0al. 2023. Rapsai: Accelerating machine learning prototyping of multimedia applications through visual programming. In Proceedings of the 2023 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 23 pages. DOI: 10.1145\/3544548.3581338"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijhcs.2020.102411"},{"key":"e_1_3_2_40_2","first-page":"2353","volume-title":"Proceedings of the 2017 IEEE International Conference on Computer Vision","author":"Fang Haoshu","year":"2017","unstructured":"Haoshu Fang, Shuqin Xie, Yu-Wing Tai, and Cewu Lu. 2017. RMPE: Regional multi-person pose estimation. In Proceedings of the 2017 IEEE International Conference on Computer Vision. IEEE Computer Society, Washington, DC, 2353\u20132362. DOI: 10.1109\/ICCV.2017.256"},{"issue":"6","key":"e_1_3_2_41_2","doi-asserted-by":"crossref","first-page":"7157","DOI":"10.1109\/TPAMI.2022.3222784","article-title":"AlphaPose: Whole-body regional multi-person pose estimation and tracking in real-time","volume":"45","author":"Fang Hao-Shu","year":"2023","unstructured":"Hao-Shu Fang, Jiefeng Li, Hongyang Tang, Chao Xu, Haoyi Zhu, Yuliang Xiu, Yong-Lu Li, and Cewu Lu. 2023. AlphaPose: Whole-body regional multi-person pose estimation and tracking in real-time. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 6 (2023), 7157\u20137173.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_42_2","first-page":"339","volume-title":"Proceedings of the 15th European Conference on Computer Vision","author":"Fischer Tobias","year":"2018","unstructured":"Tobias Fischer, Hyung Jin Chang, and Yiannis Demiris. 2018. RT-GENE: Real-time eye gaze estimation in natural environments. In Proceedings of the 15th European Conference on Computer Vision. Springer, Cham, Switzerland, 339\u2013357. DOI: 10.1007\/978-3-030-01249-6_21"},{"key":"e_1_3_2_43_2","doi-asserted-by":"crossref","first-page":"979","DOI":"10.1145\/3379337.3415592","volume-title":"Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology","author":"Fraser C. Ailie","year":"2020","unstructured":"C. Ailie Fraser, Julia M. Markel, N. James Basa, Mira Dontcheva, and Scott R. Klemmer. 2020. ReMap: Lowering the barrier to help-seeking with multimodal search. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology. ACM, New York, NY, 979\u2013986. DOI: 10.1145\/3379337.3415592"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","unstructured":"Daniel Y. Fu Will Crichton James Hong Xinwei Yao Haotian Zhang Anh Truong Avanika Narayan Maneesh Agrawala Christopher R\u00e9 and Kayvon Fatahalian. 2019. Rekall: Specifying video events using compositions of spatiotemporal labels. arXiv: 1910.02993. DOI: 10.48550\/arXiv.1910.02993","DOI":"10.48550\/arXiv.1910.02993"},{"key":"e_1_3_2_45_2","first-page":"41","volume-title":"Proceedings of the 2006 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems","author":"Gatica-Perez Daniel","year":"2006","unstructured":"Daniel Gatica-Perez. 2006. Analyzing group interactions in conversations: A review. In Proceedings of the 2006 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems. IEEE, New York, NY, 41\u201346. DOI: 10.1109\/MFI.2006.265658"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-019-12397-x"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3449084"},{"key":"e_1_3_2_48_2","first-page":"5036","volume-title":"Proceedings of the 21st Annual Conference of the International Speech Communication Association","author":"Gulati Anmol","year":"2020","unstructured":"Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et\u00a0al. 2020. Conformer: Convolution-augmented transformer for speech recognition. In Proceedings of the 21st Annual Conference of the International Speech Communication Association. ISCA, Baixas, France, 5036\u20135040. DOI: 10.21437\/Interspeech.2020-3015"},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","DOI":"10.1145\/3210539","volume-title":"Proceedings of the 2018 ACM Workshop on the Lifelog Search Challenge","author":"Gurrin Cathal","year":"2018","unstructured":"Cathal Gurrin, Klaus Schoeffmann, Hideo Joho, Duc-Tien Dang-Nguyen, Michael Riegler, and Luca Piras (Eds.). 2018. In Proceedings of the 2018 ACM Workshop on the Lifelog Search Challenge. ACM, New York, NY."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/s0166-4115(08)62386-9"},{"key":"e_1_3_2_51_2","first-page":"1","volume-title":"Proceedings of the 10th International Conference on Affective Computing and Intelligent Interaction","author":"Haut Kurtis Glenn","year":"2022","unstructured":"Kurtis Glenn Haut, Adira Blumenthal, Sarah Atterbury, Xiaofei Zhou, Wasifur Rahman, Emanuela Natali, Mohammad Rafayet Ali, and Ehsan Hoque. 2022. Assistive video filters for people with Parkinson\u2019s disease to remove tremors and adjust voice. In Proceedings of the 10th International Conference on Affective Computing and Intelligent Interaction. IEEE, New York, NY, 1\u20138. DOI: 10.1109\/ACII55700.2022.9953845"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2023.3326586"},{"key":"e_1_3_2_53_2","first-page":"770","volume-title":"Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition. IEEE Computer Society, Washington, DC, 770\u2013778. DOI: 10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_54_2","first-page":"109","volume-title":"Proceedings of the 8th International Conference on Affective Computing and Intelligent Interaction","author":"Heimerl Alexander","year":"2019","unstructured":"Alexander Heimerl, Tobias Baur, Florian Lingenfelser, Johannes Wagner, and Elisabeth Andr\u00e9. 2019. NOVA - A tool for eXplainable cooperative machine learning. In Proceedings of the 8th International Conference on Affective Computing and Intelligent Interaction. IEEE, New York, NY, 109\u2013115. DOI: 10.1109\/ACII.2019.8925519"},{"key":"e_1_3_2_55_2","doi-asserted-by":"crossref","first-page":"697","DOI":"10.1145\/2493432.2493502","volume-title":"Proceedings of the 2013 ACM International Joint Conference on Pervasive and Ubiquitous Computing","author":"Hoque Mohammed (Ehsan)","year":"2013","unstructured":"Mohammed (Ehsan) Hoque, Matthieu Courgeon, Jean-Claude Martin, Bilge Mutlu, and Rosalind W. Picard. 2013. MACH: My automated conversation coach. In Proceedings of the 2013 ACM International Joint Conference on Pervasive and Ubiquitous Computing. ACM, New York, NY, 697\u2013706. DOI: 10.1145\/2493432.2493502"},{"key":"e_1_3_2_56_2","first-page":"159","volume-title":"Proceeding of the 1999 ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Horvitz Eric","year":"1999","unstructured":"Eric Horvitz. 1999. Principles of mixed-initiative user interfaces. In Proceeding of the 1999 ACM SIGCHI Conference on Human Factors in Computing Systems. Marian G. Williams and Mark W. Altom (Eds.), ACM, New York, NY, 159\u2013166. DOI: 10.1145\/302979.303030"},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","first-page":"101","DOI":"10.18293\/VLSS2017-012","article-title":"Towards understanding successful novice example use in blocks-based programming","volume":"3","author":"Ichinco Michelle","year":"2017","unstructured":"Michelle Ichinco, Kyle J. Harms, and Caitlin Kelleher. 2017. Towards understanding successful novice example use in blocks-based programming. Journal of Visual Language and Sentient Systems 3 (2017), 101\u2013118.","journal-title":"Journal of Visual Language and Sentient Systems"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3582272"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1093\/oso\/9780198799603.003.0007"},{"key":"e_1_3_2_60_2","first-page":"582","volume-title":"Proceedings of the 34th Annual ACM Symposium on User Interface Software and Technology","author":"Li Daniel","year":"2021","unstructured":"Daniel Li, Thomas Chen, Albert Tung, and Lydia B. Chilton. 2021. Hierarchical summarization for longform spoken dialog. In Proceedings of the 34th Annual ACM Symposium on User Interface Software and Technology. ACM, New York, NY, 582\u2013597. DOI: 10.1145\/3472749.3474771"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11187-019-00205-1"},{"key":"e_1_3_2_62_2","first-page":"740","volume-title":"Proceedings of the 13th European Conference on Computer Vision","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In Proceedings of the 13th European Conference on Computer Vision. Springer, Cham, Switzerland, 740\u2013755. DOI: 10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1561\/1500000073"},{"key":"e_1_3_2_64_2","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1145\/3505284.3529959","volume-title":"Proceedings of the 2022 ACM International Conference on Interactive Media Experiences","author":"Maeda Kiyosu","year":"2022","unstructured":"Kiyosu Maeda, Riku Arakawa, and Jun Rekimoto. 2022. CalmResponses: Displaying collective audience reactions in remote communication. In Proceedings of the 2022 ACM International Conference on Interactive Media Experiences. ACM, New York, NY, 193\u2013208. DOI: 10.1145\/3505284.3529959"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41593-018-0209-y"},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1145\/3379172.3391727","volume-title":"Proceedings of the 3rd ACM Workshop on Lifelog Search Challenge","author":"Mejzl\u00edk Frantisek","year":"2020","unstructured":"Frantisek Mejzl\u00edk, Patrik Vesel\u00fd, Miroslav Kratochv\u00edl, Tom\u00e1s Soucek, and Jakub Lokoc. 2020. SOMHunter for lifelog search. In Proceedings of the 3rd ACM Workshop on Lifelog Search Challenge. ACM, New York, NY, 73\u201375. DOI: 10.1145\/3379172.3391727"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3579481"},{"key":"e_1_3_2_68_2","first-page":"394","volume-title":"Proceedings of the 22nd International Conference on MultiMedia Modeling","author":"Moumtzidou Anastasia","year":"2016","unstructured":"Anastasia Moumtzidou, Theodoros Mironidis, Evlampios E. Apostolidis, Foteini Markatopoulou, Anastasia Ioannidou, Ilias Gialampoukidis, Konstantinos Avgerinakis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, et\u00a0al. 2016. VERGE: A multimodal interactive search engine for video browsing and retrieval. In Proceedings of the 22nd International Conference on MultiMedia Modeling. Springer, Cham, Switzerland, 394\u2013399. DOI: 10.1007\/978-3-319-27674-8_39"},{"key":"e_1_3_2_69_2","doi-asserted-by":"crossref","first-page":"136","DOI":"10.1145\/2663204.2663248","volume-title":"Proceedings of the 16th International Conference on Multimodal Interaction","author":"Nihei Fumio","year":"2014","unstructured":"Fumio Nihei, Yukiko I. Nakano, Yuki Hayashi, Hung-Hsuan Huang, and Shogo Okada. 2014. Predicting influential statements in group discussions using speech and head motion information. In Proceedings of the 16th International Conference on Multimodal Interaction. ACM, New York, NY, 136\u2013143. DOI: 10.1145\/2663204.2663248"},{"key":"e_1_3_2_70_2","doi-asserted-by":"crossref","first-page":"421","DOI":"10.1145\/3136755.3136803","volume-title":"Proceedings of the 19th ACM International Conference on Multimodal Interaction","author":"Nihei Fumio","year":"2017","unstructured":"Fumio Nihei, Yukiko I. Nakano, and Yutaka Takase. 2017. Predicting meeting extracts in group discussions using multimodal convolutional neural networks. In Proceedings of the 19th ACM International Conference on Multimodal Interaction. ACM, New York, NY, 421\u2013425. DOI: 10.1145\/3136755.3136803"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0020-7373(05)80150-6"},{"key":"e_1_3_2_72_2","first-page":"171","volume-title":"Proceedings of the 2011 International Symposium on Human Interface","volume":"6772","author":"Otsuka Kazuhiro","year":"2011","unstructured":"Kazuhiro Otsuka. 2011. Multimodal conversation scene analysis for understanding people\u2019s communicative behaviors in face-to-face meetings. In Proceedings of the 2011 International Symposium on Human Interface, Vol. 6772. Springer, Cham, Switzerland, 171\u2013179. DOI: 10.1007\/978-3-642-21669-5_21"},{"key":"e_1_3_2_73_2","first-page":"51:1","volume-title":"Proceedings of the 2022 ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Park Hyanghee","year":"2022","unstructured":"Hyanghee Park, Daehwan Ahn, Kartik Hosanagar, and Joonhwan Lee. 2022. Designing fair AI in human resource management: Understanding tensions surrounding algorithmic evaluation and envisioning stakeholder-centered solutions. In Proceedings of the 2022 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 51:1\u201351:22. DOI: 10.1145\/3491102.3517672"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","unstructured":"B. V. Patel and B. B. Meshram. 2012. Content based video retrieval systems. arXiv:1205.1641. DOI: 10.48550\/arXiv.1205.1641","DOI":"10.48550\/arXiv.1205.1641"},{"key":"e_1_3_2_75_2","first-page":"421","volume-title":"Proceedings of the 2013 ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Patel Rupa A.","year":"2013","unstructured":"Rupa A. Patel, Andrea L. Hartzler, Mary Czerwinski, Wanda Pratt, Anthony L. Back, and Asta Roseway. 2013. Leveraging visual feedback from social signal processing to enhance clinicians\u2019 nonverbal skills. In Proceedings of the 2013 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 421\u2013426. DOI: 10.1145\/2468356.2468431"},{"key":"e_1_3_2_76_2","doi-asserted-by":"crossref","first-page":"181","DOI":"10.1145\/2807442.2807502","volume-title":"Proceedings of the 28th Annual ACM Symposium on User Interface Software and Technology","author":"Pavel Amy","year":"2015","unstructured":"Amy Pavel, Dan B. Goldman, Bj\u00f6rn Hartmann, and Maneesh Agrawala. 2015. SceneSkim: Searching and browsing movies using synchronized captions, scripts and plot summaries. In Proceedings of the 28th Annual ACM Symposium on User Interface Software and Technology. ACM, New York, NY, 181\u2013190. DOI: 10.1145\/2807442.2807502"},{"key":"e_1_3_2_77_2","volume-title":"Body Language","author":"Pease Allan","year":"1981","unstructured":"Allan Pease. 1981. Body Language. Sheldon Press, London, UK."},{"key":"e_1_3_2_78_2","first-page":"341","volume-title":"Proceedings of the 23rd ACM International Conference on Multimodal Interaction","author":"Penzkofer Anna","year":"2021","unstructured":"Anna Penzkofer, Philipp M\u00fcller, Felix B\u00fchler, Sven Mayer, and Andreas Bulling. 2021. ConAn: A usable tool for multimodal conversation analysis. In Proceedings of the 23rd ACM International Conference on Multimodal Interaction. ACM, New York, NY, 341\u2013351. DOI: 10.1145\/3462244.3479886"},{"key":"e_1_3_2_79_2","first-page":"460","volume-title":"Proceedings of the 2009 International Conference on Intelligent Virtual Agents","author":"Pfeifer Laura M.","year":"2009","unstructured":"Laura M. Pfeifer and Timothy W. Bickmore. 2009. Should agents speak like, um, humans? The use of conversational fillers by virtual agents. In Proceedings of the 2009 International Conference on Intelligent Virtual Agents. Springer, Cham, Swizterland, 460\u2013466. DOI: 10.1007\/978-3-642-04380-2_50"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.1145\/3484510"},{"key":"e_1_3_2_81_2","first-page":"1","volume-title":"Proceedings of the 19th IEEE International Workshop on Multimedia Signal Processing","author":"Rotman Daniel","year":"2017","unstructured":"Daniel Rotman, Dror Porat, and Gal Ashour. 2017. Robust video scene detection using multimodal fusion of optimally grouped features. In Proceedings of the 19th IEEE International Workshop on Multimedia Signal Processing. IEEE, New York, NY, 1\u20136. DOI: 10.1109\/MMSP.2017.8122267"},{"key":"e_1_3_2_82_2","first-page":"397","volume-title":"Proceedings of the 2013 IEEE International Conference on Computer Vision Workshops","author":"Sagonas Christos","year":"2013","unstructured":"Christos Sagonas, Georgios Tzimiropoulos, Stefanos Zafeiriou, and Maja Pantic. 2013. 300 faces in-the-wild challenge: The first facial landmark localization challenge. In Proceedings of the 2013 IEEE International Conference on Computer Vision Workshops. IEEE Computer Society, New York, NY, 397\u2013403. DOI: 10.1109\/ICCVW.2013.59"},{"key":"e_1_3_2_83_2","first-page":"252:1","volume-title":"Proceedings of the 2021 ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Samrose Samiha","year":"2021","unstructured":"Samiha Samrose, Daniel McDuff, Robert Sim, Jina Suh, Kael Rowan, Javier Hernandez, Sean Rintel, Kevin Moynihan, and Mary Czerwinski. 2021. MeetingCoach: An intelligent dashboard for supporting effective & inclusive meetings. In Proceedings of the 2021 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 252:1\u2013252:13. DOI: 10.1145\/3411764.3445615"},{"key":"e_1_3_2_84_2","first-page":"160:1","volume-title":"Proceedings of the ACM on Interactive, Mobile, Wearable, and Ubiquitous Technologies","volume":"1","author":"Samrose Samiha","year":"2017","unstructured":"Samiha Samrose, Ru Zhao, Jeffery White, Vivian Li, Luis Nova, Yichen Lu, Mohammad Rafayet Ali, and Mohammed E. Hoque. 2017. CoCo: Collaboration coach for understanding team dynamics during video conferencing. Proceedings of the ACM on Interactive, Mobile, Wearable, and Ubiquitous Technologies 1, 4 (2017), 160:1\u2013160:24. DOI: 10.1145\/3161186"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.1002\/9781118325001"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1037\/0021-843x.102.3.430"},{"key":"e_1_3_2_87_2","first-page":"4015","volume-title":"Proceedings of the 8th Annual Conference of the International Speech Communication Association","author":"Sloetjes Han","year":"2007","unstructured":"Han Sloetjes, Albert Russel, and Alexander Klassmann. 2007. ELAN: A free and open-source multimedia annotation tool. In Proceedings of the 8th Annual Conference of the International Speech Communication Association. ISCA, Baixas, France, 4015\u20134016."},{"key":"e_1_3_2_88_2","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1145\/3512729.3533008","volume-title":"Proceedings of the 5th Annual on Lifelog Search Challenge","author":"Spiess Florian","year":"2022","unstructured":"Florian Spiess and Heiko Schuldt. 2022. Multimodal interactive lifelog retrieval with Vitrivr-VR. In Proceedings of the 5th Annual on Lifelog Search Challenge. ACM, New York, NY, 38\u201342. DOI: 10.1145\/3512729.3533008"},{"key":"e_1_3_2_89_2","first-page":"1333","volume-title":"Proceedings of the Extended Abstracts of the 2004 ACM SIGCHI Conference on Human Factors in Computing Systems.","author":"Takemae Yoshinao","year":"2004","unstructured":"Yoshinao Takemae, Kazuhiro Otsuka, and Naoki Mukawa. 2004. Impact of video editing based on participants\u2019 gaze in multiparty conversation. In Proceedings of the Extended Abstracts of the 2004 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 1333\u20131336. DOI: 10.1145\/985921.986057"},{"key":"e_1_3_2_90_2","first-page":"286","volume-title":"Proceedings of the 20th ACM International Conference on Intelligent User Interfaces","author":"Tanveer Mohammad Iftekhar","year":"2015","unstructured":"Mohammad Iftekhar Tanveer, Emy Lin, and Mohammed (Ehsan) Hoque. 2015. Rhema: A real-time in-situ intelligent interface to help people with public speaking. In Proceedings of the 20th ACM International Conference on Intelligent User Interfaces. ACM, New York, NY, 286\u2013295. DOI: 10.1145\/2678025.2701386"},{"key":"e_1_3_2_91_2","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898717921"},{"key":"e_1_3_2_92_2","first-page":"453","volume-title":"Proceedings of the 21st ACM International Conference on Multimodal Interaction","author":"Tavabi Leili","year":"2019","unstructured":"Leili Tavabi. 2019. Multimodal machine learning for interactive mental health therapy. In Proceedings of the 21st ACM International Conference on Multimodal Interaction. ACM, New York, NY, 453\u2013456. DOI: 10.1145\/3340555.3356095"},{"key":"e_1_3_2_93_2","unstructured":"O\u2019Reilly Editorial Team. 2021. Low-code and the democratization of programming rethinking where programming is headed. Retrieved from https:\/\/www.oreilly.com\/radar\/low-code-and-the-democratization-of-programming\/"},{"key":"e_1_3_2_94_2","doi-asserted-by":"crossref","first-page":"94","DOI":"10.1145\/3536221.3556613","volume-title":"Proceeding of the 24th ACM International Conference on Multimodal Interaction","author":"Tsfasman Maria","year":"2022","unstructured":"Maria Tsfasman, Kristian Fenech, Morita Tarvirdians, Andr\u00e1s L\u00f6rincz, Catholijn M. Jonker, and Catharine Oertel. 2022. Towards creating a conversational memory for long-term meeting support: Predicting memorable moments in multi-party conversations through eye-gaze. In Proceeding of the 24th ACM International Conference on Multimodal Interaction. ACM, New York, NY, 94\u2013104. DOI: 10.1145\/3536221.3556613"},{"key":"e_1_3_2_95_2","first-page":"5998","volume-title":"Proceedings of the 2017 Annual Conference on Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 2017 Annual Conference on Neural Information Processing Systems. Curran Associates, Red Hook, NY, 5998\u20136008."},{"key":"e_1_3_2_96_2","doi-asserted-by":"publisher","DOI":"10.1109\/T-AFFC.2011.27"},{"key":"e_1_3_2_97_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2020.106897"},{"key":"e_1_3_2_98_2","first-page":"1556","volume-title":"Proceedings of the 5th International Conference on Language Resources and Evaluation","author":"Wittenburg Peter","year":"2006","unstructured":"Peter Wittenburg, Hennie Brugman, Albert Russel, Alexander Klassmann, and Han Sloetjes. 2006. ELAN: A professional framework for multimodality research. In Proceedings of the 5th International Conference on Language Resources and Evaluation. ELRA, Luxembourg, 1556\u20131559."},{"key":"e_1_3_2_99_2","doi-asserted-by":"publisher","DOI":"10.4135\/9781849208765"},{"key":"e_1_3_2_100_2","first-page":"2805","volume-title":"Proceedings of the 8th European Conference on Speech Communication and Technology","author":"Wrede Britta","year":"2003","unstructured":"Britta Wrede and Elizabeth Shriberg. 2003. Spotting \u201chot spots\u201d in meetings: Human judgments and prosodic cues. In Proceedings of the 8th European Conference on Speech Communication and Technology. ISCA, Baixas, France, 2805\u20132808. DOI: 10.21437\/EUROSPEECH.2003-747"},{"key":"e_1_3_2_101_2","first-page":"385:1","volume-title":"Proceedings of the 2022 ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Wu Tongshuang","year":"2022","unstructured":"Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. AI chains: Transparent and controllable human-AI interaction by chaining large language model prompts. In Proceedings of the 2022 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 385:1\u2013385:22. DOI: 10.1145\/3491102.3517582"},{"key":"e_1_3_2_102_2","first-page":"479:1","volume-title":"Proceedings of the Extended Abstracts of the 2023 ACM SIGCHI Conference on Human Factors in Computing Systems.","author":"Yakura Hiromu","year":"2023","unstructured":"Hiromu Yakura. 2023. A generative framework for designing interactions to overcome the gaps between humans and imperfect AIs instead of improving the accuracy of the AIs. In Proceedings of the Extended Abstracts of the 2023 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 479:1\u2013479:5. DOI: 10.1145\/3544549.3577036"},{"key":"e_1_3_2_103_2","first-page":"1208","volume-title":"Proceedings of the 30th International Joint Conference on Artificial Intelligence","author":"Yakura Hiromu","year":"2021","unstructured":"Hiromu Yakura, Yuki Koyama, and Masataka Goto. 2021. Tool- and domain-agnostic parameterization of style transfer effects leveraging pretrained perceptual metrics. In Proceedings of the 30th International Joint Conference on Artificial Intelligence. ijcai.org, Menlo Park, CA, 1208\u20131216."},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","unstructured":"Wengang Zhou Houqiang Li and Qi Tian. 2017. Recent advance in content-based image retrieval: A literature survey. arXiv:1706.06064. DOI: 10.48550\/arXiv.1706.06064","DOI":"10.48550\/arXiv.1706.06064"},{"key":"e_1_3_2_105_2","first-page":"5568","volume-title":"Proceedings of the 2017 ACM SIGCHI Conference on Human Factors in Computing Systems","author":"Zhu Yeshuang","year":"2017","unstructured":"Yeshuang Zhu, Yuntao Wang, Chun Yu, Shaoyun Shi, Yankai Zhang, Shuang He, Peijun Zhao, Xiaojuan Ma, and Yuanchun Shi. 2017. ViVo: Video-augmented dictionary for vocabulary learning. In Proceedings of the 2017 ACM SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, 5568\u20135579. DOI: 10.1145\/3025453.3025779"}],"container-title":["ACM Transactions on Interactive Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709012","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3709012","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:17:55Z","timestamp":1750295875000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709012"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,10]]},"references-count":104,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,3,31]]}},"alternative-id":["10.1145\/3709012"],"URL":"https:\/\/doi.org\/10.1145\/3709012","relation":{},"ISSN":["2160-6455","2160-6463"],"issn-type":[{"type":"print","value":"2160-6455"},{"type":"electronic","value":"2160-6463"}],"subject":[],"published":{"date-parts":[[2025,2,10]]},"assertion":[{"value":"2024-02-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}