{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T10:51:39Z","timestamp":1770893499904,"version":"3.50.1"},"reference-count":59,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2026,2,1]],"date-time":"2026-02-01T00:00:00Z","timestamp":1769904000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>With the rapid growth of online presentations, there has been an increasing need for efficient review of recorded materials. In typical presentations, speakers verbally elaborate on each slide, providing details not captured in the slides themselves. Automatically extracting and embedding these verbal explanations at their corresponding slide locations can greatly enhance the review process for audiences. This paper presents a Slide Annotation System that employs a robust hybrid two-stage detector to identify slide boundaries, extracts slide text through Optical Character Recognition (OCR), transcribes narration, and employs a multimodal Large Language Model (LLM) to generate concise, context-aware annotations that are added to their corresponding slide locations. For evaluations, the technical performance was validated on five recorded presentations, while the user experience was assessed by 37 participants. The results showed that the system achieved a macro-average F1 score of 0.879 (SD=0.024, 95% CI[0.849,0.909]) for slide segmentation and 90.0% accuracy (95% CI[74.4%,96.5%]) for annotation alignment. Subjective evaluations revealed high annotation validity and usefulness as rated by presenters, and a high System Usability Scale (SUS) score of 80.5 (SD=6.7, 95% CI[78.3,82.7]). Qualitative feedback further confirmed that the system effectively streamlined the review process, enabling users to locate key information more efficiently than standard video playback. These findings demonstrate the strong potential of the proposed system as an effective automated annotation system.<\/jats:p>","DOI":"10.3390\/a19020110","type":"journal-article","created":{"date-parts":[[2026,2,2]],"date-time":"2026-02-02T09:00:33Z","timestamp":1770022833000},"page":"110","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["A Slide Annotation System with Multimodal Analysis for Video Presentation Review"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5612-6062","authenticated-orcid":false,"given":"Amma Liesvarastranta","family":"Haz","sequence":"first","affiliation":[{"name":"Department of Information and Communication Systems, Okayama University, Okayama 700-8530, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2896-6686","authenticated-orcid":false,"given":"Komang Candra","family":"Brata","sequence":"additional","affiliation":[{"name":"Department of Information and Communication Systems, Okayama University, Okayama 700-8530, Japan"},{"name":"Department of Informatics Engineering, Universitas Brawijaya, Malang 65145, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3234-3473","authenticated-orcid":false,"given":"Nobuo","family":"Funabiki","sequence":"additional","affiliation":[{"name":"Department of Information and Communication Systems, Okayama University, Okayama 700-8530, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-8781-2018","authenticated-orcid":false,"given":"Htoo Htoo Sandi","family":"Kyaw","sequence":"additional","affiliation":[{"name":"Department of Information and Communication Systems, Okayama University, Okayama 700-8530, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9019-9775","authenticated-orcid":false,"given":"Evianita Dewi","family":"Fajrianti","sequence":"additional","affiliation":[{"name":"Human Centric Multimedia Research Laboratory, Department of Informatic and Computer Engineering, Politeknik Elektronika Negeri Surabaya, Surabaya 60111, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8157-5863","authenticated-orcid":false,"given":"Sritrusta","family":"Sukaridhoto","sequence":"additional","affiliation":[{"name":"Human Centric Multimedia Research Laboratory, Department of Informatic and Computer Engineering, Politeknik Elektronika Negeri Surabaya, Surabaya 60111, Indonesia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,2,1]]},"reference":[{"key":"ref_1","unstructured":"Adhikari Egodawele, M.H., Sedera, D., and Bui, V. (2022, January 4\u20137). A Systematic Review of Digital Transformation Literature (2013\u20132021) and the development of an overarching a-priori model to guide future research. Proceedings of the Australasian Conference on Information Systems (ACIS) 2022 Proceedings, Melbourne, Australia."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"330","DOI":"10.1037\/apl0000906","article-title":"Videoconference fatigue? Exploring changes in fatigue after videoconference meetings during COVID-19","volume":"106","author":"Bennett","year":"2021","journal-title":"J. Appl. Psychol."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ta\u015f, M., and Kiraz, A. (2023). A model for the acceptance and use of online meeting tools. Systems, 11.","DOI":"10.3390\/systems11120558"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Fabriz, S., Mendzheritskaya, J., and Stehle, S. (2021). Impact of synchronous and asynchronous settings of online teaching and learning in higher education on students\u2019 learning experience during COVID-19. Front. Psychol., 12.","DOI":"10.3389\/fpsyg.2021.733554"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Lo, N.P.K., and Wong, A.M.H. (2023). Reimagining teaching and learning in higher education in the post-COVID-19 era: The use of recorded lessons from teachers\u2019 perspectives. Proceedings of the Critical Reflections on ICT and Education: Selected Papers from the HKAECT 2023 International Conference, Hong Kong, China, 15\u201317 June 2023, Springer.","DOI":"10.1007\/978-981-99-7559-4_13"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Lee, H., Liu, M., Scriney, M., and Smeaton, A.F. (2021). Usage-Based Summaries of Learning Videos. Proceedings of the European Conference on Technology Enhanced Learning, Online, 20\u201324 September 2021, Springer.","DOI":"10.1007\/978-3-030-86436-1_46"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1631","DOI":"10.1007\/s40593-025-00481-x","article-title":"A closer look into recent video-based learning research: A comprehensive review of video characteristics, tools, technologies, and learning effectiveness","volume":"35","author":"Navarrete","year":"2025","journal-title":"Int. J. Artif. Intell. Educ."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Kry\u015bci\u0144ski, W., Keskar, N.S., McCann, B., Xiong, C., and Socher, R. (2019). Neural text summarization: A critical evaluation. arXiv.","DOI":"10.18653\/v1\/D19-1051"},{"key":"ref_9","unstructured":"Hall, M., Kirby, R.M., Li, F., Meyer, M., Pascucci, V., Phillips, J.M., Ricci, R., Van der Merwe, J., and Venkatasubramanian, S. (2013). Rethinking abstractions for big data: Why, where, how, and what. arXiv."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Haz, A.L., Funabiki, N., Fajrianti, E.D., and Sukaridhoto, S. (2023, January 15\u201317). A Study of Summarization and Keyword Extraction Function in Meeting Note Generation System from Voice Records. Proceedings of the 2023 12th International Conference on Networks, Communication and Computing, Osaka, Japan.","DOI":"10.1145\/3638837.3638853"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Mayer, R.E. (2005). The Cambridge Handbook of Multimedia Learning, Cambridge University Press.","DOI":"10.1017\/CBO9780511816819"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"ep2408","DOI":"10.30935\/ijpdll\/14694","article-title":"Universal Design for Learning and Artificial Intelligence in the Digital Era: Fostering Inclusion and Autonomous Learning","volume":"6","year":"2024","journal-title":"Int. J. Prof. Dev. Learn. Learn."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"466","DOI":"10.1080\/02619768.2020.1821184","article-title":"COVID-19 and teacher education: A literature review of online teaching and learning practices","volume":"43","author":"Carrillo","year":"2020","journal-title":"Eur. J. Teach. Educ."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"100493","DOI":"10.1016\/j.caeai.2025.100493","article-title":"How AI literacy correlates with affective, behavioral, cognitive and contextual variables: A systematic review","volume":"9","author":"Bewersdorff","year":"2025","journal-title":"Comput. Educ. Artif. Intell."},{"key":"ref_15","first-page":"1","article-title":"Towards Key Point Identification (KPI) for Lecture Videos: Approaches and Performance Evaluation","volume":"21","author":"Wang","year":"2025","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl."},{"key":"ref_16","first-page":"311","article-title":"A formal study of shot boundary detection approaches\u2014Comparative analysis","volume":"Volume 1","author":"Nankani","year":"2021","journal-title":"Soft Computing: Theories and Applications: Proceedings of SoCTA 2020"},{"key":"ref_17","first-page":"4195905","article-title":"Efficient shot boundary detection with multiple visual representations","volume":"2022","author":"Jose","year":"2022","journal-title":"Mob. Inf. Syst."},{"key":"ref_18","unstructured":"Sindel, A., Hernandez, A., Yang, S.H., Christlein, V., and Maier, A. (2022). SliTraNet: Automatic Detection of Slide Transitions in Lecture Videos using Convolutional Neural Networks. arXiv."},{"key":"ref_19","first-page":"1","article-title":"Shot boundary detection using color clustering and attention mechanism","volume":"19","author":"Yuan","year":"2023","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Che, X., Yang, H., and Meinel, C. (2013, January 21\u201325). Lecture video segmentation by automatically analyzing the synchronized slides. Proceedings of the 21st ACM international Conference on Multimedia, Barcelona, Spain.","DOI":"10.1145\/2502081.2508115"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"LeFevre, G., Hosier, J., Zhou, Y., and Gurbani, V.K. (2025). LLM Selection: Improving ASR Transcript Quality via Zero-Shot Prompting. Proceedings of the SoutheastCon 2025, Concord, NC, USA, 27\u201330 March 2025, IEEE.","DOI":"10.1109\/SoutheastCon56624.2025.10971586"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Sapena, O., and Onaindia, E. (2022). Multimodal classification of teaching activities from University lecture recordings. Appl. Sci., 12.","DOI":"10.3390\/app12094785"},{"key":"ref_23","unstructured":"Radford, A., Kim, J.W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I. (2023, January 23\u201329). Robust speech recognition via large-scale weak supervision. Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Chen, Y., Li, K., Bao, W., Patel, D., Kong, Y., Min, M.R., and Metaxas, D.N. (2024). Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment. Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September\u20134 October 2024, Springer.","DOI":"10.1007\/978-3-031-73007-8_12"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Giarelis, N., Mastrokostas, C., and Karacapilidis, N. (2023). Abstractive vs. extractive summarization: An experimental review. Appl. Sci., 13.","DOI":"10.3390\/app13137620"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Pahune, S., and Akhtar, Z. (2025). Transitioning from MLOps to LLMOps: Navigating the unique challenges of large language models. Information, 16.","DOI":"10.3390\/info16020087"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"104791","DOI":"10.1016\/j.compedu.2023.104791","article-title":"Content and quantity of highlights and annotations predict learning from multiple digital texts","volume":"199","author":"List","year":"2023","journal-title":"Comput. Educ."},{"key":"ref_28","unstructured":"Mezzetti, D. (2025, December 18). Annotateai. Available online: https:\/\/github.com\/neuml\/annotateai."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Haz, A.L., Panduman, Y.Y.F., Funabiki, N., Fajrianti, E.D., and Sukaridhoto, S. (2024). Fully Open-Source Meeting Minutes Generation Tool. Future Internet, 16.","DOI":"10.3390\/fi16110429"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Gonzalez, H., Li, J., Jin, H., Ren, J., Zhang, H., Akinyele, A., Wang, A., Miltsakaki, E., Baker, R., and Callison-Burch, C. (2023, January 13). Automatically Generated Summaries of Video Lectures May Enhance Students\u00e2\u20ac\u2122 Learning Experience. Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023), Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.bea-1.31"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Warner, J., Pavel, A., Nguyen, T., Agrawala, M., and Hartmann, B. (2023, January 27\u201331). Slidespecs: Automatic and interactive presentation feedback collation. Proceedings of the 28th International Conference on Intelligent User Interfaces, Sydney, Australia.","DOI":"10.1145\/3581641.3584035"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"100417","DOI":"10.1016\/j.caeai.2025.100417","article-title":"Retrieval-augmented generation for educational application: A systematic survey","volume":"8","author":"Li","year":"2025","journal-title":"Comput. Educ. Artif. Intell."},{"key":"ref_33","first-page":"423","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Ahuja","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1038\/s41539-025-00301-w","article-title":"On opportunities and challenges of large multimodal foundation models in education","volume":"10","author":"Avila","year":"2025","journal-title":"npj Sci. Learn."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Singh, D., Gupta, A., Jawahar, C., and Tapaswi, M. (2023, January 3\u20137). Unsupervised audio-visual lecture segmentation. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV56688.2023.00520"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Lee, D.W., Ahuja, C., Liang, P.P., Natu, S., and Morency, L.P. (2023, January 2\u20136). Lecture presentations multimodal dataset: Towards understanding multimodality in educational videos. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01838"},{"key":"ref_37","unstructured":"Wright, B., Guruvayur, V., Napolitano, L., Ozar, D., Rivera, A., Sai, A., and Tafesse, B. (2025, January 26). Using Digital Textbook and Classroom Data to Explore Multimodal (Audio, Visual, & Textual) LLM Retrieval Techniques. Proceedings of the iTextbooks 2025: Sixth Workshop on Intelligent Textbooks, Palermo, Italy."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, Z., Li, C., Zhang, M., Mei, Q., and Bendersky, M. (2024). Retrieval augmented generation or long-context llms? A comprehensive study and hybrid approach. arXiv.","DOI":"10.18653\/v1\/2024.emnlp-industry.66"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"An, S., Ma, Z., Lin, Z., Zheng, N., and Lou, J.G. (2024). Make Your LLM Fully Utilize the Context. arXiv.","DOI":"10.52202\/079017-1986"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"102274","DOI":"10.1016\/j.lindif.2023.102274","article-title":"ChatGPT for good? On opportunities and challenges of large language models for education","volume":"103","author":"Kasneci","year":"2023","journal-title":"Learn. Individ. Differ."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Grassini, S. (2023). Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT. Educ. Sci., 13.","DOI":"10.3390\/educsci13070692"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Haz, A.L., Fajrianti, E.D., Funabiki, N., and Sukaridhoto, S. (2023). A Study of Audio-to-Text Conversion Software Using Whispers Model. Proceedings of the 2023 Sixth International Conference on Vocational Education and Electrical Engineering (ICVEE), Surabaya, Indonesia, 14\u201315 October 2023, IEEE.","DOI":"10.1109\/ICVEE59738.2023.10348186"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"2003601","DOI":"10.1007\/s11704-025-50058-z","article-title":"A comprehensive taxonomy of prompt engineering techniques for large language models","volume":"20","author":"Liu","year":"2025","journal-title":"Front. Comput. Sci."},{"key":"ref_44","unstructured":"White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., and Schmidt, D.C. (2023). A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Bai, X., Wu, X., Stojkovic, I., and Tsioutsiouliklis, K. (2024, January 21\u201325). Leveraging large language models for improving keyphrase generation for contextual targeting. Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, Boise, ID, USA.","DOI":"10.1145\/3627673.3680093"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"100532","DOI":"10.1016\/j.rico.2025.100532","article-title":"A comparison study on optical character recognition models in mathematical equations and in any language","volume":"18","author":"Francis","year":"2025","journal-title":"Results Control Optim."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Alroobaea, R., and Mayhew, P.J. (2014). How many participants are really enough for usability studies?. Proceedings of the 2014 Science and Information Conference, London, UK, 27\u201329 August 2014, IEEE.","DOI":"10.1109\/SAI.2014.6918171"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"108072","DOI":"10.1016\/j.knosys.2021.108072","article-title":"Information retrieval and question answering: A case study on COVID-19 scientific literature","volume":"240","author":"Otegi","year":"2022","journal-title":"Knowl.-Based Syst."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1109\/TLT.2017.2682086","article-title":"Automatic summarization of lecture slides for enhanced student previewtechnical report and user study","volume":"11","author":"Shimada","year":"2017","journal-title":"IEEE Trans. Learn. Technol."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3583558","article-title":"From anecdotal evidence to quantitative evaluation methods: A systematic review on evaluating explainable ai","volume":"55","author":"Nauta","year":"2023","journal-title":"ACM Comput. Surv."},{"key":"ref_51","first-page":"33","article-title":"Text tiling: Segmenting text into multi-paragraph subtopic passages","volume":"23","author":"Hearst","year":"1997","journal-title":"Comput. Linguist."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Ghazimatin, A., Garmash, E., Penha, G., Sheets, K., Achenbach, M., Semerci, O., Galvez, R., Tannenberg, M., Mantravadi, S., and Narayanan, D. (2024, January 21\u201325). PODTILE: Facilitating podcast episode browsing with auto-generated chapters. Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, Boise, ID, USA.","DOI":"10.1145\/3627673.3680081"},{"key":"ref_53","unstructured":"Castellano, B. (2026, January 26). PySceneDetect: Python-Based Scene Detection Program. Available online: https:\/\/github.com\/Breakthrough\/PySceneDetect."},{"key":"ref_54","first-page":"46595","article-title":"Judging llm-as-a-judge with mt-bench and chatbot arena","volume":"36","author":"Zheng","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Aydin, O., Karaarslan, E., Erenay, F.S., and Bacanin, N. (2025). Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma. arXiv.","DOI":"10.2139\/ssrn.5133368"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Bavaresco, A., Bernardi, R., Bertolazzi, L., Elliott, D., Fern\u00e1ndez, R., Gatt, A., Ghaleb, E., Giulianelli, M., Hanna, M., and Koller, A. (2024). Llms instead of human judges? A large scale empirical study across 20 nlp evaluation tasks. arXiv.","DOI":"10.18653\/v1\/2025.acl-short.20"},{"key":"ref_57","first-page":"4","article-title":"SUS\u2014A Quick and Dirty Usability Scale","volume":"189","author":"Brooke","year":"1996","journal-title":"Usability Eval. Ind."},{"key":"ref_58","first-page":"114","article-title":"Determining what individual SUS scores mean: Adding an adjective rating scale","volume":"4","author":"Bangor","year":"2009","journal-title":"J. Usability Stud."},{"key":"ref_59","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1016\/S0079-7421(02)80005-6","article-title":"Multimedia learning","volume":"Volume 41","author":"Mayer","year":"2002","journal-title":"Psychology of Learning and Motivation"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/2\/110\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T10:01:09Z","timestamp":1770890469000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/2\/110"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,1]]},"references-count":59,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2026,2]]}},"alternative-id":["a19020110"],"URL":"https:\/\/doi.org\/10.3390\/a19020110","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,1]]}}}