{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T20:16:42Z","timestamp":1778617002483,"version":"3.51.4"},"reference-count":47,"publisher":"MDPI AG","issue":"16","license":[{"start":{"date-parts":[[2023,8,9]],"date-time":"2023-08-09T00:00:00Z","timestamp":1691539200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Science and Technology Council, Taiwan","award":["NSTC 112-2221-E-027-079-MY2"],"award-info":[{"award-number":["NSTC 112-2221-E-027-079-MY2"]}]},{"name":"University of Cincinnati, Cincinnati, OH","award":["NSTC 112-2221-E-027-079-MY2"],"award-info":[{"award-number":["NSTC 112-2221-E-027-079-MY2"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>This paper introduces a transformer encoding linker network (TELNet) for automatically identifying scene boundaries in videos without prior knowledge of their structure. Videos consist of sequences of semantically related shots or chapters, and recognizing scene boundaries is crucial for various video processing tasks, including video summarization. TELNet utilizes a rolling window to scan through video shots, encoding their features extracted from a fine-tuned 3D CNN model (transformer encoder). By establishing links between video shots based on these encoded features (linker), TELNet efficiently identifies scene boundaries where consecutive shots lack links. TELNet was trained on multiple video scene detection datasets and demonstrated results comparable to other state-of-the-art models in standard settings. Notably, in cross-dataset evaluations, TELNet demonstrated significantly improved results (F-score). Furthermore, TELNet\u2019s computational complexity grows linearly with the number of shots, making it highly efficient in processing long videos.<\/jats:p>","DOI":"10.3390\/s23167050","type":"journal-article","created":{"date-parts":[[2023,8,9]],"date-time":"2023-08-09T10:30:48Z","timestamp":1691577048000},"page":"7050","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Video Scene Detection Using Transformer Encoding Linker Network (TELNet)"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5017-3159","authenticated-orcid":false,"given":"Shu-Ming","family":"Tseng","sequence":"first","affiliation":[{"name":"Department of Electronic Engineering, National Taipei University of Technology, Taipei 106335, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhi-Ting","family":"Yeh","sequence":"additional","affiliation":[{"name":"College of Engineering and Applied Science, University of Cincinnati, Cincinnati, OH 45219, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chia-Yang","family":"Wu","sequence":"additional","affiliation":[{"name":"Department of Electronic Engineering, National Taipei University of Technology, Taipei 106335, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7269-4975","authenticated-orcid":false,"given":"Jia-Bin","family":"Chang","sequence":"additional","affiliation":[{"name":"College of Engineering and Applied Science, University of Cincinnati, Cincinnati, OH 45219, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1701-2668","authenticated-orcid":false,"given":"Mehdi","family":"Norouzi","sequence":"additional","affiliation":[{"name":"College of Engineering and Applied Science, University of Cincinnati, Cincinnati, OH 45219, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,8,9]]},"reference":[{"key":"ref_1","unstructured":"Wang, H., Neumann, J., and Choi, J. (2021). Determining Video Highlights and Chaptering. (11,172,272), U.S. Patent."},{"key":"ref_2","unstructured":"Jindal, A., and Bedi, A. (2020). Extracting Session Information from Video Content to Facilitate Seeking. (10,701,434), U.S. Patent."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Otani, M., Nakashima, Y., Rahtu, E., and Heikkila, J. (2019, January 15\u201320). Rethinking the evaluation of video summaries. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00778"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1007\/s005300050138","article-title":"Constructing table-of-content for videos","volume":"7","author":"Rui","year":"1999","journal-title":"Multimed. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"28","DOI":"10.1109\/MSP.2006.1621446","article-title":"Video shot detection and condensed representation. a review","volume":"23","author":"Cotsaces","year":"2006","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Abdulhussain, S.H., Ramli, A.R., Saripan, M.I., Mahmmod, B.M., Al-Haddad, S.A.R., and Jassim, W.A. (2018). Methods and challenges in shot boundary detection: A review. Entropy, 20.","DOI":"10.3390\/e20040214"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Chavate, S., Mishra, R., and Yadav, P. (2021, January 8\u201310). A Comparative Analysis of Video Shot Boundary Detection using Different Approaches. Proceedings of the 2021 IEEE 10th International Conference on System Modeling & Advancement in Research Trends (SMART), Moradabad, India.","DOI":"10.1109\/SMART52563.2021.9676246"},{"key":"ref_8","first-page":"119","article-title":"Video shot boundary detection: A review","volume":"Volume 2","author":"Pal","year":"2015","journal-title":"Emerging ICT for Bridging the Future\u2014Proceedings of the 49th Annual Convention of the Computer Society of India CSI"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"15623","DOI":"10.1007\/s11042-018-6959-4","article-title":"Correlation based feature fusion for the temporal video scene segmentation task","volume":"78","author":"Kishi","year":"2019","journal-title":"Multimed. Tools Appl."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Caba Heilbron, F., Escorcia, V., Ghanem, B., and Carlos Niebles, J. (2015, January 7\u201312). Activitynet: A large-scale video benchmark for human activity understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"955","DOI":"10.1109\/TMM.2016.2644872","article-title":"Recognizing and presenting the storytelling video structure with deep multimodal networks","volume":"19","author":"Baraldi","year":"2016","journal-title":"IEEE Trans. Multimed."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"17487","DOI":"10.1007\/s11042-020-10450-2","article-title":"Temporal video scene segmentation using deep-learning","volume":"80","author":"Trojahn","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Baraldi, L., Grana, C., and Cucchiara, R. (2015, January 26). A deep siamese network for scene detection in broadcast videos. Proceedings of the 23rd ACM International Conference on Multimedia, Brisbane, Australia.","DOI":"10.1145\/2733373.2806316"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Rotman, D., Yaroker, Y., Amrani, E., Barzelay, U., and Ben-Ari, R. (2020, January 12\u201316). Learnable optimal sequential grouping for video scene detection. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413612"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"3559","DOI":"10.1109\/TCSVT.2020.3042476","article-title":"Adaptive Context Reading Network for Movie Scene Detection","volume":"31","author":"Liu","year":"2020","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1163","DOI":"10.1109\/TCSVT.2011.2138830","article-title":"Temporal video segmentation to scenes using high-level audiovisual features","volume":"21","author":"Sidiropoulos","year":"2011","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"991","DOI":"10.1007\/s11760-018-1244-6","article-title":"Using deep features for video scene detection and annotation","volume":"12","author":"Protasov","year":"2018","journal-title":"Signal. Image Video Process."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Pei, Y., Wang, Z., Chen, H., Huang, B., and Tu, W. (2021, January 7\u20139). Video scene detection based on link prediction using graph convolution network. Proceedings of the 2nd ACM International Conference on Multimedia in Asia, Singapore.","DOI":"10.1145\/3444685.3446293"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1016\/j.procs.2020.08.002","article-title":"Video scenes segmentation based on multimodal genre prediction","volume":"176","author":"Bouyahi","year":"2020","journal-title":"Procedia Comput. Sci."},{"key":"ref_20","unstructured":"Rotman, D., Porat, D., and Ashour, G. (2017). Proceedings of the 2017 IEEE 19th International Workshop on Multimedia Signal Processing (MMSP), Luton, UK, 16\u201318 October 2017, IEEE."},{"key":"ref_21","unstructured":"Son, J.W., Lee, A., Kwak, C.U., and Kim, S.J. (2020). Proceedings of the 2020 IEEE\/WIC\/ACM International Joint Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT), IEEE."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Chu, W.S., Song, Y., and Jaimes, A. (2015, January 7\u201312). Video co-summarization: Video summarization by visual co-occurrence. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298981"},{"key":"ref_23","first-page":"63","article-title":"Event detection based approach for soccer video summarization using machine learning","volume":"7","author":"Zawbaa","year":"2012","journal-title":"Int. J. Multimed. Ubiquitous Eng."},{"key":"ref_24","unstructured":"Potapov, D., Douze, M., Harchaoui, Z., and Schmid, C. (2014). Proceedings of the European Conference on Computer Vision, Springer."},{"key":"ref_25","unstructured":"Sokeh, H.S., Argyriou, V., Monekosso, D., and Remagnino, P. (2018). Proceedings of the 2018 24th International Conference on Pattern Recognition (ICPR), Beijing, China, 20\u201324 August 2018, IEEE."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"6999","DOI":"10.1109\/TNNLS.2021.3084827","article-title":"A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects","volume":"33","author":"Li","year":"2022","journal-title":"IEEE Trans. Neural Networks Learn. Syst."},{"key":"ref_27","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012). Imagenet classification with deep convolutional neural networks. Adv. Neural Inf. Process. Syst., 25."},{"key":"ref_28","unstructured":"Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., and Oliva, A. (2014). Learning deep features for scene recognition using places database. Adv. Neural Inf. Process. Syst., 27."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","article-title":"3D convolutional neural networks for human action recognition","volume":"35","author":"Ji","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning spatiotemporal features with 3d convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Washington, DC, USA.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017, January 21\u201326). Quo vadis, action recognition? a new model and the kinetics dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Hara, K., Kataoka, H., and Satoh, Y. (2017, January 22\u201329). Learning spatio-temporal features with 3d residual networks for action recognition. Proceedings of the IEEE International Conference on Computer Vision Workshops, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.373"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Chen, S., Nie, X., Fan, D., Zhang, D., Bhat, V., and Hamid, R. (2021, January 20\u201325). Shot contrastive self-supervised learning for scene boundary detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00967"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wu, H., Chen, K., Luo, Y., Qiao, R., Ren, B., Liu, H., Xie, W., and Shen, L. (2022, January 18\u201324). Scene Consistency Representation Learning for Video Scene Segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01363"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"580","DOI":"10.1109\/76.767124","article-title":"Automated high-level movie segmentation for advanced video-retrieval systems","volume":"9","author":"Hanjalic","year":"1999","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"89","DOI":"10.1109\/TMM.2008.2008924","article-title":"Scene detection in videos using shot clustering and sequence alignment","volume":"11","author":"Chasanis","year":"2008","journal-title":"IEEE Trans. Multimed."},{"key":"ref_37","first-page":"1150","article-title":"Object recognition from local scale-invariant features","volume":"Volume 2","author":"Lowe","year":"1999","journal-title":"Proceedings of the 7th IEEE international Conference on Computer Vision, Kerkyra, Greece, 20\u201327 September 1999"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"2564963","DOI":"10.1155\/2018\/2564963","article-title":"Video scene detection using compact bag of visual word models","volume":"2018","author":"Haroon","year":"2018","journal-title":"Adv. Multimed."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Trojahn, T.H., Kishi, R.M., and Goularte, R. (2018, January 16\u201319). A new multimodal deep-learning model to video scene segmentation. Proceedings of the 24th Brazilian Symposium on Multimedia and the Web, Salvador, Brazil.","DOI":"10.1145\/3243082.3243108"},{"key":"ref_40","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. Adv. Neural Inf. Process. Syst., 30."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"White, S., and Smyth, P. (2005, January 21\u201423). A spectral clustering approach to finding communities in graphs. Proceedings of the 2005 SIAM International Conference on Data Mining, Newport Beach, CA, USA.","DOI":"10.1137\/1.9781611972757.25"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"888","DOI":"10.1109\/34.868688","article-title":"Normalized cuts and image segmentation","volume":"22","author":"Shi","year":"2000","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Islam, M.M., Hasan, M., Athrey, K.S., Braskich, T., and Bertasius, G. (2023, January 23\u201328). Efficient Movie Scene Detection Using State-Space Transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Portland, OR, USA.","DOI":"10.1109\/CVPR52729.2023.01798"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., and Fei-Fei, L. (2014, January 23\u201328). Large-scale video classification with convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Washington, DC, USA.","DOI":"10.1109\/CVPR.2014.223"},{"key":"ref_45","unstructured":"Movieclips (2020, September 01). Movieclips YouTube Channel. Available online: http:\/\/www.youtube.com\/user\/movieclips."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"492","DOI":"10.1109\/TMM.2002.802021","article-title":"Systematic evaluation of logical story unit segmentation","volume":"4","author":"Vendrig","year":"2002","journal-title":"IEEE Trans. Multimed."},{"key":"ref_47","unstructured":"Wikipedia Contributors (2022, February 17). Segmentation-Based Object Categorization \u2014 Wikipedia, The Free Encyclopedia. Available online: https:\/\/en.wikipedia.org\/wiki\/Segmentation-based_object_categorization."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/16\/7050\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:29:46Z","timestamp":1760128186000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/16\/7050"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,9]]},"references-count":47,"journal-issue":{"issue":"16","published-online":{"date-parts":[[2023,8]]}},"alternative-id":["s23167050"],"URL":"https:\/\/doi.org\/10.3390\/s23167050","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,9]]}}}