{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T01:19:02Z","timestamp":1779326342386,"version":"3.51.4"},"reference-count":82,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2024,6,1]],"date-time":"2024-06-01T00:00:00Z","timestamp":1717200000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>A significant number of cameras regularly generate massive amounts of data, demanding hardware, time, and labor resources to acquire, process, and monitor. Asymmetric frames within videos pose a challenge to automatic summarization of videos, making it challenging to capture key content. Developments in computer vision have accelerated the seamless capture and analysis of high-resolution video content. Video summarization (VS) has garnered considerable interest due to its ability to provide concise summaries of lengthy videos. The current literature mainly relies on a reduced set of representative features implemented using shallow sequential networks. Therefore, this work utilizes an optimal feature-assisted visual intelligence framework for representative feature selection and summarization. Initially, the empirical analysis of several features is performed, and ultimately, we adopt a fine-tuning InceptionV3 backbone for feature extraction, deviating from conventional approaches. Secondly, our strategic encoder\u2013decoder module captures complex relationships with five convolutional blocks and two convolution transpose blocks. Thirdly, we introduced a channel attention mechanism, illuminating interrelations between channels and prioritizing essential patterns to grasp complex refinement features for final summary generation. Additionally, comprehensive experiments and ablation studies validate our framework\u2019s exceptional performance, consistently surpassing state-of-the-art networks on two benchmarks (TVSum and SumMe) datasets.<\/jats:p>","DOI":"10.3390\/sym16060680","type":"journal-article","created":{"date-parts":[[2024,6,3]],"date-time":"2024-06-03T10:06:01Z","timestamp":1717409161000},"page":"680","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Effective Video Summarization Using Channel Attention-Assisted Encoder\u2013Decoder Framework"],"prefix":"10.3390","volume":"16","author":[{"given":"Faisal","family":"Alharbi","sequence":"first","affiliation":[{"name":"Quantum Technologies and Advanced Computing Institute, King Abdulaziz City for Science and Technology, Riyadh 11442, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6543-2520","authenticated-orcid":false,"given":"Shabana","family":"Habib","sequence":"additional","affiliation":[{"name":"Department of Information Technology, College of Computer, Qassim University, Buraydah 51452, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0292-7304","authenticated-orcid":false,"given":"Waleed","family":"Albattah","sequence":"additional","affiliation":[{"name":"Department of Information Technology, College of Computer, Qassim University, Buraydah 51452, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zahoor","family":"Jan","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Islamia College Peshawar, Peshawar 25000, Pakistan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1483-6207","authenticated-orcid":false,"given":"Meshari D.","family":"Alanazi","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, College of Engineering, Jouf University, Sakaka 72388, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2379-4451","authenticated-orcid":false,"given":"Muhammad","family":"Islam","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, College of Engineering, Qassim University, Buraydah 52571, Saudi Arabia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,6,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1289","DOI":"10.1007\/s11042-018-6172-5","article-title":"Visualizing the hotspots and emerging trends of multimedia big data through scientometrics","volume":"78","author":"Jin","year":"2019","journal-title":"Multimed. Tools Appl."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"2939","DOI":"10.1109\/TMM.2022.3153208","article-title":"Optimal volumetric video streaming with hybrid saliency based tiling","volume":"25","author":"Li","year":"2022","journal-title":"IEEE Trans. Multimed."},{"key":"ref_3","first-page":"81","article-title":"Digital video summarization techniques: A survey","volume":"9","author":"Workie","year":"2020","journal-title":"Int. J. Eng. Technol."},{"key":"ref_4","unstructured":"Khan, H., Huy, B.Q., Abidin, Z.U., Yoo, J., Lee, M., Seo, K.W., Hwang, D.Y., Lee, M.Y., and Suhr, J.K. (2023, January 20\u201323). A modified yolov4 network with medium-scale challenging benchmark for efficient animal detection. Proceedings of the 9th International Conference on Next Generation Computing, Danang, Vietnam."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Khan, H., Haq, I.U., Munsif, M., Khan, S.U., and Lee, M.Y. (2022). Automated wheat diseases classification framework using advanced machine learning technique. Agriculture, 12.","DOI":"10.3390\/agriculture12081226"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"27187","DOI":"10.1007\/s11042-021-10977-y","article-title":"A survey of recent work on video summarization: Approaches and techniques","volume":"80","author":"Tiwari","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1016\/j.jvcir.2018.12.009","article-title":"EVS-DK: Event video skimming using deep keyframe","volume":"58","author":"Kumar","year":"2019","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"121288","DOI":"10.1016\/j.eswa.2023.121288","article-title":"Deep multi-scale pyramidal features network for supervised video summarization","volume":"237","author":"Khan","year":"2024","journal-title":"Expert Syst. Appl."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"107567","DOI":"10.1016\/j.patcog.2020.107567","article-title":"A comprehensive survey of multi-view video summarization","volume":"109","author":"Hussain","year":"2021","journal-title":"Pattern Recognit."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"103041","DOI":"10.1109\/ACCESS.2022.3209275","article-title":"LTC-SUM: Lightweight client-driven personalized video summarization framework using 2D CNN","volume":"10","author":"Mujtaba","year":"2022","journal-title":"IEEE Access"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1109\/TII.2019.2929228","article-title":"Cloud-assisted multiview video summarization using CNN and bidirectional LSTM","volume":"16","author":"Hussain","year":"2019","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1838","DOI":"10.1109\/JPROC.2021.3117472","article-title":"Video summarization using deep neural networks: A survey","volume":"109","author":"Apostolidis","year":"2021","journal-title":"Proc. IEEE"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"694","DOI":"10.28991\/ESJ-2022-06-04-03","article-title":"External features-based approach to date grading and analysis with image processing","volume":"6","author":"Habib","year":"2022","journal-title":"Emerg. Sci. J."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zhou, K., Qiao, Y., and Xiang, T. (2018, January 2\u20137). Deep reinforcement learning for unsupervised video summarization with diversity-representativeness reward. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12255"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"207","DOI":"10.1016\/j.jvcir.2016.12.001","article-title":"Memorable and rich video summarization","volume":"42","author":"Fei","year":"2017","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2182","DOI":"10.1109\/TPAMI.2015.2511748","article-title":"Dissimilarity-based sparse subset selection","volume":"38","author":"Elhamifar","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"2711","DOI":"10.1109\/TMM.2019.2959451","article-title":"Unsupervised video summarization with cycle-consistent adversarial LSTM networks","volume":"22","author":"Yuan","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Fu, T.-J., Tai, S.-H., and Chen, H.-T. (2019, January 7\u201311). Attentive and adversarial learning for video summarization. Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA.","DOI":"10.1109\/WACV.2019.00173"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1016\/j.patrec.2010.08.004","article-title":"VSUMM: A mechanism designed to produce static video summaries and a novel evaluation method","volume":"32","author":"Lopes","year":"2011","journal-title":"Pattern Recognit. Lett."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"2126","DOI":"10.1109\/TCSVT.2018.2860797","article-title":"Action parsing-driven video summarization based on reinforcement learning","volume":"29","author":"Lei","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2654","DOI":"10.1109\/TIP.2018.2889265","article-title":"User-ranking video summarization with multi-stage spatio\u2013temporal representation","volume":"28","author":"Huang","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhang, K., Chao, W.-L., Sha, F., and Grauman, K. (2016). Video summarization with long short-term memory. Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, 11\u201314 October 2016, Springer. Proceedings, Part VII 14.","DOI":"10.1007\/978-3-319-46478-7_47"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Rochan, M., Ye, L., and Wang, Y. (2018, January 8\u201314). Video summarization using fully convolutional sequence networks. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01258-8_22"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Fajtl, J., Sokeh, H.S., Argyriou, V., Monekosso, D., and Remagnino, P. (2019). Summarizing videos with attention. Computer Vision\u2013ACCV 2018 Workshops: 14th Asian Conference on Computer Vision, Perth, Australia, 2\u20136 December 2018, Springer. Revised Selected Papers 14.","DOI":"10.1007\/978-3-030-21074-8_4"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1709","DOI":"10.1109\/TCSVT.2019.2904996","article-title":"Video summarization with attention-based encoder\u2013decoder networks","volume":"30","author":"Ji","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.neucom.2021.09.015","article-title":"Video summarization with a dual-path attentive network","volume":"467","author":"Liang","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhao, B., Li, X., and Lu, X. (2017, January 23\u201327). Hierarchical recurrent neural network for video summarization. Proceedings of the 25th ACM International Conference on Multimedia, Mountain View, CA, USA.","DOI":"10.1145\/3123266.3123328"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"105667","DOI":"10.1016\/j.engappai.2022.105667","article-title":"A review on video summarization techniques","volume":"118","author":"Meena","year":"2023","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"296","DOI":"10.1109\/TCSVT.2004.841694","article-title":"Video summarization and scene detection by graph modeling","volume":"15","author":"Ngo","year":"2005","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1718","DOI":"10.1016\/j.neucom.2009.09.022","article-title":"Feature extraction and clustering for dynamic video summarisation","volume":"73","author":"Zhou","year":"2010","journal-title":"Neurocomputing"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"298-1","DOI":"10.2352\/EI.2023.35.9.IPAS-298","article-title":"Deep learning based speech emotion recognition for Parkinson patient","volume":"35","author":"Khan","year":"2023","journal-title":"Electron. Imaging"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"120391","DOI":"10.1016\/j.eswa.2023.120391","article-title":"Deep learning based active learning technique for data annotation and improve the overall performance of classification models","volume":"228","author":"Amin","year":"2023","journal-title":"Expert Syst. Appl."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Islam, M., Aloraini, M., Aladhadh, S., Habib, S., Khan, A., Alabdulatif, A., and Alanazi, T.M. (2023). Toward a Vision-Based Intelligent System: A Stacked Encoded Deep Learning Framework for Sign Language Recognition. Sensors, 23.","DOI":"10.3390\/s23229068"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1765","DOI":"10.1109\/TNNLS.2020.2991083","article-title":"Deep attentive video summarization with distribution consistency learning","volume":"32","author":"Ji","year":"2020","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"948","DOI":"10.1109\/TIP.2020.3039886","article-title":"Dsnet: A flexible detect-to-summarize network for video summarization","volume":"30","author":"Zhu","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1016\/j.ins.2019.12.084","article-title":"Learning reinforced attentional representation for end-to-end visual tracking","volume":"517","author":"Gao","year":"2020","journal-title":"Inf. Sci."},{"key":"ref_37","first-page":"8537","article-title":"Discriminative feature learning for unsupervised video summarization","volume":"33","author":"Jung","year":"2019","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Zhao, B., Li, X., and Lu, X. (2018, January 18\u201323). Hsa-rnn: Hierarchical structure-adaptive rnn for video summarization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00773"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Habib, S., Albattah, W., Alsharekh, M.F., Islam, M., Shees, M.M., and Sherazi, H.I. (2023). Computer Network Redundancy Reduction Using Video Compression. Symmetry, 15.","DOI":"10.3390\/sym15061280"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"3652","DOI":"10.1109\/TIP.2017.2695887","article-title":"A general framework for edited video and raw video summarization","volume":"26","author":"Li","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Li, Y., Wang, L., Yang, T., and Gong, B. (2018, January 8\u201314). How local is the local diversity? reinforcing sequential determinantal point processes with dynamic ground sets for supervised video summarization. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01237-3_10"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Mahasseni, B., Lam, M., and Todorovic, S. (2017, January 21\u201326). Unsupervised video summarization with adversarial lstm networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.318"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"He, X., Hua, Y., Song, T., Zhang, Z., Xue, Z., Ma, R., Robertson, N.M., and Guan, H. (2019, January 21\u201325). Unsupervised video summarization with attentive conditional generative adversarial networks. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/3343031.3351056"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1016\/j.neucom.2016.11.011","article-title":"Graph coloring based surveillance video synopsis","volume":"225","author":"He","year":"2017","journal-title":"Neurocomputing"},{"key":"ref_45","first-page":"2793","article-title":"Reconstructive sequence-graph network for video summarization","volume":"44","author":"Zhao","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Park, J., Lee, J., Kim, I.-J., and Soh, K. (2020). Sumgraph: Video summarization via recursive graph modeling. Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, 23\u201328 August 2020, Springer. Proceedings, Part XXV 16.","DOI":"10.1007\/978-3-030-58595-2_39"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Wang, J., Bai, Y., Long, Y., Hu, B., Chai, Z., Guan, Y., and Wei, X. (2020, January 12\u201316). Query twice: Dual mixture attention meta learning for video summarization. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3414064"},{"key":"ref_48","unstructured":"Liu, Y.-T., Li, Y.-J., and Wang, Y.-C.F. (December, January 30). Transforming multi-concept attention into video summarization. Proceedings of the Asian Conference on Computer Vision, Kyoto, Japan."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"197","DOI":"10.1016\/j.neucom.2019.07.108","article-title":"Video summarization via block sparse dictionary selection","volume":"378","author":"Ma","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"522","DOI":"10.1016\/j.patcog.2014.08.002","article-title":"Video summarization via minimum sparse reconstruction","volume":"48","author":"Mei","year":"2015","journal-title":"Pattern Recognit."},{"key":"ref_51","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the NIPS 2017, Long Beach, CA, USA."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Gygli, M., Grabner, H., Riemenschneider, H., and Van Gool, L. (2014, January 6\u201312). Creating summaries from user videos. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10584-0_33"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Khan, K., Khan, R.U., Albattah, W., Nayab, D., Qamar, A.M., Habib, S., and Islam, M. (2021). Crowd Counting Using End-to-End Semantic Image Segmentation. Electronics, 10.","DOI":"10.3390\/electronics10111293"},{"key":"ref_54","unstructured":"Munsif, M., Khan, H., Khan, Z.A., Hussain, A., Ullah, F.U.M., Lee, M.Y., and Baik, S.W. (2022, January 6\u20138). Pv-anet: Attention-based network for short-term photovoltaic power forecasting. Proceedings of the The 8th International Conference on Next Generation Computing, Jeju, Republic of Korea."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Ul Amin, S., Ullah, M., Sajjad, M., Cheikh, F.A., Hijji, M., Hijji, A., and Muhammad, K. (2022). EADN: An efficient deep learning model for anomaly detection in videos. Mathematics, 10.","DOI":"10.3390\/math10091555"},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"3939","DOI":"10.32604\/csse.2023.034805","article-title":"An Efficient Attention-Based Strategy for Anomaly Detection in Surveillance Video","volume":"46","author":"Kim","year":"2023","journal-title":"Comput. Syst. Sci. Eng."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Husman, M.A., Albattah, W., Abidin, Z.Z., Mustafah, Y.M., Kadir, K., Habib, S., Islam, M., and Khan, S. (2021). Unmanned Aerial Vehicles for Crowd Monitoring and Analysis. Electronics, 10.","DOI":"10.3390\/electronics10232974"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.-Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_59","unstructured":"Hu, J., Shen, L., Albanie, S., Sun, G., and Vedaldi, A. (2018, January 3\u20138). Gather-excite: Exploiting feature context in convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Montr\u00e9al, QC, Canada."},{"key":"ref_60","first-page":"31","article-title":"Modified YOLOv4S based on Deep learning with Feature Fusion and Spatial Attention","volume":"12","author":"Hwang","year":"2021","journal-title":"J. Korea Converg. Soc."},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"83185","DOI":"10.1109\/ACCESS.2021.3086839","article-title":"A modified generative adversarial network using spatial and channel-wise attention for CS-MRI reconstruction","volume":"9","author":"Li","year":"2021","journal-title":"IEEE Access"},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"2990","DOI":"10.1109\/TMM.2020.2965434","article-title":"Spatio-temporal attention networks for action recognition and detection","volume":"22","author":"Li","year":"2020","journal-title":"IEEE Trans. Multimed."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Habib, S., Khan, I., Islam, M., Albattah, W., Alyahya, S.M., Khan, S., and Hassan, M.K. (2021, January 6\u20137). Wavelet Frequency Transformation for Specific Weeds Recognition. Proceedings of the 1st International Conference on Artificial Intelligence and Data Analytics (CAIDA), Riyadh, Saudi Arabia.","DOI":"10.1109\/CAIDA51941.2021.9425249"},{"key":"ref_64","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1007\/s00799-005-0129-9","article-title":"Keyframe-based video summarization using delaunay clustering","volume":"6","author":"Mundur","year":"2006","journal-title":"Int. J. Digit. Libr."},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Gygli, M., Chao, W.-L., Grauman, K., and Sha, F. (2014). Creating summaries from user videos. Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland, 6\u201312 September 2014, Springer. Proceedings, Part VII 13.","DOI":"10.1007\/978-3-319-10584-0_33"},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Gygli, M., Grabner, H., and Van Gool, L. (2015, January 7\u201312). Video summarization by learning submodular mixtures of objectives. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298928"},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Potapov, D., Douze, M., Harchaoui, Z., and Schmid, C. (2014). Category-specific video summarization. Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland, 6\u201312 September 2014, Springer. Proceedings, Part VI 13.","DOI":"10.1007\/978-3-319-10599-4_35"},{"key":"ref_68","unstructured":"Song, Y., Vallmitjana, J., Stent, A., and Jaimes, A. (2015, January 7\u201312). Tvsum: Summarizing web videos using titles. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_69","doi-asserted-by":"crossref","first-page":"8467","DOI":"10.1109\/TIP.2020.3016431","article-title":"An efficient fire detection method based on multiscale feature extraction, implicit deep supervision and channel attention mechanism","volume":"29","author":"Li","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_70","first-page":"640","article-title":"Fully convolutional networks for semantic segmentation","volume":"39","author":"Long","year":"2015","journal-title":"Proc. IEEE Conf. Comput. Vis. Pattern Recognit."},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Habib, S., Hussain, A., Islam, M., Khan, S., and Albattah, W. (2021, January 6\u20137). Towards Efficient Detection and Crowd Management for Law Enforcing Agencies. Proceedings of the 1st International Conference on Artificial Intelligence and Data Analytics (CAIDA), Riyadh, Saudi Arabia.","DOI":"10.1109\/CAIDA51941.2021.9425076"},{"key":"ref_72","doi-asserted-by":"crossref","first-page":"107677","DOI":"10.1016\/j.patcog.2020.107677","article-title":"Exploring global diverse attention via pairwise temporal relation for video summarization","volume":"111","author":"Li","year":"2021","journal-title":"Pattern Recognit."},{"key":"ref_73","doi-asserted-by":"crossref","first-page":"108312","DOI":"10.1016\/j.patcog.2021.108312","article-title":"Learning multiscale hierarchical attention for video summarization","volume":"122","author":"Zhu","year":"2022","journal-title":"Pattern Recognit."},{"key":"ref_74","doi-asserted-by":"crossref","unstructured":"An, Y., and Zhao, S. (2022, January 7\u20139). SHTVS: Shot-level based Hierarchical Transformer for Video Summarization. Proceedings of the 2022 the 5th International Conference on Image and Graphics Processing (ICIGP), Beijing, China.","DOI":"10.1145\/3512388.3512427"},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Jiang, H., and Mu, Y. (2022, January 18\u201324). Joint video summarization and moment localization by cross-task sample transfer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01590"},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Habib, S., Alsanea, M., Aloraini, M., Al-Rawashdeh, H.S., Islam, M., and Khan, S. (2022). An Efficient and Effective Deep Learning-Based Model for Real-Time Face Mask Detection. Sensors, 22.","DOI":"10.3390\/s22072602"},{"key":"ref_77","doi-asserted-by":"crossref","unstructured":"Apostolidis, E., Balaouras, G., Mezaris, V., and Patras, I. (2022, January 27\u201330). Summarizing videos using concentrated attention and considering the uniqueness and diversity of the video frames. Proceedings of the 2022 International Conference on Multimedia Retrieval, Newark, NJ, USA.","DOI":"10.1145\/3512527.3531404"},{"key":"ref_78","doi-asserted-by":"crossref","unstructured":"Elfeki, M., and Borji, A. (2019, January 7\u201311). Video summarization via actionness ranking. Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA.","DOI":"10.1109\/WACV.2019.00085"},{"key":"ref_79","doi-asserted-by":"crossref","first-page":"577","DOI":"10.1109\/TCSVT.2019.2890899","article-title":"A novel key-frames selection framework for comprehensive video summarization","volume":"30","author":"Huang","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_80","doi-asserted-by":"crossref","first-page":"2359","DOI":"10.1016\/j.procs.2023.01.211","article-title":"Attention over attention: An enhanced supervised video summarization approach","volume":"218","author":"Puthige","year":"2023","journal-title":"Procedia Comput. Sci."},{"key":"ref_81","doi-asserted-by":"crossref","first-page":"3629","DOI":"10.1109\/TIE.2020.2979573","article-title":"TTH-RNN: Tensor-train hierarchical recurrent neural network for video summarization","volume":"68","author":"Zhao","year":"2020","journal-title":"IEEE Trans. Ind. Electron."},{"key":"ref_82","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1016\/j.patrec.2020.12.016","article-title":"Self-attention binary neural tree for video summarization","volume":"143","author":"Fu","year":"2021","journal-title":"Pattern Recognit. Lett."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/16\/6\/680\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:52:26Z","timestamp":1760107946000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/16\/6\/680"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,1]]},"references-count":82,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2024,6]]}},"alternative-id":["sym16060680"],"URL":"https:\/\/doi.org\/10.3390\/sym16060680","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,1]]}}}