{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,22]],"date-time":"2026-03-22T07:16:54Z","timestamp":1774163814400,"version":"3.50.1"},"reference-count":54,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2020,6,10]],"date-time":"2020-06-10T00:00:00Z","timestamp":1591747200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Traditionally, searching for videos on popular streaming sites like YouTube is performed by taking the keywords, titles, and descriptions that are already tagged along with the video into consideration. However, the video content is not utilized for searching of the user\u2019s query because of the difficulty in encoding the events in a video and comparing them to the search query. One solution to tackle this problem is to encode the events in a video and then compare them to the query in the same space. A method of encoding meaning to a video could be video captioning. The captioned events in the video can be compared to the query of the user, and we can get the optimal search space for the videos. There have been many developments over the course of the past few years in modeling video-caption generators and sentence embeddings. In this paper, we exploit an end-to-end video captioning model and various sentence embedding techniques that collectively help in building the proposed video-searching method. The YouCook2 dataset was used for the experimentation. Seven sentence embedding techniques were used, out of which the Universal Sentence Encoder outperformed over all the other six, with a median percentile score of 99.51. Thus, this method of searching, when integrated with traditional methods, can help improve the quality of search results.<\/jats:p>","DOI":"10.3390\/sym12060992","type":"journal-article","created":{"date-parts":[[2020,6,16]],"date-time":"2020-06-16T00:50:49Z","timestamp":1592268649000},"page":"992","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Video Caption Based Searching Using End-to-End Dense Captioning and Sentence Embeddings"],"prefix":"10.3390","volume":"12","author":[{"given":"Akshay","family":"Aggarwal","sequence":"first","affiliation":[{"name":"Department of Computer Science &amp; Engineering, Bharati Vidyapeeth\u2019s College of Engineering, New Delhi 110063, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aniruddha","family":"Chauhan","sequence":"additional","affiliation":[{"name":"Department of Computer Science &amp; Engineering, Bharati Vidyapeeth\u2019s College of Engineering, New Delhi 110063, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6690-8500","authenticated-orcid":false,"given":"Deepika","family":"Kumar","sequence":"additional","affiliation":[{"name":"Department of Computer Science &amp; Engineering, Bharati Vidyapeeth\u2019s College of Engineering, New Delhi 110063, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mamta","family":"Mittal","sequence":"additional","affiliation":[{"name":"Department of Computer Science &amp; Engineering, G. B. Pant Govt. Engineering College, Okhla, New Delhi 110020, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5161-9311","authenticated-orcid":false,"given":"Sudipta","family":"Roy","sequence":"additional","affiliation":[{"name":"PRTTL, Washington University in Saint Louis, Saint Louis, MO 63110, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0117-8102","authenticated-orcid":false,"given":"Tai-hoon","family":"Kim","sequence":"additional","affiliation":[{"name":"School of Economics and Management, Beijing Jiaotong University, Beijing 100044, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,6,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Covington, P., Adams, J., and Sargin, E. (2016, January 7). Deep neural networks for youtube recommendations. Proceedings of the 10th ACM Conference on Recommender Systems, New York, NY, USA.","DOI":"10.1145\/2959100.2959190"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"ImageNet Large Scale Visual Recognition Challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 13\u201316). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Mittal, A., Kumar, D., Mittal, M., Saba, T., Abunadi, I., Rehman, A., and Roy, S. (2020). Detecting Pneumonia Using Convolutions and Dynamic Capsule Routing for Chest X-ray Images. Sensors, 20.","DOI":"10.3390\/s20041068"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Kim, T.-H., Solanki, V.S., Baraiya, H.J., Mitra, A., Shah, H., and Roy, S. (2020). A Smart, Sensible Agriculture System Using the Exponential Moving Average Model. Symmetry, 12.","DOI":"10.3390\/sym12030457"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Graves, A., Mohamed, A., and Hinton, G. (2013, January 26). Speech recognition with deep recurrent neural networks. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Hinton, G. (2012). Deep neural networks for acoustic modeling in speech recognition. IEEE Signal Process. Mag., 29.","DOI":"10.1109\/MSP.2012.2205597"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Guadarrama, S. (2013, January 1\u20138). Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition. Proceedings of the IEEE International Conference on Computer Vision, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.337"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1023\/A:1020346032608","article-title":"Natural language description of human activities from video images based on concept hierarchy of actions","volume":"50","author":"Kojima","year":"2002","journal-title":"Int. J. Comput. Vis."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","article-title":"3D convolutional neural networks for human action recognition","volume":"35","author":"Ji","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Venugopalan, S., Rohrbach, M., Donahue, J., Mooney, R., Darrell, T., and Saenko, K. (2015, January 13\u201316). Sequence to sequence-video to text. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.515"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Chen, S., and Jiang, Y.-G. (2019). Motion Guided Spatial Attention for Video Captioning, Association for the Advancement of Artificial Intelligence.","DOI":"10.1609\/aaai.v33i01.33018191"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Xu, J., Yao, T., Zhang, Y., and Mei, T. (2017, January 23\u201327). Learning multimodal attention LSTM networks for video captioning. Proceedings of the 25th ACM international conference on Multimedia, Mountain View, CA, USA.","DOI":"10.1145\/3123266.3123448"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wu, Z., Yao, T., Fu, Y., and Jiang, Y.-G. (2017). Deep learning for video classification and captioning. Frontiers of Multimedia Research, ACM.","DOI":"10.1145\/3122865.3122867"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Hershey, S., Hershey, S., Chaudhuri, S., Ellis, D.P., Gemmeke, J.F., Jansen, A., Moore, R.C., Plakal, M., Platt, D., and Saurous, R.A. (2017, January 5\u20137). CNN architectures for large-scale audio classification. Proceedings of the ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing-Proceedings, New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952132"},{"key":"ref_16","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 13\u201316). Learning spatiotemporal features with 3D convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Pancoast, S., and Akbacak, M. (2014, January 4\u20139). Softening quantization in bag-of-audio-words. Proceedings of the ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing-Proceedings, Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853821"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Pan, Y., Yao, T., Li, H., and Mei, T. (2017, January 21\u201326). Video captioning with transferred semantic attributes. Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.111"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Yao, L., Torabi, A., Cho, K., Ballas, N., Pal, C., Larochelle, H., and Courville, A. (2015, January 13\u201316). Describing videos by exploiting temporal structure. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.512"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Song, J., Gao, L., Guo, Z., Liu, W., Zhang, D., and Shen, H.T. (2017, January 19\u201325). Hierarchical LSTM with adjusted temporal attention for video captioning. Proceedings of the IJCAI International Joint Conference on Artificial Intelligence, Melbourne, Australia.","DOI":"10.24963\/ijcai.2017\/381"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Li, X., Zhao, B., and Lu, X. (2017, January 19\u201325). MAM-RNN: Multi-level attention model based RNN for video captioning. Proceedings of the IJCAI International Joint Conference on Artificial Intelligence, Melbourne, Australia.","DOI":"10.24963\/ijcai.2017\/307"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, H., Xu, Y., and Han, Y. (2018, January 22\u201326). Spotting and aggregating salient regions for video captioning. Proceedings of the MM 2018-Proceedings of the 2018 ACM Multimedia Conference, Seoul, Korea.","DOI":"10.1145\/3240508.3240677"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Ramanishka, V., Das, A., Park, D.H., Venugopalan, S., Hendricks, L.A., Rohrbach, M., and Saenko, K. (2016, January 15\u201319). Multimodal video description. Proceedings of the MM 2016-Proceedings of the 2016 ACM Multimedia Conference, Amsterdam, The Netherlands.","DOI":"10.1145\/2964284.2984066"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Hori, C., Hori, T., Lee, T.Y., Zhang, Z., Harsham, B., Hershey, J.R., Marks, T.K., and Sumi, K. (2017, January 22\u201329). Attention-Based Multimodal Fusion for Video Description. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.450"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Gkountakos, K., Dimou, A., Papadopoulos, G.T., and Daras, P. (2019, January 17\u201319). Incorporating Textual Similarity in Video Captioning Schemes. Proceedings of the 2019 IEEE International Conference on Engineering, Technology and Innovation (ICE\/ITMC), Sophia Antipolis, France.","DOI":"10.1109\/ICE.2019.8792602"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"100","DOI":"10.2307\/2346830","article-title":"Algorithm AS 136: A K-Means Clustering Algorithm","volume":"28","author":"Hartigan","year":"1979","journal-title":"Appl. Stat."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1016\/j.patrec.2019.11.003","article-title":"Exploring diverse and fine-grained caption for video by incorporating convolutional architecture into LSTM-based model","volume":"129","author":"Xiao","year":"2020","journal-title":"Pattern Recognit. Lett."},{"key":"ref_29","unstructured":"Pan, Y., Mei, T., Yao, T., Li, H., and Rui, Y. (July, January 26). Jointly modeling embedding and translation to bridge video and language. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"2045","DOI":"10.1109\/TMM.2017.2729019","article-title":"Video Captioning with Attention-Based LSTM and Semantic Consistency","volume":"19","author":"Gao","year":"2017","journal-title":"IEEE Trans. Multimed."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"305","DOI":"10.1016\/j.patrec.2020.03.001","article-title":"Video captioning with text-based dynamic attention and step-by-step learning","volume":"133","author":"Xiao","year":"2020","journal-title":"Pattern Recognit. Lett."},{"key":"ref_32","unstructured":"You, Q., Jin, H., Wang, Z., Fang, C., and Luo, J. (July, January 26). Image captioning with semantic attention. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_33","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013, January 2\u20134). Efficient Estimation of Word Representations in Vector Space. Proceedings of the International Conference on Learning Representations, Scottsdale, AZ, USA. Workshop Track Proceedings."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C.D. (2014, January 25\u201329). GloVe: Global vectors for word representation. Proceedings of the EMNLP 2014\u20132014 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, Doha, Qatar.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1162\/tacl_a_00051","article-title":"Enriching Word Vectors with Subword Information","volume":"5","author":"Bojanowski","year":"2017","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Peters, M. (2018, January 1\u20136). Deep Contextualized Word Representations. Proceedings of the NAACL-HLT 2018, Association for Computational Linguistics, New Orleans, LA, USA.","DOI":"10.18653\/v1\/N18-1202"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Howard, J., and Ruder, S. (2018, January 15\u201320). Universal Language Model Fine-tuning for Text Classification. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers), Association for Computational Linguistics, Melbourne, Australia.","DOI":"10.18653\/v1\/P18-1031"},{"key":"ref_38","unstructured":"Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019, January 2\u20137). {BERT}: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, MN, USA."},{"key":"ref_39","unstructured":"Arora, S., Liang, Y., and Ma, T. (2017, January 24\u201326). A simple but tough-to-beat baseline for sentence embeddings. Proceedings of the 5th International Conference on Learning Representations, Toulon, France."},{"key":"ref_40","unstructured":"Kiros, R. (2015). Skip-thought vectors. Advances in Neural Information Processing Systems 28, Curran Associates, Inc."},{"key":"ref_41","unstructured":"Logeswaran, L., and Lee, H. (May, January 30). An Efficient Framework for Learning Sentence Representations. Proceedings of the International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Ethayarajh, K. (2018, January 20). Unsupervised Random Walk Sentence Embeddings: A Strong but Simple Baseline. Proceedings of the Third Workshop on Representation Learning for {NLP}, Association for Computational Linguistics, Melbourne, Australia.","DOI":"10.18653\/v1\/W18-3012"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Conneau, A., Kiela, D., Schwenk, H., Barrault, L., and Bordes, A. (2017, January 7\u201311). Supervised Learning of Universal Sentence Representations from Natural Language Inference Data. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark.","DOI":"10.18653\/v1\/D17-1070"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Reimers, N., and Gurevych, I. (2019, January 3\u20137). Sentence-{BERT}: Sentence Embeddings using {S}iamese {BERT}-Networks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China.","DOI":"10.18653\/v1\/D19-1410"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Cer, D. (November, January 31). Universal Sentence Encoder for English. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Brussels, Belgium.","DOI":"10.18653\/v1\/D18-2029"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Zhou, L., Xu, C., and Corso, J.J. (2018, January 2\u20137). Towards automatic learning of procedures from web instructional videos. Proceedings of the 32nd AAAI Conference on Artificial Intelligence, AAAI 2018, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12342"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Rohrbach, A., Rohrbach, M., Tandon, N., and Schiele, B. (2015, January 8\u201310). A dataset for Movie Description. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298940"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Xu, J., Mei, T., Yao, T., and Rui, Y. (2016, January 27\u201330). MSR-VTT: A large video description dataset for bridging video and language. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.571"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Heilbron, F.C., and Niebles, J.C. (2014, January 1\u20134). Collecting and annotating human activities in web videos. Proceedings of the ICMR 2014-Proceedings of the ACM International Conference on Multimedia Retrieval 2014, Glasgow, Scotland.","DOI":"10.1145\/2578726.2578775"},{"key":"ref_50","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 6\u201311). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Zhou, L., Zhou, Y., Corso, J.J., Socher, R., and Xiong, C. (2018, January 18\u201322). End-to-End Dense Video Captioning with Masked Transformer. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00911"},{"key":"ref_52","first-page":"1","article-title":"IBM Research Report Bleu: A Method for Automatic Evaluation of Machine Translation","volume":"22176","author":"Papineni","year":"2001","journal-title":"Science (80-)"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Lavie, A., and Agarwal, A. (2007, January 23). METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. Proceedings of the Second Workshop on Statistical Machine Translation, Prague, Czech Republic.","DOI":"10.3115\/1626355.1626389"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Zitnick, C.L., and Parikh, D. (2015, January 8\u201310). CIDEr: Consensus-based image description evaluation. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/6\/992\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:37:27Z","timestamp":1760175447000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/6\/992"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,10]]},"references-count":54,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2020,6]]}},"alternative-id":["sym12060992"],"URL":"https:\/\/doi.org\/10.3390\/sym12060992","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,6,10]]}}}