{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:22:13Z","timestamp":1760239333241,"version":"build-2065373602"},"reference-count":56,"publisher":"MDPI AG","issue":"21","license":[{"start":{"date-parts":[[2020,11,9]],"date-time":"2020-11-09T00:00:00Z","timestamp":1604880000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The Bag-of-Words (BoW) framework has been widely used in action recognition tasks due to its compact and efficient feature representation. Various modifications have been made to this framework to increase its classification power. This often results in an increased complexity and reduced efficiency. Inspired by the success of image-based scale coded BoW representations, we propose a spatio-temporal scale coded BoW (SC-BoW) for video-based recognition. This involves encoding extracted multi-scale information into BoW representations by partitioning spatio-temporal features into sub-groups based on the spatial scale from which they were extracted. We evaluate SC-BoW in two experimental setups. We first present a general pipeline to perform real-time action recognition with SC-BoW. Secondly, we apply SC-BoW onto the popular Dense Trajectory feature set. Results showed SC-BoW representations to successfully improve performance by 2\u20137% with low added computational cost. Notably, SC-BoW on Dense Trajectories outperformed more complex deep learning approaches. Thus, scale coding is a low-cost and low-level encoding scheme that increases classification power of the standard BoW without compromising efficiency.<\/jats:p>","DOI":"10.3390\/s20216380","type":"journal-article","created":{"date-parts":[[2020,11,10]],"date-time":"2020-11-10T14:10:41Z","timestamp":1605017441000},"page":"6380","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Spatio-Temporal Scale Coded Bag-of-Words"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2226-1554","authenticated-orcid":false,"given":"Divina","family":"Govender","sequence":"first","affiliation":[{"name":"School of Engineering, University of KwaZulu-Natal, Durban 4041, South Africa"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5915-0511","authenticated-orcid":false,"given":"Jules-Raymond","family":"Tapamo","sequence":"additional","affiliation":[{"name":"School of Engineering, University of KwaZulu-Natal, Durban 4041, South Africa"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,11,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Piergiovanni, A., Angelova, A., Toshev, A., and Ryoo, M.S. (2018, January 30\u201331). Evolving space-time neural architectures for videos. Proceedings of the 2018 IEEE International Conference on Computer Vision, Instanbul, Turkey.","DOI":"10.1109\/ICCV.2019.00188"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1016\/j.cviu.2016.03.013","article-title":"Bag of Visual Words and Fusion Methods for Action Recognition: Comprehensive Study and Good Practice","volume":"150","author":"Peng","year":"2016","journal-title":"Comput. Vis. Image Understanding"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1007\/s11263-012-0594-8","article-title":"Dense trajectories and motion boundary descriptors for action recognition","volume":"103","author":"Wang","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Passalis, N., and Tefas, A. (2017, January 22\u201329). Learning Bag-of-Features Pooling for Deep Convolutional Neural Networks. Proceedings of the 2017 IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.614"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1016\/j.patrec.2017.12.024","article-title":"A Bag of Expression Framework for Improved Human Action Recognition","volume":"103","author":"Nazir","year":"2018","journal-title":"Pattern Recognit. Lett."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Shi, F., Petriu, E., and Laganiere, R. (2013, January 23\u201328). Sampling Strategies for Real-Time Action Recognition. Proceedings of the 26th IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.335"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Shi, F., Petriu, E.M., and Cordeiro, A. (2011, January 14\u201317). Human action recognition from Local Part Model. Proceedings of the 2011 IEEE International Workshop on Haptic Audio Visual Environments and Games, Qinhuangdao, China.","DOI":"10.1109\/HAVE.2011.6088408"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Van Opdenbosch, D., Oelsch, M., Garcea, A., and Steinbach, E. (2017, January 10\u201313). A joint compression scheme for local binary feature descriptors and their corresponding bag-of-words representation. Proceedings of the 2017 IEEE Visual Communications and Image Processing (VCIP), St. Petersburg, FL, USA.","DOI":"10.1109\/VCIP.2017.8305155"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2326","DOI":"10.1109\/TIP.2018.2791180","article-title":"Real-time Action Recognition with Deeply Transferred Motion Vector CNNs","volume":"27","author":"Zhang","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Chen, X., Zhou, W., Jiang, X., and Liu, Y. (2019, January 4\u20139). Real-time Human Action Recognition Based on Person Detection. Proceedings of the 2019 IEEE International Conference on Real-time Computing and Robotics (RCAR), Irkutsk, Russia.","DOI":"10.1109\/RCAR47638.2019.9043967"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"222","DOI":"10.1007\/s11263-013-0636-x","article-title":"Image Classification with the Fisher Vector: Theory and Practice","volume":"105","author":"Sanchez","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_12","first-page":"832","article-title":"Action Recognition with deep network features and dimension reduction","volume":"13","author":"Li","year":"2019","journal-title":"KSII Trans. Internet Inf. Syst."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1007\/s00138-017-0871-1","article-title":"Scale coding bag of deep features for human attribute and action recognition","volume":"29","author":"Khan","year":"2018","journal-title":"Mach. Vis. Appl."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wang, H., and Schmid, C. (2013, January 3\u20136). Action recognition with improved trajectories. Proceedings of the IEEE 2013 International Conference on Computer Vision, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.441"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1016\/j.neucom.2016.09.106","article-title":"Action recognition by saliency-based dense sampling","volume":"236","author":"Xu","year":"2017","journal-title":"Neurocomputing"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Khan, F.S., Weijer, J.v.d., Bagdanov, A.D., and Felsber, M. (2014, January 24\u201328). Scale Coding Bag-of-Words for Action Recognition. Proceedings of the 22nd International Conference on Pattern Recognition, Stockholm, Sweden.","DOI":"10.1109\/ICPR.2014.269"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Scovanner, P., Ali, S., and Shah, M. (2007, January 24\u201329). A 3-dimensional SIFT Descriptor and its Application to Action Recognition. Proccedings of the 15th ACM Conference on Multimedia, Augsburg, Germany.","DOI":"10.1145\/1291233.1291311"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Hu, J., Xia, G.S., Hu, F., Sun, H., and Zhang, L. (2015, January 26\u201331). A Comparative Study of Sampling Analysis in Scene Classification of High-resolution Remote Sensing Imagery. Proceedings of the 2015 IEEE International geoscience and remote sensing symposium (IGARSS), Milan, Italy.","DOI":"10.1109\/IGARSS.2015.7326290"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Nowak, E., Jurie, F., and Triggs, B. (2006, January 7\u201313). Sampling strategies for bag-of-features image classification. Proceedings of the 9th European Conference on Computer Vision, Graz, Austria.","DOI":"10.1007\/11744085_38"},{"key":"ref_20","unstructured":"Willems, G., Tuytelaars, T., and Van Gool, L. (2006, January 7\u201313). An Efficient Dense and Scale-invariant Spatio-temporal Interest Point Detector. Proceedings of the 9th European Conference on Computer Vision, Graz, Austria."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Yeffet, L., and Wolf, L. (October, January 27). Local trinary patterns for human action recognition. Proceedings of the 2009 IEEE 12th International Conference on Computer Vision, Kyoto, Japan.","DOI":"10.1109\/ICCV.2009.5459201"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Klaser, A., Marsza\u0142ek, M., and Schmid, C. (2008, January 1\u20134). A Spatio-temporal Descriptor Based on 3D-gradients. Proceedings of the 19th British Machine Vision Conference, Leeds, UK.","DOI":"10.5244\/C.22.99"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Sch\u00fcldt, C., Laptev, I., and Caputo, B. (2004, January 26). Recognizing Human Actions: A Local SVM Approach. Proceedings of the 17th International Conference on Pattern Recognition, Cambridge, UK.","DOI":"10.1109\/ICPR.2004.1334462"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., and Serre, T. (2011, January 6\u201313). HMDB: A large video database for human motion recognition. Proceedings of the 13th International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1109\/TCSVT.2010.2041828","article-title":"Contextual Bag-of-Words for Visual Categorization","volume":"21","author":"Li","year":"2011","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Filliat, D. (2007, January 10\u201317). A Visual Bag of Words Method for Interactive Qualitative Localization and Mapping. Proceedings of the 2007 IEEE International Conference on Robotics and Automation, Roma, Italy.","DOI":"10.1109\/ROBOT.2007.364080"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1433","DOI":"10.1109\/TIP.2017.2778561","article-title":"Contextual Bag-of-Words for Robust Visual Tracking","volume":"27","author":"Zeng","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_28","unstructured":"O\u2019Hara, S., and Draper, B.A. (2011). Introduction to the bag of features paradigm for image classification and retrieval. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Loussaief, S., and Abdelkrim, A. (2018, January 22\u201325). Deep learning vs. bag of features in machine learning for image classification. Proceedings of the 2nd International Conference on Advanced Systems and Electric Technologies, Hammamet, Tunisia.","DOI":"10.1109\/ASET.2018.8379825"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wang, H., Kl\u00e4ser, A., Schmid, C., and Liu, C.L. (2011, January 21\u201323). Action Recognition by Dense Trajectories. Proceedings of the IEEE 2011 Computer Vision and Pattern Recognition, Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995407"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1007\/s11263-005-1838-7","article-title":"On Space-Time Interest Points","volume":"64","author":"Laptev","year":"2005","journal-title":"Int. J. Comput. Vis."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Laptev, I., Marsza\u0142ek, M., Schmid, C., and Rozenfeld, B. (2008, January 23\u201328). Learning realistic human actions from movies. Proceedings of the CVPR 2008-IEEE Conference on Computer Vision and Pattern Recognition, Anchorage, AK, USA.","DOI":"10.1109\/CVPR.2008.4587756"},{"key":"ref_33","unstructured":"Lazebnik, S., Schmid, C., and Ponce, J. (2006, January 17\u201322). Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, New York, NY, USA."},{"key":"ref_34","unstructured":"Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N.D., and Weinberger, K.Q. (2014). Two-stream convolutional networks for action recognition in videos. Advances in Neural Information Processing Systems 27, Curran Associates, Inc."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Sehgal, S. (2018, January 20\u201321). Human Activity Recognition Using BPNN Classifier on HOG Features. Proceedings of the 2018 International Conference on Intelligent Circuits and Systems, Phagwara, India.","DOI":"10.1109\/ICICS.2018.00065"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., and Fei-Fei, L. (2014, January 24\u201327). Large-scale video classification with Convolutional Neural Networks. Proceedings of the 2014 IEEE conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.223"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., and Zisserman, A. (2016, January 27\u201330). Convolutional Two-stream Network Fusion for Video Action Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.213"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"8585","DOI":"10.1007\/s00521-019-04365-9","article-title":"Human action recognition with bag of visual words using different machine learning methods and hyperparameter optimization","volume":"32","author":"Aslan","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Baccouche, M., Mamalet, F., Wolf, C., Garcia, C., and Baskurt, A. (2011, January 16). Sequential deep learning for Human Action Recognition. Proceedings of the 2nd International Workshop on Human Behavior Understanding, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-642-25446-8_4"},{"key":"ref_40","first-page":"1","article-title":"Introduction to Real-time Digital Signal Processing","volume":"Volume 1","author":"Kuo","year":"2013","journal-title":"Real-Time Digital Signal Processing: Fundamentals, Implementations and Applications"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"6","DOI":"10.1109\/5.259423","article-title":"Real-time computing: A new discipline of computer science and engineering","volume":"82","author":"Shin","year":"1994","journal-title":"Proc. IEEE"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 24\u201327). Rich feature hierarchies for Accurate Object Detection and Semantic Segmentation. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (July, January 26). You only look once: Unified, real-time object detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_44","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Danelljan, M., H\u00e4ger, G., Khan, F., and Felsberg, M. (2014, January 1\u20135). Accurate scale estimation for robust visual tracking. Proceedings of the 2014 British Machine Vision Conference, Nottingham, UK.","DOI":"10.5244\/C.28.65"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"583","DOI":"10.1109\/TPAMI.2014.2345390","article-title":"High-speed tracking with kernelized correlation filters","volume":"37","author":"Henriques","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object detection with discriminatively trained part-based models","volume":"32","author":"Felzenszwalb","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1007\/s11263-006-9794-4","article-title":"Local features and kernels for classification of texture and object categories: A comprehensive study","volume":"73","author":"Zhang","year":"2007","journal-title":"Int. J. Comput. Vis."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"35","DOI":"10.7551\/mitpress\/4057.003.0004","article-title":"A primer on kernel methods","volume":"47","author":"Vert","year":"2004","journal-title":"Kernel Methods Comput. Biology"},{"key":"ref_50","unstructured":"Singh, D., Bhure, A., Mamtani, S., Mohan, C.K., and Kandi, S. (2018, January 2\u20136). Fast-BoW: Scaling Bag-of-Visual-Words Generation. Proceedings of the 2018 British Machine Vision Conference, Newcastle, UK."},{"key":"ref_51","unstructured":"Shi, J., and Tomasi, C. (1994, January 21\u201323). Good Features to Track. Proceedings of the 1994 IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Govender, D., and Tapamo, J.R. (2020, January 1\u20134). Factors Affecting the Cost to Accuracy Balance for Real-Time Video-Based Action Recognition. Proceedings of the 20th International Conference on Computational Science and Applications, University of Calgari, Calgari, Italy (held online).","DOI":"10.1007\/978-3-030-58799-4_58"},{"key":"ref_53","unstructured":"Farneb\u00e4ck, G. (July, January 29). Two-frame motion estimation based on polynomial expansion. Proceedings of the Scandinavian Conference on Image analysis, Berlin, Heidelberg."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Arandjelovi\u0107, R., and Zisserman, A. (2012, January 16\u201321). Three things everyone should know to improve object retrieval. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248018"},{"key":"ref_55","unstructured":"Basha, S., Pulabaigari, V., and Mukherjee, S. (2020). An Information-rich Sampling Technique over Spatio-Temporal CNN for Classification of Human Actions in Videos. arXiv."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Bolme, D.S., Beveridge, J.R., Draper, B.A., and Lui, Y.M. (2010, January 13\u201318). Visual object tracking using adaptive correlation filters. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539960"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/21\/6380\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:31:00Z","timestamp":1760178660000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/21\/6380"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,11,9]]},"references-count":56,"journal-issue":{"issue":"21","published-online":{"date-parts":[[2020,11]]}},"alternative-id":["s20216380"],"URL":"https:\/\/doi.org\/10.3390\/s20216380","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2020,11,9]]}}}