{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T18:33:17Z","timestamp":1775068397998,"version":"3.50.1"},"reference-count":70,"publisher":"MDPI AG","issue":"18","license":[{"start":{"date-parts":[[2022,9,16]],"date-time":"2022-09-16T00:00:00Z","timestamp":1663286400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Matsuo Institute"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Human action recognition and detection from unmanned aerial vehicles (UAVs), or drones, has emerged as a popular technical challenge in recent years, since it is related to many use case scenarios from environmental monitoring to search and rescue. It faces a number of difficulties mainly due to image acquisition and contents, and processing constraints. Since drones\u2019 flying conditions constrain image acquisition, human subjects may appear in images at variable scales, orientations, and occlusion, which makes action recognition more difficult. We explore low-resource methods for ML (machine learning)-based action recognition using a previously collected real-world dataset (the \u201cOkutama-Action\u201d dataset). This dataset contains representative situations for action recognition, yet is controlled for image acquisition parameters such as camera angle or flight altitude. We investigate a combination of object recognition and classifier techniques to support single-image action identification. Our architecture integrates YoloV5 with a gradient boosting classifier; the rationale is to use a scalable and efficient object recognition system coupled with a classifier that is able to incorporate samples of variable difficulty. In an ablation study, we test different architectures of YoloV5 and evaluate the performance of our method on Okutama-Action dataset. Our approach outperformed previous architectures applied to the Okutama dataset, which differed by their object identification and classification pipeline: we hypothesize that this is a consequence of both YoloV5 performance and the overall adequacy of our pipeline to the specificities of the Okutama dataset in terms of bias\u2013variance tradeoff.<\/jats:p>","DOI":"10.3390\/s22187020","type":"journal-article","created":{"date-parts":[[2022,9,19]],"date-time":"2022-09-19T04:49:22Z","timestamp":1663562962000},"page":"7020","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":42,"title":["Detecting Human Actions in Drone Images Using YoloV5 and Stochastic Gradient Boosting"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8108-7915","authenticated-orcid":false,"given":"Tasweer","family":"Ahmad","sequence":"first","affiliation":[{"name":"Department of Electrical and Computer Engineering, COMSATS University Islamabad, Islamabad 45550, Pakistan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6113-9696","authenticated-orcid":false,"given":"Marc","family":"Cavazza","sequence":"additional","affiliation":[{"name":"National Institute of Informatics, Tokyo 101-8430, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yutaka","family":"Matsuo","sequence":"additional","affiliation":[{"name":"Department of Engineering, The University of Tokyo, Tokyo 113-8654, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Helmut","family":"Prendinger","sequence":"additional","affiliation":[{"name":"National Institute of Informatics, Tokyo 101-8430, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,9,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Girish, D., Singh, V., and Ralescu, A. (2020, January 14\u201319). Understanding action recognition in still images. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00193"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Eweiwi, A., Cheema, M.S., Bauckhage, C., and Gall, J. (2014). Efficient pose-based action recognition. Asian Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-16814-2_28"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Wang, C., Wang, Y., and Yuille, A.L. (2013, January 23\u201328). An approach to pose-based action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.123"},{"key":"ref_4","first-page":"196","article-title":"An efficient human action recognition framework with pose-based spatiotemporal features","volume":"23","author":"Agahian","year":"2020","journal-title":"Eng. Sci. Technol. Int. J."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhao, Z., Ma, H., and You, S. (2017, January 22\u201329). Single image action recognition using semantic body part actions. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.367"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"563","DOI":"10.1016\/j.procs.2018.10.432","article-title":"Action recognition in still images using residual neural network features","volume":"143","author":"Sreela","year":"2018","journal-title":"Procedia Comput. Sci."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Liu, L., Tan, R.T., and You, S. (2018). Loss guided activation for action recognition in still images. Asian Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-030-20873-8_10"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wu, D., Sharma, N., and Blumenstein, M. (2017, January 14\u201319). Recent advances in video-based human action recognition using deep learning: A review. Proceedings of the 2017 International Joint Conference on Neural Networks (IJCNN), Anchorage, AK, USA.","DOI":"10.1109\/IJCNN.2017.7966210"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2259","DOI":"10.1007\/s10462-020-09904-8","article-title":"A survey on video-based human action recognition: Recent updates, datasets, challenges, and applications","volume":"54","author":"Pareek","year":"2021","journal-title":"Artif. Intell. Rev."},{"key":"ref_10","unstructured":"Pham, H.H., Khoudour, L., Crouzil, A., Zegers, P., and Velastin, S.A. (2022). Video-based human action recognition using deep learning: A review. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Rohrbach, M., Amin, S., Andriluka, M., and Schiele, B. (2012, January 16\u201321). A database for fine grained activity detection of cooking activities. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6247801"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Singh, B., Marks, T.K., Jones, M., Tuzel, O., and Shao, M. (2016, January 27\u201330). A multi-stream bi-directional recurrent neural network for fine-grained action detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.216"},{"key":"ref_13","unstructured":"Yeung, S., Russakovsky, O., Mori, G., and Fei-Fei, L. (July, January 26). End-to-end learning of action detection from frame glimpses in videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_14","unstructured":"Zhang, D., Shao, Y., Mei, Y., Chu, H., Zhang, X., Zhan, H., and Rao, Y. (2018, January 12\u201314). Using YOLO-based pedestrian detection for monitoring UAV. Proceedings of the Tenth International Conference on Graphics and Image Processing, Chengdu, China."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Yang, Z., Huang, Z., Yang, Y., Yang, F., and Yin, Z. (2018, January 8\u201311). Accurate specified-pedestrian tracking from unmanned aerial vehicles. In Proceeding of the IEEE 18th International Conference on Communication Technology, Chongqing, China.","DOI":"10.1109\/ICCT.2018.8600173"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Liu, M., Wang, X., Zhou, A., Fu, X., Ma, Y., and Piao, C. (2020). Uav-yolo: Small object detection on unmanned aerial vehicle perspective. Sensors, 20.","DOI":"10.3390\/s20082238"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1040","DOI":"10.1016\/j.imavis.2020.104046","article-title":"Deep learning-based object detection in low-altitude UAV datasets: A survey","volume":"104","author":"Mittal","year":"2020","journal-title":"Image Vis. Comput."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"831","DOI":"10.1016\/j.procs.2018.07.112","article-title":"YOLO based human action recognition and localization","volume":"133","author":"Shinde","year":"2018","journal-title":"Procedia Comput. Sci."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1016\/j.cviu.2014.06.014","article-title":"Evaluation of video activity localizations integrating quality and quantity measurements","volume":"127","author":"Wolf","year":"2014","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Jung, H.K., and Choi, G.S. (2022). Improved YoloV5: Efficient Object Detection Using Drone Images under Various Conditions. Appl. Sci., 12.","DOI":"10.3390\/app12147255"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Caputo, S., Castellano, G., Greco, F., Mencar, C., Petti, N., and Vessio, G. (2022). Human Detection in Drone Images Using YOLO for Search-and-Rescue Operations. International Conference of the Italian Association for Artificial Intelligence, Springer.","DOI":"10.1007\/978-3-031-08421-8_22"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_24","unstructured":"Redmon, J., and Farhadi, A. (2018). YoloV3: An incremental improvement. arXiv."},{"key":"ref_25","unstructured":"Bochkovskiy, A., Wang, C., and Liao, H.M. (2020). YoloV4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"288","DOI":"10.1109\/TPAMI.2008.284","article-title":"Human action recognition in videos using kinematic features and multiple instance learning","volume":"32","author":"Ali","year":"2008","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Cao, L., Liu, Z., and Huang, T.S. (2010, January 13\u201318). Cross-dataset action detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539875"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wang, H., and Schmid, C. (2013, January 1\u20138). Action recognition with improved trajectories. Proceedings of the IEEE International Conference on Computer Vision, Sydney, NSW, Australia.","DOI":"10.1109\/ICCV.2013.441"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Sultani, W., and Saleemi, I. (2014, January 23\u201328). Human action recognition across datasets by foreground-weighted histogram decomposition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.103"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017, January 21\u201326). Quo vadis, action recognition? a new model and the kinetics dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Zhou, X., Liu, S., Pavlakos, G., Kumar, V., and Daniilidis, K. (2018, January 21\u201325). Human motion capture using a drone. Proceedings of the IEEE International Conference on Robotics and Automation, Brisbane, QLD, Australia.","DOI":"10.1109\/ICRA.2018.8462830"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"121212","DOI":"10.1109\/ACCESS.2019.2937344","article-title":"Human action recognition in unconstrained trimmed videos using residual attention network and joints path signature","volume":"7","author":"Ahmad","year":"2019","journal-title":"J. IEEE Access"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"389","DOI":"10.1016\/j.neucom.2020.10.096","article-title":"Skeleton-based action recognition using sparse spatio-temporal GCN with edge effective resistance","volume":"423","author":"Ahmad","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1109\/TAI.2021.3076974","article-title":"Graph Convolutional Neural Network for Human Action Recognition: A Comprehensive Survey","volume":"2","author":"Ahmad","year":"2021","journal-title":"IEEE Trans. Artif. Intell."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"103186","DOI":"10.1016\/j.cviu.2021.103186","article-title":"Human action recognition in drone videos using a few aerial training examples","volume":"206","author":"Sultani","year":"2021","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_36","unstructured":"(2022, May 05). Ucf-Arg Data Set. Available online: Https:\/\/www.crcv.ucf.edu\/data\/UCF-ARG.php."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Perera, A., Wei, L., and Chahl, J. (2018, January 8\u201314). UAV-GESTURE: A dataset for UAV control and gesture recognition. Proceedings of the European Conference on Computer Vision Workshops, Munich, Germany.","DOI":"10.1007\/978-3-030-11012-3_9"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Ding, M., Li, N., Song, Z., Zhang, R., Zhang, X., and Zhou, H. (2020, January 14\u201316). A Lightweight Action Recognition Method for Unmanned-Aerial-Vehicle Video. Proceedings of the IEEE 3rd International Conference on Electronics and Communication Engineering, Xi\u2019an, China.","DOI":"10.1109\/ICECE51594.2020.9353008"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"122583","DOI":"10.1109\/ACCESS.2019.2938249","article-title":"UAV-based situational awareness system using deep learning","volume":"7","author":"Geraldes","year":"2019","journal-title":"J. IEEE Access"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"107140","DOI":"10.1016\/j.patcog.2019.107140","article-title":"Human activity recognition from UAV-captured video sequences","volume":"100","author":"Mliki","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Choi, J., Sharma, G., Chandraker, M., and Huang, J. (2020, January 1\u20135). Unsupervised and semi-supervised domain adaptation for action recognition from drones. Proceedings of the IEEE Winter Conference on Applications of Computer Vision, Snowmass, CO, USA.","DOI":"10.1109\/WACV45572.2020.9093511"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Barekatain, M., Mart\u00ed, M., Shih, H., Murray, S., Nakayama, K., Matsuo, Y., and Prendinger, H. (2017, January 21\u201326). Okutama-Action: An aerial view video dataset for concurrent human action detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.267"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015). Fast r-cnn. arXiv.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_45","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster r-cnn: Towards real-time object detection with region proposal networks. Proceedings of the Conference Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C., and Berg, A.C. (2016, January 11\u201314). Ssd: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Lin, T., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_48","unstructured":"(2022, June 05). YoloV5 Documentation. Available online: Https:\/\/docs.ultralytics.com\/."},{"key":"ref_49","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_52","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_53","unstructured":"Tan, M., and Le, Q. (2019, January 9\u201315). Efficientnet: Rethinking model scaling for convolutional neural networks. Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Lin, T., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201323). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Ghiasi, G., Lin, T., and Le, Q.V. (2019, January 13\u201319). Nas-fpn: Learning scalable feature pyramid architecture for object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR.2019.00720"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020, January 14\u201319). Efficientdet: Scalable and efficient object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_58","unstructured":"Liu, S., Huang, D., and Wang, Y. (2019). Learning spatial fusion for single-shot object detection. arXiv."},{"key":"ref_59","unstructured":"Zhao, Q., Sheng, T., Wang, Y., Tang, Z., Chen, Y., Cai, L., and Ling, H. (February, January 27). M2det: A single-shot object detector based on multi-level feature pyramid network. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Moghimi, M., Belongie, S.J., Saberian, M.J., Yang, J., Vasconcelos, N., and Li, L.J. (2016, January 19\u201322). Boosted convolutional neural networks. Proceedings of the British Machine Vision Conference, York, UK.","DOI":"10.5244\/C.30.24"},{"key":"ref_61","doi-asserted-by":"crossref","first-page":"1189","DOI":"10.1214\/aos\/1013203451","article-title":"Greedy function approximation: A gradient boosting machine","volume":"29","author":"Friedman","year":"2001","journal-title":"Ann. Stat."},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"367","DOI":"10.1016\/S0167-9473(01)00065-2","article-title":"Stochastic gradient boosting","volume":"38","author":"Friedman","year":"2002","journal-title":"Comput. Stat. Data Anal."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Chen, T., and Guestrin, C. (2016, January 13\u201317). XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM Sigkdd International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA.","DOI":"10.1145\/2939672.2939785"},{"key":"ref_64","doi-asserted-by":"crossref","first-page":"e586","DOI":"10.7717\/peerj-cs.586","article-title":"An amalgamation of YoloV4 and XGBoost for next-gen smart traffic management system","volume":"7","author":"Dave","year":"2021","journal-title":"PeerJ Comput. Sci."},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Caruana, R., and Niculescu-Mizil, A. (2006, January 25\u201329). An empirical comparison of supervised learning algorithms. Proceedings of the 23rd international conference on Machine learning, Pittsburgh, PA, USA.","DOI":"10.1145\/1143844.1143865"},{"key":"ref_66","unstructured":"(2022, March 15). sklearn.ensemble.GradientBoostingClassifier. Available online: Https:\/\/scikit-learn.org\/stable\/modules\/generated\/sklearn.ensemble.GradientBoostingClassifier.html."},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., Doll\u00e1r, P., Tu, Z., and He, K. (2017, January 21\u201326). Aggregated residual transformations for deep neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_68","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2022, January 10). Automatic Differentiation in Pytorch. Available online: Https:\/\/pytorch.org\/."},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Soleimani, A., and Nasrabadi, N.M. (2018, January 10\u201313). Convolutional neural networks for aerial multi-label pedestrian detection. Proceedings of the IEEE 21st International Conference on Information Fusion, Cambridge, UK.","DOI":"10.23919\/ICIF.2018.8455494"},{"key":"ref_70","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L. (2018, January 18\u201323). Mobilenetv2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/18\/7020\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:32:48Z","timestamp":1760142768000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/18\/7020"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,16]]},"references-count":70,"journal-issue":{"issue":"18","published-online":{"date-parts":[[2022,9]]}},"alternative-id":["s22187020"],"URL":"https:\/\/doi.org\/10.3390\/s22187020","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,16]]}}}