{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,3]],"date-time":"2025-12-03T17:50:03Z","timestamp":1764784203924,"version":"build-2065373602"},"reference-count":49,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2019,3,6]],"date-time":"2019-03-06T00:00:00Z","timestamp":1551830400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","award":["2016R1D1A1A09916581"],"award-info":[{"award-number":["2016R1D1A1A09916581"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Daegu City","award":["DG-2017-01"],"award-info":[{"award-number":["DG-2017-01"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Semi-supervised learning is known to achieve better generalisation than a model learned solely from labelled data. Therefore, we propose a new method for estimating a pedestrian pose orientation using a soft-target method, which is a type of semi-supervised learning method. Because a convolutional neural network (CNN) based pose orientation estimation requires large numbers of parameters and operations, we apply the teacher\u2013student algorithm to generate a compressed student model with high accuracy and compactness resembling that of the teacher model by combining a deep network with a random forest. After the teacher model is generated using hard target data, the softened outputs (soft-target data) of the teacher model are used for training the student model. Moreover, the orientation of the pedestrian has specific shape patterns, and a wavelet transform is applied to the input image as a pre-processing step owing to its good spatial frequency localisation property and the ability to preserve both the spatial information and gradient information of an image. For a benchmark dataset considering real driving situations based on a single camera, we used the TUD and KITTI datasets. We applied the proposed algorithm to various driving images in the datasets, and the results indicate that its classification performance with regard to the pose orientation is better than that of other state-of-the-art methods based on a CNN. In addition, the computational speed of the proposed student model is faster than that of other deep CNNs owing to the shorter model structure with a smaller number of parameters.<\/jats:p>","DOI":"10.3390\/s19051147","type":"journal-article","created":{"date-parts":[[2019,3,7]],"date-time":"2019-03-07T10:52:22Z","timestamp":1551955942000},"page":"1147","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":12,"title":["Estimation of Pedestrian Pose Orientation Using Soft Target Training Based on Teacher\u2013Student Framework"],"prefix":"10.3390","volume":"19","author":[{"given":"DuYeong","family":"Heo","sequence":"first","affiliation":[{"name":"Department of Computer Engineering, Keimyung University, Daegu 42601, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jae Yeal","family":"Nam","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Keimyung University, Daegu 42601, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7284-0768","authenticated-orcid":false,"given":"Byoung Chul","family":"Ko","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Keimyung University, Daegu 42601, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,3,6]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1142\/S0219843613500084","article-title":"Human-robot collision avoidance using a modified social force model with body pose and face orientation","volume":"10","author":"Ratsamee","year":"2013","journal-title":"Int. J. Humanoid Robot."},{"key":"ref_2","unstructured":"Choi, J., Lee, B.-J., and Zhang, B.-K. (arXiv, 2016). Human body orientation estimation using convolutional neural network, arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Chen, C., Heili, A., and Odobez, J.-M. (2011, January 6\u201313). A joint estimation of head and body orientation cues in surveillance video. Proceedings of the IEEE Conference on Computer Vision Workshops (ICCV Workshops), Barcelona, Spain.","DOI":"10.1109\/ICCVW.2011.6130342"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1872","DOI":"10.1109\/TITS.2014.2379441","article-title":"A probabilistic framework for joint pedestrian head and body orientation estimation","volume":"16","author":"Flohr","year":"2015","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Huang, C., Zhang, G., Jiang, Z., Li, C., Wang, Y., and Wang, X. (2014, January 7\u201310). Smartphone-based indoor position and orientation tracking fusing inertial and magnetic sensing. Proceedings of the International Symposium on Wireless Personal Multimedia Communications (WPMC), Sydney, Australia.","DOI":"10.1109\/WPMC.2014.7014819"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1442","DOI":"10.1109\/TCYB.2013.2272636","article-title":"Accurate estimation of human body orientation from RGB-D sensors","volume":"43","author":"Liu","year":"2013","journal-title":"IEEE Trans. Cybern."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Sharifi, A., Harati, A., and Vahedian, A. (2014, January 29\u201330). Marker based Human Pose Estimation Using Annealed Particle Swarm Optimization with Search Space Partitioning. Proceedings of the International Conference on Computer and Knowledge Engineering (ICCKE), Mashhad, Iran.","DOI":"10.1109\/ICCKE.2014.6993366"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zhou, X., Zhu, M., Pavlakos, G., Leonardos, S., Derpanis, K.G., and Daniilidis, K. (arXiv, 2018). MonoCap: Monocular Human Motion Capture using a CNN Coupled with a Geometric Prior, arXiv.","DOI":"10.1109\/TPAMI.2018.2816031"},{"key":"ref_9","unstructured":"(2019, February 21). OptiTrack for Animation. Available online: https:\/\/optitrack.com\/motion-capture-animation\/."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Ye, M., Wang, X., Yang, R., Ren, L., and Pollefeys, M. (2011, January 6\u201313). Accurate 3D Pose Estimation from a Single Depth Image. Proceedings of the International Conference on Computer Vision (ICCV), Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126310"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Enzweiler, M., and Gavrila, D.M. (2010, January 13\u201318). Integrated pedestrian classification and orientation estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540110"},{"key":"ref_12","unstructured":"Orozco, J., Gong, S., and Xiang, T. (2019, January 7\u201310). Head pose classification in crowded scenes. Proceedings of the British Machine Vision Conference (BMVC), London, UK."},{"key":"ref_13","unstructured":"Romero, A., Ballas, N., Kahou, S.E., Chassang, A., Gatta, C., and Bengio, Y. (2015, January 7\u20139). Fitnets:Hints for thin deep nets. Proceedings of the IEEE International Conference on Learning Representations (ICLR), San Diego, CA, USA."},{"key":"ref_14","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2014, January 8\u201313). Distilling the knowledge in a neural network. Proceedings of the Advances in Neural Information Processing Systems Workshop (NIPSW), Montreal, QC, Canada."},{"key":"ref_15","unstructured":"Heo, D., Nam, J.Y., and Ko, B.C. (2019, January 22\u201325). Pedestrian\u2019s orientation estimation for collision avoidance in advanced driver assistant system. Proceedings of the International Conference on Electronics, Information, and Communication (ICEIC), Auckland, New Zealand."},{"key":"ref_16","unstructured":"Shimizu, H., and Poggio, T. (2004, January 14\u201317). Direction estimation of pedestrian from multiple still images. Proceedings of the IEEE Intelligent Vehicles Symposium (IV), Parma, Italy."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/3-540-45783-6_1","article-title":"Multimodal shape tracking with point distribution models","volume":"2449","author":"Giebel","year":"2002","journal-title":"Pattern Recognit."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1109\/TPAMI.2017.2784424","article-title":"Head and body orientation estimation using convolutional random projection forests","volume":"41","author":"Lee","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ko, B. (2018). A Brief Review of Facial Emotion Recognition Based on Visual Information. Sensors, 18.","DOI":"10.3390\/s18020401"},{"key":"ref_20","unstructured":"Hara, K., Vemulapalli, R., and Chellappa, R. (arXiv, 2017). Designing deep convolutional neural networks for continuous object orientation estimation, arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"647","DOI":"10.1016\/j.neucom.2017.07.029","article-title":"Appearance based pedestrians\u2019 head pose and body orientation estimation using deep learning","volume":"272","author":"Raza","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"13763","DOI":"10.3390\/s150613763","article-title":"Classification of potential water body using Landsat 8 OLI and combination of two boosted random forest classifiers","volume":"15","author":"Ko","year":"2015","journal-title":"Sensors"},{"key":"ref_23","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference of Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_24","first-page":"1","article-title":"Wise teachers train better DNN acoustic models","volume":"10","author":"Price","year":"2016","journal-title":"EURASIP J. Audio Speech Music Process."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"48675","DOI":"10.1109\/ACCESS.2018.2867621","article-title":"Online Tracker Optimization for Multi-Pedestrian Tracking using a Moving Vehicle Camera","volume":"6","author":"Kim","year":"2018","journal-title":"IEEE Access"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Kim, S.J., Kwak, S., and Ko, B.C. (2018). Fast Pedestrian Detection in Surveillance Video Based on Soft Target Training of Shallow Random Forest. IEEE Access.","DOI":"10.1109\/ACCESS.2019.2892425"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1141","DOI":"10.1007\/s10278-011-9380-3","article-title":"X-ray image classification using random forests with local wavelet-based CS-local binary patterns","volume":"24","author":"Ko","year":"2011","journal-title":"J. Digit. Imaging"},{"key":"ref_28","unstructured":"Hosseini, S., Lee, S.H., and Cho, N.I. (arXiv, 2018). Feeding hand-crafted features for enhancing the performance of convolutional neural networks, arXiv."},{"key":"ref_29","unstructured":"(2018, December 27). Darknet Reference Model. Available online: https:\/\/pjreddie.com\/darknet\/imagenet\/#reference."},{"key":"ref_30","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20138). ImageNet Classification with Deep Convolutional Neural Networks. Proceedings of the Advances in Neural Information Processing Systems (NIPSW), Lake Tahoe, NV, USA."},{"key":"ref_31","unstructured":"Xu, B., Wang, N., Chen, T., and Li, M. (arXiv, 2015). Empirical Evaluation of Rectified Activations in Convolutional Network, arXiv."},{"key":"ref_32","unstructured":"(2018, December 27). ImageNet. Available online: http:\/\/www.image-net.org\/."},{"key":"ref_33","unstructured":"Mishina, Y., Tsuchiya, M., and Fujiyoshi, H. (2014, January 5\u20138). Boosted Random Forest. Proceedings of the International Conference on Computer Vision Theory and Applications (ICCVTA), Lisbon, Portugal."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Jeong, M., and Ko, B.C. (2018). Driver\u2019s Facial Expression Recognition in Real-Time for Safe Driving. Sensors, 18.","DOI":"10.3390\/s18124270"},{"key":"ref_35","unstructured":"Doeniconi, C., Peng, J., and Gunopulos, D. (2000, January 27\u201330). An adaptive metric machine for pattern classification. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Denver, CO, USA."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Andriluka, M., Roth, S., and Schiele, B. (2010, January 13\u201318). Monocular 3D Pose Estimation and Tracking by Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540156"},{"key":"ref_37","unstructured":"Wang, J., and Perez, L. (arXiv, 2017). The Effectiveness of Data Augmentation in Image Classification using Deep Learning, arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 18\u201320). Are we ready for autonomous driving? The KITTI vision benchmark suite. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Baltieri, D., Vezzani, R., and Cucchiara, R. (2012, January 7\u201313). People orientation recognition by mixtures of wrapped distributions on random trees. Proceedings of the European Conference on Computer Vision (ECCV), Florence, Italy.","DOI":"10.1007\/978-3-642-33715-4_20"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"904","DOI":"10.1016\/j.imavis.2014.08.002","article-title":"Partial least squares-based human upper body orientation estimation with combined detection and tracking","volume":"32","author":"Ardiyanto","year":"2014","journal-title":"Image Vis. Comput."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Fitte-Duval, L., Mekonnen, A.A., and Lerasle, F. (2015, January 11\u201314). Upper body detection and feature set evaluation for body pose classification. Proceedings of the International Conference on Computer Vision Theory and Applications (VISAPP), Berlin, Germany.","DOI":"10.5220\/0005313104390446"},{"key":"ref_42","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very Deep Convolutional Networks for Large-Scale Image Recognition. Proceedings of the IEEE International Conference on Learning Representations (ICLR), San Diego, CA, USA."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.C. (2018, January 18\u201322). Mobilenetv2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_44","unstructured":"Iandola, F.N., Han, S., Moskewicz, M.W., Ashraf, K., Dally, W.J., and Keutzer, K. (arXiv, 2016). SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5 MB model size, arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"2232","DOI":"10.1109\/TPAMI.2015.2408347","article-title":"Multi-view and 3D deformable part models","volume":"37","author":"Pepik","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_46","unstructured":"Chen, X., Kundu, K., Zhang, Z., Ma, H., Fidler, S., and Urtasun, R. (July, January 26). Monocular 3D object detection for autonomous driving. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Choi, W., Lin, Y., and Savarese, S. (2017, January 27\u201329). Subcategory-aware convolutional neural networks for object detection. Proceedings of the IEEE Winter Conference on Applications Computer Vision (WACV), Santa Rosa, CA, USA.","DOI":"10.1109\/WACV.2017.108"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"74","DOI":"10.1109\/MITS.2018.2867526","article-title":"Fast Joint Object Detection and Viewpoint Estimation for Traffic Scene Understanding","volume":"10","author":"Guindel","year":"2018","journal-title":"IEEE Intell. Transp. Syst. Mag."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A.A. (arXiv, 2018). Image-to-Image Translation with Conditional Adversarial Networks, arXiv.","DOI":"10.1109\/CVPR.2017.632"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/5\/1147\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:36:51Z","timestamp":1760186211000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/5\/1147"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,3,6]]},"references-count":49,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2019,3]]}},"alternative-id":["s19051147"],"URL":"https:\/\/doi.org\/10.3390\/s19051147","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2019,3,6]]}}}