{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T01:35:26Z","timestamp":1760232926564,"version":"build-2065373602"},"reference-count":41,"publisher":"MDPI AG","issue":"24","license":[{"start":{"date-parts":[[2022,12,9]],"date-time":"2022-12-09T00:00:00Z","timestamp":1670544000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Ministry of Science and ICT","award":["IITP-2022-2020-0-01791","NRF-2021R1F1A1062131","202201760001"],"award-info":[{"award-number":["IITP-2022-2020-0-01791","NRF-2021R1F1A1062131","202201760001"]}]},{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","award":["IITP-2022-2020-0-01791","NRF-2021R1F1A1062131","202201760001"],"award-info":[{"award-number":["IITP-2022-2020-0-01791","NRF-2021R1F1A1062131","202201760001"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Dong-eui University","award":["IITP-2022-2020-0-01791","NRF-2021R1F1A1062131","202201760001"],"award-info":[{"award-number":["IITP-2022-2020-0-01791","NRF-2021R1F1A1062131","202201760001"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In this paper, we propose an intra-picture prediction method for depth video by a block clustering through a neural network. The proposed method solves a problem that the block that has two or more clusters drops the prediction performance of the intra prediction for depth video. The proposed neural network consists of both a spatial feature prediction network and a clustering network. The spatial feature prediction network utilizes spatial features in vertical and horizontal directions. The network contains a 1D CNN layer and a fully connected layer. The 1D CNN layer extracts the spatial features for a vertical direction and a horizontal direction from a top block and a left block of the reference pixels, respectively. 1D CNN is designed to handle time-series data, but it can also be applied to find the spatial features by regarding a pixel order in a certain direction as a timestamp. The fully connected layer predicts the spatial features of the block to be coded through the extracted features. The clustering network finds clusters from the spatial features which are the outputs of the spatial feature prediction network. The network consists of 4 CNN layers. The first 3 CNN layers combine two spatial features in the vertical and horizontal directions. The last layer outputs the probabilities that pixels belong to the clusters. The pixels of the block are predicted by the representative values of the clusters that are the average of the reference pixels belonging to the clusters. For the intra prediction for various block sizes, the block is scaled to the size of the network input. The prediction result through the proposed network is scaled back to the original size. In network training, the mean square error is used as a loss function between the original block and the predicted block. A penalty for output values far from both ends is introduced to the loss function for clear network clustering. In the simulation results, the bit rate is saved by up to 12.45% under the same distortion condition compared with the latest video coding standard.<\/jats:p>","DOI":"10.3390\/s22249656","type":"journal-article","created":{"date-parts":[[2022,12,9]],"date-time":"2022-12-09T06:14:00Z","timestamp":1670566440000},"page":"9656","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Intra Prediction Method for Depth Video Coding by Block Clustering through Deep Learning"],"prefix":"10.3390","volume":"22","author":[{"given":"Dong-seok","family":"Lee","sequence":"first","affiliation":[{"name":"AI Grand ICT Research Center, Dong-eui University, Busan 47340, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6358-4985","authenticated-orcid":false,"given":"Soon-kak","family":"Kwon","sequence":"additional","affiliation":[{"name":"Department of Computer Software Engineering, Dong-eui University, Busan 47340, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,12,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1110","DOI":"10.1109\/TMM.2013.2246148","article-title":"Robust part-based hand gesture recognition using kinect sensor","volume":"15","author":"Ren","year":"2013","journal-title":"IEEE Trans. Multimed."},{"key":"ref_2","unstructured":"Li, Y., Miao, Q., Tian, K., Fan, Y., Xu, X., Li, R., and Song, J. (2016, January 4\u20138). Large-Scale Gesture Recognition with a Fusion of RGB-D Data Based on the C3D model. Proceedings of the International Conference on Pattern Recognition, Cancun, Mexico."},{"key":"ref_3","first-page":"50","article-title":"Lidar for autonomous driving: The principles, challenges, and trends for automotive lidar and perception systems","volume":"37","author":"Li","year":"2020","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Feng, D., Rosenbaum, L., and Dietmayer, K. (2018, January 4\u20137). Towards Safe Autonomous Driving: Capture Uncertainty in the Deep Neural Network for Lidar 3d Vehicle Detection. Proceedings of the International Conference on Intelligent Transportation Systems, Maui, HI, USA.","DOI":"10.1109\/ITSC.2018.8569814"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1521","DOI":"10.1109\/JPROC.2021.3062590","article-title":"MPEG immersive video coding standard","volume":"109","author":"Boyce","year":"2021","journal-title":"Proc. IEEE"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"186","DOI":"10.1016\/j.jvcir.2005.05.010","article-title":"Overview of H.264\/MPEG-4 part 10","volume":"17","author":"Kwon","year":"2006","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"3736","DOI":"10.1109\/TCSVT.2021.3101953","article-title":"Overview of the versatile video coding (VVC) standard and its applications","volume":"31","author":"Bross","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"29153","DOI":"10.1109\/ACCESS.2021.3056687","article-title":"Efficient depth data coding method based on plane modeling for intra prediction","volume":"9","author":"Lee","year":"2021","journal-title":"IEEE Access"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Hu, Y., Yang, W., Xia, S., and Liu, J. (2018, January 9\u201312). Optimized Spatial Recurrent Network for Intra Prediction in Video Coding. Proceedings of the Visual Communications and Image Processing, Taichung, Taiwan.","DOI":"10.1109\/VCIP.2018.8698658"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"3024","DOI":"10.1109\/TMM.2019.2920603","article-title":"Progressive spatial recurrent neural network for intra prediction","volume":"21","author":"Hu","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"755","DOI":"10.1109\/TETCI.2020.3028581","article-title":"Artificial neural networks in action for an automated cell-type classification of biological neural networks","volume":"5","author":"Troullinou","year":"2021","journal-title":"IEEE Trans. Emerg. Top. Comput. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"e12770","DOI":"10.2196\/12770","article-title":"Deep learning approaches to detect atrial fibrillation using photoplethysmographic signals: Algorithms development study","volume":"7","author":"Kwon","year":"2019","journal-title":"JMIR Mhealth Uhealth"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Toderici, G., O\u2019Malley, S.M., Hwang, S.J., Vincent, D., Minnen, D., Baluja, S., Covell, M., and Sukthankar, R. (2016, January 2\u20134). Variable Rate Image Compression with Recurrent Neural Networks. Proceedings of the International Conference on Learning Representations, San Juan, Puerto Rico.","DOI":"10.1109\/CVPR.2017.577"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"3292","DOI":"10.1109\/TPAMI.2020.2988453","article-title":"An end-to-end learning framework for video compression","volume":"43","author":"Lu","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Cheng, Z., Sun, H., Takeuchi, M., and Katto, J. (2018, January 24\u201327). Deep Convolutional Autoencoder-Based Lossy Image Compression. Proceedings of the Picture Coding Symposium, San Francisco, CA, USA.","DOI":"10.1109\/PCS.2018.8456308"},{"key":"ref_16","unstructured":"Ball\u00e9, J., Laparra, V., and Simoncelli, E.P. (2017, January 24\u201326). End-to-End Optimized Image Compression. Proceedings of the International Conference on Learning Representations, Toulon, France."},{"key":"ref_17","unstructured":"Theis, L., Shi, W., Cunningham, A., and Husz\u00e1r, F. (2017, January 24\u201326). Lossy Image Compression with Compressive Autoencoders. Proceedings of the International Conference on Learning Representations, Toulon, France."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Schiopu, I., Liu, Y., and Munteanu, A. (2018, January 24\u201327). CNN-Based Prediction for Lossless Coding of Photographic Images. Proceedings of the Picture Coding Symposium, San Francisco, CA, USA.","DOI":"10.1109\/PCS.2018.8456311"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1032","DOI":"10.1049\/el.2018.0889","article-title":"Residual-error prediction based on deep learning for lossless image compression","volume":"54","author":"Schiopu","year":"2018","journal-title":"Electron. Lett."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"3236","DOI":"10.1109\/TIP.2018.2817044","article-title":"Fully connected network-based intra prediction for image coding","volume":"27","author":"Li","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_21","first-page":"1816","article-title":"CNN-based intra-prediction for lossless HEVC","volume":"30","author":"Schiopu","year":"2020","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1109\/JSTSP.2020.3034768","article-title":"Intra-frame coding using a conditional autoencoder","volume":"15","author":"Brand","year":"2021","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1109\/TMM.2019.2924591","article-title":"Generative adversarial network-based intra prediction for video coding","volume":"22","author":"Zhu","year":"2020","journal-title":"IEEE Trans. Multimed."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhong, G., Wang, J., Hu, J., and Liang, F. (2021). A GAN-based video intra coding. Electronics, 10.","DOI":"10.3390\/electronics10020132"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"147","DOI":"10.33851\/JMIS.2021.8.3.147","article-title":"CNN-based fast split mode decision algorithm for versatile video coding (VVC) inter prediction","volume":"8","author":"Yeo","year":"2021","journal-title":"J. Multimed. Inf. Syst."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Yokoyama, R., Tahara, M., Takeuchi, M., Sun, H., Matsuo, Y., and Katto, J. (2020, January 4\u20136). CNN Based Optimal Intra Prediction Mode Estimation in Video Coding. Proceedings of the IEEE International Conference on Consumer Electronics, Las Vegas, NV, USA.","DOI":"10.1109\/ICCE46568.2020.9043170"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Lee, Y.W., Kim, J.H., Choi, Y.J., and Kim, B.G. (2018, January 12\u201314). CNN-Based Approach for Visual Quality Improvement on HEVC. Proceedings of the IEEE International Conference on Consumer Electronics, Las Vegas, NV, USA.","DOI":"10.1109\/ICCE.2018.8326088"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1109\/JETCAS.2018.2885981","article-title":"Emerging MPEG standards for point cloud compression","volume":"9","author":"Schwarz","year":"2019","journal-title":"IEEE J. Emerg. Sel. Top. Circuits Syst."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Garcia, D.C., and de Queiroz, R.L. (2018, January 7\u201310). Intra-Frame Context-Based Octree Coding for Point-Cloud Geometry. Proceedings of the IEEE International Conference on Image, Athens, Greece.","DOI":"10.1109\/ICIP.2018.8451802"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Kathariya, B., Li, L., Li, Z., Alvarez, J., and Chen, J. (2018, January 23\u201327). Scalable Point Cloud Geometry Coding with Binary Tree Embedded Quadtree. Proceedings of the IEEE International Conference on Multimedia and Expo, San Diego, CA, USA.","DOI":"10.1109\/ICME.2018.8486481"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"3886","DOI":"10.1109\/TIP.2017.2707807","article-title":"Motion-compensated compression of dynamic voxelized point clouds","volume":"26","author":"Chou","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Nenci, F., Spinello, L., and Stachniss, C. (2014, January 14\u201318). Effective Compression of Range Data Streams for Remote Robot Operations Using H. 264. In Proceedings of IEEE\/RSJ International Conference on Intelligent Robots and Systems, Chicago, IL, USA.","DOI":"10.1109\/IROS.2014.6943095"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, X., \u015eekercio\u011flu, Y., Drummond, T., Fr\u00e9mont, V., Natalizio, E., and Fantoni, I. (2018). Relative pose based redundancy removal: Collaborative RGB-D data transmission in mobile visual sensor networks. Sensors, 18.","DOI":"10.3390\/s18082430"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015, January 7\u201313). Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.123"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. (2012, January 7\u201313). Indoor Segmentation and Support Inference from RGBD Images. Proceedings of the European Conference on Computer Vision, Firenze, Italy.","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Lai, K., Bo, L., Ren, X., and Fox, D. (2011, January 9\u201313). A Large-Scale Hierarchical Multi-View RGB-D Object Dataset. Proceedings of the International Conference on Robotics and Automation, Shanghai, China.","DOI":"10.1109\/ICRA.2011.5980382"},{"key":"ref_37","unstructured":"Lai, K., Bo, L., and Fox, D. (June, January 31). Unsupervised Feature Learning for 3D Scene Labeling. Proceedings of the International Conference on Robotics and Automation, Hong Kong, China."},{"key":"ref_38","unstructured":"(2022, October 31). Versatile Video Coding (VVC). Available online: https:\/\/jvet.hhi.fraunhofer.de."},{"key":"ref_39","unstructured":"Ruhnke, M., Bo, L., Fox, D., and Burgard, W. (June, January 31). Hierarchical Sparse Coded Surface Models. Proceedings of the IEEE International Conference on Robotics and Automation, Hong Kong, China."},{"key":"ref_40","unstructured":"Choi, S.J., Zhou, Q.Y., and Koltun, V. (2015, January 7\u201312). Robust Reconstruction of Indoor Scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"2132","DOI":"10.1109\/LRA.2019.2900747","article-title":"A novel point cloud compression algorithm based on clustering","volume":"4","author":"Sun","year":"2019","journal-title":"IEEE Robot. Autom. Lett."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/24\/9656\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:37:14Z","timestamp":1760146634000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/24\/9656"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,9]]},"references-count":41,"journal-issue":{"issue":"24","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["s22249656"],"URL":"https:\/\/doi.org\/10.3390\/s22249656","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2022,12,9]]}}}