{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,14]],"date-time":"2026-06-14T07:41:10Z","timestamp":1781422870695,"version":"3.54.1"},"reference-count":33,"publisher":"MDPI AG","issue":"16","license":[{"start":{"date-parts":[[2019,8,17]],"date-time":"2019-08-17T00:00:00Z","timestamp":1566000000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Samsung Research Funding Center of Samsung Electronics","award":["SRFC-TB1703-02"],"award-info":[{"award-number":["SRFC-TB1703-02"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Interaction forces are traditionally predicted by a contact type haptic sensor. In this paper, we propose a novel and practical method for inferring the interaction forces between two objects based only on video data\u2014one of the non-contact type camera sensors\u2014without the use of common haptic sensors. In detail, we could predict the interaction force by observing the texture changes of the target object by an external force. For this purpose, our hypothesis is that a three-dimensional (3D) convolutional neural network (CNN) can be made to predict the physical interaction forces from video images. In this paper, we proposed a bottleneck-based 3D depthwise separable CNN architecture where the video is disentangled into spatial and temporal information. By applying the basic depthwise convolution concept to each video frame, spatial information can be efficiently learned; for temporal information, the 3D pointwise convolution can be used to learn the linear combination among sequential frames. To validate and train the proposed model, we collected large quantities of datasets, which are video clips of the physical interactions between two objects under different conditions (illumination and angle variations) and the corresponding interaction forces measured by the haptic sensor (as the ground truth). Our experimental results confirmed our hypothesis; when compared with previous models, the proposed model was more accurate and efficient, and although its model size was 10 times smaller, the 3D convolutional neural network architecture exhibited better accuracy. The experiments demonstrate that the proposed model remains robust under different conditions and can successfully estimate the interaction force between objects.<\/jats:p>","DOI":"10.3390\/s19163579","type":"journal-article","created":{"date-parts":[[2019,8,19]],"date-time":"2019-08-19T06:10:14Z","timestamp":1566195014000},"page":"3579","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":25,"title":["An Efficient Three-Dimensional Convolutional Neural Network for Inferring Physical Interaction Force from Video"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8297-5466","authenticated-orcid":false,"given":"Dongyi","family":"Kim","sequence":"first","affiliation":[{"name":"Department of Software and Computer Engineering, Ajou University, 206 Worldcup-ro, Yeongtong-gu, Suwon 16499, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hyeon","family":"Cho","sequence":"additional","affiliation":[{"name":"Department of Software and Computer Engineering, Ajou University, 206 Worldcup-ro, Yeongtong-gu, Suwon 16499, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hochul","family":"Shin","sequence":"additional","affiliation":[{"name":"Department of Software and Computer Engineering, Ajou University, 206 Worldcup-ro, Yeongtong-gu, Suwon 16499, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4198-431X","authenticated-orcid":false,"given":"Soo-Chul","family":"Lim","sequence":"additional","affiliation":[{"name":"Department of Mechanical, Robotics and Energy Engineering, Dongguk University, 30, Pildong-ro 1gil, Jung-gu, Seoul 04620, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8895-0411","authenticated-orcid":false,"given":"Wonjun","family":"Hwang","sequence":"additional","affiliation":[{"name":"Department of Software and Computer Engineering, Ajou University, 206 Worldcup-ro, Yeongtong-gu, Suwon 16499, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,8,17]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Grosu, V., Grosu, S., Vanderborght, B., Lefeber, D., and Rodriguez-Guerrero, C. (2017). Multi-axis force sensor for human\u2013robot interaction sensing in a rehabilitation robotic device. Sensors, 17.","DOI":"10.3390\/s17061294"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1109\/LRA.2015.2505061","article-title":"A conformable force\/tactile skin for physical human-robot interaction","volume":"1","author":"Cirillo","year":"2016","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_3","unstructured":"Landi, C.T., Ferraguti, F., Sabattini, L., Secchi, C., and Fantuzzi, C. (June, January 29). Admittance control parameter adaptation for physical human-robot interaction. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Singapore, Singapore."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"360","DOI":"10.1002\/rcs.1625","article-title":"Role of combined tactile and kinesthetic feedback in minimally invasive surgery","volume":"11","author":"Lim","year":"2015","journal-title":"Int. J. Med Robot. Comput. Assist. Surg."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Liu, Y., Han, H., Liu, T., Yi, J., Li, Q., and Inoue, Y. (2016). A novel tactile sensor with electromagnetic induction and its application on stick-slip interaction detection. Sensors, 16.","DOI":"10.3390\/s16040430"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhang, H., Wu, R., Li, C., Zang, X., Zhang, X., Jin, H., and Zhao, J. (2017). A force-sensing system on legs for biomimetic hexapod robots interacting with unstructured terrain. Sensors, 17.","DOI":"10.3390\/s17071514"},{"key":"ref_7","unstructured":"Lazebnik, S., Schmid, C., and Ponce, J. (2006, January 7\u201322). Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, New York, NY, USA."},{"key":"ref_8","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsk","year":"2012","journal-title":"Neural Inf. Process. Syst."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2625","DOI":"10.1109\/TPAMI.2016.2599174","article-title":"Long-term recurrent convolutional networks for visual recognition and description","volume":"39","author":"Donahue","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"429","DOI":"10.1038\/415429a","article-title":"Humans integrate visual and haptic information in a statistically optimal fashion","volume":"415","author":"Ernst","year":"2002","journal-title":"Nature"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1111\/1467-9280.00307","article-title":"Viewpoint dependence in visual and haptic object recognition","volume":"12","author":"Newell","year":"2001","journal-title":"Psychol. Sci."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Luca, M.D. (2014). Physical aspects of softness perception. Multisensory Softness, Springer.","DOI":"10.1007\/978-1-4471-6533-0_5"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Hwang, W., and Lim, S.C. (2017). Inferring interaction force from visual information without using physical force sensors. Sensors, 17.","DOI":"10.3390\/s17112455"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","article-title":"3D convolutional neural networks for human action recognition","volume":"35","author":"Ji","year":"2103","journal-title":"IEEE Trans. Pattern Recognit. Mach. Intell."},{"key":"ref_15","unstructured":"Hyeon, C., Dongyi, K., Junho, P., Kyungshik, R., and Wonjun, H. (2108, January 26\u201330). 2D barcode detection using images for drone-assisted inventory management. Proceedings of the 15th International Conference on Ubiquitous Robots and Ambient Intelligence (URAI), Honolulu, HI, USA."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"282","DOI":"10.1109\/LRA.2016.2601345","article-title":"Interaction force reconstruction for humanoid robots","volume":"2","author":"Mattioli","year":"2017","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1312","DOI":"10.1109\/LRA.2017.2666420","article-title":"Gaussian process regression for sensorless grip force estimation of cable-driven elongated surgical instruments","volume":"2","author":"Li","year":"2017","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"431","DOI":"10.1109\/TOH.2016.2640289","article-title":"Towards retrieving force feedback in robotic-assisted surgery: A supervised neuro-recurrent-vision approach","volume":"10","author":"Aviles","year":"2016","journal-title":"IEEE Trans. Haptics"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhu, Y., Jiang, C., Zhao, Y., Terzopoulos, D., and Zhu, S.C. (2016, January 27\u201330). Inferring forces and learning human utilities from videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.415"},{"key":"ref_20","unstructured":"Pham, T.H., Kheddar, A., Qammaz, A., and Argyros, A.A. (2015, January 7\u201312). Towards force sensing from vision: Observing hand-object interactions to infer manipulation forces. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"8863","DOI":"10.1109\/JSEN.2018.2868332","article-title":"Interaction force estimation using camera and electrical current without force\/torque sensor","volume":"18","author":"Wonjun","year":"2018","journal-title":"IEEE Sens. J."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"358","DOI":"10.1007\/s11263-017-0992-z","article-title":"Prediction of manipulation actions","volume":"126","author":"Wang","year":"2018","journal-title":"Int. J. Comput. Vis."},{"key":"ref_23","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., and Adam, H. (2017). MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (2018, January 18\u201322). ShuffleNet: An extremely efficient convolutional neural network for mobile devices. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_26","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very deep convolutional networks for large-scale image recognition. Proceedings of the International Conference on Learning Representations, San Diego, CA, USA."},{"key":"ref_27","unstructured":"Landola, F.N., Han, S., Moskewicz, M.W., Ashraf, K., Dally, W.J., and Keutzer, K. (2016). Squeezenet: Alexnet-level accuracy with 50\u00d7 fewer parameters and <0.5 MB model size. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017). Xception: Deep learning with depthwise separable convolutions. arXiv, 1610\u20132357.","DOI":"10.1109\/CVPR.2017.195"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ye, R., Liu, F., and Zhang, L. (2018). 3D depthwise convolution: Reducing model parameters in 3D vision tasks. arXiv.","DOI":"10.1007\/978-3-030-18305-9_15"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L. (2018, January 18\u201322). MobileNetV2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_31","unstructured":"Glorot, X., and Bengio, Y. (2015, January 13\u201315). Understanding the difficulty of training deep feedforward neural networks. Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, Sardinia, Italy."},{"key":"ref_32","unstructured":"Kingma, D.P., and Ba, J. (2017). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_33","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015, January 6\u201311). Show, attend and tell: Neural image caption generation with visual attention. Proceedings of the International Conference on Machine Learning, Lille, France."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/16\/3579\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:11:47Z","timestamp":1760188307000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/16\/3579"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,8,17]]},"references-count":33,"journal-issue":{"issue":"16","published-online":{"date-parts":[[2019,8]]}},"alternative-id":["s19163579"],"URL":"https:\/\/doi.org\/10.3390\/s19163579","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,8,17]]}}}