{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,28]],"date-time":"2025-11-28T12:27:47Z","timestamp":1764332867630,"version":"build-2065373602"},"reference-count":60,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2020,3,11]],"date-time":"2020-03-11T00:00:00Z","timestamp":1583884800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100012774","name":"Innovationsfonden","doi-asserted-by":"publisher","award":["5189-00116B","7045-00035B"],"award-info":[{"award-number":["5189-00116B","7045-00035B"]}],"id":[{"id":"10.13039\/100012774","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The challenge of getting machines to understand and interact with natural objects is encountered in important areas such as medicine, agriculture, and, in our case, slaughterhouse automation. Recent breakthroughs have enabled the application of Deep Neural Networks (DNN) directly to point clouds, an efficient and natural representation of 3D objects. The potential of these methods has mostly been demonstrated for classification and segmentation tasks involving rigid man-made objects. We present a method, based on the successful PointNet architecture, for learning to regress correct tool placement from human demonstrations, using virtual reality. Our method is applied to a challenging slaughterhouse cutting task, which requires an understanding of the local geometry including the shape, size, and orientation. We propose an intermediate five-Degree of Freedom (DoF) cutting plane representation, a point and a normal vector, which eases the demonstration and learning process. A live experiment is conducted in order to unveil issues and begin to understand the required accuracy. Eleven cuts are rated by an expert, with     8 \/ 11     being rated as acceptable. The error on the test set is subsequently reduced through the addition of more training data and improvements to the DNN. The result is a reduction in the average translation from 1.5 cm to 0.8 cm and the orientation error from 4.59\u00b0 to 4.48\u00b0. The method\u2019s generalization capacity is assessed on a similar task from the slaughterhouse and on the very different public LINEMOD dataset for object pose estimation across view points. In both cases, the method shows promising results. Code, datasets, and other materials are available in Supplementary Materials.<\/jats:p>","DOI":"10.3390\/s20061563","type":"journal-article","created":{"date-parts":[[2020,3,12]],"date-time":"2020-03-12T04:13:57Z","timestamp":1583986437000},"page":"1563","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Cutting Pose Prediction from Point Clouds"],"prefix":"10.3390","volume":"20","author":[{"given":"Mark P.","family":"Philipsen","sequence":"first","affiliation":[{"name":"Media Technology, Aalborg University, 9000 Aalborg, Denmark"},{"name":"Danish Technological Institute, Gregersensvej 9, 2630 Taastrup, Denmark"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thomas B.","family":"Moeslund","sequence":"additional","affiliation":[{"name":"Media Technology, Aalborg University, 9000 Aalborg, Denmark"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,3,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"413","DOI":"10.1038\/544413a","article-title":"Artificial Intelligence: Chess match of the century","volume":"544","author":"Hassabis","year":"2017","journal-title":"Nature"},{"key":"ref_2","unstructured":"Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. (2013). Intriguing properties of neural networks. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Alcorn, M.A., Li, Q., Gong, Z., Wang, C., Mai, L., Ku, W., and Nguyen, A. (2019, January 16\u201320). Strike (With) a Pose: Neural Networks Are Easily Fooled by Strange Poses of Familiar Objects. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00498"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"392","DOI":"10.1007\/s10278-017-9976-3","article-title":"Medical image data and datasets in the era of machine learning\u2014Whitepaper from the 2016 C-MIMI meeting dataset session","volume":"30","author":"Kohli","year":"2017","journal-title":"J. Digit. Imaging"},{"key":"ref_5","unstructured":"Animalia (2020, March 10). Meat2.0. Available online: https:\/\/www.animalia.no\/no\/animalia\/om-animalia\/arsrapporter-og-strategi\/aret-som-gikk\u20132017\/forsker-pa-framtidens-slakterier\/."},{"key":"ref_6","unstructured":"Institute, D.T. (2020, March 10). Augmented Cellular Meat Production. Available online: https:\/\/www.teknologisk.dk\/ydelser\/intelligente-robotter-skal-fastholde-koedproduktion-i-danmark\/39225."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"41","DOI":"10.4141\/cjas66-007","article-title":"Ear characteristics and performance in swine","volume":"46","author":"Boylan","year":"1966","journal-title":"Can. J. Anim. Sci."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Hinterstoisser, S., Lepetit, V., Ilic, S., Holzer, S., Bradski, G., Konolige, K., and Navab, N. (2012, January 5\u20139). Model based training, detection and pose estimation of texture-less 3D objects in heavily cluttered scenes. Proceedings of the 11th Asian Conference on Computer Vision, Daejeon, Korea.","DOI":"10.1007\/978-3-642-33885-4_60"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Hinterstoisser, S., Holzer, S., Cagniart, C., Ilic, S., Konolige, K., Navab, N., and Lepetit, V. (2011, January 6\u201313). Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenes. Proceedings of the 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126326"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Kehl, W., Manhardt, F., Tombari, F., Ilic, S., and Navab, N. (2017, January 22\u201329). SSD-6D: Making RGB-based 3D detection and 6D pose estimation great again. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.169"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Schmidt, T., Narayanan, V., and Fox, D. (2017). Posecnn: A convolutional neural network for 6D object pose estimation in cluttered scenes. arXiv.","DOI":"10.15607\/RSS.2018.XIV.019"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Wong, J.M., Kee, V., Le, T., Wagner, S., Mariottini, G.L., Schneider, A., Hamilton, L., Chipalkatty, R., Hebert, M., and Johnson, D.M.S. (2017, January 24\u201328). SegICP: Integrated Deep Semantic Segmentation and Pose Estimation. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8206470"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, C., Xu, D., Zhu, Y., Mart\u00edn, R., Lu, C., Fei, L., and Savarese, S. (2019, January 16\u201320). DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00346"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Krull, A., Brachmann, E., Michel, F., Yang, M., Gumhold, S., and Rother, C. (2015, January 13\u201316). Learning analysis-by-synthesis for 6D pose estimation in RGB-D images. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.115"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zeng, A., Yu, K.T., Song, S., Suo, D., Walker, E., Rodriguez, A., and Xiao, J. (June, January 29). Multi-view Self-supervised Deep Learning for 6D Pose Estimation in the Amazon Picking Challenge. Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore.","DOI":"10.1109\/ICRA.2017.7989165"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Tekin, B., Sinha, S.N., and Fua, P. (2018, January 12\u201318). Real-time seamless single shot 6d object pose prediction. Proceedings of the The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00038"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Hu, Y., Hugonot, J., Fua, P., and Salzmann, M. (2019, January 16\u201320). Segmentation-driven 6D object pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00350"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1177\/0278364907087172","article-title":"Robotic grasping of novel objects using vision","volume":"27","author":"Saxena","year":"2008","journal-title":"Int. J. Rob. Res."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Fischinger, D., and Vincze, M. (2012, January 7\u201312). Empty the basket-a shape based learning approach for grasping piles of unknown objects. Proceedings of the 2012 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Vilamoura, Portugal.","DOI":"10.1109\/IROS.2012.6386137"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"705","DOI":"10.1177\/0278364914549607","article-title":"Deep learning for detecting robotic grasps","volume":"34","author":"Lenz","year":"2015","journal-title":"Int. J. Rob. Res."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Pinto, L., and Gupta, A. (2016, January 16\u201321). Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours. Proceedings of the 2016 IEEE international conference on robotics and automation (ICRA), Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487517"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Kumra, S., and Kanan, C. (2017, January 24\u201328). Robotic grasp detection using deep convolutional neural networks. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8202237"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Du, G., Wang, K., and Lian, S. (2019). Vision-based Robotic Grasping from Object Localization, Pose Estimation, Grasp Detection to Motion Planning: A Review. arXiv.","DOI":"10.1007\/s10462-020-09888-5"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"688","DOI":"10.1177\/0278364918779698","article-title":"Robotic manipulation and sensing of deformable objects in domestic and industrial applications: A survey","volume":"37","author":"Sanchez","year":"2018","journal-title":"Int. J. Rob. Res."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Lin, G., Tang, Y., Zou, X., Xiong, J., and Li, J. (2019). Guava detection and pose estimation using a low-cost RGB-D sensor in the field. Sensors, 19.","DOI":"10.3390\/s19020428"},{"key":"ref_26","unstructured":"ten Pas, A., and Platt, R. (2015). Localizing antipodal grasps in point clouds. arXiv."},{"key":"ref_27","first-page":"1455","article-title":"Grasp Pose Detection in Point Clouds","volume":"13\u201314","author":"Gualtieri","year":"2017","journal-title":"SAGE J."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Dyrstad, J.S., and Mathiassen, J.R. (2017, January 5\u20138). Grasping virtual fish: A step towards robotic deep learning from demonstration in virtual reality. Proceedings of the 2017 IEEE International Conference on Robotics and Biomimetics (ROBIO), Macau, China.","DOI":"10.1109\/ROBIO.2017.8324578"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Dyrstad, J.S., \u00d8ye, E.R., Stahl, A., and Mathiassen, J.R. (2018, January 1\u20135). Teaching a Robot to Grasp Real Fish by Imitation Learning from a Human Supervisor in Virtual Reality. Proceedings of the 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain.","DOI":"10.1109\/IROS.2018.8593954"},{"key":"ref_30","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Liang, H., Ma, X., Li, S., G\u00f6rner, M., Tang, S., Fang, B., Sun, F., and Zhang, J. (2019, January 20\u201324). PointNetGPD: Detecting Grasp Configurations from Point Sets. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8794435"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Ge, L., Cai, Y., Weng, J., and Yuan, J. (2018, January 18\u201322). Hand PointNet: 3D hand pose estimation using point sets. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00878"},{"key":"ref_33","unstructured":"Wiedemeyer, T. (2020, January 10). IAI Kinect2. Available online: https:\/\/github.com\/code-iai\/iai_kinect2."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1007\/s10514-013-9366-8","article-title":"Learning of grasp selection based on shape-templates","volume":"36","author":"Herzog","year":"2014","journal-title":"Auton. Robots"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Angelova, A. (2015, January 26\u201330). Real-time grasp detection using convolutional neural networks. Proceedings of the 2015 IEEE International Conference on Robotics and Automation (ICRA), Seattle, WA, USA.","DOI":"10.1109\/ICRA.2015.7139361"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Mahler, J., Liang, J., Niyaz, S., Laskey, M., Doan, R., Liu, X., Ojea, J.A., and Goldberg, K. (2017). Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics. arXiv.","DOI":"10.15607\/RSS.2017.XIII.058"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Mahendran, S., Ali, H., and Vidal, R. (2017, January 22\u201329). 3D Pose Regression Using Convolutional Neural Networks. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCVW.2017.254"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Manhardt, F., Kehl, W., Navab, N., and Tombari, F. (2018, January 8\u201314). Deep model-based 6d pose refinement in rgb. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Gemany.","DOI":"10.1007\/978-3-030-01264-9_49"},{"key":"ref_39","unstructured":"Do, T., Cai, M., Pham, T., and Reid, I.D. (2018). Deep-6DPose: Recovering 6D Object Pose from a Single RGB Image. arXiv."},{"key":"ref_40","unstructured":"Siemens (2020, March 10). ROS#. Available online: https:\/\/github.com\/siemens\/ros-sharp."},{"key":"ref_41","unstructured":"Technologies, U. (2020, March 10). Unity. Available online: https:\/\/unity.com."},{"key":"ref_42","unstructured":"Philipsen, M.P., Wu, H., and Moeslund, T.B. (2018, January 5). Virtual Reality for Demonstrating Tool Pose. Proceedings of the 2018 Abstract from Automating Robot Experiments, Madrid, Spain."},{"key":"ref_43","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (July, January 26). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_44","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An Incremental Improvement. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Lin, T., Maire, M., Belongie, S.J., Bourdev, L.D., Girshick, R.B., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common Objects in Context. Proceedings of the 13th European Conference, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Atzmon, M., Maron, H., and Lipman, Y. (2018). Point Convolutional Neural Networks by Extension Operators. arXiv.","DOI":"10.1145\/3197517.3201301"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"021006","DOI":"10.1115\/1.4041889","article-title":"A survey on the computation of quaternions from rotation matrices","volume":"11","author":"Sarabandi","year":"2019","journal-title":"J. Mech. Rob."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Sarabandi, S., and Thomas, F. (2018, January 1\u20135). Accurate computation of quaternions from rotation matrices. Proceedings of the International Symposium on Advances in Robot Kinematics, Bologna, Italy.","DOI":"10.1007\/978-3-319-93188-3_5"},{"key":"ref_49","unstructured":"Brachmann, E., Michel, F., Krull, A., Yang, M., Gumhold, S., and Rother, C. (July, January 26). Uncertainty-driven 6d pose estimation of objects and scenes from a single rgb image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Kleppe, A., Bj\u00f8rkedal, A., Larsen, K., and Egeland, O. (2017). Automated assembly using 3D and 2D cameras. Robotics, 6.","DOI":"10.3390\/robotics6030014"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Xu, H., Chen, G., Wang, Z., Sun, L., and Su, F. (2019). RGB-D-Based Pose Estimation of Workpieces with Semantic Segmentation and Point Cloud Registration. Sensors, 19.","DOI":"10.3390\/s19081873"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Le, T.T., and Lin, C.Y. (2019). Bin-Picking for Planar Objects Based on a Deep Learning Network: A Case Study of USB Packs. Sensors, 19.","DOI":"10.3390\/s19163602"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Shotton, J., Glocker, B., Zach, C., Izadi, S., Criminisi, A., and Fitzgibbon, A. (2013, January 25\u201327). Scene coordinate regression forests for camera relocalization in RGB-D images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.377"},{"key":"ref_54","unstructured":"Hietanen, A., Latokartano, J., Foi, A., Pieters, R., Kyrki, V., Lanz, M., and K\u00e4m\u00e4r\u00e4inen, J. (2019). Benchmarking 6D Object Pose Estimation for Robotics. arXiv."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Kendall, A., Grimes, M., and Cipolla, R. (2015, January 13\u201316). Posenet: A convolutional network for real-time 6-dof camera relocalization. Proceedings of the IEEE international conference on computer vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.336"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Rad, M., and Lepetit, V. (2017, January 22\u201329). BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using Depth. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.413"},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"1357","DOI":"10.1109\/LRA.2019.2895878","article-title":"On-policy dataset synthesis for learning robot grasping policies using fully convolutional deep networks","volume":"4","author":"Satish","year":"2019","journal-title":"IEEE Rob. Autom. Lett."},{"key":"ref_58","unstructured":"Ho, J., and Ermon, S. (2016, January 5\u201310). Generative Adversarial Imitation Learning. Proceedings of the Neural Information Processing Systems 2016, Barcelona, Spain."},{"key":"ref_59","unstructured":"Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J. (2015, January 8\u201310). 3d shapenets: A deep representation for volumetric shapes. Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), Boston, MA, USA."},{"key":"ref_60","unstructured":"Chang, A.X., Funkhouser, T.A., Guibas, L.J., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., and Su, H. (2015). ShapeNet: An Information-Rich 3D Model Repository. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/6\/1563\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:06:06Z","timestamp":1760173566000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/6\/1563"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,11]]},"references-count":60,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2020,3]]}},"alternative-id":["s20061563"],"URL":"https:\/\/doi.org\/10.3390\/s20061563","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2020,3,11]]}}}