{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:53:15Z","timestamp":1760237595701,"version":"build-2065373602"},"reference-count":53,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2020,5,27]],"date-time":"2020-05-27T00:00:00Z","timestamp":1590537600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Region Lorraine and Universit\u00e9 de Lorraine","award":["1.87.02.99.216.331\/09"],"award-info":[{"award-number":["1.87.02.99.216.331\/09"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Collecting correlated scene images and camera poses is an essential step towards learning absolute camera pose regression models. While the acquisition of such data in living environments is relatively easy by following regular roads and paths, it is still a challenging task in constricted industrial environments. This is because industrial objects have varied sizes and inspections are usually carried out with non-constant motions. As a result, regression models are more sensitive to scene images with respect to viewpoints and distances. Motivated by this, we present a simple but efficient camera pose data collection method, WatchPose, to improve the generalization and robustness of camera pose regression models. Specifically, WatchPose tracks nested markers and visualizes viewpoints in an Augmented Reality- (AR) based manner to properly guide users to collect training data from broader camera-object distances and more diverse views around the objects. Experiments show that WatchPose can effectively improve the accuracy of existing camera pose regression models compared to the traditional data acquisition method. We also introduce a new dataset, Industrial10, to encourage the community to adapt camera pose regression methods for more complex environments.<\/jats:p>","DOI":"10.3390\/s20113045","type":"journal-article","created":{"date-parts":[[2020,5,28]],"date-time":"2020-05-28T12:36:58Z","timestamp":1590669418000},"page":"3045","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["WatchPose: A View-Aware Approach for Camera Pose Data Collection in Industrial Environments"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5230-7942","authenticated-orcid":false,"given":"Cong","family":"Yang","sequence":"first","affiliation":[{"name":"MAGRIT Team, INRIA\/LORIA, 54600 Nancy, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gilles","family":"Simon","sequence":"additional","affiliation":[{"name":"MAGRIT Team, INRIA\/LORIA, 54600 Nancy, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3005-4109","authenticated-orcid":false,"given":"John","family":"See","sequence":"additional","affiliation":[{"name":"Faculty of Computing and Informatics, Multimedia University, Cyberjaya 63100, Selangor, Malaysia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marie-Odile","family":"Berger","sequence":"additional","affiliation":[{"name":"MAGRIT Team, INRIA\/LORIA, 54600 Nancy, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenyong","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology, Northeast Normal University, Changchun 130000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,5,27]]},"reference":[{"key":"ref_1","first-page":"106","article-title":"A survey of industrial augmented reality","volume":"139","author":"Fernando","year":"2020","journal-title":"Comput. Ind. Eng."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Williams, B., Klein, G., and Reid, I. (2007, January 14\u201321). Real-time SLAM relocalisation. Proceedings of the 2007 IEEE 11th International Conference on Computer Vision, Rio de Janeiro, Brazil.","DOI":"10.1109\/ICCV.2007.4409115"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1561\/1100000049","article-title":"A survey of augmented reality","volume":"8","author":"Billinghurst","year":"2015","journal-title":"Found. Trends Hum.-Comput. Interact."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Sattler, T., Zhou, Q., Pollefeys, M., and Leal-Taixe, L. (2019, January 16\u201320). Understanding the Limitations of CNN-based Absolute Camera Pose Regression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00342"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Kendall, A., Grimes, M., and Cipolla, R. (2015, January 7\u201313). PoseNet: A convolutional network for real-time 6-dof camera relocalization. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.336"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Kendall, A., and Cipolla, R. (2017, January 21\u201326). Geometric loss functions for camera pose regression with deep learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.694"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Brahmbhatt, S., Gu, J., Kim, K., Hays, J., and Kautz, J. (2018, January 18\u201322). Geometry-Aware Learning of Maps for Camera Localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00277"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Melekhov, I., Ylioinas, J., Kannala, J., and Rahtu, E. (2017, January 22\u201329). Image-based localization using hourglass networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.107"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Naseer, T., and Burgard, W. (2017, January 24\u201328). Deep regression for monocular camera-based 6-dof global localization in outdoor environments. Proceedings of the International Conference on Intelligent Robots and Systems, Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8205957"},{"key":"ref_10","unstructured":"Wu, J., Ma, L., and Hu, X. (June, January 29). Delving deeper into convolutional neural networks for camera relocalization. Proceedings of the IEEE International Conference on Robotics and Automation, Singapore."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Laskar, Z., Melekhov, I., Kalia, S., and Kannala, J. (2017, January 22\u201329). Camera relocalization by computing pairwise relative poses using convolutional neural network. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.113"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Brachmann, E., Krull, A., Nowozin, S., Shotton, J., Michel, F., Gumhold, S., and Rother, C. (2017, January 21\u201326). DSAC-differentiable RANSAC for camera localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.267"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Brachmann, E., and Rother, C. (2018, January 18\u201323). Learning less is more-6d camera localization via 3d surface regression. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00489"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"He, Y., Sun, W., Huang, H., Liu, J., Fan, H., and Sun, J. (2020, January 13\u201319). PVN3D: A Deep Point-wise 3D Keypoints Voting Network for 6DoF Pose Estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01165"},{"key":"ref_15","unstructured":"Moo, Y.K., Ono, E.T.Y., Lepetit, V., Salzmann, M., and Fua, P. (2018, January 18\u201323). Learning to Find Good Correspondences. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Toft, C., Stenborg, E., Hammarstrand, L., Brynte, L., Pollefeys, M., Sattler, T., and Kahl, F. (2018, January 8\u201314). Semantic Match Consistency for Long-Term Visual Localization. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01216-8_24"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Schonberger, J.L., Pollefeys, M., Geiger, A., and Sattler, T. (2018, January 18\u201323). Semantic Visual Localization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00721"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Yogamani, S., Hughes, C., Horgan, J., Sistu, G., Varley, P., ODea, D., Uricar, M., Milz, S., Simon, M., and Amende, K. (2019, January 27\u201328). WoodScape: A multi-task, multi-camera fisheye dataset for autonomous driving. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00940"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s40537-019-0197-0","article-title":"A survey on Image Data Augmentation for Deep Learning","volume":"6","author":"Shorten","year":"2019","journal-title":"J. Big Data"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wohlhart, P., and Lepetit, V. (2015, January 7\u201312). Learning descriptors for object recognition and 3d pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298930"},{"key":"ref_21","unstructured":"Cao, Z., Sheikh, Y., and Banerjee, N.K. (2016, January 16\u201321). Real-time scalable 6DOF pose estimation for textureless objects. Proceedings of the IEEE International Conference on Robotics and Automation, Stockholm, Sweden."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Schwarz, M., Schulz, H., and Behnke, S. (2015, January 26\u201330). RGB-D object recognition and pose estimation based on pre-trained convolutional neural network features. Proceedings of the IEEE International Conference on Robotics and Automation, Seattle, WA, USA.","DOI":"10.1109\/ICRA.2015.7139363"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"625","DOI":"10.1111\/cgf.13386","article-title":"State of the Art on 3D Reconstruction with RGB-D Cameras","volume":"37","author":"Stotko","year":"2018","journal-title":"Comput. Graph. Forum"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Newcombe, R.A., Izadi, S., Hilliges, O., Molyneaux, D., Kim, D., Davison, A.J., Kohli, P., Shotton, J., Hodges, S., and Fitzgibbon, A.W. (2011, January 26\u201329). KinectFusion: Real-time dense surface mapping and tracking. Proceedings of the IEEE\/ACM International Symposium on Mixed and Augmented Reality, Basel, Switzerland.","DOI":"10.1109\/ISMAR.2011.6162880"},{"key":"ref_25","first-page":"2823","article-title":"Extending the Functionality of ARToolKit to Semi Controlled\/Uncontrolled Environment","volume":"17","author":"Rabbi","year":"2014","journal-title":"Information"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1109\/38.963459","article-title":"Recent advances in augmented reality","volume":"21","author":"Azuma","year":"2001","journal-title":"IEEE Comput. Graph. Appl."},{"key":"ref_27","unstructured":"Shavit, Y., and Ferens, R. (2019). Introduction to Camera Pose Estimation with Deep Learning. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Brachmann, E., Krull, A., Michel, F., Gumhold, S., Shotton, J., and Rother, C. (2014, January 6\u201312). Learning 6d object pose estimation using 3d object coordinates. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10605-2_35"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Tiebe, O., Yang, C., Khan, M.H., Grzegorzek, M., and Scarpin, D. (2016). Stripes-Based Object Matching. Computer and Information Science, Springer.","DOI":"10.1007\/978-3-319-40171-3_5"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Shotton, J., and Glocker, B. (2013, January 23\u201328). Scene coordinate regression forests for camera relocalization in RGB-D images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.377"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Bui, M., Baur, C., Navab, N., Ilic, S., and Albarqouni, S. (2019, January 27\u201328). Adversarial Networks for Camera Pose Regression and Refinement. Proceedings of the IEEE International Conference on Computer Vision Workshops, Seoul, Korea.","DOI":"10.1109\/ICCVW.2019.00470"},{"key":"ref_32","unstructured":"Wu, C. (July, January 29). Towards Linear-Time Incremental Structure from Motion. Proceedings of the International Conference on 3D Vision, Seattle, WA, USA."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Gu, C., and Ren, X. (2010, January 5\u201311). Discriminative Mixture-of-templates for Viewpoint Classification. Proceedings of the European Conference on Computer Vision, Crete, Greece.","DOI":"10.1007\/978-3-642-15555-0_30"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Aubry, M., Maturana, D., Efros, A.A., Russell, B.C., and Sivic, J. (2014, January 23\u201328). Seeing 3d chairs: Exemplar part-based 2d-3d alignment using a large dataset of cad models. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.487"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Masci, J., Migliore, D., Bronstein, M.M., and Schmidhuber, J. (2014). Descriptor Learning for Omnidirectional Image Matching. Registration and Recognition in Images and Videos, Springer.","DOI":"10.1007\/978-3-642-44907-9_3"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"2188","DOI":"10.1109\/TPAMI.2011.70","article-title":"Hough forests for object detection, tracking, and action recognition","volume":"33","author":"Gall","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Gupta, S., Arbel\u00e1ez, P., Girshick, R., and Malik, J. (2015, January 7\u201312). Aligning 3D models to RGB-D images of cluttered scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299105"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Kendall, A., and Cipolla, R. (2016, January 16\u201321). Modelling uncertainty in deep learning for camera relocalization. Proceedings of the IEEE International Conference on Robotics and Automation, Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487679"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Walch, F., Hazirbas, C., Leal-Taixe, L., Sattler, T., Hilsenbeck, S., and Cremers, D. (2017, January 22\u201329). Image-based localization using lstms for structured feature correlation. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.75"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Tateno, K., Kitahara, I., and Ohta, Y. (2007, January 10\u201314). A Nested Marker for Augmented Reality. Proceedings of the IEEE Virtual Reality Conference, Charlotte, NC, USA.","DOI":"10.1109\/VR.2007.352495"},{"key":"ref_42","unstructured":"Crete, F., Dolmiere, T., Ladret, P., and Nicolas, M. (February, January 29). The blur effect: Perception and estimation with a new no-reference perceptual blur metric. Proceedings of the Human Vision and Electronic Imaging, San Jose, CA, USA."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Hartley, R., and Zisserman, A. (2003). Multiple View Geometry in Computer Vision, Cambridge University Press.","DOI":"10.1017\/CBO9780511811685"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zhong, B., and Li, Y. (2019, January 5\u20137). Image Feature Point Matching Based on Improved SIFT Algorithm. Proceedings of the International Conference on Image, Vision and Computing, Xiamen, China.","DOI":"10.1109\/ICIVC47709.2019.8981329"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"1099","DOI":"10.1109\/TPAMI.2015.2477814","article-title":"Coherency Sensitive Hashing","volume":"38","author":"Korman","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_46","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster r-cnn: Towards real-time object detection with region proposal networks. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1007\/s11263-007-0090-8","article-title":"LabelMe: A database and web-based tool for image annotation","volume":"77","author":"Russell","year":"2008","journal-title":"Int. J. Comput. Vis."},{"key":"ref_48","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_49","unstructured":"Everingham, M., Gool, L.V., Williams, C.K.I., Winn, J., and Zisserman, A. (2020, May 26). The PASCAL Visual Object Classes Challenge Results. Available online: http:\/\/host.robots.ox.ac.uk\/pascal\/VOC\/voc2007\/."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., and Perona, P. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_51","unstructured":"Tian, Y., and Jiang, W. (2019). Gimbal Handheld Holder. (10,208,887), U.S. Patent."},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/30.580378","article-title":"Contrast enhancement using brightness preserving bi-histogram equalization","volume":"43","author":"Kim","year":"1997","journal-title":"IEEE Trans. Consum. Electron."},{"key":"ref_53","unstructured":"ARCore (2020, May 26). Google ARCore. Available online: https:\/\/developers.google.com\/ar\/."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/11\/3045\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:33:11Z","timestamp":1760175191000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/11\/3045"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,5,27]]},"references-count":53,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2020,6]]}},"alternative-id":["s20113045"],"URL":"https:\/\/doi.org\/10.3390\/s20113045","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2020,5,27]]}}}