{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,19]],"date-time":"2026-03-19T06:45:53Z","timestamp":1773902753105,"version":"3.50.1"},"reference-count":47,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2021,7,23]],"date-time":"2021-07-23T00:00:00Z","timestamp":1626998400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"the National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["52002285"],"award-info":[{"award-number":["52002285"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"the Shanghai Pujiang Program","award":["2020PJD075"],"award-info":[{"award-number":["2020PJD075"]}]},{"name":"the Shenzhen Future Intelligent Network Transportation System Industry Innovation Center","award":["17092530321"],"award-info":[{"award-number":["17092530321"]}]},{"name":"the Key Special Projects of the National Key R\\&amp;D Program of China","award":["2018AAA0102800"],"award-info":[{"award-number":["2018AAA0102800"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>The automated driving of agricultural machinery is of great significance for the agricultural production efficiency, yet is still challenging due to the significantly varied environmental conditions through day and night. To address operation safety for pedestrians in farmland, this paper proposes a 3D person sensing approach based on monocular RGB and Far-Infrared (FIR) images. Since public available datasets for agricultural 3D pedestrian detection are scarce, a new dataset is proposed, named as \u201cFieldSafePedestrian\u201d, which includes field images in both day and night. The implemented data augmentations of night images and semi-automatic labeling approach are also elaborated to facilitate the 3D annotation of pedestrians. To fuse heterogeneous images of sensors with non-parallel optical axis, the Dual-Input Depth-Guided Dynamic-Depthwise-Dilated Fusion network (D5F) is proposed, which assists the pixel alignment between FIR and RGB images with estimated depth information and deploys a dynamic filtering to guide the heterogeneous information fusion. Experiments on field images in both daytime and nighttime demonstrate that compared with the state-of-the-arts, the dynamic aligned image fusion achieves an accuracy gain of 3.9% and 4.5% in terms of center distance and BEV-IOU, respectively, without affecting the run-time efficiency.<\/jats:p>","DOI":"10.3390\/rs13152896","type":"journal-article","created":{"date-parts":[[2021,7,23]],"date-time":"2021-07-23T10:31:44Z","timestamp":1627036304000},"page":"2896","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["3D Pedestrian Detection in Farmland by Monocular RGB Image and Far-Infrared Sensing"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5085-7219","authenticated-orcid":false,"given":"Wei","family":"Tian","sequence":"first","affiliation":[{"name":"Institute of Intelligent Vehicles, School of Automotive Studies, Tongji University, Shanghai 201804, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7328-5508","authenticated-orcid":false,"given":"Zhenwen","family":"Deng","sequence":"additional","affiliation":[{"name":"Institute of Intelligent Vehicles, School of Automotive Studies, Tongji University, Shanghai 201804, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dong","family":"Yin","sequence":"additional","affiliation":[{"name":"Institute of Intelligent Vehicles, School of Automotive Studies, Tongji University, Shanghai 201804, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9733-6437","authenticated-orcid":false,"given":"Zehan","family":"Zheng","sequence":"additional","affiliation":[{"name":"Institute of Intelligent Vehicles, School of Automotive Studies, Tongji University, Shanghai 201804, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6723-5582","authenticated-orcid":false,"given":"Yuyao","family":"Huang","sequence":"additional","affiliation":[{"name":"Institute of Intelligent Vehicles, School of Automotive Studies, Tongji University, Shanghai 201804, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xin","family":"Bi","sequence":"additional","affiliation":[{"name":"Institute of Intelligent Vehicles, School of Automotive Studies, Tongji University, Shanghai 201804, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,7,23]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kragh, M., J\u00f8rgensen, R., and Pedersen, H. (2015, January 6\u20139). Object Detection and Terrain Classification in Agricultural Fields Using 3D Lidar Data. Proceedings of the International Conference on Computer Vision Systems, Copenhagen, Denmark.","DOI":"10.1007\/978-3-319-20904-3_18"},{"key":"ref_2","first-page":"29","article-title":"Real-time Pedestrian Detection in Orchard Based on Improved SSD","volume":"50","author":"Liu","year":"2019","journal-title":"Trans. Chin. Soc. Agric. Mach."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"126","DOI":"10.1016\/j.measurement.2016.06.067","article-title":"A new DSWTS algorithm for real-time pedestrian detection in autonomous agricultural tractors as a computer vision system","volume":"93","author":"Hamed","year":"2016","journal-title":"Measurement"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"619","DOI":"10.1080\/15389588.2019.1624731","article-title":"Improving pedestrian safety using combined HOG and Haar partial detection in mobile systems","volume":"20","author":"Alkar","year":"2019","journal-title":"Traffic Inj. Prev."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Kragh, M., Christiansen, P., and Laursen, M. (2017). FieldSAFE: Dataset for obstacle detection in agriculture. Sensors, 17.","DOI":"10.3390\/s17112579"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhu, J., Park, T., Isola, P., and Efros, A. (2017, January 22\u201329). Unpaired Image-To-Image Translation Using Cycle-Consistent Adversarial Networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.244"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.sna.2021.112566","article-title":"Sensing system of environmental perception technologies for driverless vehicle: A review of state of the art and challenges","volume":"319","author":"Chen","year":"2021","journal-title":"Sensors Actuators A Phys."},{"key":"ref_8","unstructured":"Eigen, D., Puhrsch, C., and Fergus, R. (2015, January 6\u20139). Depth map prediction from a single image using a multi-scale deep network. Proceedings of the International Conference on Neural Information Processing Systems, Copenhagen, Denmark."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2024","DOI":"10.1109\/TPAMI.2015.2505283","article-title":"Learning depth from single monocular images using deep convolutional neural fields","volume":"38","author":"Liu","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","unstructured":"Li, B., Shen, C., and Dai, Y. (2015, January 7\u201312). Depth and surface normal estimation from monocular images using regression on deep features and hierarchical CRFs. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_11","unstructured":"Iro, L., Christian, R., and Vasileios, B. (2016, January 25\u201328). Deeper depth prediction with fully convolutional residual networks. Proceedings of the 4th International Conference on 3D Vision, Stanford, CA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Fu, H., Gong, M., Wang, C., Batmanghelich, K., and Tao, D. (2018, January 18\u201323). Deep ordinal regression network for monocular depth estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00214"},{"key":"ref_13","unstructured":"Lee, J., Han, M., Ko, D., and Suh, I. (2020). From big to small: Multi-scale local planar guidance for monocular depth estimation. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zhou, T., Brown, M., Snavely, N., and Lowe, D. (2017, January 21\u201326). Unsupervised learning of depth and egomotion from video. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.700"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Yin, Z., and Shi, J. (2018, January 18\u201323). Geonet: Unsupervised learning of dense depth, optical flow and camera pose. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00212"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Godard, C., Aodha, O., Firman, M., and Brostow, G. (2019). Digging into self-supervised monocular depth estimation. arXiv.","DOI":"10.1109\/ICCV.2019.00393"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Chen, X., Kundu, K., Zhang, Z., Ma, H., Fidler, S., and Urtasun, R. (2016, January 27\u201330). Monocular 3D Object Detection for Autonomous Driving. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.236"},{"key":"ref_18","unstructured":"He, T., and Soatto, S. (February, January 27). Mono3D++: Monocular 3D Vehicle Detection with Two-Scale 3D Hypotheses and Task Priors. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Chabot, F., Chaouch, M., Rabarisoa, J., Teuliere, C., and Chateau, T. (2017, January 21\u201326). Deep MANTA: A Coarse-To-Fine Many-Task Network for Joint 2D and 3D Vehicle Analysis From Monocular Image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.198"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Manhardt, F., Kehl, W., and Gaidon, A. (2019, January 15\u201320). ROI-10D: Monocular Lifting of 2D Detection to 6D Pose and Metric Shape. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00217"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Liu, Z., Wu, Z., and Toth, R. (2020, January 14\u201319). SMOKE: Single-Stage Monocular 3D Object Detection via Keypoint Estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00506"},{"key":"ref_22","unstructured":"Zhou, X., Wang, D., and Kr\u00e4henb\u00fchl, P. (2019). Objects as Points. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Li, P., Zhao, H., Liu, P., and Cao, F. (2020). RTM3D: Real-time monocular 3D detection from object keypoints for autonomous driving. arXiv.","DOI":"10.1007\/978-3-030-58580-8_38"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Mousavian, A., Anguelov, D., Flynn, J., and Kosecka, J. (2017, January 21\u201326). 3D Bounding Box Estimation Using Deep Learning and Geometry. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.597"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Brazil, G., and Liu, X. (2019, January 16\u201320). M3D-RPN: Monocular 3D Region Proposal Network for Object Detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Long Beach, CA, USA.","DOI":"10.1109\/ICCV.2019.00938"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Xu, B., and Chen, Z. (2018, January 18\u201323). Multi-Level Fusion Based 3D Object Detection From Monocular Images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00249"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Wang, Y., Chao, W., Garg, D., Hariharan, B., Campbell, M., and Weinberger, K. (2019, January 15\u201320). Pseudo-LiDAR From Visual Depth Estimation: Bridging the Gap in 3D Object Detection for Autonomous Driving. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00864"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Qi, C., Liu, W., Wu, C., Su, H., and Guibas, L. (2018, January 18\u201323). Frustum PointNets for 3D Object Detection From RGB-D Data. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00102"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ma, X., Liu, S., Xia, Z., Zhang, H., Zeng, X., and Ouyang, W. (2020). Rethinking pseudo-LiDAR representation. arXiv.","DOI":"10.1007\/978-3-030-58601-0_19"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Ding, M., Huo, Y., Yi, H., Wang, Z., Shi, J., Lu, Z., and Luo, P. (2020, January 13\u201319). Learning Depth-Guided Convolutions for Monocular 3D Object Detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01169"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for autonomous driving? The KITTI vision benchmark suite. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Hinterstoisser, S., Lepetit, V., Ilic, S., Holzer, S., and Bradski, G. (2012, January 5\u20139). Model Based Training, Detection and Pose Estimation of Texture-Less 3D Objects in Heavily Cluttered Scenes. Proceedings of the Asian Conference on Computer Vision, Daejeon, Korea.","DOI":"10.1007\/978-3-642-33885-4_60"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Schmidt, T., Narayanan, V., and Fox, D. (2018). PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes. arXiv.","DOI":"10.15607\/RSS.2018.XIV.019"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Song, X., Wang, P., Zhou, D., Zhu, R., and Guan, C. (2019, January 15\u201320). ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00560"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. (2019). nuScenes: A multimodal dataset for autonomous driving. arXiv.","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1602","DOI":"10.1177\/0278364910384638","article-title":"The Marulan Data Sets: Multi-sensor Perception in a Natural Environment with Challenging Conditions","volume":"29","author":"Peynot","year":"2010","journal-title":"Int. J. Robot. Res."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Pezzementi, Z., Tabor, T., Hu, P., Chang, J., and Ramanan, D. (2017). Comparing Apples and Oranges: Off-Road Pedestrian Detection on the NREC Agricultural Person-Detection Dataset. arXiv.","DOI":"10.1002\/rob.21760"},{"key":"ref_38","first-page":"91","article-title":"Distinctive Image Features from Scale-Invariant Keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Robot. Res."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Yu, F., Chen, H., Wang, X., Xian, W., and Chen, Y. (2020, January 13\u201319). BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00271"},{"key":"ref_40","unstructured":"Tzuta, L. (2021, July 22). LabelImg. Git Code. Available online: https:\/\/github.com\/tzutalin\/labelImg."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1483","DOI":"10.1109\/TPAMI.2019.2956516","article-title":"Cascade R-CNN: High quality object detection and instance segmentation","volume":"43","author":"Cai","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_42","unstructured":"Sun, R., Zhao, Y., Jiang, B., Cheng, T., and Xiao, B. (2019). High-Resolution Representations for Labeling Pixels and Regions. arXiv."},{"key":"ref_43","unstructured":"MacQueen, J.B. (,  1967). Some methods for classification and analysis of multivariate observations. Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability, Berkeley, CA, USA. Available online: https:\/\/projecteuclid.org\/ebooks\/berkeley-symposium-on-mathematical-statistics-and-probability\/Proceedings%20of%20the%20Fifth%20Berkeley%20Symposium%20on%20Mathematical%20Statistics%20and%20Probability,%20Volume%201:%20Statistics\/chapter\/Some%20methods%20for%20classification%20and%20analysis%20of%20multivariate%20observations\/bsmsp\/1200512992."},{"key":"ref_44","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1145\/358669.358692","article-title":"Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography","volume":"24","author":"Fischler","year":"1981","journal-title":"Commun. ACM"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Arlia, D., and Coppola, M. (2001, January 28\u201331). Experiments in Parallel Clustering with DBSCAN. Proceedings of the International Euro-Par Conference, Manchester, UK.","DOI":"10.1007\/3-540-44681-8_46"},{"key":"ref_47","unstructured":"Wu, B., Wan, A., Yue, X., Jin, P., Zhao, S., Golmant, N., Gholaminejad, A., Gonzalez, J., and Keutzer, K. (2020). Shift: A Zero FLOP, Zero Parameter Alternative to Spatial Convolutions. arXiv."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/15\/2896\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:34:04Z","timestamp":1760164444000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/15\/2896"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,23]]},"references-count":47,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2021,8]]}},"alternative-id":["rs13152896"],"URL":"https:\/\/doi.org\/10.3390\/rs13152896","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,7,23]]}}}