{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T00:08:39Z","timestamp":1781222919478,"version":"3.54.1"},"reference-count":33,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2022,2,11]],"date-time":"2022-02-11T00:00:00Z","timestamp":1644537600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In the context of smart cities, monitoring pedestrian and vehicle movements is essential to recognize abnormal events and prevent accidents. The proposed method in this work focuses on analyzing video streams captured from a vertically installed camera, and performing contextual road user detection. The final detection is based on the fusion of the outputs of three different convolutional neural networks. We are simultaneously interested in detecting road users, their motion, and their location respecting the static environment. We use YOLOv4 for object detection, FC-HarDNet for background semantic segmentation, and FlowNet 2.0 for motion detection. FC-HarDNet and YOLOv4 were retrained with our orthophotographs dataset. The last step involves a data fusion module. The presented results show that the method allows one to detect road users, identify the surfaces on which they move, quantify their apparent velocity, and estimate their actual velocity.<\/jats:p>","DOI":"10.3390\/s22041381","type":"journal-article","created":{"date-parts":[[2022,2,11]],"date-time":"2022-02-11T05:14:43Z","timestamp":1644556483000},"page":"1381","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Contextual Detection of Pedestrians and Vehicles in Orthophotography by Fusion of Deep Learning Algorithms"],"prefix":"10.3390","volume":"22","author":[{"given":"Masoomeh Shireen","family":"Ansarnia","sequence":"first","affiliation":[{"name":"Institut Jean Lamour (UMR7198), Universit\u00e9 de Lorraine, 54052 Nancy, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2088-3598","authenticated-orcid":false,"given":"Etienne","family":"Tisserand","sequence":"additional","affiliation":[{"name":"Institut Jean Lamour (UMR7198), Universit\u00e9 de Lorraine, 54052 Nancy, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Patrick","family":"Schweitzer","sequence":"additional","affiliation":[{"name":"Institut Jean Lamour (UMR7198), Universit\u00e9 de Lorraine, 54052 Nancy, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9806-4173","authenticated-orcid":false,"given":"Mohamed Amine","family":"Zidane","sequence":"additional","affiliation":[{"name":"Institut Jean Lamour (UMR7198), Universit\u00e9 de Lorraine, 54052 Nancy, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0059-3024","authenticated-orcid":false,"given":"Yves","family":"Berviller","sequence":"additional","affiliation":[{"name":"Institut Jean Lamour (UMR7198), Universit\u00e9 de Lorraine, 54052 Nancy, France"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,2,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Barth\u00e9lemy, J., Verstaevel, N., Forehead, H., and Perez, P. (2019). Edge-Computing Video Analytics for Real-Time Traffic Monitoring in a Smart City. Sensors, 19.","DOI":"10.3390\/s19092048"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Chen, L.-C., Sheu, R.-K., Peng, W.-Y., Wu, J.-H., and Tseng, C.-H. (2020). Video-Based Parking Occupancy Detection for Smart Control System. Appl. Sci., 10.","DOI":"10.3390\/app10031079"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Rezaei, M., and Azarmi, M. (2020). DeepSOCIAL: Social Distancing Monitoring and Infection Risk Assessment in COVID-19 Pandemic. Appl. Sci., 10.","DOI":"10.1101\/2020.08.27.20183277"},{"key":"ref_4","first-page":"775","article-title":"Machine Learning Applied to Road Safety Modeling: A Systematic Literature Review","volume":"7","author":"Silva","year":"2020","journal-title":"J. Traffic Transp. Eng."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"2113","DOI":"10.1109\/TIE.2013.2266084","article-title":"Sensorless Illumination Control of a Networked LED-Lighting System Using Feedforward Neural Network","volume":"61","author":"Tran","year":"2014","journal-title":"IEEE Trans. Ind. Electron."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Dollar, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Hua, J., Hao, T., Zeng, L., and Yu, G. (2021). YOLOMask, an Instance Segmentation Algorithm Based on Complementary Fusion Network. Mathematics, 9.","DOI":"10.3390\/math9151766"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Ali, A., and Taylor, G.W. (2018, January 9\u201311). Real-Time End-to-End Action Detection with Two-Stream Networks. Proceedings of the 2018 15th Conference on Computer and Robot Vision (CRV), Toronto, ON, Canada.","DOI":"10.1109\/CRV.2018.00015"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Tran, M.-T., Dinh-Duy, T., Truong, T.-D., Ton-That, V., Do, T.-N., Luong, Q.-A., Nguyen, T.-A., Nguyen, V.-T., and Do, M.N. (2018, January 18\u201322). Traffic Flow Analysis with Multiple Adaptive Vehicle Detectors and Velocity Estimation with Landmark-Based Scanlines. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00021"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhang, S., Wang, T., Wang, C., Wang, Y., Shan, G., and Snoussi, H. (2019, January 21\u201322). Video Object Detection Base on RGB and Optical Flow Analysis. Proceedings of the 2019 2nd China Symposium on Cognitive Computing and Hybrid Intelligence (CCHI), Xi\u2019an, China.","DOI":"10.1109\/CCHI.2019.8901921"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"012003","DOI":"10.1088\/1742-6596\/1004\/1\/012003","article-title":"A Real-Time Method to Estimate Speed of Object Based on Object Detection and Optical Flow Calculation","volume":"1004","author":"Liu","year":"2018","journal-title":"J. Phys. Conf. Ser."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"987","DOI":"10.1049\/itr2.12079","article-title":"Vision-based Vehicle Speed Estimation: A Survey","volume":"15","author":"Daza","year":"2021","journal-title":"IET Intell. Transp. Syst."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"261","DOI":"10.1007\/s11263-019-01247-4","article-title":"Deep Learning for Generic Object Detection: A Survey","volume":"128","author":"Liu","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_15","unstructured":"Bochkovskiy, A., Wang, C.-Y., and Liao, H.-Y.M. (2020). YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zaidi, S.S.A., Ansari, M.S., Aslam, A., Kanwal, N., Asghar, M., and Lee, B. (2021). A Survey of Modern Deep Learning Based Object Detection Models. arXiv.","DOI":"10.1016\/j.dsp.2022.103514"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Wang, C.-Y., Bochkovskiy, A., and Liao, H.-Y.M. (2021, January 20\u201325). Scaled-YOLOv4: Scaling Cross Stage Partial Network. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01283"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_19","unstructured":"Paszke, A., Chaurasia, A., Kim, S., and Culurciello, E. (2016). ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1340","DOI":"10.1109\/TRO.2020.2974099","article-title":"MiniNet: An Efficient Semantic Segmentation ConvNet for Real-Time Robotic Applications","volume":"36","author":"Alonso","year":"2020","journal-title":"IEEE Trans. Robot."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1016\/j.isprsjprs.2019.02.006","article-title":"Semantic Segmentation of Slums in Satellite Images Using Transfer Learning on Fully Convolutional Neural Networks","volume":"150","author":"Wurm","year":"2019","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1007\/s10462-020-09854-1","article-title":"Deep Semantic Segmentation of Natural and Medical Images: A Review","volume":"54","author":"Taghanaki","year":"2021","journal-title":"Artif. Intell. Rev."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Chao, P., Kao, C.-Y., Ruan, Y.-S., Huang, C.-H., and Lin, Y.-L. (2019, January 27\u201328). HarDNet: A Low Memory Traffic Network. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00365"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Shaikh, S.H., Saeed, K., and Chaki, N. (2014). Moving Object Detection Using Background Subtraction, Springer International Publishing. Springer Briefs in Computer Science.","DOI":"10.1007\/978-3-319-07386-6"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"887","DOI":"10.1016\/j.proeng.2017.06.153","article-title":"Comparison of Background Subtraction Methods on Near Infra-Red Spectrum Video Sequences","volume":"192","author":"Hudec","year":"2017","journal-title":"Procedia Eng."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Dosovitskiy, A., Fischer, P., Ilg, E., Hausser, P., Hazirbas, C., Golkov, V., van der Smagt, P., Cremers, D., and Brox, T. (2015, January 7\u201313). FlowNet: Learning Optical Flow with Convolutional Networks. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.316"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Ilg, E., Mayer, N., Saikia, T., Keuper, M., Dosovitskiy, A., and Brox, T. (2017, January 21\u201326). FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.179"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Kong, L., Shen, C., and Yang, J. (2021). FastFlowNet: A Lightweight Network for Fast Optical Flow Estimation. arXiv.","DOI":"10.1109\/ICRA48506.2021.9560800"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Sun, D., Yang, X., Liu, M.-Y., and Kautz, J. (2018, January 18\u201323). PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00931"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Revaud, J., Weinzaepfel, P., Harchaoui, Z., and Schmid, C. (2015). EpicFlow: Edge-Preserving Interpolation of Correspondences for Optical Flow, HAL.","DOI":"10.1109\/CVPR.2015.7298720"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.-M. (2020, January 23\u201328). RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. Proceedings of the Computer Vision\u2013ECCV 2020, Glasgow, UK.","DOI":"10.1007\/978-3-030-58574-7"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1198\/10618600152418584","article-title":"The Art of Data Augmentation","volume":"10","author":"Meng","year":"2001","journal-title":"J. Comput. Graph. Stat."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"62","DOI":"10.1109\/TSMC.1979.4310076","article-title":"A Threshold Selection Method from Gray-Level Histograms","volume":"9","author":"Otsu","year":"1979","journal-title":"IEEE Trans. Syst. Man Cybern."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/4\/1381\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:18:10Z","timestamp":1760134690000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/4\/1381"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,11]]},"references-count":33,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,2]]}},"alternative-id":["s22041381"],"URL":"https:\/\/doi.org\/10.3390\/s22041381","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,11]]}}}