{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T14:40:53Z","timestamp":1777560053171,"version":"3.51.4"},"reference-count":31,"publisher":"SAGE Publications","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["AIC"],"published-print":{"date-parts":[[2024,9,18]]},"abstract":"<jats:p>Pedestrian intent prediction is an essential task for ensuring the safety of pedestrians and vehicles on the road. This task involves predicting whether a pedestrian intends to cross a road or not based on their behavior and surrounding environment. Previous studies have explored feature-based machine learning and vision-based deep learning models for this task but these methods have limitations in capturing the global spatio-temporal context and fusing different features of data effectively. To address these issues, we propose a novel hybrid framework HSTGCN for pedestrian intent prediction that combines spatio-temporal graph convolutional neural networks (STGCN) and long short-term memory (LSTM) networks. The proposed framework utilizes the strengths of both models by fusing multiple features, including skeleton pose, trajectory, height, orientation, and ego-vehicle speed, to predict their intentions accurately. The framework\u2019s performance have been evaluated on the JAAD benchmark dataset and the results show that it outperforms the state-of-the-art methods. The proposed framework has potential applications in developing intelligent transportation systems, autonomous vehicles, and pedestrian safety technologies. The utilization of multiple features can significantly improve the performance of the pedestrian intent prediction task.<\/jats:p>","DOI":"10.3233\/aic-230053","type":"journal-article","created":{"date-parts":[[2024,9,20]],"date-time":"2024-09-20T10:31:49Z","timestamp":1726828309000},"page":"549-562","source":"Crossref","is-referenced-by-count":2,"title":["Spatio-temporal deep learning framework for pedestrian intention prediction in urban traffic scenes"],"prefix":"10.1177","volume":"37","author":[{"family":"Monika","sequence":"first","affiliation":[{"name":"School of Computer and Systems Sciences, Jawaharlal Nehru University, New Delhi, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pardeep","family":"Singh","sequence":"additional","affiliation":[{"name":"School of Computer and Systems Sciences, Jawaharlal Nehru University, New Delhi, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Satish","family":"Chand","sequence":"additional","affiliation":[{"name":"School of Computer and Systems Sciences, Jawaharlal Nehru University, New Delhi, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"key":"10.3233\/AIC-230053_ref1","doi-asserted-by":"publisher","DOI":"10.1109\/IV51971.2022.9827084"},{"key":"10.3233\/AIC-230053_ref3","doi-asserted-by":"publisher","DOI":"10.1109\/ITSC.2019.8917118"},{"key":"10.3233\/AIC-230053_ref4","doi-asserted-by":"publisher","DOI":"10.1109\/ICCE-Taiwan55306.2022.9868998"},{"key":"10.3233\/AIC-230053_ref5","doi-asserted-by":"crossref","unstructured":"T.\u00a0Chen, R.\u00a0Tian and Z.\u00a0Ding, Visual reasoning using graph convolutional networks for predicting pedestrian crossing intention, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2021, pp.\u00a03103\u20133109.","DOI":"10.1109\/ICCVW54120.2021.00345"},{"key":"10.3233\/AIC-230053_ref6","doi-asserted-by":"crossref","unstructured":"L.\u00a0Fan, S.\u00a0Buch, G.\u00a0Wang, R.\u00a0Cao, Y.\u00a0Zhu, J.C.\u00a0Niebles and L.\u00a0Fei-Fei, Rubiksnet: Learnable 3d-shift for efficient video action recognition, in: Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part XIX, Springer, 2020, pp.\u00a0505\u2013521.","DOI":"10.1007\/978-3-030-58529-7_30"},{"key":"10.3233\/AIC-230053_ref8","doi-asserted-by":"publisher","DOI":"10.1109\/IVS.2018.8500413"},{"key":"10.3233\/AIC-230053_ref9","doi-asserted-by":"publisher","first-page":"851","DOI":"10.1016\/j.procs.2014.08.252","article-title":"Distributed system for crossroads traffic surveillance with prediction of incidents","volume":"35","author":"Favorskaya","year":"2014","journal-title":"Procedia computer science"},{"key":"10.3233\/AIC-230053_ref10","doi-asserted-by":"crossref","unstructured":"J.\u00a0Gesnouin, S.\u00a0Pechberti, B.\u00a0Stanciulcscu and F.\u00a0Moutarde, TrouSPI-Net: Spatio-temporal attention on parallel atrous convolutions and U-GRUs for skeletal pedestrian crossing prediction, in: 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021), IEEE, 2021, pp.\u00a01\u20137.","DOI":"10.1109\/FG52635.2021.9666989"},{"key":"10.3233\/AIC-230053_ref11","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2015.7351675"},{"key":"10.3233\/AIC-230053_ref12","doi-asserted-by":"crossref","unstructured":"C.C.\u00a0Jorge and R.J.\u00a0Rossetti, On social interactions and the emergence of autonomous vehicles, in: VEHITS, 2018, pp.\u00a0423\u2013430.","DOI":"10.5220\/0006763004230430"},{"key":"10.3233\/AIC-230053_ref13","doi-asserted-by":"crossref","unstructured":"I.\u00a0Kotseruba, A.\u00a0Rasouli and J.K.\u00a0Tsotsos, Benchmark for evaluating pedestrian action prediction, in: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 2021, pp.\u00a01258\u20131268.","DOI":"10.1109\/WACV48630.2021.00130"},{"issue":"2","key":"10.3233\/AIC-230053_ref14","doi-asserted-by":"publisher","first-page":"3485","DOI":"10.1109\/LRA.2020.2976305","article-title":"Spatiotemporal relationship reasoning for pedestrian intent prediction","volume":"5","author":"Liu","year":"2020","journal-title":"IEEE Robotics and Automation Letters"},{"key":"10.3233\/AIC-230053_ref15","doi-asserted-by":"publisher","DOI":"10.1109\/IV47402.2020.9304652"},{"key":"10.3233\/AIC-230053_ref16","doi-asserted-by":"publisher","DOI":"10.3390\/wevj13080158"},{"issue":"14","key":"10.3233\/AIC-230053_ref17","doi-asserted-by":"publisher","first-page":"15660","DOI":"10.1109\/JSEN.2021.3062762","article-title":"Pedestrian behavior analytics on dashcam videos in chaotic environments","volume":"21","author":"Mukherjee","year":"2021","journal-title":"IEEE Sensors Journal"},{"issue":"11","key":"10.3233\/AIC-230053_ref18","doi-asserted-by":"publisher","first-page":"6821","DOI":"10.1109\/TITS.2020.2995166","article-title":"Context model for pedestrian intention prediction using factored latent-dynamic conditional random fields","volume":"22","author":"Neogi","year":"2020","journal-title":"IEEE transactions on intelligent transportation systems"},{"key":"10.3233\/AIC-230053_ref19","doi-asserted-by":"crossref","unstructured":"L.\u00a0Neumann and A.\u00a0Vedaldi, Pedestrian and ego-vehicle trajectory prediction from monocular camera, in: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp.\u00a010204\u201310212.","DOI":"10.1109\/CVPR46437.2021.01007"},{"key":"10.3233\/AIC-230053_ref20","doi-asserted-by":"publisher","DOI":"10.1109\/IEEECONF51394.2020.9443552"},{"key":"10.3233\/AIC-230053_ref21","doi-asserted-by":"crossref","unstructured":"A.\u00a0Rasouli, I.\u00a0Kotseruba, T.\u00a0Kunic and J.K.\u00a0Tsotsos, Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2019, pp.\u00a06262\u20136271.","DOI":"10.1109\/ICCV.2019.00636"},{"key":"10.3233\/AIC-230053_ref22","doi-asserted-by":"crossref","unstructured":"A.\u00a0Rasouli, I.\u00a0Kotseruba and J.K.\u00a0Tsotsos, Are they going to cross? A benchmark dataset and baseline for pedestrian crosswalk behavior, in: Proceedings of the IEEE International Conference on Computer Vision Workshops, 2017, pp.\u00a0206\u2013213.","DOI":"10.1109\/ICCVW.2017.33"},{"issue":"3","key":"10.3233\/AIC-230053_ref24","doi-asserted-by":"publisher","first-page":"900","DOI":"10.1109\/TITS.2019.2901817","article-title":"Autonomous vehicles that interact with pedestrians: A survey of theory and practice","volume":"21","author":"Rasouli","year":"2019","journal-title":"IEEE transactions on intelligent transportation systems"},{"key":"10.3233\/AIC-230053_ref27","doi-asserted-by":"crossref","unstructured":"A.\u00a0Singh and U.\u00a0Suddamalla, Multi-input fusion for practical pedestrian intention prediction, in: Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2021, pp.\u00a02304\u20132311.","DOI":"10.1109\/ICCVW54120.2021.00260"},{"key":"10.3233\/AIC-230053_ref28","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S19-2128"},{"key":"10.3233\/AIC-230053_ref29","doi-asserted-by":"publisher","DOI":"10.1109\/ITSC.2019.8917510"},{"key":"10.3233\/AIC-230053_ref30","doi-asserted-by":"crossref","unstructured":"S.\u00a0Yan, Y.\u00a0Xiong and D.\u00a0Lin, Spatial temporal graph convolutional networks for skeleton-based action recognition, in: Thirty-Second AAAI Conference on Artificial Intelligence, 2018.","DOI":"10.1609\/aaai.v32i1.12328"},{"issue":"6","key":"10.3233\/AIC-230053_ref31","doi-asserted-by":"publisher","first-page":"5338","DOI":"10.1109\/TITS.2021.3053031","article-title":"Crossing or not? Context-based recognition of pedestrian crossing intention in the urban environment","volume":"23","author":"Yang","year":"2021","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"issue":"2","key":"10.3233\/AIC-230053_ref32","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1109\/TIV.2022.3162719","article-title":"Predicting pedestrian crossing intention with feature fusion and spatio-temporal attention","volume":"7","author":"Yang","year":"2022","journal-title":"IEEE Transactions on Intelligent Vehicles"},{"issue":"2","key":"10.3233\/AIC-230053_ref33","doi-asserted-by":"publisher","first-page":"1463","DOI":"10.1109\/LRA.2021.3056339","article-title":"Bitrap: Bi-directional pedestrian trajectory prediction with multi-modal goal estimation","volume":"6","author":"Yao","year":"2021","journal-title":"IEEE Robotics and Automation Letters"},{"key":"10.3233\/AIC-230053_ref34","doi-asserted-by":"crossref","unstructured":"S.\u00a0Yi, H.\u00a0Li and X.\u00a0Wang, Pedestrian behavior understanding and prediction with deep neural networks, in: Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11\u201314, 2016, Proceedings, Part I, Vol.\u00a014, Springer, 2016, pp.\u00a0263\u2013279.","DOI":"10.1007\/978-3-319-46448-0_16"},{"issue":"3","key":"10.3233\/AIC-230053_ref35","doi-asserted-by":"publisher","first-page":"2331","DOI":"10.1109\/TITS.2021.3074829","article-title":"Pedestrian crossing intention prediction at red-light using pose estimation","volume":"23","author":"Zhang","year":"2021","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"issue":"11","key":"10.3233\/AIC-230053_ref36","doi-asserted-by":"publisher","first-page":"20773","DOI":"10.1109\/TITS.2022.3177367","article-title":"ST CrossingPose: A spatial-temporal graph convolutional network for skeleton-based pedestrian crossing intention prediction","volume":"23","author":"Zhang","year":"2022","journal-title":"IEEE Transactions on Intelligent Transportation Systems"}],"container-title":["AI Communications"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/AIC-230053","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T18:28:17Z","timestamp":1777400897000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.medra.org\/servlet\/aliasResolver?alias=iospress&doi=10.3233\/AIC-230053"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9,18]]},"references-count":31,"journal-issue":{"issue":"4"},"URL":"https:\/\/doi.org\/10.3233\/aic-230053","relation":{},"ISSN":["1875-8452","0921-7126"],"issn-type":[{"value":"1875-8452","type":"electronic"},{"value":"0921-7126","type":"print"}],"subject":[],"published":{"date-parts":[[2024,9,18]]}}}