{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,17]],"date-time":"2026-05-17T07:19:09Z","timestamp":1779002349821,"version":"3.51.4"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,5,22]],"date-time":"2025-05-22T00:00:00Z","timestamp":1747872000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,5,22]],"date-time":"2025-05-22T00:00:00Z","timestamp":1747872000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Universit\u00e0 degli Studi di Roma La Sapienza"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Pattern Anal Applic"],"published-print":{"date-parts":[[2025,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>A deep understanding of pedestrian intention and crossing behaviors is crucial in applications like pedestrian attribute recognition and autonomous driving. While vehicles need to predict the movements of pedestrians accurately for safety, the recognition and re-identification systems rely on behavioral cues that help them enhance identity tracking and attribute analysis. Traditional trajectory-based methods for pedestrian intention estimation evaluate the future positions of pedestrians based on their past movements but may fail to capture their true intentions. A more effective approach will anticipate actions by analyzing underlying intent, improving the precision of pedestrian recognition and the motion prediction. Current research on estimating pedestrian intentions primarily depends on supervised learning methods. In contrast, this work introduces an unsupervised learning approach to learn intention representations. This method is based on the idea that similar intentions lead to comparable behaviors among pedestrians, and, therefore, they can be clustered. To achieve this, this paper introduces UnPIE, an unsupervised method for predicting pedestrian intentions. It utilizes Spatio-Temporal Graph Convolutional Networks to encode intentions from videos and map them into a D-dimensional latent space. The training phase incorporates Instance Recognition to increase separation between embeddings from different classes and Local Aggregation to form soft clusters of related embeddings. A supervised non-parametric classifier is used to evaluate the performance of the method. The results demonstrate that UnPIE has comparable performance with respect to supervised approaches and even surpasses them, achieving a higher Precision by about 7% on the Pedestrian Intention Estimation dataset.<\/jats:p>","DOI":"10.1007\/s10044-025-01483-0","type":"journal-article","created":{"date-parts":[[2025,5,22]],"date-time":"2025-05-22T14:31:07Z","timestamp":1747924267000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Unsupervised pedestrian intention estimation through deep neural embeddings and spatio-temporal graph convolutional networks"],"prefix":"10.1007","volume":"28","author":[{"given":"Simone","family":"Scaccia","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Francesco","family":"Pro","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Irene","family":"Amerini","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,5,22]]},"reference":[{"key":"1483_CR1","doi-asserted-by":"publisher","first-page":"120","DOI":"10.1016\/j.neucom.2022.07.085","volume":"508","author":"N Sharma","year":"2022","unstructured":"Sharma N, Dhiman C, Indu S (2022) Pedestrian intention prediction for autonomous vehicles: a comprehensive survey. Neurocomputing 508:120\u2013152. https:\/\/doi.org\/10.1016\/j.neucom.2022.07.085","journal-title":"Neurocomputing"},{"key":"1483_CR2","doi-asserted-by":"publisher","unstructured":"Rasouli A, Kotseruba I, Kunic T, Tsotsos J (2019) Pie: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction. In: 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), pp. 6261\u20136270. https:\/\/doi.org\/10.1109\/ICCV.2019.00636","DOI":"10.1109\/ICCV.2019.00636"},{"key":"1483_CR3","doi-asserted-by":"crossref","unstructured":"Ma Y, ZHU X, Cheng X, Yang R, Liu J, Manocha D (2020) AutoTrajectory: Label-free trajectory extraction and prediction from videos using dynamic points. https:\/\/arxiv.org\/abs\/2007.05719","DOI":"10.1007\/978-3-030-58601-0_38"},{"issue":"3","key":"1483_CR4","doi-asserted-by":"publisher","first-page":"3534","DOI":"10.1609\/aaai.v37i3.25463","volume":"37","author":"Z Zhang","year":"2023","unstructured":"Zhang Z, Tian R, Ding Z (2023) Trep: Transformer-based evidential prediction for pedestrian intention with uncertainty. Proceed AAAI Conf Artif Intell 37(3):3534\u20133542. https:\/\/doi.org\/10.1609\/aaai.v37i3.25463","journal-title":"Proceed AAAI Conf Artif Intell"},{"issue":"1","key":"1483_CR5","doi-asserted-by":"publisher","first-page":"61","DOI":"10.1109\/TIV.2017.2788193","volume":"3","author":"A Rasouli","year":"2018","unstructured":"Rasouli A, Kotseruba I, Tsotsos JK (2018) Understanding pedestrian behavior in complex traffic scenes. IEEE Trans Intell Veh 3(1):61\u201370. https:\/\/doi.org\/10.1109\/TIV.2017.2788193","journal-title":"IEEE Trans Intell Veh"},{"key":"1483_CR6","doi-asserted-by":"publisher","unstructured":"Naik AY, Bighashdel A, Jancura P, Dubbelman G (2022) Scene spatio-temporal graph convolutional network for pedestrian intention estimation. In: 2022 IEEE Intelligent Vehicles Symposium (IV), pp. 874\u2013881. https:\/\/doi.org\/10.1109\/IV51971.2022.9827231","DOI":"10.1109\/IV51971.2022.9827231"},{"key":"1483_CR7","doi-asserted-by":"crossref","unstructured":"Wu Z, Xiong Y, Yu SX, Lin D (2018) Unsupervised feature learning via non-parametric instance-level discrimination. CoRR abs\/1805.01978","DOI":"10.1109\/CVPR.2018.00393"},{"key":"1483_CR8","doi-asserted-by":"crossref","unstructured":"Zhuang C, Zhai AL, Yamins D (2019) Local aggregation for unsupervised learning of visual embeddings. CoRR abs\/1903.12355","DOI":"10.1109\/ICCV.2019.00610"},{"issue":"1","key":"1483_CR9","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1109\/TIT.1967.1053964","volume":"13","author":"T Cover","year":"1967","unstructured":"Cover T, Hart P (1967) Nearest neighbor pattern classification. IEEE Trans Inf Theory 13(1):21\u201327. https:\/\/doi.org\/10.1109\/TIT.1967.1053964","journal-title":"IEEE Trans Inf Theory"},{"issue":"1","key":"1483_CR10","doi-asserted-by":"publisher","first-page":"50","DOI":"10.1007\/s11263-014-0735-3","volume":"111","author":"B Zhou","year":"2014","unstructured":"Zhou B, Tang X, Wang X (2014) Learning collective crowd behaviors with dynamic pedestrian-agents. Int J Comput Vision 111(1):50\u201368. https:\/\/doi.org\/10.1007\/s11263-014-0735-3","journal-title":"Int J Comput Vision"},{"key":"1483_CR11","doi-asserted-by":"publisher","first-page":"697","DOI":"10.1007\/978-3-319-46448-0_42","volume-title":"Computer Vision - ECCV 2016","author":"L Ballan","year":"2016","unstructured":"Ballan L, Castaldo F, Alahi A, Palmieri F, Savarese S (2016) Knowledge transfer for scene-specific motion prediction. In: Leibe B, Matas J, Sebe N, Welling M (eds) Computer Vision - ECCV 2016. Springer, Cham, pp 697\u2013713"},{"issue":"5","key":"1483_CR12","doi-asserted-by":"publisher","first-page":"1803","DOI":"10.1109\/TITS.2018.2836305","volume":"20","author":"R Quintero M\u00ednguez","year":"2019","unstructured":"Quintero M\u00ednguez R, Parra Alonso I, Fernandez-Llorca D, Sotelo MA (2019) Pedestrian path, pose, and intention prediction through gaussian process dynamical models and pedestrian activity recognition. IEEE Trans Intell Transp Syst 20(5):1803\u20131814. https:\/\/doi.org\/10.1109\/TITS.2018.2836305","journal-title":"IEEE Trans Intell Transp Syst"},{"key":"1483_CR13","doi-asserted-by":"publisher","unstructured":"Karasev V, Ayvaci A, Heisele B, Soatto S (2016) Intent-aware long-term prediction of pedestrian motion. In: 2016 IEEE International Conference on Robotics and Automation (ICRA), pp. 2543\u20132549. https:\/\/doi.org\/10.1109\/ICRA.2016.7487409","DOI":"10.1109\/ICRA.2016.7487409"},{"key":"1483_CR14","unstructured":"Zhao H, Gao J, Lan T, Sun C, Sapp B, Varadarajan B, Shen Y, Shen Y, Chai Y, Schmid C, Li C, Anguelov D (2020) TNT: target-driven trajectory prediction. CoRR abs\/2008.08294"},{"key":"1483_CR15","unstructured":"Kotseruba I, Rasouli A, Tsotsos JK (2016) Joint attention in autonomous driving (JAAD). CoRR abs\/1609.04741"},{"key":"1483_CR16","doi-asserted-by":"publisher","unstructured":"Ham J-S, Kim DH, Jung N, Moon J (2023) Cipf: Crossing intention prediction network based on feature fusion modules for improving pedestrian safety. In: 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 3666\u20133675. https:\/\/doi.org\/10.1109\/CVPRW59228.2023.00374","DOI":"10.1109\/CVPRW59228.2023.00374"},{"issue":"2","key":"1483_CR17","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1109\/TIV.2022.3162719","volume":"7","author":"D Yang","year":"2022","unstructured":"Yang D, Zhang H, Yurtsever E, Redmill KA, \u00d6zg\u00fcner \u00dc (2022) Predicting pedestrian crossing intention with feature fusion and spatio-temporal attention. IEEE Trans Intell Veh 7(2):221\u2013230. https:\/\/doi.org\/10.1109\/TIV.2022.3162719","journal-title":"IEEE Trans Intell Veh"},{"key":"1483_CR18","doi-asserted-by":"publisher","unstructured":"Guo J, Ding Y, Tian A (2024) Multimodal feature fusion for pedestrian crossing intention prediction based on hybrid attention mechanism. In: 2024 6th International Conference on Internet of Things, Automation and Artificial Intelligence (IoTAAI), pp. 70\u201374. https:\/\/doi.org\/10.1109\/IoTAAI62601.2024.10692962","DOI":"10.1109\/IoTAAI62601.2024.10692962"},{"key":"1483_CR19","doi-asserted-by":"publisher","unstructured":"Xu L, You S, He G, Li Y (2024) Pedestrian-vehicle information modulation for pedestrian crossing intention prediction. IEEE Transactions on Intelligent Vehicles, 1\u201313. https:\/\/doi.org\/10.1109\/TIV.2024.3437779","DOI":"10.1109\/TIV.2024.3437779"},{"issue":"10","key":"1483_CR20","doi-asserted-by":"publisher","first-page":"13277","DOI":"10.1109\/TITS.2024.3398252","volume":"25","author":"Y Ling","year":"2024","unstructured":"Ling Y, Ma Z, Zhang Q, Xie B, Weng X (2024) Pedast-gcn: Fast pedestrian crossing intention prediction using spatial?temporal attention graph convolution networks. IEEE Trans Intell Transp Syst 25(10):13277\u201313290. https:\/\/doi.org\/10.1109\/TITS.2024.3398252","journal-title":"IEEE Trans Intell Transp Syst"},{"key":"1483_CR21","doi-asserted-by":"crossref","unstructured":"Xie C, Lin C, Zheng X, Gong B, Wu D, L\u00f3pez AM (2024) GTransPDM: a graph-embedded transformer with positional decoupling for pedestrian crossing intention prediction. https:\/\/arxiv.org\/abs\/2409.20223","DOI":"10.1109\/LSP.2025.3567249"},{"key":"1483_CR22","doi-asserted-by":"publisher","unstructured":"Li M, Liu M (2024) Deep-learning-based algorithm for classifying pedestrian behavior at crosswalks. In: Liu, B., Leng, L. (eds.) Third International Conference on Image Processing, Object Detection, and Tracking (IPODT 2024), vol. 13396, p. 133960. SPIE,???. https:\/\/doi.org\/10.1117\/12.3050750. International Society for Optics and Photonics","DOI":"10.1117\/12.3050750"},{"key":"1483_CR23","doi-asserted-by":"publisher","unstructured":"Hua C, Luo K, Wu Y, Shi R (2024) Yolo-abd: A multi-scale detection model for pedestrian anomaly behavior detection. Symmetry 16(8). https:\/\/doi.org\/10.3390\/sym16081003","DOI":"10.3390\/sym16081003"},{"key":"1483_CR24","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1007\/978-3-031-77392-1_11","volume-title":"Adv Visual Comput","author":"Y Bao","year":"2025","unstructured":"Bao Y, Saito Y, Nishio N (2025) Piepredict++: an improved pedestrian intention estimation model incorporating comprehensive environment information. In: Bebis G, Patel V, Gu J, Panetta J, Gingold Y, Johnsen K, Arefin MS, Dutta S, Biswas A (eds) Adv Visual Comput. Springer, Cham, pp 141\u2013155"},{"key":"1483_CR25","unstructured":"Rasouli A, Rohani M, Luo J (2020) Pedestrian behavior prediction via multitask learning and categorical interaction modeling. CoRR abs\/2012.03298"},{"key":"1483_CR26","unstructured":"Tang Y, Ma W (2025) INTENT: Trajectory prediction framework with intention-guided contrastive clustering. https:\/\/arxiv.org\/abs\/2503.04952"},{"key":"1483_CR27","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. https:\/\/arxiv.org\/abs\/1409.1556"},{"key":"1483_CR28","doi-asserted-by":"crossref","unstructured":"Zhuang C, She T, Andonian A, Mark MS, Yamins D (2020) Unsupervised Learning from Video with Deep Neural Embeddings. https:\/\/arxiv.org\/abs\/1905.11954","DOI":"10.1109\/CVPR42600.2020.00958"},{"key":"1483_CR29","unstructured":"Xiao T, Li S, Wang B, Lin L, Wang X (2016) End-to-end deep learning for person search. CoRR abs\/1604.01850"},{"key":"1483_CR30","unstructured":"Johnson J, Douze M, J\u00e9gou H (2017) Billion-scale similarity search with GPUs. https:\/\/arxiv.org\/abs\/1702.08734"},{"key":"1483_CR31","doi-asserted-by":"publisher","unstructured":"FRS, Kp, (1901) Liii on lines and planes of closest fit to systems of points in space. London Edinburgh and Dublin Philos Magaz J Sci 2(11):559\u2013572. https:\/\/doi.org\/10.1080\/14786440109462720","DOI":"10.1080\/14786440109462720"},{"issue":"86","key":"1483_CR32","first-page":"2579","volume":"9","author":"L Maaten","year":"2008","unstructured":"Maaten L, Hinton G (2008) Visualizing data using t-sne. J Mach Learn Res 9(86):2579\u20132605","journal-title":"J Mach Learn Res"},{"issue":"1","key":"1483_CR33","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1214\/aoms\/1177729694","volume":"22","author":"S Kullback","year":"1951","unstructured":"Kullback S, Leibler RA (1951) On information and sufficiency. Ann Math Stat 22(1):79\u201386. https:\/\/doi.org\/10.1214\/aoms\/1177729694","journal-title":"Ann Math Stat"},{"key":"1483_CR34","doi-asserted-by":"publisher","unstructured":"Rasouli A, Kotseruba I (2023) Pedformer: Pedestrian behavior prediction via cross-modal attention modulation and gated multitask learning. In: 2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 9844\u20139851. https:\/\/doi.org\/10.1109\/ICRA48891.2023.10161318","DOI":"10.1109\/ICRA48891.2023.10161318"},{"issue":"12","key":"1483_CR35","doi-asserted-by":"publisher","first-page":"14213","DOI":"10.1109\/TITS.2023.3309309","volume":"24","author":"Y Zhou","year":"2023","unstructured":"Zhou Y, Tan G, Zhong R, Li Y, Gou C (2023) Pit: Progressive interaction transformer for pedestrian crossing intention prediction. IEEE Trans Intell Transp Syst 24(12):14213\u201314225. https:\/\/doi.org\/10.1109\/TITS.2023.3309309","journal-title":"IEEE Trans Intell Transp Syst"},{"issue":"4","key":"1483_CR36","doi-asserted-by":"publisher","first-page":"9071","DOI":"10.1109\/TTE.2024.3360966","volume":"10","author":"B Yang","year":"2024","unstructured":"Yang B, Zhu J, Hu C, Yu Z, Hu H, Ni R (2024) Faster pedestrian crossing intention prediction based on efficient fusion of diverse intention influencing factors. IEEE Trans Transp Electrif 10(4):9071\u20139087. https:\/\/doi.org\/10.1109\/TTE.2024.3360966","journal-title":"IEEE Trans Transp Electrif"},{"key":"1483_CR37","doi-asserted-by":"crossref","unstructured":"Azarmi M, Rezaei M, Wang H, Glaser S (2024) PIP-Net: Pedestrian Intention Prediction in the Wild. https:\/\/arxiv.org\/abs\/2402.12810","DOI":"10.1109\/TITS.2025.3570794"}],"container-title":["Pattern Analysis and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10044-025-01483-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10044-025-01483-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10044-025-01483-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,2]],"date-time":"2025-07-02T16:39:06Z","timestamp":1751474346000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10044-025-01483-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,22]]},"references-count":37,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,6]]}},"alternative-id":["1483"],"URL":"https:\/\/doi.org\/10.1007\/s10044-025-01483-0","relation":{},"ISSN":["1433-7541","1433-755X"],"issn-type":[{"value":"1433-7541","type":"print"},{"value":"1433-755X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,22]]},"assertion":[{"value":"15 February 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 April 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 May 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval and consent to participate"}},{"value":"Yes.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}],"article-number":"108"}}