{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T16:22:32Z","timestamp":1783095752277,"version":"3.54.6"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"11","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62372074"],"award-info":[{"award-number":["62372074"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"General projects of the Natural Science Foundation of Chongqing","award":["CSTB2023NSCQ-MSX0274"],"award-info":[{"award-number":["CSTB2023NSCQ-MSX0274"]}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["2023CDJKYJH050"],"award-info":[{"award-number":["2023CDJKYJH050"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>\n                    Early traffic accident prediction using dashcam videos plays a crucial role in enhancing the safety of intelligent vehicles. Accurately predicting accidents in advance can significantly reduce traffic accidents and improve overall road safety. However, despite extensive research efforts to capture more visual information by employing different feature extraction methods within the same frame, the consistency between features within the same frame and the discrepancy between features across different frames have not been sufficiently emphasized. To address this critical issue, we introduce contrastive learning into the field of accident prediction and propose a novel feature fusion module for the deep integration of diverse features. Our method treats features from the same frame as positive pairs, neighboring frames as sub-positive pairs due to their high correlation, and features from temporally distant frames as negative pairs. This approach effectively strengthens the representation capability of the model, thereby improving overall predictive performance. Additionally, we redefine the accident prediction task by converting it into an anomaly score regression problem using soft labels. This redefinition allows the model to better quantify the likelihood of an accident, offering a more nuanced and accurate prediction. We evaluate our method comprehensively on two publicly available Dashcam Accident Dataset (DAD) and Car Crash Dataset (CCD) datasets to assess its performance. The results demonstrate that our method outperforms state-of-the-art accident prediction approaches, highlighting its potential for practical applications. Code will be available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/yangugu\/TAP_CC\">https:\/\/github.com\/yangugu\/TAP_CC<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3767737","type":"journal-article","created":{"date-parts":[[2025,9,16]],"date-time":"2025-09-16T13:16:46Z","timestamp":1758028606000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Early Traffic Accident Anticipation via Feature Consistency Representation and Soft Label Regression"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5689-1146","authenticated-orcid":false,"given":"Yuanhong","family":"Zhong","sequence":"first","affiliation":[{"name":"School of Microelectronics and Communication Engineering, Chongqing University, Chongqing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-4664-688X","authenticated-orcid":false,"given":"Ge","family":"Yan","sequence":"additional","affiliation":[{"name":"School of Microelectronics and Communication Engineering, Chongqing University, Chongqing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-7936-3851","authenticated-orcid":false,"given":"Ruyue","family":"Zhu","sequence":"additional","affiliation":[{"name":"School of Microelectronics and Communication Engineering, Chongqing University, Chongqing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-1320-9615","authenticated-orcid":false,"given":"Ping","family":"Gan","sequence":"additional","affiliation":[{"name":"School of Microelectronics and Communication Engineering, Chongqing University, Chongqing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4702-7106","authenticated-orcid":false,"given":"Xuerui","family":"Shen","sequence":"additional","affiliation":[{"name":"School of Microelectronics and Communication Engineering, Chongqing University, Chongqing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,11,21]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413827"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00441"},{"key":"e_1_3_1_4_2","unstructured":"Shaked Brody Uri Alon and Eran Yahav. 2021. How attentive are graph attention networks? arXiv:2105.14491. Retrieved from https:\/\/arxiv.org\/abs\/2105.14491"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2956516"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_1_7_2","first-page":"136","volume-title":"Proceedings of the 13th Asian Conference on Computer Vision (ACCV \u201916)","author":"Chan Fu-Hsiang","year":"2017","unstructured":"Fu-Hsiang Chan, Yu-Ting Chen, Yu Xiang, and Min Sun. 2017. Anticipating accidents in dashcam videos. In Proceedings of the 13th Asian Conference on Computer Vision (ACCV \u201916), 136\u2013153."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.2023.3283021"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01343"},{"key":"e_1_3_1_10_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Chen Pengfei","year":"2021","unstructured":"Pengfei Chen, Guangyong Chen, Junjie Ye, Pheng-Ann Heng, 2021. Noise against noise: Stochastic label noise helps combat inherent label noise. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/IVS.2004.1336352"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3307655"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2020.3044678"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR48806.2021.9412338"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3052930"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CISS50987.2021.9400257"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-023-10609-x"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2022.3155613"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIV.2023.3275543"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2021.3054625"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2022.10.007"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3680672"},{"issue":"8","key":"e_1_3_1_24_2","doi-asserted-by":"crossref","first-page":"12518","DOI":"10.1109\/TITS.2021.3115123","article-title":"Temporal shift and spatial attention-based two-stream network for traffic risk assessment","volume":"23","author":"Liu Chunsheng","year":"2022","unstructured":"Chunsheng Liu, Zijian Li, Faliang Chang, Shuang Li, and Jincan Xie. 2022. Temporal shift and spatial attention-based two-stream network for traffic risk assessment. IEEE Transactions on Intelligent Transportation Systems 23, 8 (2022), 12518\u201312530.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2024.129285"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2023.03.075"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3141044"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2023.3271395"},{"key":"e_1_3_1_29_2","unstructured":"California Department of Motor Vehicles (DMV). 2024. Autonomous Vehicle Collision Reports. Retrieved from https:\/\/www.dmv.ca.gov\/portal\/vehicle-industry-services\/autonomous-vehicles\/autonomous-vehicle-collision-reports\/"},{"key":"e_1_3_1_30_2","volume-title":"Proceedings of NIPS Autodiff Workshop, Future of Gradient-Based Machine Learning SoftwAre and Techniques","author":"Paszke Adam","year":"2017","unstructured":"Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zach DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017. Automatic differentiation in PyTorch. In Proceedings of NIPS Autodiff Workshop, Future of Gradient-Based Machine Learning SoftwAre and Techniques."},{"issue":"10","key":"e_1_3_1_31_2","article-title":"EOGT: Video anomaly detection with enhanced object information and global temporal dependency","volume":"20","author":"Pi Ruoyan","year":"2024","unstructured":"Ruoyan Pi, Peng Wu, Xiangteng He, and Yuxin Peng. 2024. EOGT: Video anomaly detection with enhanced object information and global temporal dependency. ACM Transactions on Multimedia Computing, Communications, and Applications 20, 10, Article 320 (Sep. 2024), 21 pages.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2022.3188101"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2024.3354852"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3417989"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2008.2005605"},{"key":"e_1_3_1_37_2","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Retrieved from https:\/\/arxiv.org\/abs\/1409.1556"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2022.3152527"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.110071"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00371"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV57701.2024.00736"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIV.2023.3257169"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3150763"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8794474"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3338743"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2022.3147982"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.146"},{"issue":"1","key":"e_1_3_1_48_2","doi-asserted-by":"crossref","first-page":"291","DOI":"10.1109\/TIP.2018.2867733","article-title":"Adversarial spatio-temporal learning for video deblurring","volume":"28","author":"Zhang Kaihao","year":"2018","unstructured":"Kaihao Zhang, Wenhan Luo, Yiran Zhong, Lin Ma, Wei Liu, and Hongdong Li. 2018. Adversarial spatio-temporal learning for video deblurring. IEEE Transactions on Image Processing: A Publication of the IEEE Signal Processing Society 28, 1 (2018), 291\u2013301.","journal-title":"IEEE Transactions on Image Processing: A Publication of the IEEE Signal Processing Society"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.3039798"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3190539"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3767737","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T13:06:50Z","timestamp":1763730410000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3767737"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,21]]},"references-count":49,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3767737"],"URL":"https:\/\/doi.org\/10.1145\/3767737","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,21]]},"assertion":[{"value":"2024-12-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}