{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T16:17:25Z","timestamp":1776442645755,"version":"3.51.2"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,7,17]],"date-time":"2021-07-17T00:00:00Z","timestamp":1626480000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,7,17]],"date-time":"2021-07-17T00:00:00Z","timestamp":1626480000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No.61901165"],"award-info":[{"award-number":["No.61901165"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61501199"],"award-info":[{"award-number":["61501199"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["EURASIP J. Adv. Signal Process."],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Deep learning techniques have achieved specific results in recording device source identification. The recording device source features include spatial information and certain temporal information. However, most recording device source identification methods based on deep learning only use spatial representation learning from recording device source features, which cannot make full use of recording device source information. Therefore, in this paper, to fully explore the spatial information and temporal information of recording device source, we propose a new method for recording device source identification based on the fusion of spatial feature information and temporal feature information by using an end-to-end framework. From a feature perspective, we designed two kinds of networks to extract recording device source spatial and temporal information. Afterward, we use the attention mechanism to adaptively assign the weight of spatial information and temporal information to obtain fusion features. From a model perspective, our model uses an end-to-end framework to learn the deep representation from spatial feature and temporal feature and train using deep and shallow loss to joint optimize our network. This method is compared with our previous work and baseline system. The results show that the proposed method is better than our previous work and baseline system under general conditions.<\/jats:p>","DOI":"10.1186\/s13634-021-00763-1","type":"journal-article","created":{"date-parts":[[2021,7,17]],"date-time":"2021-07-17T09:03:03Z","timestamp":1626512583000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":27,"title":["Spatial and temporal learning representation for end-to-end recording device identification"],"prefix":"10.1186","volume":"2021","author":[{"given":"Chunyan","family":"Zeng","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dongliang","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6960-509X","authenticated-orcid":false,"given":"Zhifeng","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Minghu","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Xiong","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nan","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,7,17]]},"reference":[{"key":"763_CR1","doi-asserted-by":"publisher","first-page":"593","DOI":"10.1007\/978-981-13-3600-3_56","volume-title":"Soft computing and signal processing","author":"M. Narkhede","year":"2019","unstructured":"M. Narkhede, R. Patole, in Soft computing and signal processing, ed. by J. Wang, G. R. M. Reddy, V. K. Prasad, and V. S. Reddy. Acoustic scene identification for audio authentication (SpringerSingapore, 2019), pp. 593\u2013602."},{"key":"763_CR2","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1109\/MSP.2008.931080","volume":"26","author":"R. C. Maher","year":"2009","unstructured":"R. C. Maher, Audio forensic examination. IEEE Signal Proc. Mag.26:, 84\u201394 (2009).","journal-title":"IEEE Signal Proc. Mag."},{"key":"763_CR3","first-page":"1001","volume":"2003","author":"M. Steinebach","year":"2003","unstructured":"M. Steinebach, J. Dittmann, Watermarking-based digital audio data authentication. EURASIP J. Adv. Sig. Process.2003:, 1001\u20131015 (2003).","journal-title":"EURASIP J. Adv. Sig. Process."},{"key":"763_CR4","doi-asserted-by":"publisher","first-page":"1806","DOI":"10.1109\/ICASSP.2010.5495407","volume-title":"2010 IEEE International Conference on Acoustics, Speech and Signal Processing","author":"D. Garcia-Romero","year":"2010","unstructured":"D. Garcia-Romero, C. Y. Espy-Wilson, in 2010 IEEE International Conference on Acoustics, Speech and Signal Processing. Automatic acquisition device identification from speech recordings (IEEENew Jersey, 2010), pp. 1806\u20131809."},{"key":"763_CR5","doi-asserted-by":"publisher","first-page":"777","DOI":"10.1109\/ICECA.2019.8822177","volume-title":"2019 3rd International Conference on Electronics, Communication and Aerospace Technology (ICECA)","author":"V. A. Hadoltikar","year":"2019","unstructured":"V. A. Hadoltikar, V. R. Ratnaparkhe, R. Kumar, in 2019 3rd International Conference on Electronics, Communication and Aerospace Technology (ICECA). Optimization of mfcc parameters for mobile phone recognition from audio recordings (IEEENew Jersey, 2019), pp. 777\u2013780."},{"key":"763_CR6","doi-asserted-by":"publisher","first-page":"621","DOI":"10.1109\/ChinaSIP.2014.6889318","volume-title":"2014 IEEE China Summit International Conference on Signal and Information Processing (ChinaSIP)","author":"L. Zou","year":"2014","unstructured":"L. Zou, J. Yang, T. Huang, in 2014 IEEE China Summit International Conference on Signal and Information Processing (ChinaSIP). Automatic cell phone recognition from speech recordings (IEEENew Jersey, 2014), pp. 621\u2013625."},{"key":"763_CR7","doi-asserted-by":"publisher","first-page":"625","DOI":"10.1109\/TIFS.2011.2178403","volume":"7","author":"C. Hanilci","year":"2012","unstructured":"C. Hanilci, F. Ertas, T. Ertas,. Eskidere, Recognition of brand and models of cell-phones from recorded speech signals. IEEE Trans. Inf. Forensic. Secur.7:, 625\u2013634 (2012).","journal-title":"IEEE Trans. Inf. Forensic. Secur."},{"key":"763_CR8","doi-asserted-by":"publisher","first-page":"75","DOI":"10.1016\/j.dsp.2014.08.008","volume":"35","author":"C. Hanili","year":"2014","unstructured":"C. Hanili, T. Kinnunen, Source cell-phone recognition from recorded speech using non-speech segments. Digit. Sig. Process.35:, 75\u201385 (2014).","journal-title":"Digit. Sig. Process."},{"key":"763_CR9","doi-asserted-by":"publisher","first-page":"2044","DOI":"10.1121\/1.3385386","volume":"127","author":"D. Garcia-Romero","year":"2010","unstructured":"D. Garcia-Romero, C. Espy-Wilson, Speech forensics: automatic acquisition device identification. J. Acoust. Soc. Am.127:, 2044 (2010).","journal-title":"J. Acoust. Soc. Am."},{"key":"763_CR10","doi-asserted-by":"publisher","first-page":"586","DOI":"10.1109\/ICDSP.2014.6900732","volume-title":"2014 19th International Conference on Digital Signal Processing","author":"C. Kotropoulos","year":"2014","unstructured":"C. Kotropoulos, S. Samaras, in 2014 19th International Conference on Digital Signal Processing. Mobile phone identification using recorded speech signals (IEEENew Jersey, 2014), pp. 586\u2013591."},{"key":"763_CR11","doi-asserted-by":"publisher","first-page":"158685","DOI":"10.1109\/ACCESS.2019.2950859","volume":"7","author":"G. Baldini","year":"2019","unstructured":"G. Baldini, I. Amerini, Smartphones identification through the built-in microphones with convolutional neural network. IEEE Access. 7:, 158685\u2013158696 (2019).","journal-title":"IEEE Access"},{"key":"763_CR12","doi-asserted-by":"publisher","first-page":"965","DOI":"10.1109\/TIFS.2017.2774505","volume":"13","author":"Y. Li","year":"2018","unstructured":"Y. Li, X. Zhang, X. Li, Y. Zhang, J. Yang, Q. He, Mobile phone clustering from speech recordings using deep representation and spectral clustering. IEEE Trans. Inf. Forensic. Secur.13:, 965\u2013977 (2018).","journal-title":"IEEE Trans. Inf. Forensic. Secur."},{"key":"763_CR13","doi-asserted-by":"publisher","first-page":"220980","DOI":"10.1109\/ACCESS.2020.3043142","volume":"8","author":"M. Ashraf","year":"2020","unstructured":"M. Ashraf, G. Geng, X. Wang, F. Ahmad, F. Abid, A globally regularized joint neural architecture for music classification. IEEE Access. 8:, 220980\u2013220989 (2020).","journal-title":"IEEE Access"},{"issue":"4","key":"763_CR14","doi-asserted-by":"publisher","first-page":"413","DOI":"10.1108\/IJWIS-06-2020-0038","volume":"16","author":"C. Zeng","year":"2020","unstructured":"C. Zeng, D. Zhu, Z. Wang, Z. Wang, N. Zhao, L. He, An end-to-end deep source recording device identification system for web media forensics. Int. J. Web Inf. Syst. 16(4), 413\u2013425 (2020).","journal-title":"Int. J. Web Inf. Syst."},{"key":"763_CR15","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1145\/1288869.1288879","volume-title":"Workshop on Multimedia & Security","author":"C. Kraetzer","year":"2007","unstructured":"C. Kraetzer, A. Oermann, J. Dittmann, A. Lang, in Workshop on Multimedia & Security. Digital audio forensics: a first practical evaluation on microphone and environment classification (ACM PressNew York, 2007), pp. 63\u201374."},{"key":"763_CR16","first-page":"205","volume":"9","author":"T. Qin","year":"2018","unstructured":"T. Qin, R. Wang, D. Yan, L. Lin, Source cell-phone identification in the presence of additive noise from CQT domain. Inf. (Switzerland). 9:, 205 (2018).","journal-title":"Inf. (Switzerland)"},{"key":"763_CR17","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1142\/S0218213016500160","volume":"25","author":"m. Eskidere","year":"2016","unstructured":"m. Eskidere, Source digital voice recorder identification by wavelet analysis. Int. J. Artif. Intell. Tools. 25:, 1\u201319 (2016).","journal-title":"Int. J. Artif. Intell. Tools"},{"key":"763_CR18","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1145\/2482513.2482520","volume-title":"Acm Workshop on Information Hiding & Multimedia Security","author":"C. Hanil\u00e7i","year":"2013","unstructured":"C. Hanil\u00e7i, F. Ertas, in Acm Workshop on Information Hiding & Multimedia Security. Optimizing acoustic features for source cell-phone recognition using speech signals (ACM PressNew York, 2013), pp. 141\u2013148."},{"key":"763_CR19","doi-asserted-by":"publisher","first-page":"1218","DOI":"10.1109\/ICCSP.2014.6950045","volume-title":"2014 International Conference on Communication and Signal Processing","author":"R. Aggarwal","year":"2014","unstructured":"R. Aggarwal, S. Singh, A. K. Roul, N. Khanna, in 2014 International Conference on Communication and Signal Processing. Cellphone identification using noise estimates from recorded audio (IEEENew York, 2014), pp. 1218\u20131222."},{"key":"763_CR20","doi-asserted-by":"publisher","first-page":"2079","DOI":"10.1109\/ICASSP.2016.7472043","volume-title":"2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"L. Zou","year":"2016","unstructured":"L. Zou, Q. He, J. Yang, Y. Li, in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Source cell phone matching from speech recordings by sparse representation and kiss metric (IEEENew York, 2016), pp. 2079\u20132083."},{"key":"763_CR21","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1016\/j.diin.2019.03.003","volume":"29","author":"C. Jin","year":"2019","unstructured":"C. Jin, R. Wang, D. Yan, Source smartphone identification by exploiting encoding characteristics of recorded speech. Digit. Investig.29:, 129\u2013146 (2019).","journal-title":"Digit. Investig."},{"key":"763_CR22","doi-asserted-by":"publisher","first-page":"308","DOI":"10.1109\/LSP.2006.870086","volume":"13","author":"W. M. Campbell","year":"2006","unstructured":"W. M. Campbell, D. E. Sturim, D. A. Reynolds, Support vector machines using gmm supervectors for speaker verification. IEEE Signal Proc. Lett.13:, 308\u2013311 (2006).","journal-title":"IEEE Signal Proc. Lett."},{"key":"763_CR23","doi-asserted-by":"publisher","first-page":"2875","DOI":"10.1109\/TIFS.2019.2911175","volume":"14","author":"Y. Jiang","year":"2019","unstructured":"Y. Jiang, F. H. F. Leung, Source microphone recognition aided by a kernel-based projection method. IEEE Trans. Inf. Forensic. Secur.14:, 2875\u20132886 (2019).","journal-title":"IEEE Trans. Inf. Forensic. Secur."},{"key":"763_CR24","doi-asserted-by":"publisher","first-page":"2137","DOI":"10.1109\/ICASSP.2017.7952534","volume-title":"2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Y. Li","year":"2017","unstructured":"Y. Li, X. Zhang, X. Li, X. Feng, J. Yang, A. Chen, Q. He, in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Mobile phone clustering from acquired speech recordings using deep Gaussian supervector and spectral clustering (IEEENew York, 2017), pp. 2137\u20132141."},{"key":"763_CR25","doi-asserted-by":"publisher","first-page":"605","DOI":"10.1109\/LSP.2020.2985594","volume":"27","author":"X. Lin","year":"2020","unstructured":"X. Lin, J. Zhu, D. Chen, Subband aware CNN for cell-phone recognition. IEEE Sig. Process Lett.27:, 605\u2013609 (2020).","journal-title":"IEEE Sig. Process Lett."},{"key":"763_CR26","first-page":"161","volume-title":"2002 IEEE International Conference on Acoustics, Speech, and Signal Processing","author":"W. M. Campbell","year":"2002","unstructured":"W. M. Campbell, in 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing, 1. Generalized linear discriminant sequence kernels for speaker recognition (IEEENew York, 2002), pp. 161\u2013164."},{"key":"763_CR27","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/LSENS.2019.2923590","volume":"3","author":"G. Baldini","year":"2019","unstructured":"G. Baldini, I. Amerini, C. Gentile, Microphone identification using convolutional neural networks. IEEE Sensors Lett.3:, 1\u20134 (2019).","journal-title":"IEEE Sensors Lett."},{"key":"763_CR28","doi-asserted-by":"publisher","first-page":"1761","DOI":"10.1109\/COMST.2017.2694487","volume":"19","author":"G. Baldini","year":"2017","unstructured":"G. Baldini, G. Steri, A survey of techniques for the identification of mobile phones using the physical fingerprints of the built-in components. IEEE Commun. Surv. Tutorials. 19:, 1761\u20131789 (2017).","journal-title":"IEEE Commun. Surv. Tutorials"},{"key":"763_CR29","doi-asserted-by":"publisher","first-page":"2179","DOI":"10.1109\/TIFS.2018.2812185","volume":"13","author":"D. Luo","year":"2018","unstructured":"D. Luo, P. Korus, J. Huang, Band energy difference for source attribution in audio forensics. IEEE Trans. Inf. Forensic. Secur. 13:, 2179\u20132189 (2018).","journal-title":"IEEE Trans. Inf. Forensic. Secur."},{"key":"763_CR30","doi-asserted-by":"publisher","first-page":"125","DOI":"10.1016\/j.dsp.2016.10.017","volume":"62","author":"L. Zou","year":"2016","unstructured":"L. Zou, Q. He, J. Wu, Source cell phone verification from speech recordings using sparse representation. Digital Sig. Process. 62:, 125\u2013136 (2016).","journal-title":"Digital Sig. Process"},{"key":"763_CR31","doi-asserted-by":"publisher","first-page":"1787","DOI":"10.1109\/ICASSP.2015.7178278","volume-title":"2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"L. Zou","year":"2015","unstructured":"L. Zou, Q. He, X. Feng, in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Cell phone verification from speech recordings using sparse representation (IEEENew York, 2015), pp. 1787\u20131791."},{"key":"763_CR32","doi-asserted-by":"publisher","first-page":"4311","DOI":"10.1109\/TSP.2006.881199","volume":"54","author":"M. Aharon","year":"2006","unstructured":"M. Aharon, M. Elad, A. Bruckstein, K-SVD: an algorithm for designing overcomplete dictionaries for sparse representation. IEEE Trans. Sig. Process.54:, 4311\u20134322 (2006).","journal-title":"IEEE Trans. Sig. Process."},{"key":"763_CR33","doi-asserted-by":"publisher","first-page":"2691","DOI":"10.1109\/CVPR.2010.5539989","volume-title":"2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Q. Zhang","year":"2010","unstructured":"Q. Zhang, B. Li, in 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. Discriminative K-SVD for dictionary learning in face recognition (IEEENew York, 2010), pp. 2691\u20132698."},{"key":"763_CR34","first-page":"7276","volume-title":"2018 AAAI Conference on Artificial Intelligence","author":"X. Pan","year":"2017","unstructured":"X. Pan, J. Shi, P. Luo, X. Wang, X. Tang, in 2018 AAAI Conference on Artificial Intelligence. Spatial as deep: spatial CNN for traffic scene understanding (AAAIPalo Alto, 2017), pp. 7276\u20137283."},{"key":"763_CR35","doi-asserted-by":"publisher","first-page":"770","DOI":"10.1109\/CVPR.2016.90","volume-title":"2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"K. He","year":"2016","unstructured":"K. He, X. Zhang, S. Ren, J. Sun, in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Deep residual learning for image recognition (IEEENew Jersey, 2016), pp. 770\u2013778."},{"key":"763_CR36","doi-asserted-by":"publisher","first-page":"1675","DOI":"10.1109\/TASLP.2019.2925934","volume":"27","author":"Y. Xie","year":"2019","unstructured":"Y. Xie, R. Liang, Z. Liang, C. Huang, C. Zou, B. Schuller, Speech emotion classification using attention-based LSTM. IEEE\/ACM Trans. Audio Speech Lang. Process.27:, 1675\u20131685 (2019).","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"5","key":"763_CR37","doi-asserted-by":"publisher","first-page":"308","DOI":"10.1109\/LSP.2006.870086","volume":"13","author":"W. M. Campbell","year":"2006","unstructured":"W. M. Campbell, D. E. Sturim, D. A. Reynolds, Support vector machines using GMM supervectors for speaker verification. IEEE Sign. Process. Lett.13(5), 308\u2013311 (2006). https:\/\/doi.org\/10.1109\/LSP.2006.870086.","journal-title":"IEEE Sign. Process. Lett."}],"container-title":["EURASIP Journal on Advances in Signal Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13634-021-00763-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13634-021-00763-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13634-021-00763-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,6]],"date-time":"2023-11-06T05:10:48Z","timestamp":1699247448000},"score":1,"resource":{"primary":{"URL":"https:\/\/asp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13634-021-00763-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,17]]},"references-count":37,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["763"],"URL":"https:\/\/doi.org\/10.1186\/s13634-021-00763-1","relation":{},"ISSN":["1687-6180"],"issn-type":[{"value":"1687-6180","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,7,17]]},"assertion":[{"value":"5 February 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 July 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 July 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"41"}}