{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:11:23Z","timestamp":1760235083669,"version":"build-2065373602"},"reference-count":51,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2021,7,19]],"date-time":"2021-07-19T00:00:00Z","timestamp":1626652800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China (NSFC) ;The Funds for Creative Research Groups of Higher Education of Xinjiang Uygur Autonomous Region under Grant","award":["U1903213, 61761041;XJEDU2017T002"],"award-info":[{"award-number":["U1903213, 61761041;XJEDU2017T002"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>The task of pitch estimation is an essential step in many audio signal processing applications. In this paper, we propose a data-driven pitch estimation network, the Dual Attention Network (DA-Net), which processes directly on the time-domain samples of monophonic music. DA-Net includes six Dual Attention Modules (DA-Modules), and each of them includes two kinds of attention: element-wise and channel-wise attention. DA-Net is to perform element attention and channel attention operations on convolution features, which reflects the idea of \"symmetry\". DA-Modules can model the semantic interdependencies between element-wise and channel-wise features. In the DA-Module, the element-wise attention mechanism is realized by a Convolutional Gated Linear Unit (ConvGLU), and the channel-wise attention mechanism is realized by a Squeeze-and-Excitation (SE) block. We explored three kinds of combination modes (serial mode, parallel mode, and tightly coupled mode) of the element-wise attention and channel-wise attention. Element-wise attention selectively emphasizes useful features by re-weighting the features at all positions. Channel-wise attention can learn to use global information to selectively emphasize the informative feature maps and suppress the less useful ones. Therefore, DA-Net adaptively integrates the local features with their global dependencies. The outputs of DA-Net are fed into a fully connected layer to generate a 360-dimensional vector corresponding to 360 pitches. We trained the proposed network on the iKala and MDB-stem-synth datasets, respectively. According to the experimental results, our proposed dual attention network with tightly coupled mode achieved the best performance.<\/jats:p>","DOI":"10.3390\/sym13071296","type":"journal-article","created":{"date-parts":[[2021,7,19]],"date-time":"2021-07-19T10:07:37Z","timestamp":1626689257000},"page":"1296","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Dual Attention Network for Pitch Estimation of Monophonic Music"],"prefix":"10.3390","volume":"13","author":[{"given":"Wenfang","family":"Ma","sequence":"first","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Signal Detection and Processing in Xinjiang, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ying","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Signal Detection and Processing in Xinjiang, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"},{"name":"Key Laboratory of Multilingual Information Technology in Xinjiang, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,7,19]]},"reference":[{"key":"ref_1","first-page":"155","article-title":"Medleydb: A multitrack dataset for annotation-intensive mir research","volume":"14","author":"Bittner","year":"2014","journal-title":"ISMIR"},{"key":"ref_2","unstructured":"Bosch, J., and G\u00f3mez, E. (2014, January 4\u20136). Melody extraction in symphonic classical music: A comparative study of mutual agreement between humans and algorithms. Proceedings of the 9th Conference on Interdisciplinary Musicology\u2014CIM14, Berlin, Germany."},{"key":"ref_3","unstructured":"Mauch, M., Cannam, C., Bittner, R., Fazekas, G., Salamon, J., Dai, J., Bello, J., and Dixon, S. (2021, April 15). Computer-aided melody note transcription using the Tony software: Accuracy and efficiency. Available online: https:\/\/qmro.qmul.ac.uk\/xmlui\/handle\/123456789\/7247."},{"key":"ref_4","unstructured":"Rodet, X. (2002, January 15). Synthesis and processing of the singing voice. Proceedings of the 1st IEEE Benelux Workshop on Model based Processing and Coding of Audio (MPCA-2002), Leuven, Belgium."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1002\/aris.1440370108","article-title":"Music information retrieval","volume":"37","author":"Downie","year":"2003","journal-title":"Annu. Rev. Inf. Sci. Technol."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Panteli, M., Bittner, R., Bello, J.P., and Dixon, S. (2017, January 5\u20139). Towards the characterization of singing styles in world music. Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA.","DOI":"10.1109\/ICASSP.2017.7952233"},{"key":"ref_7","unstructured":"Klapuri, A. (1998). Automatic transcription of music. [Master\u2019s Thesis, Tampere University of Technology]."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1917","DOI":"10.1121\/1.1458024","article-title":"YIN, a fundamental frequency estimator for speech and music","volume":"111","author":"Kawahara","year":"2002","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Mauch, M., and Dixon, S. (2014, January 4\u20139). pYIN: A fundamental frequency estimator using probabilistic threshold distributions. Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853678"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1638","DOI":"10.1121\/1.2951592","article-title":"A sawtooth waveform inspired pitch estimator for speech and music","volume":"124","author":"Camacho","year":"2008","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"4559","DOI":"10.1121\/1.2916590","article-title":"A spectral\/temporal method for robust fundamental frequency tracking","volume":"123","author":"Zahorian","year":"2008","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Hua, K. (2017). Nebula: F0 estimation and voicing detection by modeling the statistical properties of feature extractors. arXiv.","DOI":"10.21437\/Interspeech.2018-1258"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhang, J., Tang, J., and Dai, L.-R. (2016, January 8\u201312). RNN-BLSTM Based Multi-Pitch Estimation. Proceedings of the Interspeech, San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-117"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"518","DOI":"10.1109\/TASLP.2013.2295918","article-title":"PEFAC-a pitch estimation algorithm robust to high levels of noise","volume":"22","author":"Gonzalez","year":"2014","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Verma, P., and Schafer, R.W. (2016, January 8\u201312). Frequency Estimation from Waveforms Using Multi-Layered Neural Networks. Proceedings of the Interspeech, San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-679"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2158","DOI":"10.1109\/TASLP.2014.2363410","article-title":"Neural network based pitch tracking in very noisy speech","volume":"22","author":"Han","year":"2014","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_17","unstructured":"Bittner, R.M., McFee, B., Salamon, J., Li, P., and Bello, J.P. (2017, January 23\u201327). Deep Salience Representations for F0 Estimation in Polyphonic Music. Proceedings of the ISMIR, Suzhou, China."},{"key":"ref_18","unstructured":"Bittner, R.M., McFee, B., and Bello, J.P. (2018). Multitask learning for fundamental frequency estimation in music. arXiv."},{"key":"ref_19","unstructured":"Basaran, D., Essid, S., and Peeters, G. (2018, January 23\u201327). Main melody extraction with source-filter nmf and crnn. Proceedings of the 19th International Society for Music Information Retreival, Paris, France."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Doras, G., Esling, P., and Peeters, G. (2019, January 24\u201325). On the use of u-net for dominant melody estimation in polyphonic music. Proceedings of the 2019 International Workshop on Multilayer Music Representation and Processing (MMRP), Milano, Italy.","DOI":"10.1109\/MMRP.2019.8665373"},{"key":"ref_21","unstructured":"Lu, W.T., and Su, L. (2018, January 23\u201327). Vocal Melody Extraction with Semantic Segmentation and Audio-symbolic Domain Transfer Learning. Proceedings of the ISMIR, Paris, France."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Chen, M.-T., Li, B.-J., and Chi, T.-S. (2019, January 12\u201317). Cnn based two-stage multi-resolution end-to-end model for singing melody extraction. Proceedings of the ICASSP 2019\u20142019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683630"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Hsieh, T.-H., Su, L., and Yang, Y.-H. (2019, January 12\u201317). A streamlined encoder\/decoder architecture for melody extraction. Proceedings of the ICASSP 2019\u20142019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8682389"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Xu, S., and Shimodaira, H. (2019, January 15\u201319). Direct F0 Estimation with Neural-Network-Based Regression. Proceedings of the Interspeech, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-3267"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Airaksinen, M., Juvela, L., Alku, P., and R\u00e4s\u00e4nen, O. (2019, January 12\u201317). Data Augmentation Strategies for Neural Network F0 Estimation. Proceedings of the ICASSP 2019\u20142019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8683041"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Kim, J.W., Salamon, J., Li, P., and Bello, J.P. (2018, January 15\u201320). Crepe: A convolutional representation for pitch estimation. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461329"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Ardaillon, L., and Roebel, A. (2019, January 15\u201319). Fully-convolutional network for pitch estimation of speech signals. Proceedings of the Insterspeech 2019, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-2815"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Dong, M., Wu, J., and Luan, J. (2019, January 15\u201319). Vocal Pitch Extraction in Polyphonic Music Using Convolutional Residual Network. Proceedings of the Insterspeech 2019, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-2286"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Gfeller, B., Frank, C., Roblek, D., Sharifi, M., Tagliasacchi, M., and Velimirovi\u0107, M. (2020). Pitch Estimation Via Self-Supervision. Proceedings of the ICASSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 4\u20138 May 2020, IEEE.","DOI":"10.1109\/ICASSP40776.2020.9053798"},{"key":"ref_30","unstructured":"Dauphin, Y.N., Fan, A., Auli, M., and Grangier, D. (2017, January 6\u201311). Language modeling with gated convolutional networks. Proceedings of the International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"380","DOI":"10.1109\/TASLP.2019.2955276","article-title":"Learning complex spectral mapping with gated convolutional recurrent networks for monaural speech enhancement","volume":"28","author":"Tan","year":"2019","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1109\/TASLP.2018.2876171","article-title":"Gated residual networks with dilated convolutions for monaural speech enhancement","volume":"27","author":"Tan","year":"2018","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Shi, Z., Lin, H., Liu, L., Liu, R., Han, J., and Shi, A. (2019, January 15\u201319). Deep Attention Gated Dilated Temporal Convolutional Networks with Intra-Parallel Convolutional Modules for End-to-End Monaural Speech Separation. Proceedings of the Interspeech, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-1373"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Geng, H., Hu, Y., and Huang, H. (2020). Monaural Singing Voice and Accompaniment Separation Based on Gated Nested U-Net Architecture. Symmetry, 12.","DOI":"10.3390\/sym12061051"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"065016","DOI":"10.1063\/1.5100577","article-title":"Deep convolutional neural network based on densely connected squeeze-and-excitation blocks","volume":"9","author":"Wu","year":"2019","journal-title":"AIP Adv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Gu, J., Sun, X., Zhang, Y., Fu, K., and Wang, L. (2019). Deep residual squeeze and excitation network for remote sensing image super-resolution. Remote. Sens., 11.","DOI":"10.3390\/rs11151817"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Park, Y.J., Tuxworth, G., and Zhou, J. (2019, January 22\u201325). Insect classification using Squeeze-and-Excitation and attention modules-a benchmark study. Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP), Taipei, Taiwan.","DOI":"10.1109\/ICIP.2019.8803746"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Wu, J., Li, Q., Liang, S., and Kuang, S.-F. (2020, January 29\u201331). Convolutional Neural Network with Squeeze and Excitation Modules for Image Blind Deblurring. Proceedings of the 2020 Information Communication Technologies Conference (ICTC), Nanjing, China.","DOI":"10.1109\/ICTC49638.2020.9123259"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., and Lu, H. (2019, January 15\u201320). Dual attention network for scene segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00326"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Yu, S., Sun, X., Yu, Y., and Li, W. (2021, January 6\u201312). Frequency-temporal attention network for singing melody extraction. Proceedings of the ICASSP 2021\u2014-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413444"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Hu, Z., Luo, Y., Lin, J., Yan, Y., and Chen, J. (2019, January 10\u201316). Multi-Level Visual-Semantic Alignments with Relation-Wise Dual Attention Network for Image and Text Matching. Proceedings of the IJCAI 2019, Macao.","DOI":"10.24963\/ijcai.2019\/111"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Li, B., Ye, W., Sheng, Z., Xie, R., Xi, X., and Zhang, S. (2020, January 13\u201318). Graph enhanced dual attention network for document-level relation extraction. Proceedings of the 28th International Conference on Computational Linguistics, Barcelona, Spain.","DOI":"10.18653\/v1\/2020.coling-main.136"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"8858717","DOI":"10.1155\/2020\/8858717","article-title":"Interactive dual attention network for text sentiment classification","volume":"2020","author":"Zhu","year":"2020","journal-title":"Comput. Intell. Neurosci."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"6302","DOI":"10.1109\/JSTARS.2021.3083055","article-title":"DA-RoadNet: A Dual-Attention Network for Road Extraction from High Resolution Satellite Imagery","volume":"14","author":"Wan","year":"2021","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Chan, T.-S., Yeh, T.-C., Fan, Z.-C., Chen, H.-W., Su, L., Yang, Y.-H., and Jang, R. (2015, January 19\u201324). Vocal activity informed singing voice separation with the iKala dataset. Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, QLD, Australia.","DOI":"10.1109\/ICASSP.2015.7178063"},{"key":"ref_47","unstructured":"Salamon, J., Bittner, R.M., Bonada, J., Bosch, J.J., G\u00f3mez Guti\u00e9rrez, E., and Bello, J.P. (2017, January 23\u201327). An analysis\/synthesis framework for automatic f0 annotation of multitrack datasets. Proceedings of the ISMIR 2017 Proceedings of the 18th International Society for Music Information Retrieval Conference, Suzhou, China."},{"key":"ref_48","unstructured":"Kingman, D.P., and Ba, J. (2015, January 7\u20139). Adam: A Method for Stochastic Optimization. Conference paper. Proceedings of the 3rd International Conference for Learning Representations, San Diego, CA, USA."},{"key":"ref_49","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 6\u201311). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_50","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_51","unstructured":"Raffel, C., McFee, B., Humphrey, E.J., Salamon, J., Nieto, O., Liang, D., Ellis, D.P.W., and Raffel, C.C. (2014, January 27\u201331). mir_eval: A transparent implementation of common MIR metrics. Proceedings of the 15th International Society for Music Information Retrieval Conference, ISMIR 2014, Taipei, Taiwan."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/7\/1296\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:31:48Z","timestamp":1760164308000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/13\/7\/1296"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,19]]},"references-count":51,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2021,7]]}},"alternative-id":["sym13071296"],"URL":"https:\/\/doi.org\/10.3390\/sym13071296","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2021,7,19]]}}}