{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T12:21:22Z","timestamp":1756383682147,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":23,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,9,7]],"date-time":"2022-09-07T00:00:00Z","timestamp":1662508800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,9,7]]},"DOI":"10.1145\/3549737.3549754","type":"proceedings-article","created":{"date-parts":[[2022,9,9]],"date-time":"2022-09-09T16:29:59Z","timestamp":1662740999000},"page":"1-8","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Transformer-Based Music Language Modelling and Transcription"],"prefix":"10.1145","author":[{"given":"Christos","family":"Zonios","sequence":"first","affiliation":[{"name":"University of Ioannina, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"John","family":"Pavlopoulos","sequence":"additional","affiliation":[{"name":"Ca' Foscari, University of Venice, Italy and Athens University of Economics and Business, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Aristidis","family":"Likas","sequence":"additional","affiliation":[{"name":"University of Ioannina, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,9,9]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3366424.3383542"},{"key":"e_1_3_2_1_2_1","unstructured":"Alexei Baevski Henry Zhou Abdelrahman Mohamed and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. arXiv preprint arXiv:2006.11477(2020).  Alexei Baevski Henry Zhou Abdelrahman Mohamed and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. arXiv preprint arXiv:2006.11477(2020)."},{"key":"e_1_3_2_1_3_1","first-page":"19","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","volume":"1","author":"Baziotis Christos","year":"2019","unstructured":"Christos Baziotis , Ion Androutsopoulos , Ioannis Konstas , and Alexandros Potamianos . 2019 . SEQ\u23033: Differentiable Sequence-to-Sequence-to-Sequence Autoencoder for Unsupervised Abstractive Sentence Compression . In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, 673\u2013681. https:\/\/doi.org\/10. 18653\/v1\/N 19 - 1071 10.18653\/v1 Christos Baziotis, Ion Androutsopoulos, Ioannis Konstas, and Alexandros Potamianos. 2019. SEQ\u23033: Differentiable Sequence-to-Sequence-to-Sequence Autoencoder for Unsupervised Abstractive Sentence Compression. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, 673\u2013681. https:\/\/doi.org\/10.18653\/v1\/N19-1071"},{"key":"e_1_3_2_1_4_1","volume-title":"Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150(2020).","author":"Beltagy Iz","year":"2020","unstructured":"Iz Beltagy , Matthew\u00a0 E Peters , and Arman Cohan . 2020 . Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150(2020). Iz Beltagy, Matthew\u00a0E Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150(2020)."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2018.2869928"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10844-013-0258-3"},{"key":"e_1_3_2_1_7_1","unstructured":"Tom\u00a0B Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165(2020).  Tom\u00a0B Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165(2020)."},{"key":"e_1_3_2_1_8_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018).","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805(2018)."},{"volume-title":"Neural Networks for Automatic Polyphonic Piano Music Transcription","author":"Ender Johnathon\u00a0Michael","key":"e_1_3_2_1_9_1","unstructured":"Johnathon\u00a0Michael Ender . 2018. Neural Networks for Automatic Polyphonic Piano Music Transcription . University of Colorado Colorado Springs. Johnathon\u00a0Michael Ender. 2018. Neural Networks for Automatic Polyphonic Piano Music Transcription. University of Colorado Colorado Springs."},{"key":"e_1_3_2_1_10_1","unstructured":"Curtis Hawthorne Erich Elsen Jialin Song Adam Roberts Ian Simon Colin Raffel Jesse Engel Sageev Oore and Douglas Eck. 2017. Onsets and frames: Dual-objective piano transcription. arXiv preprint arXiv:1710.11153(2017).  Curtis Hawthorne Erich Elsen Jialin Song Adam Roberts Ian Simon Colin Raffel Jesse Engel Sageev Oore and Douglas Eck. 2017. Onsets and frames: Dual-objective piano transcription. arXiv preprint arXiv:1710.11153(2017)."},{"key":"e_1_3_2_1_11_1","unstructured":"Curtis Hawthorne Ian Simon Rigel Swavely Ethan Manilow and Jesse Engel. 2021. Sequence-to-sequence piano transcription with transformers. arXiv preprint arXiv:2107.09142(2021).  Curtis Hawthorne Ian Simon Rigel Swavely Ethan Manilow and Jesse Engel. 2021. Sequence-to-sequence piano transcription with transformers. arXiv preprint arXiv:2107.09142(2021)."},{"key":"e_1_3_2_1_12_1","unstructured":"Curtis Hawthorne Andriy Stasyuk Adam Roberts Ian Simon Cheng-Zhi\u00a0Anna Huang Sander Dieleman Erich Elsen Jesse Engel and Douglas Eck. 2018. Enabling factorized piano music modeling and generation with the MAESTRO dataset. arXiv preprint arXiv:1810.12247(2018).  Curtis Hawthorne Andriy Stasyuk Adam Roberts Ian Simon Cheng-Zhi\u00a0Anna Huang Sander Dieleman Erich Elsen Jesse Engel and Douglas Eck. 2018. Enabling factorized piano music modeling and generation with the MAESTRO dataset. arXiv preprint arXiv:1810.12247(2018)."},{"key":"e_1_3_2_1_13_1","unstructured":"Cheng-Zhi\u00a0Anna Huang Ashish Vaswani Jakob Uszkoreit Noam Shazeer Ian Simon Curtis Hawthorne Andrew\u00a0M Dai Matthew\u00a0D Hoffman Monica Dinculescu and Douglas Eck. 2018. Music transformer. arXiv preprint arXiv:1809.04281(2018).  Cheng-Zhi\u00a0Anna Huang Ashish Vaswani Jakob Uszkoreit Noam Shazeer Ian Simon Curtis Hawthorne Andrew\u00a0M Dai Matthew\u00a0D Hoffman Monica Dinculescu and Douglas Eck. 2018. Music transformer. arXiv preprint arXiv:1809.04281(2018)."},{"key":"e_1_3_2_1_14_1","unstructured":"Jong\u00a0Wook Kim and Juan\u00a0Pablo Bello. 2019. Adversarial learning for improved onsets and frames music transcription. arXiv preprint arXiv:1906.08512(2019).  Jong\u00a0Wook Kim and Juan\u00a0Pablo Bello. 2019. Adversarial learning for improved onsets and frames music transcription. arXiv preprint arXiv:1906.08512(2019)."},{"key":"e_1_3_2_1_15_1","unstructured":"Colin Raffel Noam Shazeer Adam Roberts Katherine Lee Sharan Narang Michael Matena Yanqi Zhou Wei Li and Peter\u00a0J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683(2019).  Colin Raffel Noam Shazeer Adam Roberts Katherine Lee Sharan Narang Michael Matena Yanqi Zhou Wei Li and Peter\u00a0J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683(2019)."},{"key":"e_1_3_2_1_16_1","unstructured":"Justin\u00a0J Salamon 2013. Melody extraction from polyphonic music signals. Ph.\u00a0D. Dissertation. Universitat Pompeu Fabra.  Justin\u00a0J Salamon 2013. Melody extraction from polyphonic music signals. Ph.\u00a0D. Dissertation. Universitat Pompeu Fabra."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/78.650093"},{"key":"e_1_3_2_1_18_1","unstructured":"Siddharth Sigtia Emmanouil Benetos Srikanth Cherla Tillman Weyde A Garcez and Simon Dixon. 2014. RNN-based music language models for improving automatic music transcription. (2014).  Siddharth Sigtia Emmanouil Benetos Srikanth Cherla Tillman Weyde A Garcez and Simon Dixon. 2014. RNN-based music language models for improving automatic music transcription. (2014)."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2016.2533858"},{"key":"e_1_3_2_1_20_1","unstructured":"Jonathan Sleep. 2017. Automatic music transcription with convolutional neural networks using intuitive filter shapes. (2017).  Jonathan Sleep. 2017. Automatic music transcription with convolutional neural networks using intuitive filter shapes. (2017)."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.2307\/1417526"},{"key":"e_1_3_2_1_22_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan\u00a0N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998\u20136008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan\u00a0N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998\u20136008."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Jesse Vig. 2019. A Multiscale Visualization of Attention in the Transformer Model. arxiv:1906.05714\u00a0[cs.HC]  Jesse Vig. 2019. A Multiscale Visualization of Attention in the Transformer Model. arxiv:1906.05714\u00a0[cs.HC]","DOI":"10.18653\/v1\/P19-3007"}],"event":{"name":"SETN 2022: 12th Hellenic Conference on Artificial Intelligence","acronym":"SETN 2022","location":"Corfu Greece"},"container-title":["Proceedings of the 12th Hellenic Conference on Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3549737.3549754","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3549737.3549754","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:09:55Z","timestamp":1750183795000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3549737.3549754"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,7]]},"references-count":23,"alternative-id":["10.1145\/3549737.3549754","10.1145\/3549737"],"URL":"https:\/\/doi.org\/10.1145\/3549737.3549754","relation":{},"subject":[],"published":{"date-parts":[[2022,9,7]]},"assertion":[{"value":"2022-09-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}