{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,20]],"date-time":"2026-01-20T03:00:41Z","timestamp":1768878041353,"version":"3.49.0"},"reference-count":277,"publisher":"Association for Computing Machinery (ACM)","issue":"14s","license":[{"start":{"date-parts":[[2023,7,17]],"date-time":"2023-07-17T00:00:00Z","timestamp":1689552000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2023,12,31]]},"abstract":"<jats:p>This survey documents representation approaches for classification across different modalities, from purely content-based methods to techniques utilizing external sources of structured knowledge. We present studies related to three paradigms used for representation, namely (a) low-level template-matching methods, (b) aggregation-based approaches, and (c) deep representation learning systems. We then describe existing resources of structure knowledge and elaborate on the need for enriching representations with such information. Approaches that utilize knowledge resources are presented next, organized with respect to how external information is exploited, i.e., (a) input enrichment and modification, (b) knowledge-based refinement and (c) end-to-end knowledge-aware systems. We subsequently provide a high-level discussion to summarize and compare strengths\/weaknesses of the representation\/enrichment paradigms proposed, and conclude the survey with an overview of relevant research findings and possible directions for future work.<\/jats:p>","DOI":"10.1145\/3583682","type":"journal-article","created":{"date-parts":[[2023,2,13]],"date-time":"2023-02-13T12:47:13Z","timestamp":1676292433000},"page":"1-40","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Content-based and Knowledge-enriched Representations for Classification Across Modalities: A Survey"],"prefix":"10.1145","volume":"55","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9817-7618","authenticated-orcid":false,"given":"Nikiforos","family":"Pittaras","sequence":"first","affiliation":[{"name":"National Kapodistrian University of Athens and NCSR \u201cDemokritos\u201d"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2459-589X","authenticated-orcid":false,"given":"George","family":"Giannakopoulos","sequence":"additional","affiliation":[{"name":"NCSR \u201cDemokritos\u201d"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7468-1502","authenticated-orcid":false,"given":"Panagiotis","family":"Stamatopoulos","sequence":"additional","affiliation":[{"name":"National Kapodistrian University of Athens"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7062-9424","authenticated-orcid":false,"given":"Vangelis","family":"Karkaletsis","sequence":"additional","affiliation":[{"name":"NCSR \u201cDemokritos\u201d"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,7,17]]},"reference":[{"key":"e_1_3_3_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2870052"},{"key":"e_1_3_3_3_1","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1007\/978-3-030-12767-1_5","article-title":"No free lunch theorem: A review","author":"Adam Stavros P.","year":"2019","unstructured":"Stavros P. Adam, Stamatios-Aggelos N. Alexandropoulos, Panos M. Pardalos, and Michael N. Vrahatis. 2019. No free lunch theorem: A review. Approximation and Optimization 145 (2019), 57\u201382.","journal-title":"Approximation and Optimization"},{"key":"e_1_3_3_4_1","doi-asserted-by":"crossref","first-page":"285","DOI":"10.1007\/978-3-319-14142-8_10","volume-title":"Proceedings of the Data Mining","author":"Aggarwal C. C.","year":"2015","unstructured":"C. C. Aggarwal. 2015. Data classification. In Proceedings of the Data Mining. Springer, 285\u2013344."},{"key":"e_1_3_3_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33783-3_16"},{"key":"e_1_3_3_6_1","first-page":"73","volume-title":"Proceedings of the ISMIR","author":"Allik Alo","year":"2016","unstructured":"Alo Allik, Gy\u00f6rgy Fazekas, and Mark B. Sandler. 2016. An ontology for audio features. In Proceedings of the ISMIR. 73\u201379."},{"key":"e_1_3_3_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2018.08.001"},{"key":"e_1_3_3_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2724727"},{"key":"e_1_3_3_9_1","first-page":"1","volume-title":"Proceedings of the 2012 International Conference on Information Technology and e-Services","author":"Amrouch Siham","year":"2012","unstructured":"Siham Amrouch and Sihem Mostefai. 2012. Survey on the literature of ontology mapping, alignment and merging. In Proceedings of the 2012 International Conference on Information Technology and e-Services. IEEE, 1\u20135."},{"key":"e_1_3_3_10_1","volume-title":"Comparing Latent Dirichlet Allocation and Latent Semantic Analysis as Classifiers","author":"Anaya L. H.","year":"2011","unstructured":"L. H. Anaya. 2011. Comparing Latent Dirichlet Allocation and Latent Semantic Analysis as Classifiers. ERIC."},{"key":"e_1_3_3_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2011.2125954"},{"key":"e_1_3_3_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/2390940.2390943"},{"key":"e_1_3_3_13_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0181142"},{"key":"e_1_3_3_14_1","doi-asserted-by":"crossref","unstructured":"N. Aye F. Hattori and K. Kuwabara. 2008. Use of ontologies for bridging semantic gaps in distant communication. International Conference on Innovations in Information Technology (2008) 371\u2013375.","DOI":"10.1109\/INNOVATIONS.2008.4781725"},{"key":"e_1_3_3_15_1","first-page":"2200","volume-title":"Proceedings of the Lrec","volume":"10","author":"Baccianella S.","year":"2010","unstructured":"S. Baccianella, A. Esuli, and F. Sebastiani. 2010. Sentiwordnet 3.0: An enhanced lexical resource for sentiment analysis and opinion mining. In Proceedings of the Lrec, Vol. 10. 2200\u20132204."},{"key":"e_1_3_3_16_1","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1016\/j.engappai.2014.06.012","article-title":"A novel framework for termset selection and weighting in binary text classification","volume":"35","author":"Badawi Dima","year":"2014","unstructured":"Dima Badawi and Hakan Alt\u0131n\u00e7ay. 2014. A novel framework for termset selection and weighting in binary text classification. Engineering Applications of Artificial Intelligence 35 (2014), 38\u201353.","journal-title":"Engineering Applications of Artificial Intelligence"},{"key":"e_1_3_3_17_1","unstructured":"Dzmitry Bahdanau Kyung Hyun Cho and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations (ICLR\u201915) ."},{"issue":"5","key":"e_1_3_3_18_1","first-page":"372","article-title":"Intelligent preprocessing and classification of audio signals","volume":"55","author":"Bai M. R.","year":"2007","unstructured":"M. R. Bai and M. Chen. 2007. Intelligent preprocessing and classification of audio signals. Journal of the Audio Engineering Society 55, 5 (2007), 372\u2013384.","journal-title":"Journal of the Audio Engineering Society"},{"key":"e_1_3_3_19_1","first-page":"86","volume-title":"Proceedings of the 17th International Conference on Computational Linguistics-Volume 1","author":"Baker C. F.","year":"1998","unstructured":"C. F. Baker, C. J. Fillmore, and J. B. Lowe. 1998. The berkeley framenet project. In Proceedings of the 17th International Conference on Computational Linguistics-Volume 1. Association for Computational Linguistics, Morgan Kaufmann Publishers \/ ACL, 86\u201390."},{"key":"e_1_3_3_20_1","first-page":"457","volume-title":"Proceedings of the 2014 IEEE International Conference on Systems, Man, and Cybernetics","author":"Baniya B. K.","year":"2014","unstructured":"B. K. Baniya, J. Lee, and Z. Li. 2014. Audio feature reduction and analysis for automatic music genre classification. In Proceedings of the 2014 IEEE International Conference on Systems, Man, and Cybernetics. IEEE, 457\u2013462."},{"key":"e_1_3_3_21_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1023"},{"key":"e_1_3_3_22_1","doi-asserted-by":"publisher","DOI":"10.1016\/0165-1684(89)90079-0"},{"key":"e_1_3_3_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2007.09.014"},{"key":"e_1_3_3_24_1","unstructured":"S\u00f6ren Becker Marcel Ackermann Sebastian Lapuschkin Klaus-Robert M\u00fcller and Wojciech Samek. 2018. Interpreting and Explaining Deep Neural Networks for Classification of audio signals. CoRR abs\/1807.03418."},{"key":"e_1_3_3_25_1","volume-title":"Dynamic Programming","author":"Bellman R.","year":"2013","unstructured":"R. Bellman. 2013. Dynamic Programming. Courier Corporation."},{"key":"e_1_3_3_26_1","doi-asserted-by":"publisher","DOI":"10.1561\/2200000006"},{"key":"e_1_3_3_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2011.6033302"},{"key":"e_1_3_3_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.50"},{"key":"e_1_3_3_29_1","first-page":"12","article-title":"The curse of dimensionality for local kernel machines","volume":"1258","author":"Bengio Y.","year":"2005","unstructured":"Y. Bengio, O. Delalleau, and N. Le Roux. 2005. The curse of dimensionality for local kernel machines. Techn. Rep. 1258 (2005), 12.","journal-title":"Techn. Rep."},{"key":"e_1_3_3_30_1","first-page":"153","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Bengio Y.","year":"2007","unstructured":"Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle. 2007. Greedy layer-wise training of deep networks. In Proceedings of the Advances in Neural Information Processing Systems. MIT Press, Cambridge, MA., 153\u2013160."},{"key":"e_1_3_3_31_1","first-page":"III\u2013613","volume-title":"Proceedings of the 2003 International Conference on Image Processing","author":"Benitez A. B.","year":"2003","unstructured":"A. B. Benitez and S. Chang. 2003. Image classification using multimedia knowledge networks. In Proceedings of the 2003 International Conference on Image Processing. IEEE, III\u2013613."},{"key":"e_1_3_3_32_1","first-page":"496","volume-title":"Proceedings of the Tenth International Conference on Language Resources and Evaluation","author":"Bertero D.","year":"2016","unstructured":"D. Bertero and P. Fung. 2016. Deep learning of audio and language features for humor prediction. In Proceedings of the Tenth International Conference on Language Resources and Evaluation. 496\u2013501."},{"key":"e_1_3_3_33_1","doi-asserted-by":"publisher","DOI":"10.13176\/11.191"},{"key":"e_1_3_3_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2017.2683799"},{"key":"e_1_3_3_35_1","first-page":"351","volume-title":"Proceedings of the Asian Conference on Computer Vision","author":"Binder A.","year":"2009","unstructured":"A. Binder, M. Kawanabe, and U. Brefeld. 2009. Efficient classification of images with taxonomies. In Proceedings of the Asian Conference on Computer Vision. Springer, 351\u2013362."},{"key":"e_1_3_3_36_1","first-page":"993","article-title":"Latent dirichlet allocation","volume":"3","author":"Blei D. M.","year":"2003","unstructured":"D. M. Blei, A. Y. Ng, and M. I. Jordan. 2003. Latent dirichlet allocation. Journal of Machine Learning Research 3, Jan (2003), 993\u20131022.","journal-title":"Journal of Machine Learning Research"},{"issue":"4","key":"e_1_3_3_37_1","first-page":"1","article-title":"A better decision tree: The max-cut decision tree with modified PCA improves accuracy and running time","volume":"3","author":"Bodine Jonathan","year":"2022","unstructured":"Jonathan Bodine and Dorit S. Hochbaum. 2022. A better decision tree: The max-cut decision tree with modified PCA improves accuracy and running time. SN Computer Science 3, 4 (2022), 1\u201318.","journal-title":"SN Computer Science"},{"key":"e_1_3_3_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376746"},{"key":"e_1_3_3_39_1","first-page":"2787","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Bordes A.","year":"2013","unstructured":"A. Bordes, N. Usunier, A. Garcia-Duran, J. Weston, and O. Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. In Proceedings of the Advances in Neural Information Processing Systems. 2787\u20132795."},{"key":"e_1_3_3_40_1","unstructured":"A. Borghesi F. Baldo and M. Milano. 2020. Improving deep learning models via constraint-based domain knowledge: A brief survey. arXiv:2005.10691. Retrieved from https:\/\/arxiv.org\/abs\/2005.10691."},{"key":"e_1_3_3_41_1","volume-title":"Proceedings of the 1999 International Computer Music Conference.","author":"Boyer H.","year":"1999","unstructured":"H. Boyer, X. Serra, and G. Peeters. 1999. Audio descriptors and descriptor schemes in the context of MPEG-7. In Proceedings of the 1999 International Computer Music Conference. International Computer Music Conference."},{"key":"e_1_3_3_42_1","volume-title":"The Fourier Transform and its Applications","author":"Bracewell R. N.","year":"1986","unstructured":"R. N. Bracewell and R. N. Bracewell. 1986. The Fourier Transform and its Applications. Vol. 31999. McGraw-Hill New York."},{"key":"e_1_3_3_43_1","volume-title":"Affective Norms for English Words (ANEW): Instruction Manual and Affective Ratings","author":"Bradley M. M.","year":"1999","unstructured":"M. M. Bradley and P. J. Lang. 1999. Affective Norms for English Words (ANEW): Instruction Manual and Affective Ratings. Technical Report. Technical report C-1, the center for research in psychophysiology."},{"key":"e_1_3_3_44_1","first-page":"I\u20131021","volume-title":"Proceedings of the s2002 IEEE International Conference on Acoustics, Speech, and Signal Processing","author":"Burges C. J.","year":"2002","unstructured":"C. J. Burges, J. C. Platt, and S. Jana. 2002. Extracting noise-robust features from audio data. In Proceedings of the s2002 IEEE International Conference on Acoustics, Speech, and Signal Processing. IEEE, I\u20131021."},{"key":"e_1_3_3_45_1","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.12228"},{"issue":"3","key":"e_1_3_3_46_1","first-page":"137","article-title":"A survey on knowledge compilation","volume":"10","author":"Cadoli M.","year":"1997","unstructured":"M. Cadoli and F. M. Donini. 1997. A survey on knowledge compilation. AI Communications 10, 3\u20134 (1997), 137\u2013150. Retrieved from http:\/\/content.iospress.com\/articles\/ai-communications\/aic133.","journal-title":"AI Communications"},{"key":"e_1_3_3_47_1","doi-asserted-by":"publisher","DOI":"10.5555\/1888089.1888148"},{"key":"e_1_3_3_48_1","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.11259"},{"key":"e_1_3_3_49_1","doi-asserted-by":"publisher","DOI":"10.5555\/2815662"},{"key":"e_1_3_3_50_1","volume-title":"Proceedings of the Audio Engineering Society Convention 116","author":"Cano P.","year":"2004","unstructured":"P. Cano, M. Koppenberger, P. Herrera, S. Le Groux, J. Ricard, and N. Wack. 2004. Nearest-neighbor generic sound classification with a WordNet-based taxonomy. In Proceedings of the Audio Engineering Society Convention 116. Audio Engineering Society."},{"issue":"2","key":"e_1_3_3_51_1","first-page":"1","article-title":"Comprehensive survey on distance\/similarity measures between probability density functions","volume":"1","author":"Cha S.","year":"2007","unstructured":"S. Cha. 2007. Comprehensive survey on distance\/similarity measures between probability density functions. City 1, 2 (2007), 1.","journal-title":"City"},{"key":"e_1_3_3_52_1","doi-asserted-by":"crossref","unstructured":"Simyung Chang Hyoungwoo Park Janghoon Cho Hyunsin Park Sungrack Yun and Kyuwoong Hwang. 2021. Subspectral normalization for neural audio data processing. In ICASSP 2021-2021 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP) . IEEE 850\u2013854.","DOI":"10.1109\/ICASSP39728.2021.9413522"},{"key":"e_1_3_3_53_1","first-page":"357","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Chen Chun-Fu Richard","year":"2021","unstructured":"Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. 2021. Crossvit: Cross-attention multi-scale vision transformer for image classification. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 357\u2013366."},{"issue":"15","key":"e_1_3_3_54_1","doi-asserted-by":"crossref","first-page":"10809","DOI":"10.1007\/s00521-018-3442-0","article-title":"Verbal aggression detection on Twitter comments: Convolutional neural network for short-text sentiment analysis","volume":"32","author":"Chen Junyi","year":"2020","unstructured":"Junyi Chen, Shankai Yan, and Ka-Chun Wong. 2020. Verbal aggression detection on Twitter comments: Convolutional neural network for short-text sentiment analysis. Neural Computing and Applications 32, 15 (2020), 10809\u201310818.","journal-title":"Neural Computing and Applications"},{"key":"e_1_3_3_55_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1224"},{"key":"e_1_3_3_56_1","doi-asserted-by":"crossref","first-page":"1505","DOI":"10.1109\/ICME.2008.4607732","volume-title":"Proceedings of the 2008 IEEE International Conference on Multimedia and Expo","author":"Cheng Heng-Tze","year":"2008","unstructured":"Heng-Tze Cheng, Yi-Hsuan Yang, Yu-Ching Lin, I-Bin Liao, and Homer H. Chen. 2008. Automatic chord recognition for music classification and retrieval. In Proceedings of the 2008 IEEE International Conference on Multimedia and Expo. IEEE, 1505\u20131508."},{"key":"e_1_3_3_57_1","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1007\/978-1-0716-0826-5_3","article-title":"Siamese neural networks: An overview","author":"Chicco D.","year":"2021","unstructured":"D. Chicco. 2021. Siamese neural networks: An overview. Artificial Neural Networks 2190 (2021), 73\u201394.","journal-title":"Artificial Neural Networks"},{"key":"e_1_3_3_58_1","doi-asserted-by":"crossref","unstructured":"K. Choi G. Fazekas and M. Sandler. 2016. Explaining deep convolutional neural networks on music classification. CoRR abs\/1607.02444.","DOI":"10.1109\/ICASSP.2017.7952585"},{"key":"e_1_3_3_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2912574"},{"key":"e_1_3_3_60_1","doi-asserted-by":"publisher","DOI":"10.1037\/0033-295X.82.6.407"},{"key":"e_1_3_3_61_1","doi-asserted-by":"crossref","unstructured":"N. Dalal and B. Triggs. 2005. Histograms of oriented gradients for human detection. In Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition . IEEE 886\u2013893.","DOI":"10.1109\/CVPR.2005.177"},{"key":"e_1_3_3_62_1","unstructured":"Marina Danilevsky Kun Qian Ranit Aharonov Yannis Katsis Ban Kawas and Prithviraj Sen. 2020. A Survey of the State of Explainable AI for Natural Language Processing. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing . 447\u2013459."},{"issue":"6","key":"e_1_3_3_63_1","doi-asserted-by":"crossref","first-page":"391","DOI":"10.1002\/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9","article-title":"Indexing by latent semantic analysis","volume":"41","author":"Deerwester S.","year":"1990","unstructured":"S. Deerwester, S. T. Dumais, G. W. Furnas, T. K. Landauer, and R. Harshman. 1990. Indexing by latent semantic analysis. Journal of the American Society for Information Science 41, 6 (1990), 391\u2013407.","journal-title":"Journal of the American Society for Information Science"},{"key":"e_1_3_3_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/2502081.2502171"},{"key":"e_1_3_3_65_1","first-page":"248","volume-title":"Proceedings of the Computer Vision and Pattern Recognition","author":"Deng J.","year":"2009","unstructured":"J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In Proceedings of the Computer Vision and Pattern Recognition. IEEE, 248\u2013255."},{"key":"e_1_3_3_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/97709.97725"},{"key":"e_1_3_3_67_1","first-page":"1777","volume-title":"Proceedings of the CVPR","author":"Deselaers T.","year":"2011","unstructured":"T. Deselaers and V. Ferrari. 2011. Visual and semantic similarity in ImageNet. In Proceedings of the CVPR. IEEE Computer Society, 1777\u20131784. Retrieved from http:\/\/dblp.uni-trier.de\/db\/conf\/cvpr\/cvpr2011.html#DeselaersF11."},{"key":"e_1_3_3_68_1","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin J.","year":"2019","unstructured":"J. Devlin, M. Chang, K. Lee, and K. Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 4171\u20134186."},{"key":"e_1_3_3_69_1","first-page":"647","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Donahue J.","year":"2014","unstructured":"J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell. 2014. Decaf: A deep convolutional activation feature for generic visual recognition. In Proceedings of the International Conference on Machine Learning. PMLR, 647\u2013655."},{"key":"e_1_3_3_70_1","first-page":"0210","volume-title":"Proceedings of the 2018 41st International Convention on Information and Communication Technology, Electronics and Microelectronics","author":"Do\u0161ilovi\u0107 F. K.","year":"2018","unstructured":"F. K. Do\u0161ilovi\u0107, M. Br\u010di\u0107, and N. Hlupi\u0107. 2018. Explainable artificial intelligence: A survey. In Proceedings of the 2018 41st International Convention on Information and Communication Technology, Electronics and Microelectronics. IEEE, 0210\u20130215."},{"key":"e_1_3_3_71_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly Jakob Uszkoreit and Neil Houlsby. 2021. An Image is Worth 16\u00d716 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations ."},{"issue":"1","key":"e_1_3_3_72_1","article-title":"Using WordNet for text categorization","volume":"5","author":"Elberrichi Z.","year":"2008","unstructured":"Z. Elberrichi, A. Rahmoun, and M. A. Bentaalah. 2008. Using WordNet for text categorization. International Arab Journal of Information Technology 5, 1 (2008).","journal-title":"International Arab Journal of Information Technology"},{"key":"e_1_3_3_73_1","first-page":"117","volume-title":"Proceedings of the Machine Learning Challenges Workshop","author":"Everingham M.","year":"2005","unstructured":"M. Everingham, A. Zisserman, C. K. Williams, L. Van Gool, M. Allan, C. M. Bishop, O. Chapelle, N. Dalal, T. Deselaers, and G. Dork\u00f3. 2005. The 2005 pascal visual object classes challenge. In Proceedings of the Machine Learning Challenges Workshop. Springer, 117\u2013176."},{"key":"e_1_3_3_74_1","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Favory Xavier","year":"2020","unstructured":"Xavier Favory, Konstantinos Drossos, Tuomas Virtanen, and Xavier Serra. 2020. COALA: Co-aligned autoencoders for learning semantically enriched audio representations. In Proceedings of the International Conference on Machine Learning."},{"key":"e_1_3_3_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/1871437.1871689"},{"key":"e_1_3_3_76_1","doi-asserted-by":"publisher","DOI":"10.3389\/frobt.2019.00153"},{"key":"e_1_3_3_77_1","doi-asserted-by":"publisher","DOI":"10.1515\/9783110199901.373"},{"key":"e_1_3_3_78_1","doi-asserted-by":"publisher","DOI":"10.1093\/ijl\/16.3.235"},{"issue":"2","key":"e_1_3_3_79_1","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1109\/TMM.2010.2098858","article-title":"A survey of audio-based music classification and annotation","volume":"13","author":"Fu Z.","year":"2010","unstructured":"Z. Fu, G. Lu, K. M. Ting, and D. Zhang. 2010. A survey of audio-based music classification and annotation. IEEE Transactions on Multimedia 13, 2 (2010), 303\u2013319.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_3_80_1","first-page":"758","volume-title":"Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Ganitkevitch J.","year":"2013","unstructured":"J. Ganitkevitch, B. Van Durme, and C. Callison-Burch. 2013. PPDB: The paraphrase database. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 758\u2013764."},{"key":"e_1_3_3_81_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10247-4"},{"key":"e_1_3_3_82_1","volume-title":"Proceedings of the IEEE ICASSP 2017","author":"Gemmeke J. F.","year":"2017","unstructured":"J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter. 2017. Audio set: An ontology and human-labeled dataset for audio events. In Proceedings of the IEEE ICASSP 2017. New Orleans, LA."},{"key":"e_1_3_3_83_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4615-3626-0"},{"key":"e_1_3_3_84_1","first-page":"13","volume-title":"Proceedings of the 2nd International Conference on Web Intelligence, Mining and Semantics","author":"Giannakopoulos G.","year":"2012","unstructured":"G. Giannakopoulos, P. Mavridi, G. Paliouras, G. Papadakis, and K. Tserpes. 2012. Representation models for text classification: A comparative analysis over three web document types. In Proceedings of the 2nd International Conference on Web Intelligence, Mining and Semantics. ACM, 13."},{"key":"e_1_3_3_85_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1004"},{"key":"e_1_3_3_86_1","first-page":"513","article-title":"Neighbourhood components analysis","volume":"17","author":"Goldberger J.","year":"2004","unstructured":"J. Goldberger, G. E. Hinton, S. Roweis, and R. R. Salakhutdinov. 2004. Neighbourhood components analysis. Advances in Neural Information Processing Systems 17 (2004), 513\u2013520.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_87_1","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1016\/B978-0-12-498150-8.50021-5","volume-title":"Proceedings of the Statistical Computation","author":"Golub Gene H.","year":"1969","unstructured":"Gene H. Golub. 1969. Matrix decompositions and statistical calculations. In Proceedings of the Statistical Computation. Elsevier, 365\u2013397."},{"key":"e_1_3_3_88_1","doi-asserted-by":"crossref","unstructured":"Yuan Gong Yu-An Chung and James R. Glass. 2021. AST: Audio spectrogram transformer. In Interspeech 2021 22nd Annual Conference of the Inter National Speech Communication Association (ISCA\u201921 Brno Czechia 30 August-3 September 2021) 571\u2013575.","DOI":"10.21437\/Interspeech.2021-698"},{"key":"e_1_3_3_89_1","unstructured":"Ian J. Goodfellow Jonathon Shlens and Christian Szegedy. 2015. Explaining and harnessing adversarial examples. In 3rd International Conference on Learning Representations (ICLR\u201915 San Diego CA USA May 7-9 2015) Conference Track Proceedings."},{"key":"e_1_3_3_90_1","unstructured":"Roger B. Grosse Rajat Raina Helen Kwong and Andrew Y. Ng. 2007. Shift-Invariance Sparse Coding for Audio Classification. In Proceedings of the Twenty-Third Conference on Uncertainty in Artificial Intelligence (UAI\u201907 Vancouver BC Canada July 19-22 2007) AUAI Press 149\u2013158."},{"key":"e_1_3_3_91_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.10.013"},{"key":"e_1_3_3_92_1","doi-asserted-by":"publisher","DOI":"10.1080\/00437956.1954.11659520"},{"key":"e_1_3_3_93_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.107889"},{"key":"e_1_3_3_94_1","doi-asserted-by":"crossref","unstructured":"H. He and Y. Ma. 2013. Imbalanced learning: Foundations algorithms and applications. Wiley-IEEE Press.","DOI":"10.1002\/9781118646106"},{"key":"e_1_3_3_95_1","first-page":"770","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"He K.","year":"2016","unstructured":"K. He, X. Zhang, S. Ren, and J. Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 770\u2013778."},{"key":"e_1_3_3_96_1","first-page":"131","volume-title":"Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Hershey S.","year":"2017","unstructured":"S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, and B. Seybold. 2017. CNN architectures for large-scale audio classification. In Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 131\u2013135."},{"key":"e_1_3_3_97_1","volume-title":"Distributed Representations","author":"Hinton G. E.","year":"1984","unstructured":"G. E. Hinton, J. L. McClelland, and D. E. Rumelhart. 1984. Distributed Representations. Carnegie-Mellon University Pittsburgh, PA."},{"key":"e_1_3_3_98_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_3_99_1","doi-asserted-by":"publisher","DOI":"10.1145\/1816123.1816146"},{"key":"e_1_3_3_100_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"issue":"1","key":"e_1_3_3_101_1","doi-asserted-by":"crossref","first-page":"106","DOI":"10.1113\/jphysiol.1962.sp006837","article-title":"Receptive fields, binocular interaction and functional architecture in the cat\u2019s visual cortex","volume":"160","author":"Hubel D. H.","year":"1962","unstructured":"D. H. Hubel and T. N. Wiesel. 1962. Receptive fields, binocular interaction and functional architecture in the cat\u2019s visual cortex. The Journal of Physiology 160, 1 (1962), 106\u2013154.","journal-title":"The Journal of Physiology"},{"key":"e_1_3_3_102_1","first-page":"3543","volume-title":"Proceedings of the 2016 IEEE International Conference on Image Processing","author":"Ilea Ioana","year":"2016","unstructured":"Ioana Ilea, Lionel Bombrun, Christian Germain, Romulus Terebes, Monica Borda, and Yannick Berthoumieu. 2016. Texture image classification with Riemannian Fisher vectors. In Proceedings of the 2016 IEEE International Conference on Image Processing. IEEE, 3543\u20133547."},{"key":"e_1_3_3_103_1","volume-title":"Algorithms for Clustering Data","author":"Jain A. K.","year":"1988","unstructured":"A. K. Jain and R. C. Dubes. 1988. Algorithms for Clustering Data. Prentice-Hall."},{"key":"e_1_3_3_104_1","unstructured":"Adit Jamdar Jessica Abraham Karishma Khanna and Rahul Dubey. 2015. Emotion analysis of songs based on lyrical and audio features. CoRR abs\/1506.05012 (2015)."},{"key":"e_1_3_3_105_1","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1075\/cilt.260.12jar","article-title":"Roget\u2019s thesaurus and semantic similarity","volume":"2003","author":"Jarmasz M.","year":"2004","unstructured":"M. Jarmasz and S. Szpakowicz. 2004. Roget\u2019s thesaurus and semantic similarity. Recent Advances in Natural Language Processing III: Selected Papers from RANLP 2003 (2004), 111.","journal-title":"Recent Advances in Natural Language Processing III: Selected Papers from RANLP"},{"key":"e_1_3_3_106_1","unstructured":"Mirantha Jayathilaka Tingting Mu and Uli Sattler. 2021. Ontology-based n-ball Concept Embeddings Informing Few-shot Image Classification. In Machine Learning with Symbolic Methods and Knowledge Graphs co-located with European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2021) Virtual September 17 2021 (CEUR Workshop Proceedings) Vol. 2997. CEUR-WS.org."},{"key":"e_1_3_3_107_1","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems","author":"Jim\u00e9nez A.","year":"2018","unstructured":"A. Jim\u00e9nez, B. Elizalde, and B. Raj. 2018. Sound event classification using ontology-based neural networks. In Proceedings of the Annual Conference on Neural Information Processing Systems."},{"key":"e_1_3_3_108_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2019.2921336"},{"issue":"6","key":"e_1_3_3_109_1","first-page":"1930","article-title":"A comparative study of stemming algorithms","volume":"2","author":"Jivani A. G.","year":"2011","unstructured":"A. G. Jivani. 2011. A comparative study of stemming algorithms. Int. J. Comp. Tech. Appl. 2, 6 (2011), 1930\u20131938. Retrieved from https:\/\/www.researchgate.net\/profile\/Anjali.","journal-title":"Int. J. Comp. Tech. Appl."},{"key":"e_1_3_3_110_1","volume-title":"Graph based Representations in Pattern Recognition","author":"Jolion J.","year":"2012","unstructured":"J. Jolion and W. Kropatsch. 2012. Graph based Representations in Pattern Recognition. Vol. 12. Springer Science & Business Media."},{"key":"e_1_3_3_111_1","doi-asserted-by":"crossref","first-page":"1094","DOI":"10.1007\/978-3-642-04898-2_455","volume-title":"Proceedings of the International Encyclopedia of Statistical Science","author":"Jolliffe I.","year":"2011","unstructured":"I. Jolliffe. 2011. Principal component analysis. In Proceedings of the International Encyclopedia of Statistical Science. Springer, 1094\u20131096."},{"issue":"11","key":"e_1_3_3_112_1","doi-asserted-by":"crossref","first-page":"938","DOI":"10.1002\/asi.1161","article-title":"A conceptual framework and empirical research for classifying visual descriptors","volume":"52","author":"J\u00f6rgensen C.","year":"2001","unstructured":"C. J\u00f6rgensen, A. Jaimes, A. B. Benitez, and S. Chang. 2001. A conceptual framework and empirical research for classifying visual descriptors. Journal of the American Society for Information Science and Technology 52, 11 (2001), 938\u2013947.","journal-title":"Journal of the American Society for Information Science and Technology"},{"key":"e_1_3_3_113_1","doi-asserted-by":"crossref","unstructured":"A. Joulin E. Grave P. Bojanowski and T. Mikolov. 2017. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL\u201917 Valencia Spain April 3-7 2017) Volume 2: Short Papers. Association for Computational Linguistics 427\u2013431.","DOI":"10.18653\/v1\/E17-2068"},{"key":"e_1_3_3_114_1","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324900000048"},{"key":"e_1_3_3_115_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.567"},{"key":"e_1_3_3_116_1","doi-asserted-by":"crossref","first-page":"1942","DOI":"10.1109\/ICNN.1995.488968","volume-title":"Proceedings of the ICNN\u201995-International Conference on Neural Networks","volume":"4","author":"Kennedy J.","year":"1995","unstructured":"J. Kennedy and R. Eberhart. 1995. Particle swarm optimization. In Proceedings of the ICNN\u201995-International Conference on Neural Networks, Vol. 4. IEEE, 1942\u20131948."},{"key":"e_1_3_3_117_1","doi-asserted-by":"crossref","unstructured":"S. Kim P. Georgiou and S. Narayanan. 2012. Latent acoustic topic models for unstructured audio classification. APSIPA Transactions on Signal and Information Processing 1 (2012) e6.","DOI":"10.1017\/ATSIP.2012.7"},{"key":"e_1_3_3_118_1","doi-asserted-by":"publisher","DOI":"10.1145\/1509212.1509214"},{"issue":"3","key":"e_1_3_3_119_1","doi-asserted-by":"crossref","first-page":"337","DOI":"10.1080\/09524622.2019.1606734","article-title":"Pre-processing spectrogram parameters improve the accuracy of bioacoustic classification using convolutional neural networks","volume":"29","author":"Knight E. C.","year":"2020","unstructured":"E. C. Knight, S. Poo Hernandez, E. M. Bayne, V. Bulitko, and B. V. Tucker. 2020. Pre-processing spectrogram parameters improve the accuracy of bioacoustic classification using convolutional neural networks. Bioacoustics 29, 3 (2020), 337\u2013355.","journal-title":"Bioacoustics"},{"key":"e_1_3_3_120_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.58325"},{"key":"e_1_3_3_121_1","doi-asserted-by":"crossref","unstructured":"Nikolaos Kolitsas Octavian-Eugen Ganea and Thomas Hofmann. 2018. End-to-End neural entity linking. In Proceedings of the 22nd Conference on Computational Natural Language Learning (CoNLL\u201918 Brussels Belgium October 31 - November 1 2018) Association for Computational Linguistics 519\u2013529.","DOI":"10.18653\/v1\/K18-1050"},{"key":"e_1_3_3_122_1","doi-asserted-by":"crossref","unstructured":"Ranjay Krishna Yuke Zhu Oliver Groth Justin Johnson Kenji Hata Joshua Kravitz Stephanie Chen Yannis Kalantidis Li-Jia Li David A. Shamma Michael S. Bernstein and Li Fei-Fei. 2017. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. Int. J. Comput. Vis. 123 1 (2017) 32\u201373.","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_3_123_1","first-page":"1097","volume-title":"Proceedings of the Adv. Neural Inf. Process. Syst.","author":"Krizhevsky A.","year":"2012","unstructured":"A. Krizhevsky, I. Sutskever, and G. E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Proceedings of the Adv. Neural Inf. Process. Syst.Curran Associates, Inc., 1097\u20131105. Retrieved from http:\/\/papers.nips.cc\/paper\/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf."},{"key":"e_1_3_3_124_1","first-page":"1","volume-title":"Proceedings of the 2013 International Conference on Computer Communication and Informatics","author":"Kumar V.","year":"2013","unstructured":"V. Kumar and S. Minz. 2013. Mood classifiaction of lyrics using SentiWordNet. In Proceedings of the 2013 International Conference on Computer Communication and Informatics. IEEE, 1\u20135."},{"key":"e_1_3_3_125_1","unstructured":"John D. Lafferty Andrew McCallum and Fernando C. N. Pereira. 2001. Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data. In Proceedings of the Eighteenth International Conference on Machine Learning (ICML\u201901 Williams College Williamstown MA USA June 28 - July 1 2001) Morgan Kaufmann 282\u2013289."},{"key":"e_1_3_3_126_1","volume-title":"Proceedings of the ESCOM 2009: 7th Triennial Conference of European Society for the Cognitive Sciences of Music","author":"Laurier C.","year":"2009","unstructured":"C. Laurier, O. Lartillot, T. Eerola, and P. Toiviainen. 2009. Exploring relationships between audio features and emotion in music. In Proceedings of the ESCOM 2009: 7th Triennial Conference of European Society for the Cognitive Sciences of Music."},{"key":"e_1_3_3_127_1","first-page":"2169","volume-title":"Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201906)","volume":"2","author":"Lazebnik S.","year":"2006","unstructured":"S. Lazebnik, C. Schmid, and J. Ponce. 2006. Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. In Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201906), Vol. 2. IEEE, 2169\u20132178."},{"key":"e_1_3_3_128_1","first-page":"801","volume-title":"Proceedings of the Proceedings of the Advances in Neural Information Processing Systems","author":"Lee Honglak","year":"2007","unstructured":"Honglak Lee, Alexis Battle, Rajat Raina, and Andrew Y. Ng. 2007. Efficient sparse coding algorithms. In Proceedings of the Proceedings of the Advances in Neural Information Processing Systems. 801\u2013808."},{"issue":"6","key":"e_1_3_3_129_1","doi-asserted-by":"crossref","first-page":"1406","DOI":"10.1109\/TASL.2009.2034776","article-title":"Audio-based semantic concept classification for consumer video","volume":"18","author":"Lee K.","year":"2010","unstructured":"K. Lee and D. P. Ellis. 2010. Audio-based semantic concept classification for consumer video. IEEE Transactions on Audio, Speech, and Language Processing 18, 6 (2010), 1406\u20131416.","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"key":"e_1_3_3_130_1","unstructured":"Kenton Lee Luheng He Mike Lewis and Luke Zettlemoyer. 2017. End-to-end neural coreference resolution. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201917 Copenhagen Denmark September 9-11 2017) . Association for Computational Linguistics 188\u2013197."},{"key":"e_1_3_3_131_1","doi-asserted-by":"publisher","DOI":"10.3233\/SW-140134"},{"key":"e_1_3_3_132_1","doi-asserted-by":"publisher","DOI":"10.1145\/79173.79176"},{"key":"e_1_3_3_133_1","doi-asserted-by":"crossref","first-page":"2548","DOI":"10.1109\/ICCV.2011.6126542","volume-title":"Proceedings of the 2011 International Conference on Computer Vision","author":"Leutenegger Stefan","year":"2011","unstructured":"Stefan Leutenegger, Margarita Chli, and Roland Y. Siegwart. 2011. BRISK: Binary robust invariant scalable keypoints. In Proceedings of the 2011 International Conference on Computer Vision. Ieee, 2548\u20132555."},{"key":"e_1_3_3_134_1","first-page":"81","volume-title":"Proceedings of the 3rd Annual Symposium on Document Analysis and Information Retrieval","volume":"33","author":"Lewis D. D.","year":"1994","unstructured":"D. D. Lewis and M. Ringuette. 1994. A comparison of two learning algorithms for text categorization. In Proceedings of the 3rd Annual Symposium on Document Analysis and Information Retrieval, Vol. 33. 81\u201393."},{"key":"e_1_3_3_135_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350922"},{"key":"e_1_3_3_136_1","doi-asserted-by":"crossref","unstructured":"Tian Li Xiang Chen Zhen Dong Kurt Keutzer and Shanghang Zhang. 2022. Domain-adaptive text classification with Structured Knowledge from Unlabeled Data. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence (IJCAI\u201922) . Vienna Austria ijcai.org 4216\u20134222.","DOI":"10.24963\/ijcai.2022\/585"},{"key":"e_1_3_3_137_1","first-page":"437","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Li W.","year":"2014","unstructured":"W. Li, L. Niu, and D. Xu. 2014. Exploiting privileged information from web data for image categorization. In Proceedings of the European Conference on Computer Vision. Springer, 437\u2013452."},{"key":"e_1_3_3_138_1","series-title":"Proceedings of the CICLing.","first-page":"417","volume":"3878","author":"Li W.","year":"2006","unstructured":"W. Li and M. Sun. 2006. Automatic image annotation based on WordNet and hierarchical ensembles. In Proceedings of the CICLing.Alexander F. Gelbukh (Ed.), Lecture Notes in Computer Science, Vol. 3878, Springer, 417\u2013428. Retrieved from http:\/\/dblp.uni-trier.de\/db\/conf\/cicling\/cicling2006.html#LiM06."},{"key":"e_1_3_3_139_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12559-017-9492-2"},{"key":"e_1_3_3_140_1","unstructured":"Yujia Li Daniel Tarlow Marc Brockschmidt and Richard S. Zemel. 2016. Gated graph sequence neural networks. In 4th International Conference on Learning Representations (ICLR\u201916 San Juan Puerto Rico May 2-4 2016 Conference Track Proceedings) ."},{"key":"e_1_3_3_141_1","doi-asserted-by":"publisher","DOI":"10.1093\/comjnl\/41.8.537"},{"key":"e_1_3_3_142_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:BTTJ.0000047600.45421.6d"},{"key":"e_1_3_3_143_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2021.101268"},{"key":"e_1_3_3_144_1","doi-asserted-by":"publisher","DOI":"10.5555\/2886521.2886657"},{"key":"e_1_3_3_145_1","first-page":"10012","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Liu Ze","year":"2021","unstructured":"Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 10012\u201310022."},{"key":"e_1_3_3_146_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000029664.99615.94"},{"issue":"5","key":"e_1_3_3_147_1","doi-asserted-by":"crossref","first-page":"823","DOI":"10.1080\/01431160600746456","article-title":"A survey of image classification methods and techniques for improving classification performance","volume":"28","author":"Lu D.","year":"2007","unstructured":"D. Lu and Q. Weng. 2007. A survey of image classification methods and techniques for improving classification performance. International Journal of Remote Sensing 28, 5 (2007), 823\u2013870.","journal-title":"International Journal of Remote Sensing"},{"key":"e_1_3_3_148_1","first-page":"745","volume-title":"Proceedings of the 2001 International Conference on Image Processing","author":"Luo J.","year":"2001","unstructured":"J. Luo and A. Savakis. 2001. Indoor vs outdoor classification of consumer photographs using low-level and semantic features. In Proceedings of the 2001 International Conference on Image Processing. IEEE, 745\u2013748."},{"key":"e_1_3_3_149_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.531803"},{"key":"e_1_3_3_150_1","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511809071.007"},{"key":"e_1_3_3_151_1","doi-asserted-by":"crossref","unstructured":"Kenneth Marino Ruslan Salakhutdinov and Abhinav Gupta. 2017. The more you know: Using knowledge graphs for image classification. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201917 Honolulu HI USA July 21-26 2017. IEEE Computer Society) . 20\u201328.","DOI":"10.1109\/CVPR.2017.10"},{"key":"e_1_3_3_152_1","first-page":"13","volume-title":"Proceedings of the 14th Conference Information Technologies-Applications and Theory","author":"Mar\u0161\u00edk Ladislav","year":"2014","unstructured":"Ladislav Mar\u0161\u00edk, J. Pokornyy, and Martin Ilc\u00edk. 2014. Improving music classification using harmonic complexity. In Proceedings of the 14th Conference Information Technologies-Applications and Theory. 13\u201317."},{"key":"e_1_3_3_153_1","first-page":"1","volume-title":"Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition","author":"Marszalek M.","year":"2007","unstructured":"M. Marszalek and C. Schmid. 2007. Semantic hierarchies for visual object recognition. In Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1\u20137."},{"key":"e_1_3_3_154_1","doi-asserted-by":"publisher","DOI":"10.3390\/app9040743"},{"key":"e_1_3_3_155_1","article-title":"Recurrent neural networks","volume":"5","author":"Medsker L. R.","year":"2001","unstructured":"L. R. Medsker and L. Jain. 2001. Recurrent neural networks. Design and Applications 5 (2001).","journal-title":"Design and Applications"},{"key":"e_1_3_3_156_1","doi-asserted-by":"publisher","DOI":"10.1016\/0031-3203(92)90121-X"},{"key":"e_1_3_3_157_1","doi-asserted-by":"crossref","unstructured":"Julia A. Meister Khuong An Nguyen and Zhiyuan Luo. 2022. Audio feature ranking for sound-based COVID-19 patient detection. In Progress in Artificial Intelligence - 21st EPIA Conference on Artificial Intelligence (EPIA\u201922 Lisbon Portugal August 31 - September 2 2022 Proceedings) (Lecture Notes in Computer Science) Vol. 13566. Springer 146\u201358.","DOI":"10.1007\/978-3-031-16474-3_13"},{"key":"e_1_3_3_158_1","first-page":"81","volume-title":"Proceedings of the 2019 IEEE 23rd International Conference on Computer Supported Cooperative Work in Design","author":"Menglong Cui","year":"2019","unstructured":"Cui Menglong, Ji Detao, Zeng Ting, Zhang Dehai, Xie Cheng, Chen Zhibo, and Xia Xiaoqiang. 2019. Image classification based on image knowledge graph and semantics. In Proceedings of the 2019 IEEE 23rd International Conference on Computer Supported Cooperative Work in Design. IEEE, 81\u201386."},{"key":"e_1_3_3_159_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2005.858055"},{"key":"e_1_3_3_160_1","unstructured":"P. Miettinen. 2009. Matrix decomposition methods for data mining: Computational complexity and algorithms. (2009)."},{"key":"e_1_3_3_161_1","first-page":"1792","volume-title":"Proceedings of the 10th IEEE International Conference on Computer Vision","author":"Mikolajczyk K.","year":"2005","unstructured":"K. Mikolajczyk, B. Leibe, and B. Schiele. 2005. Local features for object class recognition. In Proceedings of the 10th IEEE International Conference on Computer Vision. IEEE, 1792\u20131799."},{"key":"e_1_3_3_162_1","first-page":"3111","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Mikolov T.","year":"2013","unstructured":"T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. 2013. Distributed representations of words and phrases and their compositionality. In Proceedings of the Advances in Neural Information Processing Systems. Curran Associates, Inc., 3111\u20133119."},{"key":"e_1_3_3_163_1","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"e_1_3_3_164_1","doi-asserted-by":"publisher","DOI":"10.1145\/3439726"},{"key":"e_1_3_3_165_1","doi-asserted-by":"publisher","DOI":"10.5555\/3360093"},{"key":"e_1_3_3_166_1","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1016\/j.engappai.2019.08.025","article-title":"Poor and rich optimization algorithm: A new human-based and multi populations algorithm","volume":"86","author":"Moosavi Seyyed Hamid Samareh","year":"2019","unstructured":"Seyyed Hamid Samareh Moosavi and Vahid Khatibi Bardsiri. 2019. Poor and rich optimization algorithm: A new human-based and multi populations algorithm. Engineering Applications of Artificial Intelligence 86 (2019), 165\u2013181.","journal-title":"Engineering Applications of Artificial Intelligence"},{"key":"e_1_3_3_167_1","doi-asserted-by":"publisher","DOI":"10.3390\/app11135796"},{"key":"e_1_3_3_168_1","first-page":"216","volume-title":"Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics","author":"Navigli R.","year":"2010","unstructured":"R. Navigli and S. P. Ponzetto. 2010. BabelNet: Building a very large multilingual semantic network. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 216\u2013225."},{"issue":"1","key":"e_1_3_3_169_1","doi-asserted-by":"crossref","first-page":"27","DOI":"10.7763\/IJCCE.2014.V3.286","article-title":"Conceptual representation using wordnet for text categorization","volume":"3","author":"Nezreg H.","year":"2014","unstructured":"H. Nezreg, H. Lehbab, and H. Belbachir. 2014. Conceptual representation using wordnet for text categorization. International Journal of Computer and Communication Engineering 3, 1 (2014), 27.","journal-title":"International Journal of Computer and Communication Engineering"},{"issue":"1","key":"e_1_3_3_170_1","first-page":"1701","article-title":"Performance analysis of local binary pattern and k-nearest neighbor on image classification of fingers leaves","volume":"13","author":"Ningtyas A. D.","year":"2022","unstructured":"A. D. Ningtyas, E. B. Nababan, and S. Efendi. 2022. Performance analysis of local binary pattern and k-nearest neighbor on image classification of fingers leaves. International Journal of Nonlinear Analysis and Applications 13, 1 (2022), 1701\u20131708.","journal-title":"International Journal of Nonlinear Analysis and Applications"},{"key":"e_1_3_3_171_1","first-page":"8385","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Noh Hyeonwoo","year":"2019","unstructured":"Hyeonwoo Noh, Taehoon Kim, Jonghwan Mun, and Bohyung Han. 2019. Transfer learning via unsupervised task discovery for visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8385\u20138394."},{"key":"e_1_3_3_172_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2002.1017623"},{"key":"e_1_3_3_173_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1011139631724"},{"key":"e_1_3_3_174_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0079-6123(06)55002-2"},{"key":"e_1_3_3_175_1","doi-asserted-by":"publisher","DOI":"10.1038\/381607a0"},{"key":"e_1_3_3_176_1","doi-asserted-by":"publisher","DOI":"10.5555\/2907326.2907427"},{"key":"e_1_3_3_177_1","first-page":"11130","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Pan Lin","year":"2022","unstructured":"Lin Pan, Chung-Wei Hang, Avirup Sil, and Saloni Potdar. 2022. Improved text classification via contrastive adversarial training. In Proceedings of the AAAI Conference on Artificial Intelligence. 11130\u201311138."},{"issue":"10","key":"e_1_3_3_178_1","doi-asserted-by":"crossref","first-page":"711","DOI":"10.1016\/0262-8856(95)98753-G","article-title":"Image normalization for pattern recognition","volume":"13","author":"Pei S.","year":"1995","unstructured":"S. Pei and C. Lin. 1995. Image normalization for pattern recognition. Image and Vision computing 13, 10 (1995), 711\u2013723.","journal-title":"Image and Vision computing"},{"key":"e_1_3_3_179_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_3_180_1","first-page":"1","volume-title":"Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition","author":"Perronnin F.","year":"2007","unstructured":"F. Perronnin and C. Dance. 2007. Fisher kernels on visual vocabularies for image categorization. In Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1\u20138."},{"key":"e_1_3_3_181_1","doi-asserted-by":"crossref","unstructured":"Matthew E. Peters Mark Neumann Mohit Iyyer Matt Gardner Christopher Clark Kenton Lee and Luke Zettlemoyer. 2018. Deep Contextualized Word Representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT\u201918 New Orleans Louisiana USA June 1-6 2018) Volume 1 (Long Papers) Association for Computational Linguistics 2227\u20132237.","DOI":"10.18653\/v1\/N18-1202"},{"key":"e_1_3_3_182_1","doi-asserted-by":"crossref","unstructured":"Matthew E. Peters Mark Neumann Robert L. Logan IV Roy Schwartz Vidur Joshi Sameer Singh and Noah A. Smith. 2019. Knowledge Enhanced Contextual Word Representations. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP 2019 Hong Kong China November 3-7 2019) Association for Computational Linguistics 43\u201354.","DOI":"10.18653\/v1\/D19-1005"},{"key":"e_1_3_3_183_1","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324920000170"},{"key":"e_1_3_3_184_1","doi-asserted-by":"crossref","first-page":"102","DOI":"10.1007\/978-3-319-51811-4_9","volume-title":"Proceedings of the International Conference on Multimedia Modeling","author":"Pittaras N.","year":"2017","unstructured":"N. Pittaras, F. Markatopoulou, V. Mezaris, and I. Patras. 2017. Comparison of fine-tuning and extension strategies for deep convolutional neural networks. In Proceedings of the International Conference on Multimedia Modeling. Springer, 102\u2013114."},{"key":"e_1_3_3_185_1","first-page":"1","volume-title":"Proceedings of the 2019 International Conference on Computational Intelligence in Data Science","author":"Prasad S. Anuja","year":"2019","unstructured":"S. Anuja Prasad and Leena Mary. 2019. A comparative study of different features for vehicle classification. In Proceedings of the 2019 International Conference on Computational Intelligence in Data Science. IEEE, 1\u20135."},{"key":"e_1_3_3_186_1","first-page":"8th","volume-title":"Proceedings of the ISMIR","author":"Raimond Y.","year":"2007","unstructured":"Y. Raimond, S. A. Abdallah, M. B. Sandler, and F. Giasson. 2007. The music ontology. In Proceedings of the ISMIR. Citeseer, 8th."},{"key":"e_1_3_3_187_1","volume-title":"Proceedings of the Conf. Empir. Methods Nat. Lang. Process.","author":"Ratnaparkhi A.","year":"1996","unstructured":"A. Ratnaparkhi. 1996. A maximum entropy model for part-of-speech tagging. In Proceedings of the Conf. Empir. Methods Nat. Lang. Process."},{"key":"e_1_3_3_188_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-65981-7_12"},{"key":"e_1_3_3_189_1","doi-asserted-by":"crossref","unstructured":"Douglas A. Reynolds. 2009. Gaussian Mixture Models. In Encyclopedia of Biometrics . Springer 659\u2013663.","DOI":"10.1007\/978-0-387-73003-5_196"},{"key":"e_1_3_3_190_1","first-page":"859","volume-title":"Proceedings of the Pacific-Asia Conference on Knowledge Discovery and Data Mining","author":"Rho Seungmin","year":"2009","unstructured":"Seungmin Rho, Seheon Song, Eenjun Hwang, and Minkoo Kim. 2009. COMUS: Ontological and rule-based reasoning for music recommendation system. In Proceedings of the Pacific-Asia Conference on Knowledge Discovery and Data Mining. Springer, 859\u2013866."},{"key":"e_1_3_3_191_1","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1007\/978-3-642-20267-4_6","volume-title":"Proceedings of the International Conference on Adaptive and Natural Computing Algorithms","author":"Risojevi\u0107 Vladimir","year":"2011","unstructured":"Vladimir Risojevi\u0107, Snje\u017eana Momi\u0107, and Zdenka Babi\u0107. 2011. Gabor descriptors for aerial image classification. In Proceedings of the International Conference on Adaptive and Natural Computing Algorithms. Springer, 51\u201360."},{"key":"e_1_3_3_192_1","doi-asserted-by":"publisher","DOI":"10.1007\/0-387-25465-X_15"},{"key":"e_1_3_3_193_1","doi-asserted-by":"crossref","first-page":"2564","DOI":"10.1109\/ICCV.2011.6126544","volume-title":"Proceedings of the 2011 International Conference on Computer Vision","author":"Rublee E.","year":"2011","unstructured":"E. Rublee, V. Rabaud, K. Konolige, and G. Bradski. 2011. ORB: An efficient alternative to SIFT or SURF. In Proceedings of the 2011 International Conference on Computer Vision. IEEE, 2564\u20132571."},{"key":"e_1_3_3_194_1","doi-asserted-by":"crossref","unstructured":"D. Rumelhart G. Hinton and R. Williams. 1986. Learning representations by back-propagating errors. Nature 323 (1986) 533\u2013536.","DOI":"10.1038\/323533a0"},{"key":"e_1_3_3_195_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_3_196_1","doi-asserted-by":"publisher","DOI":"10.1037\/h0077714"},{"key":"e_1_3_3_197_1","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(88)90021-0"},{"key":"e_1_3_3_198_1","doi-asserted-by":"publisher","DOI":"10.1145\/361219.361220"},{"key":"e_1_3_3_199_1","unstructured":"H. Schmid. 1994. TreeTagger-a language independent part-of-speech tagger. (1994). Retrieved from http:\/\/www.ims.uni-stuttgart.de\/projekte\/corplex\/TreeTagger\/."},{"key":"e_1_3_3_200_1","doi-asserted-by":"publisher","DOI":"10.1145\/505282.505283"},{"key":"e_1_3_3_201_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.dsp.2007.12.004"},{"key":"e_1_3_3_202_1","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1117\/12.421083","volume-title":"Proceedings of the Data Mining and Knowledge Discovery: Theory, Tools, and Technology III","volume":"4384","author":"Sethi I. K.","year":"2001","unstructured":"I. K. Sethi, I. L. Coman, and D. Stan. 2001. Mining association rules between low-level image features and high-level concepts. In Proceedings of the Data Mining and Knowledge Discovery: Theory, Tools, and Technology III, Vol. 4384. International Society for Optics and Photonics, 279\u2013290."},{"key":"e_1_3_3_203_1","doi-asserted-by":"crossref","unstructured":"Weijia Shi Muhao Chen Pei Zhou and Kai-Wei Chang. 2019. Retrofitting Contextualized Word Embeddings with Paraphrases. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP- IJCNLP\u201919 Hong Kong China November 3-7 2019) Association for Computational Linguistics 1198\u20131203.","DOI":"10.18653\/v1\/D19-1113"},{"key":"e_1_3_3_204_1","doi-asserted-by":"crossref","unstructured":"C. Shorten and T. M. Khoshgoftaar. 2019. A survey on image data augmentation for deep learning. 6 (2019). DOI:10.1186\/s40537-019-0197-0","DOI":"10.1186\/s40537-019-0197-0"},{"key":"e_1_3_3_205_1","doi-asserted-by":"crossref","unstructured":"Leslie F. Sikos. 2017. The Semantic Gap. Description Logics in Multimedia Reasoning (2017) 51\u201366.","DOI":"10.1007\/978-3-319-54066-5_3"},{"key":"e_1_3_3_206_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-010-0175-9"},{"key":"e_1_3_3_207_1","doi-asserted-by":"crossref","first-page":"1661","DOI":"10.1109\/IJCNN.2003.1223656","volume-title":"Proceedings of the Neural Networks, 2003. Proc. Int. Jt. Conf.","volume":"3","author":"Silva C.","year":"2003","unstructured":"C. Silva and B. Ribeiro. 2003. The importance of stop word removal on recall values in text categorization. In Proceedings of the Neural Networks, 2003. Proc. Int. Jt. Conf., Vol. 3. IEEE, 1661\u20131666."},{"key":"e_1_3_3_208_1","doi-asserted-by":"crossref","unstructured":"Mattia Silvestri Michele Lombardi and Michela Milano. 2021. Injecting domain knowledge in neural networks: a controlled experiment on a constrained problem. In Integration of Constraint Programming Artificial Intelligence and Operations Research: 18th International Conference (CPAIOR\u201921 Vienna Austria July 5\u20138 2021 Proceedings 18) . Springer 266\u2013282.","DOI":"10.1007\/978-3-030-78230-6_17"},{"key":"e_1_3_3_209_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In 3rd International Conference on Learning Representations (ICLR\u201915. San Diego CA USA May 7-9 2015) Conference Track Proceedings."},{"issue":"6","key":"e_1_3_3_210_1","first-page":"545","article-title":"Optical character recognition techniques: A survey","volume":"4","author":"Singh S.","year":"2013","unstructured":"S. Singh. 2013. Optical character recognition techniques: A survey. Journal of Emerging Trends in Computing and Information Sciences 4, 6 (2013), 545\u2013550.","journal-title":"Journal of Emerging Trends in Computing and Information Sciences"},{"key":"e_1_3_3_211_1","first-page":"101104","article-title":"tax2vec: Constructing interpretable features from taxonomies for short text classification","author":"\u0160krlj B.","year":"2020","unstructured":"B. \u0160krlj, M. Martinc, J. Kralj, N. Lavra\u010d, and S. Pollak. 2020. tax2vec: Constructing interpretable features from taxonomies for short text classification. Computer Speech and Language 65 (2020), 101104.","journal-title":"Computer Speech and Language"},{"key":"e_1_3_3_212_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-021-05968-x"},{"issue":"1998","key":"e_1_3_3_213_1","article-title":"Auditory toolbox","volume":"10","author":"Slaney M.","year":"1998","unstructured":"M. Slaney. 1998. Auditory toolbox. Interval Research Corporation, Tech. Rep. 10, 1998 (1998).","journal-title":"Interval Research Corporation, Tech. Rep."},{"key":"e_1_3_3_214_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2009.03.002"},{"key":"e_1_3_3_215_1","doi-asserted-by":"publisher","DOI":"10.5120\/16899-6972"},{"key":"e_1_3_3_216_1","doi-asserted-by":"crossref","unstructured":"B. J. Sowmya Chetan and K. G. Srinivasa. 2016. Large scale multi-label text classification of a hierarchical dataset using Rocchio algorithm. In Proceedings of the 2016 International Conference on Computation System and Information Technology for Sustainable Solutions . IEEE 291\u2013296.","DOI":"10.1109\/CSITSS.2016.7779373"},{"key":"e_1_3_3_217_1","doi-asserted-by":"crossref","first-page":"552","DOI":"10.1145\/1076034.1076128","volume-title":"Proceedings of the SIGIR.","author":"Srikanth M.","year":"2005","unstructured":"M. Srikanth, J. Varner, M. Bowden, and D. I. Moldovan. 2005. Exploiting ontologies for automatic image annotation. In Proceedings of the SIGIR.Ricardo A. Baeza-Yates, Nivio Ziviani, Gary Marchionini, Alistair Moffat, and John Tait (Eds.), ACM, 552\u2013558. Retrieved from http:\/\/dblp.uni-trier.de\/db\/conf\/sigir\/sigir2005.html#SrikanthVBM05."},{"key":"e_1_3_3_218_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-018-6793-8"},{"key":"e_1_3_3_219_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2670313"},{"key":"e_1_3_3_220_1","first-page":"1","volume-title":"Proceedings of the Feature Extraction: Modern Questions and Challenges","author":"Storcheus D.","year":"2015","unstructured":"D. Storcheus, A. Rostamizadeh, and S. Kumar. 2015. A survey of modern questions and challenges in feature extraction. In Proceedings of the Feature Extraction: Modern Questions and Challenges. PMLR, 1\u201318."},{"key":"e_1_3_3_221_1","unstructured":"Carlo Strapparava and Alessandro Valitutti. 2004. Wordnet affect: An affective extension of wordnet. In Proceedings of the Lrec. Vol. 4. Lisbon 40."},{"key":"e_1_3_3_222_1","doi-asserted-by":"crossref","unstructured":"Emma Strubell Ananya Ganesh and Andrew McCallum. 2019. Energy and Policy Considerations for Deep Learning in NLP. In Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL\u201919 Florence Italy July 28- August 2 2019) Volume 1: Long Papers Association for Computational Linguistics 3645\u20133650.","DOI":"10.18653\/v1\/P19-1355"},{"key":"e_1_3_3_223_1","first-page":"29","volume-title":"Proceedings of the International Workshop on Adaptive Multimedia Retrieval","author":"Sturm B. L.","year":"2012","unstructured":"B. L. Sturm. 2012. A survey of evaluation in music genre recognition. In Proceedings of the International Workshop on Adaptive Multimedia Retrieval. Springer, 29\u201366."},{"issue":"12","key":"e_1_3_3_224_1","first-page":"1869","article-title":"Overview of textual anti-spam filtering techniques","volume":"5","author":"Subramaniam T.","year":"2010","unstructured":"T. Subramaniam, H. A. Jalab, and A. Y. Taqa. 2010. Overview of textual anti-spam filtering techniques. International Journal of Physical Sciences 5, 12 (2010), 1869\u20131882.","journal-title":"International Journal of Physical Sciences"},{"key":"e_1_3_3_225_1","doi-asserted-by":"publisher","DOI":"10.1145\/1242572.1242667"},{"key":"e_1_3_3_226_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-32381-3_16"},{"key":"e_1_3_3_227_1","first-page":"321","volume-title":"Proceedings of the ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Sun Y.","year":"2020","unstructured":"Y. Sun and S. Ghaffarzadegan. 2020. An ontology-aware framework for audio event classification. In Proceedings of the ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 321\u2013325."},{"key":"e_1_3_3_228_1","unstructured":"Yu Sun Shuohuan Wang Yu-Kun Li Shikun Feng Xuyi Chen Han Zhang Xin Tian Danxiang Zhu Hao Tian and Hua Wu. 2019. ERNIE: Enhanced representation through knowledge integration. CoRR abs\/1904.09223 (2019)."},{"key":"e_1_3_3_229_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_3_230_1","doi-asserted-by":"crossref","unstructured":"Yi Tay Mostafa Dehghani Dara Bahri and Donald Metzler. 2023. Efficient Transformers: A Survey. ACM Comput. Surv. 55 6 (2023) 109:1\u2013109:28.","DOI":"10.1145\/3530811"},{"issue":"11","key":"e_1_3_3_231_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.5815\/ijitcs.2014.11.01","article-title":"An overview of automatic audio segmentation","volume":"6","author":"Theodorou T.","year":"2014","unstructured":"T. Theodorou, I. Mporas, and N. Fakotakis. 2014. An overview of automatic audio segmentation. International Journal of Information Technology and Computer Science 6, 11 (2014), 1.","journal-title":"International Journal of Information Technology and Computer Science"},{"key":"e_1_3_3_232_1","doi-asserted-by":"crossref","unstructured":"Thirumoorthy Karpagalingam and Muneeswaran Karuppiah. 2021. Feature selection using hybrid poor and rich optimization algorithm for text classification. Pattern Recognit. Lett. 147 (2021) 63\u201370.","DOI":"10.1016\/j.patrec.2021.03.034"},{"key":"e_1_3_3_233_1","doi-asserted-by":"crossref","unstructured":"Hugo Touvron Piotr Bojanowski Mathilde Caron Matthieu Cord Alaaeldin El-Nouby Edouard Grave Gautier Izacard Armand Joulin Gabriel Synnaeve Jakob Verbeek et\u00a0al. 2022. ResMLP: Feedforward Networks for Image Classification With Data-Efficient Training. IEEE Transactions on Pattern Analysis Machine Intelligence 01 (2022) 1\u20139.","DOI":"10.1109\/TPAMI.2022.3206148"},{"key":"e_1_3_3_234_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898719574"},{"key":"e_1_3_3_235_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.proeng.2014.03.129"},{"key":"e_1_3_3_236_1","doi-asserted-by":"crossref","unstructured":"P. D. Turney and P. Pantel. 2010. From frequency to meaning : Vector space models of semantics. 37 (2010) 141\u2013188.","DOI":"10.1613\/jair.2934"},{"key":"e_1_3_3_237_1","doi-asserted-by":"publisher","DOI":"10.1561\/0600000017"},{"key":"e_1_3_3_238_1","first-page":"955","volume-title":"Proceedings of the PICMET\u201908-2008 Portland International Conference on Management of Engineering & Technology","author":"Uys J.","year":"2008","unstructured":"J. Uys, N. Du Preez, and E. Uys. 2008. Leveraging unstructured information using topic modelling. In Proceedings of the PICMET\u201908-2008 Portland International Conference on Management of Engineering & Technology. IEEE, 955\u2013961."},{"key":"e_1_3_3_239_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2012.2199972"},{"key":"e_1_3_3_240_1","doi-asserted-by":"publisher","DOI":"10.1198\/10618600152418584"},{"key":"e_1_3_3_241_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2009.06.042"},{"key":"e_1_3_3_242_1","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017), 5998\u20136008.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"1","key":"e_1_3_3_243_1","first-page":"7","article-title":"Preprocessing techniques for text mining-an overview","volume":"5","author":"Vijayarani S.","year":"2015","unstructured":"S. Vijayarani, M. J. Ilamathi, and M. Nithya. 2015. Preprocessing techniques for text mining-an overview. International Journal of Computer Science & Communication Networks 5, 1 (2015), 7\u201316.","journal-title":"International Journal of Computer Science & Communication Networks"},{"key":"e_1_3_3_244_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-006-8614-1"},{"key":"e_1_3_3_245_1","doi-asserted-by":"publisher","DOI":"10.1145\/2629489"},{"key":"e_1_3_3_246_1","first-page":"3360","volume-title":"Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Wang Jinjun","year":"2010","unstructured":"Jinjun Wang, Jianchao Yang, Kai Yu, Fengjun Lv, Thomas Huang, and Yihong Gong. 2010. Locality-constrained linear coding for image classification. In Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. IEEE, 3360\u20133367."},{"key":"e_1_3_3_247_1","unstructured":"Luyu Wang and Aaron van den Oord. 2021. Multi-format contrastive learning of audio representations. CoRRabs\/2103.06508 (2021)."},{"key":"e_1_3_3_248_1","first-page":"5670","volume-title":"Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Wang Yuxuan","year":"2017","unstructured":"Yuxuan Wang, Pascal Getreuer, Thad Hughes, Richard F. Lyon, and Rif A. Saurous. 2017. Trainable frontend for robust and far-field keyword spotting. In Proceedings of the 2017 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 5670\u20135674."},{"issue":"4","key":"e_1_3_3_249_1","doi-asserted-by":"crossref","first-page":"1191","DOI":"10.3758\/s13428-012-0314-x","article-title":"Norms of valence, arousal, and dominance for 13,915 English lemmas","volume":"45","author":"Warriner A. B.","year":"2013","unstructured":"A. B. Warriner, V. Kuperman, and M. Brysbaert. 2013. Norms of valence, arousal, and dominance for 13,915 English lemmas. Behavior Research Methods 45, 4 (2013), 1191\u20131207.","journal-title":"Behavior Research Methods"},{"key":"e_1_3_3_250_1","unstructured":"C. Whittaker B. Ryner and M. Nazif. 2010. Large-scale automatic classification of phishing pages. In Proceedings of the Network and Distributed System Security Symposium (NDSS\u201910 San Diego California USA 28th February - 3rd March 2010) The Internet Society."},{"key":"e_1_3_3_251_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2010.2045169"},{"key":"e_1_3_3_252_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213891"},{"issue":"1","key":"e_1_3_3_253_1","doi-asserted-by":"crossref","first-page":"24","DOI":"10.1007\/s11633-022-1320-9","article-title":"Weakly correlated knowledge integration for few-shot image classification","volume":"19","author":"Yang Chun","year":"2022","unstructured":"Chun Yang, Chang Liu, and Xu-Cheng Yin. 2022. Weakly correlated knowledge integration for few-shot image classification. Machine Intelligence Research 19, 1 (2022), 24\u201337.","journal-title":"Machine Intelligence Research"},{"key":"e_1_3_3_254_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-017-9319-5"},{"key":"e_1_3_3_255_1","first-page":"205","volume-title":"Proceedings of the International Conference on Signal and Information Processing, Networking and Computers","author":"Ye Jingyi","year":"2017","unstructured":"Jingyi Ye, Xiaojun Jing, and Jia Li. 2017. Sentiment analysis using modified LDA. In Proceedings of the International Conference on Signal and Information Processing, Networking and Computers. Springer, 205\u2013212."},{"key":"e_1_3_3_256_1","unstructured":"Jason Yosinski Jeff Clune Anh Nguyen Thomas Fuchs and Hod Lipson. 2015. Understanding neural networks through deep visualization. CoRR abs\/1506.06579 (2015)."},{"key":"e_1_3_3_257_1","first-page":"467","volume-title":"Proceedings of the International Conference on Computer Vision Theory and Applications","volume":"2","author":"Younes L.","year":"2012","unstructured":"L. Younes, B. Romaniuk, and E. Bittar. 2012. A comprehensive and comparative survey of the SIFT algorithm-feature detection, description, and characterization. In Proceedings of the International Conference on Computer Vision Theory and Applications, Vol. 2. SCITEPRESS, 467\u2013474."},{"key":"e_1_3_3_258_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1056"},{"key":"e_1_3_3_259_1","article-title":"Optimized audio classification and segmentation algorithm by using ensemble methods","volume":"2015","author":"Zahid Saadia","year":"2015","unstructured":"Saadia Zahid, Fawad Hussain, Muhammad Rashid, Muhammad Haroon Yousaf, and Hafiz Adnan Habib. 2015. Optimized audio classification and segmentation algorithm by using ensemble methods. Mathematical Problems in Engineering 2015 (2015), 209814\u2013209825.","journal-title":"Mathematical Problems in Engineering"},{"key":"e_1_3_3_260_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2015.09.027"},{"issue":"2","key":"e_1_3_3_261_1","doi-asserted-by":"crossref","first-page":"60","DOI":"10.5815\/ijieeb.2015.02.08","article-title":"Feature extraction or feature selection for text classification: A case study on phishing email detection","volume":"7","author":"Zareapoor Masoumeh","year":"2015","unstructured":"Masoumeh Zareapoor and K. R. Seeja. 2015. Feature extraction or feature selection for text classification: A case study on phishing email detection. International Journal of Information Engineering and Electronic Business 7, 2 (2015), 60.","journal-title":"International Journal of Information Engineering and Electronic Business"},{"key":"e_1_3_3_262_1","unstructured":"Neil Zeghidour Olivier Teboul F\u00e9lix de Chaumont Quitry and Marco Tagliasacchi. 2021. LEAF: A learnable frontend for audio classiffication. In 9th International Conference on Learning Representations (ICLR\u201921) . Virtual Event Austria OpenReview.net."},{"key":"e_1_3_3_263_1","first-page":"818","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zeiler M. D.","year":"2014","unstructured":"M. D. Zeiler and R. Fergus. 2014. Visualizing and understanding convolutional networks. In Proceedings of the European Conference on Computer Vision. Springer, 818\u2013833."},{"key":"e_1_3_3_264_1","doi-asserted-by":"crossref","first-page":"57678","DOI":"10.1109\/ACCESS.2019.2912627","article-title":"Knowledge graph-based image classification refinement","volume":"7","author":"Zhang D.","year":"2019","unstructured":"D. Zhang, M. Cui, Y. Yang, P. Yang, C. Xie, D. Liu, B. Yu, and Z. Chen. 2019a. Knowledge graph-based image classification refinement. IEEE Access 7 (2019), 57678\u201357690.","journal-title":"IEEE Access"},{"key":"e_1_3_3_265_1","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380107"},{"issue":"2","key":"e_1_3_3_266_1","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1007\/s11263-006-9794-4","article-title":"Local features and kernels for classification of texture and object categories: A comprehensive study","volume":"73","author":"Zhang J.","year":"2007","unstructured":"J. Zhang, M. Marsza\u0142ek, S. Lazebnik, and C. Schmid. 2007. Local features and kernels for classification of texture and object categories: A comprehensive study. International Journal of Computer Vision 73, 2 (2007), 213\u2013238.","journal-title":"International Journal of Computer Vision"},{"key":"e_1_3_3_267_1","doi-asserted-by":"publisher","DOI":"10.1631\/FITEE.1700808"},{"key":"e_1_3_3_268_1","doi-asserted-by":"publisher","DOI":"10.1186\/s40649-019-0069-y"},{"key":"e_1_3_3_269_1","doi-asserted-by":"crossref","first-page":"398","DOI":"10.1117\/12.325832","volume-title":"Proceedings of the Multimedia Storage and Archiving Systems III","volume":"3527","author":"Zhang T.","year":"1998","unstructured":"T. Zhang and C. J. Kuo. 1998. Hierarchical system for content-based audio classification and retrieval. In Proceedings of the Multimedia Storage and Archiving Systems III, Vol. 3527. International Society for Optics and Photonics, 398\u2013409."},{"key":"e_1_3_3_270_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2010.08.066"},{"key":"e_1_3_3_271_1","doi-asserted-by":"crossref","first-page":"710","DOI":"10.1109\/FSKD.2015.7382029","volume-title":"Proceedings of the 2015 12th International Conference on Fuzzy Systems and Knowledge Discovery","author":"Zhang Xinwei","year":"2015","unstructured":"Xinwei Zhang and Bin Wu. 2015. Short text classification based on feature extension using the n-gram model. In Proceedings of the 2015 12th International Conference on Fuzzy Systems and Knowledge Discovery. IEEE, 710\u2013716."},{"key":"e_1_3_3_272_1","doi-asserted-by":"crossref","unstructured":"Zhengyan Zhang Xu Han Zhiyuan Liu Xin Jiang Maosong Sun and Qun Liu. 2019. ERNIE: Enhanced language representation with informative entities. In Proceedings of the 57th Conference of the Association for Computational Linguistics (ACL\u201919) 1 (2019) 1441\u20131451.","DOI":"10.18653\/v1\/P19-1139"},{"key":"e_1_3_3_273_1","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482097"},{"key":"e_1_3_3_274_1","first-page":"321","volume-title":"Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Zharmagambetov Arman","year":"2022","unstructured":"Arman Zharmagambetov, Qingming Tang, Chieh-Chi Kao, Qin Zhang, Ming Sun, Viktor Rozgic, Jasha Droppo, and Chao Wang. 2022. Improved representation learning for acoustic event classification using tree-structured ontology. In Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 321\u2013325."},{"key":"e_1_3_3_275_1","volume-title":"Feature Engineering for Machine Learning: Principles and Techniques for Data Scientists","author":"Zheng A.","year":"2018","unstructured":"A. Zheng and A. Casari. 2018. Feature Engineering for Machine Learning: Principles and Techniques for Data Scientists. \u201cO\u2019Reilly Media, Inc.\u201d"},{"issue":"9","key":"e_1_3_3_276_1","doi-asserted-by":"crossref","first-page":"1779","DOI":"10.3390\/app9091779","article-title":"SURF-BRISK\u2013based image infilling method for terrain classification of a legged robot","volume":"9","author":"Zhu Yaguang","year":"2019","unstructured":"Yaguang Zhu, Chaoyu Jia, Chao Ma, and Qiong Liu. 2019. SURF-BRISK\u2013based image infilling method for terrain classification of a legged robot. Applied Sciences 9, 9 (2019), 1779.","journal-title":"Applied Sciences"},{"key":"e_1_3_3_277_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2020.3004555"},{"key":"e_1_3_3_278_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2016.02.021"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3583682","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3583682","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:27Z","timestamp":1750178787000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3583682"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,17]]},"references-count":277,"journal-issue":{"issue":"14s","published-print":{"date-parts":[[2023,12,31]]}},"alternative-id":["10.1145\/3583682"],"URL":"https:\/\/doi.org\/10.1145\/3583682","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,17]]},"assertion":[{"value":"2021-05-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-01-23","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}