{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,4]],"date-time":"2026-04-04T06:13:17Z","timestamp":1775283197348,"version":"3.50.1"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2012,8,1]],"date-time":"2012-08-01T00:00:00Z","timestamp":1343779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Czech Ministry of Education","award":["MSM0021630528"],"award-info":[{"award-number":["MSM0021630528"]}]},{"DOI":"10.13039\/501100001824","name":"Czech Science Foundation","doi-asserted-by":"publisher","award":["GP202\/12\/P567"],"award-info":[{"award-number":["GP202\/12\/P567"]}],"id":[{"id":"10.13039\/501100001824","id-type":"DOI","asserted-by":"publisher"}]},{"name":"IT4 Innovations Centre of Excellence","award":["CZ.1.05\/1.1.00\/02.0070"],"award-info":[{"award-number":["CZ.1.05\/1.1.00\/02.0070"]}]},{"DOI":"10.13039\/501100002969","name":"Technologick\u00e1 Agentura Cesk\u00e9 Republiky","doi-asserted-by":"publisher","award":["TA01011328"],"award-info":[{"award-number":["TA01011328"]}],"id":[{"id":"10.13039\/501100002969","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Czech Ministry of Trade and Commerce","award":["FR-TI1\/034"],"award-info":[{"award-number":["FR-TI1\/034"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2012,8]]},"abstract":"<jats:p>This article investigates query-by-example (QbE) spoken term detection (STD), in which the query is not entered as text, but selected in speech data or spoken. Two feature extractors based on neural networks (NN) are introduced: the first producing phone-state posteriors and the second making use of a compressive NN layer. They are combined with three different QbE detectors: while the Gaussian mixture model\/hidden Markov model (GMM\/HMM) and dynamic time warping (DTW) both work on continuous feature vectors, the third one, based on weighted finite-state transducers (WFST), processes phone lattices. QbE STD is compared to two standard STD systems with text queries: acoustic keyword spotting and WFST-based search of phone strings in phone lattices. The results are reported on four languages (Czech, English, Hungarian, and Levantine Arabic) using standard metrics: equal error rate (EER) and two versions of popular figure-of-merit (FOM). Language-dependent and language-independent cases are investigated; the latter being particularly interesting for scenarios lacking standard resources to train speech recognition systems. While the DTW and GMM\/HMM approaches produce the best results for a language-dependent setup depending on the target language, the GMM\/HMM approach performs the best dealing with a language-independent setup. As far as WFSTs are concerned, they are promising as they allow for indexing and fast search.<\/jats:p>","DOI":"10.1145\/2328967.2328971","type":"journal-article","created":{"date-parts":[[2012,9,11]],"date-time":"2012-09-11T22:21:06Z","timestamp":1347402066000},"page":"1-34","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["Comparison of methods for language-dependent and language-independent query-by-example spoken term detection"],"prefix":"10.1145","volume":"30","author":[{"given":"Javier","family":"Tejedor","sequence":"first","affiliation":[{"name":"Universidad Aut\u00b4onoma de Madrid, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michal","family":"Fap\u0161o","sequence":"additional","affiliation":[{"name":"Brno University of Technology, Czech Republic"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Igor","family":"Sz\u00f6ke","sequence":"additional","affiliation":[{"name":"Brno University of Technology, Czech Republic"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jan \u201cHonza\u201d","family":"\u010cernock\u00fd","sequence":"additional","affiliation":[{"name":"Brno University of Technology, Czech Republic"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Franti\u0161ek","family":"Gr\u00e9zl","sequence":"additional","affiliation":[{"name":"Brno University of Technology, Czech Republic"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,9,6]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"Proceedings of the ICASSP'08","author":"Akbacak M.","unstructured":"Akbacak , M. , Vergyri , D. , and Stolcke , A . 2008. Open-vocabulary spoken term detection using graphone-based hybrid recognition systems . In Proceedings of the ICASSP'08 . 5240--5243. Akbacak, M., Vergyri, D., and Stolcke, A. 2008. Open-vocabulary spoken term detection using graphone-based hybrid recognition systems. In Proceedings of the ICASSP'08. 5240--5243."},{"key":"e_1_2_2_2_1","volume-title":"Proceedings of the Workshop on Interdisciplinary Approaches to Speech Indexing and Retrieval (HLT-NAACL'04)","author":"Allauzen C.","unstructured":"Allauzen , C. , Mohri , M. , and Saraclar , M . 2004. General indexation of weighted automata: application to spoken utterance retrieval . In Proceedings of the Workshop on Interdisciplinary Approaches to Speech Indexing and Retrieval (HLT-NAACL'04) . 33--40. Allauzen, C., Mohri, M., and Saraclar, M. 2004. General indexation of weighted automata: application to spoken utterance retrieval. In Proceedings of the Workshop on Interdisciplinary Approaches to Speech Indexing and Retrieval (HLT-NAACL'04). 33--40."},{"key":"e_1_2_2_3_1","volume-title":"Proceedings of the International Conference on Implementation and Application of Automata.","volume":"4783","author":"Allauzen C.","unstructured":"Allauzen , C. , Riley , M. , Schalkwyk , J. , Skut , W. , and Mohri , M . 2007. OpenFst: A general and efficient weighted finite-state transducer library . In Proceedings of the International Conference on Implementation and Application of Automata. Vol. 4783 , 11--23. Allauzen, C., Riley, M., Schalkwyk, J., Skut, W., and Mohri, M. 2007. OpenFst: A general and efficient weighted finite-state transducer library. In Proceedings of the International Conference on Implementation and Application of Automata. Vol. 4783, 11--23."},{"key":"e_1_2_2_4_1","volume-title":"Proceedings of MediaEval'11","author":"Anguera X.","year":"2011","unstructured":"Anguera , X. 2011 . Telefonica system for the spoken web search task at MediaEval 2011 . In Proceedings of MediaEval'11 . 3--4. Anguera, X. 2011. Telefonica system for the spoken web search task at MediaEval 2011. In Proceedings of MediaEval'11. 3--4."},{"key":"e_1_2_2_5_1","volume-title":"Proceedings of ICASSP'10","author":"Anguera X.","unstructured":"Anguera , X. , Macrae , R. , and Oliver , N . 2010. Partial sequence matching using an unbounded dynamic time warping algorithm . In Proceedings of ICASSP'10 . 3582--3585. Anguera, X., Macrae, R., and Oliver, N. 2010. Partial sequence matching using an unbounded dynamic time warping algorithm. In Proceedings of ICASSP'10. 3582--3585."},{"key":"e_1_2_2_6_1","volume-title":"Proceedings of MediaEval'11","author":"Barnard E.","unstructured":"Barnard , E. , Davel , M. , Van Heerden , C. , Kleynhans , N. , and Bali , K . 2011. Phone recognition for spoken web search . In Proceedings of MediaEval'11 . 5--6. Barnard, E., Davel, M., Van Heerden, C., Kleynhans, N., and Bali, K. 2011. Phone recognition for spoken web search. In Proceedings of MediaEval'11. 5--6."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177697196"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2000.859138"},{"key":"e_1_2_2_9_1","volume-title":"Proceedings of the Congress on Evolutionary Computation. 829--835","author":"Cai L.","unstructured":"Cai , L. , Juedes , D. , and Liakhovitch , E . 2000. Evolutionary computation techniques for multiple sequence alignment . In Proceedings of the Congress on Evolutionary Computation. 829--835 . Cai, L., Juedes, D., and Liakhovitch, E. 2000. Evolutionary computation techniques for multiple sequence alignment. In Proceedings of the Congress on Evolutionary Computation. 829--835."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2009.4960494"},{"key":"e_1_2_2_11_1","volume-title":"Proceedings of the Interspeech'10","author":"Chan C.","unstructured":"Chan , C. and Lee , L . 2010. Unsupervised spoken-term detection with spoken queries using segment-based dynamic time warping . In Proceedings of the Interspeech'10 . 693--696. Chan, C. and Lee, L. 2010. Unsupervised spoken-term detection with spoken queries using segment-based dynamic time warping. In Proceedings of the Interspeech'10. 693--696."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390334.1390397"},{"key":"e_1_2_2_13_1","volume-title":"Proceedings of Mexican International Conference on Artificial Intelligence. 156--165","author":"Cuayahuitl H.","unstructured":"Cuayahuitl , H. and Serridge , B . 2002. Out-of-vocabulary word modeling and rejection for spanish keyword spotting systems . In Proceedings of Mexican International Conference on Artificial Intelligence. 156--165 . Cuayahuitl, H. and Serridge, B. 2002. Out-of-vocabulary word modeling and rejection for spanish keyword spotting systems. In Proceedings of Mexican International Conference on Artificial Intelligence. 156--165."},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.sbi.2006.04.004"},{"key":"e_1_2_2_15_1","volume-title":"Proceedings of the Workshop on Searching Spontaneous Conversational Speech (SIGIR-SSCS'07)","author":"Fiscus J. G.","unstructured":"Fiscus , J. G. , Ajot , J. , Garofolo , J. S. , and Doddington , G . 2007. Results of the 2006 spoken term detection evaluation . In Proceedings of the Workshop on Searching Spontaneous Conversational Speech (SIGIR-SSCS'07) . Fiscus, J. G., Ajot, J., Garofolo, J. S., and Doddington, G. 2007. Results of the 2006 spoken term detection evaluation. In Proceedings of the Workshop on Searching Spontaneous Conversational Speech (SIGIR-SSCS'07)."},{"key":"e_1_2_2_16_1","volume-title":"Proceedings of the ICASSP'08","author":"Grezl F.","unstructured":"Grezl , F. and Fousek , P . 2008. Optimizing bottle-neck features for LVCSR . In Proceedings of the ICASSP'08 . 4729--4732. Grezl, F. and Fousek, P. 2008. Optimizing bottle-neck features for LVCSR. In Proceedings of the ICASSP'08. 4729--4732."},{"key":"e_1_2_2_17_1","volume-title":"Proceedings of Interspeech'09","author":"Grezl F.","unstructured":"Grezl , F. , Karafiat , M. , and Burget , L . 2009. Investigation into bottle-neck features for meeting speech recognition . In Proceedings of Interspeech'09 , 2947--2950. Grezl, F., Karafiat, M., and Burget, L. 2009. Investigation into bottle-neck features for meeting speech recognition. In Proceedings of Interspeech'09, 2947--2950."},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2007.367023"},{"key":"e_1_2_2_19_1","volume-title":"Proceedings of the ICASSP'01","author":"Hazen T.","unstructured":"Hazen , T. and Bazzi , I . 2001. A comparison and combination of methods for OOV word detection and word confidence scoring . In Proceedings of the ICASSP'01 . 397--400. Hazen, T. and Bazzi, I. 2001. A comparison and combination of methods for OOV word detection and word confidence scoring. In Proceedings of the ICASSP'01. 397--400."},{"key":"e_1_2_2_20_1","volume-title":"Proceedings of ASRU'09","author":"Hazen T. J.","unstructured":"Hazen , T. J. , Shen , W. , and White , C. M . 2009. Query-by-example spoken term detection using phonetic posteriorgram templates . In Proceedings of ASRU'09 . 421--426. Hazen, T. J., Shen, W., and White, C. M. 2009. Query-by-example spoken term detection using phonetic posteriorgram templates. In Proceedings of ASRU'09. 421--426."},{"key":"e_1_2_2_21_1","volume-title":"Proceedings of the ICASSP'07","author":"Helen M.","unstructured":"Helen , M. and Virtanen , T . 2007. Query by example of audio signals using Euclidean distance between Gaussian mixture models . In Proceedings of the ICASSP'07 . 225--228. Helen, M. and Virtanen, T. 2007. Query by example of audio signals using Euclidean distance between Gaussian mixture models. In Proceedings of the ICASSP'07. 225--228."},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1155\/2010\/179303"},{"key":"e_1_2_2_23_1","volume-title":"Proceedings of Interspeech'10","author":"Jansen A.","unstructured":"Jansen , A. , Church , K. , and Hermansky , H . 2010. Towards spoken term discovery at scale with zero resources . In Proceedings of Interspeech'10 , 1676--1679. Jansen, A., Church, K., and Hermansky, H. 2010. Towards spoken term discovery at scale with zero resources. In Proceedings of Interspeech'10, 1676--1679."},{"key":"e_1_2_2_24_1","volume-title":"Proceedings of Interspeech'11","author":"Kempton T.","unstructured":"Kempton , T. , Moore , R. K. , and Hain , T . 2011. Cross-language phone recognition when the target language is not known . In Proceedings of Interspeech'11 . 3165--3168. Kempton, T., Moore, R. K., and Hain, T. 2011. Cross-language phone recognition when the target language is not known. In Proceedings of Interspeech'11. 3165--3168."},{"key":"e_1_2_2_25_1","volume-title":"Proceedings of the Conference on Speech and Computer. 156--159","author":"Kim J.","unstructured":"Kim , J. , Jung , H. , and Chung , H . 2004. A keyword spotting approach based on pseudo N-gram language model . In Proceedings of the Conference on Speech and Computer. 156--159 . Kim, J., Jung, H., and Chung, H. 2004. A keyword spotting approach based on pseudo N-gram language model. In Proceedings of the Conference on Speech and Computer. 156--159."},{"key":"e_1_2_2_26_1","volume-title":"Proceedings of Interspeech'08","author":"Lin H.","unstructured":"Lin , H. , Stupakov , A. , and Bilmes , J . 2008. Spoken keyword spotting via multi-lattice alignment . In Proceedings of Interspeech'08 , 2191--2194. Lin, H., Stupakov, A., and Bilmes, J. 2008. Spoken keyword spotting via multi-lattice alignment. In Proceedings of Interspeech'08, 2191--2194."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2009.4960724"},{"key":"e_1_2_2_28_1","volume-title":"Proceedings of Interspeech'08","author":"Mamou J.","unstructured":"Mamou , J. and Ramabhadran , B . 2008. Phonetic query expansion for spoken document retrieval . In Proceedings of Interspeech'08 , 2106--2109. Mamou, J. and Ramabhadran, B. 2008. Phonetic query expansion for spoken document retrieval. In Proceedings of Interspeech'08, 2106--2109."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1277741.1277847"},{"key":"e_1_2_2_30_1","volume-title":"Proceedings of ICASSP'97","volume":"2","author":"Manos A.","unstructured":"Manos , A. and Zue , V . 1997. A segment-based wordspotter using phonetic filler models . In Proceedings of ICASSP'97 , vol. 2 , 899--902. Manos, A. and Zue, V. 1997. A segment-based wordspotter using phonetic filler models. In Proceedings of ICASSP'97, vol. 2, 899--902."},{"key":"e_1_2_2_31_1","volume-title":"Proceedings of MediaEval'11","author":"Mantena G. V.","unstructured":"Mantena , G. V. , Bollepalli , B. , and Prahallad , K . 2011. SWS task: Articulary phonetic units and sliding DTW . In Proceedings of MediaEval'11 . 7--8. Mantena, G. V., Bollepalli, B., and Prahallad, K. 2011. SWS task: Articulary phonetic units and sliding DTW. In Proceedings of MediaEval'11. 7--8."},{"key":"e_1_2_2_32_1","volume-title":"Proceedings of ICASSP'12","author":"Metze F.","unstructured":"Metze , F. , Rajput , N. , Anguera , X. , Davel , M. , Gravier , G. , Van Heerden , C. , Mantena , G. V. , Muscariello , A. , Prahallad , K. , Szoke , I. , and Tejedor , J . 2012. The spoken web search task at MediaEval 2011 . In Proceedings of ICASSP'12 . 5165--5168. Metze, F., Rajput, N., Anguera, X., Davel, M., Gravier, G., Van Heerden, C., Mantena, G. V., Muscariello, A., Prahallad, K., Szoke, I., and Tejedor, J. 2012. The spoken web search task at MediaEval 2011. In Proceedings of ICASSP'12. 5165--5168."},{"key":"e_1_2_2_33_1","volume-title":"Proceedings of MediaEval'11","author":"Muscariello A.","unstructured":"Muscariello , A. and Gravier , G . 2011. Irisa MediaEval 2011 spoken Web search system . In Proceedings of MediaEval'11 . 9--10. Muscariello, A. and Gravier, G. 2011. Irisa MediaEval 2011 spoken Web search system. In Proceedings of MediaEval'11. 9--10."},{"key":"e_1_2_2_34_1","volume-title":"Proceedings of Interspeech'09","author":"Muscariello A.","unstructured":"Muscariello , A. , Gravier , G. , and Bimbot , F . 2009. Audio keyword extraction by unsupervised word discovery . In Proceedings of Interspeech'09 . 2843--2846. Muscariello, A., Gravier, G., and Bimbot, F. 2009. Audio keyword extraction by unsupervised word discovery. In Proceedings of Interspeech'09. 2843--2846."},{"key":"e_1_2_2_35_1","volume-title":"Proceedings of Interspeech'11","author":"Muscariello A.","unstructured":"Muscariello , A. , Gravier , G. , and Bimbot , F . 2011. Zero-resource audio-only spoken term detection based on a combination of template matching techniques . In Proceedings of Interspeech'11 . 921--924. Muscariello, A., Gravier, G., and Bimbot, F. 2011. Zero-resource audio-only spoken term detection based on a combination of template matching techniques. In Proceedings of Interspeech'11. 921--924."},{"key":"e_1_2_2_36_1","first-page":"6","article-title":"The road rally word-spotting corpora (RDRALLY1)","author":"NIST.","year":"1991","unstructured":"NIST. 1991 . The road rally word-spotting corpora (RDRALLY1) . NIST Speech Disc 6 - 1 .1. NIST. 1991. The road rally word-spotting corpora (RDRALLY1). NIST Speech Disc 6-1.1.","journal-title":"NIST Speech Disc"},{"key":"e_1_2_2_37_1","unstructured":"NIST. 2006. The spoken term detection (STD) 2006 evaluation plan. http:\/\/www.nist.gov\/speech\/tests\/std.  NIST. 2006. The spoken term detection (STD) 2006 evaluation plan. http:\/\/www.nist.gov\/speech\/tests\/std."},{"key":"e_1_2_2_38_1","volume-title":"Proceedings of the International Conference On Neural Information Processing.","author":"Ou J.","unstructured":"Ou , J. , Chen , C. , and Li , Z . 2001. Hybrid neural-network\/HMM approach for out-of-vocabulary words rejection in Mandarin place name recognition . In Proceedings of the International Conference On Neural Information Processing. Ou, J., Chen, C., and Li, Z. 2001. Hybrid neural-network\/HMM approach for out-of-vocabulary words rejection in Mandarin place name recognition. In Proceedings of the International Conference On Neural Information Processing."},{"key":"e_1_2_2_39_1","volume-title":"Proceedings of ASRU'09","author":"Parada C.","unstructured":"Parada , C. , Sethy , A. , and Ramabhadran , B . 2009. Query-by-example spoken term detection for OOV terms . In Proceedings of ASRU'09 . 404--409. Parada, C., Sethy, A., and Ramabhadran, B. 2009. Query-by-example spoken term detection for OOV terms. In Proceedings of ASRU'09. 404--409."},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2007.909282"},{"key":"e_1_2_2_41_1","volume-title":"Proceedings of Interspeech'01","author":"Schultz T.","unstructured":"Schultz , T. and Waibel , A . 2001. Experiments on cross-language acoustic modeling . In Proceedings of Interspeech'01 . 2721--2724. Schultz, T. and Waibel, A. 2001. Experiments on cross-language acoustic modeling. In Proceedings of Interspeech'01. 2721--2724."},{"key":"e_1_2_2_43_1","volume-title":"Proceedings of Interspeech'09","author":"Shen W.","unstructured":"Shen , W. , White , C. M. , and Hazen , T. J . 2009. A comparison of query-by-example methods for spoken term detection . In Proceedings of Interspeech'09 . 2143--2146. Shen, W., White, C. M., and Hazen, T. J. 2009. A comparison of query-by-example methods for spoken term detection. In Proceedings of Interspeech'09. 2143--2146."},{"key":"e_1_2_2_44_1","volume-title":"Proceedings of ICASSP'08","author":"Sinischalchi S. M.","unstructured":"Sinischalchi , S. M. , Svendsen , T. , and Lee , C. H . 2008. Toward a detector-based universal phone recognizer . In Proceedings of ICASSP'08 . 4261--4264. Sinischalchi, S. M., Svendsen, T., and Lee, C. H. 2008. Toward a detector-based universal phone recognizer. In Proceedings of ICASSP'08. 4261--4264."},{"key":"e_1_2_2_46_1","volume-title":"Proceedings of SLT'08","author":"Szoke I.","unstructured":"Szoke , I. , Burget , L. , & Cbreve ;ernocky, J., and Fapso , M . 2008a. Sub-word modeling of out of vocabulary words in spoken term detection . In Proceedings of SLT'08 . 273--276. Szoke, I., Burget, L., &Cbreve;ernocky, J., and Fapso, M. 2008a. Sub-word modeling of out of vocabulary words in spoken term detection. In Proceedings of SLT'08. 273--276."},{"key":"e_1_2_2_47_1","volume-title":"Proceedings of the Speech Search Workshop at SIGIR (SSCS'08)","author":"Szoke I.","year":"2008","unstructured":"Szoke , I. , Fapso , M. , Burget , L. , and & Cbreve ;ernocky, J. 2008 b. Hybrid word-subword decoding for spoken term detection . In Proceedings of the Speech Search Workshop at SIGIR (SSCS'08) . 42--49. Szoke, I., Fapso, M., Burget, L., and &Cbreve;ernocky, J. 2008b. Hybrid word-subword decoding for spoken term detection. In Proceedings of the Speech Search Workshop at SIGIR (SSCS'08). 42--49."},{"key":"e_1_2_2_48_1","series-title":"Lecture Notes in Computer Science","volume-title":"Machine Learning for Multimodal Interaction","author":"Szoke I.","unstructured":"Szoke , I. , Fapso , M. , Karafiat , M. , Burget , L. , Grezl , F. , Schwarz , P. , Glembek , O. , Matfjka , P. , Kopecky , J. , and & Cbreve ;ernocky, J. 2008c. Spoken term detection system based on combination of LVCSR and phonetic search . In Machine Learning for Multimodal Interaction , Lecture Notes in Computer Science , vol. 4892 , Springer , Berlin , 237--247. Szoke, I., Fapso, M., Karafiat, M., Burget, L., Grezl, F., Schwarz, P., Glembek, O., Matfjka, P., Kopecky, J., and &Cbreve;ernocky, J. 2008c. Spoken term detection system based on combination of LVCSR and phonetic search. In Machine Learning for Multimodal Interaction, Lecture Notes in Computer Science, vol. 4892, Springer, Berlin, 237--247."},{"key":"e_1_2_2_49_1","volume-title":"Proceedings of SLT'10","author":"Szoke I.","unstructured":"Szoke , I. , Grezl , F. , \u010cernocky , J. , and Fapso , M . 2010. Acoustic keyword spotter\u2014Optimization from end-user perspective . In Proceedings of SLT'10 . 177--181. Szoke, I., Grezl, F., \u010cernocky, J., and Fapso, M. 2010. Acoustic keyword spotter\u2014Optimization from end-user perspective. In Proceedings of SLT'10. 177--181."},{"key":"e_1_2_2_50_1","volume-title":"Proceedings of Interspeech'05","author":"Szoke I.","unstructured":"Szoke , I. , Schwarz , P. , Matejka , P. , Burget , L. , Karafiat , M. , Fapso , M. , and Cernocky , J . 2005. Comparison of keyword spotting approaches for informal continuous speech . In Proceedings of Interspeech'05 . 633--636. Szoke, I., Schwarz, P., Matejka, P., Burget, L., Karafiat, M., Fapso, M., and Cernocky, J. 2005. Comparison of keyword spotting approaches for informal continuous speech. In Proceedings of Interspeech'05. 633--636."},{"key":"e_1_2_2_51_1","volume-title":"Proceedings of MediaEval'11","author":"Szoke I.","unstructured":"Szoke , I. , Tejedor , J. , Fapso , M. , and Colas , J . 2011. BUT-HCT Lab approaches for spoken web search: MediaEval 2011 . In Proceedings of MediaEval'11 . 11--12. Szoke, I., Tejedor, J., Fapso, M., and Colas, J. 2011. BUT-HCT Lab approaches for spoken web search: MediaEval 2011. In Proceedings of MediaEval'11. 11--12."},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/1878101.1878106"},{"key":"e_1_2_2_54_1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo. 1863--1866","author":"Tsai W.-H.","unstructured":"Tsai , W.-H. and Wang , H . -M. 2004. A query-by-example framework to retrievemusic documents by singer . In Proceedings of the IEEE International Conference on Multimedia and Expo. 1863--1866 . Tsai, W.-H. and Wang, H.-M. 2004. A query-by-example framework to retrievemusic documents by singer. In Proceedings of the IEEE International Conference on Multimedia and Expo. 1863--1866."},{"key":"e_1_2_2_55_1","volume-title":"Proceedings of the 3rd International Conference on Music Information Retrieval (ISMIR). 31--38","author":"Tzanetakis G.","unstructured":"Tzanetakis , G. , Ermolinskyi , A. , and Cook , P . 2002. Pitch histograms in audio and symbolic music information retrieval . In Proceedings of the 3rd International Conference on Music Information Retrieval (ISMIR). 31--38 . Tzanetakis, G., Ermolinskyi, A., and Cook, P. 2002. Pitch histograms in audio and symbolic music information retrieval. In Proceedings of the 3rd International Conference on Music Information Retrieval (ISMIR). 31--38."},{"key":"e_1_2_2_56_1","volume-title":"Proceedings of Interspeech'07","author":"Vergyri D.","unstructured":"Vergyri , D. , Shafran , I. , Stolcke , A. , Gadde , R. R. , Akbacak , M. , Roark , B. , and Wang , W . 2007. The SRI\/OGI 2006 spoken term detection system . In Proceedings of Interspeech'07 . 2393--2396. Vergyri, D., Shafran, I., Stolcke, A., Gadde, R. R., Akbacak, M., Roark, B., and Wang, W. 2007. The SRI\/OGI 2006 spoken term detection system. In Proceedings of Interspeech'07. 2393--2396."},{"key":"e_1_2_2_57_1","volume-title":"The SRI 2006 spoken term detection system. In Proceedings of NIST Spoken Term Detection Workshop (STD'06)","author":"Vergyri D.","unstructured":"Vergyri , D. , Stolcke , A. , Gadde , R. R. , and Wang , W . 2006 . The SRI 2006 spoken term detection system. In Proceedings of NIST Spoken Term Detection Workshop (STD'06) . Vergyri, D., Stolcke, A., Gadde, R. R., and Wang, W. 2006. The SRI 2006 spoken term detection system. In Proceedings of NIST Spoken Term Detection Workshop (STD'06)."},{"key":"e_1_2_2_58_1","volume-title":"Proceedings of Interspeech'03","author":"Walker B. D.","unstructured":"Walker , B. D. , Lackey , B. C. , Muller , J. S. , and Schone , P. J . 2003. Language-reconfigurable universal phone recognition . In Proceedings of Interspeech'03 . 153--156. Walker, B. D., Lackey, B. C., Muller, J. S., and Schone, P. J. 2003. Language-reconfigurable universal phone recognition. In Proceedings of Interspeech'03. 153--156."},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1999.759780"},{"key":"e_1_2_2_60_1","first-page":"21","article-title":"Multiple sequence alignment using GA and NN","volume":"1","author":"Wu S.","year":"2008","unstructured":"Wu , S. , Lee , M. , Lee , Y. S. , and Gatton , T. M. 2008 . Multiple sequence alignment using GA and NN . Int. J. Signal Process. Image Process. Pattern Recogn. 1 , 21 -- 31 . Wu, S., Lee, M., Lee, Y. S., and Gatton, T. M. 2008. Multiple sequence alignment using GA and NN. Int. J. Signal Process. Image Process. Pattern Recogn. 1, 21--31.","journal-title":"Int. J. Signal Process. Image Process. Pattern Recogn."},{"key":"e_1_2_2_61_1","volume-title":"Proceedings of the International Conference on Infotec and Infonet.","volume":"3","author":"Xin L.","unstructured":"Xin , L. and Wang , B . 2001. Utterance verification for spontaneous Mandarin speech keyword spotting . In Proceedings of the International Conference on Infotec and Infonet. Vol. 3 , 397--401. Xin, L. and Wang, B. 2001. Utterance verification for spontaneous Mandarin speech keyword spotting. In Proceedings of the International Conference on Infotec and Infonet. Vol. 3, 397--401."},{"key":"e_1_2_2_62_1","unstructured":"Young S. J. Kershaw D. Odell J. Ollason D. Valtchev V. and Woodland P. 2006. The HTK Book Version 3.4. Cambridge University Press.  Young S. J. Kershaw D. Odell J. Ollason D. Valtchev V. and Woodland P. 2006. The HTK Book Version 3.4. Cambridge University Press."},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1999.757472"},{"key":"e_1_2_2_64_1","volume-title":"Proceedings of ASRU'09","author":"Zhang Y.","unstructured":"Zhang , Y. and Glass , J. R . 2009. Unsupervised spoken keyword spotting via segmental DTW on Gaussian posteriorgrams . In Proceedings of ASRU'09 . 398--403. Zhang, Y. and Glass, J. R. 2009. Unsupervised spoken keyword spotting via segmental DTW on Gaussian posteriorgrams. In Proceedings of ASRU'09. 398--403."}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2328967.2328971","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2328967.2328971","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:01:10Z","timestamp":1750276870000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2328967.2328971"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,8]]},"references-count":61,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2012,8]]}},"alternative-id":["10.1145\/2328967.2328971"],"URL":"https:\/\/doi.org\/10.1145\/2328967.2328971","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,8]]},"assertion":[{"value":"2011-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-09-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}