{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T13:12:23Z","timestamp":1783775543530,"version":"3.55.0"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2008,10,4]],"date-time":"2008-10-04T00:00:00Z","timestamp":1223078400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001665","name":"Agence Nationale de la Recherche","doi-asserted-by":"publisher","award":["ANR-08-CORD-013"],"award-info":[{"award-number":["ANR-08-CORD-013"]}],"id":[{"id":"10.13039\/501100001665","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001691","name":"Japan Society for the Promotion of Science","doi-asserted-by":"publisher","award":["(A)21240021"],"award-info":[{"award-number":["(A)21240021"]}],"id":[{"id":"10.13039\/501100001691","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Speech Lang. Process."],"published-print":{"date-parts":[[2010,8]]},"abstract":"<jats:p>Current research in text mining favors the quantity of texts over their representativeness. But for bilingual terminology mining, and for many language pairs, large comparable corpora are not available. More importantly, as terms are defined vis-\u00e0-vis a specific domain with a restricted register, it is expected that the representativeness rather than the quantity of the corpus matters more in terminology mining. Our hypothesis, therefore, is that the representativeness of the corpus is more important than the quantity and ensures the quality of the acquired terminological resources. This article tests this hypothesis on a French-Japanese bilingual term extraction task. To demonstrate how important the type of discourse is as a characteristic of the comparable corpora, we used a state-of-the-art multilingual terminology mining chain composed of two extraction programs, one in each language, and an alignment program. We evaluated the candidate translations using a reference list, and found that taking discourse type into account resulted in candidate translations of a better quality even when the corpus size was reduced by half.<\/jats:p>","DOI":"10.1145\/1839478.1839479","type":"journal-article","created":{"date-parts":[[2010,10,1]],"date-time":"2010-10-01T12:29:13Z","timestamp":1285936153000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Brains, not brawn"],"prefix":"10.1145","volume":"7","author":[{"given":"Emmanuel","family":"Morin","sequence":"first","affiliation":[{"name":"Universit\u00e9 de Nantes, Cedex, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"B\u00e9atrice","family":"Daille","sequence":"additional","affiliation":[{"name":"Universit\u00e9 de Nantes, Cedex, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Koichi","family":"Takeuchi","sequence":"additional","affiliation":[{"name":"Okayama University, Okayama, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kyo","family":"Kageura","sequence":"additional","affiliation":[{"name":"The University of Tokyo, Bunkyo-ku, Tokyo, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2008,10,4]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the ACL Workshop on Multiword Expressions: Integrating Processing. 24--31","author":"Baldwin T.","unstructured":"}} Baldwin , T. and Tanaka , T . 2004. Translation by machine of complex nominals: Getting it right . In Proceedings of the ACL Workshop on Multiword Expressions: Integrating Processing. 24--31 . }}Baldwin, T. and Tanaka, T. 2004. Translation by machine of complex nominals: Getting it right. In Proceedings of the ACL Workshop on Multiword Expressions: Integrating Processing. 24--31."},{"key":"e_1_2_1_2_1","first-page":"579","article-title":"Morphosyntaxe et genres textuels. Exploiter des donn\u00e9es morphosyntaxiques pour l'\u00e9tude statistique des genres textuels : Application au roman policier","volume":"42","author":"Beauvisage T.","year":"2001","unstructured":"}} Beauvisage , T. 2001 . Morphosyntaxe et genres textuels. Exploiter des donn\u00e9es morphosyntaxiques pour l'\u00e9tude statistique des genres textuels : Application au roman policier . Traitement Autom. Lang. 42 , 2, 579 -- 608 . }}Beauvisage, T. 2001. Morphosyntaxe et genres textuels. Exploiter des donn\u00e9es morphosyntaxiques pour l'\u00e9tude statistique des genres textuels : Application au roman policier. Traitement Autom. Lang. 42, 2, 579--608.","journal-title":"Traitement Autom. Lang."},{"key":"e_1_2_1_3_1","volume-title":"Current Issues in Computational Linguistics: In Honour of Don Walker","author":"Biber D.","unstructured":"}} Biber , D. 1994. Representativeness in corpus design . In Current Issues in Computational Linguistics: In Honour of Don Walker , A. Zampolli, N. Calzolari, and M. Palmer, Eds. Kluwer Academic Publishers , 377--407. }}Biber, D. 1994. Representativeness in corpus design. In Current Issues in Computational Linguistics: In Honour of Don Walker, A. Zampolli, N. Calzolari, and M. Palmer, Eds. Kluwer Academic Publishers, 377--407."},{"key":"e_1_2_1_4_1","volume-title":"Dimensions of Register Variation: A Cross-Linguistic Comparison","author":"Biber D.","unstructured":"}} Biber , D. 1995. Dimensions of Register Variation: A Cross-Linguistic Comparison . Cambridge University Press , Cambridge, UK . }}Biber, D. 1995. Dimensions of Register Variation: A Cross-Linguistic Comparison. Cambridge University Press, Cambridge, UK."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.4324\/9780203469255"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the 12th National Conference on Artificial Intelligence (AAAI'94)","author":"Brill E.","year":"1994","unstructured":"}} Brill , E. 1994 . Some advances in transformation-based part of speech tagging . In Proceedings of the 12th National Conference on Artificial Intelligence (AAAI'94) . 722--727. }}Brill, E. 1994. Some advances in transformation-based part of speech tagging. In Proceedings of the 12th National Conference on Artificial Intelligence (AAAI'94). 722--727."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/972470.972474"},{"key":"e_1_2_1_8_1","volume-title":"Eds. Natural Language Processing","volume":"2","author":"Cabr\u00e9 M. T.","unstructured":"}} Cabr\u00e9 , M. T. , Bagot , R. E. , and Platresi , J. V . 2001. Automatic term detection: A review of current systems. In Recent Advances in Computational Terminology, D. Bourigault, C. Jacquemin, and M.-C. L'Homme , Eds. Natural Language Processing , vol. 2 . John Benjamins, 53--88. }}Cabr\u00e9, M. T., Bagot, R. E., and Platresi, J. V. 2001. Automatic term detection: A review of current systems. In Recent Advances in Computational Terminology, D. Bourigault, C. Jacquemin, and M.-C. L'Homme, Eds. Natural Language Processing, vol. 2. John Benjamins, 53--88."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.3115\/1072228.1072239"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.3115\/990820.990842"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.3115\/1071884.1071904"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/972450.972452"},{"key":"e_1_2_1_13_1","volume-title":"The Balancing Act: Combining Symbolic and Statistical Approaches to Language","author":"Daille B.","unstructured":"}} Daille , B. 1996. Study and implementation of combined techniques for automatic extraction of terminology . In The Balancing Act: Combining Symbolic and Statistical Approaches to Language , J. Klavans and P. Resnik, Eds. The MIT Press , Cambridge, MA , 49--66. }}Daille, B. 1996. Study and implementation of combined techniques for automatic extraction of terminology. In The Balancing Act: Combining Symbolic and Statistical Approaches to Language, J. Klavans and P. Resnik, Eds. The MIT Press, Cambridge, MA, 49--66."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.3115\/1119282.1119284"},{"key":"e_1_2_1_15_1","doi-asserted-by":"crossref","unstructured":"}}Daille B. 2003b. Terminology Mining. In Information Extraction in the Web Era M. T. Pazienza Ed. Springer 29--44.  }}Daille B. 2003b. Terminology Mining. In Information Extraction in the Web Era M. T. Pazienza Ed. Springer 29--44.","DOI":"10.1007\/978-3-540-45092-4_2"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/11562214_62"},{"key":"e_1_2_1_17_1","unstructured":"}}D\u00e9jean H. and Gaussier E. 2002. Une nouvelle approche \u00e0 l'extraction de lexiques bilingues \u00e0 partir de corpus comparables. In Lexicometrica Alignement lexical dans les Corpus Multilingues 1--22.  }}D\u00e9jean H. and Gaussier E. 2002. Une nouvelle approche \u00e0 l'extraction de lexiques bilingues \u00e0 partir de corpus comparables. In Lexicometrica Alignement lexical dans les Corpus Multilingues 1--22."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.3115\/1072228.1072394"},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 6th International Conference on Computer-Assisted Information Retrieval (RIAO'00)","author":"Diab M. T.","unstructured":"}} Diab , M. T. and Finch , S . 2000. A statistical word-level translation model for comparable corpora . In Proceedings of the 6th International Conference on Computer-Assisted Information Retrieval (RIAO'00) . 1500--1508. }}Diab, M. T. and Finch, S. 2000. A statistical word-level translation model for comparable corpora. In Proceedings of the 6th International Conference on Computer-Assisted Information Retrieval (RIAO'00). 1500--1508."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/972450.972454"},{"key":"e_1_2_1_21_1","volume-title":"Transmission of Information: A Statistical Theory of Communications","author":"Fano R. M.","unstructured":"}} Fano , R. M. 1961. Transmission of Information: A Statistical Theory of Communications . MIT Press , Cambridge, MA . }}Fano, R. M. 1961. Transmission of Information: A Statistical Theory of Communications. MIT Press, Cambridge, MA."},{"key":"e_1_2_1_22_1","unstructured":"}}French-Japanese Scientific Dictionary. 1989. 4th Ed. Hakusuisha.  }}French-Japanese Scientific Dictionary. 1989. 4th Ed. Hakusuisha."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/648179.749226"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the 5th Annual Workshop on Very Large Corpora (VLC'97)","author":"Fung P.","year":"1997","unstructured":"}} Fung , P. and McKeown , K. 1997 . Finding terminology translations from non-parallel corpora . In Proceedings of the 5th Annual Workshop on Very Large Corpora (VLC'97) . 192--202. }}Fung, P. and McKeown, K. 1997. Finding terminology translations from non-parallel corpora. In Proceedings of the 5th Annual Workshop on Very Large Corpora (VLC'97). 192--202."},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the 2nd Workshop on Building and Using Comparable Corpora: From Parallel to Non-parallel Corpora (BUCC'09)","author":"Goeuriot L.","unstructured":"}} Goeuriot , L. , Morin , E. , and Daille , B . 2009. Compilation of specialized comparable corpora in French and Japanese . In Proceedings of the 2nd Workshop on Building and Using Comparable Corpora: From Parallel to Non-parallel Corpora (BUCC'09) . 55--62. }}Goeuriot, L., Morin, E., and Daille, B. 2009. Compilation of specialized comparable corpora in French and Japanese. In Proceedings of the 2nd Workshop on Building and Using Comparable Corpora: From Parallel to Non-parallel Corpora (BUCC'09). 55--62."},{"key":"e_1_2_1_26_1","volume-title":"Explorations in Automatic Thesaurus Discovery","author":"Grefenstette G.","unstructured":"}} Grefenstette , G. 1994. Explorations in Automatic Thesaurus Discovery . Kluwer Academic Publisher , Boston, MA . }}Grefenstette, G. 1994. Explorations in Automatic Thesaurus Discovery. Kluwer Academic Publisher, Boston, MA."},{"key":"e_1_2_1_27_1","unstructured":"}}Grefenstette G. 1999. The Word Wide Web as a resource for example-based machine translation tasks. In Translating and the Computer 21 (ASLIB'99).  }}Grefenstette G. 1999. The Word Wide Web as a resource for example-based machine translation tasks. In Translating and the Computer 21 (ASLIB'99)."},{"key":"e_1_2_1_28_1","volume-title":"Spotting and Discovering Terms through Natural Language Processing","author":"Jacquemin C.","unstructured":"}} Jacquemin , C. 2001. Spotting and Discovering Terms through Natural Language Processing . MIT Press , Cambridge, MA . }}Jacquemin, C. 2001. Spotting and Discovering Terms through Natural Language Processing. MIT Press, Cambridge, MA."},{"key":"e_1_2_1_29_1","volume-title":"Handbook of Logic and Language, J. van Benthem and A. ter Meulen, Eds","author":"Janssen T. M. V.","unstructured":"}} Janssen , T. M. V. 1996. Compositionality . In Handbook of Logic and Language, J. van Benthem and A. ter Meulen, Eds . Elsevier , Amsterdam , 417--473. }}Janssen, T. M. V. 1996. Compositionality. In Handbook of Logic and Language, J. van Benthem and A. ter Meulen, Eds. Elsevier, Amsterdam, 417--473."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1075\/term.10.1.02kag"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.3115\/991250.991324"},{"key":"e_1_2_1_32_1","unstructured":"}}Manning C. and Sch\u00fctze H. 1999. Foundations of Statistical Natural Language Processing. MIT Press Cambridge MA.   }}Manning C. and Sch\u00fctze H. 1999. Foundations of Statistical Natural Language Processing. MIT Press Cambridge MA."},{"key":"e_1_2_1_33_1","unstructured":"}}Matsumoto Y. Kitauchi A. Yamashita T. and Hirano Y. 1999. Japanese morphological analysis system Chasen 2.0 users manual. Tech. rep. Nara Institute of Science and Technology (NAIST).  }}Matsumoto Y. Kitauchi A. Yamashita T. and Hirano Y. 1999. Japanese morphological analysis system Chasen 2.0 users manual. Tech. rep. Nara Institute of Science and Technology (NAIST)."},{"key":"e_1_2_1_34_1","doi-asserted-by":"crossref","unstructured":"}}McEnery A. M. and Xiao R. Z. 2007. Parallel and comparable corpora: What are they up to&quest; In Incorporating Corpora: Translation and the Linguist G. M. Anderman and M. Rogers Eds. Multilingual Matters. Clevedon UK Chapter 2.  }}McEnery A. M. and Xiao R. Z. 2007. Parallel and comparable corpora: What are they up to&quest; In Incorporating Corpora: Translation and the Linguist G. M. Anderman and M. Rogers Eds. Multilingual Matters. Clevedon UK Chapter 2.","DOI":"10.21832\/9781853599873-005"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.3115\/976909.979680"},{"key":"e_1_2_1_36_1","volume-title":"Empirical Methods for Exploiting Parallel Texts","author":"Melamed I. D.","unstructured":"}} Melamed , I. D. 2001. Empirical Methods for Exploiting Parallel Texts . MIT Press , Cambridge, MA . }}Melamed, I. D. 2001. Empirical Methods for Exploiting Parallel Texts. MIT Press, Cambridge, MA."},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the 1st Workshop on Computational Terminology (COMPTERM'98)","author":"Nakagawa H.","unstructured":"}} Nakagawa , H. and Mori , T . 1998. Nested collocation and compound noun for term recognition . In Proceedings of the 1st Workshop on Computational Terminology (COMPTERM'98) . D. Bourigault, C. Jacquemin, and M.-C. L'Homme, Eds. 64--70. }}Nakagawa, H. and Mori, T. 1998. Nested collocation and compound noun for term recognition. In Proceedings of the 1st Workshop on Computational Terminology (COMPTERM'98). D. Bourigault, C. Jacquemin, and M.-C. L'Homme, Eds. 64--70."},{"key":"e_1_2_1_38_1","first-page":"523","article-title":"Flemm: Un analyseur flexionnel du fran\u00e7ais \u00e0 base de r\u00e8gles","volume":"41","author":"Namer F.","year":"2000","unstructured":"}} Namer , F. 2000 . Flemm: Un analyseur flexionnel du fran\u00e7ais \u00e0 base de r\u00e8gles . Traitement Autom. Lang. 41 , 2, 523 -- 547 . }}Namer, F. 2000. Flemm: Un analyseur flexionnel du fran\u00e7ais \u00e0 base de r\u00e8gles. Traitement Autom. Lang. 41, 2, 523--547.","journal-title":"Traitement Autom. Lang."},{"key":"e_1_2_1_39_1","volume-title":"Eds","author":"Nazarenko A.","year":"2002","unstructured":"}} Nazarenko , A. and Hamon , T. , Eds . 2002 . Structuration de terminologie. Traitement Autom. Lang . 43. }}Nazarenko, A. and Hamon, T., Eds. 2002. Structuration de terminologie. Traitement Autom. Lang. 43."},{"key":"e_1_2_1_40_1","volume-title":"Gakujutu Yogo Goki-Hyo","author":"Nomura M.","unstructured":"}} Nomura , M. and M., I. 1989. Gakujutu Yogo Goki-Hyo . National Language Research Institute , Tokyo . }}Nomura, M. and M., I. 1989. Gakujutu Yogo Goki-Hyo. National Language Research Institute, Tokyo."},{"key":"e_1_2_1_41_1","first-page":"81","article-title":"Cross-Language information retrieval: A system for comparable corpus querying. In Cross-Language Information Retrieval, G. Grefenstette, Ed. Kluwer Academic Publishers","volume":"7","author":"Peters C.","year":"1998","unstructured":"}} Peters , C. and Picchi , E. 1998 . Cross-Language information retrieval: A system for comparable corpus querying. In Cross-Language Information Retrieval, G. Grefenstette, Ed. Kluwer Academic Publishers , Chapter 7 , 81 -- 90 . }}Peters, C. and Picchi, E. 1998. Cross-Language information retrieval: A system for comparable corpus querying. In Cross-Language Information Retrieval, G. Grefenstette, Ed. Kluwer Academic Publishers, Chapter 7, 81--90.","journal-title":"Chapter"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.3115\/981658.981709"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.3115\/1034678.1034756"},{"key":"e_1_2_1_44_1","volume-title":"Proceedings of the 11th Conference of the European Chapter of the Association for Computational Linguistics (EACL'06)","author":"Robitaille X.","unstructured":"}} Robitaille , X. , Sasaki , X. , Tonoike , M. , Sato , S. , and Utsuro , S . 2006. Compiling French-Japanese terminologies from the Web . In Proceedings of the 11th Conference of the European Chapter of the Association for Computational Linguistics (EACL'06) . 225--232. }}Robitaille, X., Sasaki, X., Tonoike, M., Sato, S., and Utsuro, S. 2006. Compiling French-Japanese terminologies from the Web. In Proceedings of the 11th Conference of the European Chapter of the Association for Computational Linguistics (EACL'06). 225--232."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.3115\/1118935.1118943"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(88)90021-0"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/321439.321441"},{"key":"e_1_2_1_48_1","doi-asserted-by":"crossref","unstructured":"}}Savary A. and Jacquemin C. 2003. Reducing information variation in text. In Text- and Speech-Triggered Information Access G. Grefenstette Ed. Lecture Notes in Computer Science. Springer Verlag 141--181.  }}Savary A. and Jacquemin C. 2003. Reducing information variation in text. In Text- and Speech-Triggered Information Access G. Grefenstette Ed. Lecture Notes in Computer Science. Springer Verlag 141--181.","DOI":"10.1007\/978-3-540-45115-0_6"},{"key":"e_1_2_1_49_1","volume-title":"Proceedings of the 3rd International Workshop on Computational Terminology (COMPUTERM'04)","author":"Takeuchi K.","unstructured":"}} Takeuchi , K. , Kageura , K. , Daille , B. , and Romary , L . 2004. Construction of grammar based term extraction model for Japanese . In Proceedings of the 3rd International Workshop on Computational Terminology (COMPUTERM'04) . S. Ananadiou and P. Zweigenbaum, Eds. 91--94. }}Takeuchi, K., Kageura, K., Daille, B., and Romary, L. 2004. Construction of grammar based term extraction model for Japanese. In Proceedings of the 3rd International Workshop on Computational Terminology (COMPUTERM'04). S. Ananadiou and P. Zweigenbaum, Eds. 91--94."}],"container-title":["ACM Transactions on Speech and Language Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1839478.1839479","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1839478.1839479","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T11:22:35Z","timestamp":1750245755000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1839478.1839479"}},"subtitle":["The use of \u201csmart\u201d comparable corpora in bilingual terminology mining"],"short-title":[],"issued":{"date-parts":[[2008,10,4]]},"references-count":49,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2010,8]]}},"alternative-id":["10.1145\/1839478.1839479"],"URL":"https:\/\/doi.org\/10.1145\/1839478.1839479","relation":{},"ISSN":["1550-4875","1550-4883"],"issn-type":[{"value":"1550-4875","type":"print"},{"value":"1550-4883","type":"electronic"}],"subject":[],"published":{"date-parts":[[2008,10,4]]},"assertion":[{"value":"2009-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2010-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2008-10-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}