{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,8,22]],"date-time":"2024-08-22T11:23:24Z","timestamp":1724325804391},"reference-count":38,"publisher":"Cambridge University Press (CUP)","issue":"5","license":[{"start":{"date-parts":[[2018,7,19]],"date-time":"2018-07-19T00:00:00Z","timestamp":1531958400000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2018,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The extraction of templates such as \u2018regard X as Y\u2019 from a set of related phrases requires the identification of their internal structures. This paper presents an unsupervised approach for extracting templates on-the-fly from only tagged text by using a novel relaxed variant of the Sequence Binary Decision Diagram (SeqBDD). A SeqBDD can compress a set of sequences into a graphical structure equivalent to a minimal deterministic finite state automata, but more compact and better suited to the task of template extraction. The main contribution of this paper is a relaxed form of the SeqBDD construction algorithm that enables it to form general representations from a small amount of data. The process of compression of shared structures in the text during Relaxed SeqBDD construction, naturally induces the templates we wish to extract. Experiments show that the method is capable of high-quality extraction on tasks based on verb+preposition templates from corpora and phrasal templates from short messages from social media.<\/jats:p>","DOI":"10.1017\/s1351324918000268","type":"journal-article","created":{"date-parts":[[2018,7,19]],"date-time":"2018-07-19T06:58:27Z","timestamp":1531983507000},"page":"763-795","source":"Crossref","is-referenced-by-count":2,"title":["Extraction of templates from phrases using Sequence Binary Decision Diagrams"],"prefix":"10.1017","volume":"24","author":[{"given":"D.","family":"HIRANO","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"K.","family":"TANAKA-ISHII","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"A.","family":"FINCH","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2018,7,19]]},"reference":[{"key":"S1351324918000268_ref036","first-page":"143","article-title":"Retrieving collocations from text: Xtract","volume":"19","author":"Smadja","year":"1993","journal-title":"Computational Linguistics"},{"key":"S1351324918000268_ref033","unstructured":"Pei J. , Han J. , Mortazavi-Asl B. , Pinto H. , Chen Q. , Dayal U. , and Hsu M. 2001. Prefixspan: mining sequential patterns by prefix-projected growth. In Proceedings of the 17th International Conference on Data Engineering. IEEE Computer Society, 215\u2013224."},{"key":"S1351324918000268_ref032","unstructured":"Istv\u00e1n Nagy T. , and Vincze V. 2014. Vpctagger: Detecting verb-particle constructions with syntax-based methods. In Proceedings of the 10th Workshop on Multiword Expressions (MWE). Association for Computational Linguistics, Gothenburg, Sweden, 17\u201325."},{"key":"S1351324918000268_ref031","unstructured":"Martens S. 2010. Varro: an algorithm and toolkit for regular structure discovery in treebanks. In Proceedings of the Coling 2010: Posters. Coling 2010 Organizing Committee, Beijing, China, 810\u20138."},{"key":"S1351324918000268_ref030","first-page":"313","article-title":"Building a large annotated corpus of english: The penn treebank","volume":"19","author":"Marcus","year":"1993","journal-title":"Computational Linguistics"},{"key":"S1351324918000268_ref029","doi-asserted-by":"crossref","unstructured":"Manning C. D. , Surdeanu M. , Bauer J. , Finkel J. , Bethard S. J. , and McClosky D. 2014. The Stanford CoreNLP natural language processing toolkit. In Proceedings of the Association for Computational Linguistics System Demonstrations, 55\u201360.","DOI":"10.3115\/v1\/P14-5010"},{"key":"S1351324918000268_ref028","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511809071"},{"key":"S1351324918000268_ref027","unstructured":"Macqueen J. 1967. Some methods for classification and analysis of multivariate observations. In Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability, 281\u2013297."},{"key":"S1351324918000268_ref026","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-009-0252-9"},{"key":"S1351324918000268_ref024","doi-asserted-by":"crossref","unstructured":"Klein D. and Manning C. D. 2004. Corpus-based induction of syntactic structure: Models of dependency and constituency. In Proceedings of the 42nd Annual Meeting on Association for Computational Linguistics.","DOI":"10.3115\/1218955.1219016"},{"key":"S1351324918000268_ref020","unstructured":"Huang Z. , Xu W. , and Yu K. 2015. Bidirectional LSTM-CRF models for sequence tagging. CoRR abs\/1508.01991."},{"key":"S1351324918000268_ref015","volume-title":"Collins COBUILD Grammar Patterns 1: Verbs","author":"Francis","year":"1996"},{"key":"S1351324918000268_ref014","doi-asserted-by":"publisher","DOI":"10.1016\/j.ic.2008.12.008"},{"key":"S1351324918000268_ref013","unstructured":"Fazly A. 2006. Automatically constructing a lexicon of verb phrase idiomatic combinations. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL-2006), 337\u2013344."},{"key":"S1351324918000268_ref012","doi-asserted-by":"crossref","unstructured":"Duan J. , Lu R. , Wu W. , Hu Y. , and Tian Y. 2006. A bio-inspired approach for multi-word expression extraction. In Proceedings of the COLING\/ACL on Main Conference Poster Sessions (COLING-ACL-2006). Association for Computational Linguistics, Stroudsburg, PA, USA, 176\u201382.","DOI":"10.3115\/1273073.1273096"},{"key":"S1351324918000268_ref011","unstructured":"Denzumi S. , Yoshinaka R. , Minato S.-I. , and Arimura H. 2011. Efficient algorithms on sequence binary decision diagrams for manipulating sets of strings. Technical Report, DCS, Hokkaido U., TCS-TR-A-11-53."},{"key":"S1351324918000268_ref009","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9"},{"key":"S1351324918000268_ref008","first-page":"59","article-title":"Alignment-based extraction of multiword expressions","volume":"44","author":"Caseli","year":"2010","journal-title":"Language Resources and Evaluation: Special Issue on Multiword Expression: Hard Going or Plain Sailing"},{"key":"S1351324918000268_ref005","volume-title":"Introduction to Algorithms","author":"Cormen","year":"1990"},{"key":"S1351324918000268_ref004","doi-asserted-by":"publisher","DOI":"10.1109\/TC.1986.1676819"},{"key":"S1351324918000268_ref003","volume-title":"Handbook of Natural Language Processing","author":"Baldwin","year":"2010"},{"key":"S1351324918000268_ref002","unstructured":"Baldwin T. , Cook P. , Lui M. , Mackinlay A. , and Wang L. 2013. How noisy social media text, how diffrnt social media sources. In Proceedings of the International Joint Conference on Natural Language Processing (IJCNLP-2013). Mark: John Blitzer."},{"key":"S1351324918000268_ref001","doi-asserted-by":"publisher","DOI":"10.1007\/978-94-011-3474-3_10"},{"key":"S1351324918000268_ref019","doi-asserted-by":"crossref","unstructured":"Headden W. P. III , Johnson M. , and McClosky D. 2009. Improving unsupervised dependency parsing with richer contexts and smoothing. In Proceedings of Human Language Technologies: The 2009 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-2009). Association for Computational Linguistics, 101\u2013109.","DOI":"10.3115\/1620754.1620769"},{"key":"S1351324918000268_ref021","doi-asserted-by":"publisher","DOI":"10.1075\/scl.4"},{"key":"S1351324918000268_ref035","doi-asserted-by":"crossref","unstructured":"Sangati F. , and van Cranenburgh A. 2015. Multiword expression identification with recurring tree fragments and association measures. In Proceedings of the MWE@ NAACL-HLT, 10\u201318.","DOI":"10.3115\/v1\/W15-0902"},{"key":"S1351324918000268_ref034","doi-asserted-by":"crossref","unstructured":"Sag I. A. , Baldwin T. , Bond F. , Copestake A. , and Flickinger D. 2002. Multiword expressions: a pain in the neck for NLP. In Proceedings of the International Conference on Intelligent Text Processing and Computational Linguistics. Springer, 1\u201315.","DOI":"10.1007\/3-540-45715-1_1"},{"key":"S1351324918000268_ref022","first-page":"612","article-title":"Text: Automatic template extraction from heterogeneous web pages","volume":"23","author":"Kim","year":"2010","journal-title":"IEEE Computer Society"},{"key":"S1351324918000268_ref010","doi-asserted-by":"publisher","DOI":"10.1016\/j.dam.2014.11.022"},{"key":"S1351324918000268_ref023","doi-asserted-by":"publisher","DOI":"10.5715\/jnlp.1.21"},{"key":"S1351324918000268_ref018","doi-asserted-by":"crossref","unstructured":"Han J. , Pei J. , and Yin Y. 2000. Mining frequent patterns without candidate generation. In Proceedings of the ACM SIGMOD International Conference on Management of Data (SIGMOD-2000). ACM, 1\u201312.","DOI":"10.1145\/342009.335372"},{"key":"S1351324918000268_ref038","unstructured":"Zarrie\u00df S. and Kuhn J. 2009. Exploiting translational correspondences for pattern-independent MWE identification. In Proceedings of the Workshop on Multiword Expressions: Identification, Interpretation, Disambiguation and Applications (MWE-2009). Association for Computational Linguistics, Stroudsburg, PA, USA, 23\u201330."},{"key":"S1351324918000268_ref016","volume-title":"Collins COBUILD Grammar Patterns 2: Nouns and Adjectives","author":"Francis","year":"1998"},{"key":"S1351324918000268_ref017","unstructured":"Gimpel K. and Smith N. A. 2012. Concavity and initialization for unsupervised dependency parsing. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, 577\u2013581."},{"key":"S1351324918000268_ref025","volume-title":"The Art of Computer Programming","author":"Knuth","year":"2009"},{"key":"S1351324918000268_ref007","doi-asserted-by":"publisher","DOI":"10.1162\/089120100561601"},{"key":"S1351324918000268_ref006","doi-asserted-by":"crossref","unstructured":"Cui H. , Kan M.-Y. , and Chua T.-S. 2004. Unsupervised learning of soft patterns for generating definitions from online news. In Proceedings of the 13th International Conference on World Wide Web (WWW-2004), 90\u201399.","DOI":"10.1145\/988672.988686"},{"key":"S1351324918000268_ref037","unstructured":"Tu Y. , and Roth D. 2012. Sorting out the most confusing english phrasal verbs. In Proceedings of the 1st Joint Conference on Lexical and Computational Semantics \u2013 Volume 1: Proceedings of the Main Conference and the Shared Task, and Volume 2: Proceedings of the 6th International Workshop on Semantic Evaluation (SemEval-2012). Association for Computational Linguistics, Stroudsburg, PA, USA, pp. 65\u20139."}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324918000268","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,4,13]],"date-time":"2019-04-13T21:23:30Z","timestamp":1555190610000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324918000268\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,7,19]]},"references-count":38,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2018,9]]}},"alternative-id":["S1351324918000268"],"URL":"https:\/\/doi.org\/10.1017\/s1351324918000268","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,7,19]]}}}