{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,4,25]],"date-time":"2025-04-25T04:10:15Z","timestamp":1745554215948},"reference-count":8,"publisher":"MIT Press - Journals","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computational Linguistics"],"published-print":{"date-parts":[[2011,12]]},"abstract":"<jats:p> Noun phrases (nps) are a crucial part of natural language, and can have a very complex structure. However, this np structure is largely ignored by the statistical parsing field, as the most widely used corpus is not annotated with it. This lack of gold-standard data has restricted previous efforts to parse nps, making it impossible to perform the supervised experiments that have achieved high performance in so many Natural Language Processing (nlp) tasks. <\/jats:p><jats:p> We comprehensively solve this problem by manually annotating np structure for the entire Wall Street Journal section of the Penn Treebank. The inter-annotator agreement scores that we attain dispel the belief that the task is too difficult, and demonstrate that consistent np annotation is possible. Our gold-standard np data is now available for use in all parsers. <\/jats:p><jats:p> We experiment with this new data, applying the Collins (2003) parsing model, and find that its recovery of np structure is significantly worse than its overall performance. The parser's F-score is up to 5.69% lower than a baseline that uses deterministic rules. Through much experimentation, we determine that this result is primarily caused by a lack of lexical information. <\/jats:p><jats:p> To solve this problem we construct a wide-coverage, large-scale np Bracketing system. With our Penn Treebank data set, which is orders of magnitude larger than those used previously, we build a supervised model that achieves excellent results. Our model performs at 93.8% F-score on the simple task that most previous work has undertaken, and extends to bracket longer, more complex nps that are rarely dealt with in the literature. We attain 89.14% F-score on this much more difficult task. Finally, we implement a post-processing module that brackets nps identified by the Bikel (2004) parser. Our np Bracketing model includes a wide variety of features that provide the lexical information that was missing during the parser experiments, and as a result, we outperform the parser's F-score by 9.04%. <\/jats:p><jats:p> These experiments demonstrate the utility of the corpus, and show that many nlp applications can now make use of np structure. <\/jats:p>","DOI":"10.1162\/coli_a_00076","type":"journal-article","created":{"date-parts":[[2011,7,14]],"date-time":"2011-07-14T17:39:41Z","timestamp":1310665181000},"page":"753-809","source":"Crossref","is-referenced-by-count":6,"title":["Parsing Noun Phrases in the Penn Treebank"],"prefix":"10.1162","volume":"37","author":[{"given":"David","family":"Vadas","sequence":"first","affiliation":[{"name":"University of Sydney"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"James R.","family":"Curran","sequence":"additional","affiliation":[{"name":"University of Sydney"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","reference":[{"key":"p_2","doi-asserted-by":"publisher","DOI":"10.1162\/coli.2007.33.4.469"},{"key":"p_20","doi-asserted-by":"publisher","DOI":"10.1162\/089120103322753356"},{"issue":"4","key":"p_26","first-page":"313","volume":"19","author":"Girju Roxana","year":"2005","journal-title":"Journal of Computer Speech and Language - Special Issue on Multiword Expressions"},{"issue":"1","key":"p_30","first-page":"103","volume":"19","author":"Hindle Donald","year":"1993","journal-title":"Computational Linguistics"},{"issue":"4","key":"p_32","first-page":"613","volume":"24","author":"Johnson Mark","year":"1998","journal-title":"Computational Linguistics"},{"issue":"2","key":"p_45","first-page":"313","volume":"19","author":"Marcus Mitchell","year":"1993","journal-title":"Computational Linguistics"},{"key":"p_60","doi-asserted-by":"publisher","DOI":"10.1017\/S0022226705003713"},{"key":"p_64","doi-asserted-by":"publisher","DOI":"10.1016\/S0019-9958(67)80007-X"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mitpressjournals.org\/doi\/pdf\/10.1162\/COLI_a_00076","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,3,12]],"date-time":"2021-03-12T21:27:06Z","timestamp":1615584426000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/37\/4\/753-809\/2122"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,12]]},"references-count":8,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2011,12]]}},"alternative-id":["10.1162\/COLI_a_00076"],"URL":"https:\/\/doi.org\/10.1162\/coli_a_00076","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,12]]}}}