{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T04:39:46Z","timestamp":1759207186473,"version":"3.41.0"},"reference-count":72,"publisher":"MIT Press","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computational Linguistics"],"published-print":{"date-parts":[[2017,6]]},"abstract":"<jats:p>We propose a question answering (QA) approach for standardized science exams that both identifies correct answers and produces compelling human-readable justifications for why those answers are correct. Our method first identifies the actual information needed in a question using psycholinguistic concreteness norms, then uses this information need to construct answer justifications by aggregating multiple sentences from different knowledge bases using syntactic and lexical information. We then jointly rank answers and their justifications using a reranking perceptron that treats justification quality as a latent variable. We evaluate our method on 1,000 multiple-choice questions from elementary school science exams, and empirically demonstrate that it performs better than several strong baselines, including neural network approaches. Our best configuration answers 44% of the questions correctly, where the top justifications for 57% of these correct answers contain a compelling human-readable justification that explains the inference required to arrive at the correct answer. We include a detailed characterization of the justification quality for both our method and a strong baseline, and show that information aggregation is key to addressing the information need in complex questions.<\/jats:p>","DOI":"10.1162\/coli_a_00287","type":"journal-article","created":{"date-parts":[[2017,3,28]],"date-time":"2017-03-28T19:43:24Z","timestamp":1490730204000},"page":"407-449","source":"Crossref","is-referenced-by-count":12,"title":["Framing QA as Building and Ranking Intersentence Answer Justifications"],"prefix":"10.1162","volume":"43","author":[{"given":"Peter","family":"Jansen","sequence":"first","affiliation":[{"name":"University of Arizona"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rebecca","family":"Sharp","sequence":"additional","affiliation":[{"name":"University of Arizona"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mihai","family":"Surdeanu","sequence":"additional","affiliation":[{"name":"University of Arizona"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peter","family":"Clark","sequence":"additional","affiliation":[{"name":"Allen Institute for Artificial Intelligence"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","reference":[{"key":"bib1","doi-asserted-by":"publisher","DOI":"10.1016\/S1574-6526(07)03020-9"},{"key":"bib2","unstructured":"Banarescu, Laura, Claire Bonial, Shu Cai, Madalina Georgescu, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Martha Palmer, and Nathan Schneider. 2013. Abstract meaning representation for sembanking. In Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse, pages 178\u2013186."},{"key":"bib3","unstructured":"Baral, Chitta and Shanshan Liang. 2012. From knowledge represented in frame-based languages to declarative representation and reasoning via asp. In KR."},{"key":"bib5","unstructured":"Baral, Chitta, Nguyen Ha Vo, and Shanshan Liang. 2012. Answering why and how questions with respect to a frame-based knowledge base: a preliminary report. In ICLP (Technical Communications), pages 26\u201336."},{"key":"bib6","doi-asserted-by":"publisher","DOI":"10.1162\/coli.2008.34.1.1"},{"key":"bib7","doi-asserted-by":"publisher","DOI":"10.1162\/089120105774321091"},{"key":"bib8","doi-asserted-by":"publisher","DOI":"10.3115\/1034678.1034760"},{"key":"bib9","doi-asserted-by":"crossref","unstructured":"Berger, Adam, Rich Caruana, David Cohn, Dayne Freytag, and Vibhu Mittal. 2000. Bridging the lexical chasm: Statistical approaches to answer finding. In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research & Development on Information Retrieval, Athens, Greece.","DOI":"10.1145\/345508.345576"},{"key":"bib10","doi-asserted-by":"crossref","unstructured":"Bj\u00f6rkelund, Anders and Jonas Kuhn. 2014. Learning structured perceptrons for coreference resolution with latent antecedents and non-local features. In Proceedings of the Association for Computational Linguistics.","DOI":"10.3115\/v1\/P14-1005"},{"key":"bib12","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1067"},{"key":"bib14","doi-asserted-by":"publisher","DOI":"10.3758\/s13428-013-0403-5"},{"key":"bib15","doi-asserted-by":"crossref","unstructured":"Chen, Danqi, Jason Bolton, and Christopher D. Manning. 2016. A thorough examination of the CNN \/ Daily Mail reading comprehension task. In Proceedings of the Association for Computational Linguistics (ACL).","DOI":"10.18653\/v1\/P16-1223"},{"key":"bib16","unstructured":"Chen, Danqi and Christopher D. Manning. 2014. A fast and accurate dependency parser using neural networks. In Empirical MNLP, pages 740\u2013750."},{"key":"bib17","doi-asserted-by":"crossref","unstructured":"Clark, Peter. 2015, Elementary school science and math tests as a driver for AI: take the aristo challenge! In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, pages 4019\u20134021. Austin, TX.","DOI":"10.1609\/aaai.v29i2.19066"},{"key":"bib18","doi-asserted-by":"publisher","DOI":"10.1145\/2509558.2509565"},{"key":"bib19","doi-asserted-by":"crossref","unstructured":"Collins, Michael. 2002, Discriminative training methods for hidden Markov models: Theory and experiments with perceptron algorithms. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), EMNLP '02, pages 1\u20138, Stroudsburg, PA.","DOI":"10.3115\/1118693.1118694"},{"key":"bib20","unstructured":"Collobert, Ronan, Jason Weston, L\u00e9on Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. Natural language processing (almost) from scratch. Journal of Machine Learning Research, 2493\u20132537."},{"key":"bib21","unstructured":"Daum\u00e9 III, Hal. 2007. Frustratingly easy domain adaptation. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics, pages 256\u2013263, Prague."},{"key":"bib23","unstructured":"Dong, Li, Furu Wei, Ming Zhou, and Ke Xu. 2015. Question answering over freebase with multi-column convolutional neural networks. In Proceedings of the Association for Computational Linguistics, pages 260\u2013269."},{"key":"bib24","unstructured":"Echihabi, Abdessamad and Daniel Marcu. 2003. A noisy-channel approach to question answering. In Proceedings of the 41st Annual Meeting of the Association for Computational Linguistics-Volume 1, pages 16\u201323."},{"key":"bib25","unstructured":"Fernandes, Eraldo Rezende, C\u00edcero Nogueira Dos Santos, and Ruy Luiz Milidi\u00fa. 2012. Latent structure perceptron with feature induction for unrestricted coreference resolution. In Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL)\u2014Shared Task, pages 41\u201348."},{"key":"bib26","doi-asserted-by":"crossref","unstructured":"Ferrucci, David. A. 2012. Introduction to \u201cthis is Watson.\u201dIBM Journal of Research and Development, 56(3.4).","DOI":"10.1147\/JRD.2012.2184356"},{"key":"bib27","unstructured":"Finkel, Jenny Rose and Christopher D. Manning. 2010. Hierarchical joint learning: Improving joint parsing and named entity recognition with non-jointly labeled data. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, pages 720\u2013728."},{"key":"bib28","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00133"},{"key":"bib29","doi-asserted-by":"crossref","unstructured":"Gondek, D. C., Adam Lally, Aditya Kalyanpur, J. William Murdock, Pablo Ariel Dubou\u00e9, Lei Zhang, Yue Pan, ZM Qui, and Chris Welty, 2012. A framework for merging and ranking of answers in deepqa. IBM Journal of Research and Development, 56(3.4).","DOI":"10.1147\/JRD.2012.2188760"},{"key":"bib30","unstructured":"Graff, David, Junbo Kong, Ke Chen, and Kazuaki Maeda. 2003. English gigaword, ldc2003t05. Linguistic Data Consortium, Philadelphia."},{"key":"bib31","doi-asserted-by":"crossref","unstructured":"Harabagiu, Sanda, Dan Moldovan, Marius Pasca, Rada Mihalcea, Mihai Surdeanu, Razvan Bunescu, Roxana Girju, Vasile Rus, and Paul Morarescu. 2000. Falcon: Boosting knowledge for answer engines. In Proceedings of the Text REtrieval Conference (TREC), Gaithersburg, MD.","DOI":"10.6028\/NIST.SP.500-249.SMU"},{"key":"bib32","unstructured":"He, Hua and Jimmy Lin. 2016. Pairwise word interaction modeling with deep neural networks for semantic similarity measurement. In Proceedings of NAACL-HLT, pages 937\u2013948."},{"key":"bib33","unstructured":"Hermann, Karl Moritz, Tom\u00e1\u0161 Ko\u010disk\u00fd, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems (NIPS)."},{"key":"bib34","unstructured":"Hoffmann, Raphael, Congle Zhang, Xiao Ling, Luke Zettlemoyer, and Daniel S. Weld. 2011. Knowledge-based weak supervision for information extraction of overlapping relations. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1, pages 541\u2013550."},{"key":"bib35","doi-asserted-by":"crossref","unstructured":"Iyyer, Mohit, Jordan Boyd-Graber, Leonardo Claudino, Richard Socher, and Hal Daum\u00e9 III. 2014. A neural network for factoid question answering over paragraphs. In Empirical Methods in Natural Language Processing.","DOI":"10.3115\/v1\/D14-1070"},{"key":"bib36","doi-asserted-by":"crossref","unstructured":"Iyyer, Mohit, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daum\u00e9 III. 2015. Deep unordered composition rivals syntactic methods for text classification. In Association for Computational Linguistics.","DOI":"10.3115\/v1\/P15-1162"},{"key":"bib37","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1092"},{"key":"bib38","unstructured":"Khashabi, Daniel, Tushar Khot, Ashish Sabharwal, Peter Clark, Oren Etzioni and Dan Roth. 2016. Question answering via integer programming over semi-structured knowledge. In Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence (IJCAI-16), 1145\u20131152."},{"key":"bib39","doi-asserted-by":"publisher","DOI":"10.1162\/COLI_a_00152"},{"key":"bib40","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00219"},{"key":"bib41","unstructured":"Liang, Percy, Alexandre Bouchard-C\u00f4t\u00e9, Dan Klein, and Ben Taskar. 2006, An end-to-end discriminative approach to machine translation. In Proceedings of the 21st International Conference on Computational Linguistics and the 44th Annual Meeting of the Association for Computational Linguistics, pages 761\u2013768."},{"key":"bib42","doi-asserted-by":"publisher","DOI":"10.1162\/COLI_a_00127"},{"key":"bib43","unstructured":"MacCartney, Bill. 2009. Natural Language Inference. PhD thesis."},{"key":"bib45","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-5010"},{"key":"bib46","unstructured":"McSherry, Frank and Marc Najork. 2008. Computing information retrieval performance measures efficiently in the presence of tied scores. In 30th European Conference on IR Research (ECIR)."},{"key":"bib47","unstructured":"Mikolov, Tomas, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"bib48","doi-asserted-by":"crossref","unstructured":"Mikolov, Tomas, Martin Karafiat, Lukas Burget, Jan Cernocky, and Sanjeev Khudanpur. 2010. Recurrent neural network based language model. In Proceedings of the 11th Annual Conference of the International Speech Communication Association (INTERSPEECH 2010).","DOI":"10.1109\/ICASSP.2011.5947611"},{"key":"bib49","doi-asserted-by":"publisher","DOI":"10.1016\/j.jal.2005.12.005"},{"key":"bib50","doi-asserted-by":"publisher","DOI":"10.3115\/1073445.1073467"},{"key":"bib51","doi-asserted-by":"publisher","DOI":"10.1145\/763693.763694"},{"key":"bib52","unstructured":"Moldovan, Dan I. and Vasile Rus. 2001. Logic form transformation of Wordnet and its applicability to question answering. In Proceedings of the 39th Annual Meeting of the Association for Computational Linguistics, pages 402\u2013409."},{"key":"bib53","doi-asserted-by":"crossref","unstructured":"Moschitti, Alessandro. 2004. A study on convolution kernels for shallow semantic parsing. In Proceedings of the 42th Annual Meeting of the Association for Computational Linguistics (ACL).","DOI":"10.3115\/1218955.1218998"},{"key":"bib54","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2010.06.002"},{"key":"bib55","unstructured":"Moschitti, Alessandro, Silvia Quarteroni, Roberto Basili, and Suresh Manandhar. 2007. Exploiting syntactic and shallow semantic kernels for question\/answer classification. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics (ACL), pages 776\u2013783, Prague."},{"key":"bib56","doi-asserted-by":"crossref","unstructured":"Murdock, J. William, James Fan, Adam Lally, Hideki Shima, and B. K. Boguraev. 2012. Textual evidence gathering and analysis. IBM Journal of Research and Development, 56(3.4).","DOI":"10.1147\/JRD.2012.2187249"},{"key":"bib57","doi-asserted-by":"crossref","unstructured":"Park, Jae Hyun and W. Bruce Croft. 2015. Using key concepts in a translation model for retrieval. In Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 927\u2013930, New York, NY.","DOI":"10.1145\/2766462.2767768"},{"key":"bib59","doi-asserted-by":"crossref","unstructured":"Pradhan, Sameer S., Valerie Krugler, Steven Bethard, Wayne Ward, Daniel Jurafsky, James H. Martin, Sasha Blair-Goldensohn, Andrew Hazen Schlaikjer, Elena Filatova, Pablo Ariel Dubou\u00e9 2002. Building a foundation system for producing short answers to factual questions. In TREC.","DOI":"10.6028\/NIST.SP.500-251.qa-columbia"},{"key":"bib60","unstructured":"Riezler, Stefan, Alexander Vasserman, Ioannis Tsochantaridis, Vibhu Mittal, and Yi Liu. 2007. Statistical machine translation for query expansion in answer retrieval. In Proceedings of the 45th Annual Meeting of the Association for Computational Linguistics (ACL), pages 464\u2013471."},{"key":"bib61","doi-asserted-by":"publisher","DOI":"10.1145\/2348283.2348383"},{"key":"bib62","doi-asserted-by":"crossref","unstructured":"Severyn, Aliaksei and Alessandro Moschitti. 2013. Automatic feature engineering for answer selection and extraction. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP).","DOI":"10.18653\/v1\/D13-1044"},{"key":"bib63","unstructured":"Severyn, Aliaksei, Massimo Nicosia, and Alessandro Moschitti. 2013. Learning adaptable patterns for passage reranking. In Proceedings of the Seventeenth Conference on Computational Natural Language Learning (CoNLL)."},{"key":"bib64","unstructured":"Sharma, Arpit, Nguyen H. Vo, Somak Aditya, and Chitta Baral. 2015. Towards addressing the winograd schema challenge-building and using a semantic parser and a knowledge hunting module. In Proceedings of the International Joint Conferences on Artificial Intelligence (IJCAI)."},{"key":"bib65","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1025"},{"key":"bib66","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-005-0918-9"},{"key":"bib67","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-006-7149-y"},{"key":"bib69","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00178"},{"key":"bib70","unstructured":"Sun, Xu, Takuya Matsuzaki, Daisuke Okanohara, and Jun'ichi Tsujii. 2009. Latent variable perceptron algorithm for structured classification. In IJCAI, volume 9, pages 1236\u20131242."},{"key":"bib71","doi-asserted-by":"publisher","DOI":"10.1162\/COLI_a_00051"},{"key":"bib72","unstructured":"Tari, Luis and Chitta Baral. 2006. Using AnsProlog with link grammar and Wordnet for QA with deep reasoning. In Proceedings of the 9th International Conference on Information Technology, pages 125\u2013128."},{"key":"bib73","doi-asserted-by":"crossref","unstructured":"Tymoshenko, Kateryna and Alessandro Moschitti. 2015. Assessing the impact of syntactic and semantic structures for answer passages reranking. In Proceedings of the 24th ACM International Conference on Information and Knowledge Management (CIKM).","DOI":"10.1145\/2806416.2806490"},{"key":"bib74","unstructured":"Voorhees, Ellen M.2003. Evaluating answers to definition questions. In Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology: Companion Volume of the Proceedings of HLT-NAACL 2003\u2013short Papers - Volume 2, NAACL-Short '03, pages 109\u2013111."},{"key":"bib75","doi-asserted-by":"crossref","unstructured":"Wang, Di and Eric Nyberg. 2015. A long short-term memory model for answer sentence selection in question answering. ACL, July.","DOI":"10.3115\/v1\/P15-2116"},{"key":"bib76","doi-asserted-by":"crossref","unstructured":"Yao, Xuchen, Benjamin Van Durme, Chris Callison-Burch, and Peter Clark. 2013. Semi-Markov phrase-based monolingual alignment. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP).","DOI":"10.18653\/v1\/D13-1056"},{"key":"bib77","doi-asserted-by":"crossref","unstructured":"Yih, Wen-tau, Ming-Wei Chang, Xiaodong He, and Jianfeng Gao. 2015. Semantic parsing via staged query graph generation: Question answering with knowledge base. In Proceedings of the Association for Computational Linguistics (ACL).","DOI":"10.3115\/v1\/P15-1128"},{"key":"bib78","unstructured":"Yih, Wen-tau, Ming-Wei Chang, Christopher Meek, and Andrzej Pastusiak. 2013. Question answering using enhanced lexical semantic models. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (ACL)."},{"key":"bib79","unstructured":"Zettlemoyer, Luke S. and Michael Collins. 2007. Online learning of relaxed ccg grammars for parsing to logical form. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), pages 678\u2013687."}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mitpressjournals.org\/doi\/pdf\/10.1162\/COLI_a_00287","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:08Z","timestamp":1750186928000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/43\/2\/407-449\/1566"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,6]]},"references-count":72,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2017,6]]}},"alternative-id":["10.1162\/COLI_a_00287"],"URL":"https:\/\/doi.org\/10.1162\/coli_a_00287","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"type":"print","value":"0891-2017"},{"type":"electronic","value":"1530-9312"}],"subject":[],"published":{"date-parts":[[2017,6]]}}}