{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:09:15Z","timestamp":1750306155957,"version":"3.41.0"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2016,11,18]],"date-time":"2016-11-18T00:00:00Z","timestamp":1479427200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2017,3,31]]},"abstract":"<jats:p>Summary of a document contains words that actually contribute to the semantics of the document. Latent Semantic Analysis (LSA) is a mathematical model that is used to understand document semantics by deriving a semantic structure based on patterns of word correlations in the document. When using LSA to capture semantics from summaries, it is observed that LSA performs quite well despite being completely independent of any external sources of semantics. However, LSA can be remodeled to enhance its capability to analyze correlations within texts. By taking advantage of the model being language independent, this article presents two stages of LSA remodeling to understand document semantics in the Indian context, specifically from Hindi text summaries. One stage of remodeling is done by providing supplementary information, such as document category and domain information. The second stage of remodeling is done by using a supervised term weighting measure in the process. The remodeled LSA\u2019s performance is empirically evaluated in a document classification application by comparing the accuracies of classification to plain LSA. An improvement in the performance of LSA in the range of 4.7% to 6.2% is achieved from the remodel when compared to the plain model. The results suggest that summaries of documents efficiently capture the semantic structure of documents and is an alternative to full-length documents for understanding document semantics.<\/jats:p>","DOI":"10.1145\/2956236","type":"journal-article","created":{"date-parts":[[2016,11,18]],"date-time":"2016-11-18T15:40:54Z","timestamp":1479483654000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Understanding Document Semantics from Summaries"],"prefix":"10.1145","volume":"16","author":[{"given":"Karthik","family":"Krishnamurthi","sequence":"first","affiliation":[{"name":"Christ University, Bangalore, Karnataka, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vijayapal Reddy","family":"Panuganti","sequence":"additional","affiliation":[{"name":"GRIET, Hyderabad, Telangana, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vishnu Vardhan","family":"Bulusu","sequence":"additional","affiliation":[{"name":"JNTUHCEJ, Karimnagar, Telangana, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,11,18]]},"reference":[{"volume-title":"Retrieved","year":"2005","author":"Baker K.","key":"e_1_2_1_1_1"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1137\/1037127"},{"volume-title":"Report ORNL\/TM-13756. Computer Science and Mathematics Division, Oak Ridge National Laboratory.","year":"1999","author":"Chisholm E.","key":"e_1_2_1_3_1"},{"volume-title":"Proceedings of the Korea Information Science Society Conference (KISS\u201997)","year":"1997","author":"Cho K.","key":"e_1_2_1_4_1"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/952532.952688"},{"key":"e_1_2_1_6_1","first-page":"391","article-title":"Indexing by latent semantic analysis","volume":"4","author":"Deerwester S.","year":"1990","journal-title":"Journal of the Association for Information Science and Technology"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5087\/dad.2010.002"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526737"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.3115\/1119355.1119383"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.3115\/1220175.1220243"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.3115\/1118894.1118902"},{"key":"e_1_2_1_13_1","first-page":"540","article-title":"Web based classification of Tamil documents using ABPA","volume":"3","author":"Kanimozhi S.","year":"2012","journal-title":"International Journal of Scientific and Engineering Research"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.3115\/1117755.1117766"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.3115\/990820.990886"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0306-4573(02)00056-0"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/IALP.2013.11"},{"volume-title":"Proceedings of theInternational Oriental Conference on Asian Spoken Language Research and Evaluation. 1--6.","author":"Kumar P.","key":"e_1_2_1_19_1"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2008.110"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1080\/01638539809545028"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.4249\/scholarpedia.4356"},{"key":"e_1_2_1_23_1","unstructured":"R. Larson and M. Davis. 2002. Information Organization and Retrieval. SIMS Lecture 18: Vector Representation University of California at Berkeley.  R. Larson and M. Davis. 2002. Information Organization and Retrieval. SIMS Lecture 18: Vector Representation University of California at Berkeley."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1012491419635"},{"key":"e_1_2_1_25_1","first-page":"62","article-title":"Using class frequency for improving centroid-based text classification","volume":"2","author":"Lertnattee V.","year":"2012","journal-title":"ACEEE International Journal on Information Technology"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2007.10.042"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(89)90100-3"},{"volume-title":"Proceedings of the International Conference on Computational Linguistics. 537--544","author":"Malik M. G. A.","key":"e_1_2_1_28_1"},{"volume-title":"Proceedings of the 9th International Conference on Computer and Information Technology. 1--8.","author":"Mansur M.","key":"e_1_2_1_29_1"},{"volume-title":"Proceedings of the Conference on Recent Advances in Natural Language Processing. 150--160","author":"Mihalcea R.","key":"e_1_2_1_30_1"},{"volume-title":"Proceedings of the International Symposium on Linguistics, Quantification, and Computation (CALTS\u201905)","year":"2005","author":"Murthy K. N.","key":"e_1_2_1_31_1"},{"volume-title":"Proceedings of the Workshop on South and Southeast Asian Natural Language Processing. 109--122","year":"2012","author":"Nidhi V. G.","key":"e_1_2_1_32_1"},{"volume-title":"Proceedings of the Association for Computational Linguistics Student Research Workshop. 46--51","author":"Ogura Y.","key":"e_1_2_1_33_1"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSDA.2013.6709861"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.154"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-10847-1_35"},{"volume-title":"Proceedings of the Workshop of Computational Linguistics for South Asian Languages Expanding Synergies with Europe. 42--48","year":"2003","author":"Ramanathan A.","key":"e_1_2_1_37_1"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/IALP.2011.66"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICMLA.2006.50"},{"volume-title":"Proceedings of the International Conference on Applied Computer Science. 573--578","author":"Rishel T.","key":"e_1_2_1_40_1"},{"key":"e_1_2_1_41_1","first-page":"129","article-title":"Relevance weighting of search terms","volume":"27","author":"Robertson S. E.","year":"1976","journal-title":"Journal of the Association for Information Science and Technology"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(88)90021-0"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/361219.361220"},{"volume-title":"Proceedings of the 6th International Global Wordnet Conference. 324--329","author":"Sarmah J.","key":"e_1_2_1_44_1"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073483.1073515"},{"key":"e_1_2_1_46_1","unstructured":"A. Singal and G. Salton. 1995. Pivoted Document Length Normalization Technical Report TR95-1560. Cornell University Ithaca NY.   A. Singal and G. Salton. 1995. Pivoted Document Length Normalization Technical Report TR95-1560. Cornell University Ithaca NY."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-44949-9_23"},{"volume-title":"Proceedings of the International Conference on Artificial Intelligence. 1130--1135","author":"Soucy P.","key":"e_1_2_1_48_1"},{"key":"e_1_2_1_49_1","first-page":"9","article-title":"A model for overlapping trigram technique for Telugu script","volume":"3","author":"Vardhan B. V.","year":"2007","journal-title":"Journal of Theoretical and Applied Information Technology"},{"volume-title":"Proceedings of the International Conference on Advanced Computing Technologies. 1--5.","author":"Vispute S. R.","key":"e_1_2_1_50_1"},{"key":"e_1_2_1_51_1","first-page":"209","article-title":"Inverse-category-frequency based supervised term weighting schemes for text categorization","volume":"29","author":"Wang D.","year":"2013","journal-title":"Journal of Information Science and Engineering"},{"volume-title":"Proceedings of the Annual Conference of the Cognitive Science Society. 1112--1117","author":"Wiemer-Hastings P.","key":"e_1_2_1_52_1"},{"volume":"244","volume-title":"Knowledge and Systems Engineering. Advances in Intelligence Systems and Computing","author":"Xuan N. P.","key":"e_1_2_1_53_1"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/502585.502605"},{"volume-title":"Proceedings of the Symposium on Document Image Understanding Technology. 87--91","author":"Zukas A.","key":"e_1_2_1_55_1"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2956236","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2956236","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:39:43Z","timestamp":1750217983000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2956236"}},"subtitle":["A Case Study on Hindi Texts"],"short-title":[],"issued":{"date-parts":[[2016,11,18]]},"references-count":53,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2017,3,31]]}},"alternative-id":["10.1145\/2956236"],"URL":"https:\/\/doi.org\/10.1145\/2956236","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2016,11,18]]},"assertion":[{"value":"2014-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-11-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}