{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T10:11:20Z","timestamp":1779185480971,"version":"3.51.4"},"reference-count":80,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2026,8,31]]},"abstract":"<jats:p>\n                    Headline generation, a crucial task in summarization, aims to summarize an entire article into a concise, single line. Despite the proficiency of sequence-to-sequence encoder-decoder models and transformer-based large language models (LLMs) in text generation and summarization, generating headlines that include numerals representing the numerals in the news body remains a significant challenge. Generating a numeral-aware headline requires the ability of models to solve numerical and mathematical reasoning capabilities to infer relationships between numerals in the news body. Given the challenges in numeral-aware headline generation and numerical reasoning over numerals in news bodies, this study conducts an empirical investigation of LLMs using various strategies, including\n                    <jats:italic toggle=\"yes\">Pretrained<\/jats:italic>\n                    ,\n                    <jats:italic toggle=\"yes\">Few-shot Prompting<\/jats:italic>\n                    , and\n                    <jats:italic toggle=\"yes\">Chain-of-Thought (CoT) Prompting<\/jats:italic>\n                    , for numeral-aware headline generation and numerical reasoning for headline generation. Building upon the insights gained from our empirical study on LLMs for numeral-aware headline generation and numerical reasoning, we propose two novel approaches: instruction tuning with LLMs for numeral-aware headline generation and prompt-based masked language modeling for numerical reasoning. We conducted our experiments on the NumHG dataset. We observed that our proposed method outperforms the\n                    <jats:italic toggle=\"yes\">Pretrained<\/jats:italic>\n                    ,\n                    <jats:italic toggle=\"yes\">Few-shot Prompting<\/jats:italic>\n                    , and\n                    <jats:italic toggle=\"yes\">CoT Prompting<\/jats:italic>\n                    setups of LLMs, as well as baseline models from the literature, on both numeral-aware headline generation and numerical reasoning tasks. Observations from the experimental results reveal that our proposed Prompt-based Masked Language Modeling significantly improves the performance of small and medium-sized language models on numerical reasoning tasks. We also study the robustness of our proposed models by evaluating their performance in fact-checking numerical claims and performing numerical reasoning for numeral-aware text summarization. Our findings suggest that the proposed Prompt-based Masked Language Modeling approach is also effective for numerical claim verification and numerical reasoning for numeral-aware text summarization.\n                  <\/jats:p>","DOI":"10.1145\/3746455","type":"journal-article","created":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T09:02:18Z","timestamp":1751014938000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Prompt-based Masked Language Modeling for Numerical Reasoning"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8011-1936","authenticated-orcid":false,"given":"Sujit","family":"Kumar","sequence":"first","affiliation":[{"name":"Computer Science and Engineering, Indian Institute of Technology Guwahati, Guwahati, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-2607-787X","authenticated-orcid":false,"given":"Monika","family":"Singh","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Indian Institute of Technology Guwahati, Guwahati, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-3064-1123","authenticated-orcid":false,"given":"Abhishek","family":"Ranjan","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Indian Institute of Technology Guwahati, Guwahati, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-8727-0563","authenticated-orcid":false,"family":"Tanveen","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Indian Institute of Technology Guwahati, Guwahati, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0484-2144","authenticated-orcid":false,"given":"Sanasam Ranbir","family":"Singh","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, Indian Institute of Technology Guwahati, Guwahati, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,28]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Marah Abdin Sam Ade Jacobs Ammar Ahmad Awan Jyoti Aneja Ahmed Awadallah Hany Awadalla Nguyen Bach Amit Bahree Arash Bakhtiari Harkirat Behl et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv:2404.14219. Retrieved from https:\/\/arxiv.org\/abs\/2404.14219"},{"key":"e_1_3_2_3_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et al. 2023. Gpt-4 technical report. arXiv:2303.08774. Retrieved from https:\/\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_3_2_4_2","doi-asserted-by":"crossref","first-page":"913","DOI":"10.18653\/v1\/2024.semeval-1.131","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"Bahad Sankalp","year":"2024","unstructured":"Sankalp Bahad, Yash Bhaskar, and Parameswari Krishnamurthy. 2024. Noot noot at SemEval-2024 task 7: Numerical reasoning and headline generation. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 913\u2013917."},{"key":"e_1_3_2_5_2","first-page":"454","volume-title":"Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval \u201922)","author":"Bai Wenqiang","year":"2022","unstructured":"Wenqiang Bai, Jin Wang, and Xuejie Zhang. 2022. YNU-HPCC at SemEval-2022 task 4: Finetuning pretrained language models for patronizing and condescending language detection. In Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval \u201922), 454\u2013458."},{"key":"e_1_3_2_6_2","unstructured":"Lochan Basyal and Mihir Sanghvi. 2023. Text summarization using large language models: A comparative study of MPT-7b-instruct Falcon-7b-instruct and OpenAI Chat-GPT models. arXiv:2310.10449. Retrieved from https:\/\/arxiv.org\/abs\/2310.10449"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11912"},{"key":"e_1_3_2_8_2","unstructured":"Chung-Chi Chen Hen-Hsen Huang and Hsin-Hsi Chen. 2020. NLP in FinTech Applications: Past present and future. arXiv: 2005.01320. Retrieved from https:\/\/arxiv.org\/abs\/2005.01320"},{"key":"e_1_3_2_9_2","doi-asserted-by":"crossref","first-page":"1973","DOI":"10.1145\/3340531.3412100","volume-title":"Proceedings of the 29th ACM International Conference on Information & Knowledge Management","author":"Chen Chung-Chi","year":"2020","unstructured":"Chung-Chi Chen, Hen-Hsen Huang, and Hsin-Hsi Chen. 2020. NumClaim: Investor\u2019s fine-grained claim detection. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, 1973\u20131976."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482155"},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"6307","DOI":"10.18653\/v1\/P19-1635","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Chen Chung-Chi","year":"2019","unstructured":"Chung-Chi Chen, Hen-Hsen Huang, Hiroya Takamura, and Hsin-Hsi Chen. 2019. Numeracy-600K: Learning numeracy for detecting exaggerated information in market comments. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 6307\u20136313."},{"key":"e_1_3_2_12_2","doi-asserted-by":"crossref","first-page":"1482","DOI":"10.18653\/v1\/2024.semeval-1.213","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"Chen Chung-Chi","year":"2024","unstructured":"Chung-Chi Chen, Jian-Tao Huang, Hen-Hsen Huang, Hiroya Takamura, and Hsin-Hsi Chen. 2024. Semeval-2024 task 7: Numeral-aware language understanding and generation. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 1482\u20131491."},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","first-page":"973","DOI":"10.18653\/v1\/2024.semeval-1.141","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"Chen Kaiyuan","year":"2024","unstructured":"Kaiyuan Chen, Jin Wang, and Xuejie Zhang. 2024. YNU-HPCC at SemEval-2024 task 7: Instruction fine-tuning models for numerical understanding and generation. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 973\u2013979."},{"key":"e_1_3_2_14_2","first-page":"2754","volume-title":"Proceedings of the 25th International Joint Conference on Artificial Intelligence (IJCAI \u201916)","author":"Chen Qian","year":"2016","unstructured":"Qian Chen, Xiao-Dan Zhu, Zhen-Hua Ling, Si Wei, and Hui Jiang. 2016. Distraction-based neural networks for modeling document. In Proceedings of the 25th International Joint Conference on Artificial Intelligence (IJCAI \u201916), 2754\u20132760."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1012"},{"key":"e_1_3_2_16_2","doi-asserted-by":"crossref","first-page":"3321","DOI":"10.18653\/v1\/2023.findings-emnlp.217","volume-title":"Findings of the Association for Computational Linguistics (EMNLP \u201923)","author":"Ding Zijian","year":"2023","unstructured":"Zijian Ding, Alison Smith-Renner, Wenjuan Zhang, Joel Tetreault, and Alejandro Jaimes. 2023. Harnessing the power of LLMs: Evaluating human-AI text co-creation through the lens of news headline generation. In Findings of the Association for Computational Linguistics (EMNLP \u201923), 3321\u20133339."},{"key":"e_1_3_2_17_2","unstructured":"Qingxiu Dong Lei Li Damai Dai Ce Zheng Zhiyong Wu Baobao Chang Xu Sun Jingjing Xu and Zhifang Sui. 2022. A survey for in-context learning. arXiv:2301.00234. Retrieved from https:\/\/arxiv.org\/abs\/2301.00234"},{"key":"e_1_3_2_18_2","unstructured":"Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Amy Yang Angela Fan et al. 2024. The llama 3 herd of models. arXiv:2407.21783. Retrieved from https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_2_19_2","first-page":"47","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"Fan Yuming","year":"2024","unstructured":"Yuming Fan, Dongming Yang, and Xu He. 2024. CTYUN-AI at SemEval-2024 task 7: Boosting numerical understanding with limited data through effective data alignment. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 47\u201352."},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","first-page":"1260","DOI":"10.18653\/v1\/2024.semeval-1.183","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"Gonzalez Andres","year":"2024","unstructured":"Andres Gonzalez, Md Zobaer Hossain, and Jahedul Alam Junaed. 2024. NumDecoders at SemEval-2024 task 7: FlanT5 and GPT enhanced with CoT for numerical reasoning. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 1260\u20131268."},{"key":"e_1_3_2_21_2","unstructured":"Aaron Grattafiori Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Alex Vaughan et al. 2024. The llama 3 herd of models. arXiv:2407.21783. Retrieved from https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_3_2_22_2","first-page":"940","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"He Jianglong","year":"2024","unstructured":"Jianglong He, Saiteja Tallam, Srirama Nakshathri, Navaneeth Amarnath, Pratiba Kr, and Deepak Kumar. 2024. Infrrd. ai at SemEval-2024 task 7: RAG-based end-to-end training to generate headlines and numbers. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 940\u2013951."},{"key":"e_1_3_2_23_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Hu Edward J.","year":"2022","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=nZeVKeeFYf9"},{"key":"e_1_3_2_24_2","first-page":"12323","volume-title":"Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING \u201924)","author":"Huang Jian-Tao","year":"2024","unstructured":"Jian-Tao Huang, Chung-Chi Chen, Hen-Hsen Huang, and Hsin-Hsi Chen. 2024. NumHG: A dataset for number-focused headline generation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING \u201924), 12323\u201312329."},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","first-page":"2586","DOI":"10.18653\/v1\/2020.findings-emnlp.235","volume-title":"Findings of the Association for Computational Linguistics (EMNLP \u201920)","author":"Jiang Chengyue","year":"2020","unstructured":"Chengyue Jiang, Zhonglin Nian, Kaihao Guo, Shanbo Chu, Yinggong Zhao, Libin Shen, and Kewei Tu. 2020. Learning numeral embedding. In Findings of the Association for Computational Linguistics (EMNLP \u201920), 2586\u20132599."},{"key":"e_1_3_2_26_2","unstructured":"Vibhu Mittal Mark Kantrowitz Jade Goldstein and Jaime Carbonell. 1999. Selecting text spans for document summaries: Heuristics and metrics. In Proceedings of the AAAI\/IAAI 467\u2013473."},{"key":"e_1_3_2_27_2","first-page":"703","volume-title":"Proceedings of the Seventeenth National Conference on Artificial Intelligence and Twelfth Conference on Innovative Applications of Artificial Intelligence (AAAI\/IAAI)","author":"Knight Kevin","year":"2000","unstructured":"Kevin Knight and Daniel Marcu. 2000. Statistics-based summarization-step one: Sentence compression. In Proceedings of the Seventeenth National Conference on Artificial Intelligence and Twelfth Conference on Innovative Applications of Artificial Intelligence (AAAI\/IAAI), 703\u2013710."},{"key":"e_1_3_2_28_2","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1016\/j.patrec.2025.04.008","article-title":"Fake news detection using hashtag context","volume":"193","author":"Kumar Sujit","year":"2025","unstructured":"Sujit Kumar, Shifali Agrahari, Priyank Soni, Aayush Sachdeva, and Sanasam Ranbir Singh. 2025. Fake news detection using hashtag context. Pattern Recognition Letters 193 (2025), 43\u201349.","journal-title":"Pattern Recognition Letters"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSS.2023.3247445"},{"key":"e_1_3_2_30_2","first-page":"967","volume-title":"Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing","volume":"1","author":"Kumar Sujit","year":"2022","unstructured":"Sujit Kumar, Gaurav Kumar, and Sanasam Ranbir Singh. 2022. Detecting incongruent news articles using multi-head attention dual summarization. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing. Long Papers, Vol. 1, 967\u2013977."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSS.2024.3384698"},{"key":"e_1_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Sujit Kumar Anshul Sharma Siddharth Hemant Khincha Gargi Shroff Sanasam Ranbir Singh and Rahul Mishra. 2025. Sciclaimhunt: A large dataset for evidence-based scientific claim verification. arXiv:2502.10003. Retrieved from https:\/\/arxiv.org\/abs\/2502.10003","DOI":"10.1109\/IJCNN64981.2025.11227296"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_2_34_2","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74\u201381. Retrieved from https:\/\/www.aclweb.org\/anthology\/W04-1013"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1106"},{"key":"e_1_3_2_36_2","unstructured":"Ilya Loshchilov and Frank Hutter. 2017. Fixing weight decay regularization in adam. arXiv:1711.05101. Retrieved from https:\/\/arxiv.org\/abs\/1711.05101"},{"key":"e_1_3_2_37_2","first-page":"6493","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Luo Yutao","year":"2022","unstructured":"Yutao Luo, Menghua Lu, Gongshen Liu, and Shilin Wang. 2022. Few-shot table-to-text generation with prefix-controlled generator. In Proceedings of the 29th International Conference on Computational Linguistics, 6493\u20136504."},{"key":"e_1_3_2_38_2","unstructured":"Kathleen McKeown Judith L. Klavans Vasileios Hatzivassiloglou Regina Barzilay and Eleazar Eskin. 1999. Towards multidocument summarization by reformulation: Progress and prospects. In Proceedings of the Sixteenth National Conference on Artificial Intelligence and the Eleventh Innovative Applications of Artificial Intelligence Conference Innovative Applications of Artificial Intelligence 453\u2013460."},{"key":"e_1_3_2_39_2","first-page":"404","volume-title":"Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing","author":"Mihalcea Rada","year":"2004","unstructured":"Rada Mihalcea and Paul Tarau. 2004. Textrank: Bringing order into text. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, 404\u2013411."},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"1294","DOI":"10.1145\/3459637.3482376","volume-title":"Proceedings of the 30th ACM International Conference on Information & Knowledge Management","author":"Mishra Rahul","year":"2021","unstructured":"Rahul Mishra and Shuo Zhang. 2021. POSHAN: Cardinal POS pattern guided attention for news headline incongruence. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, 1294\u20131303."},{"key":"e_1_3_2_41_2","first-page":"3505","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","volume":"1","author":"Mishra Swaroop","year":"2022","unstructured":"Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Sachdeva, Peter Clark, Chitta Baral, and Ashwin Kalyan. 2022. NumGLUE: A suite of fundamental yet challenging mathematical reasoning tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Long Papers, Vol. 1, 3505\u20133523."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1329"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/K16-1028"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1206"},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","first-page":"218","DOI":"10.18653\/v1\/2024.semeval-1.34","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"Qian Zhen","year":"2024","unstructured":"Zhen Qian, Xiaofei Xu, and Xiuzhen Jenny Zhang. 2024. ZXQ at SemEval-2024 task 7: Fine-tuning GPT-3.5-Turbo for numerical reasoning. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 218\u2013223."},{"issue":"3","key":"e_1_3_2_46_2","first-page":"469","article-title":"Generating natural language summaries from multiple on-line sources","volume":"24","author":"Radev Dragomir","year":"1998","unstructured":"Dragomir Radev and Kathleen Mckeown. 1998. Generating natural language summaries from multiple on-line sources. Computational Linguistics 24, 3 (1998), 469\u2013500.","journal-title":"Computational Linguistics"},{"key":"e_1_3_2_47_2","unstructured":"Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI Technical Report."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","first-page":"716","DOI":"10.18653\/v1\/2024.semeval-1.103","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"Rajpoot Pawan","year":"2024","unstructured":"Pawan Rajpoot and Nut Chukamphaeng. 2024. Team NP_PROBLEM at SemEval-2024 task 7: Numerical reasoning in headline generation with preference optimization. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 716\u2013720."},{"key":"e_1_3_2_50_2","volume-title":"Proceedings of the 4th International Conference on Learning Representations (ICLR \u201916)","author":"Ranzato Marc\u2019Aurelio","year":"2016","unstructured":"Marc\u2019Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2016. Sequence level training with recurrent neural networks. In Proceedings of the 4th International Conference on Learning Representations (ICLR \u201916)."},{"key":"e_1_3_2_51_2","doi-asserted-by":"crossref","first-page":"349","DOI":"10.18653\/v1\/K19-1033","volume-title":"Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL)","author":"Ravichander Abhilasha","year":"2019","unstructured":"Abhilasha Ravichander, Aakanksha Naik, Carolyn Rose, and Eduard Hovy. 2019. EQUATE: A benchmark evaluation framework for quantitative reasoning in natural language inference. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), 349\u2013361."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3411763.3451760"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00313"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1044"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.32"},{"key":"e_1_3_2_56_2","doi-asserted-by":"crossref","first-page":"170","DOI":"10.18653\/v1\/2024.fever-1.21","volume-title":"Proceedings of the 7th Fact Extraction and VERification Workshop (FEVER)","author":"Shah Agam","year":"2024","unstructured":"Agam Shah, Arnav Hiray, Pratvi Shah, Arkaprabha Banerjee, Anushka Singh, Dheeraj Deepak Eidnani, Sahasra Chava, Bhaskar Chaudhury, and Sudheer Chava. 2024. Numerical claim detection in finance: A new financial dataset, weak-supervision model, and market analysis. In Proceedings of the 7th Fact Extraction and VERification Workshop (FEVER), 170\u2013185."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3448738"},{"key":"e_1_3_2_58_2","doi-asserted-by":"crossref","first-page":"1719","DOI":"10.18653\/v1\/2024.semeval-1.246","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"Singh Monika","year":"2024","unstructured":"Monika Singh, Sujit Kumar, Sanasam Ranbir Singh, et al. 2024. ClusterCore at SemEval-2024 task 7: Few shot prompting with large language models for numeral-aware headline generation. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 1719\u20131726."},{"key":"e_1_3_2_59_2","first-page":"5926","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Song Kaitao","year":"2019","unstructured":"Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2019. MASS: Masked sequence to sequence pre-training for language generation. In Proceedings of the International Conference on Machine Learning. PMLR, 5926\u20135936."},{"key":"e_1_3_2_60_2","first-page":"2104","volume-title":"Proceedings of the Conference on 56th Annual Meeting of the Association for Computational Linguistics (ACL \u201918)","volume":"56","author":"Spithourakis G. P.","year":"2018","unstructured":"G. P. Spithourakis and S. Riedel. 2018. Numeracy for language models: Evaluating and improving their ability to predict numbers. In Proceedings of the Conference on 56th Annual Meeting of the Association for Computational Linguistics (ACL \u201918), Long Papers, Vol. 56, Association for Computational Linguistics, 2104\u20132115."},{"key":"e_1_3_2_61_2","first-page":"779","volume-title":"Proceedings of the 20th International Conference on Natural Language Processing (ICON)","author":"Sujit Kumar","year":"2023","unstructured":"Kumar Sujit, Jaiswal Rohan, Sharma Mohit Ram, and Singh Sanasam Ranbir. 2023. Multiset dual summarization for incongruent news article detection. In Proceedings of the 20th International Conference on Natural Language Processing (ICON), 779\u2013790."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1112"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1177\/107769905303000401"},{"key":"e_1_3_2_64_2","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Yonghui Wu Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew M. Dai Anja Hauth et al. 2023. Gemini: A family of highly capable multimodal models. arXiv:2312.11805. Retrieved from https:\/\/arxiv.org\/abs\/2312.11805"},{"key":"e_1_3_2_65_2","unstructured":"Gemma Team Morgane Riviere Shreya Pathak Pier Giuseppe Sessa Cassidy Hardin Surya Bhupatiraju L\u00e9onard Hussenot Thomas Mesnard Bobak Shahriari Alexandre Ram\u00e9 et al. 2024. Gemma 2: Improving open language models at a practical size. arXiv:2408.00118. Retrieved from https:\/\/arxiv.org\/abs\/2408.00118"},{"key":"e_1_3_2_66_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288. Retrieved from https:\/\/arxiv.org\/abs\/2307.09288"},{"key":"e_1_3_2_67_2","doi-asserted-by":"crossref","unstructured":"Daphne van Zandvoort Laura Wiersema Tom Huibers Sandra van Dulmen and Sjaak Brinkkemper. 2023. Enhancing summarization performance through transformer-based prompt engineering in automated medical reporting. arXiv:2311.13274. Retrieved from https:\/\/arxiv.org\/abs\/2311.13274","DOI":"10.5220\/0012422600003657"},{"key":"e_1_3_2_68_2","first-page":"650","volume-title":"Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR \u201924)","author":"Venktesh V.","year":"2024","unstructured":"V. Venktesh, Abhijit Anand, Avishek Anand, and Vinay Setty. 2024. QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR \u201924). ACM, 650\u2013660."},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1534"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.153"},{"key":"e_1_3_2_71_2","doi-asserted-by":"crossref","first-page":"2153","DOI":"10.18653\/v1\/P19-1207","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Wang Kai","year":"2019","unstructured":"Kai Wang, Xiaojun Quan, and Rui Wang. 2019. BiSET: Bi-directional selective encoding with template for abstractive summarization. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2153\u20132162."},{"key":"e_1_3_2_72_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Wang Xuezhi","unstructured":"Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed. H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In Proceedings of the 11th International Conference on Learning Representations."},{"key":"e_1_3_2_73_2","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei Jason","year":"2022","unstructured":"Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, Vol. 35, 24824\u201324837.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_74_2","first-page":"11328","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Zhang Jingqing","year":"2020","unstructured":"Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the International Conference on Machine Learning. PMLR, 11328\u201311339."},{"key":"e_1_3_2_75_2","first-page":"7132226","article-title":"A comprehensive survey of abstractive text summarization based on deep learning","volume":"2022","author":"Zhang Mengli","year":"2022","unstructured":"Mengli Zhang, Gang Zhou, Wanting Yu, Ningbo Huang, and Wenfen Liu. 2022. A comprehensive survey of abstractive text summarization based on deep learning. Computational Intelligence and Neuroscience 2022 (2022), 7132226.","journal-title":"Computational Intelligence and Neuroscience"},{"key":"e_1_3_2_76_2","unstructured":"Shengyu Zhang Linfeng Dong Xiaoya Li Sen Zhang Xiaofei Sun Shuhe Wang Jiwei Li Runyi Hu Tianwei Zhang Fei Wu et al. 2023. Instruction tuning for large language models: A survey. arXiv:2308.10792. Retrieved from https:\/\/arxiv.org\/abs\/2308.10792"},{"key":"e_1_3_2_77_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zhang Tianyi","year":"2020","unstructured":"Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with BERT. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=SkeHuCVFDr"},{"key":"e_1_3_2_78_2","doi-asserted-by":"crossref","first-page":"261","DOI":"10.18653\/v1\/2024.semeval-1.40","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"Zhao Junzhe","year":"2024","unstructured":"Junzhe Zhao, Yingxi Wang, Huizhi Liang, and Nicolay Rusnachenko. 2024. NCL_NLP at SemEval-2024 task 7: CoT-NumHG: a CoT-based SFT training strategy with large language models for number-focused headline generation. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 261\u2013269."},{"key":"e_1_3_2_79_2","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Zhao Wei","year":"2019","unstructured":"Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019. MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP)."},{"key":"e_1_3_2_80_2","first-page":"1095","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics","volume":"1","author":"Zhou Qingyu","year":"2017","unstructured":"Qingyu Zhou, Nan Yang, Furu Wei, and Ming Zhou. 2017. Selective encoding for abstractive sentence summarization. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Long Papers, Vol. 1, 1095\u20131104."},{"key":"e_1_3_2_81_2","doi-asserted-by":"crossref","first-page":"1659","DOI":"10.18653\/v1\/2024.semeval-1.236","volume-title":"Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924)","author":"Zhunis Ali","year":"2024","unstructured":"Ali Zhunis and Hao-Yun Chuang. 2024. Challenges at SemEval 2024 task 7: Contrastive learning approach on numeral-aware language generation. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval \u201924), 1659\u20131662."}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746455","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T14:42:05Z","timestamp":1777387325000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746455"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,28]]},"references-count":80,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,8,31]]}},"alternative-id":["10.1145\/3746455"],"URL":"https:\/\/doi.org\/10.1145\/3746455","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,28]]},"assertion":[{"value":"2024-03-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-12","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}