{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T20:44:32Z","timestamp":1779223472316,"version":"3.51.4"},"reference-count":47,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2025,11,24]],"date-time":"2025-11-24T00:00:00Z","timestamp":1763942400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100004770","name":"University of Parma","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004770","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100009879","name":"Regione Emilia-Romagna","doi-asserted-by":"crossref","award":["PR FSE+ 2021-2027"],"award-info":[{"award-number":["PR FSE+ 2021-2027"]}],"id":[{"id":"10.13039\/501100009879","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Retrieval-augmented generation (RAG) enriches prompts with external knowledge, but it often relies on additional infrastructure that may be impractical in resource-constrained or offline settings. In addition, updating the internal knowledge of a language model through retraining is costly and inflexible. To address these limitations, we propose an explainable and structured prompt augmentation pipeline that enhances inputs using pre-trained models and rule-based extractors, without requiring external sources. We describe this approach as an orchestrated LLM workflow: a structured sequence in which lightweight LLM modules assume specialized roles. Specifically, (1) an extractor module identifies factual triples from input prompts by combining dependency parsing with a rule-based extraction algorithm; (2) a scorer module, based on a generic lightweight LLM, evaluates the importance of each triple via its self-attention patterns, leveraging internal beliefs to promote explainability and trustworthy cooperation with the downstream model; (3) a performer module processes the augmented prompt for downstream tasks in supervised fine-tuning or zero-shot settings. Much like in a theater staging, each module operates transparently behind the scenes to support and elevate the performer\u2019s final output. We evaluate this approach across multiple performer architectures (encoder-only, encoder-decoder, and decoder-only) and NLP tasks (multiple-choice QA, open-book QA, and summarization). Our results show that this structured augmentation with scored facts yields consistent improvements compared to baseline prompting: up to a 28.78% accuracy improvement for multiple-choice QA, up to a 9.42% BLEURT improvement for open-book QA, and up to a 18.14% ROUGE-L improvement for summarization. By decoupling knowledge scoring from task execution, our method provides a practical, interpretable, and low-cost alternative to RAG in static or knowledge-limited environments.<\/jats:p>","DOI":"10.3390\/fi17120535","type":"journal-article","created":{"date-parts":[[2025,11,24]],"date-time":"2025-11-24T13:09:25Z","timestamp":1763989765000},"page":"535","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["LLMs in Staging: An Orchestrated LLM Workflow for Structured Augmentation with Fact Scoring"],"prefix":"10.3390","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-1907-2885","authenticated-orcid":false,"given":"Giuseppe","family":"Trimigno","sequence":"first","affiliation":[{"name":"Department of Engineering and Architecture, University of Parma, 43124 Parma, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1808-4487","authenticated-orcid":false,"given":"Gianfranco","family":"Lombardo","sequence":"additional","affiliation":[{"name":"Department of Engineering and Architecture, University of Parma, 43124 Parma, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6030-9435","authenticated-orcid":false,"given":"Michele","family":"Tomaiuolo","sequence":"additional","affiliation":[{"name":"Department of Engineering and Architecture, University of Parma, 43124 Parma, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4669-512X","authenticated-orcid":false,"given":"Stefano","family":"Cagnoni","sequence":"additional","affiliation":[{"name":"Department of Engineering and Architecture, University of Parma, 43124 Parma, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3528-0260","authenticated-orcid":false,"given":"Agostino","family":"Poggi","sequence":"additional","affiliation":[{"name":"Department of Engineering and Architecture, University of Parma, 43124 Parma, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,11,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1145\/3571730","article-title":"Survey of hallucination in natural language generation","volume":"55","author":"Ji","year":"2023","journal-title":"ACM Comput. Surv."},{"key":"ref_2","first-page":"9459","article-title":"Retrieval-augmented generation for knowledge-intensive nlp tasks","volume":"33","author":"Lewis","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Izacard, G., and Grave, E. (2021, January 19\u201323). Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. Proceedings of the EACL 2021\u201416th Conference of the European Chapter of the Association for Computational Linguistics, Kyiv, Ukraine.","DOI":"10.18653\/v1\/2021.eacl-main.74"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Shuster, K., Poff, S., Chen, M., Kiela, D., and Weston, J. (2021, January 7\u201311). Retrieval Augmentation Reduces Hallucination in Conversation. Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2021, Punta Cana, Dominican Republic.","DOI":"10.18653\/v1\/2021.findings-emnlp.320"},{"key":"ref_5","unstructured":"Zhou, D., Sch\u00e4rli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., and Le, Q.V. (2023, January 1\u20135). Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_6","unstructured":"Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M. (2020, January 13\u201318). Retrieval augmented language model pre-training. Proceedings of the International Conference on Machine Learning, PMLR, Online."},{"key":"ref_7","unstructured":"Borgeaud, S., Mensch, A., Hoffmann, J., Cai, T., Rutherford, E., Millican, K., Van Den Driessche, G.B., Lespiau, J.B., Damoc, B., and Clark, A. (2022, January 17\u201323). Improving language models by retrieving from trillions of tokens. Proceedings of the International Conference on Machine Learning, PMLR, Baltimore, MD, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Petroni, F., Rockt\u00e4schel, T., Riedel, S., Lewis, P.S., Bakhtin, A., Wu, Y., and Miller, A.H. Language Models as Knowledge Bases? In Proceedings of the EMNLP\/IJCNLP, Hong Kong, China, 3\u20137 November 2019.","DOI":"10.18653\/v1\/D19-1250"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F. (2022, January 22\u201327). Knowledge Neurons in Pretrained Transformers. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Dublin, Ireland.","DOI":"10.18653\/v1\/2022.acl-long.581"},{"key":"ref_10","unstructured":"Bommasani, R., Hudson, D.A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M.S., Bohg, J., Bosselut, A., and Brunskill, E. (2021). On the Opportunities and Risks of Foundation Models. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"31249","DOI":"10.1109\/ACCESS.2024.3370444","article-title":"Language Models Fine-Tuning for Automatic Format Reconstruction of SEC Financial Filings","volume":"12","author":"Lombardo","year":"2024","journal-title":"IEEE Access"},{"key":"ref_12","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_13","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_14","unstructured":"Wang, X., Wei, J., Schuurmans, D., Le, Q.V., Chi, E.H., Narang, S., Chowdhery, A., and Zhou, D. (2023, January 1\u20135). Self-Consistency Improves Chain of Thought Reasoning in Language Models. Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_15","unstructured":"Verma, S. (2024). Contextual compression in retrieval-augmented generation for large language models: A survey. arXiv."},{"key":"ref_16","unstructured":"Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G. (2023, January 23\u201329). Pal: Program-aided language models. Proceedings of the International Conference on Machine Learning, PMLR, Honolulu, HI, USA."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"3148","DOI":"10.1109\/TAI.2025.3567369","article-title":"TrumorGPT: Graph-Based Retrieval-Augmented Large Language Model for Fact-Checking","volume":"6","author":"Hang","year":"2025","journal-title":"IEEE Trans. Artif. Intell."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"9430919","DOI":"10.1155\/2022\/9430919","article-title":"Lazy Network: A Word Embedding-Based Temporal Financial Network to Avoid Economic Shocks in Asset Pricing Models","volume":"2022","author":"Adosoglou","year":"2022","journal-title":"Complexity"},{"key":"ref_19","unstructured":"Liu, F., AlDahoul, N., Eady, G., Zaki, Y., and Rahwan, T. (2024). Self-Reflection Makes Large Language Models Safer, Less Biased, and Ideologically Neutral. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Lombardo, G., Tomaiuolo, M., Mordonini, M., Codeluppi, G., and Poggi, A. (2022). Mobility in unsupervised word embeddings for knowledge extraction\u2014The scholars\u2019 trajectories across research topics. Future Internet, 14.","DOI":"10.3390\/fi14010025"},{"key":"ref_21","first-page":"68539","article-title":"Toolformer: Language models can teach themselves to use tools","volume":"36","author":"Schick","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_22","unstructured":"Chen, W., Ma, X., Wang, X., and Cohen, W.W. (2023). Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks. arXiv."},{"key":"ref_23","first-page":"8634","article-title":"Reflexion: Language agents with verbal reinforcement learning","volume":"36","author":"Shinn","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_24","unstructured":"Du, Y., Li, S., Torralba, A., Tenenbaum, J.B., and Mordatch, I. (2024, January 21\u201327). Improving factuality and reasoning in language models through multiagent debate. Proceedings of the Forty-First International Conference on Machine Learning, Vienna, Austria."},{"key":"ref_25","first-page":"51991","article-title":"Camel: Communicative agents for \u201cmind\u201d exploration of large language model society","volume":"36","author":"Li","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Roitero, K., Wright, D., Soprano, M., Augenstein, I., and Mizzaro, S. (2025, January 13\u201318). Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking. Proceedings of the SIGIR\u201925, New York, NY, USA.","DOI":"10.1145\/3726302.3729960"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Angiani, G., Fornacciari, P., Lombardo, G., Poggi, A., and Tomaiuolo, M. (2018). Actors based agent modelling and simulation. Highlights of Practical Applications of Agents, Multi-Agent Systems, and Complexity: The PAAMS Collection, Proceedings of the International Conference on Practical Applications of Agents and Multi-Agent Systems, Toledo, Spain, 20\u201322 June 2018, Springer.","DOI":"10.1007\/978-3-319-94779-2_38"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Qi, P., Zhang, Y., Zhang, Y., Bolton, J., and Manning, C.D. (2020, January 5\u201310). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Seattle, WA, USA.","DOI":"10.18653\/v1\/2020.acl-demos.14"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Cheng, J., Liu, X., Zheng, K., Ke, P., Wang, H., Dong, Y., Tang, J., and Huang, M. (2024, January 11\u201316). Black-Box Prompt Optimization: Aligning Large Language Models without Model Training. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.176"},{"key":"ref_30","unstructured":"He, P., Gao, J., and Chen, W. (2023, January 1\u20135). DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing. Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_31","unstructured":"Diera, A., Galke, L., Karl, F., and Scherp, A. (2024). Efficient Continual Learning for Small Language Models with via a Discrete Key-Value Bottleneck. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E. (2017, January 7\u201311). RACE: Large-scale ReAding Comprehension Dataset From Examinations. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Copenhagen, Denmark.","DOI":"10.18653\/v1\/D17-1082"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Richardson, M., Burges, C.J., and Renshaw, E. (2013, January 18\u201321). MCTest: A Challenge Dataset for the Open-Domain Machine Comprehension of Text. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Seattle, WA, USA.","DOI":"10.18653\/v1\/D13-1020"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Pang, R.Y., Khashabi, D., Chen, W.t., Sabharwal, A., Clark, P., and Jia, R. (2022, January 10\u201315). QuALITY: Question Answering with Long Input Texts, Yes!. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Seattle, WA, USA.","DOI":"10.18653\/v1\/2022.naacl-main.391"},{"key":"ref_35","unstructured":"Bajaj, P., Campos, D., Craswell, N., Deng, L., Gao, J., Liu, X., Majumder, R., McNamara, A., Mitra, B., and Nguyen, T. (2016, January 17\u201319). MS MARCO: A human generated machine reading comprehension dataset. Proceedings of the Workshop on Cognitive Computing (ICRC), San Diego, CA, USA."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Cao, S., and Wang, L. (2022, January 22\u201327). HIBRIDS: Attention with Hierarchical Biases for Structure-aware Long Document Summarization. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Dublin, Ireland.","DOI":"10.18653\/v1\/2022.acl-long.58"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Huang, L., Cao, S., Parulian, N., Ji, H., and Wang, L. (2021, January 6\u201311). Efficient Attentions for Long Document Summarization. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online.","DOI":"10.18653\/v1\/2021.naacl-main.112"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Sellam, T., Das, D., and Parikh, A. (2020, January 5\u201310). BLEURT: Learning Robust Metrics for Text Generation. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.acl-main.704"},{"key":"ref_39","unstructured":"Lin, C.Y. (2004). ROUGE: A Package for Automatic Evaluation of Summaries. Text Summarization Branches Out, Proceedings of the ACL-04 Workshop, Barcelona, Spain, 25\u201326 July 2004, Association for Computational Linguistics. Anthology ID: W04-1013."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002, January 6\u201312). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, ACL\u201902, Philadelphia, PA USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Lavie, A., and Agarwal, A. (2007, January 23). Meteor: An automatic metric for MT evaluation with high levels of correlation with human judgments. Proceedings of the Second Workshop on Statistical Machine Translation, StatMT\u201907, Prague, Czech Republic.","DOI":"10.3115\/1626355.1626389"},{"key":"ref_42","unstructured":"Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., and Artzi, Y. (2019). BERTScore: Evaluating Text Generation with BERT. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Warner, B., Chaffin, A., Clavi\u00e9, B., Weller, O., Hallstr\u00f6m, O., Taghadouini, S., Gallagher, A., Biswas, R., Ladhak, F., and Aarsen, T. (2024). Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. arXiv.","DOI":"10.18653\/v1\/2025.acl-long.127"},{"key":"ref_44","unstructured":"Boizard, N., Gisserot-Boukhlef, H., Alves, D.M., Martins, A., Hammal, A., Corro, C., Hudelot, C., Malherbe, E., Malaboeuf, E., and Jourdan, F. (2025). EuroBERT: Scaling multilingual encoders for European languages. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Guo, M., Ainslie, J., Uthus, D.C., Ontanon, S., Ni, J., Sung, Y.H., and Yang, Y. (2022, January 10\u201315). LongT5: Efficient Text-To-Text Transformer for Long Sequences. Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2022, Seattle, WA, USA.","DOI":"10.18653\/v1\/2022.findings-naacl.55"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Phang, J., Zhao, Y., and Liu, P.J. (2023, January 6\u201310). Investigating Efficiently Extending Transformers for Long Input Summarization. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore.","DOI":"10.18653\/v1\/2023.emnlp-main.240"},{"key":"ref_47","unstructured":"Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., and Huang, F. (2025). Qwen2.5 Technical Report. arXiv."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/17\/12\/535\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,26]],"date-time":"2025-11-26T05:20:44Z","timestamp":1764134444000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/17\/12\/535"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,24]]},"references-count":47,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["fi17120535"],"URL":"https:\/\/doi.org\/10.3390\/fi17120535","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,24]]}}}