{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,20]],"date-time":"2026-04-20T10:07:07Z","timestamp":1776679627452,"version":"3.51.2"},"publisher-location":"Singapore","reference-count":38,"publisher":"Springer Nature Singapore","isbn-type":[{"value":"9789819570775","type":"print"},{"value":"9789819570782","type":"electronic"}],"license":[{"start":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:00:00Z","timestamp":1767225600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:00:00Z","timestamp":1767225600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026]]},"DOI":"10.1007\/978-981-95-7078-2_11","type":"book-chapter","created":{"date-parts":[[2026,4,20]],"date-time":"2026-04-20T09:25:37Z","timestamp":1776677137000},"page":"158-172","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["LMSQuant: Learnable Multiscale Post-training Quantization for\u00a0Large Language Models"],"prefix":"10.1007","author":[{"given":"Yongkang","family":"Yang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuzhe","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hong","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,4,21]]},"reference":[{"key":"11_CR1","doi-asserted-by":"crossref","unstructured":"Bisk, Y., et\u00a0al.: PIQA: reasoning about physical commonsense in natural language. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol.\u00a034, pp. 7432\u20137439 (2020)","DOI":"10.1609\/aaai.v34i05.6239"},{"key":"11_CR2","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown, T., et al.: Language models are few-shot learners. Adv. Neural. Inf. Process. Syst. 33, 1877\u20131901 (2020)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"11_CR3","doi-asserted-by":"crossref","unstructured":"Cheng, W., et al.: Optimize weight rounding via signed gradient descent for the quantization of LLMs. In: Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 11332\u201311350 (2024)","DOI":"10.18653\/v1\/2024.findings-emnlp.662"},{"key":"11_CR4","unstructured":"Clark, C., et al.: BoolQ: exploring the surprising difficulty of natural yes\/no questions. arXiv preprint arXiv:1905.10044 (2019)"},{"key":"11_CR5","unstructured":"Clark, P., et al.: Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv preprint arXiv:1803.05457 (2018)"},{"key":"11_CR6","doi-asserted-by":"crossref","unstructured":"Dettmers, T., Lewis, M., Belkada, Y., Zettlemoyer, L.: GPT3.int8(): 8-bit matrix multiplication for transformers at scale. Adv. Neural Inf. Process. Syst. 35, 30318\u201330332 (2022)","DOI":"10.52202\/068431-2198"},{"key":"11_CR7","unstructured":"Dettmers, T., et al.: SpQR: a sparse-quantized representation for near-lossless LLM weight compression. In: International Conference on Learning Representations (2024)"},{"key":"11_CR8","unstructured":"Ding, X., et\u00a0al.: CBQ: Cross-block quantization for large language models. arXiv preprint arXiv:2312.07950 (2023)"},{"key":"11_CR9","unstructured":"Frantar, E., Ashkboos, S., Hoefler, T., Alistarh, D.: OPTQ: Accurate quantization for generative pre-trained transformers. In: International Conference on Learning Representations (2023)"},{"key":"11_CR10","unstructured":"Grattafiori, A., et\u00a0al.: The Llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)"},{"key":"11_CR11","doi-asserted-by":"crossref","unstructured":"Guan, Z., et al.: APTQ: Attention-aware post-training mixed-precision quantization for large language models. In: Proceedings of the 61st ACM\/IEEE Design Automation Conference, pp.\u00a01\u20136 (2024)","DOI":"10.1145\/3649329.3658498"},{"key":"11_CR12","unstructured":"Huang, W., et al.: SliM-LLM: salience-driven mixed-precision quantization for large language models. arXiv preprint arXiv:2405.14917 (2024)"},{"key":"11_CR13","doi-asserted-by":"crossref","unstructured":"Jeon, Y., Lee, C., Park, K., Kim, H.y.: A frustratingly easy post-training quantization scheme for LLMs. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 14446\u201314461 (2023)","DOI":"10.18653\/v1\/2023.emnlp-main.892"},{"key":"11_CR14","doi-asserted-by":"crossref","unstructured":"Lang, J., Guo, Z., Huang, S.: A comprehensive study on quantization techniques for large language models. In: International Conference on Artificial Intelligence, Robotics, and Communication, pp. 224\u2013231 (2024)","DOI":"10.1109\/ICAIRC64177.2024.10899941"},{"key":"11_CR15","doi-asserted-by":"crossref","unstructured":"Lee, C., Jin, J., Kim, T., Kim, H., Park, E.: OWQ: outlier-aware weight quantization for efficient fine-tuning and inference of large language models. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol.\u00a038, pp. 13355\u201313364 (2024)","DOI":"10.1609\/aaai.v38i12.29237"},{"key":"11_CR16","unstructured":"Li, Y., et al.: BRECQ: pushing the limit of post-training quantization by block reconstruction. arXiv preprint arXiv:2102.05426 (2021)"},{"key":"11_CR17","first-page":"87","volume":"6","author":"J Lin","year":"2024","unstructured":"Lin, J., et al.: AWQ: activation-aware weight quantization for on-device LLM compression and acceleration. Proc. Mach. Learn. Syst. 6, 87\u2013100 (2024)","journal-title":"Proc. Mach. Learn. Syst."},{"key":"11_CR18","unstructured":"Liu, J., et al.: QLLM: Accurate and efficient low-bitwidth quantization for large language models. In: International Conference on Learning Representations (2024)"},{"key":"11_CR19","unstructured":"Liu, W., Ma, X., Zhang, P., Wang, Y.: CrossQuant: a post-training quantization method with smaller quantization kernel for precise large language model compression. arXiv preprint arXiv:2410.07505 (2024)"},{"key":"11_CR20","doi-asserted-by":"crossref","unstructured":"Liu, Z., et al.: LLM-QAT: data-free quantization aware training for large language models. arXiv preprint arXiv:2305.17888 (2023)","DOI":"10.18653\/v1\/2024.findings-acl.26"},{"key":"11_CR21","unstructured":"Ma, Y., et al.: AffineQuant: affine transformation quantization for large language models. In: International Conference on Learning Representations (2024)"},{"key":"11_CR22","doi-asserted-by":"crossref","unstructured":"Marcus, M., et al.: The penn treebank: annotating predicate argument structure. In: Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, 8\u201311 March 1994 (1994)","DOI":"10.3115\/1075812.1075835"},{"key":"11_CR23","unstructured":"Merity, S., Xiong, C., Bradbury, J., Socher, R.: Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843 (2016)"},{"key":"11_CR24","first-page":"1","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel, C., et al.: Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 1\u201367 (2020)","journal-title":"J. Mach. Learn. Res."},{"issue":"9","key":"11_CR25","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1145\/3474381","volume":"64","author":"K Sakaguchi","year":"2021","unstructured":"Sakaguchi, K., Bras, R.L., Bhagavatula, C., Choi, Y.: WinoGrande: an adversarial winograd schema challenge at scale. Commun. ACM 64(9), 99\u2013106 (2021)","journal-title":"Commun. ACM"},{"key":"11_CR26","unstructured":"Shabanovi, K., Wiest, L., Golkov, V., Cremers, D., Pfeil, T.: Interactions across blocks in post-training quantization of large language models. arXiv preprint arXiv:2411.03934 (2024)"},{"key":"11_CR27","unstructured":"Shao, W., et al.: OmniQuant: omnidirectionally calibrated quantization for large language models. In: International Conference on Learning Representations (2024)"},{"key":"11_CR28","unstructured":"Touvron, H., et\u00a0al.: LLaMA: open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)"},{"key":"11_CR29","unstructured":"Touvron, H., et\u00a0al.: Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)"},{"key":"11_CR30","doi-asserted-by":"publisher","first-page":"1417","DOI":"10.1007\/s00521-024-10657-6","volume":"37","author":"X Wang","year":"2025","unstructured":"Wang, X., Hu, Y., Yang, Z.: Neural network quantization: separate scaling of rows and columns in weight matrix. Neural Comput. Appl. 37, 1417\u20131428 (2025)","journal-title":"Neural Comput. Appl."},{"key":"11_CR31","doi-asserted-by":"crossref","unstructured":"Wei, X., et al.: Outlier suppression+: accurate quantization of large language models by equivalent and effective shifting and scaling. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 1648\u20131665. Association for Computational Linguistics (2023)","DOI":"10.18653\/v1\/2023.emnlp-main.102"},{"key":"11_CR32","unstructured":"Xiao, G., et a;.: SmoothQuant: accurate and efficient post-training quantization for large language models. In: International Conference on Machine Learning, pp. 38087\u201338099 (2023)"},{"key":"11_CR33","first-page":"27168","volume":"35","author":"Z Yao","year":"2022","unstructured":"Yao, Z., et al.: ZeroQuant: efficient and affordable post-training quantization for large-scale transformers. Adv. Neural. Inf. Process. Syst. 35, 27168\u201327183 (2022)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"11_CR34","unstructured":"Yue, Y., et al.: WKVQuant: quantizing weight and key\/value cache for large language models gains more. arXiv preprint arXiv:2402.12065 (2024)"},{"key":"11_CR35","doi-asserted-by":"crossref","unstructured":"Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., Choi, Y.: HellaSwag: can a machine really finish your sentence? arXiv preprint arXiv:1905.07830 (2019)","DOI":"10.18653\/v1\/P19-1472"},{"key":"11_CR36","unstructured":"Zhang, S., et\u00a0al.: OPT: open pre-trained transformer language models. arXiv preprint arXiv:2205.01068 (2022)"},{"key":"11_CR37","doi-asserted-by":"crossref","unstructured":"Zhao, J., et al.: LRQuant: learnable and robust post-training quantization for large language models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pp. 2240\u20132255. Association for Computational Linguistics (2024)","DOI":"10.18653\/v1\/2024.acl-long.122"},{"key":"11_CR38","unstructured":"Zhou, Z., et\u00a0al.: A survey on efficient inference for large language models. arXiv preprint arXiv:2404.14294 (2024)"}],"container-title":["Lecture Notes in Computer Science","PRICAI 2025: Trends in Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-981-95-7078-2_11","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,20]],"date-time":"2026-04-20T09:25:57Z","timestamp":1776677157000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-981-95-7078-2_11"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"ISBN":["9789819570775","9789819570782"],"references-count":38,"URL":"https:\/\/doi.org\/10.1007\/978-981-95-7078-2_11","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"value":"0302-9743","type":"print"},{"value":"1611-3349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026]]},"assertion":[{"value":"21 April 2026","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"While our work significantly enhances the efficiency and accessibility of powerful LLMs, it also inevitably lowers the barrier for their misuse in generating misinformation and harmful content. This dual-use nature creates an imperative to treat robust on-device safeguards not as an afterthought, but as a core principle of technological progress itself.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Considerations"}},{"value":"Despite the impressive performance demonstrated, this study has two limitations. First, due to hardware limitations, we did not apply LMSQuant to models exceeding 70B parameters. In future work, we plan to extend LMSQuant to larger models to validate its effectiveness. Second, the initialization of LMWS parameters currently relies on empirical grid search, which may result in suboptimal configurations. A promising research direction is to develop automated, distribution-aware initialization methods based on the characteristics of local weight distributions.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Limitations"}},{"value":"PRICAI","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Pacific Rim International Conference on Artificial Intelligence","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Wellington","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"New Zealand","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2025","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"17 November 2025","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"21 November 2025","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"22","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"pricai2025","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/www.pricai.org\/2025\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}