{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T18:37:29Z","timestamp":1782931049180,"version":"3.54.5"},"reference-count":23,"publisher":"SAGE Publications","issue":"2","license":[{"start":{"date-parts":[[2025,7,3]],"date-time":"2025-07-03T00:00:00Z","timestamp":1751500800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"funder":[{"name":"supported by NSFC","award":["62176155"],"award-info":[{"award-number":["62176155"]}]},{"DOI":"10.13039\/501100003399","name":"Science and Technology Commission of Shanghai Municipality","doi-asserted-by":"publisher","award":["2021SHZDZX0102"],"award-info":[{"award-number":["2021SHZDZX0102"]}],"id":[{"id":"10.13039\/501100003399","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Intelligent Data Analysis: An International Journal"],"published-print":{"date-parts":[[2026,3]]},"abstract":"<jats:p>In the rapid development of artificial intelligence, multimodal large models (VLM) have led a new wave of technological progress with their revolutionary breakthroughs in the field of natural language processing (NLP). In the wave of artificial intelligence, on-device multimodal large models (On-Device VLM) are becoming the new favourites of technological innovation with their rapid development speed and broad application prospects, and the demand for on-device inference is growing. This study conducts in-depth adaptation and optimization of multimodal large models on the Neural Network Processing Unit (NPU) based on the Qualcomm platform. By adopting the QNN (Qualcomm Neural Network) framework and model compression techniques such as quantization, pruning, knowledge distillation, low-rank factorization, and Lookahead decoding, efficient inference acceleration on Qualcomm NPU is achieved, significantly improving the model\u2019s response speed and decoding efficiency. Experimental results show that the optimized model has significant improvements in first response time and decoding speed, providing a new solution for on-device AI applications.<\/jats:p>","DOI":"10.1177\/1088467x251342172","type":"journal-article","created":{"date-parts":[[2025,7,3]],"date-time":"2025-07-03T04:28:53Z","timestamp":1751516933000},"page":"544-568","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":3,"title":["Edge-side NPU inference optimization: Adaptation research of multimodal large models on qualcomm platforms"],"prefix":"10.1177","volume":"30","author":[{"given":"Yajie","family":"Zhu","sequence":"first","affiliation":[{"name":"Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai 200000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongtao","family":"Lu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai 200000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2025,7,3]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Zhou C Kuang Z Zhan F et\u00a0al. MobileViT: lightweight vision transformer for mobile edge applications. ArXiv preprint arXiv:2302.04549 2023."},{"key":"e_1_3_2_3_2","unstructured":"Luo Z Hu X Chen X. TinyVLM: efficient vision-language model for mobile visual understanding. In: International conference on computer vision (ICCV) 2023."},{"key":"e_1_3_2_4_2","first-page":"2100","article-title":"Edge-AI vision-language models: design, challenges, and solutions","volume":"34","author":"Wang Y","year":"2023","unstructured":"Wang Y, Zhang Z, Cai W, et al. Edge-AI vision-language models: design, challenges, and solutions. IEEE Trans Neur Netw Learn Syst 2023; 34: 2100\u20132115.","journal-title":"IEEE Trans Neur Netw Learn Syst"},{"key":"e_1_3_2_5_2","first-page":"187","article-title":"Real-time visual-language processing for augmented reality","volume":"68","author":"Chen J","year":"2022","unstructured":"Chen J, Xu Q, Wang Y, et\u00a0al. Real-time visual-language processing for augmented reality. Cons Electr IEEE Trans 2022; 68: 187\u2013194.","journal-title":"Cons Electr IEEE Trans"},{"key":"e_1_3_2_6_2","unstructured":"Li Z Zhang Y Wang K et\u00a0al. Efficient vision-language transformers for mobile vision-question answering. In: Proceedings of the AAAI conference on artificial intelligence 2023."},{"key":"e_1_3_2_7_2","first-page":"35745","article-title":"Cross-modal learning for edge devices: challenges and innovations","volume":"11","author":"Gao F","year":"2023","unstructured":"Gao F, Liu X, Tang M, et\u00a0al. Cross-modal learning for edge devices: challenges and innovations. IEEE Access 2023; 11: 35745\u201335758.","journal-title":"IEEE Access"},{"key":"e_1_3_2_8_2","first-page":"104","article-title":"Lightweight vision-language interaction models for edge computing","volume":"179","author":"Huang L","year":"2023","unstructured":"Huang L, Feng Y, Shang X, et al. Lightweight vision-language interaction models for edge computing. J Parall Distrib Comput 2023; 179: 104\u2013118.","journal-title":"J Parall Distrib Comput"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3661821"},{"key":"e_1_3_2_10_2","unstructured":"Molchanov P Tyree S Karras T et\u00a0al. Pruning convolutional neural networks for resource efficient inference. In: ICLR 2017 2017."},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","unstructured":"He Y Zhang X Sun J. Channel pruning for accelerating very deep neural networks. In: ICCV 2017 2017.","DOI":"10.1109\/ICCV.2017.155"},{"key":"e_1_3_2_12_2","unstructured":"Lee N Ajanthan T Torr PHS. \u201cSNIP: single-shot network pruning based on connection sensitivity \u201d In: ICLR 2019 2019."},{"key":"e_1_3_2_13_2","doi-asserted-by":"crossref","unstructured":"Molchanov P Mallya A Tyree S et\u00a0al. \u201cImportance estimation for neural network pruning \u201d In: CVPR 2019 2019.","DOI":"10.1109\/CVPR.2019.01152"},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","unstructured":"Ilhan F Su G Tekin SF et\u00a0al. Resource-efficient transformer pruning for finetuning of large models. In: 2024 IEEE\/CVF conference on computer vision and pattern recognition (CVPR) pp.16206\u201316215. Seattle WA USA.","DOI":"10.1109\/CVPR52733.2024.01534"},{"key":"e_1_3_2_15_2","doi-asserted-by":"crossref","unstructured":"Lin TY Maire M Belongie S et\u00a0al. Microsoft coco: common objects in context. In: cocodataset 2014 https:\/\/cocodataset.org.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_16_2","first-page":"8024","article-title":"Pytorch: An imperative style, high-performance deep learning library","volume":"32","author":"Paszke A","year":"2019","unstructured":"Paszke A, Gross S, Massa F, et al. Pytorch: An imperative style, high-performance deep learning library. Adv Neural Inf Process Syst 2019 ; 32: 8024\u20138035.","journal-title":"Adv Neural Inf Process Syst"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"Bolya D Foley S Hays J et\u00a0al. Tide: A general toolbox for identifying object detection errors. In: Computer Vision\u2013ECCV 2020: 16th European conference Glasgow UK August 23\u201328 2020 Proceedings Part III 16 2020 pp.558\u2013573. Springer.","DOI":"10.1007\/978-3-030-58580-8_33"},{"key":"e_1_3_2_18_2","first-page":"512","article-title":"Cosine similarity-guided knowledge distillation for robust object detectors","volume":"626","author":"Park S","year":"2024","unstructured":"Park S, Kang D, Paik J. Cosine similarity-guided knowledge distillation for robust object detectors. Nature 2024; 626: 512\u2013520.","journal-title":"Nature"},{"key":"e_1_3_2_19_2","first-page":"111283","article-title":"Distributed constrained optimization with periodic dynamic quantization","volume":"158","author":"Jie L","year":"2023","unstructured":"Jie L, Li L, Ho DWC. Distributed constrained optimization with periodic dynamic quantization. Automatica 2023; 158:\u00a0111283.","journal-title":"Automatica"},{"key":"e_1_3_2_20_2","first-page":"1234","article-title":"Robust low-rank tensor recovery with outlier spikes and its application to hyperspectral imagery","volume":"32","author":"Lu Z","year":"2023","unstructured":"Lu Z, Zhang S, Xie Y. Robust low-rank tensor recovery with outlier spikes and its application to hyperspectral imagery. IEEE Trans Image Process 2023; 32: 1234\u20131248.","journal-title":"IEEE Trans Image Process"},{"key":"e_1_3_2_21_2","article-title":"Fast low-rank matrix factorization with adaptive sampling","author":"Chen S","year":"2024","unstructured":"Chen S, Hu X, Ye J. Fast low-rank matrix factorization with adaptive sampling. Data Min Knowl Discov 2024.","journal-title":"Data Min Knowl Discov"},{"key":"e_1_3_2_22_2","first-page":"2105","article-title":"Low-rank subspace clustering with structure-aware graph","volume":"46","author":"Wang F","year":"2024","unstructured":"Wang F, Zhang L, Liu W. Low-rank subspace clustering with structure-aware graph. IEEE Trans Pattern Anal Mach Intell 2024; 46: 2105\u20132119.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2023.01.019"},{"key":"e_1_3_2_24_2","unstructured":"LMSYS Break the sequential dependency of llm inference using lookahead decoding. LMSYS Org 2023."}],"container-title":["Intelligent Data Analysis: An International Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1088467X251342172","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1088467X251342172","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1088467X251342172","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:21:19Z","timestamp":1777454479000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1088467X251342172"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,3]]},"references-count":23,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,3]]}},"alternative-id":["10.1177\/1088467X251342172"],"URL":"https:\/\/doi.org\/10.1177\/1088467x251342172","relation":{},"ISSN":["1088-467X","1571-4128"],"issn-type":[{"value":"1088-467X","type":"print"},{"value":"1571-4128","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,3]]}}}