{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,26]],"date-time":"2025-09-26T00:20:52Z","timestamp":1758846052931,"version":"3.44.0"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62025205, 62032018, 62472358, 62322601, 62102322"],"award-info":[{"award-number":["62025205, 62032018, 62472358, 62322601, 62102322"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2025,9,3]]},"abstract":"<jats:p>With the rapid advancement of wireless technologies, in-air handwriting recognition based on radio frequency (RF) signals emerges as a promising solution for human-computer interaction. However, existing methods often constrain writing gestures to predefined planes, impose strict requirements on writing order and direction, and limit recognition tasks to a small set of digits or letters. To overcome these limitations, we propose m2VLMs, a modality mapping architecture, which serves as a bridge between the millimeter-wave (mmWave) radar sensing technologies and large vision-language models (VLMs). Based on the proposed architecture, a three-dimensional (3D) in-air handwriting word recognition system named mmPencil is developed. Specifically, we design a multi-stage spatial trajectory reconstruction algorithm that extracts frequency-domain features to identify handwriting regions and achieve high-precision reconstruction of 3D word trajectories. Furthermore, we introduce a novel spatial-to-visual mapping algorithm, which bridges the gap between spatial trajectory information captured by mmWave radar and vision-language representations, providing a foundation for cross-modality understanding in the large vision-language model. As a result, mmPencil tackles the limitations of current solutions that are heavily reliant on handwriting styles and environmental factors, expanding RF-based in-air handwriting recognition to more complex word-level scenarios.<\/jats:p>\n          <jats:p>We collect and release a 3D mmWave handwriting dataset comprising 200 distinct words (ranging from 2 to 9 letters), contributions from 12 users, and 22 different writing scenarios, totaling 7,664 samples with an overall size of 31.66 GB. Extensive experiments demonstrate that mmPencil achieves accurate and robust word recognition in real-world scenarios, remaining unaffected by variations in word categories, length, as well as writing position, range, angle, speed, size, direction, and user movement. Specifically, for 4 seen users, mmPencil achieves a recognition accuracy of approximately 97.60% across 200 word classes. Moreover, benefiting from the generalization capability of VLMs, the zero-shot recognition accuracy for 4 unseen users can also reach 92.50% across 50 word classes using training data from only 8 users, which outperforming state-of-the-art baselines.<\/jats:p>","DOI":"10.1145\/3749504","type":"journal-article","created":{"date-parts":[[2025,9,3]],"date-time":"2025-09-03T17:15:45Z","timestamp":1756919745000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["mmPencil: Toward Writing-Style-Independent In-Air Handwriting Recognition via mmWave Radar and Large Vision-Language Model"],"prefix":"10.1145","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-0284-8965","authenticated-orcid":false,"given":"Yifan","family":"Guo","sequence":"first","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, Shaanxi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2368-8947","authenticated-orcid":false,"given":"Zhu","family":"Wang","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, Shaanxi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-0420-2134","authenticated-orcid":false,"given":"Qian","family":"Qin","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, Shaanxi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-8171-0230","authenticated-orcid":false,"given":"Yangqian","family":"Lei","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, Shaanxi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-4404-4303","authenticated-orcid":false,"given":"Qiwen","family":"Gan","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, Shaanxi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3115-278X","authenticated-orcid":false,"given":"Zhuo","family":"Sun","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, Shaanxi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2094-9734","authenticated-orcid":false,"given":"Chao","family":"Chen","sequence":"additional","affiliation":[{"name":"Chongqing University, Chongqing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6097-2467","authenticated-orcid":false,"given":"Bin","family":"Guo","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, Shaanxi, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9905-3238","authenticated-orcid":false,"given":"Zhiwen","family":"Yu","sequence":"additional","affiliation":[{"name":"Northwestern Polytechnical University, Xi'an, Shaanxi, China and Harbin Engineering University, Harbin, Heilongjiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,3]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSEN.2019.2922395"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2858036.2858498"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3636534.3698835"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3699766"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3636534.3690687"},{"key":"e_1_2_2_6_1","volume-title":"MMHTSR: In-Air Handwriting Trajectory Sensing and Reconstruction Based on mmWave Radar","author":"Chen Qin","year":"2023","unstructured":"Qin Chen, Zongyong Cui, Zheng Zhou, Yu Tian, and Zongjie Cao. 2023. MMHTSR: In-Air Handwriting Trajectory Sensing and Reconstruction Based on mmWave Radar. IEEE Internet of Things Journal (2023)."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.92.22.9921"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3210240.3210345"},{"key":"e_1_2_2_9_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3351238","article-title":"ASSV: handwritten signature verification using acoustic signals","volume":"3","author":"Ding Feng","year":"2019","unstructured":"Feng Ding, Dong Wang, Qian Zhang, and Run Zhao. 2019. ASSV: handwritten signature verification using acoustic signals. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 3, 3 (2019), 1--22.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2018.8486285"},{"key":"e_1_2_2_11_1","volume-title":"Wicgesture: Meta-motion based continuous gesture recognition with wi-fi","author":"Gao Ruiyang","year":"2023","unstructured":"Ruiyang Gao, Wenwei Li, Jinyi Liu, Shuyu Dai, Mi Zhang, Leye Wang, and Daqing Zhang. 2023. Wicgesture: Meta-motion based continuous gesture recognition with wi-fi. IEEE Internet of Things Journal (2023)."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3555600"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3463504"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2021.3123694"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_2_16_1","first-page":"3","article-title":"Lora: Low-rank adaptation of large language models","volume":"1","author":"Hu Edward J","year":"2022","unstructured":"Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. 2022. Lora: Low-rank adaptation of large language models. ICLR 1, 2 (2022), 3.","journal-title":"ICLR"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/234313.234387"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372224.3419202"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613165"},{"key":"e_1_2_2_20_1","volume-title":"AirWrite: An Aerial Handwriting Trajectory Tracking and Recognition System with mmWave","author":"Lin Chi","year":"2024","unstructured":"Chi Lin, Zhouhe Sun, Asfandeyar Ahmad, Xinxin Fan, Yi Wang, Lei Wang, Xin Fan, and Guowei Wu. 2024. AirWrite: An Aerial Handwriting Trajectory Tracking and Recognition System with mmWave. IEEE Transactions on Mobile Computing (2024)."},{"key":"e_1_2_2_21_1","unstructured":"Haotian Liu Chunyuan Li Yuheng Li Bo Li Yuanhan Zhang Sheng Shen and Yong Jae Lee. 2024. LLaVA-NeXT: Improved reasoning OCR and world knowledge. https:\/\/llava-vl.github.io\/blog\/2024-01-30-llava-next\/"},{"key":"e_1_2_2_22_1","first-page":"1","article-title":"UniFi: A Unified Framework for Generalizable Gesture Recognition with Wi-Fi Signals Using Consistency-guided Multi-View Networks","volume":"7","author":"Liu Yan","year":"2024","unstructured":"Yan Liu, Anlan Yu, Leye Wang, Bin Guo, Yang Li, Enze Yi, and Daqing Zhang. 2024. UniFi: A Unified Framework for Generalizable Gesture Recognition with Wi-Fi Signals Using Consistency-guided Multi-View Networks. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 7, 4 (2024), 1--29.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/s100320200071"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272973.3273009"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3115933"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2021.3063135"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.598226"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3224313"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3066507"},{"key":"e_1_2_2_30_1","doi-asserted-by":"crossref","unstructured":"Zhihui Ren Zhu Wang Zhuo Sun Yifan Guo Wenchao Song Hualei Zhang Chao Chen Bin Guo Zhiwen Yu Xingshe Zho et al. 2024. Characterizing the through-wall sensing mechanism of wi-fi signals with a refraction-aware fresnel zone model. IEEE Transactions on Mobile Computing (2024).","DOI":"10.1109\/TMC.2024.3425847"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514236"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2011.941097"},{"key":"e_1_2_2_33_1","volume-title":"FinerSense: a Fine-grained Respiration Sensing System Based on Precise Separation of Wi-Fi Signals","author":"Song Wenchao","year":"2024","unstructured":"Wenchao Song, Zhu Wang, Yifan Guo, Zhuo Sun, Zhihui Ren, Chao Chen, Bin Guo, Zhiwen Yu, Xingshe Zhou, and Daqing Zhang. 2024. FinerSense: a Fine-grained Respiration Sensing System Based on Precise Separation of Wi-Fi Signals. IEEE Transactions on Mobile Computing (2024)."},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1049\/ip-f-2.1992.0048"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/81.948437"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2018.8486346"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3699751"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS56603.2022.00019"},{"key":"e_1_2_2_39_1","unstructured":"Peng Wang Shuai Bai Sinan Tan Shijie Wang Zhihao Fan Jinze Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge et al. 2024. Qwen2-vl: Enhancing vision-language model's perception of the world at any resolution. arXiv preprint arXiv:2409.12191 (2024)."},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3666025.3699336"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3678589"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3699726"},{"key":"e_1_2_2_43_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3631443","article-title":"LiqDetector: Enabling Container-Independent Liquid Detection with mmWave Signals Based on a Dual-Reflection Model","volume":"7","author":"Wang Zhu","year":"2024","unstructured":"Zhu Wang, Yifan Guo, Zhihui Ren, Wenchao Song, Zhuo Sun, Chao Chen, Bin Guo, and Zhiwen Yu. 2024. LiqDetector: Enabling Container-Independent Liquid Detection with mmWave Signals Based on a Dual-Reflection Model. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 7, 4 (2024), 1--24.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_2_44_1","volume-title":"Gesture-radar: A dual doppler radar based system for robust recognition and quantitative profiling of human gestures","author":"Wang Zhu","year":"2020","unstructured":"Zhu Wang, Zhiwen Yu, Xinye Lou, Bin Guo, and Liming Chen. 2020. Gesture-radar: A dual doppler radar based system for robust recognition and quantitative profiling of human gestures. IEEE transactions on human-machine systems 51, 1 (2020), 32--43."},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3534601"},{"key":"e_1_2_2_46_1","unstructured":"Haoran Wei Chenglong Liu Jinyue Chen Jia Wang Lingyu Kong Yanming Xu Zheng Ge Liang Zhao Jianjian Sun Yuang Peng et al. 2024. General ocr theory: Towards ocr-2.0 via a unified end-to-end model. (2024)."},{"key":"e_1_2_2_47_1","unstructured":"G Welch. 1995. An Introduction to the Kalman Filter. (1995)."},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3666025.3699349"},{"key":"e_1_2_2_49_1","volume-title":"Principal component analysis. Chemometrics and intelligent laboratory systems 2, 1--3","author":"Wold Svante","year":"1987","unstructured":"Svante Wold, Kim Esbensen, and Paul Geladi. 1987. Principal component analysis. Chemometrics and intelligent laboratory systems 2, 1--3 (1987), 37--52."},{"key":"e_1_2_2_50_1","first-page":"1","article-title":"mSense: Towards mobile material sensing with a single millimeter-wave radio","volume":"4","author":"Wu Chenshu","year":"2020","unstructured":"Chenshu Wu, Feng Zhang, Beibei Wang, and KJ Ray Liu. 2020. mSense: Towards mobile material sensing with a single millimeter-wave radio. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 4, 3 (2020), 1--20.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411822"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3485730.3485936"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3581791.3596839"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3638550.3641130"},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3495243.3560515"},{"key":"e_1_2_2_56_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3351273","article-title":"Acousticid: gait-based human identification using acoustic signal","volume":"3","author":"Xu Wei","year":"2019","unstructured":"Wei Xu, ZhiWen Yu, Zhu Wang, Bin Guo, and Qi Han. 2019. Acousticid: gait-based human identification using acoustic signal. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 3, 3 (2019), 1--25.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3610900"},{"key":"e_1_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397334"},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/THMS.2022.3149408"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3596237"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3170157"},{"key":"e_1_2_2_62_1","volume-title":"Rajesh K Gupta, and Jingbo Shang.","author":"Zhang Xiyuan","year":"2024","unstructured":"Xiyuan Zhang, Ranak Roy Chowdhury, Rajesh K Gupta, and Jingbo Shang. 2024. Large language models for time series: A survey. arXiv preprint arXiv:2402.01801 (2024)."},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3570361.3592515"},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2023.3265988"},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3678550"},{"key":"e_1_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3699777"},{"key":"e_1_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i28.35383"},{"key":"e_1_2_2_68_1","volume-title":"Llamafactory: Unified efficient fine-tuning of 100+ language models. arXiv preprint arXiv:2403.13372","author":"Zheng Yaowei","year":"2024","unstructured":"Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma. 2024. Llamafactory: Unified efficient fine-tuning of 100+ language models. arXiv preprint arXiv:2403.13372 (2024)."},{"key":"e_1_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307334.3326081"},{"key":"e_1_2_2_70_1","volume-title":"Tent: Connect language models with iot sensors for zero-shot activity recognition. arXiv preprint arXiv:2311.08245","author":"Zhou Yunjiao","year":"2023","unstructured":"Yunjiao Zhou, Jianfei Yang, Han Zou, and Lihua Xie. 2023. Tent: Connect language models with iot sensors for zero-shot activity recognition. arXiv preprint arXiv:2311.08245 (2023)."}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3749504","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T16:35:24Z","timestamp":1758818124000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3749504"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,3]]},"references-count":70,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,9,3]]}},"alternative-id":["10.1145\/3749504"],"URL":"https:\/\/doi.org\/10.1145\/3749504","relation":{},"ISSN":["2474-9567"],"issn-type":[{"type":"electronic","value":"2474-9567"}],"subject":[],"published":{"date-parts":[[2025,9,3]]},"assertion":[{"value":"2025-09-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}