{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T15:03:00Z","timestamp":1784732580272,"version":"3.55.0"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"8","license":[{"start":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T00:00:00Z","timestamp":1784678400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61872125"],"award-info":[{"award-number":["61872125"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2019YFB2101703"],"award-info":[{"award-number":["2019YFB2101703"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Graduate Education Innovation and Quality Improvement Program of Henan University","award":["SYLYC2023072"],"award-info":[{"award-number":["SYLYC2023072"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,8,31]]},"abstract":"<jats:p>\n                    Compressed learning (CL), integrating compressed sensing (CS) and machine learning (ML), enables direct inference from few CS measurements. However, existing CL methods either heavily rely on pre-trained models built upon extremely large-scale datasets or perform only relatively simple tasks on small datasets, restricting their scalability in real-world multimedia scenarios. To address these limitations, we propose an efficient CL framework named HTMA-CL for practical multimedia acquisition and edge intelligent processing. HTMA-CL employs CNN-based learnable sampling to realize block-based CS for high-resolution images, significantly decreasing the transmission bandwidth and storage overhead. A hierarchical tokenization module together with a deep-narrow Transformer module progressively models local and global dependencies within the measurements, enabling accurate inference directly in the compressive domain. Various task heads are constructed for performing diverse multimedia analysis tasks such as image classification and semantic segmentation. Extensive experiments demonstrate that our HTMA-CL achieves state-of-the-art performance compared to other CL methods, and nearly comparable performance to the image domain methods at a CS ratio of 10%. Our method further verifies the strong robustness against external interference in the disturbance-prone IoT multimedia environments. The source code is publicly available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/acrlife\/HTMA-CL.git\">https:\/\/github.com\/acrlife\/HTMA-CL.git<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3820057","type":"journal-article","created":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T13:41:23Z","timestamp":1780580483000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["HTMA-CL: A Hierarchical Tokenization and Multiscale Attention Framework for Compressive Domain Multimedia Inference"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5170-8744","authenticated-orcid":false,"given":"Yanhao","family":"Jing","sequence":"first","affiliation":[{"name":"School of Computer and Information Engineering, Henan University, Kaifeng, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0371-8707","authenticated-orcid":false,"given":"Xiangjun","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Mathematics and Statistics, Henan University, Kaifeng, China and School of Computer and Information Engineering, Henan University, Kaifeng, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-9282-7024","authenticated-orcid":false,"given":"Datao","family":"You","sequence":"additional","affiliation":[{"name":"School of Software, Henan University, Kaifeng, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-0685-9841","authenticated-orcid":false,"given":"Hui","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer and Information Engineering, Henan University, Kaifeng, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4516-3219","authenticated-orcid":false,"given":"Zhe","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Mathematics and Statistics, Henan University, Kaifeng, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3062-5004","authenticated-orcid":false,"given":"Haibin","family":"Kan","sequence":"additional","affiliation":[{"name":"College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5926-4276","authenticated-orcid":false,"given":"J\u00fcrgen","family":"Kurths","sequence":"additional","affiliation":[{"name":"Potsdam Institute for Climate Impact Research (PIK), Potsdam, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,22]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Robert Calderbank Sina Jafarpour and Robert Schapire. 2009. Compressed learning: Universal sparse dimensionality reduction and learning in the measurement domain. Retrieved from https:\/\/www.researchgate.net\/"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2005.862083"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2699184"},{"key":"e_1_3_1_5_2","first-page":"3213","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201916)","author":"Cordts Marius","year":"2016","unstructured":"Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. 2016. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201916), 3213\u20133223."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3649441"},{"key":"e_1_3_1_7_2","first-page":"1","article-title":"Performance bounds of compressive classification under perturbation","volume":"180","author":"Cui Yupeng","year":"2021","unstructured":"Yupeng Cui, Wenbo Xu, Yue Wang, Jiaru Lin, and Liyang Lu. 2021. Performance bounds of compressive classification under perturbation. Signal Processing 180, Article 107855 (Mar. 2021), 1\u201312.","journal-title":"Signal Processing"},{"key":"e_1_3_1_8_2","first-page":"1","volume-title":"Proceedings of the Society of Photo-Optical Instrumentation Engineers, Computational Imaging (SPIE \u201907)","volume":"6498","author":"Andrew Davenport Mark","year":"2007","unstructured":"Mark Andrew Davenport, Marco F. Duarte, Michael B. Wakin, Jason Noah Laska, Dharmpal Takhar, Kevin F. Kelly, and Richard Gordon Baraniuk. 2007. The smashed filter for compressive classification and target recognition. In Proceedings of the Society of Photo-Optical Instrumentation Engineers, Computational Imaging (SPIE \u201907), Vol. 6498, Article 64980H, 1\u201312."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2006.871582"},{"key":"e_1_3_1_10_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201920)","author":"Dosovitskiy Alexey","year":"2020","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations (ICLR \u201920)."},{"key":"e_1_3_1_11_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201925)","author":"Duan Yuchen","year":"2025","unstructured":"Yuchen Duan, Weiyun Wang, Zhe Chen, Xizhou Zhu, Lewei Lu, Tong Lu, Yu Qiao, Hongsheng Li, Jifeng Dai, and Wenhai Wang. 2025. Vision-RWKV: Efficient and scalable visual perception with RWKV-like architectures. In Proceedings of the International Conference on Learning Representations (ICLR \u201925)."},{"key":"e_1_3_1_12_2","first-page":"6877","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201921)","author":"Feng Jianfeng","year":"2021","unstructured":"Jianfeng Feng, Tao Xiang, Philip H. S. Torr, and Li Zhang. 2021. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201921), 6877\u20136886."},{"key":"e_1_3_1_13_2","first-page":"7519","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201919)","author":"He Junjun","year":"2019","unstructured":"Junjun He, Zhongying Deng, Lei Zhou, Yali Wang, and Yu Qiao. 2019. Adaptive pyramid context network for semantic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201919), 7519\u20137528."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3723358"},{"key":"e_1_3_1_16_2","first-page":"603","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201919)","author":"Huang Zilong","year":"2019","unstructured":"Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. 2019. CCNet: Criss-cross attention for semantic segmentation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201919), 603\u2013612."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2469288"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2011.2180773"},{"key":"e_1_3_1_19_2","first-page":"180","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201924)","author":"Lee Minhyun","year":"2024","unstructured":"Minhyun Lee, Song Park, Byeongho Heo, Dongyoon Han, and Hyunjung Shim. 2024. SeiT++: Masked token modeling improves storage-efficient training. In Proceedings of the European Conference on Computer Vision (ECCV \u201924), 180\u2013197."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6805"},{"key":"e_1_3_1_21_2","first-page":"50","volume-title":"IEEE Transactions on Multimedia","volume":"25","author":"Lin Xiao","year":"2023","unstructured":"Xiao Lin, Shuzhou Sun, Wei Huang, Bin Sheng, Ping Li, and David Dagan Feng. 2023. EAPT: Efficient attention pyramid transformer for image processing. IEEE Transactions on Multimedia 25 (2023), 50\u201361."},{"key":"e_1_3_1_22_2","first-page":"1913","volume-title":"Proceedings of the IEEE International Conference on Image Processing (ICIP \u201916)","author":"Lohit Suhas","year":"2016","unstructured":"Suhas Lohit, Kuldeep Kulkarni, and Pavan Turaga. 2016. Direct inference on compressive measurements using convolutional neural networks. In Proceedings of the IEEE International Conference on Image Processing (ICIP \u201916), 1913\u20131917."},{"key":"e_1_3_1_23_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201922)","author":"Mehta Sachin","year":"2022","unstructured":"Sachin Mehta and Mohammad Rastegari. 2022. MobileViT: Light-weight, general-purpose, and mobile-friendly vision transformer. In Proceedings of the International Conference on Learning Representations (ICLR \u201922)."},{"key":"e_1_3_1_24_2","first-page":"891","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201914)","author":"Mottaghi Roozbeh","year":"2014","unstructured":"Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille. 2014. The role of context for object detection and semantic segmentation in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201914), 891\u2013898."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3194001"},{"key":"e_1_3_1_26_2","first-page":"17202","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201923)","author":"Park Song","year":"2023","unstructured":"Song Park, Sanghyuk Chun, Byeongho Heo, Wonjae Kim, and Sangdoo Yun. 2023. SeiT: Storage-efficient vision training with tokens using 1% of pixel storage. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201923), 17202\u201317213."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/JCN.2013.000083"},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","unstructured":"Danial Qashqai Emad Mousavian Shahriar Baradaran Shokouhi and Sattar Mirzakuchaki. 2024. CSFNet: A cosine similarity fusion network for real-time RGB-X semantic segmentation of driving scenes. arXiv:2407.01328. Retrieved from https:\/\/arxiv.org\/abs\/2407.01328","DOI":"10.2139\/ssrn.4885740"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.74"},{"key":"e_1_3_1_31_2","first-page":"17425","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201923)","author":"Shaker Abdelrahman","year":"2023","unstructured":"Abdelrahman Shaker, Muhammad Maaz, Hanoona Rasheed, Salman Khan, Ming-Hsuan Yang, and Fahad Shahbaz Khan. 2023. SwiftFormer: Efficient additive attention for transformer-based real-time mobile vision applications. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201923), 17425\u201317436."},{"issue":"9","key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"4949","DOI":"10.1109\/TCYB.2024.3363748","article-title":"MTC-CSNet: Marrying transformer and convolution for image compressed sensing","volume":"54","author":"Shen Minghe","year":"2024","unstructured":"Minghe Shen, Hongping Gan, Chunyan Ma, Chao Ning, Hongqi Li, and Feng Liu. 2024. MTC-CSNet: Marrying transformer and convolution for image compressed sensing. IEEE Transactions on Cybernetics 54, 9 (2024), 4949\u20134961.","journal-title":"IEEE Transactions on Cybernetics"},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"6991","DOI":"10.1109\/TIP.2022.3217365","article-title":"TransCS: A transformer-based hybrid architecture for image compressed sensing","volume":"31","author":"Shen Minghe","year":"2022","unstructured":"Minghe Shen, Hongping Gan, Chao Ning, Yi Hua, and Tao Zhang. 2022. TransCS: A transformer-based hybrid architecture for image compressed sensing. IEEE Transactions on Image Processing: A Publication of the IEEE Signal Processing Society 31 (2022), 6991\u20137005.","journal-title":"IEEE Transactions on Image Processing: A Publication of the IEEE Signal Processing Society"},{"key":"e_1_3_1_34_2","first-page":"1","article-title":"TransCS-Net: A hybrid transformer-based privacy-protecting network using compressed sensing for medical image segmentation","volume":"86","author":"Tang Suigu","year":"2023","unstructured":"Suigu Tang, Chak Fong Cheang, Xiaoyuan Yu, Yanyan Liang, Qi Feng, and Zongren Chen. 2023. TransCS-Net: A hybrid transformer-based privacy-protecting network using compressed sensing for medical image segmentation. Biomedical Signal Processing and Control 86, Part A, Article 105131 (Sep. 2023), 1\u201313.","journal-title":"Biomedical Signal Processing and Control"},{"key":"e_1_3_1_35_2","first-page":"10347","volume-title":"Proceedings of the 38th International Conference on Machine Learning (PMLR \u201921)","volume":"139","author":"Touvron Hugo","year":"2021","unstructured":"Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. 2021. Training data-efficient image transformers & distillation through attention. In Proceedings of the 38th International Conference on Machine Learning (PMLR \u201921), Vol. 139, 10347\u201310357."},{"issue":"4","key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"1512","DOI":"10.1109\/TNNLS.2020.2984831","article-title":"Multilinear compressive learning","volume":"32","author":"Tran Dat Thanh","year":"2021","unstructured":"Dat Thanh Tran, Mehmet Yama\u00e7, Aysen Degerli, Moncef Gabbouj, and Alexandros Iosifidis. 2021. Multilinear compressive learning. IEEE Transactions on Neural Networks and Learning Systems 32, 4 (2021), 1512\u20131524.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_1_37_2","first-page":"5998","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS \u201917)","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS \u201917), Vol. 30, 5998\u20136008."},{"key":"e_1_3_1_38_2","first-page":"15909","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201924)","author":"Wang Ao","year":"2024","unstructured":"Ao Wang, Hui Chen, Zijia Lin, Jungong Han, and Guiguang Ding. 2024. RepViT: Revisiting mobile CNN from ViT perspective. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201924), 15909\u201315920."},{"key":"e_1_3_1_39_2","first-page":"14944","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201925)","author":"Wang Feng","year":"2025","unstructured":"Feng Wang, Jiahao Wang, Sucheng Ren, Guoyizhe Wei, Jieru Mei, Wei Shao, Yuyin Zhou, Alan Yuille, and Cihang Xie. 2025. Mamba-Reg: Vision mamba also needs registers. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201925), 14944\u201314953."},{"key":"e_1_3_1_40_2","first-page":"568","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201921)","author":"Wang Wenhai","year":"2021","unstructured":"Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201921), 568\u2013578."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2012.2189859"},{"key":"e_1_3_1_42_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201925)","author":"Yang Xingyi","year":"2025","unstructured":"Xingyi Yang and Xinchao Wang. 2025. Kolmogorov-Arnold transformer. In Proceedings of the International Conference on Learning Representations (ICLR \u201925)."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3274988"},{"key":"e_1_3_1_44_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201922)","author":"Yu Jiahui","year":"2022","unstructured":"Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. 2022. Vector-quantized image modeling with improved VQGAN. In Proceedings of the International Conference on Learning Representations (ICLR \u201922)."},{"key":"e_1_3_1_45_2","first-page":"10819","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201922)","author":"Yu Weihao","year":"2022","unstructured":"Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. 2022. MetaFormer is actually what you need for vision. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201922), 10819\u201310829."},{"key":"e_1_3_1_46_2","first-page":"558","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201921)","author":"Yuan Li","year":"2021","unstructured":"Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis E. H. Tay, Jiashi Feng, and Shuicheng Yan. 2021. Tokens-to-token ViT: Training vision transformers from scratch on ImageNet. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV \u201921), 558\u2013567."},{"key":"e_1_3_1_47_2","first-page":"2881","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Zhao Hengshuang","year":"2017","unstructured":"Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. 2017. Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917), 2881\u20132890."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-018-1140-0"},{"key":"e_1_3_1_49_2","first-page":"62429","volume-title":"Proceedings of the 41st International Conference on Machine Learning (ICML \u201924)","author":"Zhu Lianghui","year":"2024","unstructured":"Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024. Vision mamba: Efficient visual representation learning with bidirectional state space model. In Proceedings of the 41st International Conference on Machine Learning (ICML \u201924), 62429\u201362442."},{"key":"e_1_3_1_50_2","first-page":"3","article-title":"Compressed learning for image classification: A deep neural network approach","volume":"19","author":"Zisselman Eve","year":"2018","unstructured":"Eve Zisselman, Amir Adler, and Michael Elad. 2018. Compressed learning for image classification: A deep neural network approach. Handbook of Numerical Analysis 19 (2018), 3\u201317.","journal-title":"Handbook of Numerical Analysis"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3820057","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T14:26:20Z","timestamp":1784730380000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3820057"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,22]]},"references-count":49,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2026,8,31]]}},"alternative-id":["10.1145\/3820057"],"URL":"https:\/\/doi.org\/10.1145\/3820057","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,22]]},"assertion":[{"value":"2025-12-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-09","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}