{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T19:45:23Z","timestamp":1775591123054,"version":"3.50.1"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"10","license":[{"start":{"date-parts":[[2024,10,29]],"date-time":"2024-10-29T00:00:00Z","timestamp":1730160000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,10,31]]},"abstract":"<jats:p>Food image recognition has recently been given considerable attention in the multimedia field in light of its possible implications on health. The characteristics of the dispersed distribution of ingredients in food images put forward higher requirements on the long-range information extraction ability of neural networks, leading to more complex and deeper models. Nevertheless, the lightweight version of food image recognition is essential for improved implementation on end devices and sustained server-side expansion. To address this issue, we present Aggregation Feature Net (AFNet), a lightweight network that is capable of effectively capturing both global and local features from food images. In AFNet, we develop a novel convolution based on a residual model by encoding global features through row-wise and column-wise information integration. Merging aggregation block with classic local convolution yields a framework that works as the backbone of the network. Based on the efficient use of parameters by the aggregation block, we constructed a lightweight food image recognition network with fewer layers and a smaller scale, assisted by a new type of activation function. Experimental results on four popular food recognition datasets demonstrate that our approach achieves state-of-the-art performance with higher accuracy and fewer FLOPs and parameters. For example, in comparison to the current state-of-the-art model of MobileViTv2, AFNet achieved 88.4% accuracy of the top-1 level on the ETHZ Food-101 dataset, with similar parameters and FLOPs but 1.4% more accuracy. The source code will be provided in supplementary materials.<\/jats:p>","DOI":"10.1145\/3680285","type":"journal-article","created":{"date-parts":[[2024,7,22]],"date-time":"2024-07-22T15:41:37Z","timestamp":1721662897000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Lightweight Food Recognition via Aggregation Block and Feature Encoding"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0785-6007","authenticated-orcid":false,"given":"Yancun","family":"Yang","sequence":"first","affiliation":[{"name":"School of Information and Electrical Engineering, Ludong University, Yantai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6668-9208","authenticated-orcid":false,"given":"Weiqing","family":"Min","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6637-2212","authenticated-orcid":false,"given":"Jingru","family":"Song","sequence":"additional","affiliation":[{"name":"School of Information and Electrical Engineering, Ludong University, Yantai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6790-0239","authenticated-orcid":false,"given":"Guorui","family":"Sheng","sequence":"additional","affiliation":[{"name":"School of Information and Electrical Engineering, Ludong University, Yantai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1025-3955","authenticated-orcid":false,"given":"Lili","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information and Electrical Engineering, Ludong University, Yantai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1596-4326","authenticated-orcid":false,"given":"Shuqiang","family":"Jiang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,10,29]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10599-4_29"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1137\/16M1080173"},{"key":"e_1_3_1_4_2","first-page":"12021","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Jierun","year":"2023","unstructured":"Jierun Chen, Shiu-hong Kao, Hao He, Weipeng Zhuo, Song Wen, Chul-Ho Lee, and S-H Gary Chan. 2023. Run, Don\u2019t Walk: Chasing Higher FLOPS for Faster Neural Networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 12021\u201312031."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2964315"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00520"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_8_2","volume-title":"9th International Conference on Learning Representations (ICLR \u201921), Virtual Event","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image Is Worth 16 \\({\\times}\\) 16 Words: Transformers for Image Recognition at Scale. In 9th International Conference on Learning Representations (ICLR \u201921), Virtual Event. Retrieved from https:\/\/openreview.net\/forum?id=YicbFdNTTy"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01186"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00165"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2814339"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_3_1_14_2","unstructured":"Andrew G. Howard Menglong Zhu Bo Chen Dmitry Kalenichenko Weijun Wang Tobias Weyand Marco Andreetto and Hartwig Adam. 2017a. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv:1704.04861. Retrieved from http:\/\/arxiv.org\/abs\/1704.04861"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","unstructured":"Tao Huang Lang Huang Shan You Fei Wang Chen Qian and Chang Xu. 2022. LightViT: Towards Light-Weight Convolution-Free Vision Transformers. DOI: 10.48550\/arXiv.2207.05557","DOI":"10.48550\/arXiv.2207.05557"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3478563"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2929447"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654970"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2013.5"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654869"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-014-2000-8"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","unstructured":"Jiashi Li Xin Xia Wei Li Huixia Li Xing Wang Xuefeng Xiao Rui Wang Min Zheng and Xin Pan. 2022a. Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios. DOI: 10.48550\/arXiv.2207.05501","DOI":"10.48550\/arXiv.2207.05501"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","unstructured":"Yanyu Li Geng Yuan Yang Wen Eric Hu Georgios Evangelidis Sergey Tulyakov Yanzhi Wang and Jian Ren. 2022b. EfficientFormer: Vision Transformers at MobileNet Speed. DOI: 10.48550\/arXiv.2206.01191","DOI":"10.48550\/arXiv.2206.01191"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19803-8_3"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01386"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2018.00068"},{"key":"e_1_3_1_30_2","volume-title":"10th International Conference on Learning Representations (ICLR \u201922), Virtual Event","author":"Mehta Sachin","year":"2022","unstructured":"Sachin Mehta and Mohammad Rastegari. 2022a. MobileViT: Light-Weight, General-Purpose, and Mobile-Friendly Vision Transformer. In 10th International Conference on Learning Representations (ICLR \u201922), Virtual Event. Retrieved from https:\/\/openreview.net\/forum?id=vh-0sUt8HlG"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","unstructured":"Sachin Mehta and Mohammad Rastegari. 2022b. Separable Self-Attention for Mobile Vision Transformers. DOI: 10.48550\/arXiv.2206.02680","DOI":"10.48550\/arXiv.2206.02680"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00941"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3329168"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350948"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414031"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3237871"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3237871"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3552485.3554937"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","unstructured":"Junting Pan Adrian Bulat Fuwen Tan Xiatian Zhu Lukasz Dudziak Hongsheng Li Georgios Tzimiropoulos and Brais Mart\u00ednez. 2022. EdgeViTs: Competing Light-Wweight CNNs on Mobile Devices with Vision Transformers. DOI: 10.48550\/arXiv.2205.03436","DOI":"10.48550\/arXiv.2205.03436"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3063592"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","unstructured":"Xu Qin and Zhilin Wang. 2019. NASNet: A Neuron Attention Stage-By-Stage Net for Single Image Deraining. DOI: 10.48550\/arXiv.1912.03151","DOI":"10.48550\/arXiv.1912.03151"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3552484.3555754"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3552484.3555748"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414691"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.foodchem.2023.136309"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.3390\/nu16020200"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1002\/int.23050"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01625"},{"key":"e_1_3_1_50_2","volume-title":"Healthcare","author":"Tahir Ghalib Ahmed","year":"2021","unstructured":"Ghalib Ahmed Tahir and Chu Kiong Loo. 2021. A Comprehensive Survey of Image-Based Food Recognition and Volume Estimation Methods for Dietary Assessment. Healthcare, 9 (2021), 1676."},{"key":"e_1_3_1_51_2","first-page":"6105","volume-title":"Proceedings of the 36th International Conference on Machine Learning, ICML 2019, Proceedings of Machine Learning Research, Vol. 97","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc V. Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, Proceedings of Machine Learning Research, Vol. 97. Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 6105\u20136114. Retrieved from http:\/\/proceedings.mlr.press\/v97\/tan19a.html"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.3390\/math9111245"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","unstructured":"Yehui Tang Kai Han Jianyuan Guo Chang Xu Chao Xu and Yunhe Wang. 2022. GhostNetV2: Enhance Cheap Operation with Long-Range Attention. DOI: 10.48550\/arXiv.2211.12905","DOI":"10.48550\/arXiv.2211.12905"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00764"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00009"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","unstructured":"Kan Wu Jinnian Zhang Houwen Peng Mengchen Liu Bin Xiao Jianlong Fu and Lu Yuan. 2022. TinyViT: Fast Pretraining Distillation for Small Vision Transformers. DOI: 10.48550\/arXiv.2207.10666","DOI":"10.48550\/arXiv.2207.10666"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3549203"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2010.5539907"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01183"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00716"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3680285","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3680285","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:57:21Z","timestamp":1750298241000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3680285"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,29]]},"references-count":59,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2024,10,31]]}},"alternative-id":["10.1145\/3680285"],"URL":"https:\/\/doi.org\/10.1145\/3680285","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,29]]},"assertion":[{"value":"2023-07-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-29","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}