{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T10:25:31Z","timestamp":1783419931196,"version":"3.54.6"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2020,7,14]],"date-time":"2020-07-14T00:00:00Z","timestamp":1594684800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Program for Special Support of Eminent Professionals"},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61532018, 61972378, and U19B2040"],"award-info":[{"award-number":["61532018, 61972378, and U19B2040"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100018619","name":"National Program for Support of Top-notch Young Professionals","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100018619","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Beijing Natural Science Foundation","award":["L182054"],"award-info":[{"award-number":["L182054"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2020,8,31]]},"abstract":"<jats:p>This article considers the problem of few-shot learning for food recognition. Automatic food recognition can support various applications, e.g., dietary assessment and food journaling. Most existing works focus on food recognition with large numbers of labelled samples, and fail to recognize food categories with few samples. To address this problem, we propose a Multi-View Few-Shot Learning (MVFSL) framework to explore additional ingredient information for few-shot food recognition. Besides category-oriented deep visual features, we introduce ingredient-supervised deep network to extract ingredient-oriented features. As general and intermediate attributes of food, ingredient-oriented features are informative and complementary to category-oriented features, and thus they play an important role in improving food recognition. Particularly in few-shot food recognition, ingredient information can bridge the gap between disjoint training categories and test categories. To take advantage of ingredient information, we fuse these two kinds of features by first combining their feature maps from their respective deep networks and then convolving combined feature maps. Such convolution is further incorporated into a multi-view relation network, which is capable of comparing pairwise images to enable fine-grained feature learning. MVFSL is trained in an end-to-end fashion for joint optimization on two types of feature learning subnetworks and relation subnetworks. Extensive experiments on different food datasets have consistently demonstrated the advantage of MVFSL in multi-view feature fusion. Furthermore, we extend another two types of networks, namely, Siamese Network and Matching Network, by introducing ingredient information for few-shot food recognition. Experimental results have also shown that introducing ingredient information into these two networks can improve the performance of few-shot food recognition.<\/jats:p>","DOI":"10.1145\/3391624","type":"journal-article","created":{"date-parts":[[2020,7,7]],"date-time":"2020-07-07T12:38:12Z","timestamp":1594125492000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":40,"title":["Few-shot Food Recognition via Multi-view Representation Learning"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1596-4326","authenticated-orcid":false,"given":"Shuqiang","family":"Jiang","sequence":"first","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weiqing","family":"Min","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yongqiang","family":"Lyu","sequence":"additional","affiliation":[{"name":"Qingdao KingAgroot Precision Agriculture Technology Co., Ltd, Qingdao, Shandong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Linhu","family":"Liu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,7,14]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2013.2271474"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3084142"},{"key":"e_1_2_1_3_1","volume-title":"Advances in Neural Information Processing Systems","author":"Andrychowicz Marcin"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the IEEE International Conference on Data Mining Workshop. 1196--1203","author":"Ao Shuang"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2015.117"},{"key":"e_1_2_1_6_1","volume-title":"Advances in Neural Information Processing Systems","author":"Bertinetto Luca"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2015.83"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-70742-6_37"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10599-4_29"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00429"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2964315"},{"key":"e_1_2_1_12_1","unstructured":"Xin Chen Hua Zhou Yu Zhu and Liang Diao. 2017. ChineseFoodNet: A large-scale image dataset for Chinese food recognition. arXiv preprint arXiv:1705.02743.  Xin Chen Hua Zhou Yu Zhu and Liang Diao. 2017. ChineseFoodNet: A large-scale image dataset for Chinese food recognition. arXiv preprint arXiv:1705.02743."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2016.2642792"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351147"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/tpami.2006.79"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1933--1941","author":"Feichtenhofer C."},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the International Conference on Machine Learning. 1126--1135","author":"Finn Chelsea","year":"2017"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.476"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00459"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2016.2601622"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2014.2374218"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/108"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2904460"},{"key":"e_1_2_1_24_1","first-page":"1941","article-title":"SeqViews2SeqLabels: Learning 3D global features via aggregating sequential views by RNN with attention","volume":"28","author":"Han Zhizhong","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2016.2614861"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2261--2269","author":"Huang G."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2929447"},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the IEEE International Conference on Image Processing. 285--288","author":"Joutou Taichi","year":"2010"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654970"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2638728.2641339"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-04117-9_38"},{"key":"e_1_2_1_33_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2014"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the International Conference on Machine Learning","volume":"2","author":"Koch Gregory","year":"2015"},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the Annual Meeting of the Cognitive Science Society","volume":"33","author":"Lake Brenden","year":"2011"},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems. 2526--2534","author":"Lake Brenden M."},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of the International Conference on Learning Representations.","author":"Liu Yanbin","year":"2019"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.3390\/app7020189"},{"key":"e_1_2_1_39_1","unstructured":"J. Marin A. Biswas F. Ofli N. Hynes A. Salvador Y. Aytar I. Weber and A. Torralba. 2019. Recipe1M+: A dataset for learning cross-modal embeddings for cooking recipes and food images. IEEE Trans. Pattern Anal. Mach. Intell. (2019) 1. Early Access.  J. Marin A. Biswas F. Ofli N. Hynes A. Salvador Y. Aytar I. Weber and A. Torralba. 2019. Recipe1M+: A dataset for learning cross-modal embeddings for cooking recipes and food images. IEEE Trans. Pattern Anal. Mach. Intell. (2019) 1. Early Access."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2018.00068"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2015.70"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-789X.2011.00936.x"},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1233--1241","author":"Meyers Austin"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2759499"},{"key":"e_1_2_1_45_1","volume-title":"A survey on food computing. ACM Comput. Surv. 52, 5","author":"Min Weiqing","year":"2019"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2016.2639382"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350948"},{"key":"e_1_2_1_48_1","unstructured":"Tsendsuren Munkhdalai and Hong Yu. 2017. Meta networks. arXiv preprint arXiv:1703.00837.  Tsendsuren Munkhdalai and Hong Yu. 2017. Meta networks. arXiv preprint arXiv:1703.00837."},{"key":"e_1_2_1_49_1","article-title":"Deep learning for mobile multimedia: A survey","volume":"13","author":"Ota Kaoru","year":"2017","journal-title":"ACM Trans. Multimedia Comput. Commun. Appl."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3063592"},{"key":"e_1_2_1_51_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7229--7238","author":"Qiao Siyuan"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.327"},{"key":"e_1_2_1_53_1","volume-title":"Lillicrap","author":"Santoro Adam","year":"2016"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.74"},{"key":"e_1_2_1_55_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556.  Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556."},{"key":"e_1_2_1_56_1","volume-title":"Advances in Neural Information Processing Systems","author":"Snell Jake"},{"key":"e_1_2_1_57_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1199--1208","author":"Sung Flood"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/2986035.2986044"},{"key":"e_1_2_1_59_1","volume-title":"Lifelong Learning Algorithms","author":"Thrun Sebastian"},{"key":"e_1_2_1_60_1","unstructured":"Oriol Vinyals Charles Blundell Timothy Lillicrap Koray Kavukcuoglu and Daan Wierstra. 2016. Matching networks for one shot learning. In Advances in Neural Information Processing Systems. 3630--3638.  Oriol Vinyals Charles Blundell Timothy Lillicrap Koray Kavukcuoglu and Daan Wierstra. 2016. Matching networks for one shot learning. In Advances in Neural Information Processing Systems. 3630--3638."},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2015.2438717"},{"key":"e_1_2_1_62_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2249--2256","author":"Yang Shulin","year":"2010"},{"key":"e_1_2_1_63_1","volume-title":"Zeiler and Rob Fergus","author":"Matthew","year":"2013"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2567393"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10590-1_54"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.3390\/su9050856"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3391624","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3391624","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:41:41Z","timestamp":1750200101000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3391624"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,7,14]]},"references-count":66,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2020,8,31]]}},"alternative-id":["10.1145\/3391624"],"URL":"https:\/\/doi.org\/10.1145\/3391624","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,7,14]]},"assertion":[{"value":"2019-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-03-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-07-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}