{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T07:57:13Z","timestamp":1772265433470,"version":"3.50.1"},"reference-count":64,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2023,8,25]],"date-time":"2023-08-25T00:00:00Z","timestamp":1692921600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Nature Science Foundation of China","doi-asserted-by":"crossref","award":["61972378, 62125207, U1936203, U19B2040"],"award-info":[{"award-number":["61972378, 62125207, U1936203, U19B2040"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"CAAI-Huawei MindSpore Open Fund"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,1,31]]},"abstract":"<jats:p>\n            Food computing has increasingly received widespread attention in the multimedia field. As a basic task of food computing, food image retrieval has wide applications, that is, food image retrieval can help users to find the desired food from a large number of food images. Besides, the retrieved information can be applied to establish a richer database for the subsequent food content-related recommendation. Food image retrieval aims to achieve better performance on novel categories. Thus, it is worth studying to transfer the embedding ability from the training set to the unseen test set, that is, the generalization of the model. Food is influenced by various factors, such as culture and geography, leading to great differences between domains, such as Asian food and western food. Therefore, it is challenging to study the generalization of the model in food image retrieval. In this article, we improve the classical metric learning framework and propose a generalization-oriented sampling strategy, which boosts the generalization of the model by maximizing the intra-class distance from a proportion of positive pairs to avoid the excessive distance compression in the embedding space. Considering that the existing optimization process is in an opposite direction to our proposed sampling strategy, we further propose an adaptive gradient assignment policy named\n            <jats:italic>gradient-adaptive optimization<\/jats:italic>\n            , which can alleviate the intra-class distance compression during optimization by assigning different gradients to different samples. Extensive evaluation on three popular food image datasets demonstrates the effectiveness of the proposed method. We also experiment on three popular general datasets to prove that solving the problem from the generalization can also improve the performance of general image retrieval. Code is available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/Jiajun-ISIA\/Generalization-oriented-Sampling-and-Loss\">https:\/\/github.com\/Jiajun-ISIA\/Generalization-oriented-Sampling-and-Loss<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3600095","type":"journal-article","created":{"date-parts":[[2023,5,29]],"date-time":"2023-05-29T11:01:16Z","timestamp":1685358076000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Towards Food Image Retrieval via Generalization-Oriented Sampling and Loss Function Design"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-5439-0176","authenticated-orcid":false,"given":"Jiajun","family":"Song","sequence":"first","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-2550-5887","authenticated-orcid":false,"given":"Zhuo","family":"Li","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6668-9208","authenticated-orcid":false,"given":"Weiqing","family":"Min","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1596-4326","authenticated-orcid":false,"given":"Shuqiang","family":"Jiang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,8,25]]},"reference":[{"key":"e_1_3_2_2_2","first-page":"241","volume-title":"European Conference on Computer Vision","author":"Albert G.","year":"2016","unstructured":"G. Albert, A. Jon, R. Jerome, and L. Diane. 2016. Deep image retrieval: Learning global representations for image search. In European Conference on Computer Vision. 241\u2013257."},{"key":"e_1_3_2_3_2","first-page":"46","volume-title":"Italian Conference on Computational Linguistics","author":"Barlacchi G.","year":"2016","unstructured":"G. Barlacchi, A. Abad, E. Rossinelli, and A. Moschitti. 2016. Appetitoso: A search engine for restaurant retrieval based on dishes. In Italian Conference on Computational Linguistics. 46\u201350."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2014.09.044"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10599-4_29"},{"key":"e_1_3_2_6_2","first-page":"677","volume-title":"European Conference on Computer Vision","author":"Brown A.","year":"2020","unstructured":"A. Brown, W. Xie, V. Kalogeiton, and A. Zisserman. 2020. Smooth-AP: Smoothing the path towards large-scale image retrieval. In European Conference on Computer Vision. 677\u2013694."},{"key":"e_1_3_2_7_2","first-page":"451:1\u2013451:12","volume-title":"Proc. CHI Conference on Human Factors in Computing Systems","author":"Chang M.","year":"2018","unstructured":"M. Chang, L. Guillain, H. Jung, V. Hare, J. Kim, and M. Agrawala. 2018. RecipeScape: An interactive tool for analyzing cooking instructions at scale. In Proc. CHI Conference on Human Factors in Computing Systems. 451:1\u2013451:12."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3375786"},{"key":"e_1_3_2_9_2","first-page":"32","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Chen J.","year":"2016","unstructured":"J. Chen and C. Ngo. 2016. Deep-based ingredient recognition for cooking recipe retrieval. In Proceedings of the ACM International Conference on Multimedia. 32\u201341."},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"1020","DOI":"10.1145\/3240508.3240627","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Chen J.","year":"2018","unstructured":"J. Chen, C. Ngo, F. Feng, and T. Chua. 2018. Deep understanding of cooking procedure for cross-modal recipe retrieval. In Proceedings of the ACM International Conference on Multimedia. 1020\u20131028."},{"key":"e_1_3_2_11_2","first-page":"426","volume-title":"International Conference on Image Analysis and Processing","author":"Ciocca G.","year":"2017","unstructured":"G. Ciocca, P. Napoletano, and R. Schettini. 2017. Learning CNN-based features for retrieval of food images. In International Conference on Image Analysis and Processing. 426\u2013434."},{"key":"e_1_3_2_12_2","first-page":"1153","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Cui Y.","year":"2016","unstructured":"Y. Cui, F. Zhou, Y. Lin, and S. Belongie. 2016. Fine-grained categorization and dataset bootstrapping using deep metric learning with humans in the loop. In IEEE Conference on Computer Vision and Pattern Recognition. 1153\u20131162."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"issue":"2","key":"e_1_3_2_14_2","first-page":"43:1\u201343:22","article-title":"From selective deep convolutional features to compact binary representations for image retrieval","volume":"15","author":"Do T.","year":"2019","unstructured":"T. Do, T. Hoang, D. Tan, H. Le, T. Nguyen, and N. Cheung. 2019. From selective deep convolutional features to compact binary representations for image retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 15, 2 (2019), 43:1\u201343:22.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_15_2","volume-title":"International Conference on Learning Representations","author":"Dosovitskiy A.","year":"2021","unstructured":"A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations."},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compbiomed.2016.07.006"},{"key":"e_1_3_2_17_2","first-page":"1460","volume-title":"AAAI","author":"Gu G.","year":"2021","unstructured":"G. Gu, B. Ko, and H. Kim. 2021. Proxy synthesis: Learning with synthetic classes for deep metric learning. In AAAI. 1460\u20131468."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.100"},{"key":"e_1_3_2_19_2","article-title":"In defense of the triplet loss for person re-identification","author":"Hermans A.","year":"2017","unstructured":"A. Hermans, L. Beyer, and B. Leibe. 2017. In defense of the triplet loss for person re-identification. arXiv preprint arxiv: 1703.07737 (2017).","journal-title":"arXiv preprint arxiv: 1703.07737"},{"issue":"10","key":"e_1_3_2_20_2","doi-asserted-by":"crossref","first-page":"2836","DOI":"10.1109\/TMM.2018.2814339","article-title":"Personalized classifier for food image recognition","volume":"20","author":"Horiguchi S.","year":"2018","unstructured":"S. Horiguchi, S. Amano, M. Ogawa, and K. Aizawa. 2018. Personalized classifier for food image recognition. IEEE Trans. Multimedia 20, 10 (2018), 2836\u20132848.","journal-title":"IEEE Trans. Multimedia"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.242"},{"key":"e_1_3_2_22_2","first-page":"1654","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Ji Xin","year":"2017","unstructured":"Xin Ji, Wei Wang, Meihui Zhang, and Yang Yang. 2017. Cross-domain image retrieval with attention modeling. In Proceedings of the ACM International Conference on Multimedia. 1654\u20131662."},{"issue":"3","key":"e_1_3_2_23_2","doi-asserted-by":"crossref","first-page":"87:1\u201387:20","DOI":"10.1145\/3391624","article-title":"Few-shot food recognition via multi-view representation learning","volume":"16","author":"Jiang S.","year":"2020","unstructured":"S. Jiang, W. Min, Y. Lyu, and L. Liu. 2020. Few-shot food recognition via multi-view representation learning. ACM Transactions on Multimedia Computing, Communications and Applications 16, 3 (2020), 87:1\u201387:20.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_24_2","first-page":"1","volume-title":"International Conference on Learning Representations","author":"Kingma D.","year":"2015","unstructured":"D. Kingma and J. Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations. 1\u201315."},{"key":"e_1_3_2_25_2","first-page":"554","volume-title":"IEEE International Conference on Computer Vision Workshops","author":"Krause J.","year":"2013","unstructured":"J. Krause, M. Stark, J. Deng, and F. Li. 2013. 3D object representations for fine-grained categorization. In IEEE International Conference on Computer Vision Workshops. 554\u2013561."},{"issue":"4","key":"e_1_3_2_26_2","first-page":"134:1\u2013134:22","article-title":"Part-based structured representation learning for person re-identification","volume":"16","author":"Li Y.","year":"2021","unstructured":"Y. Li, H. Yao, T. Zhang, and C. Xu. 2021. Part-based structured representation learning for person re-identification. ACM Transactions on Multimedia Computing, Communications and Applications 16, 4 (2021), 134:1\u2013134:22.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_27_2","article-title":"Rethinking ranking-based loss functions: Only penalizing negative instances before positive ones is enough","author":"Li Z.","year":"2021","unstructured":"Z. Li, W. Min, J. Song, Y. Zhu, and S. Jiang. 2021. Rethinking ranking-based loss functions: Only penalizing negative instances before positive ones is enough. arXiv preprint arXiv:2102.04640 (2021).","journal-title":"arXiv preprint arXiv:2102.04640"},{"issue":"86","key":"e_1_3_2_28_2","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Maaten L.","year":"2008","unstructured":"L. Maaten and G. Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 86 (2008), 2579\u20132605.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_29_2","first-page":"2859","volume-title":"IEEE International Conference on Computer Vision","author":"Manmatha R.","year":"2017","unstructured":"R. Manmatha, Chao-Yuan Wu, Alexander Smola, and Philipp Krahenbuhl. 2017. Sampling matters in deep embedding learning. In IEEE International Conference on Computer Vision. 2859\u20132867."},{"key":"e_1_3_2_30_2","first-page":"35","volume-title":"International ACM SIGIR Conference","author":"Micael C.","year":"2018","unstructured":"C. Micael, C. R\u00e9mi, P. David, S. Laure, T. Nicolas, and C. Matthieu. 2018. Cross-modal retrieval in the cooking context: Learning semantic text-image embeddings. In International ACM SIGIR Conference. 35\u201344."},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3009620"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2958761"},{"issue":"5","key":"e_1_3_2_33_2","first-page":"36","article-title":"A survey on food computing","volume":"52","author":"Min W.","year":"2019","unstructured":"W. Min, S. Jiang, L. Liu, Y. Rui, and R. Jain. 2019. A survey on food computing. Comput. Surveys 52, 5 (2019), 36.","journal-title":"Comput. Surveys"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2016.2639382"},{"key":"e_1_3_2_35_2","first-page":"393","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Min W.","year":"2020","unstructured":"W. Min, L. Liu, Z. Wang, Z. Luo, X. Wei, X. Wei, and Jiang S.2020. ISIA Food-500: A dataset for large-scale food recognition via stacked global-local attention network. In Proceedings of the ACM International Conference on Multimedia. 393\u2013401."},{"key":"e_1_3_2_36_2","first-page":"360","volume-title":"IEEE International Conference on Computer Vision","author":"Movshovitz-Attias Y.","year":"2017","unstructured":"Y. Movshovitz-Attias, A. Toshev, T. Leung, S. Ioffe, and S. Singh. 2017. No fuss distance metric learning using proxies. In IEEE International Conference on Computer Vision. 360\u2013368."},{"key":"e_1_3_2_37_2","article-title":"Siamese network of deep Fisher-vector descriptors for image retrieval","author":"Ong E.","year":"2017","unstructured":"E. Ong, S. Husain, and M. Bober. 2017. Siamese network of deep Fisher-vector descriptors for image retrieval. arXiv preprint arXiv:1702.00338 (2017).","journal-title":"arXiv preprint arXiv:1702.00338"},{"issue":"3","key":"e_1_3_2_38_2","first-page":"36:1\u201336:21","article-title":"Mobile multi-food recognition using deep learning","volume":"13","author":"Pouladzadeh P.","year":"2017","unstructured":"P. Pouladzadeh and S. Shirmohammadi. 2017. Mobile multi-food recognition using deep learning. ACM Transactions on Multimedia Computing, Communications and Applications 13, 3s (2017), 36:1\u201336:21.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_39_2","first-page":"6449","volume-title":"IEEE International Conference on Computer Vision","author":"Qian Q.","year":"2019","unstructured":"Q. Qian, L. Shang, B. Sun, J. Hu, T. Tacoma, H. Li, and R. Jin. 2019. SoftTriple loss: Deep metric learning without triplet sampling. In IEEE International Conference on Computer Vision. 6449\u20136457."},{"key":"e_1_3_2_40_2","author":"Roth K.","year":"2019","unstructured":"K. Roth and B. Brattoli. 2019. https:\/\/github.com\/Confusezius\/Deep-Metric-Learning-Baselines. (2019).","journal-title":"https:\/\/github.com\/Confusezius\/Deep-Metric-Learning-Baselines"},{"key":"e_1_3_2_41_2","first-page":"8242","volume-title":"ICML","author":"Roth K.","year":"2020","unstructured":"K. Roth, T. Milbich, S. Sinha, P. Gupta, B. Ommer, and J. Cohen. 2020. Revisiting training strategies and generalization performance in deep metric learning. In ICML. 8242\u20138252."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.327"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"e_1_3_2_44_2","first-page":"165","volume-title":"IEEE 3rd International Conference on Multimedia Big Data","author":"Shimoda W.","year":"2017","unstructured":"W. Shimoda and K. Yanai. 2017. Learning food image similarity for food image retrieval. In IEEE 3rd International Conference on Multimedia Big Data. 165\u2013168."},{"key":"e_1_3_2_45_2","first-page":"1857","volume-title":"Advances in Neural Information Processing Systems","author":"Sohn K.","year":"2016","unstructured":"K. Sohn. 2016. Improved deep metric learning with multi-class n-pair loss objective. In Advances in Neural Information Processing Systems. 1857\u20131865."},{"key":"e_1_3_2_46_2","first-page":"4004","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Song H.","year":"2016","unstructured":"H. Song, Y. Xiang, S. Jegelka, and S. Savarese. 2016. Deep metric learning via lifted structured feature embedding. In IEEE Conference on Computer Vision and Pattern Recognition. 4004\u20134012."},{"key":"e_1_3_2_47_2","first-page":"6398","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Sun Y.","year":"2020","unstructured":"Y. Sun, C. Cheng, Y. Zhang, C. Zhang, L. Zheng, Z. Wang, and Y. Wei. 2020. Circle loss: A unified perspective of pair similarity optimization. In IEEE Conference on Computer Vision and Pattern Recognition. 6398\u20136407."},{"issue":"4","key":"e_1_3_2_48_2","doi-asserted-by":"crossref","first-page":"98:1\u201398:21","DOI":"10.1145\/3501405","article-title":"Harmonious multi-branch network for person re-identification with harder triplet loss","volume":"18","author":"Tang Z.","year":"2022","unstructured":"Z. Tang and J. Huang. 2022. Harmonious multi-branch network for person re-identification with harder triplet loss. ACM Transactions on Multimedia Computing, Communications and Applications 18, 4 (2022), 98:1\u201398:21.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_49_2","first-page":"1","volume-title":"IEEE Information Theory Workshop","author":"Tishby N.","year":"2015","unstructured":"N. Tishby and N. Zaslavsky. 2015. Deep learning and the information bottleneck principle. In IEEE Information Theory Workshop. 1\u20135."},{"key":"e_1_3_2_50_2","first-page":"4170","volume-title":"NIPS","author":"Ustinova E.","year":"2016","unstructured":"E. Ustinova and V. Lempitsky. 2016. Learning deep embeddings with histogram loss. In NIPS. 4170\u20134178."},{"key":"e_1_3_2_51_2","first-page":"6438","volume-title":"International Conference on Machine Learning","author":"Verma V.","year":"2019","unstructured":"V. Verma, A. Lamb, C. Beckham, A. Najafi, I. Mitliagkas, D. Lopez-Paz, and Y. Bengio. 2019. Manifold mixup: Better representations by interpolating hidden states. In International Conference on Machine Learning. 6438\u20136447."},{"key":"e_1_3_2_52_2","volume-title":"The Caltech-UCSD Birds-200-2011 Dataset","author":"Wah C.","year":"2011","unstructured":"C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. 2011. The Caltech-UCSD Birds-200-2011 Dataset. Technical Report CNS-TR-2011-001. California Institute of Technology."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3083109"},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","first-page":"979","DOI":"10.1145\/1367497.1367629","volume-title":"Proceedings of the ACM International Conference on World Wide Web","author":"Wang L.","year":"2008","unstructured":"L. Wang, Q. Li, N. Li, G. Dong, and Y. Yang. 2008. Substructure similarity measurement in Chinese recipes. In Proceedings of the ACM International Conference on World Wide Web. 979\u2013988."},{"issue":"1","key":"e_1_3_2_55_2","doi-asserted-by":"crossref","first-page":"33:1\u201333:19","DOI":"10.1145\/3418211","article-title":"Market2Dish: Health-aware food recommendation","volume":"17","author":"Wang W.","year":"2021","unstructured":"W. Wang, L. Duan, H. Jiang, P. Jing, X. Song, and L. Nie. 2021. Market2Dish: Health-aware food recommendation. ACM Transactions on Multimedia Computing, Communications and Applications 17, 1 (2021), 33:1\u201333:19.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"issue":"3","key":"e_1_3_2_56_2","doi-asserted-by":"crossref","first-page":"84:1\u201384:19","DOI":"10.1145\/3340262","article-title":"Eigenvector-based distance metric learning for image classification and retrieval","volume":"15","author":"Wang Z.","year":"2019","unstructured":"Z. Wang, Y. Li, R. Hong, and X. Tian. 2019. Eigenvector-based distance metric learning for image classification and retrieval. ACM Transactions on Multimedia Computing, Communications and Applications 15, 3 (2019), 84:1\u201384:19.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_57_2","first-page":"2840","volume-title":"IEEE International Conference on Computer Vision","author":"Wu C.","year":"2017","unstructured":"C. Wu, R. Manmatha, A. Smola, and P. Krahenbuhl. 2017. Sampling matters in deep embedding learning. In IEEE International Conference on Computer Vision. 2840\u20132848."},{"issue":"4","key":"e_1_3_2_58_2","first-page":"90:1\u201390:23","article-title":"Improving feature discrimination for object tracking by structural-similarity-based metric learning","volume":"18","author":"Wu J.","year":"2022","unstructured":"J. Wu, J. Jiang, M. Qi, C. Chen, and Y. Liu. 2022. Improving feature discrimination for object tracking by structural-similarity-based metric learning. ACM Transactions on Multimedia Computing, Communications and Applications 18, 4 (2022), 90:1\u201390:23.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_59_2","first-page":"254","volume-title":"IEEE International Symposium on Multimedia","author":"Xie H.","year":"2011","unstructured":"H. Xie, L. Yu, and Q. Li. 2011. A hybrid semantic item model for recipe search by example. In IEEE International Symposium on Multimedia. 254\u2013259."},{"issue":"3","key":"e_1_3_2_60_2","first-page":"92:1\u201392:18","article-title":"Lifelog image retrieval based on semantic relevance mapping","volume":"17","author":"Xu Q.","year":"2021","unstructured":"Q. Xu, A. Molino, J. Lin, F. Fang, V. Subbaraju, L. Li, and J. Lim. 2021. Lifelog image retrieval based on semantic relevance mapping. ACM Transactions on Multimedia Computing, Communications and Applications 17, 3 (2021), 92:1\u201392:18.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3001527"},{"key":"e_1_3_2_62_2","first-page":"91","volume-title":"British Machine Vision Conference","author":"Zhai A.","year":"2019","unstructured":"A. Zhai and H. Wu. 2019. Classification is a strong baseline for deep metric learning. In British Machine Vision Conference. 91."},{"issue":"1","key":"e_1_3_2_63_2","first-page":"25:1\u201325:15","article-title":"Hybrid modality metric learning for visible-infrared person re-identification","volume":"18","author":"Zhang L.","year":"2022","unstructured":"L. Zhang, H. Guo, K. Zhu, H. Qiao, G. Huang, S. Zhang, H. Zhang, J. Sun, and J. Wang. 2022. Hybrid modality metric learning for visible-infrared person re-identification. ACM Transactions on Multimedia Computing, Communications and Applications 18, 1s (2022), 25:1\u201325:15.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"issue":"1","key":"e_1_3_2_64_2","first-page":"24:1\u201324:3","article-title":"Introduction to the special issue on fine-grained visual recognition","volume":"18","author":"Zhang S.","year":"2022","unstructured":"S. Zhang, G. Li, W. Zhang, Q. Huang, Ti. Huang, M. Shah, and N. Sebe. 2022. Introduction to the special issue on fine-grained visual recognition. ACM Transactions on Multimedia Computing, Communications and Applications 18, 1s (2022), 24:1\u201324:3.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3123474"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3600095","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3600095","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:36:50Z","timestamp":1750178210000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3600095"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,25]]},"references-count":64,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,1,31]]}},"alternative-id":["10.1145\/3600095"],"URL":"https:\/\/doi.org\/10.1145\/3600095","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,25]]},"assertion":[{"value":"2022-06-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-22","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}