{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T16:38:06Z","timestamp":1760027886452},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,7]]},"abstract":"<jats:p>Visual contents, including images and videos, are dominant on the Internet today. The conventional search engine is mainly designed for textual documents, which must be extended to process and manage increasingly high volumes of visual data objects.<\/jats:p>\n          <jats:p>In this paper, we present Mixer, an effective system to identify and analyze visual contents and to extract their features for data retrievals, aiming at addressing two critical issues: (1) efficiently and timely understanding visual contents, (2) retrieving them at high precision and recall rates without impairing the performance. In Mixer, the visual objects are categorized into different classes, each of which has representative visual features. Subsystems for model production and model execution are developed. Two retrieval layers are designed and implemented for images and videos, respectively. In this way, we are able to perform aggregation retrievals of the two types in efficient ways. The experiments with Baidu's production workloads and systems show that Mixer halves the model production time and raises the feature production throughput by 9.14x. Mixer also achieves the precision and recall of video retrievals at 95% and 97%, respectively. Mixer has been in its daily operations, which makes the search engine highly scalable for visual contents at a low cost. Having observed productivity improvement of upper-level applications in the search engine, we believe our system framework would generally benefit other data processing applications.<\/jats:p>","DOI":"10.14778\/3476311.3476371","type":"journal-article","created":{"date-parts":[[2021,10,28]],"date-time":"2021-10-28T22:48:56Z","timestamp":1635461336000},"page":"2906-2917","source":"Crossref","is-referenced-by-count":6,"title":["Mixer"],"prefix":"10.14778","volume":"14","author":[{"given":"An","family":"Qin","sequence":"first","affiliation":[{"name":"Baidu, Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mengbai","family":"Xiao","sequence":"additional","affiliation":[{"name":"Shandong University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongwei","family":"Wu","sequence":"additional","affiliation":[{"name":"Baidu, Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinjie","family":"Huang","sequence":"additional","affiliation":[{"name":"Baidu, Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaodong","family":"Zhang","sequence":"additional","affiliation":[{"name":"The Ohio State University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,28]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"https:\/\/kubernetes.io\/. Accessed","year":"2021"},{"key":"e_1_2_1_2_1","volume-title":"https:\/\/developer.nvidia.com\/tensorrt. Accessed","author":"NVIDIA","year":"2021"},{"key":"e_1_2_1_3_1","first-page":"1466","volume-title":"Proceedings of the 35th IEEE International Conference on Data Engineering","author":"Anderson Michael R."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/tpami.2014.2361319"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01258-8_13"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415498"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2072298.2072484"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472621"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3323873.3325018"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/1251254.1251264"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00482"},{"key":"e_1_2_1_12_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL]","author":"Devlin Jacob","year":"2019"},{"key":"e_1_2_1_13_1","unstructured":"Yuning Du Chenxia Li Ruoyu Guo Xiaoting Yin Weiwei Liu Jun Zhou Yifan Bai Zilin Yu Yehua Yang Qingqing Dang and Haoshuang Wang. 2020. PP-OCR: A Practical Ultra Lightweight OCR System. arXiv:2009.09941 [cs.CV]  Yuning Du Chenxia Li Ruoyu Guo Xiaoting Yin Weiwei Liu Jun Zhou Yifan Bai Zilin Yu Yehua Yang Qingqing Dang and Haoshuang Wang. 2020. PP-OCR: A Practical Ultra Lightweight OCR System. arXiv:2009.09941 [cs.CV]"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE51399.2021.00092"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.379"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-51811-4_21"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557064"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.123"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/3291168.3291188"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/3360092"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2011.5946540"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00538"},{"key":"e_1_2_1_23_1","unstructured":"Jeff Johnson Matthijs Douze and Herv\u00c3I J\u00c3Tgou. 2017. Billion-Scale Similarity Search with GPUs. arXiv:1702.08734 [cs.CV]  Jeff Johnson Matthijs Douze and Herv\u00c3I J\u00c3Tgou. 2017. Billion-Scale Similarity Search with GPUs. arXiv:1702.08734 [cs.CV]"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.57"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.14778\/3372716.3372725"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the 9th Biennial Conference on Innovative Data Systems Research","author":"Kang Daniel","year":"2019"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137628.3137664"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2905741"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2017.49"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2015.7301269"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence","author":"Long Xiang","year":"2018"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00817"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.5555\/850924.851523"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183751"},{"key":"e_1_2_1_35_1","unstructured":"Abdelrahman Mohamed Dmytro Okhonko and Luke Zettlemoyer. 2020. Transformers with Convolutional Context for ASR. arXiv:1904.11660 [cs.CL]  Abdelrahman Mohamed Dmytro Okhonko and Luke Zettlemoyer. 2020. Transformers with Convolutional Context for ASR. arXiv:1904.11660 [cs.CL]"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/2733004.2733078"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2019.00195"},{"key":"e_1_2_1_38_1","volume-title":"Report","year":"2019"},{"key":"e_1_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Joseph Redmon Santosh Divvala Ross Girshick and Ali Farhadi. 2016. You Only Look Once: Unified Real-Time Object Detection. arXiv:1506.02640 [cs.CV]  Joseph Redmon Santosh Divvala Ross Girshick and Ali Farhadi. 2016. You Only Look Once: Unified Real-Time Object Detection. arXiv:1506.02640 [cs.CV]","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_2_1_40_1","unstructured":"Tim Salimans Jonathan Ho Xi Chen Szymon Sidor and Ilya Sutskever. 2017. Evolution Strategies as a Scalable Alternative to Reinforcement Learning. arXiv:1703.03864 [stat.ML]  Tim Salimans Jonathan Ho Xi Chen Szymon Sidor and Ilya Sutskever. 2017. Evolution Strategies as a Scalable Alternative to Reinforcement Learning. arXiv:1703.03864 [stat.ML]"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874021"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2072298.2072354"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2013.2271746"},{"key":"e_1_2_1_44_1","first-page":"405","volume-title":"Proceedings of the 16th European Conference on Computer Vision","author":"Wang Jianyi","year":"2020"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732967.2732976"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2019.00247"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415541"},{"key":"e_1_2_1_49_1","first-page":"2027","volume-title":"Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition","author":"Wieschollek Patrick"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00632"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/1291233.1291280"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3302424.3303971"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.226"},{"key":"e_1_2_1_54_1","unstructured":"Quanming Yao Mengshuo Wang Yuqiang Chen Wenyuan Dai Yu-Feng Li Wei-Wei Tu Qiang Yang and Yang Yu. 2019. Taking Human out of Learning Applications: A Survey on Automated Machine Learning. arXiv:1810.13306 [cs.AI]  Quanming Yao Mengshuo Wang Yuqiang Chen Wenyuan Dai Yu-Feng Li Wei-Wei Tu Qiang Yang and Yang Yu. 2019. Taking Human out of Learning Applications: A Survey on Automated Machine Learning. arXiv:1810.13306 [cs.AI]"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536206.2536210"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.14778\/2809974.2809984"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3330345.3332147"},{"key":"e_1_2_1_58_1","first-page":"1","volume-title":"Proceedings of the 36th International Symposium on Computational Geometry","author":"Zhang Simon","year":"2020"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.14778\/3372716.3372721"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3476311.3476371","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:34:10Z","timestamp":1672227250000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3476311.3476371"}},"subtitle":["efficiently understanding and retrieving visual content at web-scale"],"short-title":[],"issued":{"date-parts":[[2021,7]]},"references-count":59,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2021,7]]}},"alternative-id":["10.14778\/3476311.3476371"],"URL":"https:\/\/doi.org\/10.14778\/3476311.3476371","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,7]]}}}