{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T02:34:30Z","timestamp":1783737270401,"version":"3.55.0"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,7]]},"abstract":"<jats:p>Similarity search is a core operation of many critical applications, involving massive collections of high-dimensional (high-d) objects. Objects can be data series, text, multimedia, graphs, database tables or deep network embeddings. In this tutorial, we revisit the similarity search problem in light of the recent advances in the field and the new big data landscape. We discuss key data science applications that require efficient high-d similarity search, we survey recent approaches and share surprising insights about their strengths and weaknesses, and we discuss open research problems, including the directions of AI-driven, progressive, and distributed high-d similarity search.<\/jats:p>","DOI":"10.14778\/3476311.3476407","type":"journal-article","created":{"date-parts":[[2021,10,28]],"date-time":"2021-10-28T22:48:56Z","timestamp":1635461336000},"page":"3198-3201","source":"Crossref","is-referenced-by-count":40,"title":["New trends in high-D vector similarity search"],"prefix":"10.14778","volume":"14","author":[{"given":"Karima","family":"Echihabi","sequence":"first","affiliation":[{"name":"Mohammed VI Polytechnic University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kostas","family":"Zoumpatianos","sequence":"additional","affiliation":[{"name":"Harvard University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Themis","family":"Palpanas","sequence":"additional","affiliation":[{"name":"Universit\u00e9 de Paris &amp; IUF"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,10,28]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/645504.656414"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397536.3426358"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290353"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.14778\/3204028.3204034"},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Martin Aum\u00fcller Erik Bernhardsson and Alexander Faithfull. 2017. ANN-Benchmarks: A Benchmarking Tool for Approximate Nearest Neighbor Algorithms. In SISAP.  Martin Aum\u00fcller Erik Bernhardsson and Alexander Faithfull. 2017. ANN-Benchmarks: A Benchmarking Tool for Approximate Nearest Neighbor Algorithms. In SISAP .","DOI":"10.1007\/978-3-319-68474-1_3"},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"A. Babenko and V. Lempitsky. 2015. The Inverted Multi-Index. TPAMI 37 6 (2015).  A. Babenko and V. Lempitsky. 2015. The Inverted Multi-Index. TPAMI 37 6 (2015).","DOI":"10.1109\/TPAMI.2014.2361319"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2396761.2398596"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/93605.98741"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/645503.656271"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.14778\/3467861.3467863"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/829502.830043"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/1577069.1577092"},{"key":"e_1_2_1_13_1","volume-title":"Raftery","author":"Byers Simon","year":"1998"},{"key":"e_1_2_1_14_1","unstructured":"Deng Cai Xiuye Gu and Chaoqi Wang. 2017. A Revisit on Deep Hashings for Large-scale Content Based Image Retrieval. arXiv:1711.06016 [cs.CV]  Deng Cai Xiuye Gu and Chaoqi Wang. 2017. A Revisit on Deep Hashings for Large-scale Content Based Image Retrieval. arXiv:1711.06016 [cs.CV]"},{"key":"e_1_2_1_15_1","unstructured":"Paolo Ciaccia and Marco Patella. 2000. PAC Nearest Neighbor Queries: Approximate and Controlled Search in High-Dimensional and Metric Spaces. In ICDE.  Paolo Ciaccia and Marco Patella. 2000. PAC Nearest Neighbor Queries: Approximate and Controlled Search in High-Dimensional and Metric Spaces. In ICDE ."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/645923.671005"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/997817.997857"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1963405.1963487"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/2018783"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/3236187.3269461"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3384402"},{"key":"e_1_2_1_22_1","volume-title":"Big Sequence Management: on Scalability","author":"Echihabi Karima"},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","volume-title":"Scalable Machine Learning on High-Dimensional Vectors","author":"Echihabi Karima","DOI":"10.1145\/3405962.3405989"},{"key":"e_1_2_1_24_1","unstructured":"Karima Echihabi Kostas Zoumpatianos and Themis Palpanas. 2021. Big Sequence Management: Scaling up and Out. In EDBT.  Karima Echihabi Kostas Zoumpatianos and Themis Palpanas. 2021. Big Sequence Management: Scaling up and Out. In EDBT ."},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Karima Echihabi Kostas Zoumpatianos and Themis Palpanas. 2021. High-Dimensional Similarity Search for Scalable Data Science (ICDE).  Karima Echihabi Kostas Zoumpatianos and Themis Palpanas. 2021. High-Dimensional Similarity Search for Scalable Data Science (ICDE) .","DOI":"10.1109\/ICDE51399.2021.00268"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/3282495.3282498"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.14778\/3368289.3368303"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/191843.191925"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/354756.354820"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.14778\/3303753.3303754"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213898"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.240"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389751"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/971697.602266"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.5555\/3042573.3042582"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/2850469.2850470"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/276698.276876"},{"key":"e_1_2_1_38_1","volume-title":"Ravishankar Krishnawamy, and Rohan Kadekodi.","author":"Subramanya Suhas Jayaram","year":"2019"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.57"},{"key":"e_1_2_1_40_1","volume-title":"Torben Bach Pedersen, and Christian Thomsen","author":"Jensen Saren Kejser","year":"2017"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2006.99"},{"key":"e_1_2_1_42_1","volume-title":"Billion-scale similarity search with CPUs. arXiv preprint arXiv:1702.08734","author":"Johnson Jeff","year":"2017"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447555.3464865"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380600"},{"key":"e_1_2_1_45_1","unstructured":"W. Li Y. Zhang Y. Sun W. Wang M. Li W. Zhang and X. Lin. 2019. Approximate Nearest Neighbor Search on High Dimensional Data - Experiments Analyses and Improvement. TKDE (2019).  W. Li Y. Zhang Y. Sun W. Wang M. Li W. Zhang and X. Lin. 2019. Approximate Nearest Neighbor Search on High Dimensional Data - Experiments Analyses and Improvement. TKDE (2019)."},{"key":"e_1_2_1_46_1","volume-title":"Keogh","author":"Linardi Michele","year":"2020"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.5555\/2976040.2976144"},{"key":"e_1_2_1_48_1","unstructured":"Xiao Luo Chong Chen Huasong Zhong Hao Zhang Minghua Deng Jianqiang Huang and Xiansheng Hua. 2020 A Survey on Deep Hashing Methods. arXiv:2003.03369 [cs.CV]  Xiao Luo Chong Chen Huasong Zhong Hao Zhang Minghua Deng Jianqiang Huang and Xiansheng Hua. 2020 A Survey on Deep Hashing Methods. arXiv:2003.03369 [cs.CV]"},{"key":"e_1_2_1_49_1","unstructured":"Yury A Malkov and D. A. Yashunin 2016. Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs. CoRR abs\/1603.09320 (2016).  Yury A Malkov and D. A. Yashunin 2016. Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs. CoRR abs\/1603.09320 (2016)."},{"key":"e_1_2_1_50_1","doi-asserted-by":"crossref","unstructured":"Stanislav Morozov and Artem Babenko. 2019 Unsupervised neural quantisation for compressed-domain similarity search. In ICCV.  Stanislav Morozov and Artem Babenko. 2019 Unsupervised neural quantisation for compressed-domain similarity search. In ICCV .","DOI":"10.1109\/ICCV.2019.00313"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2814710.2814719"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-44900-1_5"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377391.3377400"},{"key":"e_1_2_1_54_1","volume-title":"Panagiota Fatourou and Themis Palpanas","author":"Peng Botao","year":"2021"},{"key":"e_1_2_1_55_1","volume-title":"SING: Sequence indexing Using GPUs. In ICDE.","author":"Peng Botao","year":"2021"},{"key":"e_1_2_1_56_1","volume-title":"Data Series Indexing on Multi-core Architectures. TKDE","author":"Peng Botao","year":"2020"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2014.27"},{"key":"e_1_2_1_58_1","unstructured":"Liudmila Prokhorenkuva and Aleksandr Shekhovtsov. 2020. Graph-based Nearest Neighbor Search: From Practice to Theory. In PMLR.  Liudmila Prokhorenkuva and Aleksandr Shekhovtsov. 2020. Graph-based Nearest Neighbor Search: From Practice to Theory. In PMLR ."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415564"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.5555\/1076819"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1921064"},{"key":"e_1_2_1_62_1","doi-asserted-by":"crossref","unstructured":"S Shekkizhar and A. Ortega 2020 Graph Construction from Data by Non-Negative Kernel Regression In ICASSP.  S Shekkizhar and A. Ortega 2020 Graph Construction from Data by Non-Negative Kernel Regression In ICASSP .","DOI":"10.1109\/ICASSP40776.2020.9054425"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735461.2735462"},{"key":"e_1_2_1_64_1","doi-asserted-by":"crossref","unstructured":"Saravanan Thirumuruganathan Shohedul Hasan Nick Koudas and Gaulam Das. 2020. Approximate query processing for data exploration using deep generative models. In ICDE. 1309--1320.  Saravanan Thirumuruganathan Shohedul Hasan Nick Koudas and Gaulam Das. 2020. Approximate query processing for data exploration using deep generative models. In ICDE . 1309--1320.","DOI":"10.1109\/ICDE48307.2020.00117"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2016.2598470"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219869"},{"key":"e_1_2_1_67_1","volume-title":"N. Sebe, and H. T. Shen.","author":"Wang J.","year":"2018"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467317"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536206.2536208"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.5555\/645924.671192"},{"key":"e_1_2_1_71_1","volume-title":"GhbalSIP","author":"Wu Hanwei"},{"key":"e_1_2_1_72_1","unstructured":"Jiaye Wu Peng Wang Ningting Pan Chen Wang Wei Wang and Jianmin Wang 2019. KV-Match: A Subsequence Matching Approach Supporting Normalisation and Time Warping. In ICDE.  Jiaye Wu Peng Wang Ningting Pan Chen Wang Wei Wang and Jianmin Wang 2019. KV-Match: A Subsequence Matching Approach Supporting Normalisation and Time Warping. In ICDE ."},{"key":"e_1_2_1_73_1","volume-title":"Massively Distributed Time Series Indexing and Querying. TKDE 32, 1","author":"Yagoubi Djamel-Edine","year":"2019"},{"key":"e_1_2_1_74_1","volume-title":"How Progressive Visualizations Affect Exploratory Analysis. TVCG 23, 8","author":"Zgraggen Emanuel","year":"2017"},{"key":"e_1_2_1_75_1","volume-title":"TARDIS: Distributed Indexing Framework for Big Time Series Data, in ICDE.","author":"Zhang L.","year":"2019"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994534"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-016-0442-5"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3476311.3476407","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T11:43:11Z","timestamp":1672227791000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3476311.3476407"}},"subtitle":["al-driven, progressive, and distributed"],"short-title":[],"issued":{"date-parts":[[2021,7]]},"references-count":77,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2021,7]]}},"alternative-id":["10.14778\/3476311.3476407"],"URL":"https:\/\/doi.org\/10.14778\/3476311.3476407","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,7]]}}}