{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T20:56:18Z","timestamp":1775595378529,"version":"3.50.1"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"name":"NSFC-RGC","award":["62461160333"],"award-info":[{"award-number":["62461160333"]}]},{"name":"National Key Research and Development Program of China","award":["2023YFB4502300"],"award-info":[{"award-number":["2023YFB4502300"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,4,2]]},"abstract":"<jats:p>\n                    Continuous\n                    <jats:italic toggle=\"yes\">Approximate Nearest Neighbor Search<\/jats:italic>\n                    (ANNS) over real-time vector data streams is an increasingly critical yet underexplored problem. In open-world settings\u2014where data distributions shift, noise accumulates, and concurrent access is common\u2014existing ANNS algorithms, originally designed for static or simplified streaming scenarios, struggle to balance ingestion latency, retrieval quality, and update efficiency. While benchmarks such as ANN-Benchmarks and Big-ANN-Benchmarks have standardized evaluation in static or large-scale settings, they fail to capture the nuanced, high-churn dynamics of real-world streams. We introduce CANDOR-Bench (\n                    <jats:bold>C<\/jats:bold>\n                    ontinuous\n                    <jats:bold>A<\/jats:bold>\n                    pproximate\n                    <jats:bold>N<\/jats:bold>\n                    earest neighbor search under\n                    <jats:bold>D<\/jats:bold>\n                    ynamic\n                    <jats:bold>O<\/jats:bold>\n                    pen-wo\n                    <jats:bold>R<\/jats:bold>\n                    ld Streams, a benchmarking framework built on Big-ANN-Benchmark to evaluate in-memory ANNS under dynamic, open-world conditions. CANDOR-Bench supports high-frequency ingestion (up to hundreds of thousands of vectors per second), adaptive drift modeling (including modality shifts), stochastic noise injection, and concurrent query-update execution\u2014all without requiring modifications to algorithm code. Across 12 datasets and 19 representative ANNS algorithms, our evaluation reveals that no single ANNS algorithm consistently delivers high recall, throughput, and update efficiency across dynamic open-world scenarios, which challenges assumptions drawn from static benchmarks. This variability reflects deeper trade-offs inherent to streaming settings. For example, smaller update batches improve data freshness but can introduce higher insertion overhead and reduce accuracy. We further observe that throughput in concurrent settings is often constrained by insertion overhead rather than query latency, which highlights a mismatch between streaming workloads and designs originally tuned for offline construction.\n                  <\/jats:p>","DOI":"10.1145\/3786630","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T17:54:13Z","timestamp":1775584453000},"page":"1-27","source":"Crossref","is-referenced-by-count":0,"title":["CANDOR-Bench: Benchmarking In-Memory Continuous ANNS under Dynamic Open-World Streams [Experiments &amp; Analysis]"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-6268-1098","authenticated-orcid":false,"given":"Mingqi","family":"Wang","sequence":"first","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8738-7268","authenticated-orcid":false,"given":"Junyao","family":"Dong","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0720-6805","authenticated-orcid":false,"given":"Zhuoyan","family":"Wu","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1514-0225","authenticated-orcid":false,"given":"Jun","family":"Liu","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-5891-0671","authenticated-orcid":false,"given":"Ruicheng","family":"Zhang","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-4008-1821","authenticated-orcid":false,"given":"Jianjun","family":"Zhao","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6189-7909","authenticated-orcid":false,"given":"Ruipeng","family":"Wan","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8552-8135","authenticated-orcid":false,"given":"Xinyan","family":"Lei","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9927-6925","authenticated-orcid":false,"given":"Shuhao","family":"Zhang","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8639-4570","authenticated-orcid":false,"given":"Bolong","family":"Zheng","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4290-1408","authenticated-orcid":false,"given":"Haikun","family":"Liu","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6302-813X","authenticated-orcid":false,"given":"Xiaofei","family":"Liao","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3934-7605","authenticated-orcid":false,"given":"Hai","family":"Jin","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Big Data Technology and System, Service Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,7]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2024. New | War Thunder Wiki. https:\/\/wiki.warthunder.com\/"},{"key":"e_1_2_1_2_1","unstructured":"AbdelrahmanMohamed129. 2023. GitHub - AbdelrahmanMohamed129\/DiskANN at farah. https:\/\/github.com\/AbdelrahmanMohamed129\/DiskANN\/tree\/farah"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.48550\/arxiv.2402.02044"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.14778\/2856318.2856324"},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Martin Aum\u00fcller Erik Bernhardsson and Alexander Faithfull. 2018. ANN-Benchmarks: A Benchmarking Tool for Approximate Nearest Neighbor Algorithms. arXiv:1807.05614 [cs.IR] https:\/\/arxiv.org\/abs\/1807.05614","DOI":"10.1007\/978-3-319-68474-1_3"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3709693"},{"key":"e_1_2_1_7_1","unstructured":"baidu. 2023. GitHub - baidu\/puck: Puck is a High-Performance ANN Search Engine. https:\/\/github.com\/baidu\/puck"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/361002.361007"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.14778\/3467861.346786"},{"key":"e_1_2_1_10_1","volume-title":"SPANN: Highly-Efficient Billion-Scale Approximate Nearest Neighbor Search. arXiv:2111.08566","author":"Chen Qi","year":"2021","unstructured":"Qi Chen, Bing Zhao, Haidong Wang, Mingqin Li, Chuanjie Liu, Zengzhong Li, Mao Yang, and Jingdong Wang. 2021. SPANN: Highly-Efficient Billion-Scale Approximate Nearest Neighbor Search. arXiv:2111.08566 (2021). doi:10.48550\/ arxiv.2111.08566"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3437963.3441810"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589282"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of 25th International Conference on Very Large Data Bases","volume":"99","author":"Gionis Aristides","year":"1999","unstructured":"Aristides Gionis, Piotr Indyk, and Rajeev Motwani. 1999. Similarity Search in High Dimensions via Hashing. In Proceedings of 25th International Conference on Very Large Data Bases, Vol. 99. 518--529."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/573304"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219885"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2022.3161156"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.14778\/3554821.3554843"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2019.2921572"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.57"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-020-01316-z"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.1979.1675439"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3284028.3284030"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2019.2909204"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413595"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.48550\/arxiv.1405.0312"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2882959"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 14th Conference on Innovative Data Systems Research.","author":"Lu Duo","year":"2025","unstructured":"Duo Lu, Siming Feng, Jonathan Zhou, Franco Solleza, Malte Schwarzkopf, and Ugur \u00c7etintemel. 2025. VectraFlow: Integrating Vectors into Stream Processing. In Proceedings of the 14th Conference on Innovative Data Systems Research."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.14778\/3717755.3717760"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2013.10.006"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2889473"},{"key":"e_1_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Christopher D. Manning Prabhakar Raghavan and Hinrich Sch\u00fctze. 2008. Introduction to Information Retrieval. (2008).","DOI":"10.1017\/CBO9780511809071"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3627535.3638475"},{"key":"e_1_2_1_33_1","unstructured":"Marqo. 2024. Context is All You Need: Multimodal Vector Search with Personalization. https:\/\/www.marqo.ai\/blog\/context-is-all-you-need-multimodal-vector-search-with-personalization Accessed: 2025-03--25."},{"key":"e_1_2_1_34_1","unstructured":"NLTK. 2009. Natural Language Toolkit \u2014 NLTK 3.4.4 Documentation. https:\/\/www.nltk.org\/"},{"key":"e_1_2_1_35_1","unstructured":"Numpy. 2024. NumPy. https:\/\/numpy.org\/"},{"key":"e_1_2_1_36_1","volume-title":"CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUs. arXiv:2308.15136 [cs.DS] https:\/\/arxiv.org\/abs\/2308.15136","author":"Ootomo Hiroyuki","year":"2024","unstructured":"Hiroyuki Ootomo, Akira Naruse, Corey Nolet, Ray Wang, Tamas Feher, and Yong Wang. 2024. CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUs. arXiv:2308.15136 [cs.DS] https:\/\/arxiv.org\/abs\/2308.15136"},{"key":"e_1_2_1_37_1","unstructured":"PAPI. 2023. PAPI. https:\/\/icl.utk.edu\/papi\/"},{"key":"e_1_2_1_38_1","unstructured":"PyTorch. 2023. PyTorch. https:\/\/pytorch.org\/"},{"key":"e_1_2_1_39_1","unstructured":"qdrant. 2023. GitHub - qdrant\/qdrant: High-Performance Massive-Scale Vector Database and Vector Search Engine for the Next Generation of AI. https:\/\/github.com\/qdrant\/qdrant"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE65448.2025.00130"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2103.00020"},{"key":"e_1_2_1_42_1","unstructured":"Harsha Vardhan Simhadri Martin Aum\u00fcller Amir Ingber Matthijs Douze George Williams Magdalen Dobson Manohar Dmitry Baranchuk Edo Liberty Frank Liu Ben Landrum Mazin Karjikar Laxman Dhulipala Meng Chen Yue Chen Rui Ma Kai Zhang Yuzheng Cai Jiayang Shi Yizhuo Chen Weiguo Zheng Zihao Wan Jie Yin and Ben Huang. 2024. Results of the Big ANN: NeurIPS'23 Competition. arXiv:2409.17424 [cs.IR] https:\/\/arxiv.org\/abs\/2409.17424"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2105.09613"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/BigData50022.2020.9377880"},{"key":"e_1_2_1_45_1","doi-asserted-by":"crossref","unstructured":"Michael Stonebraker Samuel Madden Daniel J. Abadi Stavros Harizopoulos Nabil Hachem and Pat Helland. 2018. The End of an Architectural Era: It's Time for a Complete Rewrite. In Making Databases Work: the Pragmatic Wisdom of Michael Stonebraker. 463--489.","DOI":"10.1145\/3226595.3226637"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.14778\/2556549.2556574"},{"key":"e_1_2_1_47_1","unstructured":"veaaaab. 2023. GitHub - veaaaab\/DiskANN at bigann23_streaming. https:\/\/github.com\/veaaaab\/DiskANN\/tree\/bigann23_streaming"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457550"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/TWC.2014.040914.131422"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476255"},{"key":"e_1_2_1_51_1","first-page":"3","article-title":"Graph- and Tree-Based Indexes for High-Dimensional Vector Similarity Search: Analyses, Comparisons, and Future Directions","volume":"46","author":"Wang Zeyu","year":"2023","unstructured":"Zeyu Wang, Peng Wang, Themis Palpanas, and Wei Wang. 2023. Graph- and Tree-Based Indexes for High-Dimensional Vector Similarity Search: Analyses, Comparisons, and Future Directions. IEEE Data Engineering Bulletin 46, 3 (2023), 3--21.","journal-title":"IEEE Data Engineering Bulletin"},{"key":"e_1_2_1_52_1","volume-title":"Advances in Neural Information Processing Systems 21","author":"Weiss Yair","year":"2008","unstructured":"Yair Weiss, Antonio Torralba, and Rob Fergus. 2008. Spectral Hashing. Advances in Neural Information Processing Systems 21 (2008)."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.48550\/arxiv.2407.07871"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/tkde.2018.2817526"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2502.13826"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613166"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2206.10839"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2503.00402"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00046"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2522737"},{"key":"e_1_2_1_61_1","volume-title":"Proceedings of the 4th USENIX Conference on Hot Topics in Cloud Computing","author":"Zaharia Matei","year":"2012","unstructured":"Matei Zaharia, Tathagata Das, Haoyuan Li, Scott Shenker, and Ion Stoica. 2012. Discretized Streams: An Efficient and Fault-Tolerant Model for Stream Processing on Large Clusters. In Proceedings of the 4th USENIX Conference on Hot Topics in Cloud Computing (Boston, MA) (HotCloud'12). USENIX Association, USA, 10."},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531722"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE48307.2020.00094"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.14778\/3594512.3594527"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.3390\/computers13010001"}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786630","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T19:57:31Z","timestamp":1775591851000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786630"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,2]]},"references-count":65,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,4,2]]}},"alternative-id":["10.1145\/3786630"],"URL":"https:\/\/doi.org\/10.1145\/3786630","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,2]]}}}