{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T16:33:29Z","timestamp":1781886809975,"version":"3.54.5"},"reference-count":111,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2025,7]]},"abstract":"<jats:p>Vector search systems are indispensable in large language model (LLM) serving, search engines, and recommender systems, where minimizing online search latency is essential. Among various algorithms, graph-based vector search (GVS) is particularly popular due to its high search performance and quality. However, reducing GVS latency by intra-query parallelization remains challenging due to limitations imposed by both existing hardware architectures (CPUs and GPUs) and the inherent difficulty of parallelizing graph traversals. To efficiently serve low-latency GVS, we co-design hardware and algorithm by proposing Falcon and Delayed-Synchronization Traversal (DST). Falcon is a hardware GVS accelerator that implements efficient GVS operators, pipelines these operators, and reduces memory accesses by tracking search states with an on-chip Bloom filter. DST is an efficient graph traversal algorithm that simultaneously improves search performance and quality by relaxing traversal orders to maximize accelerator utilization. Evaluation across various graphs and datasets shows that Falcon, prototyped on FPGAs, together with DST, achieves up to 4.3X and 19.5X lower latency and up to 8.0X and 26.9X improvements in energy efficiency over CPU- and GPU-based GVS systems.<\/jats:p>","DOI":"10.14778\/3749646.3749655","type":"journal-article","created":{"date-parts":[[2025,9,4]],"date-time":"2025-09-04T17:55:06Z","timestamp":1757008506000},"page":"3797-3811","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Fast Graph Vector Search via Hardware Acceleration and Delayed-Synchronization Traversal"],"prefix":"10.14778","volume":"18","author":[{"given":"Wenqi","family":"Jiang","sequence":"first","affiliation":[{"name":"Systems Group, ETH Zurich"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hang","family":"Hu","sequence":"additional","affiliation":[{"name":"Systems Group, ETH Zurich"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Torsten","family":"Hoefler","sequence":"additional","affiliation":[{"name":"SPCL, ETH Zurich"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gustavo","family":"Alonso","sequence":"additional","affiliation":[{"name":"Systems Group, ETH Zurich"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,4]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"[n.d.]. AMD Alveo SN1000 SmartNIC Accelerator Card. https:\/\/www.amd.com\/en\/products\/accelerators\/alveo\/sn1000\/a-sn1022-p4.html."},{"key":"e_1_2_1_2_1","unstructured":"[n.d.]. Faiss. https:\/\/github.com\/facebookresearch\/faiss\/."},{"key":"e_1_2_1_3_1","unstructured":"[n.d.]. Intel FPGA SmartNIC N6000-PL Platform. https:\/\/www.intel.com\/content\/www\/us\/en\/products\/details\/fpga\/platforms\/smartnic\/n6000-pl-platform.html."},{"key":"e_1_2_1_4_1","unstructured":"[n.d.]. The Memory Wall: Past Present and Future of DRAM. https:\/\/semianalysis.com\/2024\/09\/03\/the-memory-wall\/."},{"key":"e_1_2_1_5_1","unstructured":"[n.d.]. The MurmurHash family. https:\/\/github.com\/aappleby\/smhasher."},{"key":"e_1_2_1_6_1","unstructured":"[n.d.]. NVIDIA Deep Learning Recommender Model Implementation. https:\/\/github.com\/NVIDIA\/DeepLearningExamples\/tree\/master\/PyTorch\/Recommendation\/DLRM."},{"key":"e_1_2_1_7_1","unstructured":"[n.d.]. The NVIDIA GH200 Grace Hopper Superchip. https:\/\/www.nvidia.com\/en-us\/data-center\/grace-hopper-superchip."},{"key":"e_1_2_1_8_1","unstructured":"[n.d.]. SIFT ANNS dataset. http:\/\/corpus-texmex.irisa.fr\/"},{"key":"e_1_2_1_9_1","unstructured":"[n.d.]. The SPACEV Web Embedding Dataset. https:\/\/github.com\/microsoft\/SPTAG\/tree\/main\/datasets\/SPACEV1B."},{"key":"e_1_2_1_10_1","volume-title":"42nd International Conference on Very Large Data Bases","volume":"9","author":"Andr\u00e9 Fabien","year":"2016","unstructured":"Fabien Andr\u00e9, Anne-Marie Kermarrec, and Nicolas Le Scouarnec. 2016. Cache locality is not enough: High-performance nearest neighbor search with product quantization fast scan. In 42nd International Conference on Very Large Data Bases, Vol. 9. 12."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.14778\/3583140.3583166"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2055\u20132063","author":"Babenko Artem","year":"2016","unstructured":"Artem Babenko and Victor Lempitsky. 2016. Efficient indexing of billion-scale datasets of deep descriptors. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2055\u20132063."},{"key":"e_1_2_1_13_1","unstructured":"Abhimanyu Bambhaniya Ritik Raj Geonhwa Jeong Souvik Kundu Sudarshan Srinivasan Midhilesh Elavazhagan Madhu Kumar and Tushar Krishna. 2024. Demystifying Platform Requirements for Diverse LLM Inference Use Cases. arXiv:2406.01698 [cs.AR]"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/187382.187805"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/362686.362692"},{"key":"e_1_2_1_16_1","volume-title":"International conference on machine learning. PMLR, 2206\u20132240","author":"Borgeaud Sebastian","year":"2022","unstructured":"Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. In International conference on machine learning. PMLR, 2206\u20132240."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2896377.2901453"},{"key":"e_1_2_1_18_1","volume-title":"SPANN: Highly-efficient Billion-scale Approximate Nearest Neighbor Search. arXiv preprint arXiv:2111.08566","author":"Chen Qi","year":"2021","unstructured":"Qi Chen, Bing Zhao, Haidong Wang, Mingqin Li, Chuanjie Liu, Zengzhong Li, Mao Yang, and Jingdong Wang. 2021. SPANN: Highly-efficient Billion-scale Approximate Nearest Neighbor Search. arXiv preprint arXiv:2111.08566 (2021)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2019.04.033"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3323873.3325018"},{"key":"e_1_2_1_21_1","volume-title":"TPU-KNN: K Nearest Neighbor Search at Peak FLOP\/s. arXiv preprint arXiv:2206.14286","author":"Chern Felix","year":"2022","unstructured":"Felix Chern, Blake Hechtman, Andy Davis, Ruiqi Guo, David Majnemer, and Sanjiv Kumar. 2022. TPU-KNN: K Nearest Neighbor Search at Peak FLOP\/s. arXiv preprint arXiv:2206.14286 (2022)."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC42613.2021.9365803"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2959100.2959190"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735461.2735463"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/997817.997857"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the VLDB Endowment","author":"Doshi Ishita","year":"2020","unstructured":"Ishita Doshi, Dhritiman Das, Ashish Bhutani, Rajeev Kumar, Rushi Bhatt, and Niranjan Balasubramanian. 2020. LANNS: a web-scale approximate nearest neighbor lookup system. Proceedings of the VLDB Endowment (2020)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00012"},{"key":"e_1_2_1_28_1","volume-title":"Fast approximate nearest neighbor search with the navigating spreading-out graph. arXiv preprint arXiv:1707.00143","author":"Fu Cong","year":"2017","unstructured":"Cong Fu, Chao Xiang, Changxu Wang, and Deng Cai. 2017. Fast approximate nearest neighbor search with the navigating spreading-out graph. arXiv preprint arXiv:1707.00143 (2017)."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589282"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654970"},{"key":"e_1_2_1_31_1","volume-title":"Optimized product quantization","author":"Ge Tiezheng","year":"2013","unstructured":"Tiezheng Ge, Kaiming He, Qifa Ke, and Jian Sun. 2013. Optimized product quantization. IEEE transactions on pattern analysis and machine intelligence 36, 4 (2013), 744\u2013755."},{"key":"e_1_2_1_32_1","first-page":"518","article-title":"Similarity search in high dimensions via hashing","volume":"99","author":"Gionis Aristides","year":"1999","unstructured":"Aristides Gionis, Piotr Indyk, Rajeev Motwani, et al. 1999. Similarity search in high dimensions via hashing. In Vldb, Vol. 99. 518\u2013529.","journal-title":"Vldb"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3709730"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2022.3161156"},{"key":"e_1_2_1_35_1","volume-title":"Manu: A Cloud Native Vector Database Management System. arXiv preprint arXiv:2206.13843","author":"Guo Rentong","year":"2022","unstructured":"Rentong Guo, Xiaofan Luan, Long Xiang, Xiao Yan, Xiaomeng Yi, Jigao Luo, Qianya Cheng, Weizhi Xu, Jiarui Luo, Frank Liu, et al. 2022. Manu: A Cloud Native Vector Database Management System. arXiv preprint arXiv:2206.13843 (2022)."},{"key":"e_1_2_1_36_1","volume-title":"2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 982\u2013995","author":"Gupta Udit","year":"2020","unstructured":"Udit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang, Brandon Reagen, Gu-Yeon Wei, Hsien-Hsin S Lee, David Brooks, and Carole-Jean Wu. 2020. Deep-recsys: A system for optimizing end-to-end at-scale neural recommendation inference. In 2020 ACM\/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 982\u2013995."},{"key":"e_1_2_1_37_1","volume-title":"Realm: Retrieval-augmented language model pre-training. arXiv preprint arXiv:2002.08909","author":"Guu Kelvin","year":"2020","unstructured":"Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. Realm: Retrieval-augmented language model pre-training. arXiv preprint arXiv:2002.08909 (2020)."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/FPL53798.2021.00040"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00058"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403305"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2014.6927413"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457240"},{"key":"e_1_2_1_43_1","volume-title":"2023 USENIX Annual Technical Conference (USENIX ATC 23)","author":"Jang Junhyeok","year":"2023","unstructured":"Junhyeok Jang, Hanjin Choi, Hanyeoreum Bae, Seungjun Lee, Miryeong Kwon, and Myoungsoo Jung. 2023. {CXL-ANNS}:{Software-Hardware} Collaborative Memory Disaggregation and Computation for {Billion-Scale} Approximate Nearest Neighbor Search. In 2023 USENIX Annual Technical Conference (USENIX ATC 23). 585\u2013600."},{"key":"e_1_2_1_44_1","volume-title":"Ravishankar Krishnawamy, and Rohan Kadekodi.","author":"Subramanya Suhas Jayaram","year":"2019","unstructured":"Suhas Jayaram Subramanya, Fnu Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnawamy, and Rohan Kadekodi. 2019. Diskann: Fast accurate billion-point nearest neighbor search on a single node. Advances in Neural Information Processing Systems 32 (2019)."},{"key":"e_1_2_1_45_1","volume-title":"Product quantization for nearest neighbor search","author":"Jegou Herve","year":"2010","unstructured":"Herve Jegou, Matthijs Douze, and Cordelia Schmid. 2010. Product quantization for nearest neighbor search. IEEE transactions on pattern analysis and machine intelligence 33, 1 (2010), 117\u2013128."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2749447"},{"key":"e_1_2_1_47_1","first-page":"845","article-title":"MicroRec: efficient recommendation inference by hardware and data structure solutions","volume":"3","author":"Jiang Wenqi","year":"2021","unstructured":"Wenqi Jiang, Zhenhao He, Shuai Zhang, Thomas B Preu\u00dfer, Kai Zeng, Liang Feng, Jiansong Zhang, Tongxuan Liu, Yong Li, Jingren Zhou, et al. 2021. MicroRec: efficient recommendation inference by hardware and data structure solutions. Proceedings of Machine Learning and Systems 3 (2021), 845\u2013859.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467139"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3581784.3607045"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3695053.3731093"},{"key":"e_1_2_1_51_1","volume-title":"Proceedings of the VLDB Endowment 18","author":"Jiang Wenqi","year":"2025","unstructured":"Wenqi Jiang, Marco Zeller, Roger Waleffe, Torsten Hoefler, and Gustavo Alonso. 2025. Chameleon: a heterogeneous and disaggregated accelerator system for retrieval-augmented language models. Proceedings of the VLDB Endowment 18 (2025)."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3690624.3709194"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2019.2921572"},{"key":"e_1_2_1_54_1","volume-title":"Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906","author":"Karpukhin Vladimir","year":"2020","unstructured":"Vladimir Karpukhin, Barlas O\u011fuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. arXiv preprint arXiv:2004.04906 (2020)."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401075"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2022.3155956"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.5555\/3277332.3277347"},{"key":"e_1_2_1_58_1","volume-title":"Proceedings of the International Conference on Parallel and Distributed Processing Techniques and Applications (PDPTA). Citeseer, 1.","author":"LaGrone James","year":"2011","unstructured":"James LaGrone, Ayodunni Aribuki, and Barbara Chapman. 2011. A set of microbenchmarks for measuring OpenMP task overheads. In Proceedings of the International Conference on Parallel and Distributed Processing Techniques and Applications (PDPTA). Citeseer, 1."},{"key":"e_1_2_1_59_1","volume-title":"ANNA: Specialized Architecture for Approximate Nearest Neighbor Search. In 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 169\u2013183","author":"Lee Yejin","year":"2022","unstructured":"Yejin Lee, Hyunji Choi, Sunhong Min, Hyunseung Lee, Sangwon Beak, Dawoon Jeong, Jae W Lee, and Tae Jun Ham. 2022. ANNA: Specialized Architecture for Approximate Nearest Neighbor Search. In 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 169\u2013183."},{"key":"e_1_2_1_60_1","unstructured":"Charles E Leiserson. 1979. Systolic Priority Queues. Technical Report. CARNEGIE-MELLON UNIV PITTSBURGH PA DEPT OF COMPUTER SCIENCE."},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2008.130"},{"key":"e_1_2_1_62_1","first-page":"9459","article-title":"Retrieval-augmented generation for knowledge-intensive nlp tasks","volume":"33","author":"Lewis Patrick","year":"2020","unstructured":"Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\u00fcttler, Mike Lewis, Wen-tau Yih, Tim Rockt\u00e4schel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems 33 (2020), 9459\u20139474.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2019.2909204"},{"key":"e_1_2_1_64_1","volume-title":"2019 USENIX Annual Technical Conference (USENIX ATC 19)","author":"Liang Shengwen","year":"2019","unstructured":"Shengwen Liang, Ying Wang, Youyou Lu, Zhe Yang, Huawei Li, and Xiaowei Li. 2019. Cognitive {SSD}: A deep learning engine for {In-Storage} data retrieval. In 2019 USENIX Annual Technical Conference (USENIX ATC 19). 395\u2013410."},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3489517.3530560"},{"key":"e_1_2_1_66_1","volume-title":"Graph based nearest neighbor search: Promises and failures. arXiv preprint arXiv:1904.02077","author":"Lin Peng-Cheng","year":"2019","unstructured":"Peng-Cheng Lin and Wan-Lei Zhao. 2019. Graph based nearest neighbor search: Promises and failures. arXiv preprint arXiv:1904.02077 (2019)."},{"key":"e_1_2_1_67_1","volume-title":"TigerVector: Supporting Vector Search in Graph Databases for Advanced RAGs. arXiv preprint arXiv:2501.11216","author":"Liu Shige","year":"2025","unstructured":"Shige Liu, Zhifang Zeng, Li Chen, Adil Ainihaer, Arun Ramasami, Songting Chen, Yu Xu, Mingxi Wu, and Jianguo Wang. 2025. TigerVector: Supporting Vector Search in Graph Databases for Advanced RAGs. arXiv preprint arXiv:2501.11216 (2025)."},{"key":"e_1_2_1_68_1","volume-title":"JUNO: Optimizing High-Dimensional Approximate Nearest Neighbour Search with Sparsity-Aware Algorithm and Ray-Tracing Core Mapping. arXiv preprint arXiv:2312.01712","author":"Liu Zihan","year":"2023","unstructured":"Zihan Liu, Wentao Ni, Jingwen Leng, Yu Feng, Cong Guo, Quan Chen, Chao Li, Minyi Guo, and Yuhao Zhu. 2023. JUNO: Optimizing High-Dimensional Approximate Nearest Neighbour Search with Sparsity-Aware Algorithm and Ray-Tracing Core Mapping. arXiv preprint arXiv:2312.01712 (2023)."},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICFPT51103.2020.00027"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.14778\/3489496.3489506"},{"key":"e_1_2_1_71_1","volume-title":"Efficient processing of k nearest neighbor joins using mapreduce. arXiv preprint arXiv:1207.0141","author":"Lu Wei","year":"2012","unstructured":"Wei Lu, Yanyan Shen, Su Chen, and Beng Chin Ooi. 2012. Efficient processing of k nearest neighbor joins using mapreduce. arXiv preprint arXiv:1207.0141 (2012)."},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2013.10.006"},{"key":"e_1_2_1_73_1","volume-title":"Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs","author":"Malkov Yu A","year":"2018","unstructured":"Yu A Malkov and Dmitry A Yashunin. 2018. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE transactions on pattern analysis and machine intelligence 42, 4 (2018), 824\u2013836."},{"key":"e_1_2_1_74_1","volume-title":"Handbook of data structures and applications","author":"Mehta Dinesh P","unstructured":"Dinesh P Mehta and Sartaj Sahni. 2004. Handbook of data structures and applications. Chapman and Hall\/CRC."},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0196-6774(03)00076-2"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589777"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389746"},{"key":"e_1_2_1_78_1","volume-title":"Survey of vector database management systems. arXiv preprint arXiv:2310.14021","author":"Pan James Jie","year":"2023","unstructured":"James Jie Pan, Jianguo Wang, and Guoliang Li. 2023. Survey of vector database management systems. arXiv preprint arXiv:2310.14021 (2023)."},{"key":"e_1_2_1_79_1","volume-title":"Splitwise: Efficient generative llm inference using phase splitting. arXiv preprint arXiv:2311.18677","author":"Patel Pratyush","year":"2023","unstructured":"Pratyush Patel, Esha Choukse, Chaojie Zhang, \u00cd\u00f1igo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini. 2023. Splitwise: Efficient generative llm inference using phase splitting. arXiv preprint arXiv:2311.18677 (2023)."},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD51958.2021.9643528"},{"key":"e_1_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588908"},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1145\/3572848.3577527"},{"key":"e_1_2_1_83_1","doi-asserted-by":"publisher","DOI":"10.1145\/2678373.2665678"},{"key":"e_1_2_1_84_1","first-page":"10672","article-title":"Hm-ann: Efficient billion-point nearest neighbor search on heterogeneous memory","volume":"33","author":"Ren Jie","year":"2020","unstructured":"Jie Ren, Minjia Zhang, and Dong Li. 2020. Hm-ann: Efficient billion-point nearest neighbor search on heterogeneous memory. Advances in Neural Information Processing Systems 33 (2020), 10672\u201310684.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_85_1","doi-asserted-by":"publisher","DOI":"10.1145\/2815400.2815408"},{"key":"e_1_2_1_86_1","volume-title":"IEEE International Conference on","volume":"3","author":"Sivic Josef","year":"2003","unstructured":"Josef Sivic and Andrew Zisserman. 2003. Video Google: A text retrieval approach to object matching in videos. In Computer Vision, IEEE International Conference on, Vol. 3. IEEE Computer Society, 1470\u20131470."},{"key":"e_1_2_1_87_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-15286-3_16"},{"key":"e_1_2_1_88_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735461.2735462"},{"key":"e_1_2_1_89_1","volume-title":"2016 USENIX Annual Technical Conference (USENIX ATC 16)","author":"Vora Keval","year":"2016","unstructured":"Keval Vora, Guoqing Xu, and Rajiv Gupta. 2016. Load the Edges You Need: A Generic {I\/O} Optimization for Disk-based Graph Processing. In 2016 USENIX Annual Technical Conference (USENIX ATC 16). 507\u2013522."},{"key":"e_1_2_1_90_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457550"},{"key":"e_1_2_1_91_1","doi-asserted-by":"publisher","DOI":"10.1145\/3639269"},{"key":"e_1_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476255"},{"key":"e_1_2_1_93_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2749442"},{"key":"e_1_2_1_94_1","unstructured":"Yitu Wang Shiyu Li Qilin Zheng Linghao Song Zongwang Li Andrew Chang Hai Li Yiran Chen et al. 2023. In-Storage Acceleration of Graph-Traversal-Based Approximate Nearest Neighbor Search. arXiv preprint arXiv:2312.03141 (2023)."},{"key":"e_1_2_1_95_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415541"},{"key":"e_1_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.223"},{"key":"e_1_2_1_97_1","doi-asserted-by":"publisher","DOI":"10.1145\/2588555.2610500"},{"key":"e_1_2_1_98_1","volume-title":"Approximate nearest neighbor negative contrastive learning for dense text retrieval. arXiv preprint arXiv:2007.00808","author":"Xiong Lee","year":"2020","unstructured":"Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. 2020. Approximate nearest neighbor negative contrastive learning for dense text retrieval. arXiv preprint arXiv:2007.00808 (2020)."},{"key":"e_1_2_1_99_1","volume-title":"Proxima: Near-storage Acceleration for Graph-based Approximate Nearest Neighbor Search in 3D NAND. arXiv preprint arXiv:2312.04257","author":"Xu Weihong","year":"2023","unstructured":"Weihong Xu, Junwei Chen, Po-Kai Hsu, Jaeyoung Kang, Minxuan Zhou, Sumukh Pinge, Shimeng Yu, and Tajana Rosing. 2023. Proxima: Near-storage Acceleration for Graph-based Approximate Nearest Neighbor Search in 3D NAND. arXiv preprint arXiv:2312.04257 (2023)."},{"key":"e_1_2_1_100_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735479.2735492"},{"key":"e_1_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3386131"},{"key":"e_1_2_1_102_1","volume-title":"16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Zeng Chaoliang","year":"2022","unstructured":"Chaoliang Zeng, Layong Luo, Qingsong Ning, Yaodong Han, Yuhang Jiang, Ding Tang, Zilong Wang, Kai Chen, and Chuanxiong Guo. 2022. {FAERY}: An {FPGA-accelerated} Embedding-based Retrieval System. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). 841\u2013856."},{"key":"e_1_2_1_103_1","doi-asserted-by":"crossref","unstructured":"Shulin Zeng Zhenhua Zhu Jun Liu Haoyu Zhang Guohao Dai Zixuan Zhou Shuangchen Li Xuefei Ning Yuan Xie Huazhong Yang et al. 2023. DF-GAS: a Distributed FPGA-as-a-Service Architecture towards Billion-Scale Graph-based Approximate Nearest Neighbor Search. (2023).","DOI":"10.1145\/3613424.3614292"},{"key":"e_1_2_1_104_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00517"},{"key":"e_1_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS47924.2020.00057"},{"key":"e_1_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE48307.2020.00094"},{"key":"e_1_2_1_107_1","doi-asserted-by":"publisher","DOI":"10.14778\/3594512.3594527"},{"key":"e_1_2_1_108_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2882930"},{"key":"e_1_2_1_109_1","volume-title":"Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving. arXiv preprint arXiv:2401.09670","author":"Zhong Yinmin","year":"2024","unstructured":"Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024. Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving. arXiv preprint arXiv:2401.09670 (2024)."},{"key":"e_1_2_1_110_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2915234"},{"key":"e_1_2_1_111_1","doi-asserted-by":"publisher","DOI":"10.14778\/3603581.3603601"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3749646.3749655","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,5]],"date-time":"2025-09-05T03:11:37Z","timestamp":1757041897000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3749646.3749655"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7]]},"references-count":111,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2025,7]]}},"alternative-id":["10.14778\/3749646.3749655"],"URL":"https:\/\/doi.org\/10.14778\/3749646.3749655","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2025,7]]},"assertion":[{"value":"2025-09-04","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}