{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T21:09:23Z","timestamp":1775596163746,"version":"3.50.1"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,4,2]]},"abstract":"<jats:p>Many emerging AI applications demand retrieval systems that go beyond document-level relevance and capture token-level semantics. Multi-vector search addresses this need through fine-grained semantic matching between token-level embeddings of queries and documents. However, it introduces significant system challenges, including compute-intensive set-to-set scoring, complex candidate filtering, and high memory overhead from storing dense token-level embeddings. Prior systems mitigate these challenges through GPU-based similarity calculation and indexing structures (e.g., IVFPQ-GPU) but often at the cost of reduced retrieval accuracy and suffer from low GPU utilization. We present \\name, a GPU-based vector data management system that enables low-latency and high-recall multi-vector search on modern Superchip architectures. \\name achieves this through a combination of three novel optimizations: (1) \\maxivf, a GPU-native, compression-free index tailored for multi-vector search with fine-grained anchor vectors enabling scalable index construction and low-latency, high-accuracy candidate generation via CAGRA-based GPU routing; (2) \\chamferkernel, a highly-optimized GPU kernel that enables single-digit millisecond Chamfer scoring over tens of thousands of candidate documents; and (3) \\zerocomp, a tiered vector storage layer that supports on-demand, low-latency access to full-precision embeddings across Grace-Hopper NVLink-C2C interconnects. Together, these techniques enable \\name to perform multi-vector search over hundreds of millions of document token embeddings with unprecedented high recall and low latency. Compared to state-of-the-art systems like PLAID and MUVERA, \\name achieves an order-of-magnitude lower latency while significantly improving recall.<\/jats:p>","DOI":"10.1145\/3786706","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T17:54:13Z","timestamp":1775584453000},"page":"1-26","source":"Crossref","is-referenced-by-count":0,"title":["VecFlow-Chamfer: A GPU-based Data Management System for High-Performance Multi-Vector Search on Superchips"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-9860-6325","authenticated-orcid":false,"given":"Chenghao","family":"Mo","sequence":"first","affiliation":[{"name":"SSAIL Lab, UIUC, Urbana, IL, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4011-8903","authenticated-orcid":false,"given":"Ben","family":"Karsin","sequence":"additional","affiliation":[{"name":"Nvidia, Honolulu, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-6365-1853","authenticated-orcid":false,"given":"Philip","family":"Adams","sequence":"additional","affiliation":[{"name":"Microsoft, Redmond, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8165-166X","authenticated-orcid":false,"given":"Minjia","family":"Zhang","sequence":"additional","affiliation":[{"name":"SSAIL Lab, UIUC, Urbana, IL, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,7]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Advanced Micro Devices. 2023. AMD Instinct MI300A APU Architecture. Technical Report. AMD. https:\/\/www.amd.com\/en\/products\/accelerators\/instinct\/mi300\/mi300a.html"},{"key":"e_1_2_1_2_1","unstructured":"RAPIDS AI. 2025. cuVS. https:\/\/github.com\/rapidsai\/cuvs. Accessed: 2025-01-18."},{"key":"e_1_2_1_3_1","volume-title":"Efficient Indexing of Billion-Scale Datasets of Deep Descriptors. In CVPR 2016. 2055","author":"Babenko Artem","year":"2063","unstructured":"Artem Babenko and Victor S. Lempitsky. 2016. Efficient Indexing of Billion-Scale Datasets of Deep Descriptors. In CVPR 2016. 2055-2063."},{"key":"e_1_2_1_4_1","doi-asserted-by":"crossref","unstructured":"Ainesh Bakshi Piotr Indyk Rajesh Jayaram Sandeep Silwal and Erik Waingarten. 2023. A Near-Linear Time Algorithm for the Chamfer Distance. arXiv:2307.03043 [cs.DS] https:\/\/arxiv.org\/abs\/2307.03043","DOI":"10.52202\/075280-2918"},{"key":"e_1_2_1_5_1","first-page":"209","article-title":"Revisiting the Inverted Indices for Billion-Scale Approximate Nearest Neighbors","volume":"2018","author":"Baranchuk Dmitry","year":"2018","unstructured":"Dmitry Baranchuk, Artem Babenko, and Yury Malkov. 2018. Revisiting the Inverted Indices for Billion-Scale Approximate Nearest Neighbors. In ECCV 2018. 209-224.","journal-title":"ECCV"},{"key":"e_1_2_1_6_1","volume-title":"Annoy: Approximate Nearest Neighbors in C\/Python. GitHub Repository. https:\/\/github.com\/spotify\/annoy Spotify.","author":"Bernhardsson Erik","year":"2013","unstructured":"Erik Bernhardsson. 2013. Annoy: Approximate Nearest Neighbors in C\/Python. GitHub Repository. https:\/\/github.com\/spotify\/annoy Spotify."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-41062-8_28"},{"key":"e_1_2_1_8_1","unstructured":"Antoine Chaffin. 2025. Reason-ModernColBERT. https:\/\/huggingface.co\/lightonai\/Reason-ModernColBERT"},{"key":"e_1_2_1_9_1","volume-title":"SPTAG: A library for fast approximate nearest neighbor search. https:\/\/github.com\/Microsoft\/SPTAG","author":"Chen Qi","year":"2018","unstructured":"Qi Chen, Haidong Wang, Mingqin Li, Gang Ren, Scarlett Li, Jeffery Zhu, Jason Li, Chuanjie Liu, Lintao Zhang, and Jingdong Wang. 2018. SPTAG: A library for fast approximate nearest neighbor search. https:\/\/github.com\/Microsoft\/SPTAG"},{"key":"e_1_2_1_10_1","volume-title":"SPANN: Highly-efficient Billion-scale Approximate Nearest Neighbor Search. arXiv:2111.08566 [cs.DB] https:\/\/arxiv.org\/abs\/2111.08566","author":"Chen Qi","year":"2021","unstructured":"Qi Chen, Bing Zhao, Haidong Wang, Mingqin Li, Chuanjie Liu, Zengzhong Li, Mao Yang, and Jingdong Wang. 2021. SPANN: Highly-efficient Billion-scale Approximate Nearest Neighbor Search. arXiv:2111.08566 [cs.DB] https:\/\/arxiv.org\/abs\/2111.08566"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.3390\/s101211259"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2959100.2959190"},{"key":"e_1_2_1_13_1","volume-title":"MUVERA: Multi-Vector Retrieval via Fixed Dimensional Encodings. arXiv:2405.19504 [cs.DS] https:\/\/arxiv.org\/abs\/2405.19504","author":"Dhulipala Laxman","year":"2024","unstructured":"Laxman Dhulipala, Majid Hadian, Rajesh Jayaram, Jason Lee, and Vahab Mirrokni. 2024. MUVERA: Multi-Vector Retrieval via Fixed Dimensional Encodings. arXiv:2405.19504 [cs.DS] https:\/\/arxiv.org\/abs\/2405.19504"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2401.08281"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3067706"},{"key":"e_1_2_1_16_1","volume-title":"Brian Ebiyau, Eric Nyberg, and Teruko Mitamura.","author":"Gichamba Alex","year":"2024","unstructured":"Alex Gichamba, Tewodros Kederalah Idris, Brian Ebiyau, Eric Nyberg, and Teruko Mitamura. 2024. ColBERT Retrieval and Ensemble Response Scoring for Language Model Question Answering. arXiv:2408.10808 [cs.CL] https:\/\/arxiv.org\/abs\/2408.10808"},{"key":"e_1_2_1_17_1","unstructured":"Xiangnan He Lizi Liao Hanwang Zhang Liqiang Nie Xia Hu and Tat-Seng Chua. 2017. Neural Collaborative Filtering. arXiv:1708.05031 [cs.IR] https:\/\/arxiv.org\/abs\/1708.05031"},{"key":"e_1_2_1_18_1","volume-title":"Ravishankar Krishnawamy, and Rohan Kadekodi.","author":"Subramanya Suhas Jayaram","year":"2019","unstructured":"Suhas Jayaram Subramanya, Fnu Devvrit, Harsha Vardhan Simhadri, Ravishankar Krishnawamy, and Rohan Kadekodi. 2019. Diskann: Fast accurate billion-point nearest neighbor search on a single node. Advances in Neural Information Processing Systems, Vol. 32 (2019)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00405"},{"key":"e_1_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. arXiv:2004.12832 [cs.IR] https:\/\/arxiv.org\/abs\/2004.12832","DOI":"10.1145\/3397271.3401075"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00276"},{"key":"e_1_2_1_22_1","volume-title":"Tao Lei, Iftekhar Naim, Ming-Wei Chang, and Vincent Y. Zhao.","author":"Lee Jinhyuk","year":"2024","unstructured":"Jinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei, Iftekhar Naim, Ming-Wei Chang, and Vincent Y. Zhao. 2024. Rethinking the Role of Token Retrieval in Multi-Vector Retrieval. arXiv:2304.01982 [cs.CL] https:\/\/arxiv.org\/abs\/2304.01982"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/3495724.3496517"},{"key":"e_1_2_1_24_1","unstructured":"Yury A. Malkov and D. A. Yashunin. 2016. Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs. CoRR Vol. arXiv preprint abs\/1603.09320 (2016)."},{"key":"e_1_2_1_25_1","volume-title":"Milvus-docs: Conduct a hybrid search. https:\/\/github.com\/milvus-io\/milvus-docs\/blob\/v2.1.x\/site\/en\/userGuide\/search\/hybridsearch.md. Accessed","year":"2022","unstructured":"Milvus-io. 2022. Milvus-docs: Conduct a hybrid search. https:\/\/github.com\/milvus-io\/milvus-docs\/blob\/v2.1.x\/site\/en\/userGuide\/search\/hybridsearch.md. Accessed: 2025."},{"key":"e_1_2_1_26_1","volume-title":"MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv:1611.09268 [cs.CL] https:\/\/microsoft.github.io\/msmarco\/ Available at: https:\/\/microsoft.github.io\/msmarco\/.","author":"Nguyen Tri","year":"2016","unstructured":"Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv:1611.09268 [cs.CL] https:\/\/microsoft.github.io\/msmarco\/ Available at: https:\/\/microsoft.github.io\/msmarco\/."},{"key":"e_1_2_1_27_1","volume-title":"NVIDIA GH200 Grace Hopper Superchip Architecture. Whitepaper","author":"NVIDIA Corporation","unstructured":"NVIDIA Corporation. 2023. NVIDIA GH200 Grace Hopper Superchip Architecture. Whitepaper. NVIDIA Corporation. https:\/\/resources.nvidia.com\/en-us-grace-cpu\/nvidia-grace-hopper"},{"key":"e_1_2_1_28_1","volume-title":"NVIDIA GB200 Grace Blackwell Superchip. Product Technical Specification. https:\/\/www.nvidia.com\/en-us\/data-center\/gb200-nvl72\/ Accessed","author":"NVIDIA Corporation","year":"2024","unstructured":"NVIDIA Corporation. 2024. NVIDIA GB200 Grace Blackwell Superchip. Product Technical Specification. https:\/\/www.nvidia.com\/en-us\/data-center\/gb200-nvl72\/ Accessed: 2024. Available at: https:\/\/www.nvidia.com\/en-us\/data-center\/gb200-nvl72\/."},{"key":"e_1_2_1_29_1","first-page":"4236","volume-title":"CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUs. 2024 IEEE 40th International Conference on Data Engineering (ICDE)","author":"Ootomo Hiroyuki","year":"2023","unstructured":"Hiroyuki Ootomo, Akira Naruse, Corey J. Nolet, Ray Wang, Tamas B. Feh\u00e9r, and Y. Wang. 2023. CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUs. 2024 IEEE 40th International Conference on Data Engineering (ICDE) (2023), 4236-4247."},{"key":"e_1_2_1_30_1","volume-title":"Pinecone Systems","author":"Inc.","year":"2024","unstructured":"Inc. Pinecone Systems. 2024. Overview. https:\/\/docs.pinecone.io\/docs\/overview. Accessed: 2025."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557325"},{"key":"e_1_2_1_32_1","doi-asserted-by":"crossref","unstructured":"Keshav Santhanam Omar Khattab Jon Saad-Falcon Christopher Potts and Matei Zaharia. 2022b. ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction. arXiv:2112.01488 [cs.IR] https:\/\/arxiv.org\/abs\/2112.01488","DOI":"10.18653\/v1\/2022.naacl-main.272"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3729904"},{"key":"e_1_2_1_34_1","unstructured":"sigridjineth. 2025. muvera-py: Python Implementation of MUVERA (Multi-Vector Retrieval via Fixed Dimensional Encodings). https:\/\/github.com\/sigridjineth\/muvera-py."},{"key":"e_1_2_1_35_1","volume-title":"FAQ: All about the Google RankBrain algorithm. https:\/\/searchengineland.com\/faq-all-about-the-new-google-rankbrain-algorithm-234440.","author":"Sullivan Danny","year":"2018","unstructured":"Danny Sullivan. 2018. FAQ: All about the Google RankBrain algorithm. https:\/\/searchengineland.com\/faq-all-about-the-new-google-rankbrain-algorithm-234440. (2018)."},{"key":"e_1_2_1_36_1","volume-title":"BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models. arXiv:2104.08663 [cs.IR] https:\/\/arxiv.org\/abs\/2104.08663","author":"Thakur Nandan","year":"2021","unstructured":"Nandan Thakur, Nils Reimers, Andreas R\u00fcckl\u00e9, Abhishek Srivastava, and Iryna Gurevych. 2021. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models. arXiv:2104.08663 [cs.IR] https:\/\/arxiv.org\/abs\/2104.08663"},{"key":"e_1_2_1_37_1","volume-title":"Harsha Vardhan Simhadri, and Jyothi Vedurada","author":"Karthik","year":"2024","unstructured":"Karthik V., Saim Khan, Somesh Singh, Harsha Vardhan Simhadri, and Jyothi Vedurada. 2024. BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU. arXiv:2401.11324 [cs.DC] https:\/\/arxiv.org\/abs\/2401.11324"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-22747-0_7"},{"key":"e_1_2_1_39_1","volume-title":"Manning","author":"Yang Zhilin","year":"2018","unstructured":"Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. arXiv:1809.09600 [cs.CL] https:\/\/arxiv.org\/abs\/1809.09600"},{"key":"e_1_2_1_40_1","volume-title":"Phil Blunsom, and Stephen Pulman.","author":"Yu Lei","year":"2014","unstructured":"Lei Yu, Karl Moritz Hermann, Phil Blunsom, and Stephen Pulman. 2014. Deep Learning for Answer Sentence Selection. CoRR, Vol. abs\/1412.1632 (2014)."},{"key":"e_1_2_1_41_1","first-page":"552","volume-title":"GPU-accelerated Proximity Graph Approximate Nearest Neighbor Search and Construction. In 38th IEEE International Conference on Data Engineering, ICDE 2022","author":"Yu Yuanhang","year":"2022","unstructured":"Yuanhang Yu, Dong Wen, Ying Zhang, Lu Qin, Wenjie Zhang, and Xuemin Lin. 2022. GPU-accelerated Proximity Graph Approximate Nearest Neighbor Search and Construction. In 38th IEEE International Conference on Data Engineering, ICDE 2022, Kuala Lumpur, Malaysia, May 9-12, 2022. IEEE, 552-564."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357938"},{"key":"e_1_2_1_43_1","first-page":"1033","volume-title":"SONG: Approximate Nearest Neighbor Search on GPU. In 36th IEEE International Conference on Data Engineering, ICDE 2020","author":"Zhao Weijie","year":"2020","unstructured":"Weijie Zhao, Shulong Tan, and Ping Li. 2020. SONG: Approximate Nearest Neighbor Search on GPU. In 36th IEEE International Conference on Data Engineering, ICDE 2020, Dallas, TX, USA, April 20-24, 2020. IEEE, 1033-1044."}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786706","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T20:02:55Z","timestamp":1775592175000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786706"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,2]]},"references-count":43,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,4,2]]}},"alternative-id":["10.1145\/3786706"],"URL":"https:\/\/doi.org\/10.1145\/3786706","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,2]]}}}