{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T14:25:02Z","timestamp":1784643902493,"version":"3.55.0"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2022,4,12]],"date-time":"2022-04-12T00:00:00Z","timestamp":1649721600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"NSF","award":["1812537"],"award-info":[{"award-number":["1812537"]}]},{"name":"NSF I\/UCRC Center on Intelligent Storage"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Storage"],"published-print":{"date-parts":[[2022,5,31]]},"abstract":"<jats:p>Active storage devices and in-storage computing are proposed and developed in recent years to effectively reduce the amount of required data traffic and to improve the overall application performance. They are especially preferred in the compute-storage disaggregated infrastructure. In both techniques, a simple computing module is added to storage devices\/servers such that some stored data can be processed in the storage devices\/servers before being transmitted to application servers. This can reduce the required network bandwidth and offload certain computing requirements from application servers to storage devices\/servers. However, several challenges exist when designing an in-storage computing- based architecture for applications. These include what computing functions need to be offloaded, how to design the protocol between in-storage modules and application servers, and how to deal with the caching issue in application servers.<\/jats:p>\n          <jats:p>\n            HBase is an important and widely used distributed Key-Value Store. It stores and indexes key-value pairs in large files in a storage system like HDFS. However, its performance especially read performance, is impacted by the heavy traffics between HBase RegionServers and storage servers in the compute-storage disaggregated infrastructure when the available network bandwidth is limited. We propose an\n            <jats:bold>I<\/jats:bold>\n            n-\n            <jats:bold>S<\/jats:bold>\n            torage-based\n            <jats:bold>HBase<\/jats:bold>\n            architecture, called\n            <jats:bold>\n              <jats:italic>IS-HBase<\/jats:italic>\n            <\/jats:bold>\n            , to improve the overall performance and to address the aforementioned challenges. First, IS-HBase executes a data pre-processing module (\n            <jats:bold>I<\/jats:bold>\n            n-\n            <jats:bold>S<\/jats:bold>\n            torage\n            <jats:bold>S<\/jats:bold>\n            can\n            <jats:bold>N<\/jats:bold>\n            er, called\n            <jats:bold>\n              <jats:italic>ISSN<\/jats:italic>\n            <\/jats:bold>\n            ) for some read queries and returns the requested key-value pairs to RegionServers instead of returning data blocks in HFile. IS-HBase carries out compactions in storage servers to reduce the large amount of data being transmitted through the network and thus the compaction execution time is effectively reduced. Second, a set of new protocols is proposed to address the communication and coordination between HBase RegionServers at computing nodes and ISSNs at storage nodes. Third, a new self-adaptive caching scheme is proposed to better serve the read queries with fewer I\/O operations and less network traffic. According to our experiments, the IS-HBase can reduce up to 97% network traffic for read queries and the throughput (queries per second) is significantly less affected by the fluctuation of available network bandwidth. The execution time of compaction in IS-HBase is only about 6.31% \u2013 41.84% of the execution time of legacy HBase. In general, IS-HBase demonstrates the potential of adopting in-storage computing for other data-intensive distributed applications to significantly improve performance in compute-storage disaggregated infrastructure.\n          <\/jats:p>","DOI":"10.1145\/3488368","type":"journal-article","created":{"date-parts":[[2022,3,29]],"date-time":"2022-03-29T11:38:40Z","timestamp":1648553920000},"page":"1-42","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["IS-HBase: An In-Storage Computing Optimized HBase with I\/O Offloading and Self-Adaptive Caching in Compute-Storage Disaggregated Infrastructure"],"prefix":"10.1145","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6950-1776","authenticated-orcid":false,"given":"Zhichao","family":"Cao","sequence":"first","affiliation":[{"name":"University of Minnesota, Twin Cities, Minneapolis, MN"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huibing","family":"Dong","sequence":"additional","affiliation":[{"name":"University of Minnesota, Twin Cities, Minneapolis, MN"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yixun","family":"Wei","sequence":"additional","affiliation":[{"name":"University of Minnesota, Twin Cities, Minneapolis, MN"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shiyong","family":"Liu","sequence":"additional","affiliation":[{"name":"Ocean University of China, Qingdao, Shandong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David H. C.","family":"Du","sequence":"additional","affiliation":[{"name":"University of Minnesota, Twin Cities, Minneapolis, MN"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,4,12]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2014. Seagate Kinetic HDD Product Manual. Retrieved on 17 Feb. 2022 from https:\/\/www.seagate.com\/www-content\/product-content\/hdd-fam\/kinetic-hdd\/_shared\/docs\/kinetic-product-manual.pdfSamsung_Key_Value_SSD_enables_High_Performance_Scaling-0.pdf."},{"key":"e_1_3_1_3_2","unstructured":"2017. Samsung Key Value SSD enables High Performance Scaling."},{"key":"e_1_3_1_4_2","unstructured":"2019. Amazon S3. Retrieved on 17 Feb. 2022 from https:\/\/aws.amazon.com\/cn\/s3\/."},{"key":"e_1_3_1_5_2","unstructured":"2019. Apache HBase."},{"key":"e_1_3_1_6_2","unstructured":"2019. Apache HDFS Users Guide. Retrieved on 17 Feb. 2022 from https:\/\/hadoop.apache.org\/docs\/stable\/hadoop-project-dist\/hadoop-hdfs\/HdfsUserGuide.html."},{"key":"e_1_3_1_7_2","unstructured":"2019. Iptables. Retrieved on 17 Feb. 2022 from https:\/\/en.wikipedia.org\/wiki\/Iptables."},{"key":"e_1_3_1_8_2","volume-title":"12th  \\( USENIX \\)  Workshop on Hot Topics in Cloud Computing. Pages 15.","author":"Angel Sebastian","year":"2020","unstructured":"Sebastian Angel, Mihir Nanavati, and Siddhartha Sen. 2020. Disaggregation and the application. In 12th \\( USENIX \\) Workshop on Hot Topics in Cloud Computing. Pages 15."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/2254756.2254766"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/2505515.2507847"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378504"},{"key":"e_1_3_1_12_2","unstructured":"Dhruba Borthakur. 2008. HDFS architecture guide. Hadoop Apache Project 53 1\u201313 (2008) 2. https:\/\/hadoop.apache.org\/docs\/r1.2.1\/hdfs_design.pdf."},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Brad Calder Ju Wang Aaron Ogus Niranjan Nilakantan Arild Skjolsvold Sam McKelvie Yikang Xu Shashwat Srivastav Jiesheng Wu Huseyin Simitci Jaidev Haridas Chakravarthy Uddaraju Hemal Khatri Andrew Edwards Vaman Bedekar Shane Mainali Rafay Abbasi Arpit Agarwal Mian Fahim ul Haq Muhammad Ikram ul Haq Deepali Bhardwaj Sowmya Dayanand Anitha Adusumilli Marvin McNett Sriram Sankaran Kavitha Manivannan and Leonidas Rigas. 2011. Windows azure storage: A highly available cloud storage service with strong consistency. In Proceedings of the 23rd ACM Symposium on Operating Systems Principles . ACM 143\u2013157.","DOI":"10.1145\/2043556.2043571"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/BigDataService.2017.30"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.5555\/3386691.3386712"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/1365815.1365816"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/1807128.1807152"},{"key":"e_1_3_1_18_2","unstructured":"Dharmesh Desai. 2019. The Cloud Advantage: Decoupling Storage and Compute. Retrieved on 17 Feb. 2022 from https:\/\/www.qubole.com\/blog\/advantage-decoupling\/."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2465295"},{"key":"e_1_3_1_20_2","volume-title":"HBase: The Definitive Guide: Random Access to Your Planet-size Data","author":"George Lars","year":"2011","unstructured":"Lars George. 2011. HBase: The Definitive Guide: Random Access to Your Planet-size Data. \u201cO\u2019Reilly Media, Inc.\u201d"},{"key":"e_1_3_1_21_2","doi-asserted-by":"crossref","unstructured":"Boncheol Gu Andre S. Yoon Duck-Ho Bae Insoon Jo Jinyoung Lee Jonghyun Yoon Jeong-Uk Kang Moonsang Kwon Chanho Yoon Sangyeun Cho Jaeheon Jeong and Duckhyun Chang. 2016. Biscuit: A framework for near-data processing of big data workloads. In Proceedings of the ACM SIGARCH Computer Architecture News . 153\u2013165.","DOI":"10.1145\/3007787.3001154"},{"key":"e_1_3_1_22_2","first-page":"199","volume-title":"Proceedings of the 12th USENIX Conference on File and Storage Technologies.","author":"Harter Tyler","year":"2014","unstructured":"Tyler Harter, Dhruba Borthakur, Siying Dong, Amitanand Aiyer, Liyin Tang, Andrea C. Arpaci-Dusseau, and Remzi H. Arpaci-Dusseau. 2014. Analysis of HDFS under HBase: A facebook messages case study. In Proceedings of the 12th USENIX Conference on File and Storage Technologies.199\u2013212."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3267809.3267827"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994512"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSST.2013.6558444"},{"key":"e_1_3_1_26_2","volume-title":"Proceeedings of the International Workshop on Accelerating Data Management Systems.","author":"Kim Sungchan","year":"2011","unstructured":"Sungchan Kim, Hyunok Oh, Chanik Park, Sangyeun Cho, and Sang-Won Lee. 2011. Fast, energy efficient scan inside flash memory SSDs. In Proceeedings of the International Workshop on Accelerating Data Management Systems."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2019.00035"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS.2017.00072"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/s002360050048"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2016.2595566"},{"key":"e_1_3_1_31_2","unstructured":"Ben Pfaff Justin Pettit Teemu Koponen Ethan Jackson Andy Zhou Jarno Rajahalme Jesse Gross Alex Wang Joe Stringer Pravin Shelar Keith Amidon Awake Networks and Mart\u00edn Casado. 2015. The design and implementation of open vswitch. In Proceedings of the 12th USENIX Networked Systems Design and Implementation . 117\u2013130."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/513\/4\/042024"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/2.928624"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.21236\/ADA341735"},{"key":"e_1_3_1_35_2","first-page":"62","volume-title":"Proceedings of the 24th Conference on Very Large Databases","author":"Riedel Erik","year":"1998","unstructured":"Erik Riedel, Garth Gibson, and Christos Faloutsos. 1998. Active storage for large-scale data mining and multimedia applications. In Proceedings of the 24th Conference on Very Large Databases. Citeseer, 62\u201373."},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Konstantin Shvachko Hairong Kuang Sanjay Radia and Robert Chansler. 2010. The hadoop distributed file system. In Proceedings of the IEEE 26th Symposium on Mass Storage Systems and Technologies . 1\u201310.","DOI":"10.1109\/MSST.2010.5496972"},{"key":"e_1_3_1_37_2","article-title":"Co-kv: A collaborative key-value store using near-data processing to improve compaction for the lsm-tree","author":"Sun Hui","year":"2018","unstructured":"Hui Sun, Wei Liu, Jianzhong Huang, and Weisong Shi. 2018. Co-kv: A collaborative key-value store using near-data processing to improve compaction for the lsm-tree. arXiv:1807.04151. Retrieved from https:\/\/arxiv.org\/abs\/1807.04151.","journal-title":"arXiv:1807.04151."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2873579"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCSNT.2011.6182030"},{"key":"e_1_3_1_40_2","first-page":"449","volume-title":"Proceedings of the 17th USENIX Symposium on Networked Systems Design and Implementation","author":"Vuppalapati Midhul","year":"2020","unstructured":"Midhul Vuppalapati, Justin Miron, Rachit Agarwal, Dan Truong, Ashish Motivala, and Thierry Cruanes. 2020. Building an elastic query engine on disaggregated storage. In Proceedings of the 17th USENIX Symposium on Networked Systems Design and Implementation. 449\u2013462."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/2933349.2933353"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2016.2608818"}],"container-title":["ACM Transactions on Storage"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3488368","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3488368","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3488368","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:24Z","timestamp":1750188624000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3488368"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,12]]},"references-count":41,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2022,5,31]]}},"alternative-id":["10.1145\/3488368"],"URL":"https:\/\/doi.org\/10.1145\/3488368","relation":{},"ISSN":["1553-3077","1553-3093"],"issn-type":[{"value":"1553-3077","type":"print"},{"value":"1553-3093","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,4,12]]},"assertion":[{"value":"2020-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-04-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}