{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T16:22:46Z","timestamp":1781022166975,"version":"3.54.1"},"reference-count":40,"publisher":"Association for Computing Machinery (ACM)","issue":"7","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2022,3]]},"abstract":"<jats:p>Leaderless replication allows any replica to handle any type of request to achieve read scalability and high availability for distributed data stores. However, this entails burdensome coordination overhead of replication protocols, degrading write throughput. In addition, the data store still requires coordination for membership changes, making it hard to resolve server failures quickly. To this end, we present NetLR, a replicated data store architecture that supports high performance, fault tolerance, and linearizability simultaneously. The key idea of NetLR is moving the entire replication functions into the network by leveraging the switch as an on-path in-network replication orchestrator. Specifically, NetLR performs consistency-aware read scheduling, high-performance write coordination, and active fault adaptation in the network switch. Our in-network replication eliminates inter-replica coordination for writes and membership changes, providing high write performance and fast failure handling. NetLR can be implemented using programmable switches at a line rate with only 5.68% of additional memory usage. We implement a prototype of NetLR on an Intel Tofino switch and conduct extensive testbed experiments. Our evaluation results show that NetLR is the only solution that achieves high throughput and low latency and is robust to server failures.<\/jats:p>","DOI":"10.14778\/3523210.3523213","type":"journal-article","created":{"date-parts":[[2022,6,22]],"date-time":"2022-06-22T22:23:21Z","timestamp":1655936601000},"page":"1337-1349","source":"Crossref","is-referenced-by-count":12,"title":["In-network leaderless replication for distributed data stores"],"prefix":"10.14778","volume":"15","author":[{"given":"Gyuyeong","family":"Kim","sequence":"first","affiliation":[{"name":"Korea University, Seoul, South Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wonjun","family":"Lee","sequence":"additional","affiliation":[{"name":"Korea University, Seoul, South Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,6,22]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Last accessed date","author":"Cassandra Apache","year":"2022","unstructured":"[n.d.]. Apache Cassandra . https:\/\/cassandra.apache.org\/ , Last accessed date : March 25, 2022 . [n.d.]. Apache Cassandra. https:\/\/cassandra.apache.org\/, Last accessed date: March 25, 2022."},{"key":"e_1_2_1_2_1","volume-title":"Last accessed date","year":"2022","unstructured":"[n.d.]. Cavium XPliant Ethernet switch. https:\/\/www.openswitch.net\/cavium\/ , Last accessed date : March 25, 2022 . [n.d.]. Cavium XPliant Ethernet switch. https:\/\/www.openswitch.net\/cavium\/, Last accessed date: March 25, 2022."},{"key":"e_1_2_1_3_1","volume-title":"Last accessed date","year":"2022","unstructured":"[n.d.]. A fast, compliant alternative implementation of Python. https:\/\/www.pypy.org\/ , Last accessed date : March 25, 2022 . [n.d.]. A fast, compliant alternative implementation of Python. https:\/\/www.pypy.org\/, Last accessed date: March 25, 2022."},{"key":"e_1_2_1_4_1","volume-title":"Last accessed date","year":"2022","unstructured":"[n.d.]. pypacker: The fastest and simplest packet manipulation lib for Python. https:\/\/gitlab.com\/mike01\/pypacker , Last accessed date : March 25, 2022 . [n.d.]. pypacker: The fastest and simplest packet manipulation lib for Python. https:\/\/gitlab.com\/mike01\/pypacker, Last accessed date: March 25, 2022."},{"key":"e_1_2_1_5_1","volume-title":"Last accessed date","year":"2022","unstructured":"[n.d.]. RocksDB: A Persistent Key-Value Store for Flash and RAM Storage. https:\/\/rocksdb.org\/ , Last accessed date : March 25, 2022 . [n.d.]. RocksDB: A Persistent Key-Value Store for Flash and RAM Storage. https:\/\/rocksdb.org\/, Last accessed date: March 25, 2022."},{"key":"e_1_2_1_6_1","volume-title":"Last accessed date","year":"2022","unstructured":"[n.d.]. Tofino Programmable Switch. https:\/\/www.intel.com\/content\/www\/us\/en\/products\/network-io\/programmable-ethernet-switch\/tofino-series.html , Last accessed date : March 25, 2022 . [n.d.]. Tofino Programmable Switch. https:\/\/www.intel.com\/content\/www\/us\/en\/products\/network-io\/programmable-ethernet-switch\/tofino-series.html, Last accessed date: March 25, 2022."},{"key":"e_1_2_1_7_1","volume-title":"Last accessed date","year":"2022","unstructured":"2020. Advanced Congestion & Flow Control with Programmable Switches. https:\/\/opennetworking.org\/wp-content\/uploads\/2020\/04\/JK-Lee-Slide-Deck.pdf , Last accessed date : March 25, 2022 . 2020. Advanced Congestion & Flow Control with Programmable Switches. https:\/\/opennetworking.org\/wp-content\/uploads\/2020\/04\/JK-Lee-Slide-Deck.pdf, Last accessed date: March 25, 2022."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319893"},{"key":"e_1_2_1_9_1","volume-title":"Day","author":"Alsberg Peter A.","year":"1976","unstructured":"Peter A. Alsberg and John D . Day . 1976 . A Principle for Resilient Sharing of Distributed Resources. In Proc. of ICSE (San Francisco, California, USA). IEEE Computer Society Press , Washington, DC, USA, 562--570. Peter A. Alsberg and John D. Day. 1976. A Principle for Resilient Sharing of Distributed Resources. In Proc. of ICSE (San Francisco, California, USA). IEEE Computer Society Press, Washington, DC, USA, 562--570."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2656877.2656890"},{"key":"e_1_2_1_11_1","volume-title":"Du","author":"Cao Zhichao","year":"2020","unstructured":"Zhichao Cao , Siying Dong , Sagar Vemuri , and David H.C . Du . 2020 . Characterizing, Modeling , and Benchmarking RocksDB Key-Value Workloads at Facebook. In Proc. of USENIX FAST. USENIX Association , Santa Clara, CA. Zhichao Cao, Siying Dong, Sagar Vemuri, and David H.C. Du. 2020. Characterizing, Modeling, and Benchmarking RocksDB Key-Value Workloads at Facebook. In Proc. of USENIX FAST. USENIX Association, Santa Clara, CA."},{"key":"e_1_2_1_12_1","volume-title":"Redis in Action","author":"Carlson Josiah L.","unstructured":"Josiah L. Carlson . 2013. Redis in Action . Manning Publications Co. , USA. Josiah L. Carlson. 2013. Redis in Action. Manning Publications Co., USA."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1294261.1294281"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2018436.2018477"},{"key":"e_1_2_1_15_1","first-page":"12","article-title":"Mesa: Geo-Replicated, near Real-Time","volume":"7","author":"Gupta Ashish","year":"2014","unstructured":"Ashish Gupta , Fan Yang , Jason Govig , Adam Kirsch , Kelvin Chan , Kevin Lai , Shuo Wu , Sandeep Govind Dhoot , Abhilash Rajesh Kumar , Ankur Agiwal , Sanjay Bhansali , Mingsheng Hong , Jamie Cameron , Masood Siddiqi , David Jones , Jeff Shute , Andrey Gubarev , Shivakumar Venkataraman , and Divyakant Agrawal . 2014 . Mesa: Geo-Replicated, near Real-Time , Scalable Data Warehousing. Proc. VLDB Endow. 7 , 12 (Aug. 2014), 1259--1270. Ashish Gupta, Fan Yang, Jason Govig, Adam Kirsch, Kelvin Chan, Kevin Lai, Shuo Wu, Sandeep Govind Dhoot, Abhilash Rajesh Kumar, Ankur Agiwal, Sanjay Bhansali, Mingsheng Hong, Jamie Cameron, Masood Siddiqi, David Jones, Jeff Shute, Andrey Gubarev, Shivakumar Venkataraman, and Divyakant Agrawal. 2014. Mesa: Geo-Replicated, near Real-Time, Scalable Data Warehousing. Proc. VLDB Endow. 7, 12 (Aug. 2014), 1259--1270.","journal-title":"Scalable Data Warehousing. Proc. VLDB Endow."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/78969.78972"},{"key":"e_1_2_1_17_1","volume-title":"Proc. of USENIX ATC","author":"Hunt Patrick","year":"2010","unstructured":"Patrick Hunt , Mahadev Konar , Flavio P. Junqueira , and Benjamin Reed . 2010 . ZooKeeper: Wait-Free Coordination for Internet-Scale Systems . In Proc. of USENIX ATC ( Boston, MA). USENIX Association, USA, 11. Patrick Hunt, Mahadev Konar, Flavio P. Junqueira, and Benjamin Reed. 2010. ZooKeeper: Wait-Free Coordination for Internet-Scale Systems. In Proc. of USENIX ATC (Boston, MA). USENIX Association, USA, 11."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.14778\/3461535.3461551"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132747.3132764"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2011.5958223"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378496"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/279227.279229"},{"key":"e_1_2_1_23_1","volume-title":"Ports","author":"Li Jialin","year":"2016","unstructured":"Jialin Li , Ellis Michael , Naveen Kr. Sharma , Adriana Szekeres , and Dan R. K . Ports . 2016 . Just Say No to Paxos Overhead : Replacing Consensus with Network Ordering. In Proc. of USENIX OSDI (Savannah, GA, USA). USENIX Association , USA, 467--483. Jialin Li, Ellis Michael, Naveen Kr. Sharma, Adriana Szekeres, and Dan R. K. Ports. 2016. Just Say No to Paxos Overhead: Replacing Consensus with Network Ordering. In Proc. of USENIX OSDI (Savannah, GA, USA). USENIX Association, USA, 467--483."},{"key":"e_1_2_1_24_1","volume-title":"Ports","author":"Li Jialin","year":"2020","unstructured":"Jialin Li , Jacob Nelson , Ellis Michael , Xin Jin , and Dan R. K . Ports . 2020 . Pegasus : Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence Directories. In Proc. of USENIX OSDI. USENIX Association , 387--406. Jialin Li, Jacob Nelson, Ellis Michael, Xin Jin, and Dan R. K. Ports. 2020. Pegasus: Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence Directories. In Proc. of USENIX OSDI. USENIX Association, 387--406."},{"key":"e_1_2_1_25_1","volume-title":"Wenisch","author":"Lim Kevin","year":"2013","unstructured":"Kevin Lim , David Meisner , Ali G. Saidi , Parthasarathy Ranganathan , and Thomas F . Wenisch . 2013 . Thin Servers with Smart Pipes : Designing SoC Accelerators for Memcached. In Proc. of ISCA (Tel-Aviv, Israel). Association for Computing Machinery , New York, NY, USA, 36--47. Kevin Lim, David Meisner, Ali G. Saidi, Parthasarathy Ranganathan, and Thomas F. Wenisch. 2013. Thin Servers with Smart Pipes: Designing SoC Accelerators for Memcached. In Proc. of ISCA (Tel-Aviv, Israel). Association for Computing Machinery, New York, NY, USA, 36--47."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037731"},{"key":"e_1_2_1_27_1","volume-title":"Proc. of USENIX OSDI","author":"Mao Yanhua","year":"2008","unstructured":"Yanhua Mao , Flavio P. Junqueira , and Keith Marzullo . 2008 . Mencius: Building Efficient Replicated State Machines for WANs . In Proc. of USENIX OSDI ( San Diego, California). USENIX Association, USA, 369--384. Yanhua Mao, Flavio P. Junqueira, and Keith Marzullo. 2008. Mencius: Building Efficient Replicated State Machines for WANs. In Proc. of USENIX OSDI (San Diego, California). USENIX Association, USA, 369--384."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2517350"},{"key":"e_1_2_1_29_1","volume-title":"Proc. of USENIX NSDI","author":"Nishtala Rajesh","year":"2013","unstructured":"Rajesh Nishtala , Hans Fugal , Steven Grimm , Marc Kwiatkowski , Herman Lee , Harry C. Li , Ryan McElroy , Mike Paleczny , Daniel Peek , Paul Saab , David Stafford , Tony Tung , and Venkateshwaran Venkataramani . 2013 . Scaling Memcache at Facebook . In Proc. of USENIX NSDI ( Lombard, IL). USENIX Association, Berkeley, CA, USA, 385--398. Rajesh Nishtala, Hans Fugal, Steven Grimm, Marc Kwiatkowski, Herman Lee, Harry C. Li, Ryan McElroy, Mike Paleczny, Daniel Peek, Paul Saab, David Stafford, Tony Tung, and Venkateshwaran Venkataramani. 2013. Scaling Memcache at Facebook. In Proc. of USENIX NSDI (Lombard, IL). USENIX Association, Berkeley, CA, USA, 385--398."},{"key":"e_1_2_1_30_1","volume-title":"Proc. of USENIX ATC","author":"Ongaro Diego","year":"2014","unstructured":"Diego Ongaro and John Ousterhout . 2014 . In Search of an Understandable Consensus Algorithm . In Proc. of USENIX ATC ( Philadelphia, PA). USENIX Association, USA, 305--320. Diego Ongaro and John Ousterhout. 2014. In Search of an Understandable Consensus Algorithm. In Proc. of USENIX ATC (Philadelphia, PA). USENIX Association, USA, 305--320."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.14778\/1938545.1938549"},{"key":"e_1_2_1_32_1","volume-title":"Schneider","author":"Renesse Robbert Van","year":"2004","unstructured":"Robbert Van Renesse and Fred B . Schneider . 2004 . Chain Replication for Supporting High Throughput and Availability. In Proc. of USENIX OSDI. USENIX Association , San Francisco, CA, 91--104. Robbert Van Renesse and Fred B. Schneider. 2004. Chain Replication for Supporting High Throughput and Availability. In Proc. of USENIX OSDI. USENIX Association, San Francisco, CA, 91--104."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007263.3007267"},{"key":"e_1_2_1_34_1","volume-title":"Proc. of USENIX ATC","author":"Terrace Jeff","unstructured":"Jeff Terrace and Michael J. Freedman . 2009. Object Storage on CRAQ: High-Throughput Chain Replication for Read-Mostly Workloads . In Proc. of USENIX ATC ( San Diego, California). 11. Jeff Terrace and Michael J. Freedman. 2009. Object Storage on CRAQ: High-Throughput Chain Replication for Read-Mostly Workloads. In Proc. of USENIX ATC (San Diego, California). 11."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.14778\/2002938.2002939"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213957"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476273"},{"key":"e_1_2_1_38_1","volume-title":"Proc. of USENIX OSDI. USENIX Association, 191--208","author":"Yang Juncheng","unstructured":"Juncheng Yang , Yao Yue , and K. V. Rashmi . 2020. A large scale analysis of hundreds of in-memory cache clusters at Twitter . In Proc. of USENIX OSDI. USENIX Association, 191--208 . Juncheng Yang, Yao Yue, and K. V. Rashmi. 2020. A large scale analysis of hundreds of in-memory cache clusters at Twitter. In Proc. of USENIX OSDI. USENIX Association, 191--208."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137765.3137778"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.14778\/3368289.3368301"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3523210.3523213","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:49:16Z","timestamp":1672224556000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3523210.3523213"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3]]},"references-count":40,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2022,3]]}},"alternative-id":["10.14778\/3523210.3523213"],"URL":"https:\/\/doi.org\/10.14778\/3523210.3523213","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2022,3]]}}}