{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T14:26:57Z","timestamp":1775053617752,"version":"3.50.1"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2015,12]]},"abstract":"<jats:p>Data labeling is a necessary but often slow process that impedes the development of interactive systems for modern data analysis. Despite rising demand for manual data labeling, there is a surprising lack of work addressing its high and unpredictable latency. In this paper, we introduce CLAMShell, a system that speeds up crowds in order to achieve consistently low-latency data labeling. We offer a taxonomy of the sources of labeling latency and study several large crowd-sourced labeling deployments to understand their empirical latency profiles. Driven by these insights, we comprehensively tackle each source of latency, both by developing novel techniques such as straggler mitigation and pool maintenance and by optimizing existing methods such as crowd retainer pools and active learning. We evaluate CLAMShell in simulation and on live workers on Amazon's Mechanical Turk, demonstrating that our techniques can provide an order of magnitude speedup and variance reduction over existing crowdsourced labeling strategies.<\/jats:p>","DOI":"10.14778\/2856318.2856331","type":"journal-article","created":{"date-parts":[[2016,2,1]],"date-time":"2016-02-01T14:10:31Z","timestamp":1454335831000},"page":"372-383","source":"Crossref","is-referenced-by-count":30,"title":["CLAMShell"],"prefix":"10.14778","volume":"9","author":[{"given":"Daniel","family":"Haas","sequence":"first","affiliation":[{"name":"AMPLab, UC Berkeley"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiannan","family":"Wang","sequence":"additional","affiliation":[{"name":"Simon Fraser University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eugene","family":"Wu","sequence":"additional","affiliation":[{"name":"Columbia University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael J.","family":"Franklin","sequence":"additional","affiliation":[{"name":"AMPLab, UC Berkeley"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,12]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/2452376.2452377"},{"key":"e_1_2_1_2_1","volume-title":"LASM","author":"Agarwal A.","year":"2011","unstructured":"A. Agarwal Sentiment analysis of Twitter data . LASM , 2011 . A. Agarwal et al. Sentiment analysis of Twitter data. LASM, 2011."},{"key":"e_1_2_1_3_1","volume-title":"OSDI","author":"Ananthanarayanan G.","year":"2010","unstructured":"G. Ananthanarayanan Reining in the Outliers in MapReduce Clusters using Mantri . OSDI , 2010 . G. Ananthanarayanan et al. Reining in the Outliers in MapReduce Clusters using Mantri. OSDI, 2010."},{"key":"e_1_2_1_4_1","volume-title":"NSDI","author":"Ananthanarayanan G.","year":"2013","unstructured":"G. Ananthanarayanan Effective Straggler Mitigation: Attack of the Clones . NSDI , 2013 . G. Ananthanarayanan et al. Effective Straggler Mitigation: Attack of the Clones. NSDI, 2013."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2047196.2047201"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1866029.1866078"},{"key":"e_1_2_1_7_1","volume-title":"Collective Intelligence","author":"Bernstein M. S.","year":"2012","unstructured":"M. S. Bernstein , D. R. Karger , R. C. Miller , and J. Brandt . Analytic Methods for Optimizing Realtime Crowdsourcing . Collective Intelligence , 2012 . M. S. Bernstein, D. R. Karger, R. C. Miller, and J. Brandt. Analytic Methods for Optimizing Realtime Crowdsourcing. Collective Intelligence, 2012."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1866029.1866080"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/1699510.1699548"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2014.2356470"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/1622737.1622744"},{"key":"e_1_2_1_12_1","volume-title":"ICDE","author":"Das A.","year":"2014","unstructured":"A. Das Sarma et al. Crowd-powered find algorithms . ICDE , 2014 . A. Das Sarma et al. Crowd-powered find algorithms. ICDE, 2014."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1327452.1327492"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2463710"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1989323.1989331"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/2733085.2733101"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767930"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2588555.2588576"},{"key":"e_1_2_1_19_1","volume-title":"OSDI","author":"Gonzalez J. E.","year":"2012","unstructured":"J. E. Gonzalez : Distributed Graph-Parallel Computation on Natural Graphs . OSDI , 2012 . J. E. Gonzalez et al. PowerGraph: Distributed Graph-Parallel Computation on Natural Graphs. OSDI, 2012."},{"key":"e_1_2_1_20_1","volume-title":"Design of experiments for the NIPS 2003 variable selection benchmark","author":"Guyon I.","year":"2003","unstructured":"I. Guyon . Design of experiments for the NIPS 2003 variable selection benchmark , 2003 . I. Guyon. Design of experiments for the NIPS 2003 variable selection benchmark, 2003."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824122"},{"key":"e_1_2_1_22_1","unstructured":"Hadoop. http:\/\/hadoop.apache.org\/.  Hadoop. http:\/\/hadoop.apache.org\/."},{"key":"e_1_2_1_23_1","volume-title":"Visualizations of the oDesk \"oConomy\": Exploring Our World of Work. https:\/\/www.upwork.com\/blog\/2012\/07\/visualizations-of-odesk-oconomy\/","author":"Ipeirotis P.","year":"2012","unstructured":"P. Ipeirotis and J. Horton . Visualizations of the oDesk \"oConomy\": Exploring Our World of Work. https:\/\/www.upwork.com\/blog\/2012\/07\/visualizations-of-odesk-oconomy\/ , 2012 . P. Ipeirotis and J. Horton. Visualizations of the oDesk \"oConomy\": Exploring Our World of Work. https:\/\/www.upwork.com\/blog\/2012\/07\/visualizations-of-odesk-oconomy\/, 2012."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1869086.1869094"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1837885.1837906"},{"key":"e_1_2_1_26_1","volume-title":"Science","author":"Jordan M. I.","year":"2015","unstructured":"M. I. Jordan and T. M. Mitchell . Machine learning: Trends, perspectives, and prospects . Science , 2015 . M. I. Jordan and T. M. Mitchell. Machine learning: Trends, perspectives, and prospects. Science, 2015."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732951.2732953"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1978942.1979444"},{"key":"e_1_2_1_29_1","volume-title":"Iterative Learning for Reliable Crowdsourcing Systems. Advances in neural information processing systems (NIPS)","author":"Karger D. R.","year":"2011","unstructured":"D. R. Karger , S. Oh , and D. Shah . Iterative Learning for Reliable Crowdsourcing Systems. Advances in neural information processing systems (NIPS) , 2011 . D. R. Karger, S. Oh, and D. Shah. Iterative Learning for Reliable Crowdsourcing Systems. Advances in neural information processing systems (NIPS), 2011."},{"key":"e_1_2_1_30_1","unstructured":"Keystone ML. http:\/\/keystone-ml.org\/.  Keystone ML. http:\/\/keystone-ml.org\/."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1357054.1357127"},{"key":"e_1_2_1_32_1","volume-title":"Learning multiple layers of features from tiny images","author":"Krizhevsky A.","year":"2009","unstructured":"A. Krizhevsky . Learning multiple layers of features from tiny images , 2009 . A. Krizhevsky. Learning multiple layers of features from tiny images, 2009."},{"key":"e_1_2_1_33_1","volume-title":"Work & Stress","author":"Krueger G. P.","year":"2007","unstructured":"G. P. Krueger . Sustained work, fatigue, sleep loss and performance: A review of the issues . Work & Stress , 2007 . G. P. Krueger. Sustained work, fatigue, sleep loss and performance: A review of the issues. Work & Stress, 2007."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.12129"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/2535568.2448944"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1561\/9781680830910"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.14778\/2047485.2047487"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920886"},{"key":"e_1_2_1_40_1","volume-title":"MLlib: Machine Learning in Apache Spark. arXiv.org","author":"Meng X.","year":"2015","unstructured":"X. Meng MLlib: Machine Learning in Apache Spark. arXiv.org , 2015 . X. Meng et al. MLlib: Machine Learning in Apache Spark. arXiv.org, 2015."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735471.2735474"},{"key":"e_1_2_1_42_1","unstructured":"Amazon Mechanical Turk. https:\/\/www.mturk.com\/.  Amazon Mechanical Turk. https:\/\/www.mturk.com\/."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.14778\/2556549.2556555"},{"key":"e_1_2_1_44_1","volume-title":"HCOMP","author":"Parameswaran A. G.","year":"2013","unstructured":"A. G. Parameswaran : An Expressive and Accurate Crowd-Powered Search Toolkit . HCOMP , 2013 . A. G. Parameswaran et al. DataSift: An Expressive and Accurate Crowd-Powered Search Toolkit. HCOMP, 2013."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_2_1_46_1","volume-title":"Identifying Reliable Workers Swiftly. Technical report","author":"Ramesh A.","year":"2012","unstructured":"A. Ramesh Identifying Reliable Workers Swiftly. Technical report , Stanford University , 2012 . A. Ramesh et al. Identifying Reliable Workers Swiftly. Technical report, Stanford University, 2012."},{"key":"e_1_2_1_47_1","volume-title":"Active learning literature survey. Technical report","author":"Settles B.","year":"2010","unstructured":"B. Settles . Active learning literature survey. Technical report , University of Wisconsin-Madison , 2010 . B. Settles. Active learning literature survey. Technical report, University of Wisconsin-Madison, 2010."},{"key":"e_1_2_1_48_1","volume-title":"VLDB","author":"Stonebraker M.","year":"2005","unstructured":"M. Stonebraker : a column-oriented DBMS . VLDB , 2005 . M. Stonebraker et al. C-store: a column-oriented DBMS. VLDB, 2005."},{"key":"e_1_2_1_49_1","volume-title":"CIDR","author":"Stonebraker M.","year":"2013","unstructured":"M. Stonebraker Data Curation at Scale: The Data Tamer System . CIDR , 2013 . M. Stonebraker et al. Data Curation at Scale: The Data Tamer System. CIDR, 2013."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2013.6544865"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.14778\/2350229.2350263"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/2588555.2610505"},{"key":"e_1_2_1_53_1","volume-title":"RStudio","author":"Wickham H.","year":"2013","unstructured":"H. Wickham . Bin-summarise-smooth : a framework for visualising large data. Technical report , RStudio , 2013 . H. Wickham. Bin-summarise-smooth: a framework for visualising large data. Technical report, RStudio, 2013."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732951.2732964"},{"key":"e_1_2_1_55_1","volume-title":"OSDI","author":"Zaharia M.","year":"2008","unstructured":"M. Zaharia Improving MapReduce Performance in Heterogeneous Environments . OSDI , 2008 . M. Zaharia et al. Improving MapReduce Performance in Heterogeneous Environments. OSDI, 2008."},{"key":"e_1_2_1_56_1","volume-title":"NSDI","author":"Zaharia M.","year":"2012","unstructured":"M. Zaharia Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing . NSDI , 2012 . M. Zaharia et al. Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. NSDI, 2012."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2012.148"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/2856318.2856331","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:24:14Z","timestamp":1672223054000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/2856318.2856331"}},"subtitle":["speeding up crowds for low-latency data labeling"],"short-title":[],"issued":{"date-parts":[[2015,12]]},"references-count":57,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2015,12]]}},"alternative-id":["10.14778\/2856318.2856331"],"URL":"https:\/\/doi.org\/10.14778\/2856318.2856331","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2015,12]]}}}