{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T05:05:51Z","timestamp":1782709551076,"version":"3.54.5"},"reference-count":48,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2012,9,5]],"date-time":"2012-09-05T00:00:00Z","timestamp":1346803200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2013,1]]},"DOI":"10.1007\/s11263-012-0564-1","type":"journal-article","created":{"date-parts":[[2012,9,4]],"date-time":"2012-09-04T19:14:45Z","timestamp":1346786085000},"page":"184-204","source":"Crossref","is-referenced-by-count":334,"title":["Efficiently Scaling up Crowdsourced Video Annotation"],"prefix":"10.1007","volume":"101","author":[{"given":"Carl","family":"Vondrick","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Donald","family":"Patterson","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Deva","family":"Ramanan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2012,9,5]]},"reference":[{"key":"564_CR1","doi-asserted-by":"crossref","first-page":"584","DOI":"10.1145\/1015706.1015764","volume":"23","author":"A. Agarwala","year":"2004","unstructured":"Agarwala, A., Hertzmann, A., Salesin, D., & Seitz, S. (2004). Keyframe-based tracking for rotoscoping and animation. ACM Transactions on Graphics, ACM, 23, 584\u2013591.","journal-title":"ACM Transactions on Graphics, ACM"},{"key":"564_CR2","volume-title":"IEEE computer vision and pattern recognition","author":"K. Ali","year":"2011","unstructured":"Ali, K., Hasler, D., & Fleuret, F. (2011). Flowboost\u2013appearance learning from sparsely annotated video. In IEEE computer vision and pattern recognition."},{"key":"564_CR3","unstructured":"Anonymous (2012). http:\/\/www.visint.org\/ ."},{"key":"564_CR4","volume-title":"2012 AAAI spring symposium series","author":"A. Aydemir","year":"2012","unstructured":"Aydemir, A., Henell, D., Jensfelt, P., & Shilkrot, R. (2012). Kinect@ home: crowdsourcing a large 3d dataset of real environments. In 2012 AAAI spring symposium series."},{"issue":"4","key":"564_CR5","doi-asserted-by":"crossref","first-page":"685","DOI":"10.1016\/j.chb.2005.12.009","volume":"22","author":"B. Bailey","year":"2006","unstructured":"Bailey, B., & Konstan, J. (2006). On the need for attention-aware systems: measuring effects of interruption on task performance, error rate, and affective state. Computers in Human Behavior, 22(4), 685\u2013708.","journal-title":"Computers in Human Behavior"},{"issue":"10","key":"564_CR6","doi-asserted-by":"crossref","first-page":"767","DOI":"10.1073\/pnas.42.10.767","volume":"42","author":"R. Bellman","year":"1956","unstructured":"Bellman, R. (1956). Dynamic programming and Lagrange multipliers. Proceedings of the National Academy of Sciences of the United States of America, 42(10), 767.","journal-title":"Proceedings of the National Academy of Sciences of the United States of America"},{"key":"564_CR7","first-page":"626","volume-title":"CVPR 06","author":"A. Buchanan","year":"2006","unstructured":"Buchanan, A., & Fitzgibbon, A. (2006). Interactive feature tracking using kd trees and dynamic programming. In CVPR 06, Citeseer (Vol.\u00a01, pp.\u00a0626\u2013633)."},{"key":"564_CR8","unstructured":"Chen, J., Zou, W., & Ng, A. (2011). Personal communication."},{"key":"564_CR9","volume-title":"CVPR","author":"N. Dalal","year":"2005","unstructured":"Dalal, N., & Triggs, B. (2005). Histograms of oriented gradients for human detection. In CVPR."},{"key":"564_CR10","doi-asserted-by":"crossref","unstructured":"Demir\u00f6z, B., Salah, A., & Akarun, L. (2012). \u00c7evresel zeka uygulamalari i\u00e7in t\u00fcm y\u00f6nl\u00fc kamera kullanimi multi-omnidirectional cameras for ambient intelligence.","DOI":"10.1109\/SIU.2012.6204823"},{"key":"564_CR11","first-page":"710","volume-title":"Proc. CVPR","author":"J. Deng","year":"2009","unstructured":"Deng, J., Dong, W., Socher, R., Li, L., Li, K., & Fei-Fei, L. (2009). ImageNet: a large-scale hierarchical image database. In Proc. CVPR (pp.\u00a0710\u2013719)."},{"key":"564_CR12","volume-title":"CVPR workshop on advancing computer vision with humans in the loop","author":"I. Endres","year":"2010","unstructured":"Endres, I., Farhadi, A., Hoiem, D., & Forsyth, D. (2010). The benefits and challenges of collecting richer object annotations. In CVPR workshop on advancing computer vision with humans in the loop. New York: IEEE Press."},{"issue":"2","key":"564_CR13","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","volume":"88","author":"M. Everingham","year":"2010","unstructured":"Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., & Zisserman, A. (2010). The pascal visual object classes (voc) challenge. International Journal of Computer Vision, 88(2), 303\u2013338.","journal-title":"International Journal of Computer Vision"},{"key":"564_CR14","first-page":"1871","volume":"9","author":"R. Fan","year":"2008","unstructured":"Fan, R., Chang, K., Hsieh, C., Wang, X., & Lin, C. (2008). LIBLINEAR: a\u00a0library for large linear classification. Journal of Machine Learning Research, 9, 1871\u20131874.","journal-title":"Journal of Machine Learning Research"},{"key":"564_CR15","unstructured":"Felzenszwalb, P., & Huttenlocher, D. (2004). Distance transforms of sampled functions (Cornell Computing and Information Science Technical Report TR2004-1963)."},{"key":"564_CR16","first-page":"1","volume-title":"Proc. 6th IEEE international workshop on performance evaluation of tracking and surveillance","author":"R. Fisher","year":"2004","unstructured":"Fisher, R. (2004). The pets04 surveillance ground-truth data sets. In Proc. 6th IEEE international workshop on performance evaluation of tracking and surveillance (pp.\u00a01\u20135)."},{"key":"564_CR17","unstructured":"Huber, D. (2011). Personal communication."},{"key":"564_CR18","unstructured":"Kahle, B. (2010). http:\/\/www.archive.org\/details\/movies ."},{"key":"564_CR19","volume-title":"ICCV","author":"N. Kumar","year":"2009","unstructured":"Kumar, N., Berg, A. C., Belhumeur, P. N., & Nayar, S. K. (2009). Attribute and simile classifiers for face verification. In ICCV."},{"key":"564_CR20","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/CVPR.2008.4587756","volume-title":"IEEE conference on computer vision and pattern recognition, 2008. CVPR 2008","author":"I. Laptev","year":"2008","unstructured":"Laptev, I., Marszalek, M., Schmid, C., & Rozenfeld, B. (2008). Learning realistic human actions from movies. In IEEE conference on computer vision and pattern recognition, 2008. CVPR 2008 (pp.\u00a01\u20138). New York: IEEE Press."},{"key":"564_CR21","first-page":"1","volume-title":"IEEE conference on computer vision and pattern recognition, CVPR 2008","author":"C. Liu","year":"2008","unstructured":"Liu, C., Freeman, W., Adelson, E., & Weiss, Y. (2008). Human-assisted motion annotation. In IEEE conference on computer vision and pattern recognition, CVPR 2008 (pp.\u00a01\u20138)."},{"key":"564_CR22","unstructured":"Liu, W., & Lazebnik, S. (2011). Personal communication."},{"key":"564_CR23","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1145\/1054972.1055017","volume-title":"Proceedings of the SIGCHI conference on human factors in computing systems","author":"G. Mark","year":"2005","unstructured":"Mark, G., Gonzalez, V., & Harris, J. (2005). No task left behind? Examining the nature of fragmented work. In Proceedings of the SIGCHI conference on human factors in computing systems (pp.\u00a0321\u2013330). New York: ACM Press."},{"key":"564_CR24","unstructured":"Mihalcik, D., & Doermann, D. (2003). The design and implementation of ViPER (Technical report)."},{"issue":"1","key":"564_CR25","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1137\/0105003","volume":"5","author":"J. Munkres","year":"1957","unstructured":"Munkres, J. (1957). Algorithms for the assignment and transportation problems. Journal of the Society for Industrial and Applied Mathematics, 5(1), 32\u201338.","journal-title":"Journal of the Society for Industrial and Applied Mathematics"},{"key":"564_CR26","unstructured":"Oh, S. (2011). Personal communication."},{"key":"564_CR27","volume-title":"CVPR","author":"S. Oh","year":"2011","unstructured":"Oh, S., Hoogs, A., Perera, A., Cuntoor, N., Chen, C. C., Lee, J. T., Mukherjee, S., Aggarwal, J. K., Lee, H., Davis, L., Swears, E., Wang, X., Ji, Q., Reddy, K., Shah, M., Vondrick, C., Pirsiavash, H., Ramanan, D., Yuen, J., Torralba, A., Song, B., Fong, A., Roy-Chowdhury, A., & Desai, M. (2011). A\u00a0large-scale benchmark dataset for event recognition in surveillance video. In CVPR."},{"issue":"3","key":"564_CR28","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1023\/A:1011139631724","volume":"42","author":"A. Oliva","year":"2001","unstructured":"Oliva, A., & Torralba, A. (2001). Modeling the shape of the scene: a\u00a0holistic representation of the spatial envelope. International Journal of Computer Vision, 42(3), 145\u2013175.","journal-title":"International Journal of Computer Vision"},{"key":"564_CR29","volume-title":"CVPR","author":"H. Pirsiavash","year":"2012","unstructured":"Pirsiavash, H., & Ramanan, D. (2012). Detecting activities of daily living in first-person camera views. In CVPR."},{"key":"564_CR30","first-page":"1","volume-title":"IEEE 11th international conference on Computer vision, 2007. ICCV 2007","author":"D. Ramanan","year":"2007","unstructured":"Ramanan, D., Baker, S., & Kakade, S. (2007). Leveraging archival video for building face datasets. In IEEE 11th international conference on Computer vision, 2007. ICCV 2007 (pp.\u00a01\u20138). New York: IEEE Press."},{"key":"564_CR31","volume-title":"Alt.CHI session of CHI 2010 extended abstracts on human factors in computing systems","author":"J. Ross","year":"2010","unstructured":"Ross, J., Irani, L., Silberman, M. S., Zaldivar, A., & Tomlinson, B. (2010). Who are the crowdworkers? Shifting demographics in mechanical Turk. In Alt.CHI session of CHI 2010 extended abstracts on human factors in computing systems."},{"issue":"1","key":"564_CR32","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1007\/s11263-007-0090-8","volume":"77","author":"B. Russell","year":"2008","unstructured":"Russell, B., Torralba, A., Murphy, K., & Freeman, W. (2008). LabelMe: a database and web-based tool for image annotation. International Journal of Computer Vision, 77(1), 157\u2013173.","journal-title":"International Journal of Computer Vision"},{"key":"564_CR33","volume-title":"The paradox of choice: why more is less","author":"B. Schwartz","year":"2005","unstructured":"Schwartz, B. (2005). The paradox of choice: why more is less. New York: Harper Perennial."},{"key":"564_CR34","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1145\/1178677.1178722","volume-title":"Proceedings of the 8th ACM international workshop on multimedia information retrieval","author":"A. Smeaton","year":"2006","unstructured":"Smeaton, A., Over, P., & Kraaij, W. (2006). Evaluation campaigns and trecvid. In Proceedings of the 8th ACM international workshop on multimedia information retrieval (pp.\u00a0321\u2013330). New York: ACM Press."},{"key":"564_CR35","first-page":"61","volume":"51","author":"A. Sorokin","year":"2008","unstructured":"Sorokin, A., & Forsyth, D. (2008). Utility data annotation with Amazon Mechanical Turk. Urbana, 51, 61, 820.","journal-title":"Urbana"},{"issue":"11","key":"564_CR36","doi-asserted-by":"crossref","first-page":"1958","DOI":"10.1109\/TPAMI.2008.128","volume":"30","author":"A. Torralba","year":"2008","unstructured":"Torralba, A., Fergus, R., & Freeman, W. T. (2008). 80 million tiny images: a large dataset for nonparametric object and scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 30(11), 1958\u20131970.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"8","key":"564_CR37","doi-asserted-by":"crossref","first-page":"1467","DOI":"10.1109\/JPROC.2010.2050290","volume":"98","author":"A. Torralba","year":"2010","unstructured":"Torralba, A., Russell, B., & Yuen, J. (2010). Labelme: online image annotation and applications. Proceedings of the IEEE, 98(8), 1467\u20131484.","journal-title":"Proceedings of the IEEE"},{"key":"564_CR38","volume-title":"CVPR","author":"S. Vijayanarasimhan","year":"2009","unstructured":"Vijayanarasimhan, S., & Grauman, K. (2009). What\u2019s it going to cost you? Predicting effort vs. informativeness for multi-label image annotations. In CVPR."},{"key":"564_CR39","volume-title":"CVPR","author":"S. Vijayanarasimhan","year":"2010","unstructured":"Vijayanarasimhan, S., Jain, P., & Grauman, K. (2010). Far-sighted active learning on a budget for image and video recognition. In CVPR."},{"key":"564_CR40","first-page":"109","volume-title":"Proceedings of the British machine vision conference","author":"S. Vittayakorn","year":"2011","unstructured":"Vittayakorn, S., & Hays, J. (2011). Quality assessment for crowdsourced object annotations. In J. Hoey, S. McKenna, & E. Trucco (Eds.), Proceedings of the British machine vision conference (pp.\u00a0109\u2013110)."},{"key":"564_CR41","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1145\/985692.985733","volume-title":"Proceedings of the SIGCHI conference on human factors in computing systems","author":"L. Ahn Von","year":"2004","unstructured":"Von Ahn, L., & Dabbish, L. (2004). Labeling images with a computer game. In Proceedings of the SIGCHI conference on human factors in computing systems (pp.\u00a0319\u2013326). New York: ACM Press."},{"key":"564_CR42","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1145\/1124772.1124782","volume-title":"Proceedings of the SIGCHI conference on human factors in computing systems","author":"L. Ahn Von","year":"2006","unstructured":"Von Ahn, L., Liu, R., & Blum, M. (2006). Peekaboom: a game for locating objects in images. In Proceedings of the SIGCHI conference on human factors in computing systems (pp.\u00a055\u201364). New York: ACM Press."},{"key":"564_CR43","volume-title":"NIPS","author":"C. Vondrick","year":"2011","unstructured":"Vondrick, C., & Ramanan, D. (2011). Video annotation and tracking with active learning. In NIPS."},{"key":"564_CR44","doi-asserted-by":"crossref","first-page":"610","DOI":"10.1007\/978-3-642-15561-1_44","volume-title":"Computer Vision\u2013ECCV 2010","author":"C. Vondrick","year":"2010","unstructured":"Vondrick, C., Ramanan, D., & Patterson, D. (2010). Efficiently scaling up video annotation with crowdsourced marketplaces. In Computer Vision\u2013ECCV 2010 (pp.\u00a0610\u2013623)."},{"key":"564_CR45","first-page":"8","volume-title":"Neural information processing systems conference (NIPS)","author":"P. Welinder","year":"2010","unstructured":"Welinder, P., Branson, S., Belongie, S., & Perona, P. (2010). The multidimensional wisdom of crowds. In Neural information processing systems conference (NIPS) (Vol.\u00a06, p.\u00a08)."},{"key":"564_CR46","volume-title":"CVPR","author":"J. Xiao","year":"2010","unstructured":"Xiao, J., Hays, J., Ehinger, K., Oliva, A., & Torralba, A. (2010). Sun database: large-scale scene recognition from abbey to zoo. In CVPR."},{"key":"564_CR47","volume-title":"Acm computing surveys (CSUR)","author":"A. Yilmaz","year":"2006","unstructured":"Yilmaz, A., Javed, O., & Shah, M. (2006). Object tracking: a survey. In Acm computing surveys (CSUR)."},{"key":"564_CR48","volume-title":"International conference of computer vision","author":"J. Yuen","year":"2009","unstructured":"Yuen, J., Russell, B., Liu, C., & Torralba, A. (2009). LabelMe video: building a video database with human annotations. In International conference of computer vision."}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-012-0564-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11263-012-0564-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-012-0564-1","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,1,28]],"date-time":"2022-01-28T05:02:30Z","timestamp":1643346150000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11263-012-0564-1"}},"subtitle":["A Set of Best Practices for High Quality, Economical Video Labeling"],"short-title":[],"issued":{"date-parts":[[2012,9,5]]},"references-count":48,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2013,1]]}},"alternative-id":["564"],"URL":"https:\/\/doi.org\/10.1007\/s11263-012-0564-1","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,9,5]]}}}