{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T02:47:07Z","timestamp":1777603627695,"version":"3.51.4"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2014,12,29]],"date-time":"2014-12-29T00:00:00Z","timestamp":1419811200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100000781","name":"European Research Council","doi-asserted-by":"publisher","award":["SmartGeometry 335373"],"award-info":[{"award-number":["SmartGeometry 335373"]}],"id":[{"id":"10.13039\/501100000781","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000266","name":"Engineering and Physical Sciences Research Council","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2014,12,29]]},"abstract":"<jats:p>Humans describe images in terms of nouns and adjectives while algorithms operate on images represented as sets of pixels. Bridging this gap between how humans would like to access images versus their typical representation is the goal of image parsing, which involves assigning object and attribute labels to pixels. In this article we propose treating nouns as object labels and adjectives as visual attribute labels. This allows us to formulate the image parsing problem as one of jointly estimating per-pixel object and attribute labels from a set of training images. We propose an efficient (interactive time) solution. Using the extracted labels as handles, our system empowers a user to verbally refine the results. This enables hands-free parsing of an image into pixel-wise object\/attribute labels that correspond to human semantics. Verbally selecting objects of interest enables a novel and natural interaction modality that can possibly be used to interact with new generation devices (e.g., smartphones, Google Glass, livingroom devices). We demonstrate our system on a large number of real-world images with varying complexity. To help understand the trade-offs compared to traditional mouse-based interactions, results are reported for both a large-scale quantitative evaluation and a user study.<\/jats:p>","DOI":"10.1145\/2682628","type":"journal-article","created":{"date-parts":[[2015,1,5]],"date-time":"2015-01-05T13:27:09Z","timestamp":1420464429000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":44,"title":["ImageSpirit"],"prefix":"10.1145","volume":"34","author":[{"given":"Ming-Ming","family":"Cheng","sequence":"first","affiliation":[{"name":"University of Oxford, Oxford, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuai","family":"Zheng","sequence":"additional","affiliation":[{"name":"University of Oxford, Oxford, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wen-Yan","family":"Lin","sequence":"additional","affiliation":[{"name":"Oxford Brookes University, Oxford, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vibhav","family":"Vineet","sequence":"additional","affiliation":[{"name":"Oxford Brookes University, Oxford, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paul","family":"Sturgess","sequence":"additional","affiliation":[{"name":"Oxford Brookes University, Oxford, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nigel","family":"Crook","sequence":"additional","affiliation":[{"name":"Oxford Brookes University, Oxford, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Niloy J.","family":"Mitra","sequence":"additional","affiliation":[{"name":"University College London, London, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Philip","family":"Torr","sequence":"additional","affiliation":[{"name":"University of Oxford, Oxford, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2014,12,29]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2009.01645.x"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1531326.1531330"},{"key":"e_1_2_2_3_1","unstructured":"B. Berlin and P. Kay. 1991. Basic Color Terms: Their Universality and Evolution. University of California Press.  B. Berlin and P. Kay. 1991. Basic Color Terms: Their Universality and Evolution. University of California Press."},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2019627.2019639"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/800250.807503"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2001.937505"},{"key":"e_1_2_2_7_1","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV'10)","author":"Branson S.","unstructured":"S. Branson , C. Wah , F. Schroff , B. Babenko , P. Welinder , P. Perona , and S. Belongie . 2010. Visual recognition with humans in the loop . In Proceedings of the European Conference on Computer Vision (ECCV'10) . 438--451. S. Branson, C. Wah, F. Schroff, B. Babenko, P. Welinder, P. Perona, and S. Belongie. 2010. Visual recognition with humans in the loop. In Proceedings of the European Conference on Computer Vision (ECCV'10). 438--451."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778864"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1618452.1618470"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995344"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.414"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778820"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.237"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/383259.383296"},{"key":"e_1_2_2_15_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'10)","author":"Farhadi A.","unstructured":"A. Farhadi , I. Endres , and D. Hoiem . 2010. Attribute-centric recognition for cross-category generalization . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'10) . A. Farhadi, I. Endres, and D. Hoiem. 2010. Attribute-centric recognition for cross-category generalization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'10)."},{"key":"e_1_2_2_16_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'09)","author":"Farhadi A.","unstructured":"A. Farhadi , I. Endres , D. Hoiem , and D. Forsyth . 2009. Describing objects by their attributes . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'09) . 1778--1785. A. Farhadi, I. Endres, D. Hoiem, and D. Forsyth. 2009. Describing objects by their attributes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'09). 1778--1785."},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000022288.19776.77"},{"key":"e_1_2_2_18_1","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS'07)","author":"Ferrari V.","unstructured":"V. Ferrari and A. Zisserman . 2007. Learning visual attributes . In Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS'07) . V. Ferrari and A. Zisserman. 2007. Learning visual attributes. In Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS'07)."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2012.03005.x"},{"key":"e_1_2_2_20_1","doi-asserted-by":"crossref","unstructured":"S. Henderson. 2008. Augmented reality for maintenance and repair. http:\/\/www.youtube.com\/watch&quest;v=mn-zvymlSvk.  S. Henderson. 2008. Augmented reality for maintenance and repair. http:\/\/www.youtube.com\/watch&quest;v=mn-zvymlSvk.","DOI":"10.21236\/ADA475554"},{"key":"e_1_2_2_21_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'12)","author":"Khan F. S.","unstructured":"F. S. Khan , R. Anwer , J. Van De Weijer, A. Bagdanov, M. Van-Rell, and A. Lopez. 2012. Color attributes for object detection . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'12) . 3306--3313. F. S. Khan, R. Anwer, J. Van De Weijer, A. Bagdanov, M. Van-Rell, and A. Lopez. 2012. Color attributes for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'12). 3306--3313."},{"key":"e_1_2_2_22_1","unstructured":"D. Koller and N. Friedman. 2009. Probabilistic Graphical Models: Principles and Techniques. MIT Press.   D. Koller and N. Friedman. 2009. Probabilistic Graphical Models: Principles and Techniques. MIT Press."},{"key":"e_1_2_2_23_1","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS'11)","author":"Krahenbuhl P.","unstructured":"P. Krahenbuhl and V. Koltun . 2011. Efficient inference in fully connected crfs with gaussian edge potentials . In Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS'11) . P. Krahenbuhl and V. Koltun. 2011. Efficient inference in fully connected crfs with gaussian edge potentials. In Proceedings of the Conference on Advances in Neural Information Processing Systems (NIPS'11)."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995466"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2009.5459248"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.24.104"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1276377.1276381"},{"key":"e_1_2_2_28_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'09)","author":"Lampert C. H.","unstructured":"C. H. Lampert , H. Nickisch , and S. Harmeling . 2009. Learning to detect unseen object classes by between-class attribute transfer . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'09) . 951--958. C. H. Lampert, H. Nickisch, and S. Harmeling. 2009. Learning to detect unseen object classes by between-class attribute transfer. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'09). 951--958."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2468356.2479533"},{"key":"e_1_2_2_30_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV'09)","author":"Lempitsky V.","unstructured":"V. Lempitsky , P. Kohli , C. Rother , and T. Sharp . 2009. Image segmentation with a bounding box prior . In Proceedings of the IEEE International Conference on Computer Vision (ICCV'09) . 277--284. V. Lempitsky, P. Kohli, C. Rother, and T. Sharp. 2009. Image segmentation with a bounding box prior. In Proceedings of the IEEE International Conference on Computer Vision (ICCV'09). 277--284."},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.1177"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015706.1015719"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1531326.1531375"},{"key":"e_1_2_2_34_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'08)","author":"Malisiewicz T.","unstructured":"T. Malisiewicz and A. A. Efros . 2008. Recognition by association via learning per-exemplar distances . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'08) . 1--8. T. Malisiewicz and A. A. Efros. 2008. Recognition by association via learning per-exemplar distances. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'08). 1--8."},{"key":"e_1_2_2_35_1","unstructured":"Microsoft. 2012. Microsoft speech platform--Sdk. http:\/\/www.microsoft.com\/download\/details.aspx&quest;id=27226.  Microsoft. 2012. Microsoft speech platform--Sdk. http:\/\/www.microsoft.com\/download\/details.aspx&quest;id=27226."},{"key":"e_1_2_2_36_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'12)","author":"Patterson G.","unstructured":"G. Patterson and J. Hays . 2012. Sun attribute database: Discovering, annotating, and recognizing scene attributes . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'12) . 2751--2758. G. Patterson and J. Hays. 2012. Sun attribute database: Discovering, annotating, and recognizing scene attributes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'12). 2751--2758."},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1017\/S0305004100027419"},{"key":"e_1_2_2_38_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV'07)","author":"Rabinovich A.","unstructured":"A. Rabinovich , A. Vedaldi , C. Galleguillos , E. Wiewiora , and S. Belongie . 2007. Objects in context . In Proceedings of the IEEE International Conference on Computer Vision (ICCV'07) . 1--8. A. Rabinovich, A. Vedaldi, C. Galleguillos, E. Wiewiora, and S. Belongie. 2007. Objects in context. In Proceedings of the IEEE International Conference on Computer Vision (ICCV'07). 1--8."},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015706.1015720"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-007-0109-1"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.26.62"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/1073204.1073274"},{"key":"e_1_2_2_44_1","unstructured":"Sunnybrook Hospital. 2008. Xbox kinect in the hospital operating room. http:\/\/www.youtube.com\/watch&quest;v=f5Ep3oqicVU.  Sunnybrook Hospital. 2008. Xbox kinect in the hospital operating room. http:\/\/www.youtube.com\/watch&quest;v=f5Ep3oqicVU."},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015330.1015422"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126260"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-012-0574-z"},{"key":"e_1_2_2_48_1","volume-title":"Proceedings of the ECML\/PKDD Workshop on Learning from Multi-Label Data (MLD'09)","author":"Tsoumakas G.","unstructured":"G. Tsoumakas , A. Dimou , E. Spyromitros-Xioufis , V. Mezaris , I. Kompatsiaris , and I. Vlahavas . 2009. Correlation-based pruning of stacked binary relevance models for multi-label learning . In Proceedings of the ECML\/PKDD Workshop on Learning from Multi-Label Data (MLD'09) . G. Tsoumakas, A. Dimou, E. Spyromitros-Xioufis, V. Mezaris, I. Kompatsiaris, and I. Vlahavas. 2009. Correlation-based pruning of stacked binary relevance models for multi-label learning. In Proceedings of the ECML\/PKDD Workshop on Learning from Multi-Label Data (MLD'09)."},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-005-6642-x"},{"key":"e_1_2_2_50_1","article-title":"SemanticPaint: Interactive 3d labeling and learning at your fingertips. ACM","author":"Valentin J.","year":"2014","unstructured":"J. Valentin , V. Vineet , M.-M. Cheng , D. Kim , S. Izadi , J. Shotton , P. Kohli , M. Niessner , A. Criminisi , and P. Torr . 2014 . SemanticPaint: Interactive 3d labeling and learning at your fingertips. ACM Trans. Graph. (to appear). J. Valentin, V. Vineet, M.-M. Cheng, D. Kim, S. Izadi, J. Shotton, P. Kohli, M. Niessner, A. Criminisi, and P. Torr. 2014. SemanticPaint: Interactive 3d labeling and learning at your fingertips. ACM Trans. Graph. (to appear).","journal-title":"Trans. Graph. (to appear)."},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2009.154"},{"key":"e_1_2_2_52_1","first-page":"1553","article-title":"Scene segmentation with crfs learned from partially labeled images","volume":"20","author":"Verbeek J.","year":"2007","unstructured":"J. Verbeek and W. Triggs . 2007 . Scene segmentation with crfs learned from partially labeled images . Adv. Neural Inf. Process. Syst. 20 , 1553 -- 1560 . J. Verbeek and W. Triggs. 2007. Scene segmentation with crfs learned from partially labeled images. Adv. Neural Inf. Process. Syst. 20, 1553--1560.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126539"},{"key":"e_1_2_2_54_1","volume-title":"Proceedings of the 11th European Conference on Computer Vision (ECCV'10)","author":"Wang Y.","unstructured":"Y. Wang and G. Mori . 2010. A discriminative latent model of object classes and attributes . In Proceedings of the 11th European Conference on Computer Vision (ECCV'10) . 155--168. Y. Wang and G. Mori. 2010. A discriminative latent model of object classes and attributes. In Proceedings of the 11th European Conference on Computer Vision (ECCV'10). 155--168."},{"key":"e_1_2_2_55_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'10)","author":"Xiao J.","unstructured":"J. Xiao , J. Hays , K. A. Ehinger , A. Oliva , and A. Torralba . 2010. Sun database: Large-scale scene recognition from abbey to zoo . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'10) . 3485--3492. J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba. 2010. Sun database: Large-scale scene recognition from abbey to zoo. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR'10). 3485--3492."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.411"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185595"},{"key":"e_1_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778863"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2682628","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2682628","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T06:16:53Z","timestamp":1750227413000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2682628"}},"subtitle":["Verbal Guided Image Parsing"],"short-title":[],"issued":{"date-parts":[[2014,12,29]]},"references-count":58,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2014,12,29]]}},"alternative-id":["10.1145\/2682628"],"URL":"https:\/\/doi.org\/10.1145\/2682628","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,12,29]]},"assertion":[{"value":"2013-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-12-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}