{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,8]],"date-time":"2026-08-08T00:17:25Z","timestamp":1786148245721,"version":"3.56.0"},"reference-count":28,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2013,1,1]],"date-time":"2013-01-01T00:00:00Z","timestamp":1356998400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Commun. ACM"],"published-print":{"date-parts":[[2013,1]]},"abstract":"<jats:p>\n            We propose a new method to quickly and accurately predict human\n            <jats:italic>pose<\/jats:italic>\n            ---the 3D positions of body joints---from a single depth image, without depending on information from preceding frames. Our approach is strongly rooted in current object recognition strategies. By designing an intermediate representation in terms of body parts, the difficult pose estimation problem is transformed into a simpler per-pixel classification problem, for which efficient machine learning techniques exist. By using computer graphics to synthesize a very large dataset of training image pairs, one can train a classifier that estimates body part labels from test images invariant to pose, body shape, clothing, and other irrelevances. Finally, we generate confidence-scored 3D proposals of several body joints by reprojecting the classification result and finding local modes.\n          <\/jats:p>\n          <jats:p>The system runs in under 5ms on the Xbox 360. Our evaluation shows high accuracy on both synthetic and real test sets, and investigates the effect of several training parameters. We achieve state-of-the-art accuracy in our comparison with related work and demonstrate improved generalization over exact whole-skeleton nearest neighbor matching.<\/jats:p>","DOI":"10.1145\/2398356.2398381","type":"journal-article","created":{"date-parts":[[2013,1,2]],"date-time":"2013-01-02T13:23:15Z","timestamp":1357132995000},"page":"116-124","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1204,"title":["Real-time human pose recognition in parts from single depth images"],"prefix":"10.1145","volume":"56","author":[{"given":"Jamie","family":"Shotton","sequence":"first","affiliation":[{"name":"Microsoft Research, Cambridge, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Toby","family":"Sharp","sequence":"additional","affiliation":[{"name":"Microsoft Research, Cambridge, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alex","family":"Kipman","sequence":"additional","affiliation":[{"name":"Xbox Incubation"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrew","family":"Fitzgibbon","sequence":"additional","affiliation":[{"name":"Microsoft Research, Cambridge, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mark","family":"Finocchio","sequence":"additional","affiliation":[{"name":"Xbox Incubation"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrew","family":"Blake","sequence":"additional","affiliation":[{"name":"Microsoft Research, Cambridge, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mat","family":"Cook","sequence":"additional","affiliation":[{"name":"Microsoft Research, Cambridge, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Richard","family":"Moore","sequence":"additional","affiliation":[{"name":"ST-Ericsson"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2013,1]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/1896300.1896427"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.7.1545"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.993558"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1010933404324"},{"key":"e_1_2_1_5_1","unstructured":"CMU Mocap Database. http:\/\/mocap.cs.cmu.edu.  CMU Mocap Database. http:\/\/mocap.cs.cmu.edu."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.1000236"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2003.1211479"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2010.5540141"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/645314.649425"},{"key":"e_1_2_1_10_1","first-page":"38","author":"Gonzalez T.","year":"1985","unstructured":"Gonzalez , T. Clustering to minimize the maximum intercluster distance. Theor. Comp. Sci. 38 ( 1985 ). Gonzalez, T. Clustering to minimize the maximum intercluster distance. Theor. Comp. Sci. 38 (1985).","journal-title":"Theor. Comp. Sci."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2005.288"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2006.08.002"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2007.4408976"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of CVPR","author":"Ning H.","year":"2008","unstructured":"Ning , H. , Xu , W. , Gong , Y. , Huang , T.S. Discriminative learning of visual words for 3D human pose estimation . In Proceedings of CVPR ( 2008 ). Ning, H., Xu, W., Gong, Y., Huang, T.S. Discriminative learning of visual words for 3D human pose estimation. In Proceedings of CVPR (2008)."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-88688-4_32"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2010.5509559"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2006.10.016"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2003.1211504"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/946247.946721"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-88693-8_44"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/11744023_1"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2010.5543618"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/645315.649172"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2004.1315063"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2008.4587360"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/1576246.1531369"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.305"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/1775614.1775663"}],"container-title":["Communications of the ACM"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2398356.2398381","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2398356.2398381","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T08:18:47Z","timestamp":1750234727000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2398356.2398381"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,1]]},"references-count":28,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2013,1]]}},"alternative-id":["10.1145\/2398356.2398381"],"URL":"https:\/\/doi.org\/10.1145\/2398356.2398381","relation":{},"ISSN":["0001-0782","1557-7317"],"issn-type":[{"value":"0001-0782","type":"print"},{"value":"1557-7317","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,1]]},"assertion":[{"value":"2013-01-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}