{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,9,27]],"date-time":"2024-09-27T22:40:29Z","timestamp":1727476829231},"reference-count":27,"publisher":"Walter de Gruyter GmbH","issue":"1","license":[{"start":{"date-parts":[[2022,1,1]],"date-time":"2022-01-01T00:00:00Z","timestamp":1640995200000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,6,27]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Multi-person pose estimation is a challenging problem. Bottom-up methods have been greatly studied because the prediction speed of top-down methods is related to the number of people in the input image, making these methods difficult to apply in real-time environments. To solve the problems of scale sensitivity and quantization error in bottom-up methods, it is necessary to have a model that can predict multi-scale keypoints and refine quantization error. To achieve this, we propose context feature and refined network for multi-person pose estimation (CRNet), which can effectively solve the problems of scale sensitivity and quantization error in bottom-up methods. We use a multi-scale feature pyramid and context feature to achieve scale invariance of the network. We extract global and local features and then fuse them by attentional feature fusion (AFF) to obtain context feature that adapt to multi-scale keypoints. In addition, we propose an efficient refined network to solve the problem of quantization error and use multi-resolution supervised learning to further improve the prediction accuracy of CRNet. Comprehensive experiments are conducted on two benchmarks: COCO and MPII datasets. The average precision of CRNet reached 72.1 and 80.2%, respectively, surpassing most state-of-the-art methods.<\/jats:p>","DOI":"10.1515\/jisys-2022-0060","type":"journal-article","created":{"date-parts":[[2022,6,27]],"date-time":"2022-06-27T21:38:56Z","timestamp":1656365936000},"page":"780-794","source":"Crossref","is-referenced-by-count":1,"title":["CRNet: Context feature and refined network for multi-person pose estimation"],"prefix":"10.1515","volume":"31","author":[{"given":"Lanfei","family":"Zhao","sequence":"first","affiliation":[{"name":"The Higher Educational Key Laboratory for Measuring & Control Technology and Instrumentations of Heilongjiang Province, Harbin University of Science and Technology , Harbin 150080 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhihua","family":"Chen","sequence":"additional","affiliation":[{"name":"The Higher Educational Key Laboratory for Measuring & Control Technology and Instrumentations of Heilongjiang Province, Harbin University of Science and Technology , Harbin 150080 , China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"374","published-online":{"date-parts":[[2022,6,27]]},"reference":[{"key":"2022120618433921595_j_jisys-2022-0060_ref_001","doi-asserted-by":"crossref","unstructured":"Chen Y, Tian Y, He M. Monocular human pose estimation: A survey of deep learning-based methods. Computer Vis Image Underst. 2020;192:1\u201320.","DOI":"10.1016\/j.cviu.2019.102897"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_002","doi-asserted-by":"crossref","unstructured":"Newell A, Yang K, Deng J. Stacked hourglass networks for human pose estimation. Proceedings of the European Conference on Computer Vision. Amsterdam, Netherlands, Berlin: Springer; 2016, October 8\u201316. p. 483\u201399.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_003","doi-asserted-by":"crossref","unstructured":"Chu X, Yang W, Ouyang W, Ma C, Yuille AL, Wang X. Multi-context attention for human pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, Piscataway, USA: IEEE; 2017, July 21\u201326. p. 1831\u201340.","DOI":"10.1109\/CVPR.2017.601"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_004","doi-asserted-by":"crossref","unstructured":"Nie X, Feng J, Xing J, Xiao S, Yan S. Hierarchical contextual refinement networks for human pose estimation. IEEE Trans Image Process. 2018;28(2):924\u201336.","DOI":"10.1109\/TIP.2018.2872628"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_005","doi-asserted-by":"crossref","unstructured":"Wang Z, Liu G, Tian G. A parameter efficient human pose estimation method based on densely connected convolutional module. IEEE Access. 2018;6:58056\u201363.","DOI":"10.1109\/ACCESS.2018.2874307"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_006","doi-asserted-by":"crossref","unstructured":"Zhang F, Zhu X, Dai H, Ye M, Zhu C. Distribution-aware coordinate representation for human pose estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Seattle, Piscataway, USA: IEEE; 2020, June 16\u201320. p. 7093\u2013102.","DOI":"10.1109\/CVPR42600.2020.00712"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_007","doi-asserted-by":"crossref","unstructured":"Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, et al. Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision. Zurich, Switzerland, Berlin: Springer; 2014, September 5\u201312. p. 740\u201355.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_008","doi-asserted-by":"crossref","unstructured":"Andriluka M, Pishchulin L, Gehler P, Schiele B. 2d human pose estimation: New benchmark and state of the art analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Columbus, Piscataway, USA: IEEE; 2014, June 23\u201328. p. 3686\u201393.","DOI":"10.1109\/CVPR.2014.471"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_009","doi-asserted-by":"crossref","unstructured":"Papandreou G, Zhu T, Kanazawa N, Toshev A, Tompson J, Bregler C, et al. Towards accurate multi-person pose estimation in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, Piscataway, USA: IEEE; 2017, July 21\u201326. p. 4903\u201311.","DOI":"10.1109\/CVPR.2017.395"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_010","doi-asserted-by":"crossref","unstructured":"Chen Y, Wang Z, Peng Y, Zhang Z, Yu G, Sun J. Cascaded pyramid network for multi-person pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Salt Lake City, Piscataway, USA: IEEE; 2018, June 19\u201323. p. 7103\u201312.","DOI":"10.1109\/CVPR.2018.00742"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_011","doi-asserted-by":"crossref","unstructured":"Xiao B, Wu H, Wei Y. Simple baselines for human pose estimation and tracking. Proceedings of the European Conference on Computer Vision. Munich, Berlin, Germany: Springer; 2018, September 8\u201314. p. 466\u201381.","DOI":"10.1007\/978-3-030-01231-1_29"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_012","doi-asserted-by":"crossref","unstructured":"Sun K, Xiao B, Liu D, Wang J. Deep high-resolution representation learning for human pose estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Long Beach, Piscataway, USA: IEEE; 2019, June 15\u201321. p. 5693\u20135703.","DOI":"10.1109\/CVPR.2019.00584"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_013","doi-asserted-by":"crossref","unstructured":"Pishchulin L, Insafutdinov E, Tang S, Andres B, Andriluka M, Gehler PV, et al. Deepcut: Joint subset partition and labeling for multi person pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas, Piscataway, USA: IEEE; 2016, June 26\u2013July 1. p. 4929\u201337.","DOI":"10.1109\/CVPR.2016.533"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_014","doi-asserted-by":"crossref","unstructured":"Cao Z, Simon T, Wei SE, Sheikh Y. Realtime multi-person 2d pose estimation using part affinity fields. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, Piscataway, USA: IEEE; 2017, July 21\u201326. p. 7291\u20139.","DOI":"10.1109\/CVPR.2017.143"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_015","unstructured":"Newell A, Huang Z, Deng J. Associative embedding: End-to-end learning for joint detection and grouping. Adv Neural Inf Process Syst. 2017;30:2277\u201387."},{"key":"2022120618433921595_j_jisys-2022-0060_ref_016","doi-asserted-by":"crossref","unstructured":"Cheng B, Xiao B, Wang J, Shi H, Huang TS, Zhang L. Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Seattle, Piscataway, USA: IEEE; 2020, June 16\u201320. p. 5386\u201395.","DOI":"10.1109\/CVPR42600.2020.00543"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_017","unstructured":"Li J. Research on bottom-up approaches for multi-person pose estimation. PhD thesis. Hefei: University of Science and Technology of China; 2021."},{"key":"2022120618433921595_j_jisys-2022-0060_ref_018","doi-asserted-by":"crossref","unstructured":"Su K, Yu D, Xu Z, Geng X, Wang C. Multi-person pose estimation with enhanced channel-wise and spatial information. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Long Beach, Piscataway, USA: IEEE; 2019, June 15\u201321. p. 5674\u201382.","DOI":"10.1109\/CVPR.2019.00582"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_019","doi-asserted-by":"crossref","unstructured":"Dai Y, Gieseke F, Oehmcke S, Wu Y, Barnard K. Attentional feature fusion. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. (virtual), Piscataway: IEEE; 2021, January 5\u20139. p. 3560\u20139.","DOI":"10.1109\/WACV48630.2021.00360"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_020","unstructured":"Kingma DP, Ba J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980; 2014."},{"key":"2022120618433921595_j_jisys-2022-0060_ref_021","unstructured":"Tieleman T, Hinton G. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural Netw Mach Learn. 2012;4(2):26\u201331."},{"key":"2022120618433921595_j_jisys-2022-0060_ref_022","doi-asserted-by":"crossref","unstructured":"Papandreou G, Zhu T, Chen LC, Gidaris S, Tompson J, Murphy K. Personlab: Person pose estimation and instance segmentation with a bottom-up, part-based, geometric embedding model. Proceedings of the European Conference on Computer Vision. Munich, Berlin, Germany: Springer; 2018, September 8\u201314. p. 269\u201386.","DOI":"10.1007\/978-3-030-01264-9_17"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_023","doi-asserted-by":"crossref","unstructured":"Kreiss S, Bertoni L, Alahi A. Pifpaf: Composite fields for human pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Long Beach, Piscataway, USA: IEEE; 2019, June 15\u201321. p. 11977\u201386.","DOI":"10.1109\/CVPR.2019.01225"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_024","doi-asserted-by":"crossref","unstructured":"Nie X, Feng J, Zhang J, Yan S. Single-stage multi-person pose machines. Proceedings of the IEEE\/CVF International Conference on Computer Vision. Seoul, Korea, Piscataway: IEEE; 2019, October 27\u2013November 2. p. 6951\u201360.","DOI":"10.1109\/ICCV.2019.00705"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_025","doi-asserted-by":"crossref","unstructured":"Insafutdinov E, Andriluka M, Pishchulin L, Tang S, Levinkov E, Andres B, et al. Arttrack: Articulated multi-person tracking in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Honolulu, Piscataway, USA: IEEE; 2017, July 21\u201326. p. 6457\u201365.","DOI":"10.1109\/CVPR.2017.142"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_026","doi-asserted-by":"crossref","unstructured":"Duan P, Wang T, Cui M, Sang H, Sun Q. Multi-person pose estimation based on a deep convolutional neural network. J Vis Commun Image Representation. 2019;62:245\u201352.","DOI":"10.1016\/j.jvcir.2019.05.010"},{"key":"2022120618433921595_j_jisys-2022-0060_ref_027","doi-asserted-by":"crossref","unstructured":"Fieraru M, Khoreva A, Pishchulin L, Schiele B. Learning to refine human pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition workshops. Salt Lake City, Piscataway, USA: IEEE; 2018, June 19\u201323. p. 205\u201314.","DOI":"10.1109\/CVPRW.2018.00058"}],"container-title":["Journal of Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/jisys-2022-0060\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/jisys-2022-0060\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,27]],"date-time":"2024-09-27T22:26:50Z","timestamp":1727476010000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.degruyter.com\/document\/doi\/10.1515\/jisys-2022-0060\/html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,1]]},"references-count":27,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,9,16]]},"published-print":{"date-parts":[[2022,9,16]]}},"alternative-id":["10.1515\/jisys-2022-0060"],"URL":"https:\/\/doi.org\/10.1515\/jisys-2022-0060","relation":{},"ISSN":["2191-026X"],"issn-type":[{"type":"electronic","value":"2191-026X"}],"subject":[],"published":{"date-parts":[[2022,1,1]]}}}