{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T14:05:39Z","timestamp":1753884339180,"version":"3.41.2"},"reference-count":36,"publisher":"World Scientific Pub Co Pte Ltd","issue":"12","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62002160","62072238"],"award-info":[{"award-number":["62002160","62072238"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Jiangsu Key R&D Fund on Social Development","award":["BE2022789"],"award-info":[{"award-number":["BE2022789"]}]},{"name":"Scientific Research Fund of Nanjing Institute of Engineering","award":["ZKJ202003"],"award-info":[{"award-number":["ZKJ202003"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J CIRCUIT SYST COMP"],"published-print":{"date-parts":[[2023,8]]},"abstract":"<jats:p> Action recognition is a challenging task of modeling both spatial and temporal context. Numerous works focus on architectures modality and successfully make worthy progress on this task. While due to the redundancy in time and the limit of computation resources, several works focus on the efficiency study like frame sampling, some for untrimmed videos, and some for trimmed videos. With the intent of improving the effectiveness of action recognition, we propose a novel Computational Spatiotemporal Selector (CSS) to refine and reinforce the key frames with discriminative information in video. Specifically, CSS includes two modules: Temporal Adaptive Sampling (TAS) module and Spatial Frame Resolution (SFR) module. The former can refine the key frames in the temporal space for capturing the key motion information, while the latter can further zoom out some refined frames in the spatial space for eliminating the discrimination-irrelevant structural information. The proposed CSS is flexible to be embedded into most representative action recognition models. Experiments on two challenging action recognition benchmarks, i.e., ActivityNet1.3 and UCF101, show that the proposed CSS improves the performance over most existing models, not only on trimmed videos but also untrimmed videos. <\/jats:p>","DOI":"10.1142\/s0218126623502031","type":"journal-article","created":{"date-parts":[[2023,1,10]],"date-time":"2023-01-10T03:52:46Z","timestamp":1673322766000},"source":"Crossref","is-referenced-by-count":1,"title":["Learning Spatiotemporal-Selected Representations in Videos for Action Recognition"],"prefix":"10.1142","volume":"32","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3124-9461","authenticated-orcid":false,"given":"Jiachao","family":"Zhang","sequence":"first","affiliation":[{"name":"Artificial Intelligence Industrial, Technology Research Institute, Nanjing Institute of Technology, Nanjing 211167, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ying","family":"Tong","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Nanjing Institute of Technology, Nanjing 211167, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liangbao","family":"Jiao","sequence":"additional","affiliation":[{"name":"Artificial Intelligence Industrial, Technology Research Institute, Nanjing Institute of Technology, Nanjing 211167, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2023,2,23]]},"reference":[{"key":"S0218126623502031BIB001","doi-asserted-by":"publisher","DOI":"10.1142\/S021812662050190X"},{"key":"S0218126623502031BIB002","doi-asserted-by":"crossref","first-page":"2150096","DOI":"10.1142\/S0218126621500961","volume":"30","author":"Chen C.","year":"2021","journal-title":"J. Circuits Syst. Comput."},{"key":"S0218126623502031BIB003","doi-asserted-by":"crossref","first-page":"2250159","DOI":"10.1142\/S0218126622501596","volume":"31","author":"Assefa M.","year":"2022","journal-title":"J. Circuits Syst. Comput."},{"key":"S0218126623502031BIB004","doi-asserted-by":"crossref","first-page":"5281","DOI":"10.1109\/TCSVT.2022.3142771","volume":"32","author":"Shu X.","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"S0218126623502031BIB005","doi-asserted-by":"crossref","first-page":"3852","DOI":"10.1109\/TIP.2022.3175605","volume":"31","author":"Xu B.","year":"2022","journal-title":"IEEE Trans. Image Process."},{"journal-title":"IEEE Trans. Pattern Anal. Mach. Intell.","year":"2022","author":"Shu X.","key":"S0218126623502031BIB006"},{"key":"S0218126623502031BIB007","doi-asserted-by":"crossref","first-page":"2250214","DOI":"10.1142\/S0218126622502140","volume":"31","author":"Shi X.","year":"2022","journal-title":"J. Circuits Syst. Comput."},{"first-page":"2625","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Donahue J.","key":"S0218126623502031BIB008"},{"first-page":"1725","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Karpathy A.","key":"S0218126623502031BIB009"},{"key":"S0218126623502031BIB010","first-page":"20","volume-title":"European Conf. Computer Vision","author":"Wang L.","year":"2016"},{"first-page":"6469","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Yang X.","key":"S0218126623502031BIB011"},{"first-page":"6299","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Carreira J.","key":"S0218126623502031BIB012"},{"first-page":"6546","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Hara K.","key":"S0218126623502031BIB013"},{"first-page":"4489","volume-title":"Proc. IEEE In. Conf. Computer Vision","author":"Tran D.","key":"S0218126623502031BIB014"},{"first-page":"1278","volume-title":"Proc. IEEE\/CVF Conf. Computer Vision and Pattern Recognition","author":"Wu Z.","key":"S0218126623502031BIB015"},{"first-page":"6232","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision","author":"Korbar B.","key":"S0218126623502031BIB016"},{"key":"S0218126623502031BIB017","first-page":"86","volume-title":"European Conf. Computer Vision","author":"Meng Y.","year":"2020"},{"first-page":"6450","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Tran D.","key":"S0218126623502031BIB019"},{"first-page":"5552","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision","author":"Tran D.","key":"S0218126623502031BIB020"},{"first-page":"1910","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision Workshops","author":"Kopuklu O.","key":"S0218126623502031BIB021"},{"first-page":"5533","volume-title":"Proc. IEEE Int. Conf. Computer Vision","author":"Qiu Z.","key":"S0218126623502031BIB022"},{"first-page":"7794","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Wang X.","key":"S0218126623502031BIB024"},{"first-page":"305","volume-title":"Proc. European Conf. Computer Vision","author":"Xie S.","key":"S0218126623502031BIB025"},{"first-page":"7083","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision","author":"Lin J.","key":"S0218126623502031BIB026"},{"first-page":"6848","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Zhang X.","key":"S0218126623502031BIB027"},{"key":"S0218126623502031BIB029","first-page":"33","volume":"29","author":"Adelson E. H.","year":"1984","journal-title":"RCA Eng."},{"key":"S0218126623502031BIB030","first-page":"354","volume-title":"European Conf. Computer Vision","author":"Cai Z.","year":"2016"},{"first-page":"2117","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Lin T.-Y.","key":"S0218126623502031BIB031"},{"volume-title":"Int. Joint Conf. Artificial Intelligence","author":"Fan H.","key":"S0218126623502031BIB035"},{"first-page":"1513","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision","author":"Zhi Y.","key":"S0218126623502031BIB036"},{"first-page":"2678","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Yeung S.","key":"S0218126623502031BIB037"},{"first-page":"6202","volume-title":"Proc. IEEE\/CVF Int. Conf. Computer Vision","author":"Feichtenhofer C.","key":"S0218126623502031BIB038"},{"first-page":"203","volume-title":"Proc. IEEE\/CVF Conf. Computer Vision and Pattern Recognition","author":"Feichtenhofer C.","key":"S0218126623502031BIB039"},{"key":"S0218126623502031BIB040","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","volume":"13","author":"Wang Z.","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"S0218126623502031BIB041","first-page":"961","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition","author":"Caba Heilbron F.","year":"2015"},{"first-page":"3551","volume-title":"Proc. IEEE Int. Conf. Computer Vision","author":"Wang H.","key":"S0218126623502031BIB043"}],"container-title":["Journal of Circuits, Systems and Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218126623502031","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,7,11]],"date-time":"2023-07-11T08:23:47Z","timestamp":1689063827000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S0218126623502031"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,23]]},"references-count":36,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2023,8]]}},"alternative-id":["10.1142\/S0218126623502031"],"URL":"https:\/\/doi.org\/10.1142\/s0218126623502031","relation":{},"ISSN":["0218-1266","1793-6454"],"issn-type":[{"type":"print","value":"0218-1266"},{"type":"electronic","value":"1793-6454"}],"subject":[],"published":{"date-parts":[[2023,2,23]]},"article-number":"2350203"}}