{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,7]],"date-time":"2026-05-07T15:09:49Z","timestamp":1778166589906,"version":"3.51.4"},"reference-count":68,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2023,6,24]],"date-time":"2023-06-24T00:00:00Z","timestamp":1687564800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,6,24]],"date-time":"2023-06-24T00:00:00Z","timestamp":1687564800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001774","name":"University of Sydney","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001774","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2023,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Current object detectors typically have a feature pyramid (FP) module for multi-level feature fusion (MFF) which aims to mitigate the gap between features from different levels and form a comprehensive object representation to achieve better detection performance. However, they usually require heavy cross-level connections or iterative refinement to obtain better MFF result, making them complicated in structure and inefficient in computation. To address these issues, we propose a novel and efficient context modeling mechanism that can help existing FPs deliver better MFF results while reducing the computational costs effectively. In particular, we introduce a novel insight that comprehensive contexts can be decomposed and condensed into two types of representations for higher efficiency. The two representations include a locally concentrated representation and a globally summarized representation, where the former focuses on extracting context cues from nearby areas while the latter extracts general contextual representations of the whole image scene as global context cues. By collecting the condensed contexts, we employ a Transformer decoder to investigate the relations between them and each local feature from the FP and then refine the MFF results accordingly. As a result, we obtain a simple and light-weight Transformer-based Context Condensation (TCC) module, which can boost various FPs and lower their computational costs simultaneously. Extensive experimental results on the challenging MS COCO dataset show that TCC is compatible to four representative FPs and consistently improves their detection accuracy by up to 7.8% in terms of average precision and reduce their complexities by up to around 20% in terms of GFLOPs, helping them achieve state-of-the-art performance more efficiently. Code will be released at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/zhechen\/TCC\">https:\/\/github.com\/zhechen\/TCC<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s11263-023-01830-w","type":"journal-article","created":{"date-parts":[[2023,6,24]],"date-time":"2023-06-24T13:00:21Z","timestamp":1687611621000},"page":"2738-2756","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["Transformer-Based Context Condensation for Boosting Feature Pyramids in Object Detection"],"prefix":"10.1007","volume":"131","author":[{"given":"Zhe","family":"Chen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jing","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yufei","family":"Xu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dacheng","family":"Tao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,6,24]]},"reference":[{"key":"1830_CR1","unstructured":"Alayrac, J.B., Donahue. J., & Luc, P., et\u00a0al. (2022). Flamingo: A visual language model for few-shot learning. arXiv preprint arXiv:2204.14198."},{"key":"1830_CR2","unstructured":"Baker, B., Gupta, O., & Naik, N., et\u00a0al. (2016). Designing neural network architectures using reinforcement learning. In ICLR."},{"key":"1830_CR3","doi-asserted-by":"publisher","first-page":"213","DOI":"10.4324\/9781315512372-8","volume-title":"Perceptual organization","author":"I Biederman","year":"2017","unstructured":"Biederman, I. (2017). On the semantics of a glance at a scene. Perceptual organization (pp. 213\u2013253). Routledge."},{"key":"1830_CR4","doi-asserted-by":"crossref","unstructured":"Cai, Z., & Vasconcelos, N. (2018). Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 6154\u20136162).","DOI":"10.1109\/CVPR.2018.00644"},{"key":"1830_CR5","doi-asserted-by":"crossref","unstructured":"Cai, Z., Fan, Q., & Feris, R.S., et\u00a0al. (2016). A unified multi-scale deep convolutional neural network for fast object detection. In European Conference on Computer Vision (pp. 354\u2013370). Springer.","DOI":"10.1007\/978-3-319-46493-0_22"},{"key":"1830_CR6","doi-asserted-by":"crossref","unstructured":"Cao, Y., Xu, J., & Lin, S., et\u00a0al. (2019). Gcnet: Non-local networks meet squeeze-excitation networks and beyond. In Proceedings of the IEEE International Conference on Computer Vision Workshops.","DOI":"10.1109\/ICCVW.2019.00246"},{"key":"1830_CR7","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., & Synnaeve, G., et\u00a0al. (2020). End-to-end object detection with transformers. In European Conference on Computer Vision (pp. 213\u2013229). Springer.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"1830_CR8","doi-asserted-by":"crossref","unstructured":"Chen, K., Pang, J., & Wang, J., et\u00a0al. (2019a). Hybrid task cascade for instance segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 4974\u20134983).","DOI":"10.1109\/CVPR.2019.00511"},{"key":"1830_CR9","unstructured":"Chen, K., Cao, Y., & Loy, C.C., et\u00a0al. (2020). Feature pyramid grids. arXiv preprint arXiv:2004.03580."},{"key":"1830_CR10","doi-asserted-by":"crossref","unstructured":"Chen, X., & Gupta, A. (2017). Spatial memory for context reasoning in object detection. In Proceedings of the IEEE International Conference on Computer Vision (pp. 4086\u20134096).","DOI":"10.1109\/ICCV.2017.440"},{"key":"1830_CR11","doi-asserted-by":"crossref","unstructured":"Chen, Y., Fan, H., & Xu, B., et\u00a0al. (2019b). Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 3435\u20133444).","DOI":"10.1109\/ICCV.2019.00353"},{"key":"1830_CR12","doi-asserted-by":"crossref","unstructured":"Chen, Z., Huang, S., & Tao, D. (2018). Context refinement for object detection. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 71\u201386).","DOI":"10.1007\/978-3-030-01237-3_5"},{"issue":"1","key":"1830_CR13","doi-asserted-by":"publisher","first-page":"142","DOI":"10.1007\/s11263-020-01370-7","volume":"129","author":"Z Chen","year":"2021","unstructured":"Chen, Z., Zhang, J., & Tao, D. (2021). Recursive context routing for object detection. International Journal of Computer Vision, 129(1), 142\u2013160.","journal-title":"International Journal of Computer Vision"},{"key":"1830_CR14","doi-asserted-by":"crossref","unstructured":"Chen, Z., Zhang, J., & Tao, D. (2022). Recurrent glimpse-based decoder for detection with transformer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 5260\u20135269).","DOI":"10.1109\/CVPR52688.2022.00519"},{"key":"1830_CR15","unstructured":"Dai, J., Li, Y., & He, K., et\u00a0al. (2016). R-fcn: Object detection via region-based fully convolutional networks. In Advances in neural information processing systems, 29."},{"key":"1830_CR16","doi-asserted-by":"crossref","unstructured":"Dai, J., Qi, H., & Xiong, Y., et\u00a0al. (2017). Deformable convolutional networks. In Proceedings of the IEEE International Conference on Computer Vision (pp. 764\u2013773).","DOI":"10.1109\/ICCV.2017.89"},{"key":"1830_CR17","unstructured":"Dosovitskiy, A., Beyer, L., & Kolesnikov, A., et\u00a0al. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations."},{"key":"1830_CR18","unstructured":"Dosovitskiy, A., Beyer, L., & Kolesnikov, A., et\u00a0al. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR."},{"issue":"2","key":"1830_CR19","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","volume":"88","author":"M Everingham","year":"2010","unstructured":"Everingham, M., Van Gool, L., Williams, C. K., et al. (2010). The pascal visual object classes (voc) challenge. International Journal of Computer Vision, 88(2), 303\u2013338.","journal-title":"International Journal of Computer Vision"},{"key":"1830_CR20","doi-asserted-by":"crossref","unstructured":"Gao, P., Zheng, M., & Wang, X., et\u00a0al. (2021). Fast convergence of detr with spatially modulated co-attention. arXiv preprint arXiv:2101.07448.","DOI":"10.1109\/ICCV48922.2021.00360"},{"key":"1830_CR21","doi-asserted-by":"crossref","unstructured":"Ghiasi, G., Lin, T.Y., & Le, Q.V. (2019). Nas-fpn: Learning scalable feature pyramid architecture for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 7036\u20137045).","DOI":"10.1109\/CVPR.2019.00720"},{"key":"1830_CR22","doi-asserted-by":"crossref","unstructured":"Gidaris, S., & Komodakis, N. (2015). Object detection via a multi-region and semantic segmentation-aware CNN model. In Proceedings of the IEEE International Conference on Computer Vision (pp. 1134\u20131142).","DOI":"10.1109\/ICCV.2015.135"},{"key":"1830_CR23","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., & Ren, S., et\u00a0al. (2016). Deep residual learning for image recognition. In CVPR (pp 770\u2013778).","DOI":"10.1109\/CVPR.2016.90"},{"key":"1830_CR24","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., & Doll\u00e1r, P., et\u00a0al. (2017). Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision (pp. 2961\u20132969).","DOI":"10.1109\/ICCV.2017.322"},{"key":"1830_CR25","doi-asserted-by":"crossref","unstructured":"Hu, H., Gu, J., & Zhang, Z., et\u00a0al. (2018). Relation networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 3588\u20133597).","DOI":"10.1109\/CVPR.2018.00378"},{"key":"1830_CR26","unstructured":"Huang, G., Chen, D., & Li, T., et\u00a0al. (2017). Multi-scale dense convolutional networks for efficient prediction. arXiv preprint arXiv:1703.09844."},{"key":"1830_CR27","unstructured":"Jaegle, A., Borgeaud, S., & Alayrac, J.B., et\u00a0al. (2021a). Perceiver io: A general architecture for structured inputs & outputs. arXiv preprint arXiv:2107.14795."},{"key":"1830_CR28","unstructured":"Jaegle, A., Gimeno, F., & Brock, A., et\u00a0al. (2021b). Perceiver: General perception with iterative attention. In International conference on machine learning, PMLR, (pp. 4651\u20134664)."},{"key":"1830_CR29","unstructured":"Komodakis, N., & Gidaris, S. (2016). Attend refine repeat: Active box proposal generation via in-out localization. In BMVC."},{"key":"1830_CR30","doi-asserted-by":"crossref","unstructured":"Kong, T., Sun, F., & Yao, A., et\u00a0al. (2017). Ron: Reverse connection with objectness prior networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, (pp. 5936\u20135944).","DOI":"10.1109\/CVPR.2017.557"},{"key":"1830_CR31","doi-asserted-by":"crossref","unstructured":"Law, H., & Deng, J. (2018). Cornernet: Detecting objects as paired keypoints. In Proceedings of the European Conference on Computer Vision (ECCV), (pp. 734\u2013750).","DOI":"10.1007\/978-3-030-01264-9_45"},{"key":"1830_CR32","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., & Belongie, S., et\u00a0al. (2014). Microsoft coco: Common objects in context. In European Conference on Computer Vision, (pp. 740\u2013755). Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"1830_CR33","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., & Girshick, R., et\u00a0al. (2017a). Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, (pp 2117\u20132125).","DOI":"10.1109\/CVPR.2017.106"},{"key":"1830_CR34","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., & Girshick, R., et\u00a0al. (2017b). Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision, (pp. 2980\u20132988).","DOI":"10.1109\/ICCV.2017.324"},{"key":"1830_CR35","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., & Qin, H., et\u00a0al. (2018). Path aggregation network for instance segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, (pp. 8759\u20138768).","DOI":"10.1109\/CVPR.2018.00913"},{"key":"1830_CR36","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., & Erhan, D., et\u00a0al. (2016). Ssd: Single shot multibox detector. In European Conference on Computer Vision (pp. 21\u201337). Springer.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"1830_CR37","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., & Cao, Y., et\u00a0al. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 10012\u201310022).","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"1830_CR38","doi-asserted-by":"crossref","unstructured":"Meng, D., Chen, X., & Fan, Z., et\u00a0al. (2021). Conditional detr for fast training convergence. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 3651\u20133660).","DOI":"10.1109\/ICCV48922.2021.00363"},{"key":"1830_CR39","doi-asserted-by":"crossref","unstructured":"Newell, A., Yang, K., & Deng, J. (2016). Stacked hourglass networks for human pose estimation. In European Conference on Computer Vision (pp. 483\u2013499). Springer.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"1830_CR40","unstructured":"Qi, C.R., Su, H., & Mo, K., et\u00a0al. (2017). Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp 652\u2013660)."},{"key":"1830_CR41","doi-asserted-by":"crossref","unstructured":"Qiao, S., Chen, L.C., & Yuille, A. (2021). Detectors: Detecting objects with recursive feature pyramid and switchable atrous convolution. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 10213\u201310224).","DOI":"10.1109\/CVPR46437.2021.01008"},{"key":"1830_CR42","unstructured":"Ren, S., He, K., & Girshick, R., et\u00a0al. (2015). Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems 28."},{"issue":"1","key":"1830_CR43","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1037\/0033-295X.89.1.60","volume":"89","author":"DE Rumelhart","year":"1982","unstructured":"Rumelhart, D. E., & McClelland, J. L. (1982). An interactive activation model of context effects in letter perception: Ii. the contextual enhancement effect and some tests and extensions of the model. Psychological Review, 89(1), 60.","journal-title":"Psychological Review"},{"key":"1830_CR44","unstructured":"Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556."},{"key":"1830_CR45","doi-asserted-by":"crossref","unstructured":"Sun, K., Xiao, B., & Liu, D., et\u00a0al. (2019). Deep high-resolution representation learning for human pose estimation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 5693\u20135703).","DOI":"10.1109\/CVPR.2019.00584"},{"key":"1830_CR46","unstructured":"Tay, Y., Dehghani, M., & Bahri, D., et\u00a0al. (2020). Efficient transformers: A survey. arXiv preprint arXiv:2009.06732."},{"key":"1830_CR47","doi-asserted-by":"crossref","unstructured":"Tian, Z., Shen, C., & Chen, H., et\u00a0al. (2019). Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 9627\u20139636).","DOI":"10.1109\/ICCV.2019.00972"},{"key":"1830_CR48","doi-asserted-by":"publisher","first-page":"169","DOI":"10.1023\/A:1023052124951","volume":"53","author":"A Torralba","year":"2003","unstructured":"Torralba, A. (2003). Contextual priming for object detection. International Journal of Computer Vision, 53, 169\u2013191.","journal-title":"International Journal of Computer Vision"},{"key":"1830_CR49","unstructured":"Touvron, H., Cord, M., & Douze, M., et\u00a0al. (2021). Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, PMLR, (pp. 10347\u201310357)."},{"key":"1830_CR50","doi-asserted-by":"crossref","unstructured":"Uzkent, B., & Ermon, S. (2020). Learning when and where to zoom with deep reinforcement learning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 12345\u201312354).","DOI":"10.1109\/CVPR42600.2020.01236"},{"key":"1830_CR51","unstructured":"Vaswani, A., Shazeer, N., & Parmar, N., et\u00a0al. (2017). Attention is all you need. In NIPS."},{"key":"1830_CR52","unstructured":"Wang, W., Cao, Y., & Zhang, J., et\u00a0al. (2021a). Fp-detr: Detection transformer advanced by fully pre-training. In: International Conference on Learning Representations"},{"key":"1830_CR53","doi-asserted-by":"crossref","unstructured":"Wang, W., Zhang, J., & Cao, Y., et\u00a0al. (2022). Towards data-efficient detection transformers. In European Conference on Computer Vision (pp 88\u2013105). Springer.","DOI":"10.1007\/978-3-031-20077-9_6"},{"key":"1830_CR54","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., & Gupta, A., et al. (2018). Non-local neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 7794\u20137803).","DOI":"10.1109\/CVPR.2018.00813"},{"key":"1830_CR55","unstructured":"Wang, Y., Zhang, X., & Yang, T., et\u00a0al. (2021b). Anchor detr: Query design for transformer-based detector. arXiv preprint arXiv:2109.07107."},{"key":"1830_CR56","doi-asserted-by":"crossref","unstructured":"Wojke, N., & Bewley, A. (2018). Deep cosine metric learning for person re-identification. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV) (pp. 748\u2013756). IEEE.","DOI":"10.1109\/WACV.2018.00087"},{"key":"1830_CR57","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., & Doll\u00e1r, P., et\u00a0al. (2017). Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 1492\u20131500).","DOI":"10.1109\/CVPR.2017.634"},{"key":"1830_CR58","doi-asserted-by":"crossref","unstructured":"Xu, H., Yao, L., & Zhang, W., et\u00a0al. (2019). Auto-fpn: Automatic network architecture adaptation for object detection beyond classification. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 6649\u20136658).","DOI":"10.1109\/ICCV.2019.00675"},{"key":"1830_CR59","unstructured":"Xu, Y., Zhang, Q., & Zhang, J., et\u00a0al. (2021). Vitae: Vision transformer advanced by exploring intrinsic inductive bias. In Advances in Neural Information Processing Systems 34."},{"key":"1830_CR60","doi-asserted-by":"crossref","unstructured":"Yu, F., Wang, D., & Shelhamer, E., et\u00a0al. (2018). Deep layer aggregation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 2403\u20132412).","DOI":"10.1109\/CVPR.2018.00255"},{"issue":"9","key":"1830_CR61","doi-asserted-by":"publisher","first-page":"2109","DOI":"10.1109\/TPAMI.2017.2745563","volume":"40","author":"X Zeng","year":"2017","unstructured":"Zeng, X., Ouyang, W., Yan, J., et al. (2017). Crafting gbd-net for object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(9), 2109\u20132123.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1830_CR62","doi-asserted-by":"crossref","unstructured":"Zhang, N., Donahue, J., & Girshick, R., et\u00a0al. (2014). Part-based r-cnns for fine-grained category detection. In: European conference on computer vision (pp. 834\u2013849). Springer.","DOI":"10.1007\/978-3-319-10590-1_54"},{"key":"1830_CR63","doi-asserted-by":"crossref","unstructured":"Zhang, Q., Xu, Y., & Zhang, J., et\u00a0al. (2022). Vitaev2: Vision transformer advanced by exploring inductive bias for image recognition and beyond. arXiv preprint arXiv:2202.10108.","DOI":"10.1007\/s11263-022-01739-w"},{"key":"1830_CR64","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., & Qi, X., et\u00a0al. (2017). Pyramid scene parsing network. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2881\u20132890).","DOI":"10.1109\/CVPR.2017.660"},{"key":"1830_CR65","doi-asserted-by":"crossref","unstructured":"Zhao, Q., Sheng, T., & Wang, Y., et\u00a0al. (2019). M2det: A single-shot object detector based on multi-level feature pyramid network. In Proceedings of the AAAI Conference on Artificial Intelligence (pp 9259\u20139266).","DOI":"10.1609\/aaai.v33i01.33019259"},{"key":"1830_CR66","unstructured":"Zhu, X., Su, W., & Lu, L., et\u00a0al. (2020). Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159."},{"key":"1830_CR67","unstructured":"Zoph, B., & Le, Q.V. (2017). Neural architecture search with reinforcement learning. In ICLR."},{"key":"1830_CR68","doi-asserted-by":"crossref","unstructured":"Zoph, B., Vasudevan, V., & Shlens, J., et\u00a0al. (2018). Learning transferable architectures for scalable image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 8697\u20138710).","DOI":"10.1109\/CVPR.2018.00907"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-023-01830-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-023-01830-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-023-01830-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,8,19]],"date-time":"2023-08-19T02:13:49Z","timestamp":1692411229000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-023-01830-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,24]]},"references-count":68,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2023,10]]}},"alternative-id":["1830"],"URL":"https:\/\/doi.org\/10.1007\/s11263-023-01830-w","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,24]]},"assertion":[{"value":"14 July 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 May 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 June 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}