{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,27]],"date-time":"2025-12-27T10:06:57Z","timestamp":1766830017143,"version":"3.41.0"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2024,4,25]],"date-time":"2024-04-25T00:00:00Z","timestamp":1714003200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["92370119, 62376113, 62206225"],"award-info":[{"award-number":["92370119, 62376113, 62206225"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Jiangsu Science and Technology Programme","award":["BE2020006-4"],"award-info":[{"award-number":["BE2020006-4"]}]},{"name":"Natural Science Foundation of the Jiangsu Higher Education Institutions of China","award":["22KJB520039"],"award-info":[{"award-number":["22KJB520039"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,7,31]]},"abstract":"<jats:p>\n            Generalised image outpainting is an important and active research topic in computer vision, which aims to extend appealing content all-side around a given image. Existing state-of-the-art outpainting methods often rely on discrete extrapolation to extend the feature map in the bottleneck. They thus suffer from content unsmoothness, especially in circumstances where the outlines of objects in the extrapolated regions are incoherent with the input sub-images. To mitigate this issue, we design a novel bottleneck with Neural ODEs to make continuous extrapolation in latent space, which could be a plug-in for many deep learning frameworks. Our ODE-based network continuously transforms the state and makes accurate predictions by learning the incremental relationship among latent points, leading to both smooth and structured feature representation. Experimental results on three real-world datasets both applied on transformer-based and CNN-based frameworks show that our methods could generate more realistic and coherent images against the state-of-the-art image outpainting approaches. Our code is available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/PengleiGao\/Continuous-Image-Outpainting-with-Neural-ODE\">https:\/\/github.com\/PengleiGao\/Continuous-Image-Outpainting-with-Neural-ODE<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3648367","type":"journal-article","created":{"date-parts":[[2024,3,2]],"date-time":"2024-03-02T11:22:26Z","timestamp":1709378546000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Continuous Image Outpainting with Neural ODE"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1935-5752","authenticated-orcid":false,"given":"Penglei","family":"Gao","sequence":"first","affiliation":[{"name":"Department of Computer Science, University of Liverpool, Liverpool, U.K. Department of Foundational Mathematics, Xi\u2019an Jiaotong-Liverpool University, Suzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8600-2570","authenticated-orcid":false,"given":"Xi","family":"Yang","sequence":"additional","affiliation":[{"name":"Department of Intelligent Science, Xi\u2019an Jiaotong-Liverpool University, Suzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8104-5432","authenticated-orcid":false,"given":"Rui","family":"Zhang","sequence":"additional","affiliation":[{"name":"Department of Foundational Mathematics, Xi\u2019an Jiaotong-Liverpool University, Suzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3034-9639","authenticated-orcid":false,"given":"Kaizhu","family":"Huang","sequence":"additional","affiliation":[{"name":"Data Science Research Center, Duke Kunshan University, Kunshan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,4,25]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"205","volume-title":"Proceedings of the European Conference on Computer Vision Workshops","volume":"13803","author":"Cao Hu","year":"2022","unstructured":"Hu Cao, Yueyue Wang, Joy Chen, Dongsheng Jiang, Xiaopeng Zhang, Qi Tian, and Manning Wang. 2022. Swin-Unet: Unet-like pure transformer for medical image segmentation. In Proceedings of the European Conference on Computer Vision Workshops, Vol. 13803. 205\u2013218."},{"key":"e_1_3_1_3_2","first-page":"11315","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chang Huiwen","year":"2022","unstructured":"Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman. 2022. MaskGIT: Masked generative image transformer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 11315\u201311325."},{"key":"e_1_3_1_4_2","first-page":"6572","volume-title":"Advances in Annual Conference on Neural Information Processing Systems","volume":"31","author":"Chen Ricky T. Q.","year":"2018","unstructured":"Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David K. Duvenaud. 2018. Neural ordinary differential equations. In Advances in Annual Conference on Neural Information Processing Systems, Vol. 31. 6572\u20136583."},{"key":"e_1_3_1_5_2","first-page":"11431","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cheng Yen-Chi","year":"2022","unstructured":"Yen-Chi Cheng, Chieh Hubert Lin, Hsin-Ying Lee, Jian Ren, Sergey Tulyakov, and Ming-Hsuan Yang. 2022. InOut: Diverse image outpainting via GAN inversion. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 11431\u201311440."},{"key":"e_1_3_1_6_2","first-page":"636","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Cong Yuren","year":"2020","unstructured":"Yuren Cong, Hanno Ackermann, Wentong Liao, Michael Ying Yang, and Bodo Rosenhahn. 2020. NODIS: Neural ordinary differential scene understanding. In Proceedings of the European Conference on Computer Vision. Springer, 636\u2013653."},{"key":"e_1_3_1_7_2","first-page":"248","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Deng Jia","year":"2009","unstructured":"Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 248\u2013255."},{"key":"e_1_3_1_8_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_9_2","volume-title":"Advances in Annual Conference on Neural Information Processing Systems","volume":"29","author":"Dosovitskiy Alexey","year":"2016","unstructured":"Alexey Dosovitskiy and Thomas Brox. 2016. Generating images with perceptual similarity metrics based on deep networks. In Advances in Annual Conference on Neural Information Processing Systems, Vol. 29."},{"key":"e_1_3_1_10_2","first-page":"2286","volume-title":"Proceedings of the International Conference on Machine Learning","author":"D\u2019Ascoli St\u00e9phane","year":"2021","unstructured":"St\u00e9phane D\u2019Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos, Giulio Biroli, and Levent Sagun. 2021. ConViT: Improving vision transformers with soft convolutional inductive biases. In Proceedings of the International Conference on Machine Learning. PMLR, 2286\u20132296."},{"key":"e_1_3_1_11_2","first-page":"12873","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Esser Patrick","year":"2021","unstructured":"Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12873\u201312883."},{"issue":"6","key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"2296","DOI":"10.1007\/s12559-022-10029-z","article-title":"Atten-GAN: Pedestrian trajectory prediction with GAN based on attention mechanism","volume":"14","author":"Fang Fang","year":"2022","unstructured":"Fang Fang, Pengpeng Zhang, Bo Zhou, Kun Qian, and Yahui Gan. 2022. Atten-GAN: Pedestrian trajectory prediction with GAN based on attention mechanism. Cognitive Computation 14, 6 (2022), 2296\u20132305.","journal-title":"Cognitive Computation"},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.neunet.2023.02.021","article-title":"Generalised image outpainting with U-Transformer","volume":"162","author":"Gao Penglei","year":"2023","unstructured":"Penglei Gao, Xi Yang, Rui Zhang, John Y. Goulermas, Yujie Geng, Yuyao Yan, and Kaizhu Huang. 2023. Generalised image outpainting with U-Transformer. Neural Networks 162 (2023), 1\u201310.","journal-title":"Neural Networks"},{"key":"e_1_3_1_14_2","article-title":"LeViT: A vision transformer in ConvNet\u2019s clothing for faster inference","volume":"2104","author":"Graham Benjamin","year":"2021","unstructured":"Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herv\u00e9 J\u00e9gou, and Matthijs Douze. 2021. LeViT: A vision transformer in ConvNet\u2019s clothing for faster inference. CoRR abs\/2104.01136 (2021).","journal-title":"CoRR"},{"key":"e_1_3_1_15_2","first-page":"770","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 770\u2013778."},{"key":"e_1_3_1_16_2","first-page":"1732","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He Xiangyu","year":"2019","unstructured":"Xiangyu He, Zitao Mo, Peisong Wang, Yang Liu, Mingyuan Yang, and Jian Cheng. 2019. ODE-inspired network design for single image super-resolution. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1732\u20131741."},{"key":"e_1_3_1_17_2","first-page":"6626","volume-title":"Advances in Annual Conference on Neural Information Processing Systems","volume":"30","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in Annual Conference on Neural Information Processing Systems, Vol. 30. 6626\u20136637."},{"issue":"4","key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"1287","DOI":"10.1007\/s12559-022-10038-y","article-title":"FF-UNet: A u-shaped deep convolutional neural network for multimodal biomedical image segmentation","volume":"14","author":"Iqbal Ahmed","year":"2022","unstructured":"Ahmed Iqbal, Muhammad Sharif, Muhammad Attique Khan, Wasif Nisar, and Majed Alhaisoni. 2022. FF-UNet: A u-shaped deep convolutional neural network for multimodal biomedical image segmentation. Cognitive Computation 14, 4 (2022), 1287\u20131302.","journal-title":"Cognitive Computation"},{"key":"e_1_3_1_19_2","first-page":"694","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Johnson Justin","year":"2016","unstructured":"Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016. Perceptual losses for real-time style transfer and super-resolution. In Proceedings of the European Conference on Computer Vision. Springer, 694\u2013711."},{"key":"e_1_3_1_20_2","first-page":"14428","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Khrulkov Valentin","year":"2021","unstructured":"Valentin Khrulkov, Leyla Mirvakhabova, Ivan Oseledets, and Artem Babenko. 2021. Latent transformations via neural ODEs for GAN-based image editing. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 14428\u201314437."},{"key":"e_1_3_1_21_2","first-page":"2122","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Kim Kyunghun","year":"2021","unstructured":"Kyunghun Kim, Yeohun Yun, Keon-Woo Kang, Kyeongbo Kong, Siyeong Lee, and Suk-Ju Kang. 2021. Painting outside as inside: Edge guided image outpainting via bidirectional rearrangement with progressive step learning. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 2122\u20132130."},{"key":"e_1_3_1_22_2","article-title":"Adam: A method for stochastic optimization","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014).","journal-title":"arXiv preprint arXiv:1412.6980"},{"key":"e_1_3_1_23_2","first-page":"1558","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Larsen Anders Boesen Lindbo","year":"2016","unstructured":"Anders Boesen Lindbo Larsen, S\u00f8ren Kaae S\u00f8nderby, Hugo Larochelle, and Ole Winther. 2016. Autoencoding beyond pixels using a learned similarity metric. In Proceedings of the International Conference on Machine Learning. PMLR, 1558\u20131566."},{"key":"e_1_3_1_24_2","first-page":"1833","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Liang Jingyun","year":"2021","unstructured":"Jingyun Liang, Jiezhang Cao, Guolei Sun, Kai Zhang, Luc Van Gool, and Radu Timofte. 2021. SwinIR: Image restoration using swin transformer. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 1833\u20131844."},{"key":"e_1_3_1_25_2","article-title":"Geometric GAN","author":"Lim Jae Hyun","year":"2017","unstructured":"Jae Hyun Lim and Jong Chul Ye. 2017. Geometric GAN. arXiv preprint arXiv:1705.02894 (2017).","journal-title":"arXiv preprint arXiv:1705.02894"},{"key":"e_1_3_1_26_2","first-page":"1","article-title":"DS-TransUNet: Dual Swin Transformer U-Net for medical image segmentation","volume":"71","author":"Lin Ailiang","year":"2022","unstructured":"Ailiang Lin, Bingzhi Chen, Jiayu Xu, Zheng Zhang, Guangming Lu, and David Zhang. 2022. DS-TransUNet: Dual Swin Transformer U-Net for medical image segmentation. IEEE Transactions on Instrumentation and Measurement 71 (2022), 1\u201315.","journal-title":"IEEE Transactions on Instrumentation and Measurement"},{"key":"e_1_3_1_27_2","first-page":"806","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lin Han","year":"2021","unstructured":"Han Lin, Maurice Pagnucco, and Yang Song. 2021. Edge guided progressively generative image outpainting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 806\u2013815."},{"key":"e_1_3_1_28_2","article-title":"Swin Transformer: Hierarchical vision transformer using shifted windows","author":"Liu Ze","year":"2021","unstructured":"Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin Transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030 (2021).","journal-title":"arXiv preprint arXiv:2103.14030"},{"key":"e_1_3_1_29_2","first-page":"5775","article-title":"DPM-Solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps","volume":"35","author":"Lu Cheng","year":"2022","unstructured":"Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. 2022. DPM-Solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps. Advances in Annual Conference on Neural Information Processing Systems 35, 5775\u20135787.","journal-title":"Advances in Annual Conference on Neural Information Processing Systems"},{"key":"e_1_3_1_30_2","first-page":"843","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lu Chia-Ni","year":"2021","unstructured":"Chia-Ni Lu, Ya-Chu Chang, and Wei-Chen Chiu. 2021. Bridging the visual gap: Wide-range image blending. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 843\u2013851."},{"key":"e_1_3_1_31_2","article-title":"Boosting image outpainting with semantic layout prediction","author":"Ma Ye","year":"2021","unstructured":"Ye Ma, Jin Ma, Min Zhou, Quan Chen, Tiezheng Ge, Yuning Jiang, and Tong Lin. 2021. Boosting image outpainting with semantic layout prediction. arXiv preprint arXiv:2110.09267 (2021).","journal-title":"arXiv preprint arXiv:2110.09267"},{"issue":"4","key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3503927","article-title":"Scenario-aware recurrent transformer for goal-directed video captioning","volume":"18","author":"Man Xin","year":"2022","unstructured":"Xin Man, Deqiang Ouyang, Xiangpeng Li, Jingkuan Song, and Jie Shao. 2022. Scenario-aware recurrent transformer for goal-directed video captioning. ACM Transactions on Multimedia Computing, Communications, and Applications 18, 4 (2022), 1\u201317.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_33_2","first-page":"2794","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Mao Xudong","year":"2017","unstructured":"Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley. 2017. Least squares generative adversarial networks. In Proceedings of the IEEE International Conference on Computer Vision. 2794\u20132802."},{"key":"e_1_3_1_34_2","first-page":"2536","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Pathak Deepak","year":"2016","unstructured":"Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros. 2016. Context encoders: Feature learning by inpainting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2536\u20132544."},{"key":"e_1_3_1_35_2","doi-asserted-by":"crossref","first-page":"107453","DOI":"10.1016\/j.patcog.2020.107453","article-title":"Generative adversarial classifier for handwriting characters super-resolution","volume":"107","author":"Qian Zhuang","year":"2020","unstructured":"Zhuang Qian, Kaizhu Huang, Qiu-Feng Wang, Jimin Xiao, and Rui Zhang. 2020. Generative adversarial classifier for handwriting characters super-resolution. Pattern Recognition 107 (2020), 107453.","journal-title":"Pattern Recognition"},{"key":"e_1_3_1_36_2","first-page":"10684","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Rombach Robin","year":"2022","unstructured":"Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj\u00f6rn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10684\u201310695."},{"key":"e_1_3_1_37_2","article-title":"Painting outside the box: Image outpainting with GANs","author":"Sabini Mark","year":"2018","unstructured":"Mark Sabini and Gili Rusak. 2018. Painting outside the box: Image outpainting with GANs. arXiv preprint arXiv:1808.08483 (2018).","journal-title":"arXiv preprint arXiv:1808.08483"},{"key":"e_1_3_1_38_2","first-page":"2234","volume-title":"Advances in Annual Conference on Neural Information Processing Systems","volume":"29","author":"Salimans Tim","year":"2016","unstructured":"Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved techniques for training GANs. In Advances in Annual Conference on Neural Information Processing Systems, Vol. 29. 2234\u20132242."},{"key":"e_1_3_1_39_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_40_2","first-page":"3703","volume-title":"Proceedings of IEEE International Conference on Image Processing","author":"Tan Wei Ren","year":"2016","unstructured":"Wei Ren Tan, Chee Seng Chan, Hern\u00e1n E. Aguirre, and Kiyoshi Tanaka. 2016. Ceci n\u2019est pas une pipe: A deep convolutional network for fine-art paintings classification. In Proceedings of IEEE International Conference on Image Processing. IEEE, 3703\u20133707."},{"key":"e_1_3_1_41_2","article-title":"Neural ODEs for image segmentation with level sets","author":"Valle Rafael","year":"2019","unstructured":"Rafael Valle, Fitsum Reda, Mohammad Shoeybi, Patrick Legresley, Andrew Tao, and Bryan Catanzaro. 2019. Neural ODEs for image segmentation with level sets. arXiv preprint arXiv:1912.11683 (2019).","journal-title":"arXiv preprint arXiv:1912.11683"},{"key":"e_1_3_1_42_2","article-title":"Image outpainting and harmonization using generative adversarial networks","author":"Hoorick Basile Van","year":"2019","unstructured":"Basile Van Hoorick. 2019. Image outpainting and harmonization using generative adversarial networks. arXiv preprint arXiv:1912.10960 (2019).","journal-title":"arXiv preprint arXiv:1912.10960"},{"key":"e_1_3_1_43_2","first-page":"8798","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Ting-Chun","year":"2018","unstructured":"Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. 2018. High-resolution image synthesis and semantic manipulation with conditional GANs. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8798\u20138807."},{"key":"e_1_3_1_44_2","first-page":"1399","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Yi","year":"2019","unstructured":"Yi Wang, Xin Tao, Xiaoyong Shen, and Jiaya Jia. 2019. Wide-context semantic image extrapolation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1399\u20131408."},{"issue":"4","key":"e_1_3_1_45_2","first-page":"1308","article-title":"E2I: Generative inpainting from edge to image","volume":"31","author":"Xu Shunxin","year":"2020","unstructured":"Shunxin Xu, Dong Liu, and Zhiwei Xiong. 2020. E2I: Generative inpainting from edge to image. IEEE Transactions on Circuits and Systems for Video Technology 31, 4 (2020), 1308\u20131322.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_46_2","first-page":"6721","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Chao","year":"2017","unstructured":"Chao Yang, Xin Lu, Zhe Lin, Eli Shechtman, Oliver Wang, and Hao Li. 2017. High-resolution image inpainting using multi-scale neural patch synthesis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6721\u20136729."},{"key":"e_1_3_1_47_2","first-page":"10561","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yang Zongxin","year":"2019","unstructured":"Zongxin Yang, Jian Dong, Ping Liu, Yi Yang, and Shuicheng Yan. 2019. Very long natural scenery image prediction by outpainting. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 10561\u201310570."},{"key":"e_1_3_1_48_2","first-page":"153","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Yao Kai","year":"2022","unstructured":"Kai Yao, Penglei Gao, Xi Yang, Jie Sun, Rui Zhang, and Kaizhu Huang. 2022. Outpainting by queries. In Proceedings of the European Conference on Computer Vision. 153\u2013169."},{"key":"e_1_3_1_49_2","first-page":"5505","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yu Jiahui","year":"2018","unstructured":"Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S. Huang. 2018. Generative image inpainting with contextual attention. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5505\u20135514."},{"issue":"4","key":"e_1_3_1_50_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3578518","article-title":"Graph attention transformer network for multi-label image classification","volume":"19","author":"Yuan Jin","year":"2023","unstructured":"Jin Yuan, Shikai Chen, Yao Zhang, Zhongchao Shi, Xin Geng, Jianping Fan, and Yong Rui. 2023. Graph attention transformer network for multi-label image classification. ACM Transactions on Multimedia Computing, Communications and Applications 19, 4 (2023), 1\u201316.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_1_51_2","first-page":"558","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yuan Li","year":"2021","unstructured":"Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis E. H. Tay, Jiashi Feng, and Shuicheng Yan. 2021. Tokens-to-token ViT: Training vision transformers from scratch on ImageNet. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 558\u2013567."},{"issue":"1","key":"e_1_3_1_52_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3491225","article-title":"JoT-GAN: A framework for jointly training GAN and person re-identification model","volume":"18","author":"Zhao Zhongwei","year":"2022","unstructured":"Zhongwei Zhao, Ran Song, Qian Zhang, Peng Duan, and Youmei Zhang. 2022. JoT-GAN: A framework for jointly training GAN and person re-identification model. ACM Transactions on Multimedia Computing, Communications, and Applications 18, 1s (2022), 1\u201318.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_53_2","article-title":"DeepViT: Towards deeper vision transformer","volume":"2103","author":"Zhou Daquan","year":"2021","unstructured":"Daquan Zhou, Bingyi Kang, Xiaojie Jin, Linjie Yang, Xiaochen Lian, Qibin Hou, and Jiashi Feng. 2021. DeepViT: Towards deeper vision transformer. CoRR abs\/2103.11886 (2021).","journal-title":"CoRR"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3648367","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3648367","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:04:13Z","timestamp":1750291453000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3648367"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,25]]},"references-count":52,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,7,31]]}},"alternative-id":["10.1145\/3648367"],"URL":"https:\/\/doi.org\/10.1145\/3648367","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2024,4,25]]},"assertion":[{"value":"2023-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-05","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}