{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:21:27Z","timestamp":1750220487038,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":34,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,17]]},"DOI":"10.1145\/3474085.3475541","type":"proceedings-article","created":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T04:52:26Z","timestamp":1634532746000},"page":"4100-4108","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["CanvasEmb: Learning Layout Representation with Large-scale Pre-training for Graphic Design"],"prefix":"10.1145","author":[{"given":"Yuxi","family":"Xie","sequence":"first","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Danqing","family":"Huang","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinpeng","family":"Wang","sequence":"additional","affiliation":[{"name":"Meituan, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chin-Yew","family":"Lin","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/DIAL.2006.16"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2034691.2034695"},{"key":"e_1_3_2_1_3_1","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171--4186","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171--4186 . Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171--4186."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2011.175"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.139"},{"key":"e_1_3_2_1_6_1","volume-title":"Deep Residual Learning for Image Recognition. IEEE Conference on Computer Vision and Pattern Recognition(2016)","author":"He Kaiming","year":"2016","unstructured":"Kaiming He , X. Zhang , Shaoqing Ren , and Jian Sun . 2016 . Deep Residual Learning for Image Recognition. IEEE Conference on Computer Vision and Pattern Recognition(2016) , 770--778. Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. IEEE Conference on Computer Vision and Pattern Recognition(2016), 770--778."},{"key":"e_1_3_2_1_7_1","volume-title":"International Conference on Computer Vision.","author":"Jyothi Akash Abdu","year":"2019","unstructured":"Akash Abdu Jyothi , Thibaut Durand , Jiawei He , Leonid Sigal , and Greg Mori . 2019 . Layout VAE: Stochastic Scene Layout Generation from a Label Set . In International Conference on Computer Vision. Akash Abdu Jyothi, Thibaut Durand, Jiawei He, Leonid Sigal, and Greg Mori. 2019. Layout VAE: Stochastic Scene Layout Generation from a Label Set. In International Conference on Computer Vision."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294996.3295074"},{"volume-title":"Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.).","author":"Diederik","key":"e_1_3_2_1_9_1","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015 . Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.). Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58580-8_29"},{"key":"e_1_3_2_1_11_1","unstructured":"Jianan Li Jimei Yang Aaron Hertzmann Jianming Zhang and Tingfa Xu. 2019. Layout GAN: Generating Graphic Layouts with Wire frame Discriminators.  Jianan Li Jimei Yang Aaron Hertzmann Jianming Zhang and Tingfa Xu. 2019. Layout GAN: Generating Graphic Layouts with Wire frame Discriminators."},{"key":"e_1_3_2_1_12_1","volume-title":"Layout GAN: Generating Graphic Layouts with Wireframe Discriminators. In 7th International Conference on Learning Representations.","author":"Li Jianan","year":"2019","unstructured":"Jianan Li , Jimei Yang , Aaron Hertzmann , Jianming Zhang , and Tingfa Xu . 2019 . Layout GAN: Generating Graphic Layouts with Wireframe Discriminators. In 7th International Conference on Learning Representations. Jianan Li, Jimei Yang, Aaron Hertzmann, Jianming Zhang, and Tingfa Xu. 2019. Layout GAN: Generating Graphic Layouts with Wireframe Discriminators. In 7th International Conference on Learning Representations."},{"key":"e_1_3_2_1_13_1","unstructured":"Yang Li J. Amelot X. Zhou S. Bengio and S. Si. 2020. Auto Completion of User Interface Layout Design Using Transformer-Based Tree Decoders. abs\/2001.05308(2020).  Yang Li J. Amelot X. Zhou S. Bengio and S. Si. 2020. Auto Completion of User Interface Layout Design Using Transformer-Based Tree Decoders. abs\/2001.05308(2020)."},{"key":"e_1_3_2_1_14_1","volume-title":"Focal Loss for Dense Object Detection. 2017 IEEE International Conference on Computer Vision(2017)","author":"Lin Tsung-Yi","year":"2017","unstructured":"Tsung-Yi Lin , Priya Goyal , Ross Girshick , Kaiming He , and Piotr Dollar . 2017 . Focal Loss for Dense Object Detection. 2017 IEEE International Conference on Computer Vision(2017) . Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. 2017. Focal Loss for Dense Object Detection. 2017 IEEE International Conference on Computer Vision(2017)."},{"volume-title":"Evaluation of Visual Balance for Automated Layout","author":"Lok Simon","key":"e_1_3_2_1_15_1","unstructured":"Simon Lok , Steven Feiner , and Gary Ngai . 2004. Evaluation of Visual Balance for Automated Layout . Association for Computing Machinery , 101--108. Simon Lok, Steven Feiner, and Gary Ngai. 2004. Evaluation of Visual Balance for Automated Layout. Association for Computing Machinery, 101--108."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454289"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.5555\/3042817.3043083"},{"key":"e_1_3_2_1_18_1","volume-title":"READ Recursive Autoencoders for Document Layout Generation.2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops(2020)","author":"Patil A.","year":"2020","unstructured":"A. Patil , Omri Ben-Eliezer , Or Perel , and Hadar Averbuch-Elor . 2020 . READ Recursive Autoencoders for Document Layout Generation.2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops(2020) , 2316--2325. A. Patil, Omri Ben-Eliezer, Or Perel, and Hadar Averbuch-Elor. 2020. READ Recursive Autoencoders for Document Layout Generation.2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops(2020), 2316--2325."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1202"},{"key":"e_1_3_2_1_20_1","unstructured":"A. Radford. 2018. Improving Language Understanding by Generative Pre-Training.  A. Radford. 2018. Improving Language Understanding by Generative Pre-Training."},{"key":"e_1_3_2_1_21_1","unstructured":"Justus J. Randolph. 2005. Free-marginal multirater kappa (multirater k [free]): An alternative to fleiss' fixed-marginal multirater kapp. (2005).  Justus J. Randolph. 2005. Free-marginal multirater kappa (multirater k [free]): An alternative to fleiss' fixed-marginal multirater kapp. (2005)."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969239.2969250"},{"key":"e_1_3_2_1_23_1","volume-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations.","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman . 2015 . Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations. Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1348"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1774088.1774091"},{"key":"e_1_3_2_1_26_1","volume-title":"VideoBERT: A Joint Model for Video and Language Representation Learning.2019 IEEE\/CVF International Conference on Computer Vision(2019)","author":"Sun Chen","year":"2019","unstructured":"Chen Sun , Austin Myers , Carl Vondrick , Kevin Murphy , and Cordelia Schmid . 2019 . VideoBERT: A Joint Model for Video and Language Representation Learning.2019 IEEE\/CVF International Conference on Computer Vision(2019) . Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019. VideoBERT: A Joint Model for Video and Language Representation Learning.2019 IEEE\/CVF International Conference on Computer Vision(2019)."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306214.3338574"},{"volume-title":"Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing.","author":"Tai Kai Sheng","key":"e_1_3_2_1_28_1","unstructured":"Kai Sheng Tai , Richard Socher , and Christopher D. Manning . 2015. Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks . In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing. Kai Sheng Tai, Richard Socher, and Christopher D. Manning. 2015. Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1018"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.383"},{"key":"e_1_3_2_1_32_1","unstructured":"Yiheng Xu Minghao Li Lei Cui Shaohan Huang Furu Wei and Ming Zhou. 2019. LayoutLM: Pre-training of Text and Layout for Document Image Understanding.  Yiheng Xu Minghao Li Lei Cui Shaohan Huang Furu Wei and Ming Zhou. 2019. LayoutLM: Pre-training of Text and Layout for Document Image Understanding."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454804"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3322971"}],"event":{"name":"MM '21: ACM Multimedia Conference","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Virtual Event China","acronym":"MM '21"},"container-title":["Proceedings of the 29th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475541","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474085.3475541","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:49:10Z","timestamp":1750193350000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475541"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":34,"alternative-id":["10.1145\/3474085.3475541","10.1145\/3474085"],"URL":"https:\/\/doi.org\/10.1145\/3474085.3475541","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}