{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T06:10:07Z","timestamp":1784268607871,"version":"3.55.0"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2019,2,14]],"date-time":"2019-02-14T00:00:00Z","timestamp":1550102400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"973 Program of China","award":["2015CB352502"],"award-info":[{"award-number":["2015CB352502"]}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["61332015, 61772318, 61572507, 61622212, 61532003"],"award-info":[{"award-number":["61332015, 61772318, 61572507, 61622212, 61532003"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100000038","name":"NSERC","doi-asserted-by":"crossref","award":["611370"],"award-info":[{"award-number":["611370"]}],"id":[{"id":"10.13039\/501100000038","id-type":"DOI","asserted-by":"crossref"}]},{"name":"ISF","award":["2366\/16"],"award-info":[{"award-number":["2366\/16"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2019,4,30]]},"abstract":"<jats:p>\n            We present a\n            <jats:italic>generative neural network<\/jats:italic>\n            that enables us to generate plausible 3D indoor scenes in large quantities and varieties, easily and highly efficiently. Our key observation is that indoor scene structures are inherently\n            <jats:italic>hierarchical<\/jats:italic>\n            . Hence, our network is not convolutional; it is a\n            <jats:italic>recursive<\/jats:italic>\n            neural network, or RvNN. Using a dataset of annotated scene hierarchies, we train a\n            <jats:italic>variational recursive autoencoder<\/jats:italic>\n            , or RvNN-VAE, which performs scene object grouping during its encoding phase and scene generation during decoding. Specifically, a set of encoders are recursively applied to group 3D objects based on support, surround, and co-occurrence relations in a scene, encoding information about objects\u2019 spatial properties,\n            <jats:italic>semantics<\/jats:italic>\n            , and\n            <jats:italic>relative<\/jats:italic>\n            positioning with respect to other objects in the hierarchy. By training a variational autoencoder (VAE), the resulting fixed-length codes roughly follow a Gaussian distribution. A novel 3D scene can be generated hierarchically by the decoder from a randomly sampled code from the learned distribution. We coin our method GRAINS, for Generative Recursive Autoencoders for INdoor Scenes. We demonstrate the capability of GRAINS to generate plausible and diverse 3D indoor scenes and compare with existing methods for 3D scene synthesis. We show applications of GRAINS including 3D scene modeling from 2D layouts, scene editing, and semantic scene segmentation via PointNet whose performance is boosted by the large quantity and variety of 3D scenes generated by our method.\n          <\/jats:p>","DOI":"10.1145\/3303766","type":"journal-article","created":{"date-parts":[[2019,2,14]],"date-time":"2019-02-14T19:36:17Z","timestamp":1550172977000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":150,"title":["GRAINS"],"prefix":"10.1145","volume":"38","author":[{"given":"Manyi","family":"Li","sequence":"first","affiliation":[{"name":"Shandong University and Simon Fraser University, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Akshay Gadi","family":"Patil","sequence":"additional","affiliation":[{"name":"Simon Fraser University, Vancouver, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kai","family":"Xu","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Siddhartha","family":"Chaudhuri","sequence":"additional","affiliation":[{"name":"Adobe Research and IIT Bombay, Mumbai, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Owais","family":"Khan","sequence":"additional","affiliation":[{"name":"IIT Bombay, Mumbai, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ariel","family":"Shamir","sequence":"additional","affiliation":[{"name":"The Interdisciplinary Center, Herzliya, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Changhe","family":"Tu","sequence":"additional","affiliation":[{"name":"Shandong University, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Baoquan","family":"Chen","sequence":"additional","affiliation":[{"name":"Peking University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel","family":"Cohen-Or","sequence":"additional","affiliation":[{"name":"Tel Aviv University, Tel Aviv, Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hao","family":"Zhang","sequence":"additional","affiliation":[{"name":"Simon Fraser University, Vancouver, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,2,14]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1833349.1778841"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964930"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661229.2661239"},{"key":"e_1_2_1_4_1","volume-title":"Adams","author":"Duvenaud David","year":"2015","unstructured":"David Duvenaud , Dougal Maclaurin , Jorge Aguilera-Iparraguirre , Rafael G\u00f3mez-Bombarelli , Timothy Hirzel , Al\u2019an Aspuru-Guzik , and Ryan P . Adams . 2015 . Convolutional networks on graphs for learning molecular fingerprints. In Neural Information Processing Systems (NIPS) . David Duvenaud, Dougal Maclaurin, Jorge Aguilera-Iparraguirre, Rafael G\u00f3mez-Bombarelli, Timothy Hirzel, Al\u2019an Aspuru-Guzik, and Ryan P. Adams. 2015. Convolutional networks on graphs for learning molecular fingerprints. In Neural Information Processing Systems (NIPS)."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2601097.2601185"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818057"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366154"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964929"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3130800.3130805"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46466-4_29"},{"key":"e_1_2_1_11_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems. 2672--2680.   Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems. 2672--2680."},{"key":"e_1_2_1_12_1","volume-title":"Deep convolutional networks on graph-structured data. CoRR abs\/1506.05163","author":"Henaff Mikael","year":"2015","unstructured":"Mikael Henaff , Joan Bruna , and Yann LeCun . 2015. Deep convolutional networks on graph-structured data. CoRR abs\/1506.05163 ( 2015 ). Retrieved from http:\/\/arxiv.org\/abs\/1506.05163. Mikael Henaff, Joan Bruna, and Yann LeCun. 2015. Deep convolutional networks on graph-structured data. CoRR abs\/1506.05163 (2015). Retrieved from http:\/\/arxiv.org\/abs\/1506.05163."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.12694"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185551"},{"key":"e_1_2_1_15_1","volume-title":"Computer Graphics Forum","volume":"35","author":"Kermani Z. Sadeghipour","unstructured":"Z. Sadeghipour Kermani , Zicheng Liao , Ping Tan , and H. Zhang . 2016. Learning 3D scene synthesis from annotated RGB-D images . In Computer Graphics Forum , Vol. 35 . Wiley Online Library, 197--206. Z. Sadeghipour Kermani, Zicheng Liao, Ping Tan, and H. Zhang. 2016. Learning 3D scene synthesis from annotated RGB-D images. In Computer Graphics Forum, Vol. 35. Wiley Online Library, 197--206."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2461933"},{"key":"e_1_2_1_17_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba . 2014 . Adam : A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014). Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_2_1_18_1","volume-title":"Kingma and Max Welling","author":"Diederik","year":"2013","unstructured":"Diederik P. Kingma and Max Welling . 2013 . Auto-encoding variational Bayes . arXiv preprint arXiv:1312.6114 (2013). Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational Bayes. arXiv preprint arXiv:1312.6114 (2013)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073637"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275035"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2980179.2980223"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964982"},{"key":"e_1_2_1_23_1","volume-title":"Perceptrons: An Introduction to Computational Geometry","author":"Minsky Marvin","year":"1969","unstructured":"Marvin Minsky and Seymour Papert . 1969 . Perceptrons: An Introduction to Computational Geometry . MIT Press . Marvin Minsky and Seymour Papert. 1969. Perceptrons: An Introduction to Computational Geometry. MIT Press."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1141911.1141931"},{"key":"e_1_2_1_25_1","volume-title":"International Conference on Machine Learning (ICML).","author":"Niepert Mathias","year":"2016","unstructured":"Mathias Niepert , Mohamed Ahmed , and Konstantin Kutzkov . 2016 . Learning convolutional neural networks for graphs . In International Conference on Machine Learning (ICML). Mathias Niepert, Mohamed Ahmed, and Konstantin Kutzkov. 2016. Learning convolutional neural networks for graphs. In International Conference on Machine Learning (ICML)."},{"key":"e_1_2_1_26_1","unstructured":"Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in Pytorch. In Neural Information Processing Systems-Workshop (NIPS-W).  Adam Paszke Sam Gross Soumith Chintala Gregory Chanan Edward Yang Zachary DeVito Zeming Lin Alban Desmaison Luca Antiga and Adam Lerer. 2017. Automatic differentiation in Pytorch. In Neural Information Processing Systems-Workshop (NIPS-W)."},{"key":"e_1_2_1_27_1","volume-title":"Proc. of IEEE Conference on Computer Vision and Pattern Recognition. 652--660","author":"Qi Charles R.","unstructured":"Charles R. Qi , Hao Su , Kaichun Mo , and Leonidas J. Guibas . 2017. PointNet: Deep learning on point sets for 3D classification and segmentation . In Proc. of IEEE Conference on Computer Vision and Pattern Recognition. 652--660 . Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. 2017. PointNet: Deep learning on point sets for 3D classification and segmentation. In Proc. of IEEE Conference on Computer Vision and Pattern Recognition. 652--660."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00618"},{"key":"e_1_2_1_29_1","volume-title":"Ng","author":"Socher Richard","year":"2012","unstructured":"Richard Socher , Brody Huval , Bharath Bhat , Christopher D. Manning , and Andrew Y . Ng . 2012 . Convolutional-recursive deep learning for 3D object classification. In Neural Information Processing Systems (NIPS) . Richard Socher, Brody Huval, Bharath Bhat, Christopher D. Manning, and Andrew Y. Ng. 2012. Convolutional-recursive deep learning for 3D object classification. In Neural Information Processing Systems (NIPS)."},{"key":"e_1_2_1_30_1","volume-title":"International Conference on Machine Learning (ICML).","author":"Socher Richard","unstructured":"Richard Socher , Cliff C. Lin , Andrew Y. Ng , and Christopher D. Manning . 2011. Parsing natural scenes and natural language with recursive neural networks . In International Conference on Machine Learning (ICML). Richard Socher, Cliff C. Lin, Andrew Y. Ng, and Christopher D. Manning. 2011. Parsing natural scenes and natural language with recursive neural networks. In International Conference on Machine Learning (ICML)."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.28"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2380116.2380127"},{"key":"e_1_2_1_33_1","volume-title":"WaveNet: A generative model for raw audio. CoRR abs\/1609.03499","author":"van den Oord A\u00e4ron","year":"2016","unstructured":"A\u00e4ron van den Oord , Sander Dieleman , Heiga Zen , Karen Simonyan , Oriol Vinyals , Alex Graves , Nal Kalchbrenner , Andrew W. Senior , and Koray Kavukcuoglu . 2016a. WaveNet: A generative model for raw audio. CoRR abs\/1609.03499 ( 2016 ). Retrieved from http:\/\/arxiv.org\/abs\/1609.03499. A\u00e4ron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016a. WaveNet: A generative model for raw audio. CoRR abs\/1609.03499 (2016). Retrieved from http:\/\/arxiv.org\/abs\/1609.03499."},{"key":"e_1_2_1_34_1","volume-title":"Pixel recurrent neural networks. CoRR abs\/1601.06759","author":"van den Oord A\u00e4ron","year":"2016","unstructured":"A\u00e4ron van den Oord , Nal Kalchbrenner , and Koray Kavukcuoglu . 2016b. Pixel recurrent neural networks. CoRR abs\/1601.06759 ( 2016 ). Retrieved from http:\/\/arxiv.org\/abs\/1601.06759. A\u00e4ron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. 2016b. Pixel recurrent neural networks. CoRR abs\/1601.06759 (2016). Retrieved from http:\/\/arxiv.org\/abs\/1601.06759."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201362"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2011.01885.x"},{"key":"e_1_2_1_37_1","unstructured":"Paul J. Werbos. 1974. Beyond Regression: New Tools for Predicting and Analysis in the Behavioral Sciences. Ph.D. Dissertation. Harvard University.  Paul J. Werbos. 1974. Beyond Regression: New Tools for Predicting and Analysis in the Behavioral Sciences. Ph.D. Dissertation. Harvard University."},{"key":"e_1_2_1_38_1","volume-title":"Tenenbaum","author":"Wu Jiajun","year":"2016","unstructured":"Jiajun Wu , Chengkai Zhang , Tianfan Xue , William T. Freeman , and Joshua B . Tenenbaum . 2016 . Learning a probabilistic latent space of object shapes via 3D generative-adversarial modeling. In Neural Information Processing Systems (NIPS) . Jiajun Wu, Chengkai Zhang, Tianfan Xue, William T. Freeman, and Joshua B. Tenenbaum. 2016. Learning a probabilistic latent space of object shapes via 3D generative-adversarial modeling. In Neural Information Processing Systems (NIPS)."},{"key":"e_1_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Zhirong Wu Shuran Song Aditya Khosla Fisher Yu Linguang Zhang Xiaoou Tang and Jianxiong Xiao. 2015. 3D ShapeNets: A deep representation for volumetric shapes. In Computer Vision and Pattern Recognition (CVPR).  Zhirong Wu Shuran Song Aditya Khosla Fisher Yu Linguang Zhang Xiaoou Tang and Jianxiong Xiao. 2015. 3D ShapeNets: A deep representation for volumetric shapes. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2015.7298801"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2461968"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185553"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964981"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3303766","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3303766","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:53:39Z","timestamp":1750204419000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3303766"}},"subtitle":["Generative Recursive Autoencoders for INdoor Scenes"],"short-title":[],"issued":{"date-parts":[[2019,2,14]]},"references-count":42,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2019,4,30]]}},"alternative-id":["10.1145\/3303766"],"URL":"https:\/\/doi.org\/10.1145\/3303766","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,2,14]]},"assertion":[{"value":"2018-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-02-14","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}