{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T11:00:28Z","timestamp":1783335628120,"version":"3.54.6"},"reference-count":121,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2024,12,12]],"date-time":"2024-12-12T00:00:00Z","timestamp":1733961600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,12,12]],"date-time":"2024-12-12T00:00:00Z","timestamp":1733961600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2025,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>This position paper argues for the use of <jats:italic>structured generative models<\/jats:italic> (SGMs) for the understanding of static scenes. This requires the reconstruction of a 3D scene from an input image (or a set of multi-view images), whereby the contents of the image(s) are causally explained in terms of models of instantiated objects, each with their own type, shape, appearance and pose, along with global variables like scene lighting and camera parameters. This approach also requires scene models which account for the co-occurrences and inter-relationships of objects in a scene. The SGM approach has the merits that it is compositional and generative, which lead to interpretability and editability. To pursue the SGM agenda, we need models for objects and scenes, and approaches to carry out inference. We first review models for objects, which include \u201cthings\u201d (object categories that have a well defined shape), and \u201cstuff\u201d (categories which have amorphous spatial extent). We then move on to review <jats:italic>scene models<\/jats:italic> which describe the inter-relationships of objects. Perhaps the most challenging problem for SGMs is <jats:italic>inference<\/jats:italic> of the objects, lighting and camera parameters, and scene inter-relationships from input consisting of a single or multiple images. We conclude with a discussion of issues that need addressing to advance the SGM agenda.<\/jats:p>","DOI":"10.1007\/s11263-024-02316-z","type":"journal-article","created":{"date-parts":[[2024,12,12]],"date-time":"2024-12-12T13:37:23Z","timestamp":1734010643000},"page":"2845-2867","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Structured Generative Models for Scene Understanding"],"prefix":"10.1007","volume":"133","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6270-4703","authenticated-orcid":false,"given":"Christopher K. I.","family":"Williams","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,12,12]]},"reference":[{"key":"2316_CR1","doi-asserted-by":"crossref","unstructured":"Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C., & Parikh, D. (2015). VQA: Visual Question Answering. In 2015 IEEE international conference on computer vision (ICCV) (pp. 2425\u20132433).","DOI":"10.1109\/ICCV.2015.279"},{"key":"2316_CR2","doi-asserted-by":"crossref","unstructured":"Babalola, K. O., Cootes, T. F., Twining, C. J., Petrovic, V., & Taylor, C. (2008). 3D brain segmentation using active appearance models and local regressors. In D.\u00a0Metaxas, L.\u00a0Axel, G.\u00a0Fichtinger, & G.\u00a0Sz\u00e9kely (Eds.), International conference on medical image computing and computer-assisted intervention (MICCAI 2008) (Vol. 5241, pp. 401\u2013408). Springer. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-540-85988-8_48"},{"key":"2316_CR3","doi-asserted-by":"crossref","unstructured":"Bengio, Y., Louradour, J., Collobert, R., & Weston, J. (2009). Curriculum learning. In L.\u00a0Bottou & M.\u00a0Littman (Eds.), Proceedings of the twenty-sixth international conference on machine learning (ICML 2009) (pp. 41\u201348).","DOI":"10.1145\/1553374.1553380"},{"key":"2316_CR4","doi-asserted-by":"crossref","DOI":"10.1002\/9780470316870","volume-title":"Bayesian theory","author":"JM Bernardo","year":"1994","unstructured":"Bernardo, J. M., & Smith, A. F. M. (1994). Bayesian theory. Chichester: Wiley."},{"key":"2316_CR5","doi-asserted-by":"crossref","first-page":"115","DOI":"10.1037\/0033-295X.94.2.115","volume":"94","author":"I Biederman","year":"1987","unstructured":"Biederman, I. (1987). Recognition-by-components: A theory of human image understanding. Psychol. Rev., 94, 115\u2013147.","journal-title":"Psychol. Rev."},{"key":"2316_CR6","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1016\/0010-0285(82)90007-X","volume":"14","author":"I Biederman","year":"1982","unstructured":"Biederman, I., Mezzanotte, R. J., & Rabinowitz, J. C. (1982). Scene perception: Detecting and judging objects undergoing relational violations. Cogn. Psychol., 14, 143\u2013177.","journal-title":"Cogn. Psychol."},{"key":"2316_CR7","doi-asserted-by":"crossref","unstructured":"Blanz, V., & Vetter, T. (1999). A morphable model for the synthesis of 3D faces. In SIGGRAPH 99: Proceedings of the 26th annual conference on computer graphics and interactive techniques (pp.\u00a0187\u2013194). ACM Press.","DOI":"10.1145\/311535.311556"},{"key":"2316_CR8","volume-title":"Empirical model-building and response surfaces","author":"GEP Box","year":"1987","unstructured":"Box, G. E. P., & Draper, N. R. (1987). Empirical model-building and response surfaces. New York: Wiley."},{"key":"2316_CR9","volume-title":"Textures: A photographic album for artists and designers","author":"P Brodatz","year":"1966","unstructured":"Brodatz, P. (1966). Textures: A photographic album for artists and designers. New York: Dover Publications."},{"key":"2316_CR10","volume-title":"Statistical language learning","author":"E Charniak","year":"1993","unstructured":"Charniak, E. (1993). Statistical language learning. Cambridge: MIT Press."},{"key":"2316_CR11","doi-asserted-by":"crossref","unstructured":"Collins, J., Goael, S., Deng, K., & (2022). ABO: Dataset and benchmarks for real-world 3D object understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, CVPR 2022.","DOI":"10.1109\/CVPR52688.2022.02045"},{"key":"2316_CR12","doi-asserted-by":"crossref","unstructured":"Cootes, T.F., Edwards, G.J., & Taylor, C.J. (1998). Active appearance models. In H.\u00a0Burkhardt & B.\u00a0Neumann (Eds.), Proceedings of the Fifth European Conference on Computer Vision, ECCV 1998 (II, pp. 484\u2013498). Lecture Notes in Computer Science 1407. Springer.","DOI":"10.1007\/BFb0054760"},{"issue":"1","key":"2316_CR13","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1006\/cviu.1995.1004","volume":"61","author":"TF Cootes","year":"1995","unstructured":"Cootes, T. F., Taylor, C. J., Cooper, D. H., & Graham, J. (1995). Active shape models-their training and application. Computer Vision and Image Understanding, 61(1), 38\u201359.","journal-title":"Computer Vision and Image Understanding"},{"key":"2316_CR14","volume-title":"Elements of information theory","author":"TM Cover","year":"1991","unstructured":"Cover, T. M., & Thomas, J. A. (1991). Elements of information theory. New York: Wiley."},{"issue":"5","key":"2316_CR15","doi-asserted-by":"crossref","first-page":"889","DOI":"10.1162\/neco.1995.7.5.889","volume":"7","author":"P Dayan","year":"1995","unstructured":"Dayan, P., Hinton, G. E., Neal, R. M., & Zemel, R. S. (1995). The Helmholtz machine. Neural Computation, 7(5), 889\u2013904.","journal-title":"Neural Computation"},{"key":"2316_CR16","unstructured":"De\u00a0Sousa\u00a0Ribeiro, F., Duarte, K., Everett, M., Leontidis, G., & Shah, M. (2022). Learning with capsules: A survey. arXiv:2206.02664"},{"key":"2316_CR17","unstructured":"Du, Y., Durkan, C., Strudel, R., Tenenbaum, J.B., Dieleman, S., Fergus, R., & Grathwohl, W. (2023). Reduce, reuse, recycle: Compositional generation with energy-based diffusion models and MCMC. In Proceedings of the 40th international conference on machine learning, PMLR (p. 202)."},{"key":"2316_CR18","doi-asserted-by":"crossref","unstructured":"Efros, A. A., & Leung, T. K. (1999). Texture synthesis by non-parametric sampling. In Proceedings of the international conference on computer vision (ICCV 1999) (pp. 1033\u20131038).","DOI":"10.1109\/ICCV.1999.790383"},{"key":"2316_CR19","unstructured":"Eslami, S. M. A., Heess, N., Weber, T., Tassa, Y., Szepesvari, D., Kavukcuoglu, K., & Hinton, G. E. (2016). Attend, infer, repeat: Fast scene understanding with generative models. In D.\u00a0Lee, M.\u00a0Sugiyama, U.\u00a0Luxburg, I.\u00a0Guyon, & R.\u00a0Garnett (Eds.), Advances in neural information processing systems."},{"key":"2316_CR20","doi-asserted-by":"crossref","unstructured":"Eslami, S. M. A., & Williams, C. K. I. (2011). Factored shapes and appearances for parts-based object understanding. In Proceedings of the British machine vision conference 2011.","DOI":"10.5244\/C.25.18"},{"key":"2316_CR21","unstructured":"Eslami, S. M. A., & Williams, C. K. I. (2012). A generative model for parts-based object segmentation. In P.\u00a0Bartlett, F.\u00a0Pereira, C.\u00a0Burges, L.\u00a0Bottou, & K.\u00a0Weinberger (Eds.), Advances in neural information processing systems 25 (pp.\u00a0100\u2013107)."},{"issue":"2","key":"2316_CR22","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","volume":"88","author":"M Everingham","year":"2010","unstructured":"Everingham, M., Van Gool, L., Williams, C. K. I., Winn, J., & Zisserman, A. (2010). The PASCAL visual object classes (VOC) challenge. International Journal of Computer Vision, 88(2), 303\u2013338.","journal-title":"International Journal of Computer Vision"},{"issue":"9","key":"2316_CR23","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","volume":"32","author":"P Felzenszwalb","year":"2009","unstructured":"Felzenszwalb, P., Girshick, R., McAllester, D., & Ramanan, D. (2009). Object detection with discriminatively trained part based models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(9), 1627\u20131645.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"6","key":"2316_CR24","doi-asserted-by":"crossref","first-page":"381","DOI":"10.1145\/358669.358692","volume":"24","author":"MA Fischler","year":"1981","unstructured":"Fischler, M. A., & Bolles, R. C. (1981). Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6), 381\u2013395.","journal-title":"Communications of the ACM"},{"issue":"1","key":"2316_CR25","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1109\/T-C.1973.223602","volume":"22","author":"MA Fischler","year":"1973","unstructured":"Fischler, M. A., & Elschlager, R. A. (1973). The representation and matching of pictoral structures. IEEE Transactions on Computers, 22(1), 67\u201392.","journal-title":"IEEE Transactions on Computers"},{"key":"2316_CR26","volume-title":"Computer vision: A modern approach","author":"DA Forsyth","year":"2003","unstructured":"Forsyth, D. A., & Ponce, J. (2003). Computer vision: A modern approach (2nd ed.). Prentice Hall.","edition":"2"},{"issue":"1","key":"2316_CR27","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TPAMI.2003.1159942","volume":"25","author":"BJ Frey","year":"2003","unstructured":"Frey, B. J., & Jojic, N. (2003). Transformation invariant clustering using the EM algorithm. IEEE Transactions on Pattern Analysis and Machine Intelligence, 25(1), 1\u201317.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2316_CR28","volume-title":"Syntactic pattern recognition and applications","author":"KS Fu","year":"1982","unstructured":"Fu, K. S. (1982). Syntactic pattern recognition and applications. Prentice-Hall."},{"issue":"1","key":"2316_CR29","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1109\/TIP.2010.2052822","volume":"20","author":"B Galerne","year":"2011","unstructured":"Galerne, B., Gousseau, Y., & Morel, J.-M. (2011). Random phase textures: Theory and synthesis. IEEE Transactions on Image Processing, 20(1), 257\u2013267.","journal-title":"IEEE Transactions on Image Processing"},{"key":"2316_CR30","doi-asserted-by":"crossref","unstructured":"Gao, L., Sun, J.-M., Mo, K., Lai, Y.-K., Guibas, L. J., & Yang, J. (2023). SceneHGN: Hierarchical graph networks for 3D indoor scene generation with fine-grained geometry. arXiv:2302.10237","DOI":"10.1109\/TPAMI.2023.3237577"},{"key":"2316_CR31","unstructured":"Gemini Team (2023). Gemini: A family of highly capable multimodal models. Google. arXiv:2312.11805"},{"key":"2316_CR32","unstructured":"Greff, K., Kaufman, R. L., Kabra, R., Watters, N., Burgess, C., Zoran, D., & Lerchner, A. (2019). Multi-object representation learning with iterative variational inference. In ICML (97, pp. 2424\u20132433)."},{"key":"2316_CR33","volume-title":"Lectures in pattern theory: Volume 2 pattern analysis","author":"U Grenander","year":"1978","unstructured":"Grenander, U. (1978). Lectures in pattern theory: Volume 2 pattern analysis. Springer."},{"issue":"7","key":"2316_CR34","doi-asserted-by":"crossref","first-page":"444","DOI":"10.1145\/362280.362302","volume":"16","author":"PAV Hall","year":"1973","unstructured":"Hall, P. A. V. (1973). Equivalence between AND\/OR graphs and context-free grammars. Communications of the ACM, 16(7), 444\u2013445.","journal-title":"Communications of the ACM"},{"key":"2316_CR35","volume-title":"Computer vision systems","author":"AR Hanson","year":"1978","unstructured":"Hanson, A. R., & Riseman, E. M. (1978). VISIONS: A computer system for interpreting scenes. In A. R. Hanson & E. M. Riseman (Eds.), Computer vision systems. Academic Press."},{"key":"2316_CR36","doi-asserted-by":"crossref","unstructured":"Heitz, G., & Koller, D. (2008). Learning spatial context: Using stuff to find things. In Proceedings of the European conference on computer vision 2008.","DOI":"10.1007\/978-3-540-88682-2_4"},{"key":"2316_CR37","doi-asserted-by":"crossref","first-page":"1771","DOI":"10.1162\/089976602760128018","volume":"14","author":"GE Hinton","year":"2002","unstructured":"Hinton, G. E. (2002). Training products of experts by minimizing contrastive divergence. Neural Computation, 14, 1771\u20131800.","journal-title":"Neural Computation"},{"key":"2316_CR38","doi-asserted-by":"crossref","unstructured":"Hinton, G. E., Krizhevsky, A., & Wang, S. D. (2011). Transforming auto-encoders. In Proceedings ICANN 2011.","DOI":"10.1007\/978-3-642-21735-7_6"},{"key":"2316_CR39","unstructured":"Hinton, G. E., Sabour, S., & Frosst, N. (2018). Matrix capsules with EM routing. In International conference on learning representations."},{"key":"2316_CR40","unstructured":"Hu, E. J., Malkin, N., Jain, M., Everett, K., Graikos, A., & Bengio, Y. (2023). GFlowNet-EM for learning compositional latent variable models. arXiv:2302.06576"},{"key":"2316_CR41","unstructured":"Hueting, M., Reddy, P., Kim, V. G., Yumer, E., Carr, N., & Mitra, N. J. (2018). Seethrough: Finding objects in heavily occluded indoor scene images. In Proceedings of international conference on 3DVision (3DV)."},{"key":"2316_CR42","doi-asserted-by":"crossref","unstructured":"Izadinia, H., Shan, Q., & Setiz, S. M. (2017). IM2CAD. In Proceedings of the IEEE conference on computer vision and pattern recognition, CVPR 2017.","DOI":"10.1109\/CVPR.2017.260"},{"key":"2316_CR43","doi-asserted-by":"crossref","unstructured":"Jammalamadaka, N., Zisserman, A., Eichner, M., Ferrari, V., & Jawahar, C. (2012). Has my algorithm succeeded? An evaluator for human pose estimators. In A.\u00a0Fitzgibbon, S.\u00a0Lazebnik, P.\u00a0Perona, Y.\u00a0Sato, & C.\u00a0Schmid (Eds.), Computer vision\u2014ECCV 2012. Lecture notes in computer science 7574. Springer.","DOI":"10.1007\/978-3-642-33712-3_9"},{"key":"2316_CR44","doi-asserted-by":"crossref","unstructured":"Jang, W., & Agapito, L. (2021). CodeNeRF: Disentangled neural radiance fields for object categories. In Proceedings of the IEEE\/CVF international conference on computer vision (ICCV) (pp.\u00a012949\u201312958). arXiv:2109.01750","DOI":"10.1109\/ICCV48922.2021.01271"},{"key":"2316_CR45","unstructured":"Jin, Y., & Geman, S. (2006). Context and hierarchy in a probabilistic image model. In Proceedings of the CVPR 2006. IEEE Computer Society."},{"key":"2316_CR46","doi-asserted-by":"crossref","unstructured":"Johnson, J., Hariharan, B., van\u00a0der Maaten, L., Fei-Fei, L., Zitnick, C. L., & Girshick, R. (2017). CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1988\u20131997).","DOI":"10.1109\/CVPR.2017.215"},{"key":"2316_CR47","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1023\/A:1007665907178","volume":"37","author":"MI Jordan","year":"1999","unstructured":"Jordan, M. I., Ghahramani, Z., Jaakkola, L. K., & Saul, T. S. (1999). An introduction to variational methods for graphical models. Machine Learning, 37, 183\u2013233.","journal-title":"Machine Learning"},{"key":"2316_CR48","unstructured":"Kato, H., Beker, D., Morariu, M., Ando, T., Matsuoka, T., Kehl, W., & Gaidon, A. (2020). Differentiable rendering: A survey. arXiv:2006.12057"},{"key":"2316_CR49","doi-asserted-by":"crossref","unstructured":"Kim, Y., Dyer, C., & Rush, A. (2019). Compound probabilistic context-free grammars for grammar induction. In Proceedings of the 57th annual meeting of the association for computational linguistics (pp. 2369\u20132385).","DOI":"10.18653\/v1\/P19-1228"},{"key":"2316_CR50","unstructured":"Kingma, D.P., & Welling, M. (2014). Auto-encoding variational Bayes. In ICLR."},{"key":"2316_CR51","unstructured":"Kivinen, J. J., & Williams, C. K. I. (2012). Multiple texture Boltzmann machines. In Proceedings of the fifteenth international conference on artificial intelligence and statistics."},{"key":"2316_CR52","unstructured":"Kosiorek, A., Sabour, S., Teh, Y. W., & Hinton, G. E. (2019). Stacked Capsule Autoencoders. In H.\u00a0Wallach, H.\u00a0Larochelle, A.\u00a0Beygelzimer, F.\u00a0d\u2019Alch\u00e9 Buc, E.\u00a0Fox, & R.\u00a0Garnett (Eds.), Advances in neural information processing systems 32."},{"key":"2316_CR53","unstructured":"Kulkarni, T. D., Whitney, W. F., Kohli, P., & Tenenbaum, J. (2015). Deep convolutional inverse graphics network. In C.\u00a0Cortes, N.\u00a0Lawrence, D.\u00a0Lee, M.\u00a0Sugiyama, & R.\u00a0Garnett (Eds.), Advances in neural information processing systems (Vol.\u00a028)."},{"key":"2316_CR54","doi-asserted-by":"crossref","first-page":"1311","DOI":"10.1068\/p2935","volume":"28","author":"M Land","year":"1999","unstructured":"Land, M., Mennie, N., & Rusted, J. (1999). The roles of vision and eye movements in the control of activities of daily living. Perception, 28, 1311\u20131328.","journal-title":"Perception"},{"key":"2316_CR55","doi-asserted-by":"crossref","unstructured":"Lee, H., Grosse, R., Ranganath, R., & Ng, A. Y. (2009). Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations. In L.\u00a0Bottou & M.\u00a0Littman (Eds.), Proceedings of the twenty-sixth international conference on machine learning (ICML 2009).","DOI":"10.1145\/1553374.1553453"},{"key":"2316_CR56","unstructured":"Lee, J., Lee, Y., Kim, J., Kosiorek, A. R., Choi, S., & Teh, Y.-W. (2019). Set transformer: A framework for attention-based permutation-invariant neural networks. In Proceedings of the 36th international conference on machine learning (pp.\u00a03744\u20133753)."},{"issue":"2","key":"2316_CR57","first-page":"1","volume":"37","author":"M Li","year":"2019","unstructured":"Li, M., Gadi Patil, A., Xu, K., Chaudhuri, S., Khan, O., Shamir, A., & Zhang, H. (2019). GRAINS: Generative recursive autoencoders for indoor scenes. ACM Transactions on Graphics, 37(2), 1\u201316.","journal-title":"ACM Transactions on Graphics"},{"key":"2316_CR58","unstructured":"Li, N., Eastwood, C., & Fisher, R. (2020). Learning object-centric representations of multi-object scenes from multiple views. In H.\u00a0Larochelle, M.\u00a0Ranzato, R.\u00a0Hadsell, M.\u00a0Balcan, & H.\u00a0Lin (Eds.), Advances in neural information processing systems (Vol.\u00a033, pp. 5656\u20135666)."},{"key":"2316_CR59","doi-asserted-by":"crossref","unstructured":"Liu, T., Chaudhuri, S., Kim, V., Huang, Q.-X., Mitra, N.J., & Funkhouser, T. (2014). Creating consistent scene graphs using a probabilistic grammar. ACM Transactions on Graphics (Proc. of SIGGRAPH Asia), 33(6).","DOI":"10.1145\/2661229.2661243"},{"key":"2316_CR60","unstructured":"Locatello, F., Bauer, S., Lucic, M., R\u00e4tsch, G., Gelly, S., Sch\u00f6lkopf, B., & Bachem, O. (2019). Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations. In Proceedings of the 36th international conference on machine learning (ICML)."},{"key":"2316_CR61","first-page":"11525","volume":"33","author":"F Locatello","year":"2020","unstructured":"Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., & Kipf, T. (2020). Object-centric learning with slot attention. Advances in Neural Information Processing Systems, 33, 11525\u201311538.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2316_CR62","doi-asserted-by":"crossref","unstructured":"Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., & Black, M. (2015). SMPL: A skinned multi-person linear model. ACM Transactions on Graphics, 34(6).","DOI":"10.1145\/2816795.2818013"},{"key":"2316_CR63","doi-asserted-by":"crossref","unstructured":"Loper, M. M., & Black, M. J. (2014). OpenDR: An approximate differentiable renderer. Computer vision-ECCV 2014 (pp.\u00a0154\u2013169).","DOI":"10.1007\/978-3-319-10584-0_11"},{"key":"2316_CR64","volume-title":"Information theory, inference, and learning algorithms","author":"DJC MacKay","year":"2003","unstructured":"MacKay, D. J. C. (2003). Information theory, inference, and learning algorithms. Cambridge: Cambridge University Press."},{"key":"2316_CR65","doi-asserted-by":"crossref","unstructured":"Maninis, K.-K., Popov, S., Niessner, M., & Ferrari, V. (2023). CAD-Estate: Large-scale CAD model annotation in RGB videos. In Proceedings of the IEEE\/CVF international conference on computer vision (ICCV) (pp. 20132-20142).","DOI":"10.1109\/ICCV51070.2023.01847"},{"key":"2316_CR66","unstructured":"Marcus, G., Davis, E., & Aaronson, S. (2022). A very preliminary analysis of DALL-E 2. arXiv:2204.13807"},{"key":"2316_CR67","doi-asserted-by":"crossref","unstructured":"Mildenhall, B., Srinivasan, P., Tancik, M., Barron, J.T., Ramamoorthi, R., & Ng, R. (2020). NeRF: Representing scenes as neural radiance fields for view synthesis. In A.\u00a0Vedaldi, H.\u00a0Bischof, T.\u00a0Brox, & J.\u00a0Frahm (Eds.), Computer Vision\u2013ECCV 2020. Lecture Notes in Computer Science 12346. Springer.","DOI":"10.1007\/978-3-030-58452-8_24"},{"key":"2316_CR68","doi-asserted-by":"crossref","unstructured":"Moreno, P., Williams, C.K.I., Nash, C., & Kohli, P. (2016). Overcoming occlusion with inverse graphics. In H.\u00a0Gang & H.\u00a0Jegou (Eds.), Computer Vision-ECCV 2016 Workshops Proceedings Part III. LNCS 9915 (pp.\u00a0170\u2013185). Springer.","DOI":"10.1007\/978-3-319-49409-8_16"},{"key":"2316_CR69","volume-title":"Machine learning: A probabilistic perspective","author":"KP Murphy","year":"2012","unstructured":"Murphy, K. P. (2012). Machine learning: A probabilistic perspective. MIT Press."},{"key":"2316_CR70","volume-title":"Probabilistic machine learning: Advanced topics","author":"KP Murphy","year":"2023","unstructured":"Murphy, K. P. (2023). Probabilistic machine learning: Advanced topics. MIT Press."},{"issue":"4","key":"2316_CR71","doi-asserted-by":"crossref","first-page":"727","DOI":"10.1162\/neco_a_01564","volume":"35","author":"A Nazabal","year":"2023","unstructured":"Nazabal, A., Tsagkas, N., & Williams, C. K. I. (2023). Inference and learning for generative capsule models. Neural Computation, 35(4), 727\u2013761.","journal-title":"Neural Computation"},{"key":"2316_CR72","unstructured":"Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., & Chen, M. (2022). GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models. In Proceedings of the 39th international conference on machine learning, PMLR (Vol.\u00a0162)."},{"key":"2316_CR73","unstructured":"Ohta, Y., Kanade, T., & Sakai, T. (1978). An analysis system for scenes containing objects with substructures. In Proceedings of the 4th IJCPR (pp.\u00a0752\u2013754)."},{"key":"2316_CR74","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1016\/S0079-6123(06)55002-2","volume":"155","author":"A Oliva","year":"2006","unstructured":"Oliva, A., & Torralba, A. (2006). Building the gist of a scene: The role of global image features in recognition. Progress in Brain Research, 155, 23\u201336.","journal-title":"Progress in Brain Research"},{"key":"2316_CR75","unstructured":"Open AI (2023). GPT-4 Technical Report. arXiv:2303.08774"},{"key":"2316_CR76","volume-title":"Vision science: Photons to phenomenology","author":"SE Palmer","year":"1999","unstructured":"Palmer, S. E. (1999). Vision science: Photons to phenomenology. MIT Press."},{"key":"2316_CR77","first-page":"12013","volume-title":"Advances in neural information processing systems","author":"D Paschalidou","year":"2021","unstructured":"Paschalidou, D., Kar, A., Shugrina, M., Kreis, K., Geiger, A., & Fidler, S. (2021). ATISS: Autoregressive transformers for indoor scene synthesis. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, & J. W. Vaughan (Eds.), Advances in neural information processing systems (Vol. 34, pp. 12013\u201312026). Curran Associates Inc."},{"key":"2316_CR78","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1016\/0004-3702(90)90005-K","volume":"46","author":"JB Pollack","year":"1990","unstructured":"Pollack, J. B. (1990). Recursive distributed representations. Artificial Intelligence, 46, 77\u2013105.","journal-title":"Artificial Intelligence"},{"key":"2316_CR79","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511996504","volume-title":"Computer vision: Models, learning and inference","author":"SJD Prince","year":"2012","unstructured":"Prince, S. J. D. (2012). Computer vision: Models, learning and inference. Cambridge University Press."},{"issue":"9","key":"2316_CR80","doi-asserted-by":"crossref","first-page":"1537","DOI":"10.1109\/TPAMI.2008.191","volume":"31","author":"JA Quinn","year":"2009","unstructured":"Quinn, J. A., Williams, C. K. I., & McIntosh, N. (2009). Factorial switching linear dynamical systems applied to physiological condition monitoring. IEEE Transactions on Pattern Analysis and Machine Intelligence, 31(9), 1537\u20131551.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2316_CR81","unstructured":"Radford, A., Metz, L., & Chintala, S. (2016). Unsupervised representation learning with deep convolutional generative adversarial networks. In International Conference on Learning Representations (ICLR 2016)."},{"key":"2316_CR82","unstructured":"Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., & Sutskever, I. (2021). Zero-shot text-to-image generation. M.\u00a0Meila & T.\u00a0Zhang (Eds.), Proceedings of the 38th international conference on machine learning (Vol.\u00a0139, pp. 8821\u20138831). PMLR."},{"key":"2316_CR83","unstructured":"Ranzato, M., Mnih, V., & Hinton, G.E. (2010). Generating more realistic images using gated MRF\u2019s. In J.\u00a0Lafferty, C.\u00a0Williams, J.\u00a0Shawe-Taylor, R.\u00a0Zemel, & A.\u00a0Culotta (Eds.), Advances in neural information processing systems (Vol.\u00a023)."},{"key":"2316_CR84","volume-title":"Gaussian processes for machine learning","author":"CE Rasmussen","year":"2006","unstructured":"Rasmussen, C. E., & Williams, C. K. I. (2006). Gaussian processes for machine learning. Cambridge: MIT Press."},{"key":"2316_CR85","unstructured":"Rezende, D.J., Mohamed, S., & Wierstra, D. (2014). Stochastic backpropagation and approximate inference in deep generative models. In ICML."},{"key":"2316_CR86","unstructured":"Ritchie, D. (2019). Generative Models of 3D Scenes. Part of the tutorial on Learning Generative Models of 3D Structures at Eurographics 2019, https:\/\/3dstructgen.github.io"},{"key":"2316_CR87","doi-asserted-by":"crossref","unstructured":"Ritchie, D., Wang, K., & Lin, Y.-A. (2019). Fast and flexible indoor scene synthesis via deep convolutional generative models. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR 2019).","DOI":"10.1109\/CVPR.2019.00634"},{"key":"2316_CR88","unstructured":"Roberts, L. G. (1963). Machine Perception of Three-dimensional Solids (Tech. Rep. No. 315). MIT Lincoln Laboratory."},{"key":"2316_CR89","doi-asserted-by":"crossref","DOI":"10.1016\/j.patcog.2020.107369","volume":"105","author":"L Romaszko","year":"2020","unstructured":"Romaszko, L., Williams, C. K. I., & Winn, J. (2020). Learning direct optimization for scene understanding. Pattern Recognition, 105, 107369.","journal-title":"Pattern Recognition"},{"key":"2316_CR90","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE conference on computer vision and pattern recognition, CVPR 2022.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"2316_CR91","first-page":"2369","volume":"7","author":"D Ross","year":"2006","unstructured":"Ross, D., & Zemel, R. (2006). Learning parts-based representations of data. Journal of Machine Learning Research, 7, 2369\u20132397.","journal-title":"Journal of Machine Learning Research"},{"key":"2316_CR92","doi-asserted-by":"crossref","unstructured":"Roth, S., & Black, M. J. (2005). Fields of experts: A framework for learning image priors. In Proceedings of the IEEE conference on computer vision and pattern recognition, CVPR 2005 (p.\u00a0III:860\u2013867).","DOI":"10.1109\/CVPR.2005.160"},{"issue":"4","key":"2316_CR93","doi-asserted-by":"crossref","first-page":"515","DOI":"10.1070\/SM1977v032n04ABEH002404","volume":"32","author":"JA Rozanov","year":"1977","unstructured":"Rozanov, J. A. (1977). Markov random fields and stochastic partial differential equations. Math. USSR Sbornik, 32(4), 515\u2013534.","journal-title":"Math. USSR Sbornik"},{"key":"2316_CR94","doi-asserted-by":"crossref","unstructured":"R\u00fcnz, M., Li, K., Tang, M., Ma, L., Kong, C., Schmidt, T., & Newcombe, R. (2020). FroDO: From detections to 3D objects. Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR 2020).","DOI":"10.1109\/CVPR42600.2020.01473"},{"key":"2316_CR95","unstructured":"Sabour, S., Frosst, N., & Hinton, G. (2017). Dynamic routing between capsules. In Advances in neural information processing systems (pp. 3856\u20133866)."},{"key":"2316_CR96","unstructured":"Salimans, T., Karpathy, A., Chen, X., Kingma, D.P., & Bulatov, Y. (2017). PixelCNN++: A PixelCNN implementation with discretized logistic mixture likelihood and other modifications. In Proceedings of the international conference on learning representations (ICLR 2017)."},{"key":"2316_CR97","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nberger, J.L., & Frahm, J.-M. (2016). Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR 2016).","DOI":"10.1109\/CVPR.2016.445"},{"issue":"3","key":"2316_CR98","doi-asserted-by":"crossref","first-page":"703","DOI":"10.1214\/18-BA1124","volume":"14","author":"S Seth","year":"2019","unstructured":"Seth, S., Murray, I., & Williams, C. K. I. (2019). Model criticism in latent space. Bayesian Analysis, 14(3), 703\u2013725.","journal-title":"Bayesian Analysis"},{"key":"2316_CR99","doi-asserted-by":"crossref","unstructured":"Shi, M., Caesar, H., & Ferrari, V. (2017). Weakly supervised object localization using things and stuff transfer. In Proceedings of the IEEE international conference on computer vision (ICCV).","DOI":"10.1109\/ICCV.2017.366"},{"key":"2316_CR100","doi-asserted-by":"crossref","unstructured":"Song, S., Yu, F., Zeng, A., Chang, A.X., Savva, M., & Funkhouser, T. (2017). Semantic scene completion from a single depth image. In Proceedings of the IEEE conference on computer vision and pattern recognition, CVPR 2017.","DOI":"10.1109\/CVPR.2017.28"},{"issue":"7","key":"2316_CR101","doi-asserted-by":"crossref","first-page":"1370","DOI":"10.1109\/TPAMI.2013.193","volume":"36","author":"M Sun","year":"2014","unstructured":"Sun, M., Kim, B., Kohli, P., & Savarese, S. (2014). Relating things and stuff via object property interactions. IEEE Transactions Pattern Analysis and Machine Intelligence, 36(7), 1370\u20131383.","journal-title":"IEEE Transactions Pattern Analysis and Machine Intelligence"},{"key":"2316_CR102","unstructured":"Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., & Fergus, R. (2013). Intriguing properties of neural networks. arXiv:1312.6199."},{"key":"2316_CR103","volume-title":"Computer vision: Algorithms and applications","author":"R Szeliski","year":"2021","unstructured":"Szeliski, R. (2021). Computer vision: Algorithms and applications (2nd ed.). Springer.","edition":"2"},{"key":"2316_CR104","doi-asserted-by":"crossref","first-page":"1247","DOI":"10.1162\/089976600300015349","volume":"12","author":"JB Tenenbaum","year":"2000","unstructured":"Tenenbaum, J. B., & Freeman, W. T. (2000). Separating style and content with bilinear models. Neural Computation, 12, 1247\u20131283.","journal-title":"Neural Computation"},{"issue":"3","key":"2316_CR105","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1145\/1666420.1666446","volume":"53","author":"A Torralba","year":"2010","unstructured":"Torralba, A., Murphy, K. P., & Freeman, W. T. (2010). Using the forest to see the trees: Exploiting context for visual object detection and localization. Communications of the ACM, 53(3), 107\u2013114.","journal-title":"Communications of the ACM"},{"issue":"5","key":"2316_CR106","doi-asserted-by":"crossref","first-page":"657","DOI":"10.1109\/34.1000239","volume":"24","author":"Z Tu","year":"2002","unstructured":"Tu, Z., & Zhu, S.-C. (2002). Image segmentation by data-driven Markov chain Monte Carlo. IEEE Transaction on Pattern Analysis and Machine Intelligence, 24(5), 657\u2013673.","journal-title":"IEEE Transaction on Pattern Analysis and Machine Intelligence"},{"key":"2316_CR107","first-page":"27916","volume-title":"Advances in neural information processing systems","author":"GJJ van den Burg","year":"2021","unstructured":"van den Burg, G. J. J., & Williams, C. K. I. (2021). On memorization in probabilistic deep generative models. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, & J. W. Vaughan (Eds.), Advances in neural information processing systems (Vol. 34, pp. 27916\u201327928). Curran Associates Inc."},{"key":"2316_CR108","unstructured":"van\u00a0den Oord, A., Kalchbrenner, N., Espeholt, L., kavukcuoglu, k., Vinyals, O., & Graves, A. (2016). Conditional image generation with PixelCNN decoders. In D.\u00a0Lee, M.\u00a0Sugiyama, U.\u00a0Luxburg, I.\u00a0Guyon, & R.\u00a0Garnett (Eds.), Advances in neural information processing systems."},{"key":"2316_CR109","doi-asserted-by":"crossref","unstructured":"Vasilescu, M.A.O., & Terzopoulos, D. (2002). Multilinear Analysis of Image Ensembles: TensorFaces. A.\u00a0Heyden, G.\u00a0Sparr, M.\u00a0Nielsen, & P.\u00a0Johansen (Eds.), Computer Vision\u2013ECCV 2002. Lecture Notes in Computer Science 2350. Springer.","DOI":"10.1007\/3-540-47969-4_30"},{"key":"2316_CR110","unstructured":"Wang, A., Kortylewski, A., & Yuille, A. (2021). NeMo: Neural mesh models of contrastive features for robust 3D pose estimation. In Proceedings of the international conference on learning representations (ICLR 2021)."},{"key":"2316_CR111","doi-asserted-by":"crossref","first-page":"1039","DOI":"10.1162\/089976604773135096","volume":"165","author":"CKI Williams","year":"2004","unstructured":"Williams, C. K. I., & Titsias, M. K. (2004). Greedy learning of multiple objects in images using robust statistics and factorial learning. Neural Computation, 165, 1039\u20131062.","journal-title":"Neural Computation"},{"key":"2316_CR112","doi-asserted-by":"crossref","unstructured":"Xia, Y., Zhang, Y., Liu, F., Shen, W., & Yuille, A. (2020). Synthesize then compare: Detecting failures and anomalies for semantic segmentation. In A.\u00a0Vedaldi, H.\u00a0Bischof, T.\u00a0Brox, & J.\u00a0Frahm (Eds.), Computer Vision\u2014ECCV 2020. Lecture Notes in Computer Science 12346. Springer.","DOI":"10.1007\/978-3-030-58452-8_9"},{"key":"2316_CR113","unstructured":"Yao, B., Yang, X., & Zhu, S.-C. (2007). Introduction to a large-scale general purpose ground truth database: methodology, annotation tool and benchmarks. In EMMCVPR\u201907: Proceedings of the 6th international conference on energy minimization methods in computer vision and pattern recognition."},{"issue":"4","key":"2316_CR114","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2185520.2185552","volume":"31","author":"Y-T Yeh","year":"2012","unstructured":"Yeh, Y.-T., Yang, L., Watson, M., Goodman, N. D., & Hanrahan, P. (2012). Synthesizing open worlds with constraints using locally annealed reversible jump MCMC. ACM Transactions on Graphics, 31(4), 1\u201311.","journal-title":"ACM Transactions on Graphics"},{"key":"2316_CR115","doi-asserted-by":"crossref","unstructured":"Yu, L.-F., Yeung, S.-K., Tang, C.-K., Terzopoulos, D., Chan, T.F., & Osher, S.J. (2011). Make it home: Automatic optimization of furniture arrangement. ACM Transactions on Graphics, 30(4).","DOI":"10.1145\/2010324.1964981"},{"key":"2316_CR116","doi-asserted-by":"crossref","first-page":"781","DOI":"10.1007\/s11263-020-01405-z","volume":"129","author":"AL Yuille","year":"2021","unstructured":"Yuille, A. L., & Liu, C. (2021). Deep Nets: What have they ever done for Vision? International Journal of Computer Vision, 129, 781\u2013802.","journal-title":"International Journal of Computer Vision"},{"key":"2316_CR117","doi-asserted-by":"crossref","unstructured":"Zamir, A. R., Sax, A., Yeo, T., Kar, O., Cheerla, N., Suri, R., & Guibas, L. (2020). Robust learning through cross-task consistency. Proceedings of the IEEE conference on computer vision and pattern recognition, CVPR 2020.","DOI":"10.1109\/CVPR42600.2020.01121"},{"key":"2316_CR118","doi-asserted-by":"crossref","unstructured":"Zhang, X., Srinivasan, P. P., Deng, B., Debevec, P., Freeman, W. T., & Barron, J. T. (2021). NeRFactor: Neural factorization of shape and reflectance under an unknown illumination. ACM Transactions on Graphics (Proc. of SIGGRAPH Asia), 40, 6.","DOI":"10.1145\/3478513.3480496"},{"key":"2316_CR119","volume-title":"Computer vision: Stochastic grammars for parsing objects, scenes, and events","author":"S-C Zhu","year":"2021","unstructured":"Zhu, S.-C., & Huang, S. (2021). Computer vision: Stochastic grammars for parsing objects, scenes, and events. Springer."},{"issue":"4","key":"2316_CR120","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1561\/0600000018","volume":"2","author":"S-C Zhu","year":"2006","unstructured":"Zhu, S.-C., & Mumford, D. (2006). A stochastic grammar of images. Foundations and Trends in Computer Graphics and Vision, 2(4), 259\u2013362.","journal-title":"Foundations and Trends in Computer Graphics and Vision"},{"issue":"2","key":"2316_CR121","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1023\/A:1007925832420","volume":"27","author":"S-C Zhu","year":"1998","unstructured":"Zhu, S.-C., Wu, Y., & Mumford, D. (1998). Filters, Random Fields and Maximum Entropy (FRAME): Towards a unified theory for texture modeling. International Journal of Computer Vision, 27(2), 107\u2013126.","journal-title":"International Journal of Computer Vision"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-024-02316-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-024-02316-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-024-02316-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,17]],"date-time":"2025-04-17T06:05:17Z","timestamp":1744869917000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-024-02316-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,12]]},"references-count":121,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2025,5]]}},"alternative-id":["2316"],"URL":"https:\/\/doi.org\/10.1007\/s11263-024-02316-z","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,12]]},"assertion":[{"value":"4 February 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 November 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 December 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}