{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,2]],"date-time":"2026-08-02T06:45:56Z","timestamp":1785653156646,"version":"3.56.0"},"publisher-location":"Cham","reference-count":18,"publisher":"Springer Nature Switzerland","isbn-type":[{"value":"9783032316653","type":"print"},{"value":"9783032316660","type":"electronic"}],"license":[{"start":{"date-parts":[[2026,8,3]],"date-time":"2026-08-03T00:00:00Z","timestamp":1785715200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,8,3]],"date-time":"2026-08-03T00:00:00Z","timestamp":1785715200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2027]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    The reliable computational assessment of photographic composition requires features that are discriminative of spatial layout yet robust to semantic content. This paper proposes a low-level representation grounded in the assumption that composition can be understood as the flow of visual attention across geometric structure. We introduce VFCNet, which fuses saliency and edge information into a gradient vector flow (GVF) field. The model computes dual-stream GVF representations, integrates them via attention, and extracts multi-scale flow features with a DINOv3 backbone. VFCNet achieves state-of-the-art performance on the PICD benchmark (CDA-1: 0.683, CDA-2: 0.629), improving by 33.1% and 36.1% over the previous best method. We also show that a simple classifier on self-supervised DINOv3 features substantially outperforms more sophisticated, composition-specialized models. Code is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/ADadras\/VFCNet\" ext-link-type=\"uri\">https:\/\/github.com\/ADadras\/VFCNet<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/978-3-032-31666-0_35","type":"book-chapter","created":{"date-parts":[[2026,8,2]],"date-time":"2026-08-02T05:46:25Z","timestamp":1785649585000},"page":"530-544","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Semantically Stable Image Composition Analysis via\u00a0Saliency and\u00a0Gradient Vector Flow Fusion"],"prefix":"10.1007","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6474-7208","authenticated-orcid":false,"given":"Armin","family":"Dadras","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4195-1593","authenticated-orcid":false,"given":"Robert","family":"Sablatnig","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3455-8582","authenticated-orcid":false,"given":"Franziska","family":"Proksa","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7966-3602","authenticated-orcid":false,"given":"Markus","family":"Seidl","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,8,3]]},"reference":[{"key":"35_CR1","doi-asserted-by":"publisher","unstructured":"He, S., Ming, A., Zheng, S., Zhong, H., Ma, H.: EAT: an enhancer for aesthetics-oriented transformers. In: Proceedings of the 31st ACM International Conference on Multimedia, Ottawa, ON, Canada, pp. 1023\u20131032. ACM (2023). https:\/\/doi.org\/10.1145\/3581783.3611881","DOI":"10.1145\/3581783.3611881"},{"key":"35_CR2","doi-asserted-by":"crossref","unstructured":"Hong, C., Du, S., Xian, K., Lu, H., Cao, Z., Zhong, W.: Composing photos like a photographer. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 7057\u20137066 (2021)","DOI":"10.1109\/CVPR46437.2021.00698"},{"issue":"3","key":"35_CR3","doi-asserted-by":"publisher","first-page":"121","DOI":"10.1007\/s00530-024-01307-x","volume":"30","author":"Q Hou","year":"2024","unstructured":"Hou, Q., Ke, Y., Wang, K., Qin, F., Wang, Y.: Synchronous composition and semantic line detection based on cross-attention. Multimedia Syst. 30(3), 121 (2024)","journal-title":"Multimedia Syst."},{"key":"35_CR4","unstructured":"Kandinsky, W.: Point and Line to Plane: Contribution to the Analysis of the Pictorial Elements. Solomon R. Guggenheim Foundation, New York (1947). Translated by Howard Dearstyne, edited by Hilla Rebay"},{"key":"35_CR5","doi-asserted-by":"crossref","unstructured":"Ke, J., Wang, Q., Wang, Y., Milanfar, P., Yang, F.: MUSIQ: multi-scale image quality transformer. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), pp. 5148\u20135157 (2021)","DOI":"10.1109\/ICCV48922.2021.00510"},{"key":"35_CR6","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1016\/j.jvcir.2018.05.018","volume":"55","author":"JT Lee","year":"2018","unstructured":"Lee, J.T., Kim, H.U., Lee, C., Kim, C.S.: Photographic composition classification and dominant geometric element detection for outdoor scenes. J. Vis. Commun. Image Represent. 55, 91\u2013105 (2018)","journal-title":"J. Vis. Commun. Image Represent."},{"key":"35_CR7","doi-asserted-by":"crossref","unstructured":"Li, D., Zhang, J., Huang, K., Yang, M.H.: Composing good shots by exploiting mutual relations. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4213\u20134222 (2020)","DOI":"10.1109\/CVPR42600.2020.00427"},{"key":"35_CR8","doi-asserted-by":"crossref","unstructured":"Linardos, A., K\u00fcmmerer, M., Press, O., Bethge, M.: DeepGaze IIE: calibrated prediction in and out-of-domain for state-of-the-art saliency modeling. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 12919\u201312928 (2021)","DOI":"10.1109\/ICCV48922.2021.01268"},{"key":"35_CR9","doi-asserted-by":"crossref","unstructured":"She, D., Lai, Y., Yi, G., Xu, K.: Hierarchical layout-aware graph convolutional network for unified aesthetics assessment. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8475\u20138484 (2021)","DOI":"10.1109\/CVPR46437.2021.00837"},{"key":"35_CR10","unstructured":"Sim\u00e9oni, O., et al.: DINOv3 (2025). https:\/\/arxiv.org\/abs\/2508.10104"},{"key":"35_CR11","unstructured":"Su, Y., Cao, Y., Deng, J., Rao, F., Wu, Q.: Spatial-semantic collaborative cropping for user generated content (2024). https:\/\/arxiv.org\/abs\/2401.08086"},{"key":"35_CR12","doi-asserted-by":"publisher","DOI":"10.1016\/j.jvcir.2023.103751","volume":"90","author":"Y Wang","year":"2023","unstructured":"Wang, Y., Ke, Y., Wang, K., Guo, J., Yang, S.: Spatial-invariant convolutional neural network for photographic composition prediction and automatic correction. J. Vis. Commun. Image Represent. 90, 103751 (2023)","journal-title":"J. Vis. Commun. Image Represent."},{"key":"35_CR13","unstructured":"Yaseen, M.: What is YOLOv8: an in-depth exploration of the internal features of the next-generation object detector (2024). https:\/\/arxiv.org\/abs\/2408.15857"},{"key":"35_CR14","doi-asserted-by":"crossref","unstructured":"Yi, R., Tian, H., Gu, Z., Lai, Y.K., Rosin, P.L.: Towards artistic image aesthetics assessment: a large-scale dataset and a new method. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 22388\u201322397 (2023)","DOI":"10.1109\/CVPR52729.2023.02144"},{"issue":"3","key":"35_CR15","doi-asserted-by":"publisher","first-page":"1304","DOI":"10.1109\/TPAMI.2020.3024207","volume":"44","author":"H Zeng","year":"2022","unstructured":"Zeng, H., Li, L., Cao, Z., Zhang, L.: Grid anchor based image cropping: a new benchmark and an efficient model. IEEE Trans. Pattern Anal. Mach. Intell. 44(3), 1304\u20131319 (2022). https:\/\/doi.org\/10.1109\/TPAMI.2020.3024207","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"35_CR16","doi-asserted-by":"crossref","unstructured":"Zhang, B., Niu, L., Zhang, L.: Image composition assessment with saliency-augmented multi-pattern pooling. arXiv preprint arXiv:2104.03133 (2021)","DOI":"10.5244\/C.35.106"},{"key":"35_CR17","doi-asserted-by":"crossref","unstructured":"Zhao, Z., Lu, P., Peng, X., Guo, W.: Self-supervised photographic image layout representation learning, arXiv preprint arXiv:2403.03740 (2024)","DOI":"10.1109\/TMM.2025.3599102"},{"key":"35_CR18","doi-asserted-by":"crossref","unstructured":"Zhao, Z., et al.: Can machines understand composition? Dataset and benchmark for photographic image composition embedding and understanding. In: Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 14411\u201314421 (2025)","DOI":"10.1109\/CVPR52734.2025.01344"}],"container-title":["Lecture Notes in Computer Science","Pattern Recognition"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-032-31666-0_35","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,2]],"date-time":"2026-08-02T05:46:27Z","timestamp":1785649587000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-032-31666-0_35"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8,3]]},"ISBN":["9783032316653","9783032316660"],"references-count":18,"URL":"https:\/\/doi.org\/10.1007\/978-3-032-31666-0_35","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"value":"0302-9743","type":"print"},{"value":"1611-3349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,8,3]]},"assertion":[{"value":"3 August 2026","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"ICPR","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"International Conference on Pattern Recognition","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Lyon","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"France","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2026","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"17 August 2026","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"22 August 2026","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"28","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"icpr2026","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/icpr2026.org\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}