{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,20]],"date-time":"2025-11-20T19:06:19Z","timestamp":1763665579941,"version":"3.41.0"},"reference-count":68,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,10,31]],"date-time":"2024-10-31T00:00:00Z","timestamp":1730332800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Science Foundation","award":["#2232817"],"award-info":[{"award-number":["#2232817"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Appl. Percept."],"published-print":{"date-parts":[[2024,10,31]]},"abstract":"<jats:p>Diffusion models offer unprecedented image generation power given just a text prompt. While emerging approaches for controlling diffusion models have enabled users to specify the desired spatial layouts of the generated content, they cannot predict or control where viewers will pay more attention due to the complexity of human vision. Recognizing the significance of attention-controllable image generation in practical applications, we present a saliency-guided framework to incorporate the data priors of human visual attention mechanisms into the generation process. Given a user-specified viewer attention distribution, our control module conditions a diffusion model to generate images that attract viewers\u2019 attention toward the desired regions. To assess the efficacy of our approach, we performed an eye-tracked user study and a large-scale model-based saliency analysis. The results evidence that both the cross-user eye gaze distributions and the saliency models\u2019 predictions align with the desired attention distributions. Lastly, we outline several applications, including interactive design of saliency guidance, attention suppression in unwanted regions, and adaptive generation for varied display\/viewing conditions.<\/jats:p>","DOI":"10.1145\/3694969","type":"journal-article","created":{"date-parts":[[2024,9,6]],"date-time":"2024-09-06T15:45:29Z","timestamp":1725637529000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["GazeFusion: Saliency-Guided Image Generation"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0189-1776","authenticated-orcid":false,"given":"Yunxiang","family":"Zhang","sequence":"first","affiliation":[{"name":"New York University, Brooklyn, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-0866-334X","authenticated-orcid":false,"given":"Nan","family":"Wu","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1388-6974","authenticated-orcid":false,"given":"Connor Z.","family":"Lin","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9243-6885","authenticated-orcid":false,"given":"Gordon","family":"Wetzstein","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3094-5844","authenticated-orcid":false,"given":"Qi","family":"Sun","sequence":"additional","affiliation":[{"name":"New York University, Brooklyn, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,11,15]]},"reference":[{"key":"e_1_3_1_2_1","first-page":"19851","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Aberman Kfir","year":"2022","unstructured":"Kfir Aberman, Junfeng He, Yossi Gandelsman, Inbar Mosseri, David E. Jacobs, Kai Kohlhoff, Yael Pritch, and Michael Rubinstein. 2022. Deep saliency prior for reducing visual distraction. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 19851\u201319860."},{"key":"e_1_3_1_3_1","first-page":"1728","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Bain Max","year":"2021","unstructured":"Max Bain, Arsha Nagrani, G\u00fcl Varol, and Andrew Zisserman. 2021. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 1728\u20131738."},{"key":"e_1_3_1_4_1","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1016\/j.sbspro.2015.06.349","article-title":"Attributes for image content that attract consumers\u2019 attention to advertisements","volume":"195","author":"Bakar Muhammad Helmi Abu","year":"2015","unstructured":"Muhammad Helmi Abu Bakar, Mohd Asyiek Mat Desa, and Muhizam Mustafa. 2015. Attributes for image content that attract consumers\u2019 attention to advertisements. Procedia-Social and Behavioral Sciences 195 (2015), 309\u2013314.","journal-title":"Procedia-Social and Behavioral Sciences"},{"key":"e_1_3_1_5_1","unstructured":"Andreas Blattmann Tim Dockhorn Sumith Kulal Daniel Mendelevitch Maciej Kilian Dominik Lorenz Yam Levi Zion English Vikram Voleti Adam Letts Varun Jampani and Robin Rombach. 2023. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv:2311.15127. https:\/\/arxiv.org\/abs\/2311.15127"},{"key":"e_1_3_1_6_1","first-page":"12","article-title":"Salient object detection: A benchmark","volume":"24","author":"Borji Ali","year":"2015","unstructured":"Ali Borji, Ming-Ming Cheng, Huaizu Jiang, and Jia Li. 2015. Salient object detection: A benchmark. IEEE Transactions on Image Processing 24, 12 (2015), 5706\u20135722.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_7_1","unstructured":"Ali Borji and Laurent Itti. 2015. Cat2000: A large scale fixation dataset for boosting saliency research. arXiv:1505.03581. https:\/\/arxiv.org\/abs\/1505.03581"},{"key":"e_1_3_1_8_1","unstructured":"Neil Bruce and John Tsotsos. 2005. Saliency based on information maximization. In Proceedings of the 18th International Conference onNeural Information Processing Systems 155\u2013162."},{"issue":"9","key":"e_1_3_1_9_1","doi-asserted-by":"crossref","first-page":"950","DOI":"10.1167\/7.9.950","article-title":"Attention based on information maximization","volume":"7","author":"Bruce Neil","year":"2007","unstructured":"Neil Bruce and John Tsotsos. 2007. Attention based on information maximization. Journal of Vision 7, 9 (2007), 950\u2013950.","journal-title":"Journal of Vision"},{"key":"e_1_3_1_10_1","first-page":"419","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Droste Richard","year":"2020","unstructured":"Richard Droste, Jianbo Jiao, and J Alison Noble. 2020. Unified image and video saliency modeling. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 419\u2013435."},{"issue":"3","key":"e_1_3_1_11_1","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1167\/8.3.3","article-title":"Interesting objects are visually salient","volume":"8","author":"Elazary Lior","year":"2008","unstructured":"Lior Elazary and Laurent Itti. 2008. Interesting objects are visually salient. Journal of Vision 8, 3 (2008), 3\u20133.","journal-title":"Journal of Vision"},{"issue":"5","key":"e_1_3_1_12_1","first-page":"583","article-title":"Allocation of attention in the visual field","volume":"11","author":"Eriksen Charles W","year":"1985","unstructured":"Charles W Eriksen and Yei-yu Yeh. 1985. Allocation of attention in the visual field. Journal of Experimental Psychology: Human Perception and Performance 11, 5 (1985), 583.","journal-title":"Journal of Experimental Psychology: Human Perception and Performance"},{"key":"e_1_3_1_13_1","first-page":"4473","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fosco Camilo","year":"2020","unstructured":"Camilo Fosco, Anelise Newman, Pat Sukhum, Yun Bin Zhang, Nanxuan Zhao, Aude Oliva, and Zoya Bylinskii. 2020. How much time do you have? Modeling multi-duration saliency. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 4473\u20134482."},{"issue":"3","key":"e_1_3_1_14_1","doi-asserted-by":"crossref","first-page":"433","DOI":"10.1177\/0018720819889533","article-title":"Does using multiple computer monitors for office tasks affect user experience? a systematic review","volume":"63","author":"Gallagher Kaitlin M.","year":"2021","unstructured":"Kaitlin M. Gallagher, Laura Cameron, Diana De Carvalho, and Madison Boule. 2021. Does using multiple computer monitors for office tasks affect user experience? a systematic review. Human Factors 63, 3 (2021), 433\u2013449.","journal-title":"Human Factors"},{"key":"e_1_3_1_15_1","first-page":"1220","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Golestaneh S. Alireza","year":"2022","unstructured":"S. Alireza Golestaneh, Saba Dadsetan, and Kris M Kitani. 2022. No-reference image quality assessment via transformers, relative ranking, and self-consistency. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 1220\u20131230."},{"key":"e_1_3_1_16_1","unstructured":"Yuwei Guo Ceyuan Yang Anyi Rao Yaohui Wang Yu Qiao Dahua Lin and Bo Dai. 2023. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv:2307.04725. https:\/\/arxiv.org\/abs\/2307.04725"},{"key":"e_1_3_1_17_1","first-page":"545","volume-title":"Proceedings of the 19th International Conference on Neural Information Processing Systems","author":"Harel Jonathan","year":"2006","unstructured":"Jonathan Harel, Christof Koch, and Pietro Perona. 2006. Graph-based visual saliency. In Proceedings of the 19th International Conference on Neural Information Processing Systems, 545\u2013552."},{"issue":"1","key":"e_1_3_1_18_1","doi-asserted-by":"crossref","first-page":"18434","DOI":"10.1038\/s41598-021-97879-z","article-title":"Deep saliency models learn low-, mid-, and high-level features to predict scene attention","volume":"11","author":"Hayes Taylor R.","year":"2021","unstructured":"Taylor R. Hayes and John M. Henderson. 2021. Deep saliency models learn low-, mid-, and high-level features to predict scene attention. Scientific Reports 11, 1 (2021), 18434.","journal-title":"Scientific Reports"},{"key":"e_1_3_1_19_1","volume-title":"International Conference on Learning Representations","author":"Hu Edward J.","year":"2021","unstructured":"Edward J. Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen. 2021. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations."},{"key":"e_1_3_1_20_1","first-page":"262","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Huang Xun","year":"2015","unstructured":"Xun Huang, Chengyao Shen, Xavier Boix, and Qi Zhao. 2015. Salicon: Reducing the semantic gap in saliency prediction by adapting deep neural networks. In Proceedings of the IEEE International Conference on Computer Vision, 262\u2013270."},{"issue":"3","key":"e_1_3_1_21_1","doi-asserted-by":"crossref","first-page":"194","DOI":"10.1038\/35058500","article-title":"Computational modelling of visual attention","volume":"2","author":"Itti Laurent","year":"2001","unstructured":"Laurent Itti and Christof Koch. 2001. Computational modelling of visual attention. Nature Reviews Neuroscience 2, 3 (2001), 194\u2013203.","journal-title":"Nature Reviews Neuroscience"},{"issue":"11","key":"e_1_3_1_22_1","doi-asserted-by":"crossref","first-page":"1254","DOI":"10.1109\/34.730558","article-title":"A model of saliency-based visual attention for rapid scene analysis","volume":"20","author":"Itti Laurent","year":"1998","unstructured":"Laurent Itti, Christof Koch, and Ernst Niebur. 1998. A model of saliency-based visual attention for rapid scene analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence 20, 11 (1998), 1254\u20131259.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_23_1","doi-asserted-by":"crossref","first-page":"103887","DOI":"10.1016\/j.imavis.2020.103887","article-title":"Eml-net: An expandable multi-layer network for saliency prediction","volume":"95","author":"Jia Sen","year":"2020","unstructured":"Sen Jia and Neil D. B. Bruce. 2020. Eml-net: An expandable multi-layer network for saliency prediction. Image and Vision Computing 95 (2020), 103887.","journal-title":"Image and Vision Computing"},{"key":"e_1_3_1_24_1","first-page":"602","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Jiang Lai","year":"2018","unstructured":"Lai Jiang, Mai Xu, Tie Liu, Minglang Qiao, and Zulin Wang. 2018. Deepvs: A deep learning based video saliency prediction approach. In Proceedings of the European Conference on Computer Vision (ECCV), 602\u2013617."},{"key":"e_1_3_1_25_1","first-page":"16509","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Jiang Lai","year":"2021","unstructured":"Lai Jiang, Mai Xu, Xiaofei Wang, and Leonid Sigal. 2021. Saliency-guided image translation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 16509\u201316518."},{"key":"e_1_3_1_26_1","first-page":"1072","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Jiang Ming","year":"2015","unstructured":"Ming Jiang, Shengsheng Huang, Juanyong Duan, and Qi Zhao. 2015. Salicon: Saliency in context. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1072\u20131080."},{"key":"e_1_3_1_27_1","first-page":"2106","volume-title":"Proceedings of the IEEE 12th International Conference on Computer Vision","author":"Judd Tilke","year":"2009","unstructured":"Tilke Judd, Krista Ehinger, Fr\u00e9do Durand, and Antonio Torralba. 2009. Learning to predict where humans look. In Proceedings of the IEEE 12th International Conference on Computer Vision. IEEE, 2106\u20132113."},{"issue":"1","key":"e_1_3_1_28_1","doi-asserted-by":"crossref","first-page":"315","DOI":"10.1146\/annurev.neuro.23.1.315","article-title":"Mechanisms of visual attention in the human cortex","volume":"23","author":"Kastner Sabine","year":"2000","unstructured":"Sabine Kastner and Leslie G. Ungerleider. 2000. Mechanisms of visual attention in the human cortex. Annual Review of Neuroscience 23, 1 (2000), 315\u2013341.","journal-title":"Annual Review of Neuroscience"},{"key":"e_1_3_1_29_1","doi-asserted-by":"crossref","unstructured":"Levon Khachatryan Andranik Movsisyan Vahram Tadevosyan Roberto Henschel Zhangyang Wang Shant Navasardyan and Humphrey Shi. 2023. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. arXiv:2303.13439. https:\/\/arxiv.org\/abs\/2303.13439","DOI":"10.1109\/ICCV51070.2023.01462"},{"issue":"5","key":"e_1_3_1_30_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3131275","article-title":"Bubbleview: An interface for crowdsourcing image importance maps and tracking visual attention","volume":"24","author":"Kim Nam Wook","year":"2017","unstructured":"Nam Wook Kim, Zoya Bylinskii, Michelle A Borkin, Krzysztof Z Gajos, Aude Oliva, Fredo Durand, and Hanspeter Pfister. 2017. Bubbleview: An interface for crowdsourcing image importance maps and tracking visual attention. ACM Transactions on Computer-Human Interaction 24, 5 (2017), 1\u201340.","journal-title":"ACM Transactions on Computer-Human Interaction"},{"key":"e_1_3_1_31_1","volume-title":"Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915)","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915)."},{"key":"e_1_3_1_32_1","first-page":"1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR \u201915)","author":"K\u00fcmmerer M.","year":"2014","unstructured":"M. K\u00fcmmerer, L. Theis, and M. Bethge. 2014. Deep Gaze I: Boosting saliency prediction with feature maps trained on ImageNet. In Proceedings of the International Conference on Learning Representations (ICLR \u201915), 1\u201312."},{"key":"e_1_3_1_33_1","first-page":"16054","volume-title":"Proceedings of the National Academy of Sciences","volume":"112","author":"K\u00fcmmerer Matthias","year":"2015","unstructured":"Matthias K\u00fcmmerer, Thomas S. A. Wallis, and Matthias Bethge. 2015. Information-theoretic model comparison unifies saliency metrics. Proceedings of the National Academy of Sciences 112, 52 (2015), 16054\u201316059."},{"key":"e_1_3_1_34_1","first-page":"770","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Kummerer Matthias","year":"2018","unstructured":"Matthias Kummerer, Thomas SA Wallis, and Matthias Bethge. 2018. Saliency benchmarking made easy: Separating models, maps and metrics. In Proceedings of the European Conference on Computer Vision (ECCV), 770\u2013787."},{"key":"e_1_3_1_35_1","first-page":"4789","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Kummerer Matthias","year":"2017","unstructured":"Matthias Kummerer, Thomas S. A. Wallis, Leon A. Gatys, and Matthias Bethge. 2017. Understanding low-and high-level contributions to fixation prediction. In Proceedings of the IEEE International Conference on Computer Vision, 4789\u20134798."},{"key":"e_1_3_1_36_1","unstructured":"Junnan Li Dongxu Li Silvio Savarese and Steven Hoi. 2023a. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv:2301.12597. https:\/\/arxiv.org\/abs\/2301.12597"},{"key":"e_1_3_1_37_1","first-page":"22511","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Yuheng","year":"2023","unstructured":"Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee. 2023b. Gligen: Open-set grounded text-to-image generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 22511\u201322521."},{"key":"e_1_3_1_38_1","first-page":"740","volume-title":"Proceedings of the 13th European Conference on Computer Vision (ECCV \u201924)","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In Proceedings of the 13th European Conference on Computer Vision (ECCV \u201924). Springer, 740\u2013755."},{"key":"e_1_3_1_39_1","first-page":"12919","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Linardos Akis","year":"2021","unstructured":"Akis Linardos, Matthias K\u00fcmmerer, Ori Press, and Matthias Bethge. 2021. DeepGaze IIE: Calibrated prediction in and out-of-domain for state-of-the-art saliency modeling. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 12919\u201312928."},{"key":"e_1_3_1_40_1","unstructured":"Shilong Liu Zhaoyang Zeng Tianhe Ren Feng Li Hao Zhang Jie Yang Chunyuan Li Jianwei Yang Hang Su Jun Zhu and Lei Zhang. 2023. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv:2303.05499. https:\/\/arxiv.org\/abs\/2303.05499"},{"issue":"5","key":"e_1_3_1_41_1","doi-asserted-by":"crossref","first-page":"2003","DOI":"10.1109\/TVCG.2022.3150502","article-title":"Scangan360: A generative model of realistic scanpaths for 360 images","volume":"28","author":"Martin Daniel","year":"2022","unstructured":"Daniel Martin, Ana Serrano, Alexander W Bergman, Gordon Wetzstein, and Belen Masia. 2022. Scangan360: A generative model of realistic scanpaths for 360 images. IEEE Transactions on Visualization and Computer Graphics 28, 5 (2022), 2003\u20132013.","journal-title":"IEEE Transactions on Visualization and Computer Graphics"},{"issue":"3","key":"e_1_3_1_42_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1531326.1531361","article-title":"Eye-catching crowds: Saliency based selective variation","volume":"28","author":"McDonnell Rachel","year":"2009","unstructured":"Rachel McDonnell, Mich\u00e9al Larkin, Benjam\u00edn Hern\u00e1ndez, Isaac Rudomin, and Carol O\u2019Sullivan. 2009. Eye-catching crowds: Saliency based selective variation. ACM Transactions on Graphics 28, 3 (2009), 1\u201310.","journal-title":"ACM Transactions on Graphics"},{"key":"e_1_3_1_43_1","first-page":"343","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Mejjati Youssef A.","year":"2020","unstructured":"Youssef A. Mejjati, Celso F. Gomez, Kwang In Kim, Eli Shechtman, and Zoya Bylinskii. 2020. Look here! A parametric learning based approach to redirect visual attention. In Proceedings of the European Conference on Computer Vision, 343\u2013361."},{"key":"e_1_3_1_44_1","first-page":"186","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Mahdi S.","year":"2023","unstructured":"S. Mahdi, H. Miangoleh, Zoya Bylinskii, Eric Kee, Eli Shechtman, and Ya\u011fiz Aksoy. 2023. Realistic saliency guided image enhancement. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 186\u2013194."},{"key":"e_1_3_1_45_1","first-page":"2394","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Min Kyle","year":"2019","unstructured":"Kyle Min and Jason J. Corso. 2019. Tased-net: Temporally-aggregating spatial encoder-decoder network for video saliency detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2394\u20132403."},{"key":"e_1_3_1_46_1","unstructured":"Junting Pan Cristian Canton Ferrer Kevin McGuinness Noel E. O\u2019Connor Jordi Torres Elisa Sayrol and Xavier Giro-i Nieto. 2017. Salgan: Visual saliency prediction with generative adversarial networks. arXiv:1701.01081. https:\/\/arxiv.org\/abs\/1701.01081"},{"issue":"6","key":"e_1_3_1_47_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2980179.2982422","article-title":"Directing user attention via visual flow on web designs","volume":"35","author":"Pang Xufang","year":"2016","unstructured":"Xufang Pang, Ying Cao, Rynson W. H. Lau, and Antoni B. Chan. 2016. Directing user attention via visual flow on web designs. ACM Transactions on Graphics 35, 6 (2016), 1\u201311.","journal-title":"ACM Transactions on Graphics"},{"key":"e_1_3_1_48_1","first-page":"18","article-title":"Components of bottom-up gaze allocation in natural images","volume":"45","author":"Peters Robert J.","year":"2005","unstructured":"Robert J. Peters, Asha Iyer, Laurent Itti, and Christof Koch. 2005. Components of bottom-up gaze allocation in natural images. Vision Research 45, 18 (2005), 2397\u20132416.","journal-title":"Vision Research"},{"key":"e_1_3_1_49_1","unstructured":"Ryan Po Wang Yifan Vladislav Golyanik Kfir Aberman Jonathan T. Barron Amit H. Bermano Eric Ryan Chan Tali Dekel Aleksander Holynski Angjoo Kanazawa C. Karen Liu Lingjie Liu Ben Mildenhall Matthias Nie\u00dfner Bj\u00f6rn Ommer Christian Theobalt Peter Wonka and Gordon Wetzstein. 2023. State of the art on diffusion models for visual computing. arXiv:2310.07204. https:\/\/arxiv.org\/abs\/2310.07204"},{"key":"e_1_3_1_50_1","first-page":"8748","article-title":"Learning transferable visual models from natural language supervision","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763.","journal-title":"Proceedings of the International Conference on Machine Learning."},{"issue":"1","key":"e_1_3_1_51_1","doi-asserted-by":"crossref","first-page":"49","DOI":"10.1207\/s1532785xmep0101_4","article-title":"The effects of screen size and message content on attention and arousal","volume":"1","author":"Reeves Byron","year":"1999","unstructured":"Byron Reeves, Annie Lang, Eun Young Kim, and Deborah Tatar. 1999. The effects of screen size and message content on attention and arousal. Media Psychology 1, 1 (1999), 49\u201367.","journal-title":"Media Psychology"},{"key":"e_1_3_1_52_1","first-page":"25","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops","author":"Ren Jianqiang","year":"2015","unstructured":"Jianqiang Ren, Xiaojin Gong, Lu Yu, Wenhui Zhou, and Michael Ying Yang. 2015. Exploiting global priors for RGB-D saliency detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 25\u201332."},{"key":"e_1_3_1_53_1","first-page":"10684","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Rombach Robin","year":"2022","unstructured":"Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj\u00f6rn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 10684\u201310695."},{"issue":"4","key":"e_1_3_1_54_1","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1037\/aca0000025","article-title":"How attention is driven by film edits: A multimodal experience","volume":"9","author":"Shimamura Arthur P.","year":"2015","unstructured":"Arthur P. Shimamura, Brendan I. Cohn-Sheehy, Brianna L. Pogue, and Thomas A. Shimamura. 2015. How attention is driven by film edits: A multimodal experience. Psychology of Aesthetics, Creativity, and the Arts 9, 4 (2015), 417.","journal-title":"Psychology of Aesthetics, Creativity, and the Arts"},{"issue":"4","key":"e_1_3_1_55_1","doi-asserted-by":"crossref","first-page":"1633","DOI":"10.1109\/TVCG.2018.2793599","article-title":"Saliency in VR: How do people explore virtual environments?","volume":"24","author":"Sitzmann Vincent","year":"2018","unstructured":"Vincent Sitzmann, Ana Serrano, Amy Pavel, Maneesh Agrawala, Diego Gutierrez, Belen Masia, and Gordon Wetzstein. 2018. Saliency in VR: How do people explore virtual environments? IEEE Transactions on Visualization and Computer Graphics 24, 4 (2018), 1633\u20131642.","journal-title":"IEEE Transactions on Visualization and Computer Graphics"},{"key":"e_1_3_1_56_1","first-page":"1407","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Sun Peng","year":"2021","unstructured":"Peng Sun, Wenhu Zhang, Huanyu Wang, Songyuan Li, and Xi Li. 2021. Deep RGB-D saliency detection with depth-sensitive attention and automatic multi-modal fusion. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1407\u20131417."},{"issue":"1","key":"e_1_3_1_57_1","doi-asserted-by":"crossref","first-page":"4553","DOI":"10.1038\/s41467-020-18360-5","article-title":"Accelerating eye movement research via accurate and affordable smartphone eye tracking","volume":"11","author":"Valliappan Nachiappan","year":"2020","unstructured":"Nachiappan Valliappan, Na Dai, Ethan Steinberg, Junfeng He, Kantwon Rogers, Venky Ramachandran, Pingmei Xu, Mina Shojaeizadeh, Li Guo, Kai Kohlhoff, and Vidhya Navalpakkam. 2020. Accelerating eye movement research via accurate and affordable smartphone eye tracking. Nature Communications 11, 1 (2020), 4553.","journal-title":"Nature Communications"},{"key":"e_1_3_1_58_1","first-page":"4894","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Wenguan","year":"2018","unstructured":"Wenguan Wang, Jianbing Shen, Fang Guo, Ming-Ming Cheng, and Ali Borji. 2018. Revisiting video saliency: A large-scale benchmark and a new model. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 4894\u20134903."},{"issue":"1","key":"e_1_3_1_59_1","doi-asserted-by":"crossref","first-page":"220","DOI":"10.1109\/TPAMI.2019.2924417","article-title":"Revisiting video saliency prediction in the deep learning era","volume":"43","author":"Wang Wenguan","year":"2019","unstructured":"Wenguan Wang, Jianbing Shen, Jianwen Xie, Ming-Ming Cheng, Haibin Ling, and Ali Borji. 2019. Revisiting video saliency prediction in the deep learning era. IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 1 (2019), 220\u2013237.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"4","key":"e_1_3_1_60_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3626235","article-title":"Diffusion models: A comprehensive survey of methods and applications","volume":"56","author":"Yang Ling","year":"2023","unstructured":"Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2023. Diffusion models: A comprehensive survey of methods and applications. ACM Computing Surveys 56, 4 (2023), 1\u201339.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_1_61_1","doi-asserted-by":"crossref","first-page":"1383","DOI":"10.1145\/3343031.3350990","volume-title":"Proceedings of the 27th ACM International Conference on Multimedia","author":"Yang Sheng","year":"2019","unstructured":"Sheng Yang, Qiuping Jiang, Weisi Lin, and Yongtao Wang. 2019. SGDNet: An end-to-end saliency-guided deep neural network for no-reference image quality assessment. In Proceedings of the 27th ACM International Conference on Multimedia, 1383\u20131391."},{"key":"e_1_3_1_62_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TIM.2021.3108538","article-title":"A measurement for distortion induced saliency variation in natural images","volume":"70","author":"Yang Xiaohan","year":"2021","unstructured":"Xiaohan Yang, Fan Li, and Hantao Liu. 2021. A measurement for distortion induced saliency variation in natural images. IEEE Transactions on Instrumentation and Measurement 70 (2021), 1\u201314.","journal-title":"IEEE Transactions on Instrumentation and Measurement"},{"key":"e_1_3_1_63_1","unstructured":"Hu Ye Jun Zhang Sibo Liu Xiao Han and Wei Yang. 2023. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv:2308.06721. https:\/\/arxiv.org\/abs\/2308.06721"},{"issue":"9","key":"e_1_3_1_64_1","first-page":"5761","article-title":"Uncertainty inspired RGB-D saliency detection","volume":"44","author":"Zhang Jing","year":"2021","unstructured":"Jing Zhang, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemeh Saleh, Sadegh Aliakbarian, and Nick Barnes. 2021. Uncertainty inspired RGB-D saliency detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 9 (2021), 5761\u20135779.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_65_1","first-page":"3836","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhang Lvmin","year":"2023","unstructured":"Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023b. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 3836\u20133847."},{"issue":"10","key":"e_1_3_1_66_1","doi-asserted-by":"crossref","first-page":"4270","DOI":"10.1109\/TIP.2014.2346028","article-title":"VSI: A visual saliency-induced index for perceptual image quality assessment","volume":"23","author":"Zhang Lin","year":"2014","unstructured":"Lin Zhang, Ying Shen, and Hongyu Li. 2014. VSI: A visual saliency-induced index for perceptual image quality assessment. IEEE Transactions on Image Processing 23, 10 (2014), 4270\u20134281.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_67_1","first-page":"1","volume-title":"Proceedings of the ACM SIGGRAPH 2023 Conference Proceedings","author":"Zhang Yunxiang","year":"2023","unstructured":"Yunxiang Zhang, Kenneth Chen, and Qi Sun. 2023a. Toward optimized VR\/AR ergonomics: Modeling and predicting user neck muscle contraction. In Proceedings of the ACM SIGGRAPH 2023 Conference Proceedings, 1\u201312."},{"key":"e_1_3_1_68_1","first-page":"488","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Zhang Ziheng","year":"2018","unstructured":"Ziheng Zhang, Yanyu Xu, Jingyi Yu, and Shenghua Gao. 2018. Saliency detection in 360 videos. In Proceedings of the European Conference on Computer Vision (ECCV), 488\u2013503."},{"key":"e_1_3_1_69_1","unstructured":"Shihao Zhao Dongdong Chen Yen-Chun Chen Jianmin Bao Shaozhe Hao Lu Yuan and Kwan-Yee K. Wong. 2023. Uni-ControlNet: All-in-one control to text-to-image diffusion models. arXiv:2305.16322. https:\/\/arxiv.org\/abs\/2305.16322"}],"container-title":["ACM Transactions on Applied Perception"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3694969","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3694969","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:07Z","timestamp":1750295887000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3694969"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,31]]},"references-count":68,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,10,31]]}},"alternative-id":["10.1145\/3694969"],"URL":"https:\/\/doi.org\/10.1145\/3694969","relation":{},"ISSN":["1544-3558","1544-3965"],"issn-type":[{"type":"print","value":"1544-3558"},{"type":"electronic","value":"1544-3965"}],"subject":[],"published":{"date-parts":[[2024,10,31]]},"assertion":[{"value":"2024-08-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-28","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-11-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}