{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T06:04:38Z","timestamp":1785305078685,"version":"3.55.0"},"reference-count":78,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T00:00:00Z","timestamp":1782172800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T00:00:00Z","timestamp":1782172800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100002341","name":"Academy of Finland","doi-asserted-by":"crossref","award":["353139"],"award-info":[{"award-number":["353139"]}],"id":[{"id":"10.13039\/501100002341","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100002341","name":"Academy of Finland","doi-asserted-by":"crossref","award":["362409"],"award-info":[{"award-number":["362409"]}],"id":[{"id":"10.13039\/501100002341","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Existing multimodal audio generation models often lack precise user control, which limits their applicability in professional Foley workflows. In particular, these models focus on the entire video and do not provide precise methods for prioritizing a specific object within a scene, generating unnecessary background sounds, or focusing on the wrong objects. To address this gap, we introduce the novel task of video object segmentation-aware audio generation, which explicitly conditions sound synthesis on object-level segmentation maps. We present SAGANet, a new multimodal generative model that enables controllable audio generation for musical instruments by leveraging visual segmentation masks along with video and textual cues. Our model provides users with fine-grained and visually localized control over audio generation. To support this task and further research on segmentation-aware Foley, we propose Segmented Music Solos, a benchmark dataset of musical instrument performance videos with segmentation information. Our method demonstrates substantial improvements over current state-of-the-art methods and sets a new standard for controllable, high-fidelity Foley synthesis for musical audio. Code, samples, and Segmented Music Solos are available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"http:\/\/saganet.notion.site\" ext-link-type=\"uri\">http:\/\/saganet.notion.site<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s11263-026-02911-2","type":"journal-article","created":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T07:46:24Z","timestamp":1782200784000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Video Object Segmentation-Aware Audio Generation"],"prefix":"10.1007","volume":"134","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-2361-8067","authenticated-orcid":false,"given":"Ilpo","family":"Viertola","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Vladimir","family":"Iashin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Esa","family":"Rahtu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,6,23]]},"reference":[{"key":"2911_CR1","doi-asserted-by":"crossref","unstructured":"Alayrac, J- B., Donahue, J., Luc, P., Miech, A., Barr, I., & Hasson, Y.. others (2022). Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems (neurips) (pp. 23716\u201323736).","DOI":"10.52202\/068431-1723"},{"key":"2911_CR2","doi-asserted-by":"crossref","unstructured":"Chang, H., Zhang, H., Jiang, L., Liu, C., & Freeman, W. T. (2022). Maskgit: Masked generative image transformer. Computer vision and pattern recognition conference (cvpr) (pp. 11315\u201311325)","DOI":"10.1109\/CVPR52688.2022.01103"},{"key":"2911_CR3","doi-asserted-by":"crossref","unstructured":"Chen, C., Peng, P., Baid, A., Xue, Z., Hsu, W. N., Harwath, D., & Grauman, K. (2024). Action2sound: Ambient-aware generation of action sounds from egocentric videos. European conference on computer vision (eccv) (pp. 277\u2013295)","DOI":"10.1007\/978-3-031-72897-6_16"},{"key":"2911_CR4","doi-asserted-by":"crossref","unstructured":"Chen, H., Xie, W., Vedaldi, A., & Zisserman, A. (2020). Vggsound: A large-scale audio-visual dataset. Proceedings of international conference on acoustics, speech and signal processing (icassp) (pp. 721\u2013725)","DOI":"10.1109\/ICASSP40776.2020.9053174"},{"key":"2911_CR5","doi-asserted-by":"crossref","unstructured":"Chen, Z., Seetharaman, P., Russell, B., Nieto, O., Bourgin, D., Owens, A., & Salamon, J. (2025). Video-guided foley sound generation with multimodal controls. Computer vision and pattern recognition conference (cvpr) (pp. 18770\u201318781)","DOI":"10.1109\/CVPR52734.2025.01749"},{"key":"2911_CR6","doi-asserted-by":"crossref","unstructured":"Cheng, H. K., Ishii, M., Hayakawa, A., Shibuya, T., Schwing, A., & Mitsufuji, Y. (2025). Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis. Computer vision and pattern recognition conference (cvpr) (pp. 28901\u201328911)","DOI":"10.1109\/CVPR52734.2025.02691"},{"key":"2911_CR7","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L- J., Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. Computer vision and pattern recognition conference (cvpr) (pp. 248\u2013255).","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"2911_CR8","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T. & others (2021). An image is worth 16x16 words: Transformers for image recognition at scale. International conference on learning representations (iclr)."},{"key":"2911_CR9","unstructured":"Esser, P., Kulal, S., Blattmann, A., Entezari, R., M\u00fcller, J., Saini, H. & others (2024). Scaling rectified flow transformers for high-resolution image synthesis. International conference on machine learning (icml)."},{"key":"2911_CR10","doi-asserted-by":"crossref","unstructured":"Fragkiadaki, K., Arbelaez, P., Felsen, P., & Malik, J. (2015). Learning to segment moving objects in videos. International conference on computer vision (cvpr) (pp. 4083\u20134090)","DOI":"10.1109\/CVPR.2015.7299035"},{"key":"2911_CR11","doi-asserted-by":"crossref","unstructured":"Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., & Ritter, M. (2017). Audio set: An ontology and human-labeled dataset for audio events. Proceedings of international conference on acoustics, speech and signal processing (icassp) (pp. 776\u2013780)","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"2911_CR12","doi-asserted-by":"crossref","unstructured":"Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K. V., Joulin, A., & Misra, I. (2023). Imagebind: One embedding space to bind them all. Computer vision and pattern recognition conference (cvpr) (pp. 15180\u201315190)","DOI":"10.1109\/CVPR52729.2023.01457"},{"key":"2911_CR13","doi-asserted-by":"crossref","unstructured":"Gong, Y., Chung, Y. A., & Glass, J. (2021). AST: Audio Spectrogram Transformer. Interspeech (pp. 571\u2013575)","DOI":"10.21437\/Interspeech.2021-698"},{"key":"2911_CR14","unstructured":"Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., & Hochreiter, S. (2017). Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems (neurips) (pp. 6626\u20136637)"},{"key":"2911_CR15","doi-asserted-by":"publisher","unstructured":"Ho, J., & Salimans, T. (2022). Classifier-Free Diffusion Guidance. arXiv e-prints, arXiv:2207.12598, https:\/\/doi.org\/10.48550\/arXiv.2207.12598","DOI":"10.48550\/arXiv.2207.12598"},{"key":"2911_CR16","doi-asserted-by":"crossref","unstructured":"Hou, S., Liu, S., Yuan, R., Xue, W., Shan, Y., Zhao, M., & Zhang, C. (2025). Editing music with melody and text: Using controlnet for diffusion transformer. Proceedings of international conference on acoustics, speech and signal processing (icassp) (pp. 1\u20135)","DOI":"10.1109\/ICASSP49660.2025.10890309"},{"key":"2911_CR17","unstructured":"Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., & Chen, W. (2022). Lora: Low-rank adaptation of large language models International conference on learning representations (iclr."},{"key":"2911_CR18","doi-asserted-by":"crossref","unstructured":"Iashin, V., & Rahtu, E. (2021). Taming visually guided sound generation. British machine vision conference (bmvc)","DOI":"10.5244\/C.35.336"},{"key":"2911_CR19","unstructured":"Iashin, V., Xie, W., Rahtu, E., & Zisserman, A. (2022). Sparse in space and time: Audio-visual synchronisation with trainable selectors. British machine vision conference (bmvc)"},{"key":"2911_CR20","doi-asserted-by":"crossref","unstructured":"Iashin, V., Xie, W., Rahtu, E., & Zisserman, A. (2024). Synchformer: Efficient synchronization from sparse cues. Proceedings of international conference on acoustics, speech and signal processing (icassp) (pp. 5325\u20135329)","DOI":"10.1109\/ICASSP48485.2024.10448489"},{"key":"2911_CR21","doi-asserted-by":"crossref","unstructured":"Jeong, Y., Kim, Y., Chun, S., & Lee, J. (2025). Read, watch and scream! sound generation from text and video. Aaai conference on artificial intelligence (aaai) (pp. 17590\u201317598)","DOI":"10.1609\/aaai.v39i17.33934"},{"key":"2911_CR22","unstructured":"Jocher, G., Chaurasia, A., Stoken, A., Borovec, J., NanoCode012, Kwon, Y. & Jain, M. (2022). ultralytics\/yolov5: v7.0 - YOLOv5 SOTA Realtime Instance Segmentation. Zenodo."},{"key":"2911_CR23","unstructured":"Kim, C.D., Kim, B., Lee, H., & Kim, G. (2019). Audiocaps: Generating captions for audios in the wild. Proceedings of conference of the north american chapter of the association for computational linguistics: Human language technologies (pp. 119\u2013132)."},{"key":"2911_CR24","doi-asserted-by":"publisher","unstructured":"Kim, G., Martinez, A., Su, Y- C., Jou, B., Lezama, J., Gupta, A. & Somandepalli, K. (2024). A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual Generation. arXiv e-prints, arXiv:2405.13762, https:\/\/doi.org\/10.48550\/arXiv.2405.13762","DOI":"10.48550\/arXiv.2405.13762"},{"key":"2911_CR25","unstructured":"Kingma, D. P., & Welling, M. (2013). Auto-encoding variational bayes. International conference on learning representations (iclr)"},{"key":"2911_CR26","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L. & others (2023). Segment anything. Computer vision and pattern recognition conference (cvpr) (pp. 4015\u20134026).","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"2911_CR27","doi-asserted-by":"crossref","unstructured":"Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., & Plumbley, M.D. (2020). Panns: Large-scale pretrained audio neural networks for audio pattern recognition. Transactions on Audio, Speech, and Language Processing (TASLPRO), 2880\u20132894","DOI":"10.1109\/TASLP.2020.3030497"},{"key":"2911_CR28","doi-asserted-by":"crossref","unstructured":"Koutini, K., Schl\u00fcter, J., Eghbal-zadeh, H., & Widmer, G. (2022). Efficient training of audio transformers with patchout. Interspeech 2022, 23rd annual conference of the international speech communication association, incheon, korea, 18-22 september 2022 (pp. 2753\u20132757). ISCA.","DOI":"10.21437\/Interspeech.2022-227"},{"key":"2911_CR29","unstructured":"Labs, B.F., Batifol, S., Blattmann, A., Boesel, F., Consul, S., Diagne, C. & Smith, L. (2025). Flux.1 kontext: Flow matching for in-context image generation and editing in latent space. https:\/\/arxiv.org\/abs\/2506.15742"},{"key":"2911_CR30","unstructured":"Lee, S. G., Ping, W., Ginsburg, B., Catanzaro, B., & Yoon, S. (2023). BigVGAN: A universal neural vocoder with large-scale training International conference on learning representations (iclr."},{"issue":"2","key":"2911_CR31","doi-asserted-by":"publisher","first-page":"522","DOI":"10.1109\/TMM.2018.2856090","volume":"21","author":"B Li","year":"2018","unstructured":"Li, B., Liu, X., Dinesh, K., Duan, Z., & Sharma, G. (2018). Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications. Transactions on Multimedia, 21(2), 522\u2013535.","journal-title":"Transactions on Multimedia"},{"key":"2911_CR32","unstructured":"Li, J., Li, D., Xiong, C., & Hoi, S. (2022). Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. International conference on machine learning (icml) (pp. 12888\u201312900)."},{"key":"2911_CR33","unstructured":"Li, T., Huang, B., Zhuang, X., Jia, D., Chen, J., Wang, Y. & Wang, Y. (2025). Sounding that object: Interactive object-aware image to audio generation. Icml."},{"key":"2911_CR34","doi-asserted-by":"publisher","unstructured":"Lian, L., Ding, Y., Ge, Y., Liu, S., Mao, H., Li, B. & Cui, Y. (2025). Describe Anything: Detailed Localized Image and Video Captioning. arXiv e-prints, arXiv:2504.16072, https:\/\/doi.org\/10.48550\/arXiv.2504.16072","DOI":"10.48550\/arXiv.2504.16072"},{"key":"2911_CR35","unstructured":"Likert, R. (1932). A technique for the measurement of attitudes. Archives of psychology"},{"key":"2911_CR36","unstructured":"Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., & Le, M. (2022). Flow matching for generative modeling International conference on learning representations (iclr."},{"key":"2911_CR37","unstructured":"Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D. & Plumbley, M.D. (2023). AudioLDM: Text-to-audio generation with latent diffusion models. International conference on machine learning (icml) (p.21450-21474)."},{"key":"2911_CR38","doi-asserted-by":"crossref","unstructured":"Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J. & others (2024). Grounding dino: Marrying dino with grounded pre-training for open-set object detection. European conference on computer vision (eccv) (pp. 38\u201355).","DOI":"10.1007\/978-3-031-72970-6_3"},{"key":"2911_CR39","doi-asserted-by":"publisher","unstructured":"Liu, X., Gong, C., & Liu, Q. (2022). Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. arXiv e-prints, arXiv:2209.03003, https:\/\/doi.org\/10.48550\/arXiv.2209.03003","DOI":"10.48550\/arXiv.2209.03003"},{"key":"2911_CR40","doi-asserted-by":"crossref","unstructured":"Liu, X., Su, K., & Shlizerman, E. (2024). Tell what you hear from what you see - video to audio generation through text. A.\u00a0Globerson et\u00a0al. (Eds.), Advances in neural information processing systems (neurips) (Vol.\u00a037, pp. 101337\u2013101366). Curran Associates, Inc.","DOI":"10.52202\/079017-3213"},{"key":"2911_CR41","unstructured":"Loshchilov, I., & Hutter, F. (2017). Decoupled weight decay regularization International conference on learning representations (iclr."},{"key":"2911_CR42","doi-asserted-by":"crossref","unstructured":"Luo, S., Yan, C., Hu, C., & Zhao, H. (2023). Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models. Advances in neural information processing systems (neurips) (pp. 48855\u201348876)","DOI":"10.52202\/075280-2121"},{"key":"2911_CR43","doi-asserted-by":"crossref","unstructured":"Mei, X., Meng, C., Liu, H., Kong, Q., Ko, T., Zhao, C. & Wang, W. (2024). Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research. Transactions on Audio, Speech, and Language Processing","DOI":"10.1109\/TASLP.2024.3419446"},{"key":"2911_CR44","doi-asserted-by":"crossref","unstructured":"Mei, X., Nagaraja, V., Le Lan, G., Ni, Z., Chang, E., Shi, Y., & Chandra, V. (2024). Foleygen: Visually-guided audio generation. International workshop on machine learning for signal processing (mlsp) (pp. 1\u20136)","DOI":"10.1109\/MLSP58920.2024.10734721"},{"key":"2911_CR45","doi-asserted-by":"publisher","unstructured":"Mo, S., Shi, J., & Tian, Y. (2024). Text-to-Audio Generation Synchronized with Videos. arXiv e-prints, arXiv:2403.07938, https:\/\/doi.org\/10.48550\/arXiv.2403.07938","DOI":"10.48550\/arXiv.2403.07938"},{"key":"2911_CR46","doi-asserted-by":"crossref","unstructured":"Montesinos, J. F., Slizovskaia, O., & Haro, G. (2020). Solos: A dataset for audio-visual music analysis. International workshop on multimedia signal processing (mmsp) (pp. 1\u20136)","DOI":"10.1109\/MMSP48831.2020.9287124"},{"key":"2911_CR47","doi-asserted-by":"crossref","unstructured":"Pascual, S., Yeh, C., Tsiamas, I., & Serr\u00e0, J. (2024). Masked generative video-to-audio transformers with enhanced synchronicity. European conference on computer vision (eccv) (pp. 247\u2013264)","DOI":"10.1007\/978-3-031-73021-4_15"},{"key":"2911_CR48","doi-asserted-by":"crossref","unstructured":"Peebles, W., & Xie, S. (2023). Scalable diffusion models with transformers. Computer vision and pattern recognition conference (cvpr (pp. 4195\u20134205)","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"2911_CR49","doi-asserted-by":"publisher","unstructured":"Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A. & Du, Y. (2024). Movie Gen: A Cast of Media Foundation Models. arXiv e-prints, arXiv:2410.13720, https:\/\/doi.org\/10.48550\/arXiv.2410.13720","DOI":"10.48550\/arXiv.2410.13720"},{"key":"2911_CR50","doi-asserted-by":"publisher","unstructured":"Pont-Tuset, J., Perazzi, F., Caelles, S., Arbel\u00e1ez, P., Sorkine-Hornung, A., & Van Gool, L. (2017). The 2017 DAVIS Challenge on Video Object Segmentation. arXiv e-prints, arXiv:1704.00675, https:\/\/doi.org\/10.48550\/arXiv.1704.00675","DOI":"10.48550\/arXiv.1704.00675"},{"key":"2911_CR51","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S. & others (2021). Learning transferable visual models from natural language supervision. International conference on machine learning (icml) (pp. 8748\u20138763)."},{"key":"2911_CR52","unstructured":"Radiocommunication, I. (2007). Relative timing of sound and vision for broadcasting"},{"key":"2911_CR53","doi-asserted-by":"publisher","unstructured":"Ravi, N., Gabeur, V., Hu, Y- T., Hu, R., Ryali, C., Ma, T. & Feichtenhofer, C. (2024). SAM 2: Segment Anything in Images and Videos. arXiv e-prints, arXiv:2408.00714, https:\/\/doi.org\/10.48550\/arXiv.2408.00714","DOI":"10.48550\/arXiv.2408.00714"},{"key":"2911_CR54","doi-asserted-by":"publisher","unstructured":"Ren, T., Jiang, Q., Liu, S., Zeng, Z., Liu, W., Gao, H. & Zhang, L. (2024). Grounding DINO 1.5: Advance the \u201cEdge\u201d of Open-Set Object Detection. arXiv e-prints, arXiv:2405.10300, https:\/\/doi.org\/10.48550\/arXiv.2405.10300","DOI":"10.48550\/arXiv.2405.10300"},{"key":"2911_CR55","doi-asserted-by":"publisher","unstructured":"Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H. & Zhang, L. (2024). Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks. arXiv e-prints, arXiv:2401.14159, https:\/\/doi.org\/10.48550\/arXiv.2401.14159","DOI":"10.48550\/arXiv.2401.14159"},{"key":"2911_CR56","doi-asserted-by":"crossref","unstructured":"Ruan, L., Ma, Y., Yang, H., He, H., Liu, B., Fu, J., & Guo, B. (2023). Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation. Computer vision and pattern recognition conference (cvpr) (pp. 10219\u201310228)","DOI":"10.1109\/CVPR52729.2023.00985"},{"key":"2911_CR57","unstructured":"Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., & Chen, X. (2016). Improved techniques for training gans. Advances in neural information processing systems (neurips)"},{"key":"2911_CR58","doi-asserted-by":"crossref","unstructured":"Sheffer, R., A., & Y. (2023). I hear your true colors: Image guided audio generation. Proceedings of international conference on acoustics, speech and signal processing (icassp) (pp. 1\u20135)","DOI":"10.1109\/ICASSP49357.2023.10096023"},{"key":"2911_CR59","unstructured":"Song, K., Tan, X., Qin, T., Lu, J., & Liu, T. Y. (2020). Mpnet: Masked and permuted pre-training for language understanding. Advances in neural \u0131nformation processing systems (neurips) (pp. 16857\u201316867)"},{"issue":"3","key":"2911_CR60","doi-asserted-by":"publisher","first-page":"185","DOI":"10.1121\/1.1915893","volume":"8","author":"SS Stevens","year":"1937","unstructured":"Stevens, S. S., Volkmann, J., & Newman, E. B. (1937). A scale for the measurement of the psychological magnitude pitch. The Journal of the Acoustical Society of America, 8(3), 185\u2013190.","journal-title":"The Journal of the Acoustical Society of America"},{"key":"2911_CR61","unstructured":"Su, K., Liu, X., & Shlizerman, E. (2024). From vision to audio and beyond: a unified model for audio-visual representation and generation nternational conference on machine learning (icml."},{"key":"2911_CR62","doi-asserted-by":"crossref","unstructured":"Tang, Z., Yang, Z., Khademi, M., Liu, Y., Zhu, C., & Bansal, M. (2024). Codi-2: In-context interleaved and interactive any-to-any generation. Computer vision and pattern recognition conference (cvpr) (pp. 27425\u201327434)","DOI":"10.1109\/CVPR52733.2024.02589"},{"key":"2911_CR63","doi-asserted-by":"crossref","unstructured":"Tang, Z., Yang, Z., Zhu, C., Zeng, M., & Bansal, M. (2023). Any-to-any generation via composable diffusion. Advances in neural information processing systems (neurips) (pp. 16083\u201316099)","DOI":"10.52202\/075280-0707"},{"key":"2911_CR64","doi-asserted-by":"publisher","unstructured":"Tong, A., Fatras, K., Malkin, N., Huguet, G., Zhang, Y., Rector-Brooks, J. & Bengio, Y. (2023). Improving and generalizing flow-based generative models with minibatch optimal transport. arXiv e-prints, arXiv:2302.00482, https:\/\/doi.org\/10.48550\/arXiv.2302.00482","DOI":"10.48550\/arXiv.2302.00482"},{"key":"2911_CR65","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems (neurips)"},{"key":"2911_CR66","doi-asserted-by":"crossref","unstructured":"Viertola, I., Iashin, V., & Rahtu, E. (2025). Temporally aligned audio for video with autoregression. Proceedings of international conference on acoustics, speech and signal processing (icassp) (pp. 1\u20135)","DOI":"10.1109\/ICASSP49660.2025.10890587"},{"key":"2911_CR67","doi-asserted-by":"crossref","unstructured":"Wang, H., Ma, J., Pascual, S., Cartwright, R., & Cai, W. (2024). V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models. Aaai conference on artificial intelligence (aaai) (pp. 15492\u201315501)","DOI":"10.1609\/aaai.v38i14.29475"},{"key":"2911_CR68","doi-asserted-by":"crossref","unstructured":"Wang, Y., Guo, W., Huang, R., Huang, J., Wang, Z., You, F. & Zhao, Z. (2024). Frieren: Efficient video-to-audio generation network with rectified flow matching. Advances in neural information processing systems (neurips) (pp. 128118\u2013128138).","DOI":"10.52202\/079017-4068"},{"key":"2911_CR69","doi-asserted-by":"crossref","unstructured":"Wu, S- L., Donahue, C., Watanabe, S., & Bryan, N.J. (2024). Music controlnet: Multiple time-varying controls for music generation. Transactions on Audio, Speech, and Language Processing,32, 2692\u20132703.","DOI":"10.1109\/TASLP.2024.3399026"},{"key":"2911_CR70","doi-asserted-by":"publisher","unstructured":"Xiao, B., Wu, H., Xu, W., Dai, X., Hu, H., Lu, Y. & Yuan, L. (2023). Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks. arXiv e-prints,arXiv:2311.06242, https:\/\/doi.org\/10.48550\/arXiv.2311.06242","DOI":"10.48550\/arXiv.2311.06242"},{"key":"2911_CR71","doi-asserted-by":"crossref","unstructured":"Xing, Y., He, Y., Tian, Z., Wang, X., & Chen, Q. (2024). Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners. Conference on computer vision and pattern recognition (cvpr) (pp. 7151\u20137161)","DOI":"10.1109\/CVPR52733.2024.00683"},{"key":"2911_CR72","doi-asserted-by":"publisher","unstructured":"Xu, M., Li, C., Tu, X., Ren, Y., Chen, R., Gu, Y. & Yu, D. (2024). Video-to-Audio Generation with Hidden Alignment. arXiv e-prints, arXiv:2407.07464, https:\/\/doi.org\/10.48550\/arXiv.2407.07464","DOI":"10.48550\/arXiv.2407.07464"},{"key":"2911_CR73","doi-asserted-by":"crossref","unstructured":"Zhang, L., Rao, A., & Agrawala, M. (2023). Adding conditional control to text-to-image diffusion models. Conference on computer vision and pattern recognition (cvpr) (pp. 3836\u20133847)","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"2911_CR74","doi-asserted-by":"crossref","unstructured":"Zhao, H., Gan, C., Ma, W. C., & Torralba, A. (2019). The sound of motions. International conference on computer vision (cvpr) (pp. 1735\u20131744)","DOI":"10.1109\/ICCV.2019.00182"},{"key":"2911_CR75","doi-asserted-by":"crossref","unstructured":"Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J. & Torralba, A. (2018). The sound of pixels. European conference on computer vision (eccv) (pp. 570\u2013586).","DOI":"10.1007\/978-3-030-01246-5_35"},{"key":"2911_CR76","unstructured":"Zhou, J., Shen, X., Wang, J., Zhang, J., Sun, W., Zhang, J. & Zhong, Y. (2024). Audio-visual segmentation with semantics. International Journal of Computer Vision (IJCV), 1\u201321"},{"key":"2911_CR77","doi-asserted-by":"crossref","unstructured":"Zhou, J., Wang, J., Zhang, J., Sun, W., Zhang, J., Birchfield, S. & Zhong, Y. (2022). Audio\u2013visual segmentation. European conference on computer vision (eccv) (pp. 386\u2013403).","DOI":"10.1007\/978-3-031-19836-6_22"},{"key":"2911_CR78","doi-asserted-by":"publisher","unstructured":"Ziv, A., Gat, I., Le Lan, G., Remez, T., Kreuk, F., D\u00e9fossez, A. & Adi, Y. (2024). Masked Audio Generation using a Single Non-Autoregressive Transformer. arXiv e-prints, arXiv:2401.04577, https:\/\/doi.org\/10.48550\/arXiv.2401.04577","DOI":"10.48550\/arXiv.2401.04577"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02911-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-026-02911-2","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02911-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T05:50:02Z","timestamp":1785304202000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-026-02911-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,23]]},"references-count":78,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["2911"],"URL":"https:\/\/doi.org\/10.1007\/s11263-026-02911-2","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,23]]},"assertion":[{"value":"15 January 2026","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 May 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 June 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"326"}}