{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T11:37:37Z","timestamp":1786621057134,"version":"3.56.0"},"reference-count":75,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T00:00:00Z","timestamp":1786579200000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T00:00:00Z","timestamp":1786579200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Computer\u2010Aided Design (CAD) boosts modern manufacturing, yet design reuse remains constrained by the absence of large, openly available CAD repositories with rich multi\u2010modal annotations suitable for search\/retrieval. Recent large\u2010scale efforts to annotate public datasets rely on hash\u2010based redundancy removal that leaves no semantic structure, and on captioning by Vision\u2010Language Models (VLMs) using rendered images alone, which struggles to capture geometric and procedural information. We introduce MM\u2010CAD, a multi\u2010modal CAD dataset designed to level\u2010up retrieval and retrieval\u2010augmented generation models for engineering geometry, comprising two complementary parts. MM\u2010CAD:A brings 33,816 unique CAD models from eleven widely used benchmark datasets under a common identifier scheme, with isometric renderings, point clouds, and humanly\u2010curated multi\u2010level text captions, and 4,376 real hand\u2010drawn user sketches among others. MM\u2010CAD:B curates 192,626 models from the 1M\u2010model ABC corpus through a seven\u2010stage pipeline centered on Manifold\u2010Aware Adaptive Sampling (MAAS), which organizes models into semantically coherent neighborhoods rather than merely removing duplicates, directly supplying the hard negatives that contrastive retrieval training requires. Every retained model is annotated through a metadata\u2010grounded pipeline that conditions caption generation on parsed construction sequences rather than rendered views alone, producing three\u2010level text descriptions, multi\u2010level contour sketches, a hierarchical application taxonomy, and photorealistic in\u2010context images that largely preserve source CAD geometry, a modality not previously available at this scale on CAD data. We further introduce a joint retrieval architecture that aligns sketch, text, image, B\u2010Rep, and point cloud encoders in a single latent space through Matryoshka\u2010nested contrastive objectives, establishing the first unified cross\u2010modal retrieval benchmark for large\u2010scale CAD.<\/jats:p>","DOI":"10.1111\/cgf.70523","type":"journal-article","created":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T11:11:01Z","timestamp":1786619461000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["MM\u2010CAD: A Multi\u2010Modal CAD Dataset and Benchmark for Cross\u2010Modal Geometric Learning"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-3838-0945","authenticated-orcid":false,"given":"Anush","family":"Bharathi","sequence":"first","affiliation":[{"name":"Indian Institute of Technology Madras  India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7825-6830","authenticated-orcid":false,"given":"Ananthakrishnan","family":"A","sequence":"additional","affiliation":[{"name":"Indian Institute of Technology Madras  India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0182-977X","authenticated-orcid":false,"given":"Ramanathan","family":"Muthuganapathy","sequence":"additional","affiliation":[{"name":"Indian Institute of Technology Madras  India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2026,8,13]]},"reference":[{"key":"e_1_2_9_2_2","doi-asserted-by":"publisher","DOI":"10.24132\/JWSCG.2025-11"},{"key":"e_1_2_9_3_2","unstructured":"AnilR. et al.:Gemini: A family of highly capable multimodal models 2025. URL:https:\/\/arxiv.org\/abs\/2312.11805 arXiv:2312.11805. 5"},{"key":"e_1_2_9_4_2","unstructured":"AlamM. F. AhmedF.: GenCAD: Image-conditioned computer-aided design generation with transformer-based contrastive representation and diffusion priors.arXiv preprint arXiv:2409.16294(2024). 12"},{"key":"e_1_2_9_5_2","unstructured":"AbbasA. TirumalaK. SimigD. GanguliS. MorcosA. S.: SemDeDup: Data-efficient learning at web-scale through semantic deduplication.arXiv preprint arXiv:2303.09540(2023). 2"},{"key":"e_1_2_9_6_2","unstructured":"Black Forest Labs:FLUX.2 [klein]: Towards interactive visual intelligence 2026.https:\/\/bfl.ai\/blog\/flux2-klein-towards-interactive-visual-intelligence. 2 8"},{"key":"e_1_2_9_7_2","doi-asserted-by":"crossref","unstructured":"ChenT. et al.:Img2cad: Conditioned 3d cad model generation from single image with structured visual geometry 2024. URL:https:\/\/arxiv.org\/abs\/2410.03417 arXiv:2410.03417. 12","DOI":"10.1109\/TII.2025.3584476"},{"key":"e_1_2_9_8_2","unstructured":"DosovitskiyA. et al.: An image is worth 16x16 words: Transformers for image recognition at scale. InInternational Conference on Learning Representations (ICLR)(2021). 10"},{"key":"e_1_2_9_9_2","doi-asserted-by":"crossref","unstructured":"DaiY. et al.: BRepFormer: Transformer-based B-rep geometric feature recognition. InProceedings of the 2025 International Conference on Multimedia Retrieval (ICMR)(2025). doi:10.1145\/3731715.3733283. 10","DOI":"10.1145\/3731715.3733283"},{"key":"e_1_2_9_10_2","unstructured":"DengL. et al.: BrepLLM: Native boundary representation understanding with large language models.arXiv preprint arXiv:2512.16413(2025). 3"},{"key":"e_1_2_9_11_2","unstructured":"EmundsC. PauenN. et al.: IFCNet: A benchmark dataset for IFC entity classification. InProceedings of the 28th International Workshop on Intelligent Computing in Engineering (EG-ICE)(2021). 2 3"},{"key":"e_1_2_9_12_2","unstructured":"GrillJ.-B. et al.: Bootstrap your own latent: A new approach to self-supervised learning. InAdvances in Neural Information Processing Systems (NeurIPS)(2020). 6"},{"key":"e_1_2_9_13_2","unstructured":"GattiP. et al.: Composite sketch+text queries for retrieving objects with elusive names and complex interactions. InAAAI(2024). 2 10"},{"key":"e_1_2_9_14_2","unstructured":"Google DeepMind: EmbeddingGemma: Powerful and lightweight text representations.arXiv preprint arXiv:2509.20354(2025). 10"},{"key":"e_1_2_9_15_2","doi-asserted-by":"crossref","unstructured":"HeK. et al.: Deep residual learning for image recognition. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)(2016). 6","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_9_16_2","unstructured":"HeidariN. et al.: Geometric deep learning for computer-aided design: A survey.arXiv preprint arXiv:2402.17695(2024). URL:https:\/\/arxiv.org\/abs\/2402.17695. 1"},{"key":"e_1_2_9_17_2","unstructured":"JayantiS. et al.: Developing an engineering shape benchmark for cad models.Computer-Aided Design(2006). URL:https:\/\/www.sciencedirect.com\/science\/article\/pii\/S001044850600100X"},{"key":"e_1_2_9_17_3","doi-asserted-by":"crossref","unstructured":"doi:10.1016\/j.cad.2006.06.007. 2 3","DOI":"10.1016\/j.cad.2006.06.007"},{"key":"e_1_2_9_18_2","unstructured":"JohnsonJ. DouzeM. J\u00e9gouH.:Billion-scale similarity search with gpus 2017. URL:https:\/\/arxiv.org\/abs\/1702.08734 arXiv:1702.08734. 6 11"},{"key":"e_1_2_9_19_2","unstructured":"KimS. et al.: A large-scale annotated mechanical components benchmark for classification and retrieval tasks with deep neural networks. InProceedings of the 16th European Conference on Computer Vision (ECCV)(2020). 2 3 7"},{"key":"e_1_2_9_20_2","doi-asserted-by":"crossref","unstructured":"KhanM. S. et al.: Text2CAD: Generating sequential CAD models from beginner-to-expert level text prompts. InAdvances in Neural Information Processing Systems (NeurIPS)(2024). 2 12","DOI":"10.52202\/079017-0242"},{"key":"e_1_2_9_21_2","doi-asserted-by":"crossref","unstructured":"KusupatiA. BhattG. RegeA. WallingfordM. SinhaA. RamanujanV. Howard-SnyderW. ChenK. KakadeS. JainP. FarhadiA.: Matryoshka representation learning. InAdvances in Neural Information Processing Systems (NeurIPS)(2022). 11","DOI":"10.52202\/068431-2192"},{"key":"e_1_2_9_22_2","unstructured":"KazhdanM. FunkhouserT. RusinkiewiczS.: Rotation invariant spherical harmonic representation of 3D shape descriptors. InProceedings of the Eurographics\/ACM SIGGRAPH Symposium on Geometry Processing (SGP)(2003) pp.156\u2013165. 6 9"},{"key":"e_1_2_9_23_2","doi-asserted-by":"crossref","unstructured":"KochS. MatveevA. JiangZ. WilliamsF. ArtemovA. BurnaevE. AlexaM. ZorinD. PanozzoD.: ABC: A big CAD model dataset for geometric deep learning. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2019). 2 3","DOI":"10.1109\/CVPR.2019.00983"},{"key":"e_1_2_9_24_2","doi-asserted-by":"crossref","unstructured":"LiM. et al.: Photo-sketching: Inferring contour drawings from images. InIEEE Winter Conference on Applications of Computer Vision (WACV)(2019). 2 4","DOI":"10.1109\/WACV.2019.00154"},{"key":"e_1_2_9_25_2","doi-asserted-by":"crossref","unstructured":"LiC. et al.: Free2cad: Parsing freehand drawings into cad commands.ACM Transactions on Graphics(2022). URL:https:\/\/doi.org\/10.1145\/3528223.3530133","DOI":"10.1145\/3528223.3530133"},{"key":"e_1_2_9_25_3","doi-asserted-by":"crossref","unstructured":"doi:10.1145\/3528223.3530133. 2 12","DOI":"10.1145\/3528223.3530133"},{"key":"e_1_2_9_26_2","first-page":"44860","volume-title":"Advances in Neural Information Processing Systems","author":"Liu M.","year":"2023"},{"key":"e_1_2_9_27_2","doi-asserted-by":"crossref","unstructured":"LiX. et al.: Cad translator: An effective drive for text to 3d parametric computer-aided design generative modeling. InProceedings of the 32nd ACM International Conference on Multimedia(2024). URL:https:\/\/doi.org\/10.1145\/3664647.3681549","DOI":"10.1145\/3664647.3681549"},{"key":"e_1_2_9_27_3","doi-asserted-by":"crossref","unstructured":"doi:10.1145\/3664647.3681549. 2","DOI":"10.1145\/3664647.3681549"},{"key":"e_1_2_9_28_2","doi-asserted-by":"crossref","unstructured":"LiJ. et al.:Cad-llama: Leveraging large language models for computer-aided design parametric 3d model generation 2025. URL:https:\/\/arxiv.org\/abs\/2505.04481 arXiv:2505.04481. 2","DOI":"10.1109\/CVPR52734.2025.01730"},{"key":"e_1_2_9_29_2","unstructured":"LiuY. et al.:B-repler: Language-guided editing of cad models 2025. URL:https:\/\/arxiv.org\/abs\/2508.10201 arXiv:2508.10201. 2"},{"key":"e_1_2_9_30_2","unstructured":"LiuY. et al.:Hola: B-rep generation using a holistic latent representation 2025. URL:https:\/\/arxiv.org\/abs\/2504.14257 arXiv:2504.14257. 3"},{"key":"e_1_2_9_31_2","doi-asserted-by":"crossref","first-page":"103926","DOI":"10.1016\/j.cad.2025.103926","article-title":"Cadinstruct: A multimodal dataset for natural language-guided cad program synthesis","volume":"188","author":"Lv C.","year":"2025","journal-title":"Computer-Aided Design"},{"key":"e_1_2_9_31_3","doi-asserted-by":"crossref","unstructured":"doi:https:\/\/doi.org\/10.1016\/j.cad.2025.103926. 2","DOI":"10.1016\/j.cad.2025.103926"},{"key":"e_1_2_9_32_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.cviu.2014.10.006","article-title":"A comparison of 3D shape retrieval methods based on a large-scale benchmark supporting multimodal queries","volume":"131","author":"Li B.","year":"2015","journal-title":"Computer Vision and Image Understanding"},{"key":"e_1_2_9_33_2","doi-asserted-by":"crossref","unstructured":"MandaB. et al.: CADSketchNet - an annotated sketch dataset for 3d cad model retrieval with deep neural networks.Computers & Graphics(2021). URL:https:\/\/doi.org\/10.1016\/j.cag.2021.07.001","DOI":"10.1016\/j.cag.2021.07.001"},{"key":"e_1_2_9_33_3","doi-asserted-by":"crossref","unstructured":"doi:10.1016\/j.cag.2021.07.001. 2","DOI":"10.1016\/j.cag.2021.07.001"},{"key":"e_1_2_9_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3055826"},{"key":"e_1_2_9_35_2","unstructured":"OpenAI et al.:Gpt-4 technical report 2024. URL:https:\/\/arxiv.org\/abs\/2303.08774 arXiv:2303.08774. 5"},{"key":"e_1_2_9_36_2","doi-asserted-by":"crossref","unstructured":"QiZ. DongR. ZhangS. GengH. HanC. GeZ. YiL. MaK.: ShapeLLM: Universal 3D object understanding for embodied interaction. InProceedings of the European Conference on Computer Vision (ECCV)(2024). 3 6","DOI":"10.1007\/978-3-031-72775-7_13"},{"key":"e_1_2_9_37_2","doi-asserted-by":"crossref","first-page":"104","DOI":"10.1016\/j.cag.2022.07.009","article-title":"SHREC'22 track: Sketch-based 3D shape retrieval in the wild","volume":"107","author":"Qin J.","year":"2022","journal-title":"Computers & Graphics"},{"key":"e_1_2_9_38_2","unstructured":"RadfordA. et al.: Learning transferable visual models from natural language supervision. InProceedings of the 38th International Conference on Machine Learning (ICML)(2021). 2"},{"key":"e_1_2_9_39_2","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1109\/SMI.2004.1314504","volume-title":"Proceedings of the Shape Modeling International 2004","author":"Shilane P.","year":"2004"},{"key":"e_1_2_9_40_2","doi-asserted-by":"crossref","unstructured":"SangkloyP. et al.: A sketch is worth a thousand words: Image retrieval with text and sketch. InProceedings of the European Conference on Computer Vision (ECCV)(2022). 11","DOI":"10.1007\/978-3-031-19839-7_15"},{"issue":"2","key":"e_1_2_9_41_2","doi-asserted-by":"crossref","first-page":"461","DOI":"10.1214\/aos\/1176344136","article-title":"Estimating the Dimension of a Model","volume":"6","author":"Schwarz G.","year":"1978","journal-title":"Annals of Statistics"},{"key":"e_1_2_9_42_2","doi-asserted-by":"crossref","unstructured":"SorscherB. GeirhosR. ShekharS. GanguliS. MorcosA. S.: Beyond neural scaling laws: Beating power law scaling via data pruning. InAdvances in Neural Information Processing Systems (NeurIPS)(2022). 2","DOI":"10.52202\/068431-1419"},{"key":"e_1_2_9_43_2","doi-asserted-by":"crossref","unstructured":"SennrichR. HaddowB. BirchA.: Neural machine translation of rare words with subword units. InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL)(2016). 9","DOI":"10.18653\/v1\/P16-1162"},{"key":"e_1_2_9_44_2","doi-asserted-by":"crossref","unstructured":"Simo-SerraE. et al.: Learning to simplify: Fully convolutional networks for rough sketch cleanup.ACM Transactions on Graphics(2016). URL:https:\/\/doi.org\/10.1145\/2897824.2925972","DOI":"10.1145\/2897824.2925972"},{"key":"e_1_2_9_44_3","doi-asserted-by":"crossref","unstructured":"doi:10.1145\/2897824.2925972. 4","DOI":"10.1145\/2897824.2925972"},{"issue":"11","key":"e_1_2_9_45_2","doi-asserted-by":"crossref","first-page":"2153","DOI":"10.1109\/TPAMI.2015.2408351","article-title":"An efficient algorithm for calculating the exact hausdorff distance","volume":"37","author":"Taha A. A.","year":"2015","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_2_9_46_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-019-41695-z"},{"key":"e_1_2_9_46_3","doi-asserted-by":"crossref","unstructured":"doi:10.1038\/s41598-019-41695-z. 8","DOI":"10.1038\/s41598-019-41695-z"},{"key":"e_1_2_9_47_2","doi-asserted-by":"crossref","unstructured":"UsamaM. et al.: NURBGen: High-fidelity text-to-CAD generation through LLM-driven NURBS modeling.arXiv preprint arXiv:2511.06194(2025). doi:10.48550\/arXiv.2511.06194. 2 5 12","DOI":"10.1609\/aaai.v40i12.37922"},{"key":"e_1_2_9_48_2","unstructured":"WuZ. et al.: 3D ShapeNets: A deep representation for volumetric shapes. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)(2015). 2 3"},{"issue":"4","key":"e_1_2_9_49_2","article-title":"Fusion 360 gallery: A dataset and environment for programmatic CAD construction from human design sequences","volume":"40","author":"Willis K. D. D.","year":"2021","journal-title":"ACM Transactions on Graphics (Proc. SIGGRAPH)"},{"key":"e_1_2_9_50_2","doi-asserted-by":"crossref","unstructured":"WuR. et al.: Deepcad: A deep generative network for computer-aided design models. InProceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)(2021). 2 3","DOI":"10.1109\/ICCV48922.2021.00670"},{"key":"e_1_2_9_51_2","unstructured":"WangW. et al.: InternVL3.5: Advancing open-source multimodal models in versatility reasoning and efficiency.arXiv preprint arXiv:2508.18265(2025). 8"},{"key":"e_1_2_9_52_2","doi-asserted-by":"crossref","unstructured":"WuJ. et al.: CMT: A cascade MAR with topology predictor for multimodal conditional CAD generation. InProceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)(2025). 2 5 12","DOI":"10.1109\/ICCV51701.2025.00659"},{"key":"e_1_2_9_53_2","doi-asserted-by":"crossref","unstructured":"WuS. KhasahmadiA. H. KatzM. JayaramanP. K. PuY. WillisK. LiuB.: CadVLM: Bridging language and vision in the generation of parametric CAD sketches. InProceedings of the European Conference on Computer Vision (ECCV)(2024). 12","DOI":"10.1007\/978-3-031-72897-6_21"},{"key":"e_1_2_9_54_2","doi-asserted-by":"crossref","unstructured":"WuZ. SongS. KhoslaA. YuF. ZhangL. TangX. XiaoJ.: 3D ShapeNets: A deep representation for volumetric shapes. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)(2015) pp.1912\u20131920. 2 3","DOI":"10.1109\/CVPR.2015.7298801"},{"issue":"5","key":"e_1_2_9_55_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3326362","article-title":"Dynamic graph CNN for learning on point clouds","volume":"38","author":"Wang Y.","year":"2019","journal-title":"ACM Transactions on Graphics"},{"key":"e_1_2_9_56_2","doi-asserted-by":"crossref","first-page":"12041","DOI":"10.1109\/TII.2024.3413358","article-title":"Parametric primitive analysis of CAD sketches with vision transformer","volume":"20","author":"Wang X.","year":"2024","journal-title":"IEEE Transactions on Industrial Informatics"},{"key":"e_1_2_9_57_2","doi-asserted-by":"crossref","unstructured":"XueL. et al.: ULIP: Learning a unified representation of language images and point clouds for 3D understanding. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2023). 6","DOI":"10.1109\/CVPR52729.2023.00120"},{"key":"e_1_2_9_58_2","unstructured":"XuJ. et al.:Cad-mllm: Unifying multimodality-conditioned cad generation with mllm 2025. URL:https:\/\/arxiv.org\/abs\/2411.04954 arXiv:2411.04954. 2"},{"key":"e_1_2_9_59_2","volume-title":"Proceedings of the SIGGRAPH Asia 2025 Conference Papers","author":"Xu X.","year":"2025"},{"key":"e_1_2_9_59_3","doi-asserted-by":"crossref","unstructured":"doi:10.1145\/3757377.3763814. 2 5","DOI":"10.1145\/3757377.3763814"},{"key":"e_1_2_9_60_2","unstructured":"XiaoS. LiuZ. ZhangP. MuennighoffN.: C-Pack: Packed resources for general Chinese embeddings.arXiv preprint arXiv:2309.07597(2023). 8"},{"key":"e_1_2_9_61_2","doi-asserted-by":"crossref","unstructured":"YuX. TangL. et al.: Point-bert: Pre-training 3d point cloud transformers with masked point modeling. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)(2022). 6","DOI":"10.1109\/CVPR52688.2022.01871"},{"key":"e_1_2_9_62_2","unstructured":"ZhouQ. et al.: Thingi10k: A dataset of 10 000 3d-printing models.ArXiv abs\/1605.04797(2016). URL:https:\/\/api.semanticscholar.org\/CorpusID:39867743. 3"},{"key":"e_1_2_9_63_2","unstructured":"ZhouQ.-Y. et al.:Open3d: A modern library for 3d data processing 2018. URL:https:\/\/arxiv.org\/abs\/1801.09847 arXiv:1801.09847. 4"},{"key":"e_1_2_9_64_2","doi-asserted-by":"crossref","unstructured":"ZhouS. et al.: Cadparser: A learning approach of sequence modeling for b-rep cad. InProceedings of the Thirty-Second International Joint Conference on Artificial Intelligence (IJCAI)(2023). URL:https:\/\/doi.org\/10.24963\/ijcai.2023\/200","DOI":"10.24963\/ijcai.2023\/200"},{"key":"e_1_2_9_64_3","doi-asserted-by":"crossref","unstructured":"doi:10.24963\/ijcai.2023\/200. 2 3","DOI":"10.24963\/ijcai.2023\/200"},{"key":"e_1_2_9_65_2","doi-asserted-by":"crossref","unstructured":"ZhangR. GuoZ. ZhangW. LiK. MiaoX. CuiB. QiaoY. GaoP. LiH.: PointCLIP: Point cloud understanding by CLIP. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2022). 2","DOI":"10.1109\/CVPR52688.2022.00836"},{"key":"e_1_2_9_66_2","doi-asserted-by":"crossref","unstructured":"ZhaiX. MustafaB. KolesnikovA. BeyerL.: Sigmoid loss for language image pre-training. InProceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)(2023). 10","DOI":"10.1109\/ICCV51070.2023.01100"},{"key":"e_1_2_9_67_2","doi-asserted-by":"crossref","unstructured":"ZhangL. RaoA. AgrawalaM.: Adding conditional control to text-to-image diffusion models. InProceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)(2023). 2","DOI":"10.1109\/ICCV51070.2023.00355"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70523","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70523","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70523","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T11:11:29Z","timestamp":1786619489000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70523"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8,13]]},"references-count":75,"alternative-id":["10.1111\/cgf.70523"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70523","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,8,13]]},"assertion":[{"value":"2026-08-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70523"}}