{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T06:51:07Z","timestamp":1782283867801,"version":"3.54.5"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T00:00:00Z","timestamp":1777593600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T00:00:00Z","timestamp":1778630400000},"content-version":"vor","delay-in-days":12,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004462","name":"Consiglio Nazionale Delle Ricerche","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004462","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2026,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The growing adoption of prompt-based generative models has raised concerns over the unauthorized use of proprietary data, as such models may memorize and replicate training content. To address this issue, we introduce ProCAP, a novel Membership Inference Attack approach based on a prompt-driven auditing framework.Given a proprietary dataset and a target generative model, ProCAP trains an auxiliary model to craft prompts that trigger the target model to produce outputs revealing potential violations of the proprietary data.Unlike current literature, ProCAP is automatic, fully black-box, model-agnostic, and designed to operate in settings with limited or no knowledge of the training process.To reduce the computational cost of training the prompt generator, we adopt an optimization strategy that filters high-loss samples, i.e., those less likely to have been memorized. Our approach can then \u201cspecialize\u201d the learning phase on the most informative data regions. We validate ProCAP across different scenarios, by using both real and synthetic data. Results demonstrate its effectiveness in recognizing unauthorized data usages with strong accuracy-efficiency trade-offs.<\/jats:p>","DOI":"10.1007\/s10994-026-07010-4","type":"journal-article","created":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T12:56:04Z","timestamp":1778676964000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Automated Membership Inference via Prompt-Based Attacks in Generative Models"],"prefix":"10.1007","volume":"115","author":[{"given":"Daniela","family":"Gallo","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Angelica","family":"Liguori","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ettore","family":"Ritacco","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Luca","family":"Caviglione","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fabrizio","family":"Durante","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Giuseppe","family":"Manco","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,5,13]]},"reference":[{"issue":"5","key":"7010_CR1","doi-asserted-by":"publisher","first-page":"792","DOI":"10.1214\/aop\/1176996548","volume":"2","author":"AA Balkema","year":"1974","unstructured":"Balkema, A. A., & de Haan, L. (1974). Residual life time at great age. The Annals of Probability, 2(5), 792\u2013804.","journal-title":"The Annals of Probability"},{"key":"7010_CR2","doi-asserted-by":"crossref","unstructured":"Carlini, N., Chien, S., Nasr, M., Song, S., Terzis, A., & Tramer, F. (2022). Membership inference attacks from first principles. In: 2022 IEEE symposium on security and privacy, IEEE.","DOI":"10.1109\/SP46214.2022.9833649"},{"key":"7010_CR3","unstructured":"Carlini, N., Hayes, J., Nasr, M., Jagielski, M., Sehwag, V., Tramer, F., & Wallace, E. (2023a). Extracting training data from diffusion models. In: USENIX security symposium."},{"key":"7010_CR4","unstructured":"Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., & Zhang, C. (2023b). Quantifying memorization across neural language models. In: The eleventh international conference on learning representations."},{"key":"7010_CR5","unstructured":"Carlini, N., Tram\u00e9r, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., & Raffel, C. (2021). Extracting training data from large language models. In: USENIX security symposium."},{"key":"7010_CR6","doi-asserted-by":"crossref","unstructured":"Chang, K., Cramer, M., Soni, S., & Bamman, D. (2023). Speak, memory: An archaeology of books known to ChatGPT\/GPT-4. In: Proceedings of the 2023 conference on empirical methods in natural language processing, pp. 7312\u20137327.","DOI":"10.18653\/v1\/2023.emnlp-main.453"},{"key":"7010_CR8","doi-asserted-by":"crossref","unstructured":"Chen, D., Yu, N., Zhang, Y., & Fritz, M. (2020). GAN-Leaks: A taxonomy of membership inference attacks against generative models. In: Proceedings of the 2020 ACM SIGSAC conference on computer and communications security, pp. 343\u2013362.","DOI":"10.1145\/3372297.3417238"},{"key":"7010_CR7","doi-asserted-by":"crossref","unstructured":"Chen, Z. (2024). Catch me if you can: Detecting unauthorized data use in training deep learning models. In: Proceedings of the 2024 on ACM SIGSAC conference on computer and communications security, pp. 5098\u20135100.","DOI":"10.1145\/3658644.3690858"},{"key":"7010_CR9","doi-asserted-by":"crossref","unstructured":"Deistler, M., & Scherrer, W. (2022). Time series models. Lecture Notes in Statistics.","DOI":"10.1007\/978-3-031-13213-1"},{"key":"7010_CR10","doi-asserted-by":"crossref","unstructured":"Devroye, L., Gy\u00f6rfi, L., &amp; Lugosi, G. (1996). A probabilistic theory of pattern recognition, stochastic modelling and applied probability, vol 31.","DOI":"10.1007\/978-1-4612-0711-5"},{"key":"7010_CR11","unstructured":"Eldan, R., & Li, Y. (2023). TinyStories: How small can language models be and still speak coherent English? arXiv:abs\/2305.07759."},{"key":"7010_CR12","doi-asserted-by":"crossref","unstructured":"Fischer, J. E. (2023). Generative AI considered harmful. In: Proceedings of the 5th international conference on conversational user interfaces, pp. 1\u20135.","DOI":"10.1145\/3571884.3603756"},{"issue":"11","key":"7010_CR13","doi-asserted-by":"publisher","first-page":"139","DOI":"10.1145\/3422622","volume":"63","author":"IJ Goodfellow","year":"2020","unstructured":"Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., & Bengio, Y. (2020). Generative adversarial networks. Communications of the ACM, 63(11), 139\u2013144.","journal-title":"Communications of the ACM"},{"key":"7010_CR14","unstructured":"Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. In: Proceedings of the 34th conference on neural information processing systems."},{"key":"7010_CR15","doi-asserted-by":"crossref","unstructured":"Huang, Z., Gong, N. Z., & Reiter, M. K. (2024). A general framework for data-use auditing of ML models. In: Proceedings of the 2024 on ACM SIGSAC conference on computer and communications security, pp. 1300\u20131314.","DOI":"10.1145\/3658644.3690226"},{"issue":"4","key":"7010_CR16","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3287051","volume":"2","author":"H Jin","year":"2018","unstructured":"Jin, H., Liu, M., Dodhia, K., Li, Y., Srivastava, G., Fredrikson, M., Agarwal, Y., & Hong, J. I. (2018). Why are they collecting my data? Inferring the purposes of network traffic in mobile apps. Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies, 2(4), 1\u201327.","journal-title":"Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies"},{"key":"7010_CR17","unstructured":"Kim, J., Kim, M., & Mozafari, B. (2023). Provable memorization capacity of transformers. In: The eleventh international conference on learning representations."},{"key":"7010_CR18","unstructured":"Kingma, D. P., & Welling, M. (2014). Auto-encoding variational bayes. In: Proceedings of the international conference on learning representations."},{"issue":"1","key":"7010_CR19","doi-asserted-by":"publisher","first-page":"38","DOI":"10.1214\/aos\/1193342380","volume":"1","author":"L LeCam","year":"1973","unstructured":"LeCam, L. (1973). Convergence of estimates under dimensionality restrictions. The Annals of Statistics, 1(1), 38\u201353.","journal-title":"The Annals of Statistics"},{"key":"7010_CR20","doi-asserted-by":"crossref","unstructured":"Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed ,A., Levy, O., Stoyanov, V., & Zettlemoyer, L. (2019). BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv:abs\/1910.13461.","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"7010_CR23","unstructured":"Li, B., Wei, Y., Fu, Y., Wang, Z., Li, Y., Zhang, J., Wang, R., & Zhang, T. (2024a). Towards reliable verification of unauthorized data usage in personalized text-to-image diffusion models. arXiv:abs\/2410.10437."},{"key":"7010_CR21","unstructured":"Li, H., Deng, G., Liu, Y., Wang, K., Li, Y., Zhang, T., Liu, Y., Xu, G., Xu, G., & Wang, H. (2024b). Digger: Detecting copyright content mis-usage in large language model training. arXiv:abs\/2401.00676."},{"key":"7010_CR22","doi-asserted-by":"crossref","unstructured":"Li, Z., Wang, C., Wang, S., & Cuiyun, G. (2023). Protecting intellectual property of large language model-based code generation APIs via watermarks. In: Proceedings of the 2023 ACM SIGSAC conference on computer and communications security, pp. 2336\u20132350.","DOI":"10.1145\/3576915.3623120"},{"key":"7010_CR24","doi-asserted-by":"crossref","unstructured":"Lin, C. Y., & Och, F. J. (2004). Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In: Proceedings of the 42nd annual meeting of the association for computational linguistics, pp. 605\u2013612.","DOI":"10.3115\/1218955.1219032"},{"key":"7010_CR25","doi-asserted-by":"crossref","unstructured":"Liu, Y., Huang, J., & Li, Y., et al. (2025). Generative AI model privacy: A survey. Artificial Intelligence Review,58(33).","DOI":"10.1007\/s10462-024-11024-6"},{"key":"7010_CR26","doi-asserted-by":"crossref","unstructured":"Maini, P., Jia, H., Papernot, N., & Dziedzic, N. (2024). LLM dataset inference: Did you train on my dataset? arXiv:abs\/2406.06443.","DOI":"10.52202\/079017-3941"},{"key":"7010_CR27","unstructured":"Maini, P., Yaghini, M., & Papernot, N. (2021). Dataset inference: Ownership resolution in machine learning. In: International conference on learning representations."},{"issue":"2","key":"7010_CR28","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3440754","volume":"54","author":"C Meurisch","year":"2021","unstructured":"Meurisch, C., & M\u00fchlh\u00e4user, M. (2021). Data protection in AI services: A survey. ACM Computing Surveys, 54(2), 1\u201338.","journal-title":"ACM Computing Surveys"},{"key":"7010_CR29","doi-asserted-by":"crossref","unstructured":"Nasr, M., Shokri, R., & Houmansadr, A. (2019). Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning. In: 2019 IEEE symposium on security and privacy, pp 739\u2013753.","DOI":"10.1109\/SP.2019.00065"},{"key":"7010_CR30","doi-asserted-by":"publisher","DOI":"10.1016\/j.cosrev.2020.100312","volume":"38","author":"MM Ogonji","year":"2020","unstructured":"Ogonji, M. M., Okeyo, G., & Wafula, J. M. (2020). A survey on privacy and security of internet of things. Computer Science Review, 38, Article 100312.","journal-title":"Computer Science Review"},{"key":"7010_CR31","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., & Chintala, S. (2019). PyTorch: An imperative style, high-performance deep learning library. In: 33rd Conference on neural information processing systems."},{"key":"7010_CR32","doi-asserted-by":"crossref","unstructured":"Prechelt, L. (2012). Early stopping \u2013 But when? In: Neural networks: Tricks of the trade: Second Edition. pp. 53\u201367","DOI":"10.1007\/978-3-642-35289-8_5"},{"key":"7010_CR33","volume-title":"Data science for business","author":"F Provost","year":"2013","unstructured":"Provost, F., & Fawcett, T. (2013). Data science for business. O\u2019Reilly."},{"key":"7010_CR34","doi-asserted-by":"crossref","unstructured":"Severin, I. C. (2020a). Head posture monitor based on 3 IMU sensors: Consideration toward healthcare application. In: Proceedings of the international conference on e-health and bioengineering, pp. 1\u20134.","DOI":"10.1109\/EHB50910.2020.9280106"},{"key":"7010_CR35","doi-asserted-by":"crossref","unstructured":"Severin, I. C. (2020b). The head posture system based on 3 inertial sensors and machine learning models: Offline analyze. In: Proceedings of the 3rd international seminar on research of information technology and intelligent systems, pp. 672\u2013676.","DOI":"10.1109\/ISRITI51436.2020.9315418"},{"key":"7010_CR36","doi-asserted-by":"crossref","unstructured":"Shi, Y., Sagduyu, Y. E., Davaslioglu, K., & Li, J. H. (2018). Active deep learning attacks under strict rate limitations for online API calls. In: 2018 IEEE international symposium on technologies for homeland security, pp. 1\u20136.","DOI":"10.1109\/THS.2018.8574124"},{"key":"7010_CR37","doi-asserted-by":"crossref","unstructured":"Shokri, R., Stronati, M., Song, C., & Shmatikov, V. (2017). Membership inference attacks against machine learning models. In: 2017 IEEE symposium on security and privacy, pp. 3\u201318.","DOI":"10.1109\/SP.2017.41"},{"key":"7010_CR38","doi-asserted-by":"crossref","unstructured":"Sirisuriya, S. D. S. (2023). Importance of web scraping as a data source for machine learning algorithms-review. In: 2023 IEEE 17th international conference on industrial and information systems, pp. 134\u2013139.","DOI":"10.1109\/ICIIS58898.2023.10253502"},{"key":"7010_CR39","first-page":"45","volume":"41","author":"BL Sobel","year":"2017","unstructured":"Sobel, B. L. (2017). Artificial intelligence\u2019s fair use crisis. Colum JL & Arts, 41, 45.","journal-title":"Colum JL & Arts"},{"key":"7010_CR40","unstructured":"Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., & Ganguli, S. (2015). Deep unsupervised learning using nonequilibrium thermodynamics. In: International conference on machine learning, pp. 2256\u20132265."},{"key":"7010_CR41","doi-asserted-by":"crossref","unstructured":"Somepalli, G., Singla, V., Goldblum, M., Geiping, J., & Goldstein., T. (2023). Diffusion art or digital forgery? Investigating data replication in diffusion models. In: Conference on computer vision and pattern recognition, pp. 6048\u20136058.","DOI":"10.1109\/CVPR52729.2023.00586"},{"issue":"1","key":"7010_CR42","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1145\/3606274.3606279","volume":"25","author":"R Tang","year":"2023","unstructured":"Tang, R., Feng, Q., Liu, N., Yang, F., & Hu, X. (2023). Did you train on my dataset? Towards public dataset protection with clean label backdoor watermarking. SIGKDD Explor Newsl, 25(1), 43\u201353.","journal-title":"SIGKDD Explor Newsl"},{"key":"7010_CR43","doi-asserted-by":"crossref","unstructured":"Tsybakov, A. B. (2008). Introduction to nonparametric estimation, 1st edn.","DOI":"10.1007\/978-0-387-79052-7_1"},{"key":"7010_CR44","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. & Polosukhin, I. (2017). Attention is all you need. In: Proceedings of the 31st conference on neural information processing systems, pp. 5998\u20136008."},{"issue":"4","key":"7010_CR45","doi-asserted-by":"publisher","first-page":"501","DOI":"10.1007\/s10687-020-00393-0","volume":"23","author":"E Vignotto","year":"2020","unstructured":"Vignotto, E., & Engelke, S. (2020). Extreme value theory for anomaly detection-the GPD classifier. Extremes, 23(4), 501\u2013520.","journal-title":"Extremes"},{"key":"7010_CR46","doi-asserted-by":"crossref","unstructured":"Wei, J. T. Z., Wang, R. Y., & Jia, R. (2024). Proving membership in LLM pretraining data via data watermarks. arXiv: abs\/2402.10892.","DOI":"10.18653\/v1\/2024.findings-acl.788"},{"key":"7010_CR47","doi-asserted-by":"crossref","unstructured":"Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P. S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A. & Biles, C. (2022). Taxonomy of risks posed by language models. In: Proceedings of the 2022 ACM conference on fairness, accountability, and transparency, pp. 214\u2013229.","DOI":"10.1145\/3531146.3533088"},{"issue":"4","key":"7010_CR48","doi-asserted-by":"publisher","first-page":"332","DOI":"10.1198\/000313006X152243","volume":"60","author":"L Wilkinson","year":"2006","unstructured":"Wilkinson, L. (2006). Revising the Pareto chart. The American Statistician, 60(4), 332\u2013334.","journal-title":"The American Statistician"},{"key":"7010_CR49","unstructured":"Wu, H., & Cao, Y. (2025). Membership inference attacks on large-scale models: A survey. arXiv:abs\/2503.19338."},{"issue":"5","key":"7010_CR50","doi-asserted-by":"publisher","first-page":"1250","DOI":"10.1109\/JIOT.2017.2694844","volume":"4","author":"Y Yang","year":"2017","unstructured":"Yang, Y., Wu, L., Yin, G., Li, L., & Zhao, H. (2017). A survey on security and privacy issues in internet-of-things. IEEE Internet of Things Journal, 4(5), 1250\u20131258.","journal-title":"IEEE Internet of Things Journal"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-026-07010-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-026-07010-4","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-026-07010-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T06:35:36Z","timestamp":1782282936000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-026-07010-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5]]},"references-count":50,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5]]}},"alternative-id":["7010"],"URL":"https:\/\/doi.org\/10.1007\/s10994-026-07010-4","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5]]},"assertion":[{"value":"18 April 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 December 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 February 2026","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical Approval and Consent to Participate"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for Publication"}}],"article-number":"124"}}