{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,21]],"date-time":"2026-04-21T01:12:04Z","timestamp":1776733924647,"version":"3.51.2"},"reference-count":33,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2022,7,16]],"date-time":"2022-07-16T00:00:00Z","timestamp":1657929600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>Due to the expansion of the internet, we encounter various types of big data such as web documents or sensing data. Compared to traditional small data such as experimental samples, big data provide more chances to find hidden and novel patterns with big data analysis using statistics and machine learning algorithms. However, as the use of big data increases, problems also occur. One of them is a zero-inflated problem in structured data preprocessed from big data. Most count values are zeros because a specific word is found in only some documents. In particular, since most of the patent data are in the form of a text document, they are more affected by the zero-inflated problem. To solve this problem, we propose a generation of synthetic samples using statistical inference and tree structure. Using patent document and simulation data, we verify the performance and validity of our proposed method. In this paper, we focus on patent keyword analysis as text big data analysis, and we encounter the zero-inflated problem just like other text data.<\/jats:p>","DOI":"10.3390\/fi14070211","type":"journal-article","created":{"date-parts":[[2022,7,17]],"date-time":"2022-07-17T21:00:28Z","timestamp":1658091628000},"page":"211","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Zero-Inflated Patent Data Analysis Using Generating Synthetic Samples"],"prefix":"10.3390","volume":"14","author":[{"given":"Daiho","family":"Uhm","sequence":"first","affiliation":[{"name":"Department of Mathematics, University of Arkansas\u2014Fort Smith, Fort Smith, AR 72913, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1961-0055","authenticated-orcid":false,"given":"Sunghae","family":"Jun","sequence":"additional","affiliation":[{"name":"Department of Big Data and Statistics, Cheongju University, Chungbuk 28503, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,7,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Cameron, A.C., and Trivedi, P.K. (2013). Regression Analysis of Count Data, Cambridge University Press. [2nd ed.].","DOI":"10.1017\/CBO9781139013567"},{"key":"ref_2","first-page":"431","article-title":"Zero-Inflated Poisson and Negative Binomial Regressions for Technology Analysis","volume":"10","author":"Kim","year":"2016","journal-title":"Int. J. Softw. Eng. Appl."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1186\/s40488-021-00121-4","article-title":"A comparison of zero-inflated and hurdle models for modeling zero-inflated count data","volume":"8","author":"Feng","year":"2021","journal-title":"J. Stat. Distrib. Appl."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1007\/s13744-019-00729-x","article-title":"Modeling overdispersion, autocorrelation, and zero-inflated count data via generalized additive models and Bayesian statistics in an Aphid population study","volume":"49","author":"Carvalho","year":"2020","journal-title":"Neotrop. Entomol."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Hilbe, J.M. (2011). Negative Binomial Regression, Cambridge University Press. [2nd ed.].","DOI":"10.1017\/CBO9780511973420"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Hilbe, J.M. (2014). Modeling Count Data, Cambridge University Press.","DOI":"10.1017\/CBO9781139236065"},{"key":"ref_7","unstructured":"Hunt, D., Nguyen, L., and Rodgers, M. (2007). Patent Searching Tools & Techniques, Wiley."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Roper, A.T., Cunningham, S.W., Porter, A.L., Mason, T.W., Rossini, F.A., and Banks, J. (2011). Forecasting and Management of Technology, John Wiley & Sons.","DOI":"10.1002\/9781118047989"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"103379","DOI":"10.1016\/j.compind.2020.103379","article-title":"Patent infringement analysis using a text mining technique based on SAO structure","volume":"125","author":"Kim","year":"2021","journal-title":"Comput. Ind."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wang, H.C., Chi, Y.C., and Hsin, P.L. (2018). Constructing patent maps using text mining to sustainably detect potential technological opportunities. Sustainability, 10.","DOI":"10.3390\/su10103729"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"467","DOI":"10.1080\/09537325.2012.674669","article-title":"Patent Text Mining and Informetric-based Patent Technology Morphological Analysis: An Empirical Study","volume":"24","author":"Feng","year":"2021","journal-title":"Technol. Anal. Strateg. Manag."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1","DOI":"10.18637\/jss.v025.i05","article-title":"Text mining infrastructure in R","volume":"25","author":"Feinerer","year":"2008","journal-title":"J. Stat. Softw."},{"key":"ref_13","unstructured":"Feinerer, I., and Hornik, K. (2022, March 01). Package \u2018tm\u2019 Ver. 0.7\u20138, Text Mining Package. Available online: https:\/\/cran.microsoft.com\/web\/packages\/tm\/tm.pdf."},{"key":"ref_14","unstructured":"R Development Core Team (2022, March 01). R: A Language and Environment for Statistical Computing, R Foundation for Statistical Computing. Available online: http:\/\/www.R-project.org."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1","DOI":"10.18637\/jss.v074.i11","article-title":"synthpop: Bespoke Creation of Synthetic Data in R","volume":"74","author":"Nowok","year":"2016","journal-title":"J. Stat. Softw."},{"key":"ref_16","unstructured":"Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep Learning, MIT Press."},{"key":"ref_17","unstructured":"Han, J., Kamber, M., and Pei, J. (2012). Data Mining: Concepts and Techniques, Morgan Kaufmann. [3rd ed.]."},{"key":"ref_18","unstructured":"Nowok, B., Raab, G.M., Snoke, J., Dibben, C., and Nowok, M.B. (2022, March 01). Package \u2018synthpop\u2019 Ver. 1.7\u20130, Generating Synthetic Versions of Sensitive Microdata for Statistical Disclosure Control. Available online: https:\/\/cran.r-project.org\/web\/packages\/synthpop\/synthpop.pdf."},{"key":"ref_19","first-page":"67","article-title":"Practical Data Synthesis for Large Samples","volume":"7","author":"Raab","year":"2018","journal-title":"J. Priv. Confid."},{"key":"ref_20","first-page":"441","article-title":"Using CART to Generate Partially Synthetic Public Use Microdata","volume":"21","author":"Reiter","year":"2005","journal-title":"J. Off. Stat."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"663","DOI":"10.1111\/rssa.12358","article-title":"General and specific utility measures for synthetic data","volume":"181","author":"Snoke","year":"2018","journal-title":"J. R. Stat. Soc. Ser. A"},{"key":"ref_22","unstructured":"Bruce, P., Bruce, A., and Gedeck, P. (2020). Practical Statistics for Data Scientists, O\u2019Reilly Media."},{"key":"ref_23","unstructured":"Murphy, K.P. (2012). Machine Learning: A Probabilistic Perspective, MIT Press."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Theodoridis, S. (2015). Machine Learning A Bayesian and Optimization Perspective, Elsevier.","DOI":"10.1016\/B978-0-12-801522-3.00012-4"},{"key":"ref_25","unstructured":"Montgomery, D.C., Peck, E.A., and Vining, G.G. (2012). Introduction to Linear Regression Analysis, John Wiley & Sons."},{"key":"ref_26","unstructured":"(2022, March 01). USPTO, The United States Patent and Trademark Office, Available online: http:\/\/www.uspto.gov."},{"key":"ref_27","unstructured":"(2022, March 01). KIPRIS, Korea Intellectual Property Rights Information Service. Available online: www.kipris.or.kr."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"3204","DOI":"10.1016\/j.eswa.2013.11.018","article-title":"Document Clustering Method Using Dimension Reduction and Support Vector Clustering to Overcome Sparseness","volume":"41","author":"Jun","year":"2014","journal-title":"Expert Syst. Appl."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"3129","DOI":"10.1080\/00949655.2014.953534","article-title":"Simultaneous generation of multivariate mixed data with Poisson and normal marginals","volume":"85","author":"Amatya","year":"2015","journal-title":"J. Stat. Comput. Simul."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"2241","DOI":"10.1080\/03610918.2015.1039854","article-title":"PoisNor: An R package for generation of multivariate data with Poisson and normal marginals","volume":"46","author":"Amatya","year":"2017","journal-title":"Commun. Stat. Simul. Comput."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"173","DOI":"10.32614\/RJ-2021-007","article-title":"RNGforGPD: An R Package for Generation of Univariate and Multivariate Generalized Poisson Data","volume":"12","author":"Li","year":"2020","journal-title":"R J."},{"key":"ref_32","unstructured":"Li, H., Chen, R., Nguyen, H., Chung, Y., Gao, R., and Demirtas, H. (2022, March 01). Package \u2018RNGforGPD\u2019 Ver. 1.1.0, Random Number Generation for Generalized Poisson Distribution. Available online: https:\/\/cran.r-project.org\/web\/packages\/RNGforGPD\/RNGforGPD.pdf."},{"key":"ref_33","first-page":"57","article-title":"A multivariate generalization of the generalized Poisson distribution. ASTIN Bulletin","volume":"30","author":"Vernic","year":"2000","journal-title":"J. Int. Actuar. Assoc."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/7\/211\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:51:54Z","timestamp":1760140314000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/7\/211"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,16]]},"references-count":33,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2022,7]]}},"alternative-id":["fi14070211"],"URL":"https:\/\/doi.org\/10.3390\/fi14070211","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,16]]}}}