{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,20]],"date-time":"2025-10-20T22:05:00Z","timestamp":1760997900541,"version":"build-2065373602"},"reference-count":35,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,3,10]],"date-time":"2023-03-10T00:00:00Z","timestamp":1678406400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Center for TDB Advanced Data Analysis and Modeling, Tokyo Institute of Technology"},{"name":"EIKOKU DATABANK, Ltd."}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>We introduce a new non-black-box method of extracting multiple areas in a high-dimensional big data space where data points that satisfy specific conditions are highly concentrated. First, we extract one-dimensional areas where the data that satisfy specific conditions are mostly gathered by using the Bayesian method. Second, we construct higher-dimensional areas where the densities of focused data points are higher than the simple combination of the results for one dimension, and then we verify the results through data validation. Third, we apply this method to estimate the set of significant factors shared in successful firms with growth rates in sales at the top 1% level using 156-dimensional data of corporate financial reports for 12 years containing about 320,000 firms. We also categorize high-growth firms into 15 groups of different sets of factors.<\/jats:p>","DOI":"10.3390\/e25030488","type":"journal-article","created":{"date-parts":[[2023,3,13]],"date-time":"2023-03-13T04:04:00Z","timestamp":1678680240000},"page":"488","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Extraction of Important Factors in a High-Dimensional Data Space: An Application for High-Growth Firms"],"prefix":"10.3390","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7632-5050","authenticated-orcid":false,"given":"Takuya","family":"Wada","sequence":"first","affiliation":[{"name":"Department of Mathematical and Computing Science, School of Computing, Tokyo Institute of Technology, Yokohama 226-8502, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4452-5517","authenticated-orcid":false,"given":"Hideki","family":"Takayasu","sequence":"additional","affiliation":[{"name":"Institute of Innovative Research, Tokyo Institute of Technology, Yokohama 226-8502, Japan"},{"name":"Sony Computer Science Laboratories, Tokyo 141-0022, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2467-614X","authenticated-orcid":false,"given":"Misako","family":"Takayasu","sequence":"additional","affiliation":[{"name":"Department of Mathematical and Computing Science, School of Computing, Tokyo Institute of Technology, Yokohama 226-8502, Japan"},{"name":"Institute of Innovative Research, Tokyo Institute of Technology, Yokohama 226-8502, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"131","DOI":"10.3233\/IDA-1997-1302","article-title":"Feature selection for classification","volume":"1","author":"Dash","year":"1997","journal-title":"Intell. Data Anal."},{"key":"ref_2","unstructured":"Kira, K., and Rendell, L.A. (1992). Machine Learning Proceedings 1992, Elsevier."},{"key":"ref_3","first-page":"1157","article-title":"An introduction to variable and feature selection","volume":"3","author":"Guyon","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"2507","DOI":"10.1093\/bioinformatics\/btm344","article-title":"A review of feature selection techniques in bioinformatics","volume":"23","author":"Saeys","year":"2007","journal-title":"Bioinformatics"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1109\/34.574797","article-title":"Feature selection: Evaluation, application, and small sample performance","volume":"19","author":"Jain","year":"1997","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","first-page":"51","article-title":"A comparative study on feature selection and classification methods using gene expression profiles and proteomic patterns","volume":"13","author":"Liu","year":"2002","journal-title":"Genome Inform."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/A:1010933404324","article-title":"Random forests","volume":"45","author":"Breiman","year":"2001","journal-title":"Mach. Learn."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"2225","DOI":"10.1016\/j.patrec.2010.03.014","article-title":"Variable selection using random forests","volume":"31","author":"Genuer","year":"2010","journal-title":"Pattern Recognit. Lett."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Vapnik, V. (1999). The Nature of Statistical Learning Theory, Springer Science & Business Media.","DOI":"10.1007\/978-1-4757-3264-1"},{"key":"ref_10","unstructured":"Grandvalet, Y., and Canu, S. (2002). Adaptive scaling for feature selection in SVMs. Adv. Neural Inf. Process. Syst., 15."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"764","DOI":"10.1093\/aje\/kwt312","article-title":"Comparison of random forest and parametric imputation models for imputing missing data using MICE: A CALIBER study","volume":"179","author":"Shah","year":"2014","journal-title":"Am. J. Epidemiol."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"206","DOI":"10.1038\/s42256-019-0048-x","article-title":"Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead","volume":"1","author":"Rudin","year":"2019","journal-title":"Nat. Mach. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1152","DOI":"10.1214\/aos\/1176342871","article-title":"Mixtures of Dirichlet processes with applications to Bayesian nonparametric problems","volume":"2","author":"Antoniak","year":"1974","journal-title":"Ann. Stat."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1038\/nrg1318","article-title":"The Bayesian revolution in genetics","volume":"5","author":"Beaumont","year":"2004","journal-title":"Nat. Rev. Genet."},{"key":"ref_15","first-page":"151","article-title":"Bayesian methods for analysis of stock mixtures from genetic characters","volume":"99","author":"Pella","year":"2001","journal-title":"Fish. Bull."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"703","DOI":"10.1590\/0102-311X00144013","article-title":"Trends in epidemiology in the 21st century: Time to adopt Bayesian methods","volume":"30","author":"Martinez","year":"2014","journal-title":"Cad. Sa\u00fade P\u00fablica"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"509","DOI":"10.1111\/j.1461-0248.2004.00603.x","article-title":"Bayesian inference in ecology","volume":"7","author":"Ellison","year":"2004","journal-title":"Ecol. Lett."},{"key":"ref_18","first-page":"422","article-title":"Bayesian estimation of seismic hazards in Iran","volume":"20","author":"Yazdani","year":"2013","journal-title":"Sci. Iran."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yamada, K., Takayasu, H., and Takayasu, M. (2018). Estimation of economic indicator announced by government from social big data. Entropy, 20.","DOI":"10.3390\/e20110852"},{"key":"ref_20","first-page":"19","article-title":"A survey on similarity measures in text mining","volume":"3","author":"Vijaymeena","year":"2016","journal-title":"Mach. Learn. Appl. Int. J."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"567","DOI":"10.2307\/2098588","article-title":"The relationship between firm growth, size, and age: Estimates for 100 manufacturing industries","volume":"35","author":"Evans","year":"1987","journal-title":"J. Ind. Econ."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/0304-405X(95)00842-3","article-title":"Leverage, investment, and firm growth","volume":"40","author":"Lang","year":"1996","journal-title":"J. Financ. Econ."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"2107","DOI":"10.1111\/0022-1082.00084","article-title":"Law, finance, and firm growth","volume":"53","author":"Maksimovic","year":"1998","journal-title":"J. Financ."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"292","DOI":"10.2307\/3069456","article-title":"A multidimensional model of venture growth","volume":"44","author":"Baum","year":"2001","journal-title":"Acad. Manag. J."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"e00107","DOI":"10.1016\/j.jbvi.2018.e00107","article-title":"Is firm growth random? A machine learning perspective","volume":"11","author":"Kolkman","year":"2019","journal-title":"J. Bus. Ventur. Insights"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1016\/S0883-9026(02)00080-0","article-title":"Arriving at the high-growth firm","volume":"18","author":"Delmar","year":"2003","journal-title":"J. Bus. Ventur."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1111\/j.2517-6161.1996.tb02080.x","article-title":"Regression shrinkage and selection via the lasso","volume":"58","author":"Tibshirani","year":"1996","journal-title":"J. R. Stat. Soc. Ser. (Methodol.)"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"541","DOI":"10.1007\/s11187-019-00203-3","article-title":"Catching Gazelles with a Lasso: Big data techniques for the prediction of high-growth firms","volume":"55","author":"Coad","year":"2020","journal-title":"Small Bus. Econ."},{"key":"ref_29","unstructured":"Teikoku Databank Ltd (2023, January 31). Our Profile and History. Available online: https:\/\/www.tdb-en.jp\/company\/profile.html."},{"key":"ref_30","unstructured":"O\u2019Neill, M.E. (2023, January 30). PCG: A Family of Simple Fast Space-Efficient Statistically Good Algorithms for Random Number Generation. ACM Transactions on Mathematical Software. Available online: https:\/\/www.pcg-random.org\/pdf\/toms-oneill-pcg-family-v1.02.pdf."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"236","DOI":"10.1080\/01621459.1963.10500845","article-title":"Hierarchical grouping to optimize an objective function","volume":"58","author":"Ward","year":"1963","journal-title":"J. Am. Stat. Assoc."},{"key":"ref_32","unstructured":"Sakurai, H. (2021). Financial Accounting Lecture, Chuokeizai-Sha Holdings, Inc.. [22nd ed.]. (In Japanese)."},{"key":"ref_33","unstructured":"Haykin, S. (1998). Neural Networks: A Comprehensive Foundation, Prentice Hall PTR."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"657","DOI":"10.1086\/261480","article-title":"Tests of alternative theories of firm growth","volume":"95","author":"Evans","year":"1987","journal-title":"J. Political Econ."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1016\/0883-9026(91)90028-C","article-title":"Continued entrepreneurship: Ability, need, and opportunity as determinants of small firm growth","volume":"6","author":"Davidsson","year":"1991","journal-title":"J. Bus. Ventur."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/25\/3\/488\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:52:32Z","timestamp":1760122352000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/25\/3\/488"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,10]]},"references-count":35,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["e25030488"],"URL":"https:\/\/doi.org\/10.3390\/e25030488","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2023,3,10]]}}}