{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:16:06Z","timestamp":1760235366095,"version":"build-2065373602"},"reference-count":42,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2021,8,13]],"date-time":"2021-08-13T00:00:00Z","timestamp":1628812800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>In this paper, we investigate the problem of classifying feature vectors with mutually independent but non-identically distributed elements that take values from a finite alphabet set. First, we show the importance of this problem. Next, we propose a classifier and derive an analytical upper bound on its error probability. We show that the error probability moves to zero as the length of the feature vectors grows, even when there is only one training feature vector per label available. Thereby, we show that for this important problem at least one asymptotically optimal classifier exists. Finally, we provide numerical examples where we show that the performance of the proposed classifier outperforms conventional classification algorithms when the number of training data is small and the length of the feature vectors is sufficiently high.<\/jats:p>","DOI":"10.3390\/e23081045","type":"journal-article","created":{"date-parts":[[2021,8,13]],"date-time":"2021-08-13T09:22:38Z","timestamp":1628846558000},"page":"1045","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["On Supervised Classification of Feature Vectors with Independent and Non-Identically Distributed Elements"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7731-4329","authenticated-orcid":false,"given":"Farzad","family":"Shahrivari","sequence":"first","affiliation":[{"name":"Electrical and Computer Systems Engineering, Monash University, Alliance Ln, Clayton, VIC 3168, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nikola","family":"Zlatanov","sequence":"additional","affiliation":[{"name":"Electrical and Computer Systems Engineering, Monash University, Alliance Ln, Clayton, VIC 3168, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,8,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Shalev-Shwartz, S., and Ben-David, S. (2014). Understanding Machine Learning: From Theory to Algorithms, Cambridge University Press.","DOI":"10.1017\/CBO9781107298019"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Quinlan, J.R. (1983). Learning Efficient Classification Procedures and Their Application to Chess End Games, Springer.","DOI":"10.1016\/B978-0-08-051054-5.50019-4"},{"key":"ref_3","unstructured":"Breiman, L., Friedman, J.H., Olshen, R.A., and Stone, C.J. (1984). Classification and Regression Trees, Wadsworth and Brooks."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Boser, B.E., Guyon, I.M., and Vapnik, V.N. (1992, January 27\u201329). A training algorithm for optimal margin classifiers. Proceedings of the Fifth Annual Workshop on Computational Learning Theory, San Jose, CA, USA.","DOI":"10.1145\/130385.130401"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1007\/BF00994018","article-title":"Support-vector networks","volume":"20","author":"Cortes","year":"1995","journal-title":"Mach. Learn."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Lallich, S., Teytaud, O., and Prudhomme, E. (2007). Association Rule Interestingness: Measure and Statistical Validation, Springer. Studies in Computational Intelligence.","DOI":"10.1007\/978-3-540-44918-8_11"},{"key":"ref_7","unstructured":"Langley, P., Iba, W., and Thompson, K. (1992, January 14\u201316). An analysis of bayesian classifiers. Proceedings of the Tenth National Conference on Artificial Intelligence, San Jose, CA, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Aha, D.W. (1997). Lazy learning. Artificial Intelligent Review, Kluwer Academic Publishers.","DOI":"10.1007\/978-94-017-2053-3"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Bishop, C.M. (1995). Neural Networks for Pattern Recognition, Oxford University Press.","DOI":"10.1093\/oso\/9780198538493.001.0001"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Hastie, T., Tibshirani, R., and Friedman, J. (2009). The Elements of Statistical Learning: Data Mining, Springer.","DOI":"10.1007\/978-0-387-84858-7"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1988","DOI":"10.1109\/18.476321","article-title":"The asymptotics of posterior entropy and error probability for bayesian estimation","volume":"41","author":"Kanaya","year":"1995","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"401","DOI":"10.1109\/18.32134","article-title":"Asymptotically optimal classification for multiple tests with empirically observed statistics","volume":"35","author":"Gutman","year":"1989","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/TIT.2017.2757496","article-title":"Arimotornyi conditional entropy and bayesian hypothesis testing","volume":"64","author":"Sason","year":"2018","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_14","unstructured":"Wang, C., She, Z., and Cao, L. (2013, January 2\u20139). Coupled attribute analysis on numerical data. Proceedings of the Twenty-Third International Joint Conference on Artificial Intelligence, Beijing, China."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"781","DOI":"10.1109\/TNNLS.2014.2325872","article-title":"Coupled attribute similarity learning on categorical data","volume":"26","author":"Wang","year":"2015","journal-title":"IEEE Trans. Neural Networks Learn. Syst."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wang, C., Cao, L., Wang, M., Li, J., Wei, W., and Ou, Y. (2011, January 24\u201328). Coupled nominal similarity in unsupervised learning. Proceedings of the 20th ACM International Conference on Information and Knowledge Management, Scotland, UK.","DOI":"10.1145\/2063576.2063715"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Liu, C., and Cao, L. (2015). A coupled k-nearest neighbor algorithm for multi-label classification. PAKDD, Springer International Publishing.","DOI":"10.1007\/978-3-319-18038-0_14"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Liu, C., Cao, L., and Yu, P. (2014, January 6\u201311). A hybrid coupled k-nearest neighbor algorithm on imbalance data. Proceedings of the International Joint Conference on Neural Networks, Beijing, China.","DOI":"10.1109\/IJCNN.2014.6889798"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liu, C., Cao, L., and Yu, P.S. (2014, January 8\u201313). Coupled fuzzy k-nearest neighbors classification of imbalanced non-iid categorical data. Proceedings of the 2014 International Joint Conference on Neural Networks IJCNN, Beijing, China.","DOI":"10.1109\/IJCNN.2014.6889773"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1358","DOI":"10.1093\/comjnl\/bxt084","article-title":"Non-IIDness Learning in Behavioral and Social Data","volume":"57","author":"Cao","year":"2013","journal-title":"Comput. J."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1016\/j.ipm.2014.08.007","article-title":"Coupling learning of complex interactions","volume":"51","author":"Cao","year":"2015","journal-title":"Inf. Process. Manag."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"699","DOI":"10.1109\/TSMCB.2010.2086060","article-title":"Combined mining discovering informative knowledge in complex data","volume":"41","author":"Cao","year":"2011","journal-title":"IEEE Trans. Syst. Man Cybern. Part B (Cybern.)"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Getoor, L., and Taskar, B. (2007). Introduction to Statistical Relational Learning (Adaptive Computation and Machine Learning), The MIT Press.","DOI":"10.7551\/mitpress\/7432.001.0001"},{"key":"ref_24","first-page":"499","article-title":"A subspace algorithm for certain blind identification problems","volume":"43","author":"Loubaton","year":"2006","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_25","first-page":"231","article-title":"Linear and nonlinear ica based on mutual information\u2014The misep method","volume":"84","author":"Almeida","year":"2004","journal-title":"IEEE Trans. Signal Process."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"411430","DOI":"10.1016\/S0893-6080(00)00026-5","article-title":"Independent component analysis: Algorithms and applications","volume":"13","author":"Hyvarinen","year":"2000","journal-title":"Neural Netw."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"53425359","DOI":"10.1109\/TIT.2011.2145090","article-title":"Independent component analysis over galois fields of prime order","volume":"57","author":"Yeredor","year":"2011","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"31683181","DOI":"10.1109\/TSP.2011.2144975","article-title":"Binary independent component analysis with or mixtures","volume":"59","author":"Nguyen","year":"2011","journal-title":"IEEE Trans. Signal Processing"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"295311","DOI":"10.1162\/neco.1989.1.3.295","article-title":"Unsupervised learning","volume":"1","author":"Barlow","year":"1989","journal-title":"Neural Comput."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Nakhaeizadeh, G., Taylor, C.C., and Kunisch, G. (1997). Dynamic Supervised Learning: Some Basic Issues and Application Aspects. Classification and Knowledge Organization, Springer.","DOI":"10.1007\/978-3-642-59051-1_13"},{"key":"ref_31","unstructured":"Koppen, M. (2000). The Curse of Dimensionality, SAGE."},{"key":"ref_32","unstructured":"Duda, R.O., Hart, P.E., and Stork, D.G. (2000). Pattern Classification, Wiley-Interscience. [2nd ed.]."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"437","DOI":"10.1109\/34.824819","article-title":"Statistical pattern recognition: A review","volume":"22","author":"Jain","year":"2000","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"44","DOI":"10.4018\/jdwm.2012040103","article-title":"Classifying very high-dimensional data with random forests built from small subspaces","volume":"8","author":"Xu","year":"2012","journal-title":"Int. J. Data Warehous. Min."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"769","DOI":"10.1016\/j.patcog.2012.09.005","article-title":"Stratified sampling for feature subspace selection in random forests for high dimensional data","volume":"46","author":"Ye","year":"2013","journal-title":"Pattern Recognit."},{"key":"ref_36","unstructured":"Kouiroukidis, N., and Evangelidis, G. (October, January 30). The effects of dimensionality curse in high dimensional knn search. Proceedings of the 15th Panhellenic Conference on Informatics, Kastoria, Greece."},{"key":"ref_37","unstructured":"Matusita, K. (1967). Classification based on distance in multivariate gaussian cases. Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability; Volume 1: Statistics, University of California Press."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"253285","DOI":"10.1023\/A:1013912006537","article-title":"Logistic regression, adaboost and bregman distances","volume":"48","author":"Collins","year":"2002","journal-title":"Mach. Learn."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Deisenroth, M.P., Faisal, A.A., and Ong, C.S. (2020). Mathematics for Machine Learning, Cambridge University Press.","DOI":"10.1017\/9781108679930"},{"key":"ref_40","unstructured":"Lebanon, G., and Lafferty, J. (2002, January 8\u201312). Cranking: Combining rankings using conditional probability models on permutations. Proceedings of the 19th International Conference on Machine Learning, San Francisco, CA, USA."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"713","DOI":"10.1214\/aoms\/1177728178","article-title":"On the distribution of the number of successes in independent trials","volume":"27","author":"Hoeffding","year":"1956","journal-title":"Ann. Math. Statist."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1080\/01621459.1963.10500830","article-title":"Probability inequalities for sums of bounded random variables","volume":"58","author":"Hoeffding","year":"1963","journal-title":"J. Am. Stat. Assoc."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/8\/1045\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T06:45:39Z","timestamp":1760165139000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/8\/1045"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,13]]},"references-count":42,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2021,8]]}},"alternative-id":["e23081045"],"URL":"https:\/\/doi.org\/10.3390\/e23081045","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2021,8,13]]}}}