{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T01:10:19Z","timestamp":1779325819137,"version":"3.51.4"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[1994,1,1]],"date-time":"1994-01-01T00:00:00Z","timestamp":757382400000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[1994,1]]},"DOI":"10.1007\/bf00993163","type":"journal-article","created":{"date-parts":[[2005,1,9]],"date-time":"2005-01-09T17:18:07Z","timestamp":1105291087000},"page":"83-113","source":"Crossref","is-referenced-by-count":42,"title":["Bounds on the sample complexity of Bayesian learning using information theory and the VC dimension"],"prefix":"10.1007","volume":"14","author":[{"given":"David","family":"Haussler","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"Kearns","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robert E.","family":"Schapire","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","reference":[{"issue":"3","key":"CR1","doi-asserted-by":"crossref","first-page":"233","DOI":"10.5802\/aif.938","volume":"33","author":"P. Assouad","year":"1983","unstructured":"Assouad, P. (1983). Densit\u00e9 et dimension.Annales de l'Institut Fourier, 33(3):233?282.","journal-title":"Annales de l'Institut Fourier"},{"key":"CR2","first-page":"1224","volume":"13","author":"J. M. Barzdin","year":"1972","unstructured":"Barzdin, J. M. and Freivald, R. V. (1972). On the prediction of general recursive functions.Soviet Mathematics-Doklady, 13:1224?1228.","journal-title":"Soviet Mathematics-Doklady"},{"issue":"1","key":"CR3","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1162\/neco.1989.1.1.151","volume":"1","author":"E. Baum","year":"1989","unstructured":"Baum, E. and Haussler, D. (1989). What size net gives valid generalization?Neural Computation, 1(1):151?160.","journal-title":"Neural Computation"},{"issue":"4","key":"CR4","doi-asserted-by":"crossref","first-page":"929","DOI":"10.1145\/76359.76371","volume":"36","author":"A. Blumer","year":"1989","unstructured":"Blumer, A., Ehrenfeucht, A., Haussler, D., and Warmuth, M. K. (1989). Learnability and the Vapnik-Chervonenkis dimension.Journal of the Association for Computing Machinery, 36(4):929?965.","journal-title":"Journal of the Association for Computing Machinery"},{"key":"CR5","volume-title":"A Theory of Learning Classification Rules","author":"W. Buntine","year":"1990","unstructured":"Buntine, W. (1990).A Theory of Learning Classification Rules. PhD thesis, University of Technology, Sydney."},{"key":"CR6","unstructured":"Buntine, W. and Weigend, A. (1991). Bayesian back propagation. Unpublished manuscript."},{"issue":"3","key":"CR7","doi-asserted-by":"crossref","first-page":"453","DOI":"10.1109\/18.54897","volume":"36","author":"B. Clarke","year":"1990","unstructured":"Clarke, B. and Barron, A. (1990). Information-theoretic asymptotics of Bayes methods.IEEE Transactions on Information Theory, 36(3):453?471.","journal-title":"IEEE Transactions on Information Theory"},{"key":"CR8","unstructured":"Clarke, B. and Barron, A. (1991). Entropy, risk and the Bayesian central limit theorem. manuscript."},{"key":"CR9","doi-asserted-by":"crossref","unstructured":"Cover, T. and Thomas, J. (1991).Elements of Information Theory. Wiley.","DOI":"10.1002\/0471200611"},{"key":"CR10","first-page":"877","volume":"1","author":"J. Denker","year":"1987","unstructured":"Denker, J., Schwartz, D., Wittner, B., Solla, S., Howard, R., Jackel, L., and Hopfield, J. (1987). Automatic learning, rule extraction and generalization.Complex Systems, 1:877?922.","journal-title":"Complex Systems"},{"key":"CR11","doi-asserted-by":"crossref","unstructured":"DeSantis, A., Markowski, G., and Wegman, M. N. (1988). Learning probabilistic prediction functions. InProceedings of the 1988 Workshop on Computational Learning Theory, pages 312?328. Morgan Kaufmann.","DOI":"10.1109\/SFCS.1988.21929"},{"key":"CR12","unstructured":"Duda, R. O. and Hart, P. E. (1973).Pattern Classification and Scene Analysis. Wiley."},{"key":"CR13","first-page":"2","volume":"1097","author":"R. M. Dudley","year":"1984","unstructured":"Dudley, R. M. (1984). A course on empirical processes.Lecture Notes in Mathematics, 1097:2?142.","journal-title":"Lecture Notes in Mathematics"},{"key":"CR14","unstructured":"Fano, R. (1952). Class notes for course 6.574. Technical report, Massachusetts Institute of Technology."},{"key":"CR15","unstructured":"Gyorgyi, G. and Tishby, N. (1990). In Thuemann, K. and Koeberle, R., editors,Neural Networks and Spin Glasses. World Scientific."},{"key":"CR16","unstructured":"Haussler, D. (1991). Sphere packing numbers for subsets of the Booleann-cube with bounded Vapnik-Chervonenkis dimension. Technical Report UCSC-CRL-91-41, University of Calif. Computer Research Laboratory, Santa Cruz, CA."},{"key":"CR17","unstructured":"Haussler, D., Littlestone, N., and Warmuth, M. (1990). Predicting {0, 1}-functions on randomly drawn points. Technical Report UCSC-CRL-90-54, University of California Santa Cruz, Computer Research Laboratory. To appear in Information and Computation."},{"key":"CR18","unstructured":"Littlestone, N. (1989).Mistake Bounds and Logarithmic Linear-threshold Learning Algorithms. PhD thesis, University of California Santa Cruz."},{"key":"CR19","doi-asserted-by":"crossref","unstructured":"Littlestone, N., Long, P. M., and Warmuth, M. K. (1991). On-line learning of linear functions. InProceedings of the Twenty Third Annual ACM Symposium on Theory of Computing, pages 465?475.","DOI":"10.1145\/103418.103467"},{"key":"CR20","doi-asserted-by":"crossref","unstructured":"Littlestone, N. and Warmuth, M. (1989). The weighted majority algorithm. Technical Report UCSC-CRL-89-16, Computer Research Laboratory, University of Santa Cruz.","DOI":"10.1109\/SFCS.1989.63487"},{"key":"CR21","unstructured":"MacKay, D. (1992).Bayesian Methods for Adaptive Models. PhD thesis, California Institute of Technology."},{"key":"CR22","first-page":"381","volume":"22","author":"P. Massart","year":"1986","unstructured":"Massart, P. (1986). Rates of convergence in the central limit theorem for empirical processes.Annales de l'Institut Henri Poincar\u00e9 Probabilites et Statistiques, 22:381?423.","journal-title":"Annales de l'Institut Henri Poincar\u00e9 Probabilites et Statistiques"},{"issue":"3","key":"CR23","doi-asserted-by":"crossref","first-page":"438","DOI":"10.1137\/0221029","volume":"21","author":"B. K. Natarajan","year":"1992","unstructured":"Natarajan, B. K. (1992). Probably approximate learning over classes of distributions.SIAM Journal on Computing, 21(3):438?449.","journal-title":"SIAM Journal on Computing"},{"key":"CR24","doi-asserted-by":"crossref","unstructured":"Opper, M. and Haussler, D. (1991). Calculation of the learning curve of Bayes optimal classification algorithm for learning a perceptron with noise. InProceedings of the Fourth Annual Workshop on Computational Learning Theory, pages 75?87. Morgan Kaufmann.","DOI":"10.1016\/B978-1-55860-213-7.50011-0"},{"issue":"4","key":"CR25","first-page":"349","volume":"9","author":"M. J. Pazzani","year":"1992","unstructured":"Pazzani, M. J. and Sarrett, W. (1992). A framework for average case analysis of conjunctive learning algorithms.Machine Learning, 9(4):349?372.","journal-title":"Machine Learning"},{"key":"CR26","volume-title":"Probability Theory","author":"A. Renyi","year":"1970","unstructured":"Renyi, A. (1970).Probability Theory. North Holland, Amsterdam."},{"key":"CR27","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1016\/0097-3165(72)90019-2","volume":"13","author":"N. Sauer","year":"1972","unstructured":"Sauer, N. (1972). On the density of families of sets.Journal of Combinatorial Theory (Series A), 13:145?147.","journal-title":"Journal of Combinatorial Theory (Series A)"},{"key":"CR28","series-title":"Technical Report","volume-title":"Bounding sample size with the Vapnik-Chervonenkis dimension","author":"J. Shawe-Taylor","year":"1989","unstructured":"Shawe-Taylor, J., Anthony, M., and Biggs, N. (1989). Bounding sample size with the Vapnik-Chervonenkis dimension. Technical Report CSD-TR-618, University of London, Surrey, England."},{"key":"CR29","doi-asserted-by":"crossref","first-page":"1683","DOI":"10.1103\/PhysRevLett.65.1683","volume":"65","author":"H. Sompolinsky","year":"1990","unstructured":"Sompolinsky, H., Tishby, N., and Seung, H. (1990). Learning from examples in large neural networks.Physical Review Letters, 65:1683?1686.","journal-title":"Physical Review Letters"},{"key":"CR30","doi-asserted-by":"crossref","first-page":"169","DOI":"10.1007\/BF00322017","volume":"78","author":"M. Talagrand","year":"1988","unstructured":"Talagrand, M. (1988). Donsker classes of sets.Probability Theory and Related Fields, 78:169?191.","journal-title":"Probability Theory and Related Fields"},{"key":"CR31","doi-asserted-by":"crossref","unstructured":"Tishby, N., Levin, E., and Solla, S. (1989). Consistent inference of probabilities in layered networks: predictions and generalizations. InIJCNN International Joint Conference on Neural Networks, volume II, pages 403?409. IEEE.","DOI":"10.1109\/IJCNN.1989.118274"},{"issue":"11","key":"CR32","doi-asserted-by":"crossref","first-page":"1134","DOI":"10.1145\/1968.1972","volume":"27","author":"L. G. Valiant","year":"1984","unstructured":"Valiant, L. G. (1984). A theory of the learnable.Communications of the ACM, 27(11):1134?42.","journal-title":"Communications of the ACM"},{"key":"CR33","unstructured":"Vapnik, V. N. (1979).Theorie der Zeichenerkennung. Akademie-Verlag."},{"key":"CR34","volume-title":"Estimation of Dependences Based on Empirical Data","author":"V. N. Vapnik","year":"1982","unstructured":"Vapnik, V. N. (1982).Estimation of Dependences Based on Empirical Data. Springer-Verlag, New York."},{"issue":"2","key":"CR35","doi-asserted-by":"crossref","first-page":"264","DOI":"10.1137\/1116025","volume":"16","author":"V. N. Vapnik","year":"1971","unstructured":"Vapnik, V. N. and Chervonenkis, A. Y. (1971). On the uniform convergence of relative frequencies of events to their probabilities.Theory of Probability and its Applications, 16(2):264?80.","journal-title":"Theory of Probability and its Applications"},{"key":"CR36","doi-asserted-by":"crossref","unstructured":"Vovk, V. (1990). Aggregating strategies. InProceedings of the Third Annual Workshop on Computational Learning Theory, pages 371?383. Morgan Kaufmann.","DOI":"10.1016\/B978-1-55860-146-8.50032-1"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/BF00993163.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/BF00993163\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/BF00993163","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,4,29]],"date-time":"2019-04-29T22:58:41Z","timestamp":1556578721000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/BF00993163"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[1994,1]]},"references-count":36,"journal-issue":{"issue":"1","published-print":{"date-parts":[[1994,1]]}},"alternative-id":["BF00993163"],"URL":"https:\/\/doi.org\/10.1007\/bf00993163","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[1994,1]]}}}