{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T17:42:14Z","timestamp":1781545334836,"version":"3.54.5"},"reference-count":34,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2023,8,10]],"date-time":"2023-08-10T00:00:00Z","timestamp":1691625600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,8,10]],"date-time":"2023-08-10T00:00:00Z","timestamp":1691625600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Stat Comput"],"published-print":{"date-parts":[[2023,10]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In this work, minibatch MCMC sampling for feedforward neural networks is made more feasible. To this end, it is proposed to sample subgroups of parameters via a blocked Gibbs sampling scheme. By partitioning the parameter space, sampling is possible irrespective of layer width. It is also possible to alleviate vanishing acceptance rates for increasing depth by reducing the proposal variance in deeper layers. Increasing the length of a non-convergent chain increases the predictive accuracy in classification tasks, so avoiding vanishing acceptance rates and consequently enabling longer chain runs have practical benefits. Moreover, non-convergent chain realizations aid in the quantification of predictive uncertainty. An open problem is how to perform minibatch MCMC sampling for feedforward neural networks in the presence of augmented data.<\/jats:p>","DOI":"10.1007\/s11222-023-10285-5","type":"journal-article","created":{"date-parts":[[2023,8,10]],"date-time":"2023-08-10T12:07:28Z","timestamp":1691669248000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Approximate blocked Gibbs sampling for Bayesian neural networks"],"prefix":"10.1007","volume":"33","author":[{"given":"Theodore","family":"Papamarkou","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,8,10]]},"reference":[{"key":"10285_CR1","unstructured":"Alexos, A., Boyd, A.J., Mandt, S.: Structured stochastic gradient MCMC. In: Proceedings of the 39th International Conference on Machine Learning, vol. 162, pp. 414\u2013434. PMLR, Baltimore (2022)"},{"key":"10285_CR2","unstructured":"Andrieu, C., de Freitas, J.F.G., Doucet, A.: Sequential Bayesian Estimation and Model Selection Applied to Neural Networks, Cambridge (1999)"},{"key":"10285_CR3","unstructured":"Andrieu, C., de Freitas, N., Doucet, A.: Reversible jump MCMC simulated annealing for neural networks. In: Proceedings of the 16th Conference on Uncertainty in Artificial Intelligence, pp. 11\u201318 (2000)"},{"issue":"4","key":"10285_CR4","doi-asserted-by":"publisher","first-page":"343","DOI":"10.1007\/s11222-008-9110-y","volume":"18","author":"C Andrieu","year":"2008","unstructured":"Andrieu, C., Thoms, J.: A tutorial on adaptive MCMC. Stat. Comput. 18(4), 343\u2013373 (2008)","journal-title":"Stat. Comput."},{"key":"10285_CR5","unstructured":"Bardenet, R., Doucet, A., Holmes, C.: Towards scaling up Markov chain Monte Carlo: an adaptive subsampling approach. In: Proceedings of the 31st International Conference on Machine Learning, vol. 32, pp. 405\u2013413. PMLR (2014)"},{"issue":"28","key":"10285_CR6","first-page":"1","volume":"18","author":"A Bouchard-C\u00f4t\u00e9","year":"2017","unstructured":"Bouchard-C\u00f4t\u00e9, A., Doucet, A., Roth, A.: Particle Gibbs split-merge sampling for Bayesian inference in mixture models. J. Mach. Learn. Res. 18(28), 1\u201339 (2017)","journal-title":"J. Mach. Learn. Res."},{"key":"10285_CR7","unstructured":"Chen, T., Fox, E., Guestrin, C.: Stochastic gradient Hamiltonian Monte Carlo. In: Proceedings of the 31st International Conference on Machine Learning, vol. 32, pp. 1683\u20131691. PMLR (2014)"},{"key":"10285_CR8","unstructured":"de Freitas, N.: Bayesian methods for neural networks. PhD thesis, University of Cambridge (1999)"},{"key":"10285_CR9","doi-asserted-by":"crossref","unstructured":"de Freitas, N., Andrieu, C., H\u00f8jen-S\u00f8rensen, P., Niranjan, M., Gee, A.: Sequential Monte Carlo methods for neural networks, pp. 359\u2013379. Springer, New York (2001)","DOI":"10.1007\/978-1-4757-3437-9_17"},{"issue":"2","key":"10285_CR10","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1111\/j.1467-9868.2010.00765.x","volume":"73","author":"M Girolami","year":"2011","unstructured":"Girolami, M., Calderhead, B.: Riemann manifold Langevin and Hamiltonian Monte Carlo methods. J. Roy. Stat. Soc. Ser. B (Stat. Methodol.) 73(2), 123\u2013214 (2011)","journal-title":"J. Roy. Stat. Soc. Ser. B (Stat. Methodol.)"},{"key":"10285_CR11","unstructured":"Gong, W., Li, Y., Hern\u00e1ndez-Lobato, J.M.: Meta-learning for stochastic gradient MCMC. In: International Conference on Learning Representations. PMLR (2019)"},{"key":"10285_CR12","unstructured":"Grathwohl, W., Swersky, K., Hashemi, M., Duvenaud, D., Maddison, C.: Oops i took a gradient: scalable sampling for discrete distributions. In: Proceedings of the 38th International Conference on Machine Learning, vol. 139, pp. 3831\u20133841. PMLR (2021)"},{"key":"10285_CR13","volume-title":"The Elements of Statistical Learning: Data Mining, Inference and Prediction","author":"T Hastie","year":"2016","unstructured":"Hastie, T., Tibshirani, R., Friedman, J.: The Elements of Statistical Learning: Data Mining, Inference and Prediction, 2nd edn. Springer, New York (2016)","edition":"2"},{"key":"10285_CR14","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition, pp. 770\u2013778 (2016)","DOI":"10.1109\/CVPR.2016.90"},{"key":"10285_CR15","unstructured":"Izmailov, P., Vikram, S., Hoffman, M.D., Wilson, A.G.G.: What are Bayesian neural network posteriors really like? In: Proceedings of the 38th International Conference on Machine Learning, vol. 139, pp. 4629\u20134640. PMLR, Vienna (2021)"},{"key":"10285_CR16","unstructured":"Krizhevsky, A., Hinton, G.: Learning multiple layers of features from tiny images. Technical report. University of Toronto, Toronto (2009)"},{"issue":"1","key":"10285_CR17","doi-asserted-by":"publisher","first-page":"66","DOI":"10.1214\/11-AAP806","volume":"23","author":"K \u0141atuszy\u0144ski","year":"2013","unstructured":"\u0141atuszy\u0144ski, K., Roberts, G.O., Rosenthal, J.S.: Adaptive Gibbs samplers and related MCMC methods. Ann. Appl. Probab. 23(1), 66\u201398 (2013)","journal-title":"Ann. Appl. Probab."},{"issue":"11","key":"10285_CR18","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","volume":"86","author":"Y Lecun","year":"1998","unstructured":"Lecun, Y., Bottou, L., Bengio, Y., Haffner, P.: Gradient-based learning applied to document recognition. Proc. IEEE 86(11), 2278\u20132324 (1998)","journal-title":"Proc. IEEE"},{"issue":"157","key":"10285_CR19","first-page":"1","volume":"22","author":"T Matsubara","year":"2021","unstructured":"Matsubara, T., Oates, C.J., Briol, F.-X.: The ridgelet prior: a covariance function approach to prior specification for Bayesian neural networks. J. Mach. Learn. Res. 22(157), 1\u201357 (2021)","journal-title":"J. Mach. Learn. Res."},{"key":"10285_CR20","volume-title":"Perceptrons","author":"ML Minsky","year":"1988","unstructured":"Minsky, M.L., Papert, S.A.: Perceptrons. MIT Press, Cambridge (1988)"},{"key":"10285_CR21","unstructured":"Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., Ng, A.Y.: Reading digits in natural images with unsupervised feature learning. In: NIPS Workshop on Deep Learning and Unsupervised Feature Learning (2011)"},{"issue":"3","key":"10285_CR22","doi-asserted-by":"publisher","first-page":"425","DOI":"10.1214\/21-STS840","volume":"37","author":"T Papamarkou","year":"2022","unstructured":"Papamarkou, T., Hinkle, J., Young, M.T., Womble, D.: Challenges in Markov chain Monte Carlo for Bayesian neural networks. Stat. Sci. 37(3), 425\u2013442 (2022)","journal-title":"Stat. Sci."},{"issue":"2","key":"10285_CR23","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1111\/1467-9868.00070","volume":"59","author":"GO Roberts","year":"1997","unstructured":"Roberts, G.O., Sahu, S.K.: Updating schemes, correlation structure, blocking and parameterization for the Gibbs sampler. J. Roy. Stat. Soc. Ser. B (Stat. Methodol.) 59(2), 291\u2013317 (1997)","journal-title":"J. Roy. Stat. Soc. Ser. B (Stat. Methodol.)"},{"issue":"6","key":"10285_CR24","doi-asserted-by":"publisher","first-page":"386","DOI":"10.1037\/h0042519","volume":"65","author":"F Rosenblatt","year":"1958","unstructured":"Rosenblatt, F.: The perceptron: a probabilistic model for information storage and organization in the brain. Psychol. Rev. 65(6), 386 (1958)","journal-title":"Psychol. Rev."},{"key":"10285_CR25","unstructured":"Saul, L., Jordan, M.: Exploiting tractable substructures in intractable networks. In: Advances in Neural Information Processing Systems, vol. 8. MIT Press, Denver (1995)"},{"issue":"1","key":"10285_CR26","doi-asserted-by":"publisher","first-page":"128","DOI":"10.1214\/088342304000000099","volume":"19","author":"DM Titterington","year":"2004","unstructured":"Titterington, D.M.: Bayesian methods for neural networks and related models. Stat. Sci. 19(1), 128\u2013139 (2004)","journal-title":"Stat. Sci."},{"issue":"74","key":"10285_CR27","first-page":"1","volume":"23","author":"B-H Tran","year":"2022","unstructured":"Tran, B.-H., Rossi, S., Milios, D., Filippone, M.: All you need is a good functional prior for Bayesian deep learning. J. Mach. Learn. Res. 23(74), 1\u201356 (2022)","journal-title":"J. Mach. Learn. Res."},{"issue":"6","key":"10285_CR28","doi-asserted-by":"publisher","first-page":"1648","DOI":"10.1109\/TSP.2019.2894825","volume":"67","author":"M Vono","year":"2019","unstructured":"Vono, M., Dobigeon, N., Chainais, P.: Split-and-augmented Gibbs sampler-application to large-scale inference problems. IEEE Trans. Signal Process. 67(6), 1648\u20131661 (2019)","journal-title":"IEEE Trans. Signal Process."},{"issue":"25","key":"10285_CR29","first-page":"1","volume":"23","author":"M Vono","year":"2022","unstructured":"Vono, M., Paulin, D., Doucet, A.: Efficient MCMC sampling with dimension-free convergence rate using ADMM-type splitting. J. Mach. Learn. Res. 23(25), 1\u201369 (2022)","journal-title":"J. Mach. Learn. Res."},{"key":"10285_CR30","unstructured":"Welling, M., Teh, Y.W.: Bayesian learning via stochastic gradient Langevin dynamics. In: Proceedings of the 28th International Conference on Machine Learning, pp. 681\u2013688 (2011)"},{"key":"10285_CR31","unstructured":"Wenzel, F., Roth, K., Veeling, B., Swiatkowski, J., Tran, L., Mandt, S., Snoek, J., Salimans, T., Jenatton, R., Nowozin, S.: How good is the Bayes posterior in deep neural networks really? In: Proceedings of the 37th International Conference on Machine Learning, vol. 119, pp. 10248\u201310259. PMLR, Vienna (2020)"},{"key":"10285_CR32","doi-asserted-by":"crossref","unstructured":"Wiese, J.G., Wimmer, L., Papamarkou, T., Bischl, B., G\u00fcnnemann, S., R\u00fcgamer, D.: Towards efficient MCMC sampling in Bayesian neural networks by exploiting symmetry. In: European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases. Springer, Turin (2023)","DOI":"10.1007\/978-3-031-43412-9_27"},{"key":"10285_CR33","unstructured":"Xiao, H., Rasul, K., Vollgraf, R.: Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747 (2017)"},{"key":"10285_CR34","unstructured":"Zhang, R., Li, C., Zhang, J., Chen, C., Wilson, A.G.: Cyclical stochastic gradient MCMC for Bayesian deep learning. In: International Conference on Learning Representations (2020)"}],"container-title":["Statistics and Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11222-023-10285-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11222-023-10285-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11222-023-10285-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,12,18]],"date-time":"2023-12-18T16:23:20Z","timestamp":1702916600000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11222-023-10285-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,10]]},"references-count":34,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2023,10]]}},"alternative-id":["10285"],"URL":"https:\/\/doi.org\/10.1007\/s11222-023-10285-5","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-2444402\/v1","asserted-by":"object"}]},"ISSN":["0960-3174","1573-1375"],"issn-type":[{"value":"0960-3174","type":"print"},{"value":"1573-1375","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,10]]},"assertion":[{"value":"4 January 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 July 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 August 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"119"}}