{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T23:33:08Z","timestamp":1780356788887,"version":"3.54.1"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"name":"National Science Foundation","award":["CNS-1616947"],"award-info":[{"award-number":["CNS-1616947"]}]},{"name":"National Science Foundation CAREER","award":["CCF-2045973"],"award-info":[{"award-number":["CCF-2045973"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>\n            Markov-Chain Monte-Carlo (MCMC) algorithms offer a general framework for performing interpretable inference but have high overheads due to the computational complexity of the sampling process and the large number of samples required to produce an accurate result. Computer Vision is a common class of workloads that can be performed using MCMC methods. As computer vision workloads trend toward high-resolution real-time inference, it becomes challenging to perform inference in contexts such as edge computing, which operates under strict power and area budgets. Previous work explores hardware techniques for efficient sampling; however, MCMC algorithms still require many samples. We reduce the overheads of Gibbs Sampling, an MCMC algorithm, using an approach we call mixed-resolution sampling. This approach uses low-resolution inference to provide a starting point for full-resolution sampling. We evaluate this approach on three important computer vision tasks: stereo matching, optical flow, and blind source separation. Mixed-resolution sampling reduces root mean square error (RMSE) by an average of 19.6% for stereo-matching tasks, 13% for optical flow tasks, and 6.3% for blind source separation relative to traditional Gibbs Sampling. To enable real-time, explainable MCMC inference under edge power constraints, we exploit the structure of mixed-resolution sampling to architect and implement a hardware-software co-designed accelerator architecture, BigLittleMCA (\n            <jats:underline>Big<\/jats:underline>\n            -\n            <jats:underline>Little<\/jats:underline>\n            <jats:underline>MC<\/jats:underline>\n            MC\n            <jats:underline>A<\/jats:underline>\n            ccelerator). BigLittleMCA is a tiled MCMC accelerator architecture that uses a small sampler for low-resolution sampling and a large sampler for full-resolution sampling. Our results show that the architecture sustains real-time 720p inference at 30 FPS (frames per second) using 48.5% less power than prior work.\n          <\/jats:p>","DOI":"10.1145\/3736171","type":"journal-article","created":{"date-parts":[[2025,5,20]],"date-time":"2025-05-20T07:20:27Z","timestamp":1747725627000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["BigLittleMCA: A Spatially-Optimal Tiled Hardware Accelerator for MCMC Image Processing"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4792-1910","authenticated-orcid":false,"given":"Chris","family":"Kjellqvist","sequence":"first","affiliation":[{"name":"Computer Science, Duke University","place":["Durham, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3574-3440","authenticated-orcid":false,"given":"Lisa","family":"Wills","sequence":"additional","affiliation":[{"name":"Computer Science and ECE, Duke University","place":["Durham, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1893-5464","authenticated-orcid":false,"given":"Alvin","family":"Lebeck","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Duke University","place":["Durham, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,17]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2013. big.LITTLE Technology: The Future of Mobile. Retrieved from https:\/\/armkeil.blob.core.windows.net\/developer\/Files\/pdf\/white-paper\/big-little-technology-the-future-of-mobile.pdf"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00126"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1080\/17415977.2021.1880398"},{"key":"e_1_3_1_5_2","unstructured":"Alphabet. 2023. Edge TPU. Retrieved from https:\/\/cloud.google.com\/edge-tpu"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/2228360.2228584"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2007.4408903"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304019"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00054836"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-021-00423-x"},{"key":"e_1_3_1_11_2","unstructured":"Ramin Bashizade Xiangyu Zhang Sayan Mukherjee and Alvin R. Lebeck. 2021. Accelerating markov random field inference with uncertainty quantification. arXiv:2108.00570. Retrieved from https:\/\/arxiv.org\/abs\/2108.00570 (2021)."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA53966.2022.00012"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_1_14_2","first-page":"431","volume-title":"Proceedings of the Machine Learning and Systems","volume":"1","author":"Chin Ting-Wu","year":"2019","unstructured":"Ting-Wu Chin, Ruizhou Ding, and Diana Marculescu. 2019. AdaScale: Towards real-time video object detection using adaptive scaling. In Proceedings of the Machine Learning and Systems, A. Talwalkar, V. Smith, and M. Zaharia (Eds.), Vol. 1. 431\u2013441. Retrieved from https:\/\/proceedings.mlsys.org\/paper\/2019\/file\/b1d10e7bafa4421218a51b1e1f1b0ba2-Paper.pdf"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.3390\/s17071680"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(90)90060-D"},{"key":"e_1_3_1_17_2","unstructured":"Intel Corporation. 2022. Intel\u00ae Movidius\u2122 Myriad\u2122 X Vision Processing Unit (VPU) with Neural Compute Engine. Retrieved from https:\/\/intel.ly\/3jiGwpl"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","unstructured":"Jifeng Dai Haozhi Qi Yuwen Xiong Yi Li Guodong Zhang Han Hu and Yichen Wei. 2017. Deformable Convolutional Networks. DOI:10.48550\/ARXIV.1703.06211","DOI":"10.48550\/ARXIV.1703.06211"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","unstructured":"Javier Felip Nilesh Ahuja and Omesh Tickoo. 2019. Tree Pyramidal Adaptive Importance Sampling. DOI:10.48550\/ARXIV.1912.08434","DOI":"10.48550\/ARXIV.1912.08434"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSI.2023.3337529"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1214\/ss\/1177011136"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","unstructured":"Ross Girshick. 2015. Fast R-CNN. DOI:10.48550\/ARXIV.1504.08083","DOI":"10.48550\/ARXIV.1504.08083"},{"key":"e_1_3_1_23_2","unstructured":"Google. 2022. Retrieved from https:\/\/coral.ai\/docs\/edgetpu\/benchmarks\/"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISLPED58423.2023.10244461"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10578-9_23"},{"key":"e_1_3_1_26_2","first-page":"5704","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Hu Yinlin","year":"2016","unstructured":"Yinlin Hu, Rui Song, and Yunsong Li. 2016. Efficient coarse-to-fine patchmatch for large displacement optical flow. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5704\u20135712."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS51168.2021.9635869"},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1007\/978-3-030-69532-3_4","volume-title":"Proceedings of the Computer Vision \u2013 ACCV 2020","author":"Huang Bowen","year":"2021","unstructured":"Bowen Huang, Jinjia Zhou, Xiao Yan, Ming\u2019e Jing, Rentao Wan, and Yibo Fan. 2021. CS-MCNet: A video compressive sensing reconstruction network with interpretable motion compensation. In Proceedings of the Computer Vision \u2013 ACCV 2020, Hiroshi Ishikawa, Cheng-Lin Liu, Tomas Pajdla, and Jianbo Shi (Eds.). Springer International Publishing, Cham, 54\u201367."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00936"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2015.4"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1006\/ijhc.1995.1029"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1086\/302524"},{"key":"e_1_3_1_33_2","article-title":"On-chip networks: Second edition","author":"Jerger Natalie Enright","year":"2017","unstructured":"Natalie Enright Jerger, Tushar Krishna, and Li-Shiuan Peh. 2017. On-chip networks: Second edition. Morgan and Claypool Publishers (2017).","journal-title":"Morgan and Claypool Publishers"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASIC.2007.4415786"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","unstructured":"Norman P. Jouppi Cliff Young Nishant Patil David Patterson Gaurav Agrawal Raminder Bajwa Sarah Bates Suresh Bhatia Nan Boden Al Borchers Al Borchers Rick Boyle Pierre-luc Cantin Clifford Chao Chris Clark Jeremy Coriell Mike Daley Matt Dau Jeffrey Dean Ben Gelb Tara Vazir Ghaemmaghami Rajendra Gottipati William Gulland Robert Hagmann C. Richard Ho Doug Hogberg John Hu Robert Hundt Dan Hurt Julian Ibarz Aaron Jaffey Alek Jaworski Alexander Kaplan Harshit Khaitan Daniel Killebrew Andy Koch Naveen Kumar Steve Lacy James Laudon James Law Diemthu Le Chris Leary Zhuyuan Liu Kyle Lucke Alan Lundin Gordon MacKean Adriana Maggiore Maire Mahony Kieran Miller Rahul Nagarajan Ravi Narayanaswami Ray Ni Kathy Nix Thomas Norrie Mark Omernick Narayana Penukonda Andy Phelps Jonathan Ross Matt Ross Amir Salek Emad Samadiani Chris Severn Gregory Sizikov Matthew Snelham Jed Souter Dan Steinberg Andy Swing Mercedes Tan Gregory Thorson Bo Tian Horia Toma Erick Tuttle Vijay Vasudevan Richard Walter Walter Wang Eric Wilcox and Doe Hyun Yoon.2017. In-Datacenter Performance Analysis of a Tensor Processing Unit. In Proceedings of the 44th Annual International Symposium on Computer Architecture. DOI:10.48550\/ARXIV.1704.04760","DOI":"10.48550\/ARXIV.1704.04760"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRITO.2015.7359323"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2009.2012905"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2009.5414190"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2023.3238584"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-0761-0"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2019.00075"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2019.00033"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/34.161350"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-405888-0.00007-6"},{"issue":"5","key":"e_1_3_1_45_2","article-title":"BlockNet: A deep neural network for block-based motion estimation using representative matching","volume":"12","author":"Lee Junggi","year":"2020","unstructured":"Junggi Lee, Kyeongbo Kong, Gyujin Bae, and Woo-Jin Song. 2020. BlockNet: A deep neural network for block-based motion estimation using representative matching. Symmetry 12, 5 (2020). https:\/\/www.mdpi.com\/about\/announcements\/784","journal-title":"Symmetry"},{"key":"e_1_3_1_46_2","article-title":"Monte carlo strategies in scientific computing","author":"Liu Jun S.","year":"2001","unstructured":"Jun S. Liu. 2001. Monte carlo strategies in scientific computing. Springer Publishing Company, Incorporated (2001).","journal-title":"Springer Publishing Company, Incorporated"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2016.2630682"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL50879.2020.00054"},{"key":"e_1_3_1_49_2","unstructured":"Hasan Mahmud Mashrur M. Morshed and Md Kamrul Hasan. 2021. A deep learning-based multimodal depth-aware dynamic hand gesture recognition system. arXiv:2107.02543. Retrieved from https:\/\/arxiv.org\/abs\/2107.02543. (2021)."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.456"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2014.7478823"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/DCABES.2018.00035"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2007.383191"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/SMBV.2001.988771"},{"issue":"387","key":"e_1_3_1_55_2","doi-asserted-by":"crossref","first-page":"609","DOI":"10.1080\/01621459.1984.10478087","article-title":"Smoothness priors and nonlinear regression","volume":"79","author":"Shiller Robert J","year":"1984","unstructured":"Robert J Shiller. 1984. Smoothness priors and nonlinear regression. J. Amer. Statist. Assoc. 79, 387 (1984), 609\u2013615.","journal-title":"J. Amer. Statist. Assoc."},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.vlsi.2017.02.002"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2019.8875669"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.55"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00196"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00034"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446697"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3736171","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,17]],"date-time":"2025-09-17T13:43:36Z","timestamp":1758116616000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3736171"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,17]]},"references-count":60,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3736171"],"URL":"https:\/\/doi.org\/10.1145\/3736171","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,17]]},"assertion":[{"value":"2024-07-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-11","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}