{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T10:57:44Z","timestamp":1777546664469,"version":"3.51.4"},"reference-count":17,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2009,9,1]],"date-time":"2009-09-01T00:00:00Z","timestamp":1251763200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["CNS-0725354"],"award-info":[{"award-number":["CNS-0725354"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2009,9]]},"abstract":"<jats:p>Lithography simulation, an essential step in design for manufacturability (DFM), is still far from computationally efficient. Most leading companies use large clusters of server computers to achieve acceptable turn-around time. Thus coprocessor acceleration is very attractive for obtaining increased computational performance with a reduced power consumption. This article describes the implementation of a customized accelerator on FPGA using a polygon-based simulation model. An application-specific memory partitioning scheme is designed to meet the bandwidth requirements for a large number of processing elements. Deep loop pipelining and ping-pong buffer based function block pipelining are also implemented in our design. Initial results show a 15X speedup versus the software implementation running on a microprocessor, and more speedup is expected via further performance tuning. The implementation also leverages state-of-art C-to-RTL synthesis tools. At the same time, we also identify the need for manual architecture-level exploration for parallel implementations. Moreover, we implement the algorithm on NVIDIA GPUs using the CUDA programming environment, and provide some useful comparisons for different kinds of accelerators.<\/jats:p>","DOI":"10.1145\/1575774.1575776","type":"journal-article","created":{"date-parts":[[2009,10,6]],"date-time":"2009-10-06T18:18:59Z","timestamp":1254853139000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":28,"title":["FPGA-Based Hardware Acceleration of Lithographic Aerial Image Simulation"],"prefix":"10.1145","volume":"2","author":[{"given":"Jason","family":"Cong","sequence":"first","affiliation":[{"name":"University of California, Los Angeles"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yi","family":"Zou","sequence":"additional","affiliation":[{"name":"University of California, Los Angeles"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2009,9]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of SPIE: Optical Microlithography XVIII.","volume":"5754","author":"Cao Y.","unstructured":"Cao , Y. , Lu , Y.-W. , Chen , L. , and Ye , J . 2004. Optimized hardware and software for fast full-chip simulation . In Proceedings of SPIE: Optical Microlithography XVIII. Vol. 5754 , 407--414. Cao, Y., Lu, Y.-W., Chen, L., and Ye, J. 2004. Optimized hardware and software for fast full-chip simulation. In Proceedings of SPIE: Optical Microlithography XVIII. Vol. 5754, 407--414."},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of SPIE: Optical\/Laser Microlithography VIII.","volume":"2440","author":"Cobb N. B.","unstructured":"Cobb , N. B. and Zakhor , A . 1995. Fast, low-complexity mask design . In Proceedings of SPIE: Optical\/Laser Microlithography VIII. Vol. 2440 , T. A. Brunner, Ed. 313--327. Cobb, N. B. and Zakhor, A. 1995. Fast, low-complexity mask design. In Proceedings of SPIE: Optical\/Laser Microlithography VIII. Vol. 2440, T. A. Brunner, Ed. 313--327."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1344671.1344683"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the IEEE International Symposium on Circuits and Systems (ISCAS\u201999)","author":"Doggett M.","unstructured":"Doggett , M. and Meissner , M . 1999. A memory addressing and access design for real time volume rendering . In Proceedings of the IEEE International Symposium on Circuits and Systems (ISCAS\u201999) . 344--347. Doggett, M. and Meissner, M. 1999. A memory addressing and access design for real time volume rendering. In Proceedings of the IEEE International Symposium on Circuits and Systems (ISCAS\u201999). 344--347."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2004.840301"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1117\/12.584441"},{"key":"e_1_2_1_8_1","unstructured":"Mencer O. and Clapp R. G. 2007. Accelerating 2D FFTs and convolutions for seismic processing. Brief notes Maxeler Technologies.  Mencer O. and Clapp R. G. 2007. Accelerating 2D FFTs and convolutions for seismic processing. Brief notes Maxeler Technologies."},{"key":"e_1_2_1_9_1","volume-title":"Datasheet of Calibre nmOPC","author":"Mentor","unstructured":"Mentor . 2004. Datasheet of Calibre nmOPC . Mentor Graphics Corporation . Mentor. 2004. Datasheet of Calibre nmOPC. Mentor Graphics Corporation."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1364\/JOSAA.11.002438"},{"key":"e_1_2_1_11_1","unstructured":"Podlozhnyuk V. 2007. FFT-based 2D convolution. NVIDIA white paper.  Podlozhnyuk V. 2007. FFT-based 2D convolution. NVIDIA white paper."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2004.835148"},{"key":"e_1_2_1_13_1","first-page":"283","article-title":"FPGA implementations of fast fourier transforms for real-time signal and image processing. IEEE Proc. Vision, Image","volume":"152","author":"Uzun I.","year":"2005","unstructured":"Uzun , I. , Amira , A. , and Bouridane , A. 2005 . FPGA implementations of fast fourier transforms for real-time signal and image processing. IEEE Proc. Vision, Image , Signal Process. 152 , 3, 283 -- 296 . Uzun, I., Amira, A., and Bouridane, A. 2005. FPGA implementations of fast fourier transforms for real-time signal and image processing. IEEE Proc. Vision, Image, Signal Process. 152, 3, 283--296.","journal-title":"Signal Process."},{"key":"e_1_2_1_14_1","volume-title":"-C","author":"Wang Y.-T.","year":"2006","unstructured":"Wang , Y.-T. , Tsai , C.-M. , and Chang , F . -C . 2006 . Lithographic simulations using graphical processing units. United States Patent Application 20060242618. Wang, Y.-T., Tsai, C.-M., and Chang, F.-C. 2006. Lithographic simulations using graphical processing units. United States Patent Application 20060242618."},{"key":"e_1_2_1_15_1","volume-title":"Optical Imaging in Projection Microlithography","author":"Wong A. K.-K.","unstructured":"Wong , A. K.-K. 2005. Optical Imaging in Projection Microlithography . SPIE Press , Bellingham, WA . Wong, A. K.-K. 2005. Optical Imaging in Projection Microlithography. SPIE Press, Bellingham, WA."},{"key":"e_1_2_1_16_1","volume-title":"Private communication","author":"Wong A. K.-K.","unstructured":"Wong , A. K.-K. 2007. Private communication . Magma Design Automation Inc . Wong, A. K.-K. 2007. Private communication. Magma Design Automation Inc."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1117\/12.485433"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design (ICCAD\u201907)","author":"Yu P.","unstructured":"Yu , P. and Pan , D. Z . 2007. A novel intensity based optical proximity correction algorithm with speedup in lithography simulation . In Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design (ICCAD\u201907) . 854--859. Yu, P. and Pan, D. Z. 2007. A novel intensity based optical proximity correction algorithm with speedup in lithography simulation. In Proceedings of the IEEE\/ACM International Conference on Computer-Aided Design (ICCAD\u201907). 854--859."}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1575774.1575776","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1575774.1575776","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T12:23:08Z","timestamp":1750249388000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1575774.1575776"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,9]]},"references-count":17,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2009,9]]}},"alternative-id":["10.1145\/1575774.1575776"],"URL":"https:\/\/doi.org\/10.1145\/1575774.1575776","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,9]]},"assertion":[{"value":"2008-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2008-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-09-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}