{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T15:36:56Z","timestamp":1780673816239,"version":"3.54.1"},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,7]]},"abstract":"<jats:p>8-bit integer inference, as a promising direction in reducing both the latency and storage of deep neural networks, has made great progress recently. On the other hand, previous systems still rely on 32-bit floating point for certain functions in complex models (e.g., Softmax in Transformer), and make heavy use of quantization and de-quantization. In this work, we show that after a principled modification on the Transformer architecture, dubbed Integer Transformer, an (almost) fully 8-bit integer inference algorithm Scale Propagation could be derived. De-quantization is adopted when necessary, which makes the network more efficient. Our experiments on WMT16 En&lt;-&gt;Ro, WMT14 En&lt;-&gt;De and En-&gt;Fr translation tasks as well as the WikiText-103 language modelling task show that the fully 8-bit Transformer system achieves comparable performance with the floating point baseline but requires nearly 4x less memory footprint.<\/jats:p>","DOI":"10.24963\/ijcai.2020\/520","type":"proceedings-article","created":{"date-parts":[[2020,7,8]],"date-time":"2020-07-08T12:12:10Z","timestamp":1594210330000},"page":"3759-3765","source":"Crossref","is-referenced-by-count":16,"title":["Towards Fully 8-bit Integer Inference for the Transformer Model"],"prefix":"10.24963","author":[{"given":"Ye","family":"Lin","sequence":"first","affiliation":[{"name":"Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanyang","family":"Li","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tengbo","family":"Liu","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tong","family":"Xiao","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang, China"},{"name":"NiuTrans Research, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tongran","family":"Liu","sequence":"additional","affiliation":[{"name":"CAS Key Laboratory of Behavioral Science, Institute of Psychology, CAS, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jingbo","family":"Zhu","sequence":"additional","affiliation":[{"name":"Northeastern University, Shenyang, China"},{"name":"NiuTrans Research, Shenyang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"10584","event":{"name":"Twenty-Ninth International Joint Conference on Artificial Intelligence and Seventeenth Pacific Rim International Conference on Artificial Intelligence {IJCAI-PRICAI-20}","theme":"Artificial Intelligence","location":"Yokohama, Japan","acronym":"IJCAI-PRICAI-2020","number":"28","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"start":{"date-parts":[[2020,7,11]]},"end":{"date-parts":[[2020,7,17]]}},"container-title":["Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2020,7,9]],"date-time":"2020-07-09T02:15:41Z","timestamp":1594260941000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2020\/520"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2020,7]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2020\/520","relation":{},"subject":[],"published":{"date-parts":[[2020,7]]}}}