{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T18:39:17Z","timestamp":1770143957607,"version":"3.49.0"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,12,1]],"date-time":"2021-12-01T00:00:00Z","timestamp":1638316800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,12,13]],"date-time":"2021-12-13T00:00:00Z","timestamp":1639353600000},"content-version":"vor","delay-in-days":12,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Background<\/jats:title><jats:p>Clinical notes are documents that contain detailed information about the health status of patients. Medical codes generally accompany them. However, the manual diagnosis is costly and error-prone. Moreover, large datasets in clinical diagnosis are susceptible to noise labels because of erroneous manual annotation. Therefore, machine learning has been utilized to perform automatic diagnoses. Previous state-of-the-art (SOTA) models used convolutional neural networks to build document representations for predicting medical codes. However, the clinical notes are usually long-tailed. Moreover, most models fail to deal with the noise during code allocation. Therefore, denoising mechanism and long-tailed classification are the keys to automated coding at scale.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>In this paper, a new joint learning model is proposed to extend our attention model for predicting medical codes from clinical notes. On the MIMIC-III-50 dataset, our model outperforms all the baselines and SOTA models in all quantitative metrics. On the MIMIC-III-full dataset, our model outperforms in the macro-F1, micro-F1, macro-AUC, and precision at eight compared to the most advanced models. In addition, after introducing the denoising mechanism, the convergence speed of the model becomes faster, and the loss of the model is reduced overall.<\/jats:p><\/jats:sec><jats:sec><jats:title>Conclusions<\/jats:title><jats:p>The innovations of our model are threefold: firstly, the code-specific representation can be identified by adopted the self-attention mechanism and the label attention mechanism. Secondly, the performance of the long-tailed distributions can be boosted by introducing the joint learning mechanism. Thirdly, the denoising mechanism is suitable for reducing the noise effects in medical code prediction. Finally, we evaluate the effectiveness of our model on the widely-used MIMIC-III datasets and achieve new SOTA results.<\/jats:p><\/jats:sec>","DOI":"10.1186\/s12859-021-04520-x","type":"journal-article","created":{"date-parts":[[2021,12,13]],"date-time":"2021-12-13T12:13:54Z","timestamp":1639397634000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["JLAN: medical code prediction via joint learning attention networks and denoising mechanism"],"prefix":"10.1186","volume":"22","author":[{"given":"Xingwang","family":"Li","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yijia","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Faiz ul","family":"Islam","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Deshi","family":"Dong","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Wei","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mingyu","family":"Lu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,12,13]]},"reference":[{"issue":"1","key":"4520_CR1","doi-asserted-by":"publisher","first-page":"24","DOI":"10.1038\/s41591-018-0316-z","volume":"25","author":"A Esteva","year":"2019","unstructured":"Esteva A, Robicquet A, Ramsundar B, Kuleshov V, DePristo M, Chou K, et al. A guide to deep learning in healthcare. Nat Med. 2019;25(1):24\u20139.","journal-title":"Nat Med"},{"key":"4520_CR2","doi-asserted-by":"crossref","unstructured":"Xie\u2020 P, Shi\u00a7 H, Ming Z, Xing\u2020 E, editors. A neural architecture for automated ICD coding. Meeting of the Association for Computational Linguistics; 2018.","DOI":"10.18653\/v1\/P18-1098"},{"issue":"1","key":"4520_CR3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/sdata.2016.35","volume":"3","author":"AE Johnson","year":"2016","unstructured":"Johnson AE, Pollard TJ, Shen L, Li-Wei HL, Feng M, Ghassemi M, et al. MIMIC-III, a freely accessible critical care database. Sci Data. 2016;3(1):1\u20139.","journal-title":"Sci Data"},{"key":"4520_CR4","unstructured":"Zhang C, Be Ngio S, Hardt M, Recht B, Vinyals O. Understanding deep learning requires rethinking generalization. 2016."},{"key":"4520_CR5","unstructured":"Thulasidasan S, Bhattacharya T, Bilmes J, Chennupati G, Mohd-Yusof J. Combating label noise in deep learning using abstention. arXiv preprint arXiv:1905.10964. 2019."},{"issue":"3","key":"4520_CR6","doi-asserted-by":"publisher","first-page":"204","DOI":"10.1136\/adc.2007.128132","volume":"93","author":"JE Sheppard","year":"2008","unstructured":"Sheppard JE, Weidner LC, Zakai S, Fountain-Polley S, Williams J. Ambiguous abbreviations: an audit of abbreviations in paediatric note keeping. Arch Dis Child. 2008;93(3):204\u20136.","journal-title":"Arch Dis Child"},{"key":"4520_CR7","doi-asserted-by":"crossref","unstructured":"Farkas R, Szarvas G, editors. Automatic construction of rule-based ICD-9-CM coding systems. BMC Bioinform; 2008: Springer.","DOI":"10.1186\/1471-2105-9-S3-S10"},{"key":"4520_CR8","doi-asserted-by":"crossref","unstructured":"Li F, Yu H. ICD Coding from clinical text using multi-filter residual convolutional neural network. 2019.","DOI":"10.1609\/aaai.v34i05.6331"},{"key":"4520_CR9","unstructured":"Byrd J, Lipton Z, editors. What is the effect of importance weighting in deep learning? International Conference on Machine Learning; 2019: PMLR."},{"key":"4520_CR10","doi-asserted-by":"crossref","unstructured":"Zhou B, Cui Q, Wei X-S, Chen Z-M, editors. BBN: Bilateral-branch network with cumulative learning for long-tailed visual recognition. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition; 2020.","DOI":"10.1109\/CVPR42600.2020.00974"},{"key":"4520_CR11","doi-asserted-by":"publisher","first-page":"112887","DOI":"10.1016\/j.eswa.2019.112887","volume":"140","author":"RS Sreepada","year":"2020","unstructured":"Sreepada RS, Patra BK. Mitigating long tail effect in recommendations using few shot learning technique. Expert Syst Appl. 2020;140:112887.","journal-title":"Expert Syst Appl"},{"issue":"1","key":"4520_CR12","first-page":"1","volume":"27","author":"H Azarbonyad","year":"2020","unstructured":"Azarbonyad H, Dehghani M, Marx M, Kamps J. Learning to rank for multi-label text classification: combining different sources of information. Nat Lang Eng. 2020;27(1):1\u201323.","journal-title":"Nat Lang Eng"},{"key":"4520_CR13","first-page":"1","volume":"99","author":"H Dong","year":"2020","unstructured":"Dong H, Wang W, Huang K, Coenen F. Automated social text annotation with joint multi-label attention networks. IEEE Trans Neural Netw Learn Syst. 2020;99:1\u201315.","journal-title":"IEEE Trans Neural Netw Learn Syst"},{"issue":"1","key":"4520_CR14","doi-asserted-by":"publisher","first-page":"89","DOI":"10.1017\/S1351324920000029","volume":"27","author":"H Azarbonyad","year":"2021","unstructured":"Azarbonyad H, Dehghani M, Marx M, Kamps J. Learning to rank for multi-label text classification: combining different sources of information. Nat Lang Eng. 2021;27(1):89\u2013111.","journal-title":"Nat Lang Eng"},{"key":"4520_CR15","unstructured":"Shi H, Xie P, Hu Z, Zhang M, Xing EP. Towards automated ICD coding using deep learning. 2017."},{"key":"4520_CR16","unstructured":"Baumel T, Nassour-Kassis J, Elhadad M, Elhadad N. Multi-label classification of patient notes a case study on ICD code assignment. 2017."},{"key":"4520_CR17","doi-asserted-by":"crossref","unstructured":"Wang G, Li C, Wang W, Zhang Y, Shen D, Zhang X, et al. Joint embedding of words and labels for text classification. arXiv preprint arXiv:1805.04174. 2018.","DOI":"10.18653\/v1\/P18-1216"},{"key":"4520_CR18","doi-asserted-by":"crossref","unstructured":"Mullenbach J, Wiegreffe S, Duke J, Sun J, Eisenstein J, editors. Explainable prediction of medical codes from clinical text. In: Proceedings of the 2018 conference of the north american chapter of the association for computational linguistics: Human Language Technologies, Volume 1 (Long Papers); 2018.","DOI":"10.18653\/v1\/N18-1100"},{"key":"4520_CR19","doi-asserted-by":"crossref","unstructured":"Bai T, Vucetic S. Improving medical code prediction from clinical text via incorporating online knowledge sources. The World Wide Web Conference; San Francisco, CA, USA: Association for Computing Machinery; 2019. p. 72\u201382.","DOI":"10.1145\/3308558.3313485"},{"key":"4520_CR20","unstructured":"Mikolov T, Sutskever I, Chen K, Corrado G, Dean J. Distributed representations of words and phrases and their compositionality arXiv: 1310.4546v1[cs.CL] 16 Oct 2013. 2013."},{"key":"4520_CR21","doi-asserted-by":"crossref","unstructured":"Murphy GS, Kopman AF. Neostigmine as an antagonist of residual block: best practices do not guarantee predictable results. BJA Br J Anaesthesia. 2018;121:S0007091218303842.","DOI":"10.1016\/j.bja.2018.05.003"},{"key":"4520_CR22","unstructured":"Zhou P, Qi Z, Zheng S, Xu J, Bao H, Xu B. Text classification improved by integrating bidirectional LSTM with two-dimensional max pooling. arXiv preprint arXiv:1611.06639. 2016."},{"key":"4520_CR23","unstructured":"Lin Z, Feng M, Santos CNd, Yu M, Xiang B, Zhou B, et al. A structured self-attentive sentence embedding. arXiv preprint arXiv:1703.03130. 2017."},{"key":"4520_CR24","doi-asserted-by":"crossref","unstructured":"Tan Z, Wang M, Xie J, Chen Y, Shi X, editors. Deep semantic role labeling with self-attention. In: Proceedings of the AAAI conference on artificial intelligence; 2018.","DOI":"10.1609\/aaai.v32i1.11928"},{"key":"4520_CR25","unstructured":"Raja S, Tuwani R. Adversarial attacks against deep learning systems for ICD-9 code assignment. 2020."},{"key":"4520_CR26","doi-asserted-by":"crossref","unstructured":"Wang W, Feng F, He X, Nie L, Chua T-S, editors. Denoising implicit feedback for recommendation. In: Proceedings of the 14th ACM international conference on web search and data mining; 2021.","DOI":"10.1145\/3437963.3441800"},{"key":"4520_CR27","unstructured":"Arazo E, Ortego D, Albert P, O'Connor N, McGuinness K, editors. Unsupervised label noise modeling and loss correction. In: International conference on machine learning; 2019: PMLR."},{"key":"4520_CR28","doi-asserted-by":"crossref","unstructured":"Han S, Lim C, Cha B, Lee J, editors. An empirical study for class imbalance in extreme multi-label text classification. In: 2021 IEEE international conference on big data and smart computing (BigComp); 2021: IEEE.","DOI":"10.1109\/BigComp51126.2021.00073"},{"key":"4520_CR29","unstructured":"Nichol A, Dhariwal P. Improved denoising diffusion probabilistic models. arXiv preprint arXiv:2102.09672. 2021."},{"key":"4520_CR30","doi-asserted-by":"crossref","unstructured":"Sch\u00fctze H, Manning CD, Raghavan P. Introduction to information retrieval: Cambridge University Press Cambridge; 2008.","DOI":"10.1017\/CBO9780511809071"},{"key":"4520_CR31","unstructured":"Kingma D, Ba J. Adam: a method for stochastic optimization. Computer Science. 2014."},{"issue":"2","key":"4520_CR32","doi-asserted-by":"publisher","first-page":"e0192360","DOI":"10.1371\/journal.pone.0192360","volume":"13","author":"S Gehrmann","year":"2018","unstructured":"Gehrmann S, Dernoncourt F, Li Y, Carlson ET, Celi LAG. Comparing deep learning and concept extraction based methods for patient phenotyping from clinical narratives. PLoS ONE. 2018;13(2):e0192360.","journal-title":"PLoS ONE"},{"key":"4520_CR33","doi-asserted-by":"crossref","unstructured":"Xie X, Xiong Y, Yu PS, Zhu Y, editors. Ehr coding with multi-scale feature attention and structured knowledge graph propagation. In: Proceedings of the 28th ACM international conference on information and knowledge management; 2019.","DOI":"10.1145\/3357384.3357897"},{"key":"4520_CR34","unstructured":"Berg Rvd, Kipf TN, Welling M. Graph convolutional matrix completion. arXiv preprint arXiv:1706.02263. 2017."},{"key":"4520_CR35","doi-asserted-by":"crossref","unstructured":"Cho K, Van Merri\u00ebnboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, et al. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078. 2014.","DOI":"10.3115\/v1\/D14-1179"},{"key":"4520_CR36","doi-asserted-by":"crossref","unstructured":"Croce D, Castellucci G, Basili R, editors. Gan-bert: generative adversarial learning for robust text classification with a bunch of labeled examples. In: Proceedings of the 58th annual meeting of the association for computational linguistics; 2020.","DOI":"10.18653\/v1\/2020.acl-main.191"},{"issue":"3","key":"4520_CR37","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3439726","volume":"54","author":"S Minaee","year":"2021","unstructured":"Minaee S, Kalchbrenner N, Cambria E, Nikzad N, Chenaghlu M, Gao J. Deep learning\u2013based text classification: a comprehensive review. ACM Comput Surv (CSUR). 2021;54(3):1\u201340.","journal-title":"ACM Comput Surv (CSUR)"},{"key":"4520_CR38","doi-asserted-by":"crossref","unstructured":"Xin J, Tang R, Yu Y, Lin J, editors. BERxiT: Early Exiting for BERT with Better fine-tuning and extension to regression. In: Proceedings of the 16th conference of the European chapter of the association for computational linguistics: Main Volume; 2021.","DOI":"10.18653\/v1\/2021.eacl-main.8"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04520-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-021-04520-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04520-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,18]],"date-time":"2023-01-18T03:29:56Z","timestamp":1674012596000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-021-04520-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12]]},"references-count":38,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["4520"],"URL":"https:\/\/doi.org\/10.1186\/s12859-021-04520-x","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12]]},"assertion":[{"value":"15 May 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 December 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 December 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"No ethics approval was required for the study.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not Applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"590"}}