{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T18:04:30Z","timestamp":1783015470820,"version":"3.54.6"},"reference-count":43,"publisher":"Oxford University Press (OUP)","issue":"6","license":[{"start":{"date-parts":[[2021,8,24]],"date-time":"2021-08-24T00:00:00Z","timestamp":1629763200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"DOI":"10.13039\/501100001747","name":"HKBU","doi-asserted-by":"publisher","award":["SDF19-0402-P02"],"award-info":[{"award-number":["SDF19-0402-P02"]}],"id":[{"id":"10.13039\/501100001747","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["2020QNA7003"],"award-info":[{"award-number":["2020QNA7003"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004731","name":"Zhejiang Provincial Natural Science Foundation","doi-asserted-by":"publisher","award":["LZ19H300001"],"award-info":[{"award-number":["LZ19H300001"]}],"id":[{"id":"10.13039\/501100004731","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Key R&D Program of Zhejiang Province","award":["2020C03010"],"award-info":[{"award-number":["2020C03010"]}]},{"DOI":"10.13039\/100005196","name":"Bureau of Justice Assistance","doi-asserted-by":"publisher","award":["kq2001034"],"award-info":[{"award-number":["kq2001034"]}],"id":[{"id":"10.13039\/100005196","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,11,5]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Computational methods have become indispensable tools to accelerate the drug discovery process and alleviate the excessive dependence on time-consuming and labor-intensive experiments. Traditional feature-engineering approaches heavily rely on expert knowledge to devise useful features, which could be costly and sometimes biased. The emerging deep learning (DL) methods deliver a data-driven method to automatically learn expressive representations from complex raw data. Inspired by this, researchers have attempted to apply various deep neural network models to simplified molecular input line entry specification (SMILES) strings, which contain all the composition and structure information of molecules. However, current models usually suffer from the scarcity of labeled data. This results in a low generalization ability of SMILES-based DL models, which prevents them from competing with the state-of-the-art computational methods. In this study, we utilized the BiLSTM (bidirectional long short term merory) attention network (BAN) in which we employed a novel multi-step attention mechanism to facilitate the extracting of key features from the SMILES strings. Meanwhile, SMILES enumeration was utilized as a data augmentation method in the training phase to substantially increase the number of labeled data and enlarge the probability of mining more patterns from complex SMILES. We again took advantage of SMILES enumeration in the prediction phase to rectify model prediction bias and provide a more accurate prediction. Combined with the BAN model, our strategies can greatly improve the performance of latent features learned from SMILES strings. In 11 canonical absorption, distribution, metabolism, excretion and toxicity-related tasks, our method outperformed the state-of-the-art approaches.<\/jats:p>","DOI":"10.1093\/bib\/bbab327","type":"journal-article","created":{"date-parts":[[2021,7,27]],"date-time":"2021-07-27T11:08:46Z","timestamp":1627384126000},"source":"Crossref","is-referenced-by-count":49,"title":["Learning to SMILES: BAN-based strategies to improve latent representation learning from molecules"],"prefix":"10.1093","volume":"22","author":[{"given":"Cheng-Kun","family":"Wu","sequence":"first","affiliation":[{"name":"State Key Laboratory of High-Performance Computing, College of Computer, National University of Defense Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiao-Chen","family":"Zhang","sequence":"additional","affiliation":[{"name":"The College of Computer, National University of Defense Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhi-Jiang","family":"Yang","sequence":"additional","affiliation":[{"name":"Xiangya School of Pharmaceutical Sciences, Central South University, Hunan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ai-Ping","family":"Lu","sequence":"additional","affiliation":[{"name":"Institute for Advancing Translational Medicine in Bone and Joint Diseases, School of Chinese Medicine, Hong Kong Baptist University, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7227-2580","authenticated-orcid":false,"given":"Ting-Jun","family":"Hou","sequence":"additional","affiliation":[{"name":"College of Pharmaceutical Sciences, Zhejiang University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3604-3785","authenticated-orcid":false,"given":"Dong-Sheng","family":"Cao","sequence":"additional","affiliation":[{"name":"Xiangya School of Pharmaceutical Sciences, Central South University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2021,8,24]]},"reference":[{"key":"2021110815092936600_ref1","doi-asserted-by":"crossref","first-page":"1275","DOI":"10.3389\/fphar.2018.01275","article-title":"QSAR-based virtual screening: advances and applications in drug discovery","volume":"9","author":"Neves","year":"2018","journal-title":"Front Pharmacol"},{"key":"2021110815092936600_ref2","doi-asserted-by":"crossref","first-page":"1106","DOI":"10.2174\/1389200218666170607102104","article-title":"Recent advances of computational modeling for predicting drug metabolism: a perspective","volume":"18","author":"Kar","year":"2017","journal-title":"Curr Drug Metab"},{"key":"2021110815092936600_ref3","doi-asserted-by":"crossref","first-page":"1604","DOI":"10.1093\/bib\/bbz176","article-title":"Biomedical data and computational models for drug repositioning: a comprehensive review","volume":"22","author":"Luo","year":"2021","journal-title":"Brief Bioinform"},{"key":"2021110815092936600_ref4","first-page":"74","article-title":"A practical overview of quantitative structure-activity relationship","volume":"8","author":"Nantasenamat","year":"2009","journal-title":"EXCLI J"},{"key":"2021110815092936600_ref5","doi-asserted-by":"crossref","first-page":"595","DOI":"10.1007\/s10822-016-9938-8","article-title":"Molecular graph convolutions: moving beyond fingerprints","volume":"30","author":"Kearnes","year":"2016","journal-title":"J Comput Aided Mol Des"},{"key":"2021110815092936600_ref6","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1016\/j.eswa.2016.12.008","article-title":"Automatic selection of molecular descriptors using random forest: application to drug discovery","volume":"72","author":"Cano","year":"2017","journal-title":"Expert Syst Appl"},{"key":"2021110815092936600_ref7","doi-asserted-by":"crossref","first-page":"2641","DOI":"10.4155\/fmc-2018-0076","article-title":"A review of ligand-based virtual screening web tools and screening algorithms in large molecular databases in the age of big data","volume":"10","author":"Banegas-Luna","year":"2018","journal-title":"Future Med Chem"},{"key":"2021110815092936600_ref8","doi-asserted-by":"crossref","first-page":"487","DOI":"10.1186\/s12859-016-1353-6","article-title":"LBSizeCleav: improved support vector machine (SVM)-based prediction of Dicer cleavage sites using loop\/bulge length","volume":"17","author":"Bao","year":"2016","journal-title":"BMC Bioinformatics"},{"key":"2021110815092936600_ref9","volume-title":"Advances in Kernel Methods-Support Vector Learning","author":""},{"key":"2021110815092936600_ref10","doi-asserted-by":"crossref","first-page":"2449","DOI":"10.1093\/bioinformatics\/bty087","article-title":"A new approach for interpreting random forest models and its application to the biology of ageing","volume":"34","author":"Fabris","year":"2018","journal-title":"Bioinformatics"},{"key":"2021110815092936600_ref11","doi-asserted-by":"crossref","first-page":"197","DOI":"10.1007\/s11749-016-0481-7","article-title":"A random forest guided tour","volume":"25","author":"Biau","year":"2016","journal-title":"TEST"},{"key":"2021110815092936600_ref12","first-page":"785","volume-title":"Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, United States, 2016","author":"Chen"},{"key":"2021110815092936600_ref13","article-title":"Bioactive molecule prediction using extreme gradient boosting","volume-title":"Molecules","author":""},{"key":"2021110815092936600_ref14","doi-asserted-by":"crossref","first-page":"4977","DOI":"10.1021\/jm4004285","article-title":"QSAR modeling: Where have you been? Where are you going to?","volume":"57","author":"Cherkasov","year":"2014","journal-title":"J Med Chem"},{"key":"2021110815092936600_ref15","doi-asserted-by":"crossref","first-page":"bbab152","DOI":"10.1093\/bib\/bbab152","article-title":"MG-BERT: leveraging unsupervised atomic representation learning for molecular property prediction","volume":"5","author":"Zhang","year":"2021","journal-title":"Brief Bioinform"},{"key":"2021110815092936600_ref16","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1155\/2021\/6611777","article-title":"Could graph neural networks learn better molecular representation for drug discovery? A comparison study of descriptor-based and graph-based models","volume":"13","author":"Jiang","year":"2021","journal-title":"J Chem"},{"key":"2021110815092936600_ref17","doi-asserted-by":"crossref","first-page":"8749","DOI":"10.1021\/acs.jmedchem.9b00959","article-title":"Pushing the boundaries of molecular representation for drug discovery with the graph attention mechanism","volume":"63","author":"Xiong","year":"2019","journal-title":"J Med Chem"},{"key":"2021110815092936600_ref18","doi-asserted-by":"crossref","first-page":"129","DOI":"10.2174\/157340991002140708105124","article-title":"QSAR multi-target in drug discovery: a review","volume":"10","author":"Zanni","year":"2014","journal-title":"Curr Comput Aided Drug Des"},{"key":"2021110815092936600_ref19","volume-title":"the 26th Annual Conference on Neural Information Processing Systems, Lake Tahoe, Nevada, USA, 2012","author":"Krizhevsky"},{"key":"2021110815092936600_ref20","first-page":"770","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 2016","author":"He"},{"key":"2021110815092936600_ref21","article-title":"Google\u2019s neural machine translation system: bridging the gap between human and machine translation","author":"Wu","year":"2016"},{"key":"2021110815092936600_ref22","article-title":"Bert: pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA","author":"Devlin"},{"key":"2021110815092936600_ref23","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"2021110815092936600_ref24","article-title":"Learning to smile(s)","author":"Jastrz\u0119bski","year":"2016"},{"key":"2021110815092936600_ref25","first-page":"1263","volume-title":"International Conference on Machine Learning. Sydney, NSW, Australia, 2017","author":"Gilmer"},{"key":"2021110815092936600_ref26","doi-asserted-by":"publisher","DOI":"10.1186\/s13321-020-00423-w","article-title":"Transformer-CNN: Swiss knife for QSAR modeling and interpretation","volume-title":"J Cheminform","author":"Karpov"},{"key":"2021110815092936600_ref27","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1021\/ci00057a005","article-title":"SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules","volume":"28","author":"Weininger","year":"1988","journal-title":"J Chem Inform Comput Sci"},{"key":"2021110815092936600_ref28","doi-asserted-by":"crossref","first-page":"97","DOI":"10.1021\/ci00062a008","article-title":"SMILES. 2. Algorithm for generation of unique SMILES notation","volume":"29","author":"Weininger","year":"1989","journal-title":"J Chem Inform Comput Sci"},{"key":"2021110815092936600_ref29","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput"},{"key":"2021110815092936600_ref30","article-title":"SMILES enumeration as data augmentation for neural network modeling of molecules","author":"Bjerrum","year":"2017"},{"key":"2021110815092936600_ref31","first-page":"2508","article-title":"Survey of convolutional neural network","volume":"36","author":"Li","year":"2016","journal-title":"J Comput Appl"},{"key":"2021110815092936600_ref32","first-page":"1","article-title":"Randomized SMILES strings improve the quality of molecular generative models","volume":"11","author":"Ar\u00fas-Pous","year":"2019","journal-title":"J Chem"},{"key":"2021110815092936600_ref33","article-title":"Attention is all you need","author":"Vaswani","journal-title":"Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems, Long Beach, CA, USA, 2017"},{"key":"2021110815092936600_ref34","first-page":"1","article-title":"Transformer-CNN: Swiss knife for QSAR modeling and interpretation","volume":"12","author":"Karpov","year":"2020","journal-title":"J Chem"},{"key":"2021110815092936600_ref35","doi-asserted-by":"crossref","first-page":"5441","DOI":"10.1039\/C8SC00148K","article-title":"Large-scale comparison of machine learning methods for drug target prediction on ChEMBL","volume":"9","author":"Mayr","year":"2018","journal-title":"Chem Sci"},{"key":"2021110815092936600_ref36","first-page":"1480","volume-title":"Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. San Diego California, USA, 2016","author":"Yang"},{"key":"2021110815092936600_ref37","article-title":"Graph attention networks","volume-title":"International Conference on Learning Representations, Vancouver, BC, Canada, 2018","author":"Veli\u010dkovi\u0107"},{"key":"2021110815092936600_ref38","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1186\/s13321-018-0283-x","article-title":"ADMETlab: a platform for systematic ADMET evaluation based on a comprehensively collected ADMET database","volume":"10","author":"Dong","year":"2018","journal-title":"J Chem"},{"key":"2021110815092936600_ref39","doi-asserted-by":"crossref","DOI":"10.1093\/nar\/gkab255","article-title":"ADMETlab 2.0: an integrated online platform for accurate and comprehensive predictions of ADMET properties","volume":"49","author":"Xiong","year":"2021","journal-title":"Nucleic Acids Res"},{"key":"2021110815092936600_ref40","article-title":"Order matters: sequence to sequence for sets","author":"Vinyals","year":"2015"},{"key":"2021110815092936600_ref41","first-page":"1929","article-title":"Dropout: a simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J Mach Learn Res"},{"key":"2021110815092936600_ref42","article-title":"Layer normalization","author":"Ba","year":"2015","journal-title":"arXiv Preprint arXiv:1506:01057"},{"key":"2021110815092936600_ref43","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1021\/acscentsci.6b00367","article-title":"Low data drug discovery with one-shot learning","volume":"3","author":"Altae-Tran","year":"2017","journal-title":"ACS Cent Sci"}],"container-title":["Briefings in Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bib\/article-pdf\/22\/6\/bbab327\/41090371\/bbab327.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bib\/article-pdf\/22\/6\/bbab327\/41090371\/bbab327.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,11,8]],"date-time":"2021-11-08T15:31:49Z","timestamp":1636385509000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bib\/article\/doi\/10.1093\/bib\/bbab327\/6356874"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,24]]},"references-count":43,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2021,11,5]]}},"URL":"https:\/\/doi.org\/10.1093\/bib\/bbab327","relation":{},"ISSN":["1467-5463","1477-4054"],"issn-type":[{"value":"1467-5463","type":"print"},{"value":"1477-4054","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2021,11]]},"published":{"date-parts":[[2021,8,24]]},"article-number":"bbab327"}}