{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T00:38:07Z","timestamp":1784075887440,"version":"3.55.0"},"reference-count":45,"publisher":"Oxford University Press (OUP)","issue":"11","license":[{"start":{"date-parts":[[2023,11,1]],"date-time":"2023-11-01T00:00:00Z","timestamp":1698796800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["32370657"],"award-info":[{"award-number":["32370657"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62072435"],"award-info":[{"award-number":["62072435"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["82130055"],"award-info":[{"award-number":["82130055"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["32271297"],"award-info":[{"award-number":["32271297"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,11,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>N-linked glycosylation is a frequently occurring post-translational protein modification that serves critical functions in protein folding, stability, trafficking, and recognition. Its involvement spans across multiple biological processes and alterations to this process can result in various diseases. Therefore, identifying N-linked glycosylation sites is imperative for comprehending the mechanisms and systems underlying glycosylation. Due to the inherent experimental complexities, machine learning and deep learning have become indispensable tools for predicting these sites.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>In this context, a new approach called EMNGly has been proposed. The EMNGly approach utilizes pretrained protein language model (Evolutionary Scale Modeling) and pretrained protein structure model (Inverse Folding Model) for features extraction and support vector machine for classification. Ten-fold cross-validation and independent tests show that this approach has outperformed existing techniques. And it achieves Matthews Correlation Coefficient, sensitivity, specificity, and accuracy of 0.8282, 0.9343, 0.8934, and 0.9143, respectively on a benchmark independent test set.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btad650","type":"journal-article","created":{"date-parts":[[2023,11,6]],"date-time":"2023-11-06T17:26:31Z","timestamp":1699291591000},"source":"Crossref","is-referenced-by-count":36,"title":["EMNGly: predicting N-linked glycosylation sites using the language models for feature extraction"],"prefix":"10.1093","volume":"39","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-8004-2001","authenticated-orcid":false,"given":"Xiaoyang","family":"Hou","sequence":"first","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology , Beijing 100190, China"},{"name":"University of Chinese Academy of Sciences , Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Wang","sequence":"additional","affiliation":[{"name":"Syneron Technology , Guangzhou 510000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4119-4238","authenticated-orcid":false,"given":"Dongbo","family":"Bu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology , Beijing 100190, China"},{"name":"University of Chinese Academy of Sciences , Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yaojun","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Information and Electrical Engineering, China Agricultural University , Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0351-3720","authenticated-orcid":false,"given":"Shiwei","family":"Sun","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Information Processing, Institute of Computing Technology , Beijing 100190, China"},{"name":"University of Chinese Academy of Sciences , Beijing 100049, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2023,11,1]]},"reference":[{"key":"2023110721575948500_btad650-B1","doi-asserted-by":"crossref","first-page":"74","DOI":"10.1016\/j.tibs.2009.10.001","article-title":"N-glycan structures: recognition and processing in the ER","volume":"35","author":"Aebi","year":"2010","journal-title":"Trends Biochem Sci"},{"key":"2023110721575948500_btad650-B2","doi-asserted-by":"crossref","first-page":"3096","DOI":"10.1021\/ja01039a051","article-title":"Feline gastrin. An example of peptide sequence analysis by mass spectrometry","volume":"91","author":"Agarwal","year":"1969","journal-title":"J Am Chem Soc"},{"key":"2023110721575948500_btad650-B3","doi-asserted-by":"crossref","first-page":"774","DOI":"10.2174\/1574893615666210108094847","article-title":"Intelligent techniques analysis for glycosylation site prediction","volume":"16","author":"Alkuhlani","year":"2021","journal-title":"CBIO"},{"key":"2023110721575948500_btad650-B4","doi-asserted-by":"crossref","first-page":"12702","DOI":"10.1109\/ACCESS.2022.3146395","article-title":"Pustackngly: positive-unlabeled and stacking learning for n-linked glycosylation site prediction","volume":"10","author":"Alkuhlani","year":"2022","journal-title":"IEEE Access"},{"key":"2023110721575948500_btad650-B5","doi-asserted-by":"crossref","first-page":"95","DOI":"10.1093\/glycob\/cwh004","article-title":"Biases and complex patterns in the residues flanking protein n-glycosylation sites","volume":"14","author":"Ben-Dor","year":"2004","journal-title":"Glycobiology"},{"key":"2023110721575948500_btad650-B6","doi-asserted-by":"crossref","first-page":"e40155","DOI":"10.1371\/journal.pone.0040155","article-title":"GlycoPP: a webserver for prediction of n-and o-glycosites in prokaryotic protein sequences","volume":"7","author":"Chauhan","year":"2012","journal-title":"PLoS One"},{"key":"2023110721575948500_btad650-B7","doi-asserted-by":"crossref","first-page":"e67008","DOI":"10.1371\/journal.pone.0067008","article-title":"In silico platform for prediction of N-, O- and C-glycosites in eukaryotic protein sequences","volume":"8","author":"Chauhan","year":"2013","journal-title":"PLoS One"},{"key":"2023110721575948500_btad650-B9","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1007\/BF00994018","article-title":"Support-vector networks","volume":"20","author":"Cortes","year":"1995","journal-title":"Mach Learn"},{"key":"2023110721575948500_btad650-B10","doi-asserted-by":"crossref","first-page":"433","DOI":"10.1093\/protein\/3.5.433","article-title":"Sequence differences between glycosylated and non-glycosylated Asn-X-Thr\/Ser acceptor sites: implications for protein engineering","volume":"3","author":"Gavel","year":"1990","journal-title":"Protein Eng"},{"key":"2023110721575948500_btad650-B4288288","doi-asserted-by":"crossref","first-page":"310","DOI":"10.1142\/9789812799623_0029","article-title":"Prediction of glycosylation across the human proteome and the correlation to protein function","author":"Gupta","year":"2001","journal-title":"Biocomputing 2002, Worldscientific"},{"key":"2023110721575948500_btad650-B12","doi-asserted-by":"crossref","first-page":"500","DOI":"10.1186\/1471-2105-9-500","article-title":"Prediction of glycosylation sites using random forests","volume":"9","author":"Hamby","year":"2008","journal-title":"BMC Bioinformatics"},{"key":"2023110721575948500_btad650-B13","doi-asserted-by":"crossref","first-page":"672","DOI":"10.1016\/j.cell.2010.11.008","article-title":"Glycomics hits the big time","volume":"143","author":"Hart","year":"2010","journal-title":"Cell"},{"key":"2023110721575948500_btad650-B14","doi-asserted-by":"crossref","first-page":"115528","DOI":"10.1016\/j.carbpol.2019.115528","article-title":"Identification of carbohydrate peripheral epitopes important for recognition by positive-ion MALDI multistage mass spectrometry","volume":"229","author":"Huang","year":"2020","journal-title":"Carbohydr Polym"},{"key":"2023110721575948500_btad650-B15","doi-asserted-by":"crossref","first-page":"116122","DOI":"10.1016\/j.carbpol.2020.116122","article-title":"Multistage mass spectrometry with intelligent precursor selection for N-glycan branching pattern analysis","volume":"237","author":"Huang","year":"2020","journal-title":"Carbohydr Polym"},{"key":"2023110721575948500_btad650-B16","first-page":"811","article-title":"Gips-mix for accurate identification of isomeric components in glycan mixtures using intelligent group-opting strategy","volume":"95","author":"Huang","year":"2022","journal-title":"Anal Chem"},{"key":"2023110721575948500_btad650-B17","doi-asserted-by":"crossref","first-page":"217","DOI":"10.1016\/j.compbiolchem.2019.03.015","article-title":"De novo glycan structural identification from mass spectra using tree merging strategy","volume":"80","author":"Ju","year":"2019","journal-title":"Comput Biol Chem"},{"key":"2023110721575948500_btad650-B18","doi-asserted-by":"crossref","first-page":"1957","DOI":"10.1038\/sj.emboj.7601087","article-title":"Definition of the bacterial n-glycosylation site consensus sequence","volume":"25","author":"Kowarik","year":"2006","journal-title":"EMBO J"},{"key":"2023110721575948500_btad650-B19","doi-asserted-by":"crossref","first-page":"515","DOI":"10.1016\/j.sbi.2009.06.004","article-title":"Glycoprotein folding, quality control and ER-associated degradation","volume":"19","author":"Lederkremer","year":"2009","journal-title":"Curr Opin Struct Biol"},{"key":"2023110721575948500_btad650-B20","doi-asserted-by":"crossref","first-page":"1411","DOI":"10.1093\/bioinformatics\/btu852","article-title":"Glycomine: a machine learning-based approach for predicting N-, C- and O-linked glycosylation in the human proteome","volume":"31","author":"Li","year":"2015","journal-title":"Bioinformatics"},{"key":"2023110721575948500_btad650-B21","doi-asserted-by":"crossref","first-page":"34595","DOI":"10.1038\/srep34595","article-title":"GlycoMinestruct: a new bioinformatics tool for highly accurate mapping of the human N-linked and O-linked glycoproteomes by incorporating structural features","volume":"6","author":"Li","year":"2016","journal-title":"Sci Rep"},{"key":"2023110721575948500_btad650-B22","doi-asserted-by":"crossref","first-page":"1658","DOI":"10.1093\/bioinformatics\/btl158","article-title":"Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences","volume":"22","author":"Li","year":"2006","journal-title":"Bioinformatics"},{"key":"2023110721575948500_btad650-B23","first-page":"199","article-title":"Lectins as a tool for glycan profiling","volume":"1273","author":"Liu","year":"2015","journal-title":"Methods Mol Biol"},{"key":"2023110721575948500_btad650-B24","doi-asserted-by":"crossref","first-page":"713","DOI":"10.1038\/nchembio.437","article-title":"A systematic approach to protein glycosylation analysis: a path through the maze","volume":"6","author":"Marino","year":"2010","journal-title":"Nat Chem Biol"},{"key":"2023110721575948500_btad650-B26","doi-asserted-by":"crossref","first-page":"e1002285","DOI":"10.1371\/journal.pcbi.1002285","article-title":"Integrating bioinformatics tools to handle glycosylation","volume":"7","author":"Mazola","year":"2011","journal-title":"PLoS Comput Biol"},{"key":"2023110721575948500_btad650-B27","doi-asserted-by":"crossref","first-page":"448","DOI":"10.1038\/nrm3383","article-title":"Vertebrate protein glycosylation: diversity, synthesis and function","volume":"13","author":"Moremen","year":"2012","journal-title":"Nat Rev Mol Cell Biol"},{"key":"2023110721575948500_btad650-B28","doi-asserted-by":"crossref","first-page":"855","DOI":"10.1016\/j.cell.2006.08.019","article-title":"Glycosylation in cellular mechanisms of health and disease","volume":"126","author":"Ohtsubo","year":"2006","journal-title":"Cell"},{"key":"2023110721575948500_btad650-B29","doi-asserted-by":"crossref","first-page":"7314","DOI":"10.3390\/molecules26237314","article-title":"DeepNGlypred: a deep neural network-based approach for human N-linked glycosylation site prediction","volume":"26","author":"Pakhrin","year":"2021","journal-title":"Molecules"},{"key":"2023110721575948500_btad650-B30","doi-asserted-by":"crossref","first-page":"411","DOI":"10.1093\/glycob\/cwad033","article-title":"LMNglyPred: prediction of human n-linked glycosylation sites using embeddings from a pre-trained protein language model","volume":"33","author":"Pakhrin","year":"2023","journal-title":"Glycobiology"},{"key":"2023110721575948500_btad650-B31","doi-asserted-by":"crossref","first-page":"16933","DOI":"10.1038\/s41598-022-21366-2","article-title":"Improving protein succinylation sites prediction using embeddings from protein language model","volume":"12","author":"Pokharel","year":"2022","journal-title":"Sci Rep"},{"key":"2023110721575948500_btad650-B32","doi-asserted-by":"crossref","first-page":"e2016239118","DOI":"10.1073\/pnas.2016239118","article-title":"Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences","volume":"118","author":"Rives","year":"2021","journal-title":"Proc Natl Acad Sci USA"},{"key":"2023110721575948500_btad650-B33","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1186\/s12014-019-9254-0","article-title":"N-glycositeatlas: a database resource for mass spectrometry-based human N-linked glycoprotein and glycosylation site mapping","volume":"16","author":"Sun","year":"2019","journal-title":"Clin Proteomics"},{"key":"2023110721575948500_btad650-B34","doi-asserted-by":"crossref","first-page":"14412","DOI":"10.1021\/acs.analchem.8b03967","article-title":"Toward automated identification of glycan branching patterns using multistage mass spectrometry with intelligent precursor selection","volume":"90","author":"Sun","year":"2018","journal-title":"Anal Chem"},{"key":"2023110721575948500_btad650-B35","doi-asserted-by":"crossref","first-page":"4140","DOI":"10.1093\/bioinformatics\/btz215","article-title":"SPRINT-Gly: predicting N- and O-linked glycosylation sites of human and mouse proteins by using sequence and predicted structural properties","volume":"35","author":"Taherzadeh","year":"2019","journal-title":"Bioinformatics"},{"key":"2023110721575948500_btad650-B36","doi-asserted-by":"crossref","first-page":"941","DOI":"10.1093\/bioinformatics\/btab801","article-title":"NetSolP: predicting protein solubility in escherichia coli using language models","volume":"38","author":"Thumuluri","year":"2022","journal-title":"Bioinformatics"},{"key":"2023110721575948500_btad650-B0059246","first-page":"785","article-title":"Xgboost: A scalable tree boosting system","author":"Tianqi","year":"2016","journal-title":"Proceedings of the 22nd Acm Sigkdd International Conference on Knowledge Discovery and Data Mining"},{"key":"2023110721575948500_btad650-B37","doi-asserted-by":"crossref","first-page":"533","DOI":"10.1001\/jama.2016.7653","article-title":"Logistic regression: relating patient characteristics to outcomes","volume":"316","author":"Tolles","year":"2016","journal-title":"JAMA"},{"key":"2023110721575948500_btad650-B38","doi-asserted-by":"crossref","first-page":"D204","DOI":"10.1093\/nar\/gku989","article-title":"UniProt: a hub for protein information","volume":"43","author":"UniProt Consortium","year":"2015","journal-title":"Nucleic Acids Res"},{"key":"2023110721575948500_btad650-B40","volume-title":"Essentials of Glycobiology","author":"Varki","year":"2009","edition":"2nd ed"},{"key":"2023110721575948500_btad650-B41","doi-asserted-by":"crossref","first-page":"103649","DOI":"10.1016\/j.jprot.2020.103649","article-title":"Identification of glycan branching patterns using multistage mass spectrometry with spectra tree analysis","volume":"217","author":"Wang","year":"2020","journal-title":"J Proteomics"},{"key":"2023110721575948500_btad650-B42","doi-asserted-by":"crossref","first-page":"723149","DOI":"10.3389\/fchem.2021.723149","article-title":"HepParser: an intelligent software program for deciphering low-molecular-weight heparin based on mass spectrometry","volume":"9","author":"Wang","year":"2021","journal-title":"Front Chem"},{"key":"2023110721575948500_btad650-B43","doi-asserted-by":"crossref","first-page":"2991","DOI":"10.1093\/bioinformatics\/btz056","article-title":"Best-first search guided multistage mass spectrometry-based glycan identification","volume":"35","author":"Wang","year":"2019","journal-title":"Bioinformatics"},{"key":"2023110721575948500_btad650-B44","doi-asserted-by":"crossref","first-page":"91R","DOI":"10.1093\/glycob\/cwj099","article-title":"Asparagine-linked protein glycosylation: from eukaryotic to prokaryotic systems","volume":"16","author":"Weerapana","year":"2006","journal-title":"Glycobiology"},{"key":"2023110721575948500_btad650-B46","first-page":"2022","author":"Yang","year":"2022"},{"key":"2023110721575948500_btad650-B47","doi-asserted-by":"crossref","first-page":"881","DOI":"10.1016\/j.chembiol.2008.07.016","article-title":"Mass spectrometry and the emerging field of glycomics","volume":"15","author":"Zaia","year":"2008","journal-title":"Chem Biol"},{"key":"2023110721575948500_btad650-B48","doi-asserted-by":"crossref","first-page":"3183","DOI":"10.1016\/j.jmb.2016.02.030","article-title":"Glycosylation quality control by the Golgi structure","volume":"428","author":"Zhang","year":"2016","journal-title":"J Mol Biol"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/11\/btad650\/52799569\/btad650.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/39\/11\/btad650\/52799569\/btad650.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,7]],"date-time":"2023-11-07T22:39:38Z","timestamp":1699396778000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/doi\/10.1093\/bioinformatics\/btad650\/7335841"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,1]]},"references-count":45,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2023,11,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btad650","relation":{},"ISSN":["1367-4811"],"issn-type":[{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,11,1]]},"published":{"date-parts":[[2023,11,1]]},"article-number":"btad650"}}