{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T00:26:45Z","timestamp":1777854405401,"version":"3.51.4"},"reference-count":34,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2006,2,1]],"date-time":"2006-02-01T00:00:00Z","timestamp":1138752000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Journal of Information Science"],"published-print":{"date-parts":[[2006,2]]},"abstract":"<jats:p>In molecular biology, DNA sequence matching is one of the most crucial operations. Since DNA databases contain a huge volume of sequences, fast indexes are essential for efficient processing of DNA sequence matching. In this paper, we first point out the problems of the suffix tree, an index structure widely-used for DNA sequence matching, in respect of storage overhead, search performance, and difficulty in seamless integration with DBMS. Then, we propose a new index structure that resolves such problems. The proposed index structure consists of two parts: the primary part realizes the trie as binary bit-string representation without any pointers, and the secondary part helps fast access to the trie's leaf nodes that need to be accessed for post-processing. We also suggest efficient algorithms based on that index for DNA sequence matching. To verify the superiority of the proposed approach, we conduct performance evaluation via a series of experiments. The results reveal that the proposed approach, which requires smaller storage space, can be a few orders of magnitude faster than the suffix tree.<\/jats:p>","DOI":"10.1177\/0165551506059229","type":"journal-article","created":{"date-parts":[[2006,2,13]],"date-time":"2006-02-13T04:52:34Z","timestamp":1139806354000},"page":"88-104","source":"Crossref","is-referenced-by-count":3,"title":["An efficient approach for sequence matching in large DNA databases"],"prefix":"10.1177","volume":"32","author":[{"given":"Jung-Im","family":"Won","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sanghyun","family":"Park","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Yonsei University, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jee-Hee","family":"Yoon","sequence":"additional","affiliation":[{"name":"Division of Information Engineering and Telecommunications, Hallym                         University, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sang-Wook","family":"Kim","sequence":"additional","affiliation":[{"name":"College of Information and Communications, Hanyang University, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2006,2,1]]},"reference":[{"key":"atypb1","volume-title":"Developing Bioinformatics Computer Skills","author":"C. Gibas","year":"2001"},{"key":"atypb2","volume-title":"Bioinformatics: Sequence and Genome Analysis","author":"D.W. Mount","year":"2001"},{"key":"atypb3","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/17.2.180"},{"key":"atypb4","doi-asserted-by":"publisher","DOI":"10.1093\/nar\/26.1.1"},{"key":"atypb5","doi-asserted-by":"publisher","DOI":"10.1016\/S0022-2836(05)80360-2"},{"key":"atypb6","doi-asserted-by":"publisher","DOI":"10.1093\/nar\/25.17.3389"},{"key":"atypb7","doi-asserted-by":"publisher","DOI":"10.1016\/0022-2836(81)90087-5"},{"key":"atypb8","doi-asserted-by":"publisher","DOI":"10.1142\/2418"},{"key":"atypb9","doi-asserted-by":"publisher","DOI":"10.1007\/BFb0029808"},{"key":"atypb10","doi-asserted-by":"publisher","DOI":"10.1093\/nar\/27.11.2369"},{"key":"atypb11","doi-asserted-by":"publisher","DOI":"10.1093\/nar\/29.22.4633"},{"key":"atypb12","doi-asserted-by":"publisher","DOI":"10.1007\/s007780200064"},{"issue":"1","key":"atypb13","first-page":"205","volume":"1","author":"G. Navarro","year":"2000","journal-title":"Journal of Discrete Algorithms"},{"key":"atypb14","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-48452-3_13"},{"key":"atypb15","doi-asserted-by":"publisher","DOI":"10.1016\/B978-012722442-8\/50085-9"},{"key":"atypb16","volume-title":"The A* Search and Applications to Sequence Alignment","author":"K. Kelly","year":"1996"},{"key":"atypb17","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/15.5.426"},{"key":"atypb18","doi-asserted-by":"publisher","DOI":"10.1002\/spe.535"},{"key":"atypb19","first-page":"175","volume-title":"Proceedings of the 12th Genome Informatics (GIW01) 2001","author":"K. Sadakane","year":"2001"},{"key":"atypb20","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btg310"},{"key":"atypb21","volume-title":"Fundamentals of Data Structures in C","author":"E. Horowitz","year":"1993"},{"key":"atypb22","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/17.5.419"},{"key":"atypb23","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/18.3.440"},{"key":"atypb24","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.1993.341106"},{"issue":"3","key":"atypb25","first-page":"273","volume":"11","author":"C. Fondrat","year":"1995","journal-title":"Computer Applications in the Biosciences"},{"issue":"1","key":"atypb26","first-page":"63","volume":"14","author":"H.E. Williams","year":"2002","journal-title":"IEEE TKDE"},{"key":"atypb27","first-page":"351","volume-title":"Proceedings of the 27th International Conference on Very Large Databases (VLDB01) 2001","author":"T. Kaheci","year":"2001"},{"key":"atypb28","doi-asserted-by":"publisher","DOI":"10.1137\/0222058"},{"key":"atypb29","doi-asserted-by":"publisher","DOI":"10.1145\/1457838.1457895"},{"key":"atypb30","volume-title":"The Art of Computer Programming 3: Sorting and searching","author":"D.E. Knuth","year":"1973"},{"key":"atypb31","doi-asserted-by":"publisher","DOI":"10.1109\/69.536247"},{"key":"atypb32","doi-asserted-by":"publisher","DOI":"10.1109\/HICSS.1994.323593"},{"key":"atypb33","volume-title":"Genbank"},{"key":"atypb34","doi-asserted-by":"publisher","DOI":"10.1145\/93597.98741"}],"container-title":["Journal of Information Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0165551506059229","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/0165551506059229","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T23:07:05Z","timestamp":1777504025000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/0165551506059229"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,2]]},"references-count":34,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2006,2]]}},"alternative-id":["10.1177\/0165551506059229"],"URL":"https:\/\/doi.org\/10.1177\/0165551506059229","relation":{},"ISSN":["0165-5515","1741-6485"],"issn-type":[{"value":"0165-5515","type":"print"},{"value":"1741-6485","type":"electronic"}],"subject":[],"published":{"date-parts":[[2006,2]]}}}