{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:29:27Z","timestamp":1750307367968,"version":"3.41.0"},"reference-count":13,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2010,7,20]],"date-time":"2010-07-20T00:00:00Z","timestamp":1279584000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGSOFT Softw. Eng. Notes"],"published-print":{"date-parts":[[2010,7,20]]},"abstract":"<jats:p>CRAYSE1 is a SEarch WHIle CRAwl application, intended to perform fast searching of text in web pages. A Web crawler is a computer program that browses the World Wide Web in a methodical, automated manner. This process is also called spidering. Search engines, use spidering as a means of providing up-to-date data. Most of the existing web-crawlers archive the contents of the web starting from the input URL. Search engines index the results of web-crawlers and then perform searching when queried. As such, the searching is not performed while crawling. Hence such softwares can not be used for general use by web browsers. Also, the existing search mechanism in web browsers, search only on the current page and not recursively through all the links present in that page. In order to overcome such disadvantages, we propose in this paper to implement a web crawler that searches for a pattern efficiently and recursively through all the links including pdf links while crawling. CRAYSE can be used as a general purpose open source software by web browsers. It can also be used for offine searching. Further, the applications that require selective archival of web pages (based on the presence of a key word), can deploy CRAYSE for efficient search operations. This paper focusses on the design and implementation of CRAYSE and its demonstration through web applications.<\/jats:p>","DOI":"10.1145\/1811226.1811236","type":"journal-article","created":{"date-parts":[[2010,7,22]],"date-time":"2010-07-22T18:52:11Z","timestamp":1279824731000},"page":"1-8","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["CRAYSE"],"prefix":"10.1145","volume":"35","author":[{"given":"V.","family":"Radhakishan","sequence":"first","affiliation":[{"name":"National Institute of Technology Tiruchirappalli, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yaser","family":"Farook","sequence":"additional","affiliation":[{"name":"National Institute of Technology Tiruchirappalli, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"S.","family":"Selvakumar","sequence":"additional","affiliation":[{"name":"National Institute of Technology Tiruchirappalli, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2010,7,20]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Java online documentation http:\/\/java.sun.com\/j2se\/1.4.2\/docs\/api\/  Java online documentation http:\/\/java.sun.com\/j2se\/1.4.2\/docs\/api\/"},{"key":"e_1_2_1_2_1","unstructured":"Wikipedia http:\/\/en.wikipedia.org\/wiki\/web crawler.  Wikipedia http:\/\/en.wikipedia.org\/wiki\/web crawler."},{"key":"e_1_2_1_3_1","first-page":"10","volume-title":"Baeza-Yates. Scheduling algorithms for web crawling","author":"Castillo Carlos","year":"2004","unstructured":"Carlos Castillo , Mauricio Mar\u00edn , Andrea Rodr\u00edguez , and Ricardo A . Baeza-Yates. Scheduling algorithms for web crawling . pages 10 -- 17 , 2004 . Carlos Castillo, Mauricio Mar\u00edn, Andrea Rodr\u00edguez, and Ricardo A. Baeza-Yates. Scheduling algorithms for web crawling. pages 10--17, 2004."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0169-7552(98)00108-1"},{"key":"e_1_2_1_5_1","first-page":"923","volume-title":"Introduction to algorithms","author":"Cormen Thomas H.","year":"2006","unstructured":"Thomas H. Cormen , Charles E. Leiserson , Ronald L. Rivest , and Clifford Stein . Introduction to algorithms . pages 923 -- 931 , 2006 . Thomas H. Cormen, Charles E. Leiserson, Ronald L. Rivest, and Clifford Stein. Introduction to algorithms. pages 923--931, 2006."},{"key":"e_1_2_1_6_1","first-page":"689","volume-title":"Thinking in java","author":"Eckel Bruce","year":"2000","unstructured":"Bruce Eckel . Thinking in java . pages 689 -- 823 , 2000 . Bruce Eckel. Thinking in java. pages 689--823, 2000."},{"key":"e_1_2_1_7_1","first-page":"106","volume-title":"An adaptive model for optimizing performance of an incremental web crawler","author":"Edwards Jenny","year":"2001","unstructured":"Jenny Edwards , Kevin S. McCurley , and John A. Tomlin . An adaptive model for optimizing performance of an incremental web crawler . pages 106 -- 113 , 2001 . Jenny Edwards, Kevin S. McCurley, and John A. Tomlin. An adaptive model for optimizing performance of an incremental web crawler. pages 106--113, 2001."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1019213109274"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1137\/0206024"},{"key":"e_1_2_1_10_1","volume-title":"Jeffrey Scott Vitter, and Ramesh C. Agarwal. Characterizing web document change. 2118:133--144","author":"Lim Lipyeow","year":"2001","unstructured":"Lipyeow Lim , MinWang, Sriram Padmanabhan , Jeffrey Scott Vitter, and Ramesh C. Agarwal. Characterizing web document change. 2118:133--144 , 2001 . Lipyeow Lim, MinWang, Sriram Padmanabhan, Jeffrey Scott Vitter, and Ramesh C. Agarwal. Characterizing web document change. 2118:133--144, 2001."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/371920.371965"},{"key":"e_1_2_1_12_1","first-page":"587","volume-title":"Java complete reference","author":"Schildt Herbert","year":"2002","unstructured":"Herbert Schildt . Java complete reference . pages 587 -- 626 , 2002 . Herbert Schildt. Java complete reference. pages 587--626, 2002."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/876875.879017"}],"container-title":["ACM SIGSOFT Software Engineering Notes"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1811226.1811236","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1811226.1811236","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T11:22:46Z","timestamp":1750245766000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1811226.1811236"}},"subtitle":["design and implementation of efficient text search algorithm in a web crawler"],"short-title":[],"issued":{"date-parts":[[2010,7,20]]},"references-count":13,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2010,7,20]]}},"alternative-id":["10.1145\/1811226.1811236"],"URL":"https:\/\/doi.org\/10.1145\/1811226.1811236","relation":{},"ISSN":["0163-5948"],"issn-type":[{"type":"print","value":"0163-5948"}],"subject":[],"published":{"date-parts":[[2010,7,20]]},"assertion":[{"value":"2010-07-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}