Domain Specific Parallel Crawler Architecture
DOI:
https://doi.org/10.37591/rtpc.v3i1.697Abstract
The World Wide Web is an interlinked gathering of billions of reports organized utilizing HTML. Because of the developing and dynamic nature of the web, it has turned into a test to cross all URLs in the web archives and handle these URLs, so it has ended up basic to parallelize a creeping prepare. The crawler process is further being parallelized in the shape nature of crawler specialists that parallel download data from the web. This paper proposes a novel engineering of parallel crawler, which depends on area particular creeping, makes slithering assignment more compelling, versatile and load-sharing among the distinctive crawlers which parallel download website pages identified with various spaces particular URLs.
Keywords: URL (uniform resource locator), URI (uniform resource identifier), crawl work
Cite this Article
Himanshu Verma. Domain Specific Parallel Crawler Architecture. Recent Trends in Parallel Computing. 2016; 3(1): 17–21p.
References
Burner M. Crawling towards Eternity: Building An Archive of The World Wide Web. In Web Techniques Magazine. 1997; 2(5): 37–40p.
Yadav D, Sharma AK, Gupta JP, Garg N, Mahajan A. Architecture for Parallel Crawling and Algorithm for Change Detection in Web Pages. In Proceedings of the 10th international Conference on information Technology (December 17 - 20, 2007). ICIT. IEEE Computer Society, Washington, DC. 2007: 258–264p.
Yu C, Lin S. Parallel Crawling and Capturing for On-Line Auction. In Proceedings of the IEEE ISI Paisi, Paccf, and SOCO international Workshops on intelligence and Security informatics (Taipei, Taiwan, June 17 - 17, 2008) Springer-Verlag, Berlin, Heidelberg. 2008; 5075: 455–466p.
Balamurugan Newlin, Rajkumar, Preethi J. Design and Implementation of a New Model Web Crawler with Enhanced Reliability. 2008.
Brin S, Page L. The anatomy of a large-scale hypertextual Web search engine. In Computer Networks and ISDN Systems. 1998; 30(1–7): 107–117p.
Cho J, Garcia-Molina H. Parallel Crawlers. In WWW’02. 11th International World Wide Web Conference. 2002.
Junghoo Cho, Hector Garcia–Molina. The Evolution of the Web and implementation for an incremental crawler. Prc. of VLDB Conf. 2000.
Heydon A, Najork M. Mercator: A scalable, extensible Web crawler. In World Wide Web. 1999; 2(4): 219–229p.
Downloads
Published
Issue
Section
License
Declaration and Copyright Transfer Form
(to be completed by authors)
I/ We, the undersigned author(s) of the submitted manuscript, hereby declare, that the above manuscript which is submitted for publication in the STM Journals(s), is not published already in part or whole (except in the form of abstract) in any journal or magazine for private or public circulation, and, is not under consideration of publication elsewhere.
- I/We will not withdraw the manuscript after 1 week of submission as I have read the Author Guidelines and will adhere to the guidelines.
- I/We Author(s ) have niether given nor will give this manuscript elsewhere for publishing after submitting in STM Journal(s).
- I/ We have read the original version of the manuscript and am/ are responsible for the thought contents embodied in it. The work dealt in the manuscript is my/ our own, and my/ our individual contribution to this work is significant enough to qualify for authorship.
- I/We also agree to the authorship of the article in the following order:
Author’s name
1. ________________
2. ________________
3. ________________
4. ________________
| We Author(s) tick this box and would request you to consider it as our signature as we agree to the terms of this Copyright Notice, which will apply to this submission if and when it is published by this journal. |