An Implementation of Extracting Data and Mining Extracted Data from Web Pages
DOI:
https://doi.org/10.37591/jospa.v1i1.65Keywords:
Web, data, mashup, data clusteringAbstract
Web contains various information of particular object, which could be relevant as well as non-relevant are called as data records. It is necessary to extract relevant information from web pages. Web data extraction is the system which is used for extracting data from various web pages. Data present on web pages are in un-structured format. In the process of data extraction, we convert un-structured data into structured format. This paper contains web data extraction system, stages of making a mashup and the data mining concept for data clustering. Mashup is the process which provides functionality such as data retrieval, data source modeling, data cleaning/filtering, data integration, and data visualization. We use “Xtractorz” system for data extraction and mashup for data records. By using data mining, we analyze data from different sources and using text mining we cluster all the information. And we also use those data for providing value added services
References
Knoblock CA, Lerman K, Minton S, et al. Accurately and reliably extracting data from the web: A machine learning approach. Intelligent Exploration of the Web. Springer-Verlag, Berkeley, CA; 2003.
Chamberlin D, et al. (Eds). XQuery: A query language for XML. http://www.w3.org, 2001.
Huynh D, Mazzocchi S, Karger D. Piggy bank: Experience the semantic web inside your web browser. In: Proc. of ISWC. 2005.
Google Map Facility, http://maps.google.com, last accessed 12 October 2009.
Wong Jeffrey, Hong Jason I. Making Mashups with Marmite: Towards End-User Programming for the Web. Human-Computer Interaction Institute, Carnegie Mellon University, Pittsburgh, last downloaded 12 October 2009.
Kapow Technologies. Kapow Mashup Server 6.3 Robomaker User Guide. http://www.kapowtech.com, last accessed 12 October 2009.
Lerman K, Plangrasopchok A, Knoblock CA. Semantic labeling of online information sources. In: Pavel Shaiko (Ed.). IJSWIS, Special Issue on Ontology Matching. 2007.
Lee Y, Sayyadian M, Doan A, et al. eTuner: Tuning schema matching software using synthetic scenarios. VLDB Journal, Special Issue 2006.
Lixto Technologies. Lixto Visual Developer. http://www.lixto.com, last accessed 12 October 2009.
Downloads
Published
Issue
Section
License
Declaration and Copyright Transfer Form
(to be completed by authors)
I/ We, the undersigned author(s) of the submitted manuscript, hereby declare, that the above manuscript which is submitted for publication in the STM Journals(s), is not published already in part or whole (except in the form of abstract) in any journal or magazine for private or public circulation, and, is not under consideration of publication elsewhere.
- I/We will not withdraw the manuscript after 1 week of submission as I have read the Author Guidelines and will adhere to the guidelines.
- I/We Author(s ) have niether given nor will give this manuscript elsewhere for publishing after submitting in STM Journal(s).
- I/ We have read the original version of the manuscript and am/ are responsible for the thought contents embodied in it. The work dealt in the manuscript is my/ our own, and my/ our individual contribution to this work is significant enough to qualify for authorship.
- I/We also agree to the authorship of the article in the following order:
Author’s name
1. ________________
2. ________________
3. ________________
4. ________________
| We Author(s) tick this box and would request you to consider it as our signature as we agree to the terms of this Copyright Notice, which will apply to this submission if and when it is published by this journal. |