Data Extraction Using NLP for Unstructured Text Categorization
DOI:
https://doi.org/10.37591/josettt.v4i3.1319Abstract
Abstract
In this research project, I tried to crawl the web and files to create a dataset so that it can be fetched to FRL (Fuzzy rough set-based semi-supervised learning algorithm). The approach used in the project is with the help of semi-supervised learning that made use of unlabeled data for training typically a small amount of labeled data with a large amount of unlabeled data. We de ne and use various Information extraction and Web data mining techniques to categorize nouns and find their context while tokenizing it with nouns. The mentioned techniques are application of Natural Language Processing part-of-speech(pos) tagging. The Dataset is then created by categorizing these nouns and phrases with the context in a tree like structure which is termed as chunking in Natural Language Processing(NLP). It helps extraction of the nouns through particular toolkit library available in programming languages. After the categorization is done we process the data labeled out and give it to FRL so that it can rank them accordingly and the text categorization is complete.
Keywords: Natural Language Processing, POS tagging, Web Information Extraction, Rough Sets, Advanced Machine Learning, Semi-Supervised learning, Data Categorization
Cite this Article
Waarengeye Varun Vikram. Data Extraction Using NLP for Unstructured Text Categorization. Journal of Software Engineering Tools & Technology Trends. 2017; 4(3): 30–39p.
Downloads
Published
Issue
Section
License
Declaration and Copyright Transfer Form
(to be completed by authors)
I/ We, the undersigned author(s) of the submitted manuscript, hereby declare, that the above manuscript which is submitted for publication in the STM Journals(s), is not published already in part or whole (except in the form of abstract) in any journal or magazine for private or public circulation, and, is not under consideration of publication elsewhere.
- I/We will not withdraw the manuscript after 1 week of submission as I have read the Author Guidelines and will adhere to the guidelines.
- I/We Author(s ) have niether given nor will give this manuscript elsewhere for publishing after submitting in STM Journal(s).
- I/ We have read the original version of the manuscript and am/ are responsible for the thought contents embodied in it. The work dealt in the manuscript is my/ our own, and my/ our individual contribution to this work is significant enough to qualify for authorship.
- I/We also agree to the authorship of the article in the following order:
Author’s name
1. ________________
2. ________________
3. ________________
4. ________________
| We Author(s) tick this box and would request you to consider it as our signature as we agree to the terms of this Copyright Notice, which will apply to this submission if and when it is published by this journal. |