A Novel Approach for Performance Analysis of Dynamic Information Integration
DOI:
https://doi.org/10.37591/jospa.v6i1.2074Abstract
Abstract
Data analysis and management was not a problem few years ago as the amount of data generated was not as huge as to cause any complexity. But in the recent years the amount of data being generated has increased exponentially. Thus management and analysis has become a problem since the traditional database management systems were not designed to handle such large amounts of data. The relational database management systems still can handle large data sets, but it increases the complexity thus making it a difficult task. Hadoop comes into picture as a solution to this problem. Hive which is an open source data warehouse built on the Hadoop framework provides a solution to handle large datasets. It provides an SQL tongue called Hive Query Language (HQL) for querying and processing of large sets of data. But the problem with RDBMS is that it takes more time when queries are applied to convert data in rows to columns, i.e., horizontal to vertical. This limitation of Relational Database Management System (RDBMS) is improved by using the combination of UDF and HQL in Hive. With the help of this approach, many calculations that are outside the scope of built in RDBMS operations and functions in Hive like query many columns, combine several column values into one and transformations that are taking more time in RDBMS, can be solved easily. In the proposed work seven different data sets are taken from web for experimental results. Aggregating queries using RDBMS and Hive are run on these data sets with the combination of UDF. The results obtained on these data sets show that combination of UDF with HQL is better than RDBMS when aggregation queries are fired on horizontal data and to join many columns in one. In the proposed work datasets of varying sizes have been analyzed using RDBMS, i.e., MySQL and MS SQL and then using Hive. Different comparison has been done which shows the advantage of using Hive over RDBMS.
Keywords: MySQL, MS SQL, Hadoop, Hive, UDF, HQL, HiveServer2, Dlimit, Processing Big Data
Cite this Article
Vikash Kumar Garg, Ashish Oberoi, Manish Arora. A Novel Approach for Performance Analysis of Dynamic Information Integration. Journal of Advances in Shell Programming. 2019; 6(1): 27–41p.
Downloads
Published
Issue
Section
License
Declaration and Copyright Transfer Form
(to be completed by authors)
I/ We, the undersigned author(s) of the submitted manuscript, hereby declare, that the above manuscript which is submitted for publication in the STM Journals(s), is not published already in part or whole (except in the form of abstract) in any journal or magazine for private or public circulation, and, is not under consideration of publication elsewhere.
- I/We will not withdraw the manuscript after 1 week of submission as I have read the Author Guidelines and will adhere to the guidelines.
- I/We Author(s ) have niether given nor will give this manuscript elsewhere for publishing after submitting in STM Journal(s).
- I/ We have read the original version of the manuscript and am/ are responsible for the thought contents embodied in it. The work dealt in the manuscript is my/ our own, and my/ our individual contribution to this work is significant enough to qualify for authorship.
- I/We also agree to the authorship of the article in the following order:
Author’s name
1. ________________
2. ________________
3. ________________
4. ________________
| We Author(s) tick this box and would request you to consider it as our signature as we agree to the terms of this Copyright Notice, which will apply to this submission if and when it is published by this journal. |