Improved data clustering methods and integrated A-FP algorithm for crop yield prediction

P. Suvitha Vani,S. Rathi
DOI: https://doi.org/10.1007/s10619-021-07350-1
IF: 0.974
2021-07-15
Distributed and Parallel Databases
Abstract:<p class="a-plus-plus">Big data analysis is the process of gathering, managing and analyzing a large volume of data to determine patterns and other valuable information. Agricultural data can be a significant area of big data applications. The big data analysis for agricultural data can comprise the various data from both internal systems and outside sources like weather data, soil data, and crop data. Though big data analysis has led to advances in different industries, it has not yet been extensively used in agriculture. Several machine learning techniques are developed to cluster the data for the prediction of crop yield. However, it has low accuracy and low quality of the clustering. To improve clustering accuracy with less complexity, a Proximity Likelihood Maximization Data Clustering (PLMDC) technique is developed for both sparse and densely distributed agricultural big data to enhance the accuracy of crop yield prediction for farmers. In this process, unnecessary data is cleansed from the sparse and dense based agricultural data using a logical linear regression model. After that, the presented clustering method is executed depending on the similarity and weight-based Manhattan distance. The genetic algorithm (GA) is applied with a good fitness function to select the features from the clustered data. Finally, the decision support system is computed by the A-FP growth algorithm to predict the crop yields according to their selected features such as weather features and crop features. The results of the proposed PLMDC technique are better in case of clustering accuracy of both spare and densely distributed data with minimum time and space complexity. Based on the results observations, the PLMDC technique is more efficient than the existing methods.</p>
computer science, information systems, theory & methods
What problem does this paper attempt to address?