MCIP: Mining Crop Image Data on PySpark Data Frame Using Feature Selection and Cluster-Based Techniques

6Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

Abstract

In India, the yearly economic losses incurred due to crop-related issues, including pests and diseases, surpass a staggering amount of $500 billion. Leaf blight constitutes a significant determinant in the remarkable economic ramifications, primarily affecting farmers engaged in cultivating forage and grain sorghum, who bear the brunt of its consequences. Numerous crops are affected by this disease, including maize, rice, tomato, potato, millet, and onion. However, crop variety, different disease kinds, and environmental conditions make early disease identification difficult. Due to the variety of crops and diseases, existing approaches for disease classification and prediction need more broad applicability. These techniques use image preprocessing and segmentation to handle datasets with specified inputs and outputs, frequently resulting in data loss and erroneous categorization. Additionally, they ignore specialized datasets. To address these challenges, this study proposes an innovative approach leveraging the PySpark-based mining crop image data (MCIP) framework. MCIP employs Principal Component Analysis (PCA) to extract relevant features, subsequently utilized by the K-means algorithm to identify distinct subgroups. This approach, initially demonstrated on potato leaves, proves valuable for disease identification. Notably, MCIP isn't restricted to potatoes; it's adept at detecting diseases across agricultural crops. To validate, an experiment was conducted on a rice disease dataset. Evaluation metrics, including Accuracy, Silhouette score, speed, and F1 score, affirm MCIP's robustness. Impressively, MCIP exhibits exceptional speed and accuracy, nearly achieving 100% accuracy. This innovative model signifies a significant advancement over existing techniques, offering a promising solution to the pressing issue of crop disease management.

Cite

CITATION STYLE

APA

Chaudhary, Y., & Pathak, H. (2023). MCIP: Mining Crop Image Data on PySpark Data Frame Using Feature Selection and Cluster-Based Techniques. International Journal of Experimental Research and Review, 34, 106–119. https://doi.org/10.52756/ijerr.2023.v34spl.011

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free