Deep Learning for Plant Leaf Disease Detection: A Comprehensive Review of Techniques and Early Detection Applications
T. Thilagavathi1, Dr.V. Jude Nirmal2
1Department of Artificial Intelligence, 2Department of Computer Science, St. Joseph’s College (Autonomous), Affiliated to Bharathidasan University, Tiruchirappalli 620002, India;
ISSN: 2583-5343
International Journal of Information Texchnology, Research & Applications, Vol. 5 No. 3: September 2026

Article Info
Article history:Received June 10, 2026; Revised July 30, 2026; Accepted August 2, 2026

Keywords:Plant Disease Categorization; Convolutional Neural Network; Segmentation; Feature Extraction; Classification
ABSTRACT

Due to the growing global population and the ensuing rise in food consumption, agriculture continues to be every nation’s top priority and most recent research topic. Particularly in applications involving visual data, audio data, and text analysis, developments in autonomous learning and feature extraction have drawn both academics and industry have shown a lot of interest. Agriculture plant protection, more specifically plant disease detection, has drawn a lot of interest among these applications. The subjectivity involved in disease spot feature selection can be addressed by using methods like Deep Learning (DL) and Machine Learning (ML) methods, which will result in more objective feature extraction and quick technological adoption. This review includes the works published between the year 2014 to 2023 that were acquired from dependable databases like Scopus and Web of Science. One hundred and three peer-reviewed articles were examined using search terms like "plant leaf disease diagnosis", "ML and DL methods." In this review, categorization techniques for plant diseases are systematically compared, with a focus on ML algorithms exactly Convolutional Neural Network (CNN) algorithm. The various aspects, such as experimental setup, metrics are considered for disease classification, the processing methods for each algorithm including image segmentation and feature extraction are considered in this study. Researchers looking to recognize certain plant diseases using data-driven methodologies will find the systematic comparison of techniques to be very helpful in highlighting the advantages and disadvantages of various approaches. Furthermore, this study points towards promising future research directions, including the exploration of novel ML and DL architectures, integration of multi-modal data for improved accuracy, and the development of robust models to tackle complex disease scenarios. By leveraging the knowledge presented in this study, researchers can further enhance the detection of disease methods and back to the advancement area of agriculture.
This is an open access article under the CC BY-SA license.
image: e_a965531e138c_e_91aacb5cfe7d_CC_BY-SA_icon_svg.png
           

Corresponding Author:
T. Thilagavathi
Department of Artificial Intelligence,
St. Joseph’s College (Autonomous),
Affiliated to Bharathidasan University,
Tiruchirappalli 620002, India;
Email: thilagamca96@gmail.com

Introduction

Many farmers and agricultural researchers are concerned about the identification of plant diseases. Diseases can readily spread from one plant to the entire crop, posing numerous threats to both the agriculture yield and the overall economy [1]. Statistics show that plant illnesses brought on by nematodes, viruses, bacteria, and fungi result in an annual economic cost of $220 billion globally. [2]. Plant diseases also negatively affect agricultural output, increasing the number of people experiencing hunger. In 2024, it is projected that there will be 5 million hunger-related fatalities, which is ten times more than the COVID-19 fatalities over the same period [3]. The majority of these deaths occur in less developed nations and areas. If plant diseases are not quickly detected, food scarcity will increase. In recent years, the importance of detecting plant disease has increased.
ML and DL technologies in agriculture are effective for early disease detection, aiding farmers in preserving crop yields and preventing disease spread. Techniques like fluorescence imaging and DNA/RNA-based affinity biosensors have been widely used to assess leaf diseases, but they face issues such as inadequacy, inconsistency, and scalability. ML and DL methods address these challenges, providing solutions that benefit agriculture. Image processing procedures play a crucial role in accurately recognizing and categorizing diseases at early stages [4, 5, 6].
Recent studies on plant disease detection and classification have primarily focused on the utilization of either ML or DL algorithms. These algorithms play a crucial role in identifying and categorizing leaf diseases. The suggested study incorporates new developments in DL models for plant disease detection [7]. The research highlights existing gaps in the literature, emphasizing the need for a better understanding of observed symptoms during leaf disease classification [8].
The agricultural decision-making process heavily relies on early identification, serving as the foundation for successful anticipation and disease control. Identifying the source of plant diseases has become increasingly significant in recent years. Symptoms of plant diseases are often visible on the plant itself, such as lesions or markings on leaves, stems, and fruits. Diagnosis of anomalies is relatively simple in most cases due to the distinctive visual patterns associated with each illness or pest situation. Leaves, being the primary source for diagnosis, typically exhibit the initial indications of plant diseases ([9]).
image: e_2b269c02354f_fig1.png
Figure 1. The flow of image processing
Figure 1 illustrates that image processing entails a sequence of systematic steps designed to manipulate and analyze digital images. These steps aim to extract valuable information, enhance visual quality, or accomplish specific tasks.

1.1 The Review’s Contribution

The following is a summary of this contribution:

Related Work

LiLi et al. [25] discussed the identification of plant leaf diseases, exploring cutting-edge methods such as deep learning, hyperspectral imaging, and image processing. They delved into strategies to enhance classification accuracy, including the construction of extensive datasets, data augmentation, transfer learning methods, and the display of CNN activation maps. The importance of hyperspectral imaging in disease detection, especially with small samples, was emphasized.
Javid Ahmad Wani et al. [26] examined strategies for autonomous agricultural disease diagnosis, focusing on the potential use of Machine Learning models for the early detection of plant illnesses. The study covered four different crops (tomato, rice, potato, and apple) and evaluated various Machine Learning and Deep Learning-based classification models while addressing a large number of online datasets for the detection of plant diseases.
Muhammad Hammad Saleem et al. [27] provided an overview of the literature on the use of Deep Learning (DL) models for plant disease detection and classification. Several DL structures and visualization techniques for recognizing and classifying plant diseases were discussed. The study also demonstrated how the accuracy and effectiveness of DL algorithms could be evaluated using the Plant Village dataset.
Hasan et al. [28] examined farmers’ historical battles with pests and diseases in agriculture. While chemical pesticides were effective, their widespread usage endangered human health and the environment. The emphasis shifted to sustainable agricultural production increases, calling for the early identification of pests and pathogens to control them effectively. Recent developments in remote sensing and image analysis technology made it possible to identify infections and pests precisely, non-intrusively, and in real-time.
Ayat Mohammad et al.[29] emphasized the importance of infectious plant illnesses produced by numerous pathogenic bacteria, significantly affecting plant quality and production losses. Pathogen detection and treatment were critical for global food security and sustainable agriculture. Due to restrictions such as portability, time consumption, and specialized user needs, traditional methods for plant disease control were being superseded with electronic monitoring (E-monitoring). E-sensors could be used in conjunction with traditional approaches to improve disease control.
Poornima Singh Thakur et al. [30] underlined the global effect of crop diseases and the difficulties in manual inspection, resulting in significant crop losses. Smart agriculture solutions, based on vision-based machine learning algorithms, were being investigated for pest and plant disease control. The suitability of these technologies for real-time applications and their implementation in IoT-based smart agriculture solutions remained mainly unexplored.
C. Jacklin et al. [31] discussed a detailed analysis of several methodologies used in plant disease diagnosis utilizing machine learning and deep learning approaches. The report also emphasized the importance of agriculture in India, as well as the necessity to improve agricultural productivity by avoiding and treating diseases caused by bacteria, fungi, and viruses. The traditional Support Vector Classification methodology was utilized to detect plant leaf disease, while another method focused on tomato plant leaf disease diagnosis utilizing deep learning methodologies. It also covered the MLP (Multi-Layer Perceptron) idea and its application in classification.
Aanis Ahmad and others [32] addressed seven important issues, including the need for, accessibility of, and usability of datasets, imaging sensors and platforms for data collection, deep learning methods, generalization of deep learning models, disease severity estimation using deep learning, human accuracy comparison, and open research topics. It also identified open research gaps to assist further development and deployment of tools to improve plant disease detection and give disease management to help farmers.
Vinicius Bischoff et al. [33] conducted an initial search using electronic databases and retrieved 668 candidate studies. They discussed dominant frameworks in the literature for modeling disease detection systems and presented a comparative analysis of related works.
According to Zhang et al. [34], there were three ways in which GPDCNN outperformed both the standard CNN and the AlexNet models. By adopting a global pooling layer, the spatial resolution could be regained without adding further parameters to the training process. The proposed model was evaluated with datasets containing samples of six common diseases found on cucumber leaves, and the findings showed that the model properly diagnosed all of the diseases.
Atila et al.[35] developed the Efficient Net deep learning architecture, comparing its results to those of other top-tier techniques for recognizing illnesses in leaf plants. Models were trained using data from the Plant Village dataset. Instead of the original dataset’s 55,448 photos, 61,486 images from the larger dataset were used to train all the models. Deep learning models, such as the Efficient Net architecture, were taught using transfer learning. All model layers became customizable during the transfer learning phase. The Efficient Net architecture’s B5 and B4 models surpassed all others in terms of accuracy and precision on both the original and upgraded datasets.
Adedoja et al. [36] designed and developed a smart mobile detection of plant diseases deployed on a smartphone. The system used images of diseased leaves to diagnose the specific ailment using a mobile version of the NASNet CNN architecture. There was an Android and iOS app that could be used to snap pictures of plant leaves. A web service took the diagnosis provided by the CNN model and ran it through the system. A NASNet-Mobile CNN model was used to identify plant diseases in the images after receiving the plant leaf pictures from the customized mobile app over a web service. The illness detection accuracy of the NASNet-Mobile CNN model was 99.31%. of 95.48%, according to
Ozguven & Adem [37], who evaluated it on the 155 images used for training and testing. The recommended method also stated that the performance of the Faster R-CNN architecture might be enhanced by altering CNN settings for each image and set of areas to be detected. The proposed methodology fared better than the cutting-edge methodologies described in the aforementioned literature for the relevant parameters.
Tarek et al. [38] reviewed multiple deep-learning models trained on the ImageNet dataset. According to tests, MobileNetV3 Small achieved an accuracy of 98.99%, while MobileNetV3 Large reached an accuracy of 99.81%. The speed at which each model could make an accurate prediction was evaluated on a desktop computer using images of tomato leaves. To create an IoT device for identifying diseases in tomato leaves, the models from that location were also placed on a Raspberry Pi 4.
Zhang et al. [39] introduced a new hybrid clustering-based segmentation method for images of plant diseased leaves. The convergence of the expectation maximization (EM) approach was sped up by first dividing the full-color leaf image into a collection of compact and essentially uniform super pixels. These "super pixels" acted as reliable clustering signals that directed the segmentation of images of plant diseases.
Chen et al. [40] utilized photos of sick leaves to train a deep CNN for classifying tea plant illnesses. They created LeafNet, which automatically classified tea plant illnesses from photos. The sickness-detection abilities of the three classifiers were then assessed and compared. The LeafNet approach outperformed the SVM and MLP processes in diagnosing illnesses in tea leaves, achieving a classification accuracy of 90.16% versus 60.62%.
Chen et al. [41] developed an android-based CNN based on the AlexNet modification architecture to predict tomato illnesses using images of leaf surfaces. A prediction model was created by combining 4,585 testing data and 18,345 training data. Ten labels, each composed of 64*64 RGB pixels, were created to help identify diseases damaging tomato leaves. The optimal model was developed using the Adam optimizer with specific settings: a strict cross-entropy loss function, a realization rate of 0.0005, and a batch size of 128.
Storey et al. [42]created an instance-segmentation-based Mask R-CNN implementation. They built, assessed, and compared ResNet-50, MobileNetV3-Large, and MobileNetV3-Large-Mobile Mask R-CNN networks to evaluate their performance in image-based tasks, including object recognition, segmentation, and disease detection. The study demonstrated that the maximum accuracy was achieved by using a Mask R-CNN model trained on top of a ResNet-50 foundation, especially for identifying extremely minute rust disease particles on the leaves.
Lilhore et al. [43] proposed the ECNN model, which used a unique block processing feature to handle asymmetrical pictures. The suggested ECNN model incorporated a gamma correction function to overcome the color mismatch issue. Batch normalization and global average election polling were used to streamline the variable selection process and speed up computation. The study utilized 6,256 images of cassava leaves from five different disease groups from a freely accessible online Cassava image library, improving the accuracy of categorization with a larger and more varied sample of images.
Ali et al.[44] proposed a technique for identifying agricultural diseases (FF-PCA-LDA). Hybrid and deep features were manually detected and extracted from RGB images with the help of TL-ResNet50. The procedure’s efficacy was tested in a case study involving the identification of diseases in potato crop leaves. The dataset, not utilized during training, achieved an impressive 98.20% accuracy. This strategy outperformed other strategies by a wide margin and eliminated the need for human leaf segmentation due to its better discriminating and learning capabilities.
Ochago et al. [45] provided picture classification algorithms and feature extraction methods. The investigator upgraded the dataset of maize leaf diseases by visiting the Kaggle website. Filtered images of maize diseases were fed into a machine learning classification system. Photographs of both infected and uninfected maize were used to emphasize the contrasts between the two. A neural network trained on Histograms of Oriented Gradients outperformed KAZE, Oriented FAST, and Rotating BRIEF in terms of classification accuracy.
Chen et al. [46]described strategies to enhance the network model’s performance in target recognition. A training set and a test set were produced from the rubber tree disease database to evaluate the model’s effectiveness. The enhanced YOLOv5 network achieved an overall accuracy improvement of 70% over its predecessor. Compared to both the baseline YOLOv5 network and the YOLOX nano network models, plant diseases were effectively recognized in their natural environments using this updated model for future disease prevention and management efforts.
Thenchao et al. [47]utilized transfer learning and a deep residual network to identify and categorize strawberry diseases. The G-ResNet50, an improved version of the ResNet50, used a specialized loss function to focus on the worst illness instances. Using the pre-trained weight parameters from the Plant Village dataset, the ResNet50 model and the G-ResNet50 model trained with dropout regularization and batch regularization to further enhance the network model. Images of plants with leaf spot disease, strawberries with anthracnose, and healthy plants with powdery mildew were considered in this study.
Ye et al. [48] suggested employing attention processes and convolutional neural networks to recognize cassava leaves. The PDRNet (plant disease detection network) was designed to combat the massive dissemination of similar picture sets. To better accomplish its goals, this network combined a spatial attention mechanism and a channel attention mechanism. The pre-trained weights were included in the proposed model using a transfer learning technique, and the model’s identification accuracy was improved via fine-grained parameter tweaking and dynamic modifications to the learning rate. According to several different assessments, the suggested approach demonstrated good accuracy among CNN-based algorithms, achieving 99.56%.
Li et al. [49] discussed plant disease detection, pairing AlexNet with a modified version of Inception-V4 to produce a hybrid that boosts productivity without significantly compromising performance. The suggested model outperformed AlexNet, VGG11, Zenit, and VGG16, according to experimental findings on the expanded Plant Village dataset. The proposed model had the maximum accuracy for apples (96.5%), followed by maize (95.5%), tomatoes (94.8%), and grapes (92.3%). The F1 score for maize (0.938), tomatoes (0.91), grapes (0.94), and apples (0.924).
Research [25–49] shows that conventional methods are mostly used for growing specific types of crops. Despite this, there is a dearth of research into the topic, which drives the creation of self-sufficient AI models for disease detection in sunflower leaves.

2.1 Literature Sources and Search Technique

For this review, databases such as Google Scholar were searched for retrieving relevant articles published from 2014 to 2024.
Table 1. Search Terms
Search Term Set of Keywords
Plant disease Plant disease diagnosis
Machine learning techniques Decision Tree, Artificial Neural Network, Naive Bayes ensemble learning, random forest, support vector machine, and k-nearest neighbours clustering.
Machine Machine Learning
Deep Learning Techniques Deep Belief Networks, Deep Neural Networks, Deep Convolutional Neural Networks, Recurrent Neural Networks, Deep Autoencoders, Long Short-Term Memory Deep Reinforcement Learning, Deep Boltzmann Machine, and Extreme Learning Machine
Deep Deep Convolutional Neural Network, Deep Learning

Materials and Methods

3.1 Dataset Description

Data is essential for deep learning models; even the most sophisticated models are ineffective without high-quality input data. The common data split is often 70% for training, 15% for validation, and 15% for testing. A typical deep learning dataset includes a training set, a validation set used for tuning hyperparameters, and a test set to assess performance. Evaluating a deep learning model involves testing it with a sample from the dataset, and this study covers crucial datasets.

3.1.1 Plant Village Dataset and BIFROST Dataset

The Plant Village open dataset currently includes 54,309 photographs of plant diseases. These images encompass 14 different varieties, including cherry, grape, orange, peach, bell pepper, potato, and raspberry. Additionally, there are disease photographs for soybeans, pumpkins, strawberries, tomatoes, and corn. The corn category specifically contains 26 disease photographs, along with 12 images depicting healthy crop leaves. The diseases are further categorized into 17 fungi, 4 bacteria, 2 mycoses, 2 viruses, and 1 mite disease.
For review tasks, a dataset available on Kaggle (https://www.kaggle.com/datasets). The dataset was retrieved on February 12, 2023.
Table 2. Kaggle and BIFROST public plant leaf datasets
Name Sum of Images Kind of View Classes Task Source References
PlantVillage Dataset 162,916 Unchanging background 38 Image classification Kaggle Marwan Adnan et al.[62], 2020
New Plant Diseases Dataset 87,000 Field data 38 Image classification Kaggle Lawrence et al.[61], 2020
Flowers Recognition 4242 Field data 4 Image classification Kaggle Husnul et al.[63], 2020
Plant Seedings Dataset 5539 Field data 12 Target detection BIFROST Jinzhulu et al. [64], 2021
Theed Detection in Soybean Crops 15,336 Unchanging background 4 Target detection Kaggle Ana Corceiro et al. [60], 2023

3.1.2 Pathology Dataset

CVPR-2020-FGVC7 "Plant Pathology Challenge" (https://www.kaggle.com/c/plantpathology - 2020 fgvc7) comprises 3,651 annotated RGB photos, from this dataset 1,200 shows apple scabs, 1,399, 187, and 865 leaves images with various diseases.

3.1.3 Digi pathos

Recently, Digi pathos, a massive new dataset for plant illnesses [18]. It has 46,513 photos covering 171 diseases with 21 different crops. Only 2 326 photos in this collection depict illnesses of leaves, and those there either obtained in a lab with a consistent background or in the wild against a more complex backdrop. Cropped photos of disease lesions make up the remaining 44,187 photographs.

3.1.4 PlantDoc

Recently, a smaller library of plant disease photos named as PlantDoc. This dataset has 2598 images covering 17 diseases with 13 distinct crop types. While the majority of the photographs are captured in outdoors, there are other images with more standardized backdrops. Variations in picture-capturing conditions can be used to develop robust deep learning-based sickness diagnosis algorithms. Some of the photos in this dataset include many sick leaves or entire crops, making it problematic for models to acquire relevant disease traits. PlantDoc is also sketched, with a few examples of each type. Table 2 outlines the information encompassed in the PlantDoc dataset.

3.1.5 NLB Dataset

UAS-based airborne photography, a camera mounted on a pole used to collect the field dataset of 18,222 photos in (NLB) afflicted maize, prevalent foliar corn[19]. The NLB collection are real-field pictures annotated with 105, and 735 lesions. Images of maize leaves infected with several illnesses are included in this collection, making it useful for diagnosing more than one disease. Therefore, this dataset is more useful for object recognition and find NLB lesions or for DL to differentiate between healthy and sick maize plants. In addition, the testing pictures provided by this dataset may be used to evaluate the deep learning models’ capacity to generalize to new illness identification datasets.

3.1.6 RoCoLe: coffee disease dataset

The severity of coffee-specific illnesses was previously unknown, but a new annotated dataset was gathered [20]. The dataset includes 1560 photographs taken using a 5-megapixel camera in an Ecuadorian field with 390 coffee plants. Each coffee plant was photographed four times, and then manually annotated with ground truth data using a free software program.

3.1.7 Rice disease dataset

This is a collection of 3,355 photos affected with rice diseases. There are four distinct types of diseases present: leaf blast, brown spot, hispa, and healthy. However, the photographs in the collection are taken in a studio with a white background. As a result, it will be challenging to diagnosing rice illnesses in the field using such a dataset.

3.1.8 Cassava disease dataset

Recent field-based picture acquisition [21] adds to the growing body of data on cassava diseases. This dataset includes 5656 photos labelled as a disease. The dataset had complex backgrounds because it was collected in the field. This dataset may be used to train algorithms for diagnosing cassava infections in the leaf.

3.1.9 CD&S (corn disease & severity) dataset

The dataset includes 4455 photos of the three most common [22] corn diseases so that deep learning models can be trained to distinguish between them in the field. Three variants of original dataset are generated for the photos used in the training set by erasing the backdrop, swapping the field background for a black one, and swapping it for a white one, respectively. There are also five degrees of NLS disease severity depicted in visuals that may be used to train severity estimation algorithms.

3.2 Pre-Processing Techniques

Captured photographs of plant leaves often contain noise, unwanted backdrops, inadequate lighting, and other imperfections. Using classification methods on these raw images can result in inaccurate findings. Therefore, it is essential to preprocess the images before inputting them into a Convolutional Neural Network (CNN). This preprocessing step is crucial for reducing training time and enhancing accuracy in categorization. Various pre-processing techniques such as image scaling,color-to-grayscale conversion, normalization, enhancement, cropping, and Region of Interest (ROI) extraction are employed. Figure 2 illustrates some of the most common types of pre-processing procedures.
image: e_3a7868d6a223_fig2.png
Figure 2. Types of pre-processing
Evidence from literature reviews shows that several pre-processing methods are available for use in each of these settings. Pre-processing techniques must be used for a given dataset to be properly and successfully classified by a CNN model. The training process can be simplified and sped up by converting RGB pictures to grayscale [23], the single-Color channel is less computing and needs manifold Color channels. Incorporating PCA reduces the number of dimensions that data exists in. Whitening techniques such as Zero-Phase Component Analysis (ZCA) are analogous to PCA. It is used to draw attention to the relevant characteristics and structures to facilitate learning [17]. To emphasize the ROI, photos are cropped [24]. Before applying the correlation, coefficient technique is used to segment, a region, and stretch the contrast. Consequently, the visual quality of infected area is improved.

3.3 Image segmentation

Image segmentation is a computer vision and image processing task that involves dividing an image into distinct regions or segments based on certain characteristics. When it comes to the segmentation process, K-means is the most widely used technique. This technique has successfully identified the disease affected areas in the analyzed leaf photos. However, it has several drawbacks, such as a challenging k-value prediction phase. Because of this, it is difficult for the researcher to reliably provide the value of k for the big data sets during the training process. When employing the segmentation technique known as Active contour, it becomes possible to pinpoint the diseased area on a leaf. The segmentation procedure, while worthwhile, is time-consuming. The Otsu segmentation technique, which is much quicker than Active contour, has been utilized to circumvent this issue. Even yet, the threshold value is still determined by human intervention. Since the threshold value was computed automatically, a better histogram has been employed to identify the unhealthy area in the leaf picture. However, because of the binary approach’s lowered threshold value, it affects illness detection rates.
Table 3. Existing Strengths and Weakness of the Segmentation Technique
References Weaknesses Segmentation Technique Strengths
Majji et al. [65], 2021 The components of the item that are located outside of the box will be disregarded if the bounding box is too tiny. GAACO If the bounding box is not large enough, the portions of the object that are located outside of the box will be disregarded.
Xin Xu et al. [66], 2020 Prediction of the K-value is difficult. K-means It is an effective way for locating the infected tissue.
Story et al. [42], 2019 Low processing time Active Contour High accuracy
Ali et al. [44], 2022 Individual threshold values have to be determined by hand. Otsu Faster in computation
Li et al. [49], 2022 Reduces disease Improved Histogram Automatic generation of a threshold value.
Ye et al. [48], 2022 Detection rate. Genetic Algorithm Fully automatic
Zhang et al. [39], 2019 Takes more processing time. Grab cut No singular Contextual is desirable. It is functional and more stable than before. There is no necessity for the existence of special circumstances.
The segmentation procedure now includes the Genetic Algorithm to prevent this issue. It’s fully automated, it can properly detect illnesses, and it can function under any environmental situations like sunshine and darkness. A key drawback is the longer processing time. Grab cut is an algorithm that has been implemented during segmentation to prevent this from happening. It works and is stronger than before. There is no need for any special circumstances, and no prerequisite knowledge is expected. However, because of the initial narrow bounding box, the complete item is not covered. GAACO has been included in the segmentation phase to prevent this issue with procedure combinations, and it successfully extracts the diseased area of the leaf picture. To advance classification accuracy, investigators should use a variety of image segmentation procedures on the pictures and then use the best feature values obtained from images to make classifications.

3.4 Feature extraction

Each algorithm employed to solve the plant disease classification problem must go through a feature extraction phase. This feature extraction method, though, might not be required for DL models. As a result, a framework is now utilized by the great majority of plant disease detection and classification systems. [93, 94] Table 4 summarises the aspects that have been addressed by the various feature extraction methods.
Table 4. The Feature Extraction Procedures
References Features Procedure
Xu J.L et al. [25] Color, Texture (CCM)
Zhang et al. [39] Color (CCM)
Huang et al. [75] Texture (GLCM)
Adjabi et al. [74] Shape (MER)
Sun et al. [15] Color, Shape, Texture CCM, GLCM
Lu et al. [76] Other features (SIFT)

3.5 Classification

One specialized subfield of AI that focuses on machine learning is deep learning. It relies on the ability to identify patterns and draw inferences from data. The type of deep learning chosen depends on the availability of labels, with machine learning methods utilizing labelled training data, semi-supervised learning, and unsupervised learning (which uses no labels at all) being more commonly employed. Convolutional neural networks have proven successful in detecting the colors, textures of lesions, and other crucial plant characteristics. Deep learning emerges as a viable approach for documenting plant diseases.
The merits and shortcomings of various ML classifiers outlined in this study are summarized in Table 5. Training and evaluating ML systems generally take more time, and not all ML techniques scale well to massive datasets. To overcome these challenges, researchers have developed this system using DL methods.
Machine learning algorithms play a significant role in plant disease diagnosis as they can be trained to detect and classify diseases based on input data such as images of plant leaves showing disease symptoms [89, 90, 91, 92]. Here are some common techniques considered in this review.
image: e_e690a9a152c7_fig3.png
Figure 3. Machine learning techniques for diagnosing plant diseases considered in this review
Table 5. Existing strengths and Weaknesses of machine learning techniques
References ML Algorithms Weaknesses Strengths
Ayat mohammad et al., [29], 2022 SVM A sluggish segmentation would need more time both in training and assessment. It works poorly with massive datasets. There is optimism that it can be harnessed to accurately detect and identify weeds.
Mohammad et al. [52], 2019 KNN Picking the right k-value is tricky. For data with several dimensions, it is useless. No training is required prior to usage, leading to a reduction in both recognition time and computational complexity.
Adem et al. [37], 2019 NB It works on offline data only. A fast, accurate, and simple classifier that can handle a sizable data set.
Chittabarni et al. [53], 2023 BPNN More training period is obligatory High accuracy.
Rahul Sharma et al. [54], 2021 DT Overfitting is problematic and needs more training period. High accuracy.
Panuwat Mekha et al. [55], 2021 RF It needs more training time. Accurate to a high degree, can deal with missing data, but violates the overfitting issue.
XceptionNet stands out as one of the extensively employed DL models for classifying plant diseases, alongside other notable models such as ResNet and Inceptionv3. Current research suggests that augmenting the number of layers in a DL model can enhance its performance. Despite having a considerable number of layers and an overwhelming number of parameters, certain deep learning models, including XceptionNet, exhibit superior classification techniques. The latest DL models, such as Inceptionv3, ResNet, and XceptionNet, achieve impressive depth reduction with parameters, resulting in classification accuracies of 99.76%, 96%, and 98.7%, respectively. GoogLeNet, with more layers and fewer parameters, achieves a classification accuracy of 99.35%.
To overcome the limitations outlined in Table 6, researchers are encouraged to construct systems using a combination of ML and DL algorithms. The ability to learn intricate patterns from large datasets makes these techniques particularly effective for plant diagnosis [84]. Deep learning algorithms, frequently employed for diagnosing plant diseases, are highlighted in [101].
image: e_d505067ce199_fig4.png
Figure 4. Tree Illustration of Deep Learning Models for Plant Disease Diagnosis
Table 6. Assessment of Deep Learning Procedures with Limitations.
References Deep leaning algorithms Limitations Number of Layers Parameters in millions Classification accuracy (%)
Zhang et al. [34], 2019 AlexNet Over-fitting problem. 8 60 95.5
Aanis Ahmad et al. [32], 2023 VGGNet Requiring a lot of computation and being difficult to implement on systems with few resources. 16 138 99.53
Sachin et al. [50], 2020 GoogLeNet There is a chance that information will be lost when feature space in the buried layers shrinks. 22 5 99.35
Zhi et al. [47], 2022 Inceptionv3 Architecture design is difficult to grasp. 48 24 99.76
Poornima singh thakur et al. [30], 2022 ResNet There may be many layers that combine to produce very little or longer information. 50 26 96
Prabira et al.[51], 2020 XceptionNet The computational cost is more. 71 23 98.7
The primary limitations, identified through a comparison of data in Tables 5 and 6, apply universally to all the ML/DL algorithms investigated. These constraints include:
  1. Poor addressing of visual symmetries in symptoms of illnesses.
  2. Employment of CNN architecture to train DL models with higher accuracy on smaller datasets resulting in inaccurate outcomes.
  3. Insufficient availability of additional photos for effectively training and evaluating each DL model under examination

3.5.1 Optimization Techniques

Improving a Convolutional Neural Network model’s performance is crucial, hence it’s essential to use an appropriate optimization strategy. In Table 10, some of the most popular optimization strategies are compared.
Table 7. The Benefits and Drawbacks of Optimization technique
References Name of Optimizer Disadvantages Advantages
Vijay pal singh et al. [59], 2021 BGD The calculation of gradients throughout the whole dataset calls for a significant amount of memory. As the weights are altered following the calculation of the gradient on the entire dataset, the amount of time needed to converge to minimum values are increases. Easy to compute, device, and understand.
Rajasekaran et al.[58], 2020 SGD SGD use the plenty of hyperparameters and iterations. This makes it extremely dependent on the scale applied to the features. It might still fire after reaching a local minimum. Simple to put into action. Effective when working with data sets of a significant magnitude. Because it performs updates more often than batch gradient descent, it converges much more quickly. Because the values of loss functions do not have to be stored, less memory is required to operate them.
Hesham Tarek et al. [56], 2022 AdaGrad The high computational cost stems from the necessity of computing the second-order derivative. Training progresses slowly since the learning rate is always falling. Automatically adjusts the learning rate based on different training variables, eliminating the need for manual intervention. So while dealing with sparse data, the system performs well under such conditions
Mohamed Bound et al. [57], 2022 RMSProp The learning rate is still handcrafted. Pseudo-curvature data makes a powerful optimizer even more so. As a result of its proficiency with stochastic goals, it may be used for min-batch training.
Chen et al.[41], 2022 Adam Costly computationally. Adam can close the gap between two points with an incredible speed.

3.6 Frameworks

With advancements in AI, deep learning can now be applied across various domains, utilizing a diverse range of tools and platforms. Table 8 provides a quick summary of several popular frameworks and how they are often applied.
Table 8. Comparative analysis of CNN frameworks
References Framework CUDA Open Source Programming Language Used for Development OpenCL Support Interface OpenMP Support
Bastien et al. [67] 2012 Theano Support Yes Python Under development Python Yes
Chollet et al. [68], 2017 Keras Yes Yes Python TensorFlow as backend R, Python Yes
Jia Y et al. [69], 2014 Caffe Yes Yes C++ Under development C++, MATL B, Python Yes
Ferentinos et al [70], 2018 Torch Yes Yes C, Lua Third-party implementations Lua, Lua JIT, C, C++/OpenCL Yes
Amara et al. [71], 2017 deep learning4j Yes Yes C++, java No Kotlin, Java, Scala, Python, Clojure Yes
Bhatt et al. [72], 2017 TensorFlow Yes TensorFlow No No No Linux, macOS,
Shakoor et al. [73], 2017 DL Matlab Toolbox No MATLAB, Java, C, C++ No MATLAB No

Performance Metrics Used in Results

The publications surveyed made use of a wide range of performance indicators [99, 100]. Precision and recall, F1 score, are frequently used as markers of judgment when evaluating these algorithms/architectures. Each of these measures, as well as its definition, description, used in the survey, are detailed in Table 9.
Table 9. Performance evaluation metrics used in related work
References Description Performance Evaluation Metrics
Ochago et al. [45], 2022 This is the accuracy is used to classifications. 82%
Wenchao, [47], 2022 The accuracy of a classification system is measured by how many instances belong to the class for which classification was made (the "true positive" rate). 93.67%
Jinzhulu et al. [64], 2021 F-measure delivers a single score that balances both concerns in one number 91%
Daniyaa et al. [77], 2019 One statistical indicator of the extent to which an independent variable account for the variance in a dependent variable is the determination coefficient. 0.08%
Rajasekaran et al. [58], 2022 It is a standard deviation of the errors that occur when a prediction is made on a dataset. 0.83%

Discussion

Modern agricultural practices must be improved with cutting-edge technologies in order to sustainably feed the world’s growing population [85]. There are many subfields available within the broader field of "smart agriculture," such as sensor design and implementation, integrity, expert system design and implementation, and AI-based decision-making. The number of AI and deep learning-based efforts in the agricultural sector has skyrocketed during the last two decades. The study’s findings show that deep learning techniques perform better and are more useful in this area than traditional methods [87, 88].

5.1 Compensations and Disadvantages of Deep Learning

Deep learning’s strength is in its ability to automatically derive features and fine-tune them to get the desired result. One other perk is that the same neural network-based strategy may be used for several tasks and forms of data. What’s more, the architecture of deep learning is malleable enough to be used for solving different challenges down the road [95, 96]. When it comes to generalization, deep learning also performs admirably. Although deep learning models take longer to train than conventional methods, but testing is much quicker. One of the noteworthy downsides of deep learning is that it requires large datasets to deliver greater performance, even if data augmentation techniques are applied. The agricultural research community suffers from a dearth of publicly available datasets, forcing many scientists to create their picture collections [97].

5.2 Future of Deep Learning in Agriculture

Since every place has its unique climate, wildlife, and topography, agriculture is one of the more intricate fields of use. Therefore, technology to identify the relevant factors and examine the data is desperately needed. Considering the dynamic nature of the data, this demands a massive amount of time and resources to study. To this end, deep learning is a crucial technology since it employs the right to do these tasks. Given information about the environment, such as climatic characteristics, soil types, and weather patterns, a decision-making algorithm can construct a probabilistic model. The ability to accurately and quickly identify various illnesses is crucial for early tracking to prevent food or money losses. This study can analyse a decade’s worth of diseased plant photos and pinpoint the exact disease and its severity. One major benefit of using a deep learning model is that it allows the program to generate the desired feature without any human intervention. The unpredictability and rapid evolution of the real-world workplace are best prepared for via unsupervised learning. This is becoming more significant due to the growth of the Internet of Things and the fact that the great majority of data generated by machines and people today is unstructured and unclassified [86]. As compared to more conventional approaches, such as ANN, SVM, RF, etc., the results from using deep learning are far more promising. The deep learning model automatic feature extraction is more effective than manual feature extraction.

Open Challenges – Plant Disease Diagnosis

Numerous issues remained unresolved in the field of plant disease diagnosis.These difficulties aim to improve the accuracy, speed, and effectiveness of disease detection and diagnosis in order to lessen the impact of plant diseases on agriculture and food security. Here are a few examples of common difficulties:
  1. Early Detection and Monitoring: One of the primary challenges is the early detection and continuous monitoring of plant diseases. Developing technologies that can identify subtle symptoms and signs of diseases in their early stages can help to prevent their rapid spread and minimize crop losses.
  2. Data Collection and Standardization: Collecting accurate and comprehensive data for disease diagnosis is essential. However, there is often a lack of standardized data collection methods and formats. Developing consistent data collection protocols and sharing standardized datasets can enhance the accuracy of diagnostic tools.
  3. Variability in Symptoms: Plant diseases can exhibit a wide range of symptoms, and these symptoms can vary based on factors like environmental conditions, pathogen strains, and host plant varieties. Creating diagnostic tools that can accurately identify diseases despite symptom variability is a challenge.
  4. Integration of Multiple Data Sources: Effective disease diagnosis often requires the integration of data from various sources, such as visual observations, sensor data, molecular analyses, and remote sensing. Developing methods to combine and interpret these diverse data types can lead to more accurate diagnoses.
  5. Automated and High-Throughput Analysis: Traditional disease diagnosis methods can be time-consuming and labor-intensive. Developing automated and high-throughput techniques, such as robotics, drones, and AI-powered image analysis, can expedite the process and allow for rapid screening of large agricultural areas.
  6. Pathogen Identification and Classification: Identifying the specific pathogens responsible for plant diseases is crucial for implementing targeted management strategies. However, many pathogens are difficult to culture and identify. Advances in molecular techniques, like DNA sequencing, can aid in accurate pathogen identification.
  7. Real-Time Disease Monitoring: Establishing real-time disease monitoring systems that provide instant updates on disease presence and progression can help farmers make timely decisions regarding disease management and mitigation strategies.
  8. User-Friendly Tools for Farmers: While advanced diagnostic technologies exist, they may not always be accessible or user-friendly for farmers, particularly in remote or resource-constrained areas. Developing user-friendly, cost-effective, and portable diagnostic tools that can be used by farmers directly can improve disease management at the grassroots level.
image: e_dced1c33a706_fig5.png
Figure 5. Open Challenges - Plant Disease Diagnosis
Table 10. Open Challenges – Plant Disease Diagnosis
References Dataset Quality and size Automated disease detection Robustness to environmental variations Interpretability and explain the ability Integration with precision agriculture Privacy and data security Transferability to new plant species and diseases
Hedun Chun et al. [79], 2016 ✓ ✓ × ✓ × × ✓
Xinda liu et al. [80], 2021 ✓ × × ✓ × × ×
Haiqing wang et al. [81], 2022 ✓ ✓ ✓ × × × ✓
Hellenic et al. [82], 2018 ✓ × × ✓ × × ×
Sasikala et al. [83], 2021 ✓ × × × ✓ × ✓

Open Research Topics

It is crucial to correctly identify diseases and assess their severity when diagnosing and monitoring plant health. To improve security, several researchers have applied various image-processing, methods after the 2015 release of the Plant Village dataset [102]. There have been several successful attempts by researchers to develop algorithms and models for detection of plant disease as detailed in this article. End-to-end plant disease management systems are lacking; however, this is due to research gaps and restrictions.
Because each user must train the model in their arena under-regulated or pre-defined picture attainment settings, a new strategy is needed for identifying plant diseases using deep learning. Most of the examined studies can correctly identify illnesses, However, the trained models did not perform well when applied to images from different datasets or in different conditions. Therefore, it is crucial to verify that a trained model is detecting illness from all datasets and situations to develop a reliable model for effectively detecting diseases.
Current deep learning-based systems currently lack the capacity to discriminate between various diseases [98]. The majority of plant disease detection research has been on developing deep learning models for a particular disease in a single crop or a limited set of diseases. These methods may be successful in detecting diseases that are present in the training set, but they will miss diseases that are absent from the dataset. This problem needs to be overcome to establish a disease diagnosis system for precision agriculture.
Several studies have used a variety of different criteria for estimating severity. However, to train deep learning models that can reliably estimate the severity of illnesses, a standardized definition must be employed. When used for severity estimate, picture classification fails because it cannot precisely locate plant disease lesions and symptoms. Object detection gets beyond image classification’s shortcomings and aids in pinpointing illness spots. However, object identification can only generate rectangular bounding frames, thus healthy tissue outside of disease lesions and symptom hotspots may be flagged as suspicious. That’s why segmentation is the best method for determining where infections are concentrated on a leaf relative to the whole leaf.
Finally, proper management approaches to enhance crop output depend on the timing of disease discovery throughout the season.

Future Research Directions

Certainly, with advances in technology and study of detecting and classifying plant diseases has advanced quickly. The utilization of a dataset, data segmentation, deep learning model architecture, training/validation, performance measures, and visualization approaches are future research directions for the detection and classification of plant diseases.
1. Dataset Collection and Augmentation:
Gathering a comprehensive and diverse dataset comprising images of both healthy and diseased plants and enhancing the size and diversity of the dataset by employing techniques such as rotation, flipping, scaling, and colour changes.
2. Data Splitting:
Separating the dataset into the training, validation, and testing subsets. A typical division may be as follows: 15% for testing, 15% for validation, and 70% for training.
3. Deep Learning Model Architecture:
Choosing a suitable pre-existing convolutional neural network (CNN) architecture like ResNet, VGG, or Inception as a base and fine-tuning the architecture to fit the problem. Layers can be added or modified to make it more suitable for plant disease classification.
4. Model Training and Validation:
Creating a deep learning model that is similar to transfer learning using the training dataset and utilizing the validation dataset to fine-tune hyperparameters, preventing overfitting, and selecting the best model.
5. Performance Metrics:
Selecting relevant evaluation metrics like Accuracy, precision, recall, ROC curve, F1-score metrics. Additionally, considering the confusion matrices to visualize the performance across different classes.
6. Visualization Techniques:
Visualizing the model’s training progress with learning curves, showing how loss and accuracy change over epochs. Creating heatmaps to highlight the areas of plant images that the model focuses on during classification. Use t-SNE or PCA to visualize high-dimensional feature embeddings in a lower-dimensional space.
image: e_77c8f205884f_fig6.png
Figure 6. Future Research Directions - Plant Disease Diagnosis

Conclusions

This research delves into the fundamental concepts of machine learning and conducts a thorough analysis of existing research on techniques for identifying plant leaf diseases. The study reveals promising potential for accurately diagnosing plant leaf diseases, contingent upon the availability of sufficient training data. Additionally, it underscores the advantages of disease detection in small sample plants, offering valuable insights into the field of agriculture.
However, it is crucial to acknowledge the challenges within the current deep learning frameworks proposed in the literature. These studies often demonstrate variations in performance across different datasets, indicating a lack of robustness in the models. To advance knowledge in this field, there is a critical need for more robust deep-learning models capable of handling the diverse range of disease datasets encountered in real-world scenarios.
To propel the field of deep learning-based disease recognition forward, the suggestion is to create a sizable dataset focusing specifically on sunflower leaves in natural situations. Such a dataset would address the limitations of existing datasets like the Plant Village, which consists of photos in controlled environments and may not adequately represent the diversity of plant diseases in real-world conditions.
This study contributes to the understanding of disease detection techniques, emphasizing the need for more robust models and datasets to drive progress in this area. By overcoming the identified challenges, there is a significant opportunity to enhance the efficiency of deep learning models in plant leaf disease recognition, ultimately benefiting agricultural practices and plant health management.

References