View Article

  • A Comprehensive Review Of Image Processing Techniques: Classical Methods, Deep Learning, Applications, And Recent Advances

  • 1Telangana Tribal Welfare Residential Degree College (Boys), Boath@Adilabad, Telangana, India.
    2TGTWRDC GIRLS UTNOOR, Utnoor, Adilabad, Telangana, India.
    3Government Degree College (Arts and Commerce), Adilabad, Telangana, India.

Abstract

Image processing has evolved from conventional pixel-level operations to sophisticated artificial intelligence (AI)-driven approaches capable of automated visual understanding, interpretation, and decision support. This review provides a comprehensive overview of image processing techniques, their methodological evolution, major applications, and recent advances. The review examines fundamental image representation and quality assessment, preprocessing, image enhancement, restoration, segmentation, feature extraction, and image compression, followed by modern deep learning approaches based on convolutional neural networks (CNNs), transfer learning, Vision Transformers, generative AI, and foundation models. A structured narrative literature review and thematic comparative analysis were used to organize the literature according to processing technique, learning paradigm, computational requirements, interpretability, and deployment potential. Classical techniques remain valuable because of their computational efficiency, simplicity, and suitability for controlled image-processing tasks, whereas deep learning approaches enable automatic hierarchical feature learning and improved performance for complex visual analysis. Recent developments in Transformers, explainable AI (XAI), generative models, and vision foundation models have further expanded the capabilities of image processing by supporting contextual understanding, visual explanation, image synthesis, and generalization across tasks and domains. Applications across medical imaging, agriculture, remote sensing, autonomous transportation, industrial inspection, security and surveillance, multimedia, environmental monitoring, and biometrics demonstrate the broad relevance of these technologies. The review also highlights persistent challenges involving data requirements, computational cost, model interpretability, robustness, privacy, domain generalization, and real-time deployment. Future research is expected to increasingly integrate efficient AI architectures, explainability, edge computing, multimodal learning, and foundation models to develop scalable and trustworthy intelligent image-processing systems.

Keywords

Image Processing, Image Enhancement, Image Segmentation, Deep Learning, Convolutional Neural Networks, Vision Transformers, Explainable AI, Generative AI, Foundation Models, Multimodal Vision.

Introduction

× Popup Image

1.1 Background of Digital Image Processing

Digital image processing is a multidisciplinary field concerned with the computational acquisition, enhancement, transformation, analysis, and interpretation of digital images. It integrates concepts from computer science, electrical engineering, mathematics, signal processing, artificial intelligence (AI), and machine learning to convert raw visual information into meaningful data for human interpretation and automated decision-making. A digital image can be represented as a two-dimensional array of discrete picture elements, or pixels, where each pixel contains intensity or color information. The spatial organization of these pixels allows computational systems to perform operations such as enhancement, restoration, segmentation, feature extraction, classification, and recognition (Gonzalez & Woods, 2024).

The development of digital image processing has progressed considerably over several decades. During the 1960s, early applications primarily focused on image enhancement and restoration, particularly for satellite imagery and medical images. During the 1970s and 1980s, improvements in digital computing enabled more sophisticated filtering, edge detection, segmentation, and image analysis techniques. During the 1990s, feature extraction, pattern recognition, and statistical approaches became increasingly important for automated visual interpretation. Subsequently, the 2000s witnessed greater integration of machine-learning algorithms into image classification and recognition. The 2010s marked a major transition toward deep learning, particularly Convolutional Neural Networks (CNNs), which enabled automatic learning of hierarchical visual features. More recently, the 2020s have witnessed the rapid development of Vision Transformers (ViTs), generative models, multimodal systems, and foundation models, extending image processing from task-specific analysis toward more general visual understanding. The existing literature reviewed in this study reflects this progression from conventional operators toward deep learning and foundation-model-based approaches (Goodfellow et al., 2024; Gajanan et al., 2026).

1.2 Evolution from Classical to AI-Based Image Processing

Traditional image processing generally depends on explicitly designed mathematical operations and handcrafted features. Common approaches include histogram-based enhancement, spatial filtering, thresholding, edge detection, region-based segmentation, and feature descriptors such as Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), Oriented FAST and Rotated BRIEF (ORB), and Histogram of Oriented Gradients (HOG). These methods remain useful because they can operate with relatively limited computational resources and, in many cases, provide easily interpretable processing steps. SIFT, SURF, ORB, and HOG, for example, transform visual patterns into numerical descriptors that can subsequently be used for recognition and matching tasks.

The emergence of deep learning changed this paradigm by enabling models to learn features directly from image data. CNNs automatically learn hierarchical representations through convolution, activation, and pooling operations. ResNet introduced residual connections to facilitate the training of deeper networks (He et al., 2016), while U-Net demonstrated the effectiveness of encoder-decoder architectures for image segmentation (Ronneberger et al., 2015).

The subsequent development of Vision Transformers extended image analysis by applying self-attention mechanisms to image patches. Transformer-based architectures can model relationships between different regions of an image and have therefore become an important direction in computer vision. Generative models, including Generative Adversarial Networks (GANs) and diffusion models, further expanded image processing toward image synthesis, restoration, transformation, and data generation. Foundation models represent another major development, providing pretrained visual capabilities that can potentially be adapted across multiple downstream tasks. Segment Anything is an important example of this direction (Kirillov et al., 2023).

Thus, the evolution of image processing should not be understood as a complete replacement of classical techniques by AI. Instead, contemporary systems increasingly combine conventional preprocessing with learned representations. Traditional methods can provide efficient preprocessing and enhancement, while deep models can perform complex classification, segmentation, and recognition. Transfer learning and optimized architectures further provide opportunities to reduce training requirements and computational costs (Ankatwar, 2025; Yadav & Gajanan, 2026).

1.3 Importance of Image Processing in Modern Applications

Image processing has become an important component of modern scientific, industrial, agricultural, and intelligent systems. In medical imaging, enhancement, segmentation, and classification are used with MRI, CT, and X-ray images to support the analysis of anatomical structures and disease-related patterns. CNN-based approaches have substantially expanded automated medical image analysis (Litjens et al., 2017).

In agriculture, image processing supports crop monitoring, disease detection, pest identification, precision farming, and plant-health assessment. Drone, satellite, and multispectral imagery can provide information for identifying diseases, nutrient deficiencies, and other crop stresses. Recent work has particularly examined CNNs, transfer learning, multispectral imaging, and explainable AI for cotton and other agricultural applications (Ankatwar & Dhawale, 2024, 2025, 2026; Gajanan & Preetham, 2026).

Remote sensing uses image classification, segmentation, and feature extraction for land-use mapping, environmental monitoring, disaster assessment, and urban analysis. Autonomous vehicles employ image processing for detecting pedestrians, vehicles, lanes, and traffic signs, while object-detection methods provide important foundations for automated visual recognition (Zhao et al., 2019). Security and surveillance systems use image recognition, object detection, tracking, and biometric analysis. In industrial inspection, computer vision is applied to identify scratches, cracks, dimensional abnormalities, and manufacturing defects.

1.4 Research Gap

Although image processing has developed rapidly, an integrated understanding of classical and AI-based approaches remains important. Conventional techniques are frequently discussed according to individual operations such as enhancement, restoration, segmentation, or feature extraction, whereas modern studies increasingly focus on deep architectures, Transformers, explainability, and foundation models. A comprehensive perspective therefore needs to connect pixel-level processing with higher-level visual understanding.

A further consideration is that accuracy alone does not fully characterize an image-processing system. Computational complexity, memory requirements, data requirements, inference latency, interpretability, robustness, and deployment environment can substantially affect practical applicability. This is particularly relevant for mobile, embedded, agricultural, healthcare, and edge-computing environments. The reviewed literature also indicates increasing interest in explainable AI, lightweight architectures, federated learning, foundation models, and multimodal vision systems.

Accordingly, this review considers image processing as a continuum extending from classical image operators to deep learning, Transformer architectures, generative AI, and foundation models, while also examining their applications, computational characteristics, and emerging research challenges.

1.5 Aim and Objectives

The aim of this review is to systematically organize and critically examine classical and AI-driven image processing techniques, compare their computational characteristics and application domains, and identify emerging research directions involving explainable AI, lightweight networks, edge computing, multimodal systems, and foundation models.

The specific objectives are to:

  1. review fundamental concepts and image-processing operations;
  2. analyze image enhancement and restoration techniques;
  3. examine traditional and deep-learning-based segmentation approaches;
  4. review handcrafted and automatically learned feature representations;
  5. examine CNN, transfer-learning, and Transformer architectures;
  6. compare conventional and AI-based image-processing approaches;
  7. analyze major applications across different domains;
  8. examine explainability and deployment considerations; and
  9. identify current challenges and future research directions.

Figure 1. Evolution of Image Processing from Classical Algorithms to Foundation Models

Figure 1 should visually represent the technological progression of image processing from pixel-level computational operations to increasingly sophisticated AI-based visual understanding. Unlike the existing Figure 1, which primarily presents a general processing workflow, the revised figure should emphasize the historical and technological evolution of the field. The current manuscript already illustrates the conventional progression from image acquisition through preprocessing, enhancement, segmentation, feature extraction, classification, and real-world applications.

2. Materials and Methods

2.1 Review Design

The present study adopted a structured narrative literature review with thematic and comparative analysis to examine the development of digital image processing from conventional computational techniques to contemporary artificial intelligence (AI)-based approaches. The review was structured to cover the major stages of image processing, including image representation, enhancement, restoration, segmentation, feature extraction, compression, deep learning, Transformer-based methods, generative models, explainable artificial intelligence (XAI), and foundation models. Particular attention was given to the relationship between methodological characteristics, computational requirements, application domains, and practical deployment. The organization of the review follows the broad progression identified in the source paper, from classical image-processing operators toward CNNs, Transformers, generative models, and foundation-model approaches.

The review is narrative rather than a formal systematic review, because the available source manuscript does not report a registered protocol, complete database-search record, PRISMA flow diagram, or quantitative meta-analysis. Therefore, no numerical pooled effect size or statistical meta-analysis was performed.

2.2 Literature Sources

Relevant literature was considered from major scholarly information sources and publisher platforms, including IEEE Xplore, ScienceDirect, SpringerLink, PubMed, ACM Digital Library, Google Scholar, and relevant journals published by major academic publishers. The literature base was supplemented by authoritative books and conference proceedings where they provided foundational or technically important contributions.

The principal search concepts included “digital image processing,” “image enhancement,” “image restoration,” “image segmentation,” “feature extraction,” “CNN image processing,” “Vision Transformer,” “deep learning image analysis,” “explainable AI image processing,” “foundation models computer vision,” and “image processing agriculture.” The existing paper provides a substantial reference base covering image processing fundamental, deep learning, medical imaging, remote sensing, agriculture, explainable AI, transfer learning, and foundation models.

2.3 Inclusion and Exclusion Criteria

Literature was considered relevant when it addressed one or more major aspects of digital image processing or computer vision. The inclusion criteria comprised studies that:

  • described image-processing techniques or computational image analysis;
  • presented computer vision or pattern-recognition methodologies;
  • investigated CNN, transfer learning, Transformer, or related AI approaches;
  • examined practical image-processing applications;
  • introduced important methodological developments;
  • provided methodological, comparative, or application-related information; and
  • represented peer-reviewed articles, authoritative books, or recognized conference publications.

The exclusion criteria comprised duplicate publications, sources unrelated to image processing, non-technical opinion material, and publications that did not provide sufficient methodological information for meaningful classification or comparison. Duplicate versions of the same study were also excluded from thematic consideration.

2.4 Thematic Classification

The selected literature was organized into major methodological themes to facilitate comparison. These categories were derived from the technical structure of the reviewed paper and its reference literature.

Theme

Major Methods/Approaches

Image representation

RGB, HSV, YCbCr, grayscale

Enhancement

HE, CLAHE, contrast stretching, gamma correction

Restoration

Wiener, median, adaptive filtering

Segmentation

Thresholding, Otsu, K-Means, U-Net, SAM

Feature extraction

SIFT, SURF, ORB, HOG

Compression

JPEG, JPEG2000, PNG

Deep learning

CNN, ResNet, EfficientNet

Transformer

ViT, Swin Transformer

Explainability

Grad-CAM, XAI

Generative AI

GAN, diffusion models

Foundation models

SAM, multimodal/foundation systems

Deployment

Edge, mobile, IoT

This classification enabled the comparison of traditional handcrafted approaches with automatically learned visual representations and newer general-purpose vision architectures.

2.5 Comparative Evaluation Framework

The reviewed approaches were compared according to accuracy, precision, recall, F1-score, computational complexity, dataset requirements, interpretability, processing speed, memory requirements, application domain, and deployment suitability. These measures were treated as comparative criteria reported or discussed in the literature rather than as results generated from a new experimental dataset.

For classification studies, the principal evaluation measures are defined as:

Accuracy

Precision

Recall

F1-Score

where TP represents true positives, TN true negatives, FP false positives, and FN false negatives. These measures provide complementary information regarding classification performance and are commonly used when comparing image-analysis approaches.

For image-quality assessment, the source paper also identifies Mean Square Error (MSE), Peak Signal-to-Noise Ratio (PSNR), and Structural Similarity Index Measure (SSIM) as important evaluation measures (Gonzalez & Woods, 2024; Wang et al., 2004).

Methodological note: Because this is a review study, the comparative tables and graphs in the Results and Discussion section should clearly distinguish reported values extracted from published studies from qualitative synthesis or illustrative information. This avoids presenting hypothetical values as original experimental findings.

3. Results and Discussion

The thematic analysis of the reviewed literature demonstrates that digital image processing has developed from relatively simple pixel-level operations to complex AI-driven systems capable of learning visual representations, performing semantic segmentation, generating images, and supporting multimodal visual reasoning. The reviewed approaches can broadly be classified into image representation, enhancement and restoration, segmentation, feature extraction, compression, deep learning, Transformer-based processing, explainable AI, generative AI, and foundation-model approaches. The analysis also indicates that no single technique is universally suitable for all applications; rather, the selection of an image-processing method depends on image characteristics, computational resources, dataset availability, required accuracy, interpretability, and deployment environment.

3.1 Fundamental Image Representation

Digital images are composed of pixels arranged according to spatial dimensions such as width and height. Each pixel represents the visual information at a particular spatial location. The amount of information represented by an image depends on factors including spatial resolution, bit depth, and the number of color channels. Higher spatial resolution can preserve finer visual details but generally increases storage and computational requirements. Similarly, greater bit depth provides a larger range of intensity values and can preserve more subtle variations in image information (Gonzalez & Woods, 2024). The existing paper describes a digital image as a two-dimensional array of pixels and emphasizes the importance of resolution and color representation in image analysis.

Different color representations are useful for different image-processing tasks. RGB represents images through red, green, and blue channels and is widely used in cameras and displays. HSV separates hue, saturation, and value and can be advantageous for color-based segmentation. YCbCr separates luminance from chrominance and is widely associated with image and video compression. Grayscale represents an image through a single intensity channel, reducing computational and memory requirements.

Model

Main characteristic

Advantages

Applications

RGB

Three color channels

Simple and widely supported

Cameras, displays, image acquisition

HSV

Hue, saturation, and value

Separates color information from intensity

Color segmentation, object detection

YCbCr

Luminance and chrominance

Efficient representation for compression

JPEG, digital video

Grayscale

Single intensity channel

Low computational and memory requirements

Medical imaging, documents, edge detection

Table 1. Comparison of Image Representation Models

Thus, image representation is not merely a preliminary technical choice; it influences the effectiveness and computational requirements of subsequent enhancement, segmentation, feature extraction, and classification operations.

3.2 Image Enhancement and Restoration

Image enhancement improves the visual quality or visibility of important image characteristics, whereas image restoration attempts to recover information degraded by noise, blur, or other imaging processes. Common enhancement approaches include Histogram Equalization (HE), Contrast Limited Adaptive Histogram Equalization (CLAHE), contrast stretching, and gamma correction.

Histogram Equalization redistributes intensity values over the available dynamic range and can improve the contrast of low-contrast images. CLAHE extends this principle by operating on local regions and limiting excessive contrast amplification. Contrast stretching expands a restricted intensity range, while gamma correction applies a nonlinear transformation to adjust image brightness (Gonzalez & Woods, 2024). The existing Figure 2 demonstrates the progression from an original image through low contrast to an enhanced image and includes corresponding intensity histograms.

Image restoration addresses degradation using mathematical or statistical models. Gaussian noise can originate from electronic sensor and transmission effects, salt-and-pepper noise appears as isolated dark and bright pixels, and speckle noise is commonly associated with imaging systems such as radar and ultrasound. Wiener filtering is useful for restoration under suitable noise assumptions, median filtering is particularly useful for impulse noise, and adaptive filters can adjust their behavior according to local image characteristics.

Figure 2. Before and After Image Enhancement

The figure should contain the corresponding histograms to demonstrate the redistribution of pixel intensities.

3.3 Image Segmentation

Segmentation divides an image into meaningful regions and is fundamental to object identification and analysis. Conventional approaches include global thresholding, Otsu thresholding, region growing, watershed segmentation, and K-Means clustering. These methods differ in the assumptions they make about intensity, color, connectivity, or region similarity. Otsu's method automatically determines a threshold based on between-class variance, while K-Means groups pixels according to similarity.

Deep learning introduced more adaptive segmentation approaches, including methods reviewed specifically for agricultural image segmentation (Lei et al., 2024). U-Net, for example, uses an encoder-decoder architecture to combine contextual information with spatial localization and has become an important architecture for biomedical segmentation (Ronneberger et al., 2015). More recently, foundation-model approaches have expanded segmentation beyond individual task-specific models. Segment Anything (SAM) introduced a promptable segmentation framework designed to generalize across a wide range of image segmentation tasks (Kirillov et al., 2023).

Consequently, segmentation has evolved from manually designed thresholds and region rules toward learned and increasingly general-purpose visual representations.

Figure 3. Comparative Image Segmentation Workflow

This figure should visually compare thresholding/Otsu/K-Means with U-Net and SAM rather than presenting them as equivalent algorithms.

3.4 Feature Extraction

Feature extraction converts visual information into numerical representations that can be used for classification, recognition, matching, or analysis. Traditional feature descriptors include SIFT, SURF, ORB, and HOG, while image characteristics such as edges, shape, and texture can also serve as important descriptors. SIFT provides scale- and rotation-invariant local features, SURF emphasizes computational efficiency, ORB provides a comparatively lightweight alternative, and HOG represents local gradient and edge orientations.

Deep learning changed feature extraction by allowing features to be learned automatically from training data. CNNs progressively learn low-level, intermediate, and high-level representations, reducing the need for manually designed descriptors. Transformer architectures further extend representation learning by modelling contextual relationships between image regions.

Feature method

Feature type

Computational demand

Typical use

SIFT

Local, scale-invariant

Medium

Object matching

SURF

Local

Medium

Recognition and matching

ORB

Local

Low

Real-time vision

HOG

Shape/gradient

Low–Medium

Object detection

CNN

Learned hierarchical features

High

Image classification

Transformer

Global contextual representation

High–Very High

Image understanding

Table 2. Traditional and Deep Feature Extraction Approaches

Traditional descriptors therefore remain relevant where computational resources or training data are limited, while learned representations are particularly useful for complex visual patterns.

3.5 Image Compression

Image compression reduces storage requirements and transmission bandwidth. It can be classified into lossless and lossy compression. Lossless methods preserve the original information, whereas lossy methods remove information considered less important for visual representation.

JPEG is a widely used lossy compression standard associated with the Discrete Cosine Transform (DCT), while JPEG2000 uses wavelet-based processing. PNG provides lossless compression and is particularly useful when preservation of image information or transparency is required.

Compression efficiency can be expressed as:

Compression Ratio=Original Data SizeCompressed Data Size

Higher compression can reduce storage and transmission requirements, but excessive compression can introduce artifacts and reduce image fidelity. Therefore, image compression involves a practical balance between compression ratio, storage efficiency, transmission requirements, and image quality.

3.6 Deep Learning-Based Image Processing

Deep learning has substantially changed image processing by enabling automatic feature learning from large image datasets. CNNs use convolutional filters to extract spatial features, activation functions such as ReLU to introduce nonlinear transformations, pooling to reduce spatial dimensions, and fully connected or classification layers to produce predictions. The existing Figure 4 illustrates this general CNN workflow.

ResNet introduced residual connections that allow information to pass through shortcut pathways and facilitate the training of deeper networks (He et al., 2016). Efficient architectures subsequently focused on balancing predictive performance with computational requirements. EfficientNet introduced compound scaling across network depth, width, and input resolution (Tan & Le, 2019), while MobileNet architectures were designed specifically for resource-constrained environments. MobileNetV3 incorporated hardware-aware architecture search and optimization for mobile deployment (Howard et al., 2019).

Transfer learning is particularly relevant when task-specific datasets are limited. A pretrained model can provide reusable visual representations that can subsequently be adapted to a target application. The reviewed literature identifies transfer learning and optimized CNN training as important approaches for improving image-processing performance while reducing training requirements (Ankatwar, 2025; Yadav & Gajanan, 2026).

3.7 Vision Transformers and Explainable AI

Vision Transformers (ViTs) introduced a different approach to image representation by dividing an image into patches and processing those patches as a sequence. Self-attention enables relationships between image regions to be modelled, providing a mechanism for capturing broader contextual information (Dosovitskiy et al., 2021). Swin Transformer further developed hierarchical Transformer representations and has become an important architecture for visual recognition and dense prediction (Liu et al., 2022).

The increasing complexity of deep models has simultaneously increased interest in Explainable AI (XAI). Techniques such as Grad-CAM, saliency maps, feature visualization, and SHAP can assist researchers and users in understanding model predictions (Samek et al., 2021). Grad-CAM uses gradients associated with a target class to generate class-discriminative localization maps, allowing important image regions influencing a CNN prediction to be visualized (Selvaraju et al., 2017). XAI is especially relevant in sensitive domains such as medical diagnosis and agricultural disease detection, where understanding the visual evidence supporting a prediction can complement conventional performance metrics. Recent plant-disease XAI literature also emphasizes the importance of evaluating explanation methods and their validation practices (Ismail et al., 2026).

Figure 4. Explainable AI Workflow

The figure 4 illustrates the workflow of explainable AI, showing how an input image is processed by a deep learning model, followed by prediction, gradient analysis, generation of visual explanations such as Grad-CAM heatmaps, and human interpretation of the model’s decision.

3.8 Generative AI and Foundation Models

Generative AI has expanded image processing beyond analysis toward image generation and transformation. GANs employ generator and discriminator networks and can be applied to synthetic-image generation, augmentation, and image-to-image translation. Diffusion models use a progressive noise-removal process to generate images and have become an important generative modelling paradigm (Ho et al., 2020).

Foundation models represent another major transition. Rather than developing an independent model for every individual task, foundation models are pretrained on large-scale data and can potentially be adapted to multiple downstream applications. The foundation-model literature identifies both opportunities and challenges associated with this approach, including generalization, computational requirements, and responsible deployment (Bommasani et al., 2022). SAM provides an important computer-vision example through promptable segmentation (Kirillov et al., 2023).

3.9 Applications of Image Processing

The reviewed literature demonstrates broad applicability across scientific and industrial domains.

In medical imaging, image enhancement, segmentation, and classification support MRI, CT, and X-ray analysis (Litjens et al., 2017). In remote sensing, image processing supports land-use mapping, environmental monitoring, disaster assessment, and infrastructure analysis.

In agriculture, image processing supports crop monitoring, disease detection, pest identification, and precision farming. CNNs, transfer learning, multispectral imaging, XAI, and deep-learning-based segmentation methods (Lei et al., 2024) are increasingly being investigated for crop-health assessment, including cotton disease detection (Ankatwar & Dhawale, 2024, 2025, 2026; Gajanan & Preetham, 2026).

Other important applications include autonomous vehicles, where image processing assists in detecting pedestrians, vehicles, lanes, and traffic signs; security, where visual recognition supports biometric authentication and surveillance; and industrial inspection, where computer vision identifies manufacturing defects and quality problems.

3.10 Comparative Analysis

The overall literature indicates a transition from computationally simple, explicitly designed algorithms toward data-driven models with increasing representational capacity. However, the relationship between accuracy, computational cost, data requirements, explainability, and deployment must be considered when selecting a method.

Method

Learning type

Accuracy potential*

Complexity*

Data requirement*

Explainability*

Deployment suitability*

Histogram Equalization

Classical

Moderate

Low

Low

High

Excellent

Otsu

Classical

Moderate

Low

Low

High

Excellent

K-Means

Unsupervised

Moderate

Medium

Low–Medium

Medium–High

Good

U-Net

Deep learning

High

High

High

Moderate

Moderate

CNN

Deep learning

High

High

High

Moderate

Moderate

EfficientNet

Deep learning

High

Medium–High

High

Moderate

Good

MobileNet

Deep learning

High

Low–Medium

High

Moderate

Excellent

ViT

Transformer

High

Very High

High

Low–Moderate

Moderate

Swin Transformer

Transformer

High

High

High

Low–Moderate

Moderate

SAM

Foundation model

High potential

Very High

Very high pretraining requirement

Moderate

Hardware dependent

Table 3. Comparative Analysis of Major Image Processing Approaches

*Qualitative synthesis of the reviewed literature, not original experimental measurements.

The comparison shows that classical methods generally require fewer computational resources and can be easier to interpret. Their effectiveness, however, may depend strongly on image quality, illumination, noise, and background complexity. Deep learning methods can learn more complex representations but normally require greater computational resources and suitable training data. The current paper similarly identifies conventional methods as valuable for preprocessing and resource-constrained environments while highlighting the increasing importance of deep learning, transfer learning, and optimized architectures.

3.11 Performance Visualization and Graphs

The existing manuscript contains a performance graph comparing thresholding, Otsu, K-Means, region growing, watershed, and U-Net using accuracy, precision, recall, and F1-score. However, the manuscript explicitly identifies the numerical values in this graph as illustrative and states that they should be replaced with quantitative results obtained from the reviewed studies.

Figure 5. Qualitative Computational Complexity of Image Processing Techniques

The figure presents a qualitative comparison of computational complexity across classical image-processing methods, feature extraction techniques, CNNs, Transformers, and foundation models, ranging from low to very high computational demand.

Figure 6. Evolution of Image Processing Technologies

The figure illustrates the progression of image processing from classical techniques and machine learning to CNNs, transfer learning, Vision Transformers, generative AI, and foundation models, highlighting the increasing capabilities and broader applications of modern visual-processing systems.

Figure 7. Major Application Domains of Image Processing

The figure 7 illustrates the major application areas of image processing, including healthcare, agriculture, remote sensing, industrial inspection, security and surveillance, autonomous transportation, multimedia, environmental monitoring, and biometrics.

3.12 Challenges

Several challenges remain despite the rapid development of image-processing technologies. Deep learning systems frequently depend on sufficiently large and representative labelled datasets. Dataset imbalance, annotation cost, limited diversity, and domain shift can influence model generalization. These issues become particularly important when models are transferred from laboratory or benchmark datasets to real-world environments.

Computational requirements are another challenge. Large CNNs, Transformers, and foundation models may require substantial GPU memory and processing capacity, making lightweight CNNs and AI edge devices important considerations for practical deployment (Sun et al., 2024). Such requirements can restrict deployment on smartphones, embedded devices, agricultural sensors, and other edge platforms.

Interpretability also remains important. High predictive performance does not automatically explain why a model produces a particular prediction. Consequently, XAI methods are increasingly incorporated into visual analysis systems. Privacy and security are particularly relevant for medical and surveillance applications, while unreliable network connectivity can limit the practicality of cloud-dependent systems.

3.13 Future Research Directions

The literature indicates several important research directions, including lightweight CNNs, EfficientNet and MobileNet architectures, edge AI, explainable AI, federated learning, multispectral image processing, multimodal vision, vision-language models (Zhang et al., 2024), foundation models, generative data augmentation, diffusion models, real-time image processing, and privacy-preserving computer vision. The existing review similarly identifies lightweight architectures, edge computing, federated learning, foundation models, and multimodal systems as emerging areas for future development.

Future systems are therefore likely to move toward hybrid architectures that combine classical preprocessing with efficient deep learning and Transformer-based representations. In agriculture, for example, image enhancement and segmentation can be combined with lightweight classification models and XAI for deployment on edge devices. In healthcare, segmentation and classification can be integrated with explainability mechanisms to provide visual evidence alongside predictions. Vision-language and other multimodal foundation-model approaches may further reduce the need to construct independent models for every visual task (Zhang et al., 2024), although their computational and reliability requirements require continued investigation.

Figure 8. Future Research Framework for Intelligent Image Processing

The framework illustrates the integration of image acquisition, preprocessing, efficient AI models, explainable AI, edge deployment, multimodal learning, foundation models, and real-time intelligent decision support for future image-processing systems.

CONCLUSION

Digital image processing has undergone a substantial transformation from conventional pixel-level operations to sophisticated artificial intelligence-driven systems capable of automated visual analysis and semantic understanding. Classical approaches initially focused on image enhancement, restoration, filtering, segmentation, feature extraction, and compression, while subsequent developments in machine learning and deep learning enabled increasingly automated interpretation of visual information. The reviewed literature demonstrates that this evolution has expanded image processing across medical imaging, agriculture, remote sensing, autonomous vehicles, security surveillance, and industrial inspection.

Despite the rapid development of AI-based approaches, classical image-processing techniques remain relevant. Histogram equalization, CLAHE, filtering, thresholding, handcrafted feature extraction, and compression methods provide relatively simple and computationally efficient solutions for many image-processing tasks. Their lower computational requirements and comparatively transparent processing mechanisms make them particularly useful for preprocessing and applications with constrained computational resources. Consequently, conventional techniques should not be regarded solely as outdated alternatives but as important components of contemporary image-processing pipelines.

The development of deep learning has substantially expanded automated image classification, detection, and segmentation. CNNs can learn hierarchical visual representations directly from image data, while architectures such as ResNet and U-Net have addressed important challenges in deep feature learning and image segmentation. Efficient architectures and transfer-learning approaches further provide mechanisms for adapting pretrained visual representations and reducing computational requirements. The reviewed literature indicates that optimized CNNs and transfer learning are particularly relevant for applications where labelled datasets or computational resources may be limited.

More recently, Vision Transformers, Swin Transformers, Segment Anything, generative models, and foundation-model approaches have extended the capabilities of image analysis toward global contextual representation, large-scale pretraining, promptable segmentation, image generation, and broader task transfer. These developments indicate a transition from highly task-specific image-processing systems toward more general visual computing frameworks.

Future image-processing research should therefore evaluate more than predictive accuracy. Explainability, computational efficiency, robustness, privacy, memory requirements, inference latency, dataset quality, and real-world deployment are important considerations for developing reliable visual systems. This is particularly relevant for edge devices, mobile platforms, healthcare applications, agricultural monitoring, and other environments where computational and operational constraints are significant. The increasing role of XAI, lightweight architectures, edge computing, federated learning, and multimodal vision supports this broader evaluation perspective.

Overall, the future of image processing is likely to involve hybrid and integrated approaches combining classical preprocessing, efficient deep learning, Transformer architectures, explainable AI, generative models, edge computing, and foundation models. Such integration can provide a balance between computational efficiency, visual representation, interpretability, scalability, and application-specific requirements. Continued research should focus on developing robust and efficient systems that can operate across diverse imaging conditions while maintaining appropriate levels of transparency, privacy, and practical usability. Thus, the progression from classical image processing to intelligent visual systems represents not the disappearance of traditional methods, but their integration with increasingly capable AI technologies to address the growing complexity of real-world image-analysis problems.

DECLARATION OF AI USE

This manuscript was prepared through the intellectual contributions of the author(s) to the study design, data analysis, interpretation of results, and scholarly content. Grammarly and ChatGPT were used solely to assist with grammatical refinement and reference formatting during manuscript revision. The author(s) accept full responsibility for the accuracy, integrity, and final content of the manuscript.

REFERENCES

  1. Ankatwar, G. (2025). A comprehensive study of CNN training techniques for image processing. The Academic, 3(8), 705–711. https://doi.org/10.5281/zenodo.17129739
  2. Ankatwar, G., & Dhawale, C. (2024, November). Enhancing crop disease detection systems with explainable AI techniques for deep learning models using spectral imaging. In 2024 2nd International Conference on Emerging Trends in Engineering and Medical Sciences (ICETEMS) (pp. 312–319). IEEE.
  3. Ankatwar, G., & Dhawale, C. (2025). Multispectral imaging and CNN architectures for cotton leaf disease classification: A comprehensive review. International Journal for Multidisciplinary Research, 7(3).
  4. Ankatwar, G., & Dhawale, C. (2026a, January). Automated cotton leaf disease detection: A step towards smart and sustainable agriculture. In 2026 International Conference on Intelligent and Innovative Technologies in Computing, Electrical and Electronics (IITCEE) (pp. 1–6). IEEE.
  5. Ankatwar, G., & Dhawale, C. (2026b, January). Enhancing cotton crop health monitoring through transfer learning and image preprocessing. In 2026 International Conference on Intelligent and Innovative Technologies in Computing, Electrical and Electronics (IITCEE) (pp. 1–6). IEEE.
  6. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., ... Liang, P. (2022). On the opportunities and risks of foundation models. Stanford Center for Research on Foundation Models.
  7. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An image is worth 16 × 16 words: Transformers for image recognition at scale. International Conference on Learning Representations.
  8. Gajanan, A., & Preetham, N. (2026). Deep learning for smart agriculture: A comprehensive review of CNN architectures, multispectral imaging, explainable AI and transfer learning for crop disease detection. Asian Journal of Research in Computer Science, 19(3), 128–143.
  9. Gajanan, A., Preetham, N., Rani, G. L., Venkatreddy, A. T., & Yadav, K. K. (2026). Image processing from handcrafted operators to foundation models: A critical review of methods, applications and recent advances. Advances in Research, 27(5), 67–90.
  10. Gonzalez, R. C., & Woods, R. E. (2024). Digital image processing (5th ed.). Pearson.
  11. Goodfellow, I., Bengio, Y., & Courville, A. (2024). Deep learning (Updated ed.). MIT Press.
  12. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 770–778).
  13. Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33, 6840–6851.
  14. Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudev, V., Le, Q. V., & Adam, H. (2019). Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 1314–1324).
  15. Ismail, U. I., Chua, H. N., Nordin, R., & Lee, H. L. (2026). Explainable artificial intelligence for plant disease detection: A scoping review of methods, datasets, and evaluation practices. Computers and Electronics in Agriculture. https://doi.org/10.1016/j.compag.2026.112338
  16. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 4015–4026).
  17. Lei, L., Yang, Q., Yang, L., Shen, T., Wang, R., et al. (2024). Deep learning implementation of image segmentation in agricultural applications: A comprehensive review. Artificial Intelligence Review, 57, 149.
  18. Litjens, G., Kooi, T., Bejnordi, B. E., Setio, A. A. A., Ciompi, F., Ghafoorian, M., van der Laak, J. A. W. M., van Ginneken, B., & Sánchez, C. I. (2017). A survey on deep learning in medical image analysis. Medical Image Analysis, 42, 60–88.
  19. Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022). Swin Transformer V2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 12009–12019).
  20. Long, Y., Gong, Y., Xiao, Z., & Liu, Q. (2023). Remote sensing image interpretation using deep learning: A comprehensive review. ISPRS Journal of Photogrammetry and Remote Sensing, 198, 347–371.
  21. Preetham, N., Venkatreddy, A. T., Yadav, K. K., Gajanan, A., & Rani, G. L. (2026). Multimodal emotion recognition using explainable hybrid CNN-Transformer networks. Asian Journal of Research in Computer Science, 19(7), 12–30.
  22. Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention (MICCAI) (pp. 234–241).
  23. Samek, W., Montavon, G., Vedaldi, A., Hansen, L. K., & Müller, K.-R. (Eds.). (2021). Explainable AI: Interpreting, explaining and visualizing deep learning. Springer.
  24. Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision (pp. 618–626).
  25. Sun, K., Wang, X., Miao, X., & Zhao, Q. (2024). A review of AI edge devices and lightweight CNN and LLM deployment. Neurocomputing.
  26. Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning, 97, 6105–6114.
  27. Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4), 600–612. https://doi.org/10.1109/TIP.2003.819861
  28. Yadav, K. K., & Gajanan, A. (2026). Transfer learning for image processing: A review and practical considerations. Asian Journal of Research in Computer Science, 19(2), 192–203.
  29. Zhang, J., Huang, J., Jin, S., & Lu, S. (2024). Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8), 5625–5644. https://doi.org/10.1109/TPAMI.2024.3369699
  30. Zhao, Z.-Q., Zheng, P., Xu, S.-T., & Wu, X. (2019). Object detection with deep learning: A review. IEEE Transactions on Neural Networks and Learning Systems, 30(11), 3212–3232. https://doi.org/10.1109/TNNLS.2018.2876865

Reference

  1. Ankatwar, G. (2025). A comprehensive study of CNN training techniques for image processing. The Academic, 3(8), 705–711. https://doi.org/10.5281/zenodo.17129739
  2. Ankatwar, G., & Dhawale, C. (2024, November). Enhancing crop disease detection systems with explainable AI techniques for deep learning models using spectral imaging. In 2024 2nd International Conference on Emerging Trends in Engineering and Medical Sciences (ICETEMS) (pp. 312–319). IEEE.
  3. Ankatwar, G., & Dhawale, C. (2025). Multispectral imaging and CNN architectures for cotton leaf disease classification: A comprehensive review. International Journal for Multidisciplinary Research, 7(3).
  4. Ankatwar, G., & Dhawale, C. (2026a, January). Automated cotton leaf disease detection: A step towards smart and sustainable agriculture. In 2026 International Conference on Intelligent and Innovative Technologies in Computing, Electrical and Electronics (IITCEE) (pp. 1–6). IEEE.
  5. Ankatwar, G., & Dhawale, C. (2026b, January). Enhancing cotton crop health monitoring through transfer learning and image preprocessing. In 2026 International Conference on Intelligent and Innovative Technologies in Computing, Electrical and Electronics (IITCEE) (pp. 1–6). IEEE.
  6. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., ... Liang, P. (2022). On the opportunities and risks of foundation models. Stanford Center for Research on Foundation Models.
  7. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An image is worth 16 × 16 words: Transformers for image recognition at scale. International Conference on Learning Representations.
  8. Gajanan, A., & Preetham, N. (2026). Deep learning for smart agriculture: A comprehensive review of CNN architectures, multispectral imaging, explainable AI and transfer learning for crop disease detection. Asian Journal of Research in Computer Science, 19(3), 128–143.
  9. Gajanan, A., Preetham, N., Rani, G. L., Venkatreddy, A. T., & Yadav, K. K. (2026). Image processing from handcrafted operators to foundation models: A critical review of methods, applications and recent advances. Advances in Research, 27(5), 67–90.
  10. Gonzalez, R. C., & Woods, R. E. (2024). Digital image processing (5th ed.). Pearson.
  11. Goodfellow, I., Bengio, Y., & Courville, A. (2024). Deep learning (Updated ed.). MIT Press.
  12. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 770–778).
  13. Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33, 6840–6851.
  14. Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudev, V., Le, Q. V., & Adam, H. (2019). Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 1314–1324).
  15. Ismail, U. I., Chua, H. N., Nordin, R., & Lee, H. L. (2026). Explainable artificial intelligence for plant disease detection: A scoping review of methods, datasets, and evaluation practices. Computers and Electronics in Agriculture. https://doi.org/10.1016/j.compag.2026.112338
  16. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollár, P., & Girshick, R. (2023). Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 4015–4026).
  17. Lei, L., Yang, Q., Yang, L., Shen, T., Wang, R., et al. (2024). Deep learning implementation of image segmentation in agricultural applications: A comprehensive review. Artificial Intelligence Review, 57, 149.
  18. Litjens, G., Kooi, T., Bejnordi, B. E., Setio, A. A. A., Ciompi, F., Ghafoorian, M., van der Laak, J. A. W. M., van Ginneken, B., & Sánchez, C. I. (2017). A survey on deep learning in medical image analysis. Medical Image Analysis, 42, 60–88.
  19. Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022). Swin Transformer V2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 12009–12019).
  20. Long, Y., Gong, Y., Xiao, Z., & Liu, Q. (2023). Remote sensing image interpretation using deep learning: A comprehensive review. ISPRS Journal of Photogrammetry and Remote Sensing, 198, 347–371.
  21. Preetham, N., Venkatreddy, A. T., Yadav, K. K., Gajanan, A., & Rani, G. L. (2026). Multimodal emotion recognition using explainable hybrid CNN-Transformer networks. Asian Journal of Research in Computer Science, 19(7), 12–30.
  22. Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention (MICCAI) (pp. 234–241).
  23. Samek, W., Montavon, G., Vedaldi, A., Hansen, L. K., & Müller, K.-R. (Eds.). (2021). Explainable AI: Interpreting, explaining and visualizing deep learning. Springer.
  24. Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision (pp. 618–626).
  25. Sun, K., Wang, X., Miao, X., & Zhao, Q. (2024). A review of AI edge devices and lightweight CNN and LLM deployment. Neurocomputing.
  26. Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning, 97, 6105–6114.
  27. Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: From error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4), 600–612. https://doi.org/10.1109/TIP.2003.819861
  28. Yadav, K. K., & Gajanan, A. (2026). Transfer learning for image processing: A review and practical considerations. Asian Journal of Research in Computer Science, 19(2), 192–203.
  29. Zhang, J., Huang, J., Jin, S., & Lu, S. (2024). Vision-language models for vision tasks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8), 5625–5644. https://doi.org/10.1109/TPAMI.2024.3369699
  30. Zhao, Z.-Q., Zheng, P., Xu, S.-T., & Wu, X. (2019). Object detection with deep learning: A review. IEEE Transactions on Neural Networks and Learning Systems, 30(11), 3212–3232. https://doi.org/10.1109/TNNLS.2018.2876865

Photo
Gajanan Ankatwar
Corresponding author

Government Degree College (Arts and Commerce), Adilabad, Telangana, India.

Photo
Preetham Narote
Co-author

Telangana Tribal Welfare Residential Degree College (Boys), Boath@Adilabad, Telangana, India.

Photo
Saroja Rachakonda
Co-author

TGTWRDC GIRLS UTNOOR, Utnoor, Adilabad, Telangana, India.

Photo
Krunal Yadav K.
Co-author

Government Degree College (Arts and Commerce), Adilabad, Telangana, India.

Preetham Narote1, Saroja Rachakonda2, Krunal Yadav K.3, Gajanan Ankatwar3*, A Comprehensive Review Of Image Processing Techniques: Classical Methods, Deep Learning, Applications, And Recent Advances, Int. J. Sci. R. Tech., 2026, 3 (9), 775-798. https://doi.org/10.5281/zenodo.23120059

Related Articles
AIVERSE: A Unified AI-Powered Image Intelligence Platform Using Deep Learning, C...
Anurag Dhondge, Jayraj Patil, Sushant Karle, Rudresh Kankrej...
Image Classification System...
Shivam Kumar, Shraddha Kashid , Aditya Pardeshi, Enrique Anthony...