Image ProcessingUnit 15 min read

Image Processing Basics: Fundamentals, Types, and Applications

Unit 1 of Image Processing introduces core concepts like digital image formation, types of images (grayscale, color, binary), image processing systems, and applications in real-world domains such as medical imaging, satellite remote sensing, and computer vision. This note covers definitions, workflows, and comparisons


1. What is an Image?

An image is a 2D function where and are spatial coordinates, and the amplitude represents the intensity (e.g., brightness or color) at each point (pixel). Images can be classified based on their dimensionality and color representation:

1.1 Types of Images

classDiagram
    class Image {
        <<abstract>>
        +dimensions: int
        +colorModel: string
    }
    class GrayscaleImage {
        +intensity: 0-255
    }
    class ColorImage {
        +model: RGB/CMYK/HSV
    }
    class BinaryImage {
        +threshold: 0 or 1
    }
  • Grayscale Image: Single-channel (0–255), e.g., X-ray scans.
  • Color Image: Multi-channel (RGB, CMYK), e.g., photos from smartphones.
  • Binary Image: Black-and-white (0 or 1), e.g., OCR (Optical Character Recognition) outputs.

2. Digital Image Formation

  • Resolution: Number of pixels (e.g., 1920×1080 for Full HD).
  • Bit Depth: Bits per pixel (e.g., 8-bit = 256 grayscale levels, 24-bit = 16.7M colors).
  • Dynamic Range: Ratio of brightest to darkest detectable intensity.

Worked Example: Pixel Calculation A 4K image has dimensions 3840×2160. Calculate:

  1. Total pixels = pixels.
  2. File size (8-bit grayscale) = byte = 8.29 MB.
  3. File size (24-bit RGB) = bytes = 24.88 MB.

3. Image Processing Systems

An image processing system consists of:

  1. Input: Sensor (camera, scanner) or synthetic data (3D rendering).
  2. Processing: Algorithms (enhancement, segmentation, recognition).
  3. Output: Display, storage, or further analysis.
flowchart LR
    A["Input: Camera/Scanner"] --> B["Preprocessing: Noise Reduction"]
    B --> C["Enhancement: Contrast Adjustment"]
    C --> D["Segmentation: Object Extraction"]
    D --> E["Recognition: Classification"]
    E --> F["Output: Display/Storage"]

In the Real World

  • eSewa: Uses image processing to verify fingerprint scans (binary segmentation) for user authentication.
  • Pathao: Applies edge detection (Canny filter) to analyze driver routes from satellite images for optimal delivery paths.
  • NTC Traffic Cameras: Use motion detection (frame differencing) to identify traffic violations in real time.

4. Applications of Image Processing

Domain Application Key Technique
Medical Tumor detection in MRI scans Segmentation + Machine Learning
Satellite Land cover classification Supervised Learning (SVM, CNN)
Security Face recognition (e.g., Ncell SIM cards) Eigenfaces + PCA
Automotive Lane detection (self-driving cars) Hough Transform + Edge Detection
Retail Barcode scanning (Daraz, MegaMart) OCR (Optical Character Recognition)

5. Workflow of Image Processing

A typical pipeline:

  1. Acquisition: Capture image (e.g., via webcam or medical scanner).
  2. Preprocessing: Correct distortions (e.g., noise removal).
  3. Enhancement: Improve visual quality (e.g., histogram equalization).
  4. Segmentation: Partition image into regions (e.g., foreground/background).
  5. Description/Recognition: Extract features (e.g., SIFT for object matching).

Worked Example: Traffic Sign Recognition (Nepal)

  1. Input: Blurry traffic sign photo from a dashboard camera.
  2. Preprocessing: Apply Gaussian blur to reduce noise.
  3. Enhancement: Sharpen using unsharp masking.
  4. Segmentation: Use Otsu’s thresholding to binarize the sign.
  5. Recognition: Train a CNN to classify signs (e.g., "Stop" vs. "Speed Limit").
graph LR
    A["Input: Blurry Sign"] --> B["Gaussian Blur"]
    B --> C["Unsharp Masking"]
    C --> D["Otsu Thresholding"]
    D --> E["CNN Classification"]
    E --> F["Output: 'Stop Sign'"]

6. Challenges in Image Processing

Challenge Cause Solution
Noise (Gaussian/Salt&Pepper) Low-light conditions, sensor errors Median filter, Wiener deconvolution
Low Resolution Compression artifacts Super-resolution techniques
Illumination Variations Changing light conditions Histogram equalization
Occlusion Objects partially hidden Multi-view stereo reconstruction

Exam Tip

  1. Definitions: Know the difference between grayscale, color, and binary images (e.g., "Binary images use a threshold to separate foreground/background").
  2. Calculations: Practice pixel count and file size problems (e.g., "A 1024×768 RGB image has 24-bit depth. Calculate its size in MB").
  3. Applications: Link techniques to real-world examples:
    • eSewa: Binary segmentation for fingerprint verification.
    • Pathao: Edge detection for route optimization.
  4. Diagrams: Draw the image processing pipeline (input → preprocessing → enhancement → segmentation → recognition → output) in exams.
  5. Common Pitfalls:
    • Confusing resolution (pixels) with bit depth (color levels).
    • Forgetting that color images require 3 channels (RGB), not 1.

Key Formulae to Memorize

  1. Total pixels = .
  2. File size (bytes) = .
  3. Dynamic range = .

Based on the TU BIT syllabus for Image Processing, unit 1.

Discussion

Loading…