Image ProcessingUnit 66 min read

Image Representation, Description & Recognition: Features, Models & Matching

Unit 6 of Image Processing covers how to mathematically represent images (pixels, transforms), extract meaningful features (edges, textures, histograms), describe objects (shape descriptors, signatures), and recognize patterns (template matching, neural networks). Includes real-world applications in OCR, facial recogni

Key Concepts and Definitions

Image Representation

  • Spatial domain: Direct pixel values (e.g., ).
  • Frequency domain: Transformed using Fourier, Wavelet, or DCT (used in JPEG compression).

Worked Example: Pixel Representation

Consider a 2×2 grayscale image:

[100, 150]
[200, 50]
  • Spatial representation: , , etc.
  • Frequency domain: Apply 2D Discrete Cosine Transform (DCT) to decompose into frequency components.
graph LR
    A["Original Image (Spatial)"] -->|"DCT"| B["Frequency Domain"]
    B --> C["Low Freq (DC)"] & D["High Freq (AC)"]
    C -->|"JPEG Compression"| E["Quantized"]

Feature Extraction

Features are distinctive attributes extracted to describe an image or object. Common features:

  1. Edges: Detected using Sobel, Prewitt, or Canny operators.
  2. Corners: Harris corner detector.
  3. Textures: Gray-Level Co-occurrence Matrix (GLCM).
  4. Color Histograms: Distribution of pixel intensities.

Worked Example: Edge Detection (Sobel Operator)

For a 3×3 grayscale patch:

[-1, 0, 1]
[-2, 0, 2]
[-1, 0, 1]

Apply Sobel kernels:

  • Horizontal edge (Gx):
  • Vertical edge (Gy):
  • Magnitude:

Image Description

Describing images involves quantifying features into a compact form for comparison or recognition.

Shape Descriptors

  1. Moment Invariants: Hu moments (7 invariants to scale, rotation, and translation).
  2. Fourier Descriptors: Represent contours using Fourier series.
  3. Chain Codes: Directional encoding of object boundaries.

Worked Example: Hu Moments

For a binary image, compute central moments : Then compute Hu moments (normalized invariants). For example, the first Hu moment:

Color Descriptors

  • HSV/HSL: Separate hue, saturation, and value for robust color matching.
  • Dominant Color: K-means clustering to find primary colors in an image.

Image Recognition

Recognition involves matching extracted features to known patterns or classes.

Template Matching

Compare a template image with a sub-image using: Limitation: Works only for rigid objects with no deformation.

Structural Matching

Use graphs or trees to represent object parts and their relationships (e.g., syntactic pattern recognition).

Statistical Methods

  • k-Nearest Neighbors (k-NN): Classify based on feature similarity.
  • Support Vector Machines (SVM): Find optimal hyperplanes for classification.

Neural Networks

Convolutional Neural Networks (CNNs) are state-of-the-art for image recognition:

  1. Convolutional Layers: Extract features using filters.
  2. Pooling Layers: Downsample to reduce dimensionality.
  3. Fully Connected Layers: Classify based on learned features.
graph TD
    A["Input Image"] --> B["Convolutional Layer"]
    B --> C["ReLU Activation"]
    C --> D["Pooling Layer"]
    D --> E["Fully Connected"]
    E --> F["Softmax Output"]

Worked Example: CNN for Handwritten Digit Recognition (MNIST)

  1. Input: 28×28 grayscale image.
  2. Conv Layer 1: 32 filters (3×3), stride 1 → 28×28×32.
  3. Pooling: Max-pooling (2×2) → 14×14×32.
  4. Conv Layer 2: 64 filters → 14×14×64.
  5. Flatten: 14×14×64 → 12544.
  6. Fully Connected: 128 neurons → Softmax (10 classes).

Applications in the Real World

1. eSewa (Nepal)

  • Feature Extraction: Uses edge detection (Canny) to read QR codes on receipts for digital payments.
  • Recognition: Template matching to verify QR codes against the database.

2. Pathao (Ride-Hailing App)

  • Object Detection: CNNs detect pedestrians, vehicles, and traffic signs in real-time using camera feeds from driver phones.
  • Feature Matching: Match detected objects to pre-trained models for navigation and safety alerts.

3. Nepal Police Facial Recognition

  • Description: Uses Local Binary Patterns (LBP) to describe facial textures.
  • Recognition: k-NN or SVM to match faces in crowds against a database of criminals.

4. Daraz (E-Commerce)

  • Image Search: Extracts color histograms and SIFT features to match product images in the catalog.
  • OCR: Template matching for reading text on product labels.

5. Medical Imaging (e.g., Kathmandu’s Patan Hospital)

  • Segmentation + Description: Watershed algorithm segments tumors; Hu moments describe their shape for diagnosis.
  • Recognition: CNNs classify X-rays as normal/abnormal (e.g., pneumonia detection).

Comparison Table: Feature Extraction Methods

Method Description Pros Cons Applications
Sobel/Canny Edge detection using gradients Fast, simple Sensitive to noise Object detection
Harris Corner Detects corners using intensity Robust to rotation Computationally expensive Feature matching
SIFT/SURF Scale-invariant feature transform Invariant to scale/rotation Slow for real-time Image stitching
GLCM Texture analysis using co-occurrence Captures spatial relationships High dimensionality Medical imaging
Hu Moments Shape descriptors using moments Rotation/scale invariant Sensitive to noise Object recognition

Exam Tip

  1. Understand the Math: Know how to compute Sobel, Hu moments, and DCT manually for small images.
  2. Diagrams: Always draw feature extraction pipelines (e.g., CNN layers) in exams.
  3. Real-World Links: Relate questions to apps like eSewa (QR recognition) or Pathao (object detection).
  4. Pros/Cons: Compare methods (e.g., SIFT vs. Harris corners) in tables.
  5. Code Snippets: Be ready to write pseudocode for template matching or k-NN classification.
  6. Visuals: Sketch edge-detected images or CNN architectures when asked about recognition.

Based on the TU BIT syllabus for Image Processing, unit 6.

Discussion

Loading…