Image ProcessingUnit 66 min read
Image Representation, Description & Recognition: Features, Models & Matching
Unit 6 of Image Processing covers how to mathematically represent images (pixels, transforms), extract meaningful features (edges, textures, histograms), describe objects (shape descriptors, signatures), and recognize patterns (template matching, neural networks). Includes real-world applications in OCR, facial recogni
Key Concepts and Definitions
Image Representation
- Spatial domain: Direct pixel values (e.g., ).
- Frequency domain: Transformed using Fourier, Wavelet, or DCT (used in JPEG compression).
Worked Example: Pixel Representation
Consider a 2×2 grayscale image:
[100, 150]
[200, 50]
- Spatial representation: , , etc.
- Frequency domain: Apply 2D Discrete Cosine Transform (DCT) to decompose into frequency components.
graph LR
A["Original Image (Spatial)"] -->|"DCT"| B["Frequency Domain"]
B --> C["Low Freq (DC)"] & D["High Freq (AC)"]
C -->|"JPEG Compression"| E["Quantized"]Feature Extraction
Features are distinctive attributes extracted to describe an image or object. Common features:
- Edges: Detected using Sobel, Prewitt, or Canny operators.
- Corners: Harris corner detector.
- Textures: Gray-Level Co-occurrence Matrix (GLCM).
- Color Histograms: Distribution of pixel intensities.
Worked Example: Edge Detection (Sobel Operator)
For a 3×3 grayscale patch:
[-1, 0, 1]
[-2, 0, 2]
[-1, 0, 1]
Apply Sobel kernels:
- Horizontal edge (Gx):
- Vertical edge (Gy):
- Magnitude:
Image Description
Describing images involves quantifying features into a compact form for comparison or recognition.
Shape Descriptors
- Moment Invariants: Hu moments (7 invariants to scale, rotation, and translation).
- Fourier Descriptors: Represent contours using Fourier series.
- Chain Codes: Directional encoding of object boundaries.
Worked Example: Hu Moments
For a binary image, compute central moments : Then compute Hu moments (normalized invariants). For example, the first Hu moment:
Color Descriptors
- HSV/HSL: Separate hue, saturation, and value for robust color matching.
- Dominant Color: K-means clustering to find primary colors in an image.
Image Recognition
Recognition involves matching extracted features to known patterns or classes.
Template Matching
Compare a template image with a sub-image using: Limitation: Works only for rigid objects with no deformation.
Structural Matching
Use graphs or trees to represent object parts and their relationships (e.g., syntactic pattern recognition).
Statistical Methods
- k-Nearest Neighbors (k-NN): Classify based on feature similarity.
- Support Vector Machines (SVM): Find optimal hyperplanes for classification.
Neural Networks
Convolutional Neural Networks (CNNs) are state-of-the-art for image recognition:
- Convolutional Layers: Extract features using filters.
- Pooling Layers: Downsample to reduce dimensionality.
- Fully Connected Layers: Classify based on learned features.
graph TD
A["Input Image"] --> B["Convolutional Layer"]
B --> C["ReLU Activation"]
C --> D["Pooling Layer"]
D --> E["Fully Connected"]
E --> F["Softmax Output"]Worked Example: CNN for Handwritten Digit Recognition (MNIST)
- Input: 28×28 grayscale image.
- Conv Layer 1: 32 filters (3×3), stride 1 → 28×28×32.
- Pooling: Max-pooling (2×2) → 14×14×32.
- Conv Layer 2: 64 filters → 14×14×64.
- Flatten: 14×14×64 → 12544.
- Fully Connected: 128 neurons → Softmax (10 classes).
Applications in the Real World
1. eSewa (Nepal)
- Feature Extraction: Uses edge detection (Canny) to read QR codes on receipts for digital payments.
- Recognition: Template matching to verify QR codes against the database.
2. Pathao (Ride-Hailing App)
- Object Detection: CNNs detect pedestrians, vehicles, and traffic signs in real-time using camera feeds from driver phones.
- Feature Matching: Match detected objects to pre-trained models for navigation and safety alerts.
3. Nepal Police Facial Recognition
- Description: Uses Local Binary Patterns (LBP) to describe facial textures.
- Recognition: k-NN or SVM to match faces in crowds against a database of criminals.
4. Daraz (E-Commerce)
- Image Search: Extracts color histograms and SIFT features to match product images in the catalog.
- OCR: Template matching for reading text on product labels.
5. Medical Imaging (e.g., Kathmandu’s Patan Hospital)
- Segmentation + Description: Watershed algorithm segments tumors; Hu moments describe their shape for diagnosis.
- Recognition: CNNs classify X-rays as normal/abnormal (e.g., pneumonia detection).
Comparison Table: Feature Extraction Methods
| Method | Description | Pros | Cons | Applications |
|---|---|---|---|---|
| Sobel/Canny | Edge detection using gradients | Fast, simple | Sensitive to noise | Object detection |
| Harris Corner | Detects corners using intensity | Robust to rotation | Computationally expensive | Feature matching |
| SIFT/SURF | Scale-invariant feature transform | Invariant to scale/rotation | Slow for real-time | Image stitching |
| GLCM | Texture analysis using co-occurrence | Captures spatial relationships | High dimensionality | Medical imaging |
| Hu Moments | Shape descriptors using moments | Rotation/scale invariant | Sensitive to noise | Object recognition |
Exam Tip
- Understand the Math: Know how to compute Sobel, Hu moments, and DCT manually for small images.
- Diagrams: Always draw feature extraction pipelines (e.g., CNN layers) in exams.
- Real-World Links: Relate questions to apps like eSewa (QR recognition) or Pathao (object detection).
- Pros/Cons: Compare methods (e.g., SIFT vs. Harris corners) in tables.
- Code Snippets: Be ready to write pseudocode for template matching or k-NN classification.
- Visuals: Sketch edge-detected images or CNN architectures when asked about recognition.
Based on the TU BIT syllabus for Image Processing, unit 6.
Discussion
Loading…