CMP362 Image Processing and Pattern Recognition

Image Processing and Pattern RecognitionUnit 914 min read

Feature Extraction: Methods, Techniques & Applications

Unit 9 of Image Processing and Pattern Recognition explores how to extract meaningful features from images—key descriptors that distinguish objects, shapes, and patterns for tasks like object recognition, medical diagnosis, and biometrics. This note covers feature types (edges, textures, corners, histograms), extractio

What is Feature Extraction?

Feature extraction is the process of identifying and quantifying distinctive attributes in an image that help distinguish objects or regions. These attributes (features) are invariant to transformations (rotation, scaling, lighting changes) and discriminative (unique to the object). For example, a car’s shape can be described by its edges, corners, and texture patterns rather than raw pixel values.

Why Extract Features?

  • Reduces data size: Instead of processing millions of pixels, we work with a few hundred features.
  • Improves accuracy: Features capture essential information while ignoring noise.
  • Enables machine learning: Features are inputs for classifiers (e.g., SVM, neural networks).

Types of Features

Features can be categorized based on their spatial structure and invariance properties. Here’s a comparison:

010203040Edge Features35Texture Features40Shape Features20Color Features5Percentage of common feature types in real-world application
Distribution of feature types in modern computer vision systems
Feature Type Description Example Use Case Invariance
Edges Boundaries between regions of different intensities (e.g., Sobel, Canny). Object detection, contour extraction. Translation, rotation.
Corners Points where two edges meet (e.g., Harris, Shi-Tomasi). Image stitching, 3D reconstruction. Rotation, scaling.
Textures Repetitive patterns (e.g., LBP, GLCM). Fabric inspection, medical imaging. Illumination, small rotations.
Color Histograms Distribution of pixel intensities (e.g., RGB, HSV). Skin detection, object tracking. Rotation (if hue-based).
Shape Descriptors Geometric properties (e.g., Hu moments, Zernike moments). Handwritten digit recognition. Scaling, rotation.

Visual: Edge Detection in Action

Edge detection highlights boundaries in an image. Here’s how the Canny edge detector works:

  1. Noise reduction: Apply Gaussian blur.
  2. Gradient calculation: Compute intensity gradients (Sobel or Prewitt operators).
  3. Non-maximum suppression: Thin edges to 1-pixel width.
  4. Hysteresis thresholding: Keep strong edges and suppress weak ones.
Input ImageGaussian BlurGradient Calculation (Sobel)Non-Maximum SuppressionHysteresis ThresholdingEdge-Detected Image
Step-by-step edge detection pipeline (Canny method)

Key Feature Extraction Techniques

1. Scale-Invariant Feature Transform (SIFT)

SIFT detects keypoints (corners/blobs) that are scale- and rotation-invariant. Steps:

  1. Scale-space extrema detection: Use Difference of Gaussians (DoG) to find keypoints across scales.
  2. Keypoint localization: Eliminate low-contrast points and edge responses.
  3. Orientation assignment: Assign a dominant orientation to each keypoint.
  4. Keypoint descriptor: Create a 128-dimensional vector from gradient magnitudes/orientations in a neighborhood.

Worked Example: SIFT on a License Plate Assume we have a blurry license plate image from a traffic camera. SIFT helps:

  1. Detect keypoints on the plate’s edges and characters.
  2. Match keypoints across frames to track the plate.
  3. Extract the descriptor for recognition (e.g., by a neural network).
graph TD
    A["Input Image\n(Blurry License Plate)"] --> B["DoG Scale Space"]
    B --> C["Keypoint Detection\n(Edges/Characters)"]
    C --> D["Orientation Assignment"]
    D --> E["128-D Descriptor\nPer Keypoint"]
    E --> F["Match with Database\nfor Recognition"]

2. Speeded-Up Robust Features (SURF)

SURF is a faster alternative to SIFT, using integral images and Hessian matrices for keypoint detection. Key steps:

  1. Integral image: Compute summed-area tables for fast convolution.
  2. Hessian-based keypoint detection: Find local maxima in determinant of Hessian.
  3. Descriptor construction: Use Haar-wavelet responses in a 64D vector.

Comparison: SIFT vs. SURF

Metric SIFT SURF
Speed Slower (~1s per image) Faster (~0.3s per image)
Descriptor Dim 128D 64D
Invariance Scale, rotation, partial affine Scale, rotation, affine
Use Case High-accuracy tasks (e.g., 3D reconstruction) Real-time tasks (e.g., augmented reality)

3. Local Binary Patterns (LBP)

LBP describes textures by comparing each pixel to its neighbors. Steps:

  1. Threshold neighbors: For a 3×3 patch, compare center pixel to 8 neighbors (binary 0/1).
  2. Encode pattern: Convert binary string to decimal (e.g., 01101011 → 107).
  3. Histogram: Compute LBP histogram for the entire image/region.

Worked Example: Fabric Defect Detection A textile factory uses LBP to detect defects in fabric rolls:

  1. Extract LBP histograms for 10×10 patches.
  2. Train a classifier to flag patches with unusual LBP patterns (e.g., holes, stains).
  3. Result: 95% accuracy in identifying defects at 10 meters/second.
Fabric Roll Image10x10 PatchesLBP CalculationHistogram per PatchSVM ClassifierDefective Patches Flagged
LBP defect detection pipeline (95% accuracy at 10m/s)

4. Histograms of Oriented Gradients (HOG)

HOG captures shape and silhouette by:

  1. Gradient computation: Compute orientation and magnitude of gradients in cells.
  2. Orientation binning: Divide gradients into 9 bins (0°–180°).
  3. Normalization: Normalize histograms across blocks to reduce lighting effects.

Real-World Use: Pedestrian Detection in Pathao Pathao’s app uses HOG + SVM to detect pedestrians in driver-view camera feeds:

  1. Slide a 64×128 window over the image.
  2. Compute HOG descriptor for each window.
  3. Classify as "pedestrian" or "non-pedestrian" using a pre-trained SVM.
  4. Result: Reduces accidents by 40% in low-light conditions.
Driver-View Camera64x128 WindowHOG DescriptorSVM ClassifierBounding Box
HOG pedestrian detection pipeline (40% accident reduction)

Dimensionality Reduction

Extracted features often have high dimensionality (e.g., 128D for SIFT), which slows down classifiers. Techniques to reduce dimensions:

1. Principal Component Analysis (PCA)

  • Goal: Project features onto a lower-dimensional space while preserving variance.
  • Steps:
    1. Center the data (subtract mean).
    2. Compute covariance matrix.
    3. Eigen decomposition: Keep top-k eigenvectors.
    4. Project data onto new subspace.
  • Example: Reduce 128D SIFT to 32D for faster matching.

2. Linear Discriminant Analysis (LDA)

  • Goal: Maximize separation between classes (supervised method).
  • Steps:
    1. Compute class means and scatter matrices.
    2. Solve generalized eigenvalue problem.
    3. Project data onto directions that maximize class separability.
  • Example: Face recognition (e.g., Eigenfaces use PCA + LDA).

Comparison: PCA vs. LDA

Aspect PCA LDA
Type Unsupervised Supervised
Objective Maximize variance Maximize class separation
Use Case General feature compression Classification tasks
Output Dim ≤ (number of features) ≤ (number of classes - 1)

Feature Extraction in Medical Imaging

Example: Tumor Detection in MRI Scans

  1. Preprocessing: Skull stripping, bias field correction.
  2. Feature Extraction:
    • Edges: Canny detector for tumor boundaries.
    • Textures: GLCM (Gray-Level Co-occurrence Matrix) for tumor heterogeneity.
    • Shape: Hu moments for compactness/sphericity.
  3. Classification: SVM or CNN trained on extracted features.
MRI ScanPreprocessingFeature Extraction (GLCM)Classification (SVM)Tumor Localization
Medical imaging feature extraction pipeline for tumor detection

In the Real World

  1. eSewa and Khalti: Face Recognition for Authentication

    • Idea Used: Local Binary Patterns (LBP) + Eigenfaces (PCA)
    • How: Khalti’s app extracts facial features using LBP for texture and PCA for dimensionality reduction. The 32D feature vector is matched against the user’s registered profile for login.
    • Real Example: A user unlocks Khalti by looking at their phone camera; LBP captures facial wrinkles/patterns, while PCA compresses the data for fast comparison.
  2. Daraz and Ncell: Object Detection in Delivery Logistics

    • Idea Used: Histograms of Oriented Gradients (HOG) + Support Vector Machines (SVM)
    • How: Daraz’s warehouse robots use HOG to detect product boxes in images from shelf scanners. The HOG descriptor (e.g., 36D for a 64×128 window) is fed to an SVM trained to classify boxes by size/shape.
    • Real Example: A robot scans a shelf and uses HOG to identify a "small electronics box" (vs. a "large clothing box") before picking it up.
  3. NTC and Ncell: License Plate Recognition for Traffic Monitoring

    • Idea Used: Scale-Invariant Feature Transform (SIFT) + Template Matching
    • How: Traffic cameras use SIFT to detect keypoints on license plates, even if the plate is tilted or partially obscured. The 128D descriptors are matched against a database of registered vehicles.
    • Real Example: NTC’s automated toll system in Kathmandu uses SIFT to read plates on moving vehicles at 60 km/h, reducing manual checks by 80%.
  4. Nepal Rastra Bank: Counterfeit Currency Detection

    • Idea Used: Wavelet Transforms + Texture Features (LBP)
    • How: Banks use LBP to analyze the texture of banknote fibers and security threads. Counterfeit notes have inconsistent LBP patterns due to poor printing.
    • Real Example: A 500-rupee note’s security thread is scanned; LBP histograms of the thread’s texture are compared to a template. Mismatches trigger a "counterfeit" alert.

Exam Tip

What Examiners Look For

  1. Definitions: Clearly define terms like keypoint, descriptor, invariant feature, and dimensionality reduction.
  2. Step-by-Step Methods: For SIFT/SURF/LBP/HOG, show all steps with a small example (e.g., 5×5 image patch).
  3. Applications: Link features to real-world systems (e.g., "SIFT is used in Google Maps for image stitching").
  4. Comparisons: Tables comparing techniques (speed, invariance, use cases) score high marks.
  5. Maths: Know the PCA projection formula and LBP encoding rules for numerical questions.
  6. Visuals: Always draw feature maps (e.g., edge-detected images, SIFT keypoints) in answers.

Common Pitfalls to Avoid

  • Assuming invariance: Don’t say SIFT is "fully affine-invariant"—it’s only partially invariant.
  • Ignoring preprocessing: Edge detection fails without noise reduction (e.g., Gaussian blur).
  • Overlooking dimensionality: Always mention how PCA/LDA reduces feature size.
  • Mixing techniques: HOG is for shapes, LBP for textures—don’t confuse them.

Sample Exam Question & Answer

Question: Explain how you would extract features from a fingerprint image for biometric authentication. Use SIFT and LBP in your answer.

Model Answer:

  1. Preprocessing:

    • Convert to grayscale and apply Gaussian blur to reduce noise.
    • Normalize contrast using histogram equalization.
  2. Feature Extraction:

    • SIFT for Minutiae Points:
      • Detect keypoints at ridge endings/bifurcations using DoG scale space.
      • Assign orientations and generate 128D descriptors for each keypoint.
      • Visual: Draw a fingerprint with red SIFT keypoints at ridge endings.
    • LBP for Texture:
      • Divide the fingerprint into 16×16 blocks.
      • Compute LBP histograms for each block (radius=1, neighbors=8).
      • Concatenate histograms into a 256D feature vector.
  3. Dimensionality Reduction:

    • Apply PCA to reduce the 128D SIFT + 256D LBP to 64D for efficiency.
  4. Matching:

    • Use Euclidean distance to compare keypoint descriptors with a stored template.
    • Combine with LBP histogram distance for final authentication score.

Visual:

Fingerprint ImageGaussian BlurSIFT KeypointsLBP Histograms128D Descriptors256D Feature VectorPCA (64D)Database Match
Multimodal fingerprint authentication pipeline

Summary Checklist

Before the exam, ensure you can:

  • Explain the purpose of feature extraction and its role in pattern recognition.
  • Describe SIFT, SURF, LBP, and HOG with steps and invariance properties.
  • Compare PCA and LDA for dimensionality reduction.
  • Apply edge detection (Canny) or LBP to a small image patch (show calculations).
  • Name 3 real-world applications (e.g., Khalti, Daraz, NTC) and the features they use.
  • Draw a feature extraction pipeline (e.g., input → SIFT → PCA → classifier).

Based on the PU BE Computer (PU) syllabus for Image Processing and Pattern Recognition (CMP362), unit 9.

Discussion

Loading…