Image Processing and Pattern RecognitionUnit 914 min read
Feature Extraction: Methods, Techniques & Applications
Unit 9 of Image Processing and Pattern Recognition explores how to extract meaningful features from images—key descriptors that distinguish objects, shapes, and patterns for tasks like object recognition, medical diagnosis, and biometrics. This note covers feature types (edges, textures, corners, histograms), extractio
What is Feature Extraction?
Feature extraction is the process of identifying and quantifying distinctive attributes in an image that help distinguish objects or regions. These attributes (features) are invariant to transformations (rotation, scaling, lighting changes) and discriminative (unique to the object). For example, a car’s shape can be described by its edges, corners, and texture patterns rather than raw pixel values.
Why Extract Features?
- Reduces data size: Instead of processing millions of pixels, we work with a few hundred features.
- Improves accuracy: Features capture essential information while ignoring noise.
- Enables machine learning: Features are inputs for classifiers (e.g., SVM, neural networks).
Types of Features
Features can be categorized based on their spatial structure and invariance properties. Here’s a comparison:
| Feature Type | Description | Example Use Case | Invariance |
|---|---|---|---|
| Edges | Boundaries between regions of different intensities (e.g., Sobel, Canny). | Object detection, contour extraction. | Translation, rotation. |
| Corners | Points where two edges meet (e.g., Harris, Shi-Tomasi). | Image stitching, 3D reconstruction. | Rotation, scaling. |
| Textures | Repetitive patterns (e.g., LBP, GLCM). | Fabric inspection, medical imaging. | Illumination, small rotations. |
| Color Histograms | Distribution of pixel intensities (e.g., RGB, HSV). | Skin detection, object tracking. | Rotation (if hue-based). |
| Shape Descriptors | Geometric properties (e.g., Hu moments, Zernike moments). | Handwritten digit recognition. | Scaling, rotation. |
Visual: Edge Detection in Action
Edge detection highlights boundaries in an image. Here’s how the Canny edge detector works:
- Noise reduction: Apply Gaussian blur.
- Gradient calculation: Compute intensity gradients (Sobel or Prewitt operators).
- Non-maximum suppression: Thin edges to 1-pixel width.
- Hysteresis thresholding: Keep strong edges and suppress weak ones.
Key Feature Extraction Techniques
1. Scale-Invariant Feature Transform (SIFT)
SIFT detects keypoints (corners/blobs) that are scale- and rotation-invariant. Steps:
- Scale-space extrema detection: Use Difference of Gaussians (DoG) to find keypoints across scales.
- Keypoint localization: Eliminate low-contrast points and edge responses.
- Orientation assignment: Assign a dominant orientation to each keypoint.
- Keypoint descriptor: Create a 128-dimensional vector from gradient magnitudes/orientations in a neighborhood.
Worked Example: SIFT on a License Plate Assume we have a blurry license plate image from a traffic camera. SIFT helps:
- Detect keypoints on the plate’s edges and characters.
- Match keypoints across frames to track the plate.
- Extract the descriptor for recognition (e.g., by a neural network).
graph TD
A["Input Image\n(Blurry License Plate)"] --> B["DoG Scale Space"]
B --> C["Keypoint Detection\n(Edges/Characters)"]
C --> D["Orientation Assignment"]
D --> E["128-D Descriptor\nPer Keypoint"]
E --> F["Match with Database\nfor Recognition"]2. Speeded-Up Robust Features (SURF)
SURF is a faster alternative to SIFT, using integral images and Hessian matrices for keypoint detection. Key steps:
- Integral image: Compute summed-area tables for fast convolution.
- Hessian-based keypoint detection: Find local maxima in determinant of Hessian.
- Descriptor construction: Use Haar-wavelet responses in a 64D vector.
Comparison: SIFT vs. SURF
| Metric | SIFT | SURF |
|---|---|---|
| Speed | Slower (~1s per image) | Faster (~0.3s per image) |
| Descriptor Dim | 128D | 64D |
| Invariance | Scale, rotation, partial affine | Scale, rotation, affine |
| Use Case | High-accuracy tasks (e.g., 3D reconstruction) | Real-time tasks (e.g., augmented reality) |
3. Local Binary Patterns (LBP)
LBP describes textures by comparing each pixel to its neighbors. Steps:
- Threshold neighbors: For a 3×3 patch, compare center pixel to 8 neighbors (binary 0/1).
- Encode pattern: Convert binary string to decimal (e.g.,
01101011→ 107). - Histogram: Compute LBP histogram for the entire image/region.
Worked Example: Fabric Defect Detection A textile factory uses LBP to detect defects in fabric rolls:
- Extract LBP histograms for 10×10 patches.
- Train a classifier to flag patches with unusual LBP patterns (e.g., holes, stains).
- Result: 95% accuracy in identifying defects at 10 meters/second.
4. Histograms of Oriented Gradients (HOG)
HOG captures shape and silhouette by:
- Gradient computation: Compute orientation and magnitude of gradients in cells.
- Orientation binning: Divide gradients into 9 bins (0°–180°).
- Normalization: Normalize histograms across blocks to reduce lighting effects.
Real-World Use: Pedestrian Detection in Pathao Pathao’s app uses HOG + SVM to detect pedestrians in driver-view camera feeds:
- Slide a 64×128 window over the image.
- Compute HOG descriptor for each window.
- Classify as "pedestrian" or "non-pedestrian" using a pre-trained SVM.
- Result: Reduces accidents by 40% in low-light conditions.
Dimensionality Reduction
Extracted features often have high dimensionality (e.g., 128D for SIFT), which slows down classifiers. Techniques to reduce dimensions:
1. Principal Component Analysis (PCA)
- Goal: Project features onto a lower-dimensional space while preserving variance.
- Steps:
- Center the data (subtract mean).
- Compute covariance matrix.
- Eigen decomposition: Keep top-k eigenvectors.
- Project data onto new subspace.
- Example: Reduce 128D SIFT to 32D for faster matching.
2. Linear Discriminant Analysis (LDA)
- Goal: Maximize separation between classes (supervised method).
- Steps:
- Compute class means and scatter matrices.
- Solve generalized eigenvalue problem.
- Project data onto directions that maximize class separability.
- Example: Face recognition (e.g., Eigenfaces use PCA + LDA).
Comparison: PCA vs. LDA
| Aspect | PCA | LDA |
|---|---|---|
| Type | Unsupervised | Supervised |
| Objective | Maximize variance | Maximize class separation |
| Use Case | General feature compression | Classification tasks |
| Output Dim | ≤ (number of features) | ≤ (number of classes - 1) |
Feature Extraction in Medical Imaging
Example: Tumor Detection in MRI Scans
- Preprocessing: Skull stripping, bias field correction.
- Feature Extraction:
- Edges: Canny detector for tumor boundaries.
- Textures: GLCM (Gray-Level Co-occurrence Matrix) for tumor heterogeneity.
- Shape: Hu moments for compactness/sphericity.
- Classification: SVM or CNN trained on extracted features.
In the Real World
eSewa and Khalti: Face Recognition for Authentication
- Idea Used: Local Binary Patterns (LBP) + Eigenfaces (PCA)
- How: Khalti’s app extracts facial features using LBP for texture and PCA for dimensionality reduction. The 32D feature vector is matched against the user’s registered profile for login.
- Real Example: A user unlocks Khalti by looking at their phone camera; LBP captures facial wrinkles/patterns, while PCA compresses the data for fast comparison.
Daraz and Ncell: Object Detection in Delivery Logistics
- Idea Used: Histograms of Oriented Gradients (HOG) + Support Vector Machines (SVM)
- How: Daraz’s warehouse robots use HOG to detect product boxes in images from shelf scanners. The HOG descriptor (e.g., 36D for a 64×128 window) is fed to an SVM trained to classify boxes by size/shape.
- Real Example: A robot scans a shelf and uses HOG to identify a "small electronics box" (vs. a "large clothing box") before picking it up.
NTC and Ncell: License Plate Recognition for Traffic Monitoring
- Idea Used: Scale-Invariant Feature Transform (SIFT) + Template Matching
- How: Traffic cameras use SIFT to detect keypoints on license plates, even if the plate is tilted or partially obscured. The 128D descriptors are matched against a database of registered vehicles.
- Real Example: NTC’s automated toll system in Kathmandu uses SIFT to read plates on moving vehicles at 60 km/h, reducing manual checks by 80%.
Nepal Rastra Bank: Counterfeit Currency Detection
- Idea Used: Wavelet Transforms + Texture Features (LBP)
- How: Banks use LBP to analyze the texture of banknote fibers and security threads. Counterfeit notes have inconsistent LBP patterns due to poor printing.
- Real Example: A 500-rupee note’s security thread is scanned; LBP histograms of the thread’s texture are compared to a template. Mismatches trigger a "counterfeit" alert.
Exam Tip
What Examiners Look For
- Definitions: Clearly define terms like keypoint, descriptor, invariant feature, and dimensionality reduction.
- Step-by-Step Methods: For SIFT/SURF/LBP/HOG, show all steps with a small example (e.g., 5×5 image patch).
- Applications: Link features to real-world systems (e.g., "SIFT is used in Google Maps for image stitching").
- Comparisons: Tables comparing techniques (speed, invariance, use cases) score high marks.
- Maths: Know the PCA projection formula and LBP encoding rules for numerical questions.
- Visuals: Always draw feature maps (e.g., edge-detected images, SIFT keypoints) in answers.
Common Pitfalls to Avoid
- Assuming invariance: Don’t say SIFT is "fully affine-invariant"—it’s only partially invariant.
- Ignoring preprocessing: Edge detection fails without noise reduction (e.g., Gaussian blur).
- Overlooking dimensionality: Always mention how PCA/LDA reduces feature size.
- Mixing techniques: HOG is for shapes, LBP for textures—don’t confuse them.
Sample Exam Question & Answer
Question: Explain how you would extract features from a fingerprint image for biometric authentication. Use SIFT and LBP in your answer.
Model Answer:
Preprocessing:
- Convert to grayscale and apply Gaussian blur to reduce noise.
- Normalize contrast using histogram equalization.
Feature Extraction:
- SIFT for Minutiae Points:
- Detect keypoints at ridge endings/bifurcations using DoG scale space.
- Assign orientations and generate 128D descriptors for each keypoint.
- Visual: Draw a fingerprint with red SIFT keypoints at ridge endings.
- LBP for Texture:
- Divide the fingerprint into 16×16 blocks.
- Compute LBP histograms for each block (radius=1, neighbors=8).
- Concatenate histograms into a 256D feature vector.
- SIFT for Minutiae Points:
Dimensionality Reduction:
- Apply PCA to reduce the 128D SIFT + 256D LBP to 64D for efficiency.
Matching:
- Use Euclidean distance to compare keypoint descriptors with a stored template.
- Combine with LBP histogram distance for final authentication score.
Visual:
Summary Checklist
Before the exam, ensure you can:
- Explain the purpose of feature extraction and its role in pattern recognition.
- Describe SIFT, SURF, LBP, and HOG with steps and invariance properties.
- Compare PCA and LDA for dimensionality reduction.
- Apply edge detection (Canny) or LBP to a small image patch (show calculations).
- Name 3 real-world applications (e.g., Khalti, Daraz, NTC) and the features they use.
- Draw a feature extraction pipeline (e.g., input → SIFT → PCA → classifier).
Based on the PU BE Computer (PU) syllabus for Image Processing and Pattern Recognition (CMP362), unit 9.
Discussion
Loading…