Image Processing and Pattern RecognitionUnit 815 min read
Image Segmentation: Techniques, Algorithms & Applications
Unit 8 of Image Processing and Pattern Recognition covers image segmentation—partitioning an image into meaningful regions/objects using spatial, spectral, or contextual cues. Learn thresholding, edge-based, region-based, and clustering methods, their mathematical foundations, and real-world applications in medical ima
TAKEAWAYS:
- Image segmentation divides an image into homogeneous regions based on pixel intensity, texture, or higher-level features (e.g., edges, color).
- Thresholding (global/local) and edge detection (Sobel, Canny) are basic but powerful methods for simple segmentation tasks.
- Region-growing and clustering (e.g., K-means, watershed) handle complex scenes but require careful parameter tuning.
- Deep learning (U-Net, Mask R-CNN) dominates modern segmentation due to its ability to learn hierarchical features from data.
- Evaluation metrics (Dice coefficient, IoU) quantify segmentation accuracy, critical for real-world deployment.
- Applications span medical diagnosis (tumor detection), autonomous driving (lane/pedestrian segmentation), and agriculture (crop health monitoring).
1. Introduction to Image Segmentation
Why Segment Images?
- Object detection: Identify and isolate objects (e.g., faces in a crowd).
- Medical imaging: Segment tumors, organs, or blood vessels in MRI/CT scans.
- Autonomous systems: Separate lanes, pedestrians, or obstacles in self-driving cars.
- Augmented reality: Track and manipulate real-world objects in AR apps.
Types of Segmentation
| Type | Description | Example Use Case |
|---|---|---|
| Pixel-level | Each pixel is assigned to a region. | Medical imaging (tissue classification). |
| Super-pixel | Groups of pixels (super-pixels) are segmented. | Object tracking in videos. |
| Instance-level | Segments individual objects (e.g., separate cars in a parking lot). | Autonomous driving. |
| Semantic-level | Labels regions by class (e.g., "sky," "road") without distinguishing instances. | Scene understanding in robotics. |
2. Fundamental Approaches to Segmentation
A. Thresholding-Based Segmentation
Thresholding converts a grayscale image into binary (or multi-level) by comparing pixel intensities to a threshold T.
Methods
Global Thresholding:
- Single threshold T for the entire image.
- Formula:
- Limitations: Fails in non-uniform lighting (e.g., shadows).
Local (Adaptive) Thresholding:
- Computes T for small regions (e.g., using mean or median of a window).
- Example: Otsu’s method (automatically selects T to maximize inter-class variance).
Worked Example: Global Thresholding on a Medical Image
Scenario: Segment a blood cell from a microscope image (assume background is darker). Given:
- Grayscale image f(x,y) with pixel values:
[50, 60, 70, 80, 100, 120, 130, 140, 150]. - Choose T = 100 (empirically or via Otsu’s method).
Steps:
- Apply threshold:
- Pixels ≥100 →
1(foreground: blood cell). - Pixels <100 →
0(background).
- Pixels ≥100 →
- Result:
[0, 0, 0, 0, 0, 1, 1, 1, 1]
Visualization:
graph LR
A["Original Image\n(8-bit grayscale)"] -->|"Threshold T=100"| B["Binary Image\nForeground=1"]
B -->|"Post-processing"| C["Cleaned Segmentation\n(Morphological ops)"]Real-World Tie-In:
- eSewa’s OTP verification: Thresholding can segment handwritten digits (e.g., separating a
7from noise in a scanned OTP image). However, adaptive thresholding is preferred for varying lighting.
B. Edge-Based Segmentation
Edges mark boundaries between regions. Methods:
Gradient-Based (Sobel, Prewitt, Laplacian):
- Compute gradients to detect intensity changes.
- Sobel Operator: Edge magnitude: .
Canny Edge Detector (3-step process):
- Smooth (Gaussian blur).
- Compute gradients (Sobel).
- Non-maximum suppression + hysteresis thresholding.
Worked Example: Sobel Edge Detection
Scenario: Detect edges in a simple shape (e.g., a square on a plain background). Given:
[0, 0, 0, 0, 0;
0,255,255,255,0;
0,255,255,255,0;
0,255,255,255,0;
0, 0, 0, 0, 0]
Steps:
- Apply and to each 3×3 patch.
- Compute .
- Threshold G to get edges (e.g., T = 100).
Result:
[0, 0, 0, 0, 0;
0,255, 0,255,0;
0, 0, 0, 0,0;
0,255, 0,255,0;
0, 0, 0, 0, 0]
Visualization:
graph LR
A["Original Image\n(Square)"] -->|"Sobel Operator"| B["Gradient Magnitude"]
B -->|"Thresholding"| C["Edge Map\n(Boundaries only)"]Real-World Tie-In:
- Pathao’s delivery route optimization: Edge detection segments roads from satellite images to map delivery paths. For example, detecting a pothole (edge of a dark hole) helps avoid accidents.
C. Region-Based Segmentation
Groups pixels into regions based on similarity (e.g., intensity, texture).
Methods
Region Growing:
- Start with a seed pixel, iteratively add neighboring pixels if they meet a similarity criterion (e.g., intensity difference < T).
- Challenge: Over-segmentation if T is too small.
Split-and-Merge:
- Split: Divide image into squares; merge if regions are similar.
- Merge: Combine adjacent regions if they satisfy a homogeneity criterion.
Watershed Algorithm:
- Treats the image as a topographic surface; "water" fills basins until dams (edges) form.
- Problem: Over-segmentation (solved by markers or gradient-based watershed).
Worked Example: Region Growing
Scenario: Segment a circle from noise. Given:
- Image with a bright circle (intensity = 200) on a dark background (intensity = 50).
- Seed pixel at center (200), T = 30.
Steps:
- Start at seed (200). Compare neighbors:
- Neighbor A: 190 → |200–190| = 10 < 30 → add to region.
- Neighbor B: 40 → |200–40| = 160 > 30 → ignore.
- Repeat until no more pixels meet the criterion.
Visualization:
graph TD
A["Seed Pixel\n(200)"] -->|"Grow Region"| B["Region 1\n(All pixels ≥170)"]
B -->|"Stop"| C["Final Segmented\nCircle"]Real-World Tie-In:
- NTC’s traffic monitoring: Region growing segments vehicles from CCTV footage. For example, in Kathmandu’s busy Thapathali, a camera captures a frame where cars are brighter than the road. By setting a seed in a car and growing the region, NTC can count vehicles in real time.
D. Clustering-Based Segmentation
Groups pixels into K clusters (e.g., using K-means) based on feature vectors (RGB, texture).
K-means Segmentation
- Initialize: Choose K centroids randomly.
- Assign: Each pixel → nearest centroid.
- Update: Recompute centroids as mean of assigned pixels.
- Repeat until convergence.
Worked Example: K-means on a Color Image
Scenario: Segment a simple image into 3 regions (e.g., sky, grass, object). Given:
- RGB pixels:
[(255,255,255), (0,128,0), (100,100,100)](white, green, gray). - K = 3.
Steps:
- Initialize centroids:
C1 = (255,255,255),C2 = (0,128,0),C3 = (100,100,100). - Assign pixels to nearest centroid (Euclidean distance in RGB space).
- Update centroids (no change in this case).
Result:
- Cluster 1: White pixels.
- Cluster 2: Green pixels.
- Cluster 3: Gray pixels.
Visualization:
graph LR
A["RGB Image"] -->|"K-means<br/>K=3"| B["Cluster 1\n(White)"]
A --> C["Cluster 2\n(Green)"]
A --> D["Cluster 3\n(Gray)"]Real-World Tie-In:
- Daraz’s product categorization: K-means segments product images by color dominant in the background (e.g., white for electronics, brown for furniture). This helps in auto-tagging products for search.
E. Deep Learning for Segmentation
Modern methods use Convolutional Neural Networks (CNNs) to learn hierarchical features.
Key Architectures
| Method | Description | Example |
|---|---|---|
| FCN (Fully Convolutional Network) | End-to-end pixel-wise classification. | Semantic segmentation. |
| U-Net | Encoder-decoder with skip connections (used in medical imaging). | Tumor segmentation in MRI. |
| Mask R-CNN | Extends Faster R-CNN for instance segmentation (bounds + mask). | Autonomous driving (object + shape). |
How U-Net Works
- Encoder: Downsample (contract) to extract features.
- Decoder: Upsample (expand) with skip connections to preserve spatial info.
- Output: Pixel-wise class probabilities.
Visualization:
graph LR
A["Input Image"] --> B["Encoder\n(Feature Extraction)"]
B --> C["Bottleneck\n(Deep Features)"]
C --> D["Decoder\n(Upsampling)"]
D --> E["Skip Connections\n(Preserve Location)"]
E --> F["Segmentation Mask"]Real-World Tie-In:
- Google’s DeepMind in healthcare: U-Net segments retinal blood vessels from OCT scans to detect glaucoma. A single mis-segmented vessel can change diagnosis from healthy to diseased.
3. Evaluation Metrics for Segmentation
Quantify accuracy using:
Pixel Accuracy:
- Limitation: Biased toward dominant classes.
Dice Similarity Coefficient (DSC):
- Measures overlap between predicted (A) and ground truth (B) regions.
Intersection over Union (IoU):
- Higher IoU = better segmentation.
Worked Example: DSC Calculation
Scenario: Compare a predicted tumor region (A) with ground truth (B).
- A = 100 pixels, B = 120 pixels, A ∩ B = 90 pixels. Calculation: \text{DSC} = \frac{2 \times 90}{100 + 120} = \frac{180}{220} \approx 0.818 \quad (\text{81.8%})
4. Applications of Image Segmentation
| Domain | Application | Method Used |
|---|---|---|
| Medical Imaging | Tumor detection in MRI/CT scans. | U-Net, Watershed. |
| Autonomous Vehicles | Lane/pedestrian segmentation. | Mask R-CNN, Canny edges. |
| Agriculture | Crop disease detection (e.g., leaf spots). | K-means, Thresholding. |
| Retail | Product background removal (e.g., Daraz listings). | GrabCut (interactive segmentation). |
| Security | Face recognition in surveillance. | Edge detection + template matching. |
In the Real World
eSewa’s OTP Verification:
- Idea Used: Adaptive thresholding + morphological operations.
- How: When you enter an OTP on eSewa, the system scans your handwritten digit (e.g.,
7) from a photo. Adaptive thresholding separates the digit from noise, while morphological closing fills small gaps (e.g., a broken7). This ensures the OTP is read correctly even if your handwriting is messy.
Pathao’s Delivery Route Optimization:
- Idea Used: Edge detection (Canny) + region growing.
- How: Pathao’s app uses satellite images to detect roads (edges) and buildings (regions). For example, in Pokhara’s busy Lakeside, Canny edges highlight roads, while region growing segments buildings to avoid. This reduces delivery time by 20% by planning the shortest path.
Nepal Rastra Bank’s Currency Validation:
- Idea Used: Color clustering (K-means) + texture analysis.
- How: To detect counterfeit notes, the bank uses K-means to segment ink colors and analyze texture (e.g., holograms). A real 1000-rupee note has distinct blue and green clusters in its security thread, while a fake note may lack these clusters entirely.
5. Challenges and Limitations
- Computational Cost: Deep learning methods (e.g., U-Net) require GPUs.
- Parameter Sensitivity: Thresholding/region growing fails with poor lighting.
- Over/Under-Segmentation: Watershed may produce too many regions; merging is needed.
- Real-Time Constraints: Edge devices (e.g., drones) need lightweight methods (e.g., Sobel + thresholding).
Exam Tip
What to Expect in TU/PU Exams
Theory Questions (30%):
- Define thresholding, region growing, and watershed algorithm.
- Compare global vs. adaptive thresholding (table format).
- Explain why deep learning (U-Net) outperforms traditional methods for medical imaging.
Problem-Solving (40%):
- Given: A 3×3 grayscale image and a threshold T. Task: Apply global thresholding and draw the binary output.
- Given: A Sobel operator and an image patch. Task: Compute gradients and identify edges.
- Given: A segmented image with DSC = 0.7. Task: Calculate the number of overlapping pixels if ground truth has 200 pixels.
Application-Based (20%):
- Scenario: "How would you segment blood cells in a microscope image?" Answer: Use Otsu’s thresholding for global segmentation, followed by morphological opening to remove noise.
- Scenario: "Design a system to count vehicles in Kathmandu traffic."
Answer:
- Capture frame (CCTV).
- Apply Canny edge detection to find vehicle boundaries.
- Use region growing to segment each vehicle.
- Count regions.
Short Answer (10%):
- "What is the difference between semantic and instance segmentation?"
- "Why does watershed segmentation often over-segment images?"
How to Score Full Marks
- For theory: Use bullet points + diagrams (e.g., show the Sobel operator matrix).
- For calculations: Show every step (e.g., DSC formula with substituted values).
- For applications: Tie answers to real-world examples (e.g., eSewa, Pathao).
- Diagrams: Always draw before/after segmentation (e.g., original → binary mask).
Practice Questions
Apply local thresholding (mean-based) to the following 3×3 image with window size = 3×3:
[50, 60, 70; 80, 90, 100; 110, 120, 130]Assume T = mean of the window.
Given a Canny edge detector output, how would you fill gaps in detected edges? Name two methods.
Compare K-means segmentation and region growing in a table (criteria: speed, parameter sensitivity, output quality).
Key Formulas to Memorize
| Concept | Formula |
|---|---|
| Global Thresholding | |
| Sobel Gradient | |
| Dice Coefficient | |
| IoU |
Final Visual Summary
mindmap
root((Image Segmentation))
-> Approaches
-> Thresholding
-> Global
-> Adaptive (Otsu)
-> Edge-Based
-> Sobel
-> Canny
-> Region-Based
-> Region Growing
-> Watershed
-> Clustering
-> K-means
-> Deep Learning
-> U-Net
-> Mask R-CNN
-> Applications
-> Medical (Tumor Segmentation)
-> Autonomous Vehicles (Lane Detection)
-> Retail (Product Background Removal)
-> Evaluation
-> DSC
-> IoUBased on the PU BE Computer (PU) syllabus for Image Processing and Pattern Recognition (CMP362), unit 8.
Discussion
Loading…