Image ProcessingUnit 89 min read
Advanced Image Processing: Compression, Morphology, Transforms & AI
Unit 8 of Image Processing covers cutting-edge techniques—image compression standards (JPEG, JPEG2000, H.264), morphological operations (erosion/dilation), Fourier transforms (DFT vs FFT), neural networks for segmentation, and real-world pipelines—with visual workflows, math traces, and Nepalese examples like Ncell’s v
TAKEAWAYS
- Compression standards (JPEG, JPEG2000, H.264) trade off quality vs. size using DCT/FFT transforms—JPEG2000 uses wavelet transforms for better lossy compression, while H.264 adds motion compensation for video.
- Morphological operations (erosion/dilation) clean images via structuring elements (e.g., removing noise from Kathmandu traffic camera feeds) but distort shapes if overused.
- Fourier transforms (DFT/FFT) convert images to frequency space—FFT’s O(N log N) speed enables real-time filters (e.g., WhatsApp’s stickers use FFT for blur effects).
- Neural networks (CNNs, GANs) now outperform traditional methods for segmentation (e.g., Pathao’s ride-sharing app uses U-Net to detect potholes from street images).
- Exam hotspots: Compare restoration vs enhancement, derive histogram equalization, and sketch compression block diagrams—always link to real tools (e.g., GIMP’s filters use FFT).
1. Image Compression Standards: How They Work
Compression reduces file size while preserving visual quality. Standards differ in lossy vs lossless, transforms used, and applications.
Key Standards & Their Math
| Standard | Type | Transform Used | Key Feature | Example Use Case |
|---|---|---|---|---|
| JPEG | Lossy | DCT (Discrete Cosine) | 8×8 block processing, chroma subsampling | Daraz product images, Facebook photos |
| JPEG2000 | Lossy/Lossless | Wavelet (DWT) | Better compression for high-res (e.g., medical scans) | NTC’s satellite imagery |
| PNG | Lossless | None (run-length) | Transparency support | WhatsApp emoji, logos |
| H.264/AVC | Lossy (video) | DCT + motion compensation | Reduces temporal redundancy | YouTube videos, Ncell’s video calls |
Why DCT? The Discrete Cosine Transform converts an 8×8 pixel block into frequency coefficients. Most energy concentrates in low-frequency coefficients (smooth areas), so we discard high-frequency noise.
F(u,v) = \frac{1}{4} C(u)C(v) \sum_{x=0}^7 \sum_{y=0}^7 f(x,y) \cos\left[\frac{\pi u(2x+1)}{16}\right] \cos\left[\frac{\pi v(2y+1)}{16}\right]
Visual: JPEG’s 8×8 block DCT coefficients (low frequencies = smooth areas, high = edges).
Worked Example: JPEG Compression Pipeline
Input: A 512×512 RGB image (150KB uncompressed). Steps:
- Divide into 8×8 blocks → 4096 blocks.
- Apply DCT to each block → frequency coefficients.
- Quantize coefficients (divide by a quality matrix, e.g.,
Q[0,0]=1,Q[7,7]=100). - Zigzag scan → run-length encoding.
- Entropy coding (Huffman) → final
.jpg(~15KB at 80% quality).
Real World:
- Daraz’s product images: Uses JPEG for thumbnails (fast loading) and JPEG2000 for high-res zoomable images.
- Ncell’s video calls: H.264 compresses video streams in real time by predicting motion between frames (like a flipbook).
2. Morphological Operations: Cleaning Images with Shapes
Morphology uses structuring elements (kernels) to erode/dilate images. Critical for noise removal, edge detection, and object separation.
Core Operations
| Operation | Effect | Math Formula | Example Use Case |
|---|---|---|---|
| Erosion | Shrinks bright regions | Removing salt-and-pepper noise | |
| Dilation | Expands bright regions | Filling gaps in handwritten digits | |
| Opening | Erosion → Dilation | Removes small objects | Cleaning Kathmandu traffic camera images |
| Closing | Dilation → Erosion | Fills small holes | Restoring torn historical photos |
Structuring Elements:
- Disk: Smooths circular objects.
- Cross: Detects lines (e.g., grid removal).
- Custom shapes: For specific objects (e.g., a "T" shape for license plates).
Visual: Erosion vs dilation on a binary image of a "K" character.
Worked Example: Noise Removal in Traffic Cameras
Problem: A NTC traffic camera image has salt-and-pepper noise (random black/white pixels). Solution: Apply opening (erosion + dilation) with a 3×3 square kernel. Steps:
- Erode → Removes noise but shrinks the road.
- Dilate → Restores road size while keeping noise gone.
Code Snippet (Python/OpenCV):
import cv2
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (3,3))
cleaned = cv2.morphologyEx(noisy_image, cv2.MORPH_OPEN, kernel)
3. Fourier Transforms: From Pixels to Frequencies
The Discrete Fourier Transform (DFT) converts an image into frequency components, while Fast Fourier Transform (FFT) computes it efficiently.
DFT vs FFT
| Feature | DFT | FFT |
|---|---|---|
| Complexity | O(N²) | O(N log N) |
| Use Case | Small images, theoretical | Real-time filters (e.g., WhatsApp) |
| Output | Complex coefficients (magnitude + phase) | Same, but faster |
Visual: DFT of a simple image (a sine wave + noise).
Worked Example: Blurring with FFT
Goal: Apply a low-pass filter to remove high-frequency noise (e.g., WhatsApp’s "blur effect"). Steps:
- Compute 2D FFT of the image.
- Create a filter mask (e.g., Gaussian low-pass):
H(u,v) = e^{-\frac{D^2(u,v)}{2D_0^2}}, \quad D(u,v) = \sqrt{(u-M/2)^2 + (v-N/2)^2} - Multiply FFT by the mask.
- Inverse FFT → blurred image.
Real World:
- WhatsApp stickers: Uses FFT to apply artistic filters (e.g., oil-paint effect) in real time.
- Medical imaging: FFT helps detect tumors by isolating frequency bands in MRI scans.
4. Neural Networks for Image Processing
Traditional methods (filters, morphology) are now outperformed by Convolutional Neural Networks (CNNs) and Generative Adversarial Networks (GANs).
Key Architectures
| Model | Task | Example Use Case |
|---|---|---|
| U-Net | Segmentation | Pathao’s pothole detection |
| GANs | Super-resolution | Ncell’s old photo restoration |
| Autoencoders | Denoising | Daraz’s product image cleanup |
Visual: U-Net architecture for segmentation.
```figure
{"type":"network","nodes":["Input Image","Encoder (Conv Layers)","Bottleneck","Decoder (Transposed Conv)","Segmented Output"],"edges":[["Input Image","Encoder (Conv Layers)"],["Encoder (Conv Layers)","Bottleneck"],["Bottleneck","Decoder (Transposed Conv)"],["Decoder (Transposed Conv)","Segmented Output"],["Encoder (Conv Layers)","Decoder (Transposed Conv)","Skip Connection"]],"directed":true,"caption":"U-Net architecture with skip connections (simplified)"}
Worked Example: Pathao’s Pothole Detection
Problem: Detect potholes in street images for autonomous ride-hailing. Solution: Train a U-Net on labeled images (pothole = 1, road = 0). Steps:
- Input: 256×256 RGB image.
- Encoder: 4 conv layers (32→64→128→256 filters).
- Bottleneck: 512 filters.
- Decoder: Transposed conv → 1-channel output (binary mask).
- Output: Pixel-wise classification (e.g., white = pothole).
Real Picture:
5. Exam Tip: How to Score Full Marks
- Diagrams > Words: Always sketch block diagrams (e.g., compression pipeline) or DFT plots. Label every box/arrow.
- Example: For JPEG, draw:
Input Image → [8×8 Blocks] → DCT → Quantization → Huffman → Output
- Example: For JPEG, draw:
- Math Traces: Show one step of a calculation (e.g., DCT or histogram equalization). Examiners reward partial credit for setup.
- Real-World Links: Tie every concept to a Nepalese tool:
- FFT → WhatsApp filters.
- Morphology → Traffic camera noise removal.
- JPEG2000 → NTC satellite images.
- Differentiate Clearly: Use tables for comparisons (e.g., restoration vs enhancement).
- Avoid Memorization: Explain why a method works (e.g., "DCT works because natural images are smooth").
Final Visual Summary:
```mermaid
mindmap
root((Advanced Image Processing))
JPEG
DCT
Quantization
Huffman
Morphology
Erosion
Dilation
Opening/Closing
FFT
O(N log N)
Low-pass Filtering
Neural Networks
U-Net
GANs
Real World
Daraz: Compression
Pathao: Segmentation
Ncell: Video Codec
Based on the TU BCA syllabus for Image Processing, unit 8.
Discussion
Loading…