Elective Image Processing

Image ProcessingUnit 89 min read

Advanced Image Processing: Compression, Morphology, Transforms & AI

Unit 8 of Image Processing covers cutting-edge techniques—image compression standards (JPEG, JPEG2000, H.264), morphological operations (erosion/dilation), Fourier transforms (DFT vs FFT), neural networks for segmentation, and real-world pipelines—with visual workflows, math traces, and Nepalese examples like Ncell’s v

TAKEAWAYS

  • Compression standards (JPEG, JPEG2000, H.264) trade off quality vs. size using DCT/FFT transforms—JPEG2000 uses wavelet transforms for better lossy compression, while H.264 adds motion compensation for video.
  • Morphological operations (erosion/dilation) clean images via structuring elements (e.g., removing noise from Kathmandu traffic camera feeds) but distort shapes if overused.
  • Fourier transforms (DFT/FFT) convert images to frequency space—FFT’s O(N log N) speed enables real-time filters (e.g., WhatsApp’s stickers use FFT for blur effects).
  • Neural networks (CNNs, GANs) now outperform traditional methods for segmentation (e.g., Pathao’s ride-sharing app uses U-Net to detect potholes from street images).
  • Exam hotspots: Compare restoration vs enhancement, derive histogram equalization, and sketch compression block diagrams—always link to real tools (e.g., GIMP’s filters use FFT).

1. Image Compression Standards: How They Work

Compression reduces file size while preserving visual quality. Standards differ in lossy vs lossless, transforms used, and applications.

Key Standards & Their Math

Standard Type Transform Used Key Feature Example Use Case
JPEG Lossy DCT (Discrete Cosine) 8×8 block processing, chroma subsampling Daraz product images, Facebook photos
JPEG2000 Lossy/Lossless Wavelet (DWT) Better compression for high-res (e.g., medical scans) NTC’s satellite imagery
PNG Lossless None (run-length) Transparency support WhatsApp emoji, logos
H.264/AVC Lossy (video) DCT + motion compensation Reduces temporal redundancy YouTube videos, Ncell’s video calls

Why DCT? The Discrete Cosine Transform converts an 8×8 pixel block into frequency coefficients. Most energy concentrates in low-frequency coefficients (smooth areas), so we discard high-frequency noise.

F(u,v) = \frac{1}{4} C(u)C(v) \sum_{x=0}^7 \sum_{y=0}^7 f(x,y) \cos\left[\frac{\pi u(2x+1)}{16}\right] \cos\left[\frac{\pi v(2y+1)}{16}\right]

Visual: JPEG’s 8×8 block DCT coefficients (low frequencies = smooth areas, high = edges).


Worked Example: JPEG Compression Pipeline

Input: A 512×512 RGB image (150KB uncompressed). Steps:

  1. Divide into 8×8 blocks → 4096 blocks.
  2. Apply DCT to each block → frequency coefficients.
  3. Quantize coefficients (divide by a quality matrix, e.g., Q[0,0]=1, Q[7,7]=100).
  4. Zigzag scan → run-length encoding.
  5. Entropy coding (Huffman) → final .jpg (~15KB at 80% quality).

Real World:

  • Daraz’s product images: Uses JPEG for thumbnails (fast loading) and JPEG2000 for high-res zoomable images.
  • Ncell’s video calls: H.264 compresses video streams in real time by predicting motion between frames (like a flipbook).

2. Morphological Operations: Cleaning Images with Shapes

Morphology uses structuring elements (kernels) to erode/dilate images. Critical for noise removal, edge detection, and object separation.

Core Operations

Operation Effect Math Formula Example Use Case
Erosion Shrinks bright regions Removing salt-and-pepper noise
Dilation Expands bright regions Filling gaps in handwritten digits
Opening Erosion → Dilation Removes small objects Cleaning Kathmandu traffic camera images
Closing Dilation → Erosion Fills small holes Restoring torn historical photos

Structuring Elements:

  • Disk: Smooths circular objects.
  • Cross: Detects lines (e.g., grid removal).
  • Custom shapes: For specific objects (e.g., a "T" shape for license plates).

Visual: Erosion vs dilation on a binary image of a "K" character.


Worked Example: Noise Removal in Traffic Cameras

Problem: A NTC traffic camera image has salt-and-pepper noise (random black/white pixels). Solution: Apply opening (erosion + dilation) with a 3×3 square kernel. Steps:

  1. Erode → Removes noise but shrinks the road.
  2. Dilate → Restores road size while keeping noise gone.
Original ImageNoiseErosionDilationCleaned Image
Morphological noise removal sequence for traffic images

Code Snippet (Python/OpenCV):

import cv2
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (3,3))
cleaned = cv2.morphologyEx(noisy_image, cv2.MORPH_OPEN, kernel)

3. Fourier Transforms: From Pixels to Frequencies

The Discrete Fourier Transform (DFT) converts an image into frequency components, while Fast Fourier Transform (FFT) computes it efficiently.

DFT vs FFT

Feature DFT FFT
Complexity O(N²) O(N log N)
Use Case Small images, theoretical Real-time filters (e.g., WhatsApp)
Output Complex coefficients (magnitude + phase) Same, but faster

Visual: DFT of a simple image (a sine wave + noise).


Worked Example: Blurring with FFT

Goal: Apply a low-pass filter to remove high-frequency noise (e.g., WhatsApp’s "blur effect"). Steps:

  1. Compute 2D FFT of the image.
  2. Create a filter mask (e.g., Gaussian low-pass):
    H(u,v) = e^{-\frac{D^2(u,v)}{2D_0^2}}, \quad D(u,v) = \sqrt{(u-M/2)^2 + (v-N/2)^2}
    
  3. Multiply FFT by the mask.
  4. Inverse FFT → blurred image.
-8-6-4-224680.20.40.60.81xyOriginal signalBlurred signal (low-pass filtered)
Frequency domain blurring via low-pass filtering (FFT example)

Real World:

  • WhatsApp stickers: Uses FFT to apply artistic filters (e.g., oil-paint effect) in real time.
  • Medical imaging: FFT helps detect tumors by isolating frequency bands in MRI scans.

4. Neural Networks for Image Processing

Traditional methods (filters, morphology) are now outperformed by Convolutional Neural Networks (CNNs) and Generative Adversarial Networks (GANs).

Key Architectures

Model Task Example Use Case
U-Net Segmentation Pathao’s pothole detection
GANs Super-resolution Ncell’s old photo restoration
Autoencoders Denoising Daraz’s product image cleanup

Visual: U-Net architecture for segmentation.

```figure
{"type":"network","nodes":["Input Image","Encoder (Conv Layers)","Bottleneck","Decoder (Transposed Conv)","Segmented Output"],"edges":[["Input Image","Encoder (Conv Layers)"],["Encoder (Conv Layers)","Bottleneck"],["Bottleneck","Decoder (Transposed Conv)"],["Decoder (Transposed Conv)","Segmented Output"],["Encoder (Conv Layers)","Decoder (Transposed Conv)","Skip Connection"]],"directed":true,"caption":"U-Net architecture with skip connections (simplified)"}

Worked Example: Pathao’s Pothole Detection

Problem: Detect potholes in street images for autonomous ride-hailing. Solution: Train a U-Net on labeled images (pothole = 1, road = 0). Steps:

  1. Input: 256×256 RGB image.
  2. Encoder: 4 conv layers (32→64→128→256 filters).
  3. Bottleneck: 512 filters.
  4. Decoder: Transposed conv → 1-channel output (binary mask).
  5. Output: Pixel-wise classification (e.g., white = pothole).

Real Picture:



5. Exam Tip: How to Score Full Marks

  1. Diagrams > Words: Always sketch block diagrams (e.g., compression pipeline) or DFT plots. Label every box/arrow.
    • Example: For JPEG, draw:
      Input Image → [8×8 Blocks] → DCT → Quantization → Huffman → Output
      
  2. Math Traces: Show one step of a calculation (e.g., DCT or histogram equalization). Examiners reward partial credit for setup.
  3. Real-World Links: Tie every concept to a Nepalese tool:
    • FFT → WhatsApp filters.
    • Morphology → Traffic camera noise removal.
    • JPEG2000 → NTC satellite images.
  4. Differentiate Clearly: Use tables for comparisons (e.g., restoration vs enhancement).
  5. Avoid Memorization: Explain why a method works (e.g., "DCT works because natural images are smooth").

Final Visual Summary:

```mermaid
mindmap
  root((Advanced Image Processing))
    JPEG
      DCT
      Quantization
      Huffman
    Morphology
      Erosion
      Dilation
      Opening/Closing
    FFT
      O(N log N)
      Low-pass Filtering
    Neural Networks
      U-Net
      GANs
    Real World
      Daraz: Compression
      Pathao: Segmentation
      Ncell: Video Codec

Based on the TU BCA syllabus for Image Processing, unit 8.

Discussion

Loading…