Elective Image Processing

Image ProcessingUnit 69 min read

Image Compression: Techniques, Standards & Real-World Applications

Unit 6 of Image Processing explores how to reduce image file sizes while preserving quality, covering lossless/lossy methods, compression standards (JPEG, PNG, GIF), and mathematical foundations like DCT and Huffman coding—with real-world examples from eSewa, Daraz, and Ncell.

TAKEAWAYS:

  • Image compression trades off storage vs. quality using lossless (reversible) or lossy (irreversible) techniques.
  • Standards like JPEG (DCT-based), PNG (lossless), and GIF (palette-based) dominate real-world use.
  • DCT transforms (JPEG) and Huffman coding (entropy coding) are core mathematical tools.
  • Rate-distortion theory explains why lossy compression sacrifices quality for smaller files.
  • eSewa’s QR code generation and Daraz’s product thumbnails rely on compression to save bandwidth.
  • Exam focus: Block diagrams, compression ratios, and distinguishing between standards.

Core Concepts: Why Compress Images?

Digital images are large because they store millions of pixels, each with color/brightness values. Compression reduces this data while keeping the image usable. Two key trade-offs:

  1. Lossless: No quality loss (e.g., PNG, ZIP files).
  2. Lossy: Sacrifices quality for smaller size (e.g., JPEG, MP3).
0255075100Raw Image (8-bit RGB)100PNG (Lossless)85JPEG (High Quality)60JPEG (Low Quality)20File Size (relative %)
Typical size reduction from raw to compressed formats (8-bit RGB example)
Pixel ValuesRaw Image DataCompressionLossless (PNG, TIFF)Lossy (JPEG, MP3)Reversible (No Quality Loss)Irreversible (Quality Trade-off)Medical/Documents (eSewa QR)Photos/Videos (Daraz Thumbnails)
Trade-offs in image compression: Lossless vs. Lossy methods with local examples

1. Mathematical Foundations

A. Pixel Relationships and Redundancy

Pixels are not independent: adjacent pixels often have similar colors (spatial redundancy) or repeated patterns (spectral redundancy). Compression exploits this.

Example: In a blue sky image, most pixels are similar shades of blue. Storing one "blue" value + small adjustments saves space.

B. Transform Coding (DCT)

The Discrete Cosine Transform (DCT) converts image blocks (e.g., 8×8 pixels) into frequency components. High-frequency components (edges, noise) are discarded in lossy compression.

**Visual**: A grayscale 8×8 block → its DCT coefficients (most energy in top-left corner).

Key Idea:

  • Low-frequency coefficients (top-left) = smooth areas (sky, walls).
  • High-frequency coefficients (bottom-right) = edges/noise (discarded in JPEG).

2. Compression Techniques

A. Lossless Methods

Method How It Works Example Use Case
Run-Length Encoding (RLE) Replaces sequences of identical pixels with (value, count). Fax machines, simple line art.
Huffman Coding Assigns shorter codes to frequent pixel values. PNG, ZIP files.
Lempel-Ziv-Welch (LZW) Replaces repeated patterns with indices. GIF, TIFF.

Worked Example: Huffman Coding Given pixel frequencies:

Pixel Value Frequency
0 45
1 12
2 18
3 25

Steps:

  1. Build Huffman tree (rarest first):
    
    
  2. Assign codes:
    • 0 (most frequent) → 00
    • 3 → 01
    • 2 → 11
    • 1 (least frequent) → 10
  3. Compressed output: Replace each pixel with its code.
    • Original: 0, 0, 3, 2, 1 → Compressed: 00 00 01 11 10.

Compression Ratio:

  • Original bits: 5 pixels × 8 bits = 40 bits.
  • Compressed bits: 5 codes (avg 2 bits) = 10 bits.
  • Ratio: 4:1.

B. Lossy Methods

Method How It Works Example Use Case
DCT (JPEG) Discards high-frequency DCT coefficients. Photos, Daraz product images.
Wavelet (JPEG2000) Multi-resolution analysis. Medical imaging, archives.
Vector Quantization Groups similar pixels into "codebooks". Low-bitrate video.

Worked Example: JPEG DCT Compression

  1. Divide image into 8×8 blocks.
  2. Apply DCT → get frequency coefficients.
  3. Quantize: Divide coefficients by a matrix (e.g., Q = [[16,11,...],[12,14,...]]).
  4. Zigzag scan: Reorder coefficients for entropy coding.
  5. Huffman coding: Encode remaining data.

Visual:


Real-World Tie-In:

  • Daraz’s product thumbnails use JPEG to balance file size (fast loading) and quality (recognizable products).
  • Ncell’s mobile data plans save bandwidth by compressing images in WhatsApp/Instagram.

3. Compression Standards

Standard Type Key Features Use Cases
JPEG Lossy DCT + Huffman, 8-bit color. Photos, web images.
PNG Lossless RLE + Huffman, 24-bit color. Logos, screenshots, eSewa QR.
GIF Lossless Palette-based (256 colors), animation. Simple animations, memes.
WebP Hybrid Lossy/lossless, smaller than JPEG. Google Chrome, YouTube thumbnails.
HEIF Lossy High efficiency, Apple devices. iPhone photos, NEPSE stock charts.

Comparison Table:

Feature JPEG PNG GIF WebP
Lossy? Yes No No Yes/No
Color Depth 24-bit 24/32-bit 8-bit 24/32-bit
Transparency No Yes Yes Yes
Animation No No Yes Yes
File Size Small Large Small Smallest

Exam Tip: Memorize JPEG = lossy (photos), PNG = lossless (logos), GIF = animation.


4. Real-World Applications

A. eSewa’s QR Code Generation

  • Compression Idea: QR codes use error correction (redundancy) and lossless encoding to fit data into a small grid.
  • How:
    1. User data (e.g., payment details) is encoded into a binary matrix.
    2. Error correction blocks are added (like Huffman coding for robustness).
    3. The image is compressed to fit into the QR’s limited pixels.
flowchart LR
    A["User Data"] --> B["Binary Encoding"]
    B --> C["Error Correction<br/>(Hamming Codes)"]
    C --> D["QR Grid<br/>(Compressed)"]
    D --> E["Printed QR"]

B. Daraz’s Image Search Optimization

  • Problem: Millions of product images must load quickly.
  • Solution:
    • JPEG compression reduces file size by 80–90%.
    • Resizing thumbnails to 300×300 pixels.
    • CDN caching stores compressed versions globally.
Image DatabaseFeature ExtractionIndexing (Hashing)QueryResults
Daraz’s approximate nearest-neighbor search pipeline for product images

Worked Example:

  • Original image: 2048×1536 pixels × 24 bits = 9.2 MB.
  • After JPEG (90% quality): ~1.2 MB.
  • After resizing to 300×300: ~30 KB.
  • Bandwidth saved: 99.7%!

C. Ncell’s Mobile Data Efficiency

  • Challenge: Users on limited data plans.
  • Techniques:
    • Progressive JPEG: Loads low-res first, then sharpens.
    • WebP: Smaller than JPEG for same quality.
    • Lazy loading: Only compresses images when scrolled into view.

5. Advanced Topics

A. Rate-Distortion Theory

  • Goal: Maximize compression while keeping distortion (quality loss) below a threshold.
  • Formula: Where:
    • = bit rate (compression ratio).
    • = distortion (MSE between original and compressed).
    • = mutual information.

Visual:


B. Vector Quantization (VQ)

  • Idea: Replace pixel values with indices to a codebook of typical patterns.
  • Example: A codebook for skies might have 10 entries (e.g., "light blue," "dark blue").
  • Used in: Low-bitrate video (e.g., old webcams).

Exam Tip: How to Score Full Marks

  1. Block Diagrams: Always draw the JPEG compression pipeline (DCT → Quantization → Huffman) when asked about standards.
  2. Calculations:
    • For compression ratio, use:
    • For Huffman coding, show the tree construction and code assignment.
  3. Differentiate Standards:
    • JPEG: Lossy, DCT, photos.
    • PNG: Lossless, RLE, logos.
    • GIF: Palette-based, animation.
  4. Real-World Links:
    • eSewa QR: Error correction + compression.
    • Daraz images: JPEG + resizing.
    • Ncell data: Progressive loading + WebP.
  5. Avoid Common Mistakes:
    • ❌ Saying "JPEG is lossless" (it’s not!).
    • ❌ Forgetting to mention quantization in DCT.
    • ❌ Confusing Huffman (entropy coding) with RLE (run-length).

Final Visual Summary:


Based on the TU BCA syllabus for Image Processing, unit 6.

Discussion

Loading…