Image ProcessingUnit 69 min read
Image Compression: Techniques, Standards & Real-World Applications
Unit 6 of Image Processing explores how to reduce image file sizes while preserving quality, covering lossless/lossy methods, compression standards (JPEG, PNG, GIF), and mathematical foundations like DCT and Huffman coding—with real-world examples from eSewa, Daraz, and Ncell.
TAKEAWAYS:
- Image compression trades off storage vs. quality using lossless (reversible) or lossy (irreversible) techniques.
- Standards like JPEG (DCT-based), PNG (lossless), and GIF (palette-based) dominate real-world use.
- DCT transforms (JPEG) and Huffman coding (entropy coding) are core mathematical tools.
- Rate-distortion theory explains why lossy compression sacrifices quality for smaller files.
- eSewa’s QR code generation and Daraz’s product thumbnails rely on compression to save bandwidth.
- Exam focus: Block diagrams, compression ratios, and distinguishing between standards.
Core Concepts: Why Compress Images?
Digital images are large because they store millions of pixels, each with color/brightness values. Compression reduces this data while keeping the image usable. Two key trade-offs:
- Lossless: No quality loss (e.g., PNG, ZIP files).
- Lossy: Sacrifices quality for smaller size (e.g., JPEG, MP3).
1. Mathematical Foundations
A. Pixel Relationships and Redundancy
Pixels are not independent: adjacent pixels often have similar colors (spatial redundancy) or repeated patterns (spectral redundancy). Compression exploits this.
Example: In a blue sky image, most pixels are similar shades of blue. Storing one "blue" value + small adjustments saves space.
B. Transform Coding (DCT)
The Discrete Cosine Transform (DCT) converts image blocks (e.g., 8×8 pixels) into frequency components. High-frequency components (edges, noise) are discarded in lossy compression.
**Visual**: A grayscale 8×8 block → its DCT coefficients (most energy in top-left corner).
Key Idea:
- Low-frequency coefficients (top-left) = smooth areas (sky, walls).
- High-frequency coefficients (bottom-right) = edges/noise (discarded in JPEG).
2. Compression Techniques
A. Lossless Methods
| Method | How It Works | Example Use Case |
|---|---|---|
| Run-Length Encoding (RLE) | Replaces sequences of identical pixels with (value, count). |
Fax machines, simple line art. |
| Huffman Coding | Assigns shorter codes to frequent pixel values. | PNG, ZIP files. |
| Lempel-Ziv-Welch (LZW) | Replaces repeated patterns with indices. | GIF, TIFF. |
Worked Example: Huffman Coding Given pixel frequencies:
| Pixel Value | Frequency |
|---|---|
| 0 | 45 |
| 1 | 12 |
| 2 | 18 |
| 3 | 25 |
Steps:
- Build Huffman tree (rarest first):
- Assign codes:
0(most frequent) →003→012→111(least frequent) →10
- Compressed output: Replace each pixel with its code.
- Original:
0, 0, 3, 2, 1→ Compressed:00 00 01 11 10.
- Original:
Compression Ratio:
- Original bits: 5 pixels × 8 bits = 40 bits.
- Compressed bits: 5 codes (avg 2 bits) = 10 bits.
- Ratio: 4:1.
B. Lossy Methods
| Method | How It Works | Example Use Case |
|---|---|---|
| DCT (JPEG) | Discards high-frequency DCT coefficients. | Photos, Daraz product images. |
| Wavelet (JPEG2000) | Multi-resolution analysis. | Medical imaging, archives. |
| Vector Quantization | Groups similar pixels into "codebooks". | Low-bitrate video. |
Worked Example: JPEG DCT Compression
- Divide image into 8×8 blocks.
- Apply DCT → get frequency coefficients.
- Quantize: Divide coefficients by a matrix (e.g.,
Q = [[16,11,...],[12,14,...]]). - Zigzag scan: Reorder coefficients for entropy coding.
- Huffman coding: Encode remaining data.
Visual:
Real-World Tie-In:
- Daraz’s product thumbnails use JPEG to balance file size (fast loading) and quality (recognizable products).
- Ncell’s mobile data plans save bandwidth by compressing images in WhatsApp/Instagram.
3. Compression Standards
| Standard | Type | Key Features | Use Cases |
|---|---|---|---|
| JPEG | Lossy | DCT + Huffman, 8-bit color. | Photos, web images. |
| PNG | Lossless | RLE + Huffman, 24-bit color. | Logos, screenshots, eSewa QR. |
| GIF | Lossless | Palette-based (256 colors), animation. | Simple animations, memes. |
| WebP | Hybrid | Lossy/lossless, smaller than JPEG. | Google Chrome, YouTube thumbnails. |
| HEIF | Lossy | High efficiency, Apple devices. | iPhone photos, NEPSE stock charts. |
Comparison Table:
| Feature | JPEG | PNG | GIF | WebP |
|---|---|---|---|---|
| Lossy? | Yes | No | No | Yes/No |
| Color Depth | 24-bit | 24/32-bit | 8-bit | 24/32-bit |
| Transparency | No | Yes | Yes | Yes |
| Animation | No | No | Yes | Yes |
| File Size | Small | Large | Small | Smallest |
Exam Tip: Memorize JPEG = lossy (photos), PNG = lossless (logos), GIF = animation.
4. Real-World Applications
A. eSewa’s QR Code Generation
- Compression Idea: QR codes use error correction (redundancy) and lossless encoding to fit data into a small grid.
- How:
- User data (e.g., payment details) is encoded into a binary matrix.
- Error correction blocks are added (like Huffman coding for robustness).
- The image is compressed to fit into the QR’s limited pixels.
flowchart LR
A["User Data"] --> B["Binary Encoding"]
B --> C["Error Correction<br/>(Hamming Codes)"]
C --> D["QR Grid<br/>(Compressed)"]
D --> E["Printed QR"]B. Daraz’s Image Search Optimization
- Problem: Millions of product images must load quickly.
- Solution:
- JPEG compression reduces file size by 80–90%.
- Resizing thumbnails to 300×300 pixels.
- CDN caching stores compressed versions globally.
Worked Example:
- Original image: 2048×1536 pixels × 24 bits = 9.2 MB.
- After JPEG (90% quality): ~1.2 MB.
- After resizing to 300×300: ~30 KB.
- Bandwidth saved: 99.7%!
C. Ncell’s Mobile Data Efficiency
- Challenge: Users on limited data plans.
- Techniques:
- Progressive JPEG: Loads low-res first, then sharpens.
- WebP: Smaller than JPEG for same quality.
- Lazy loading: Only compresses images when scrolled into view.
5. Advanced Topics
A. Rate-Distortion Theory
- Goal: Maximize compression while keeping distortion (quality loss) below a threshold.
- Formula:
Where:
- = bit rate (compression ratio).
- = distortion (MSE between original and compressed).
- = mutual information.
Visual:
B. Vector Quantization (VQ)
- Idea: Replace pixel values with indices to a codebook of typical patterns.
- Example: A codebook for skies might have 10 entries (e.g., "light blue," "dark blue").
- Used in: Low-bitrate video (e.g., old webcams).
Exam Tip: How to Score Full Marks
- Block Diagrams: Always draw the JPEG compression pipeline (DCT → Quantization → Huffman) when asked about standards.
- Calculations:
- For compression ratio, use:
- For Huffman coding, show the tree construction and code assignment.
- Differentiate Standards:
- JPEG: Lossy, DCT, photos.
- PNG: Lossless, RLE, logos.
- GIF: Palette-based, animation.
- Real-World Links:
- eSewa QR: Error correction + compression.
- Daraz images: JPEG + resizing.
- Ncell data: Progressive loading + WebP.
- Avoid Common Mistakes:
- ❌ Saying "JPEG is lossless" (it’s not!).
- ❌ Forgetting to mention quantization in DCT.
- ❌ Confusing Huffman (entropy coding) with RLE (run-length).
Final Visual Summary:
Based on the TU BCA syllabus for Image Processing, unit 6.
Discussion
Loading…