Multimedia SystemUnit 410 min read
Image & Video Compression: Techniques, Trade-offs & Real-world Use
Unit 4 of Multimedia System explores how image and video compression reduces file sizes while preserving quality, covering lossless (RLE, Huffman) and lossy (JPEG, MPEG) methods, perceptual models, and hardware acceleration. Includes comparisons, trade-offs, and Nepalese/global applications like eSewa’s video thumbnail
TAKEAWAYS:
- Lossy vs. lossless: Lossy compression (JPEG, MPEG) sacrifices minor data for huge size gains (e.g., 90% reduction), while lossless (RLE, Huffman) preserves 100% data but offers modest savings (e.g., 50%).
- Perceptual models: Humans ignore high-frequency details (e.g., fine textures in grass) and redundant colors—JPEG exploits this via DCT and quantization.
- Video compression: MPEG uses inter-frame prediction (I/P/B frames) to store only changes between frames, cutting bandwidth by 90%+ for 4K streams.
- Trade-offs: Higher compression → worse quality → more artifacts (blockiness, blur). Real-world examples: eSewa’s thumbnails use JPEG (lossy) for fast loading; Ncell’s video calls use MPEG (lossy) for smooth streaming.
- Hardware matters: GPUs/TPUs accelerate compression (e.g., YouTube’s VP9 codec runs on Google’s TPUs).
- Standards matter: JPEG for photos, MPEG-4 for videos, WebP for web—each balances speed, quality, and compatibility.
1. Why Compress Multimedia?
Multimedia files (images, videos) are huge:
- A 1-minute 4K video at 60 fps = ~1.5 GB uncompressed.
- A single 8MP photo = ~20 MB raw (16-bit color depth).
- Problem: Networks (e.g., NTC’s 10 Mbps), storage (e.g., Daraz’s servers), and devices (e.g., Pathao’s phones) can’t handle this.
Solution: Compression reduces size while keeping perceived quality intact. Key metrics:
- Compression ratio = Uncompressed size / Compressed size.
- PSNR (Peak Signal-to-Noise Ratio): Measures quality loss (higher = better).
- Bitrate: Bits per second (e.g., 10 Mbps for 4K Netflix).
2. Lossless Compression: No Data Lost
Definition: Reversibly reduces redundancy without losing any data. Used for medical images, CAD files, or archival photos.
A. Run-Length Encoding (RLE)
How it works:
- Replace consecutive identical values with a count + value pair.
- Example:
"aaaaabbbbcccdde"→(5,a)(4,b)(3,c)(2,d)(1,e).
Worked Example: Traffic Light Sequence
Assume a 10-second traffic light cycle recorded as RGB values every second:
Red(255,0,0), Red(255,0,0), Green(0,255,0), ...,
RLE compresses repeated Red frames into (2,Red) → 80% smaller.
Limitations:
- Only works for simple patterns (e.g., fax machines, BMP files).
- Not effective for complex images (e.g., photos).
graph LR
A["Uncompressed: R R R G G B B"] -->|"RLE"| B["(3,R)(2,G)(2,B)"]
B -->|"Decompress"| AB. Huffman Coding
How it works:
- Assign shorter codes to frequent symbols, longer to rare ones.
- Example: For
"aaaaabbbbcccdde":a(5x) →0,b(4x) →10,c(3x) →110,d(2x) →1110,e(1x) →1111.
Advantages:
- Works for any data (text, audio, images).
- Optimal for lossless compression (theoretical limit: entropy).
Disadvantage:
- Slow for real-time apps (e.g., video calls).
Real-world use:
- PNG images (lossless web format) use Huffman + LZ77.
- ZIP files combine Huffman with LZ77.
3. Lossy Compression: Sacrifice Quality for Size
Definition: Permanently discards perceptually irrelevant data. Used for photos (JPEG), videos (MPEG), music (MP3).
A. JPEG for Images
How it works (8-step process):
- Color space conversion: RGB → YCbCr (luma + chroma).
- Humans see luma (Y) better than color (CbCr), so chroma is subsampled (e.g., 4:2:0).
- Downsampling: Reduce resolution (e.g., 8MP → 2MP).
- DCT (Discrete Cosine Transform): Split image into 8×8 blocks and convert to frequency domain.
- High frequencies = textures/details (e.g., grass blades). . Quantization: Discard high-frequency coefficients (e.g., keep only top 10%).
- Zig-zag scan: Reorder coefficients for Huffman coding.
- Entropy coding: Huffman/RLE on quantized data.
- Huffman coding: Assign short codes to frequent coefficients.
Artifacts:
- Blocking: Visible 8×8 grid at high compression.
- Blurring: Loss of high-frequency details.
Worked Example: eSewa Thumbnail
- Original: 5MB PNG (lossless).
- JPEG at 80% quality: 500 KB (90% smaller).
- Trade-off: Slight blur but loads instantly on mobile.
flowchart TD
A["RGB Image"] --> B["YCbCr\n(Subsample CbCr)"]
B --> C["8×8 DCT Blocks"]
C --> D["Quantize\n(Discard high freq)"]
D --> E["Zig-zag Scan"]
E --> F["Huffman\nEncode"]
F --> G["JPEG File"]B. MPEG for Video
How it works:
- Temporal compression: Store only changes between frames (unlike JPEG, which treats each frame independently).
- Frame types:
- I-frame: Full frame (like JPEG), used as reference.
- P-frame: Stores differences from previous I/P frame (predictive).
- B-frame: Stores differences from past/future frames (bidirectional).
- Motion compensation: Track object movement (e.g., a car in traffic) and store only the delta.
- DCT + Quantization: Same as JPEG for each frame.
Example: Daraz Streaming
- Uncompressed: 1080p60 = 1.5 Gbps.
- MPEG-4 (H.264): 5 Mbps (300× smaller).
- Why? Uses I/P/B frames + motion compensation.
sequenceDiagram
participant I as I-Frame
participant P as P-Frame
participant B as B-Frame
I->>P: "Store changes from me"
P->>B: "Store changes from I/P"
B->>P: "Bidirectional\n(uses future frame)"4. Perceptual Models: Why Lossy Works
Humans can’t perceive all details:
| Feature | Perceptual Limit | Exploited by... |
|---|---|---|
| Color resolution | 8–10 bits per channel (24-bit RGB) | JPEG’s 4:2:0 subsampling |
| Spatial resolution | ~1 arcminute (30 cycles/degree) | DCT’s high-frequency discard |
| Temporal resolution | ~24 fps (film standard) | MPEG’s motion compensation |
| Loudness | 120 dB dynamic range (but we hear 0–120) | MP3’s psychoacoustics |
Example: Kathmandu Traffic Video
- Uncompressed: 4K @ 30 fps = 10 Gbps.
- MPEG-4 (H.265): 2 Mbps (5000× smaller).
- Why? Motion compensation ignores tiny car movements; DCT blurs distant buildings.
5. Real-World Applications in Nepal
| Company/Product | Compression Technique | Use Case | Why It Matters |
|---|---|---|---|
| eSewa | JPEG (lossy) | User profile photos, receipt thumbnails | Faster loading on low-bandwidth networks |
| Ncell Video Call | H.264 (MPEG-4) | Real-time 720p streaming | Reduces data usage by 95% |
| Daraz Live Stream | H.265 (HEVC) | Product demos | 50% smaller than H.264 at same quality |
| NEPSE Stock Charts | PNG (lossless) + WebP | High-frequency trading graphs | Preserves all data for analysis |
| WhatsApp Status | VP9 (Google’s codec) | 1080p videos | Balances quality and upload speed |
Worked Example: Pathao Driver’s Video Feed
- Problem: Driver’s phone must stream 360p @ 15 fps to the app in real-time.
- Solution: H.264 with GOP (Group of Pictures) every 1 second (I-frame + 2 P-frames).
- Result: 500 Kbps instead of 10 Mbps → works on 2G networks.
6. Trade-offs: Lossy vs. Lossless
| Criteria | Lossless (e.g., PNG, ZIP) | Lossy (e.g., JPEG, MPEG) |
|---|---|---|
| Compression | 2:1 to 5:1 | 10:1 to 100:1 |
| Quality | 100% original | Artifacts (blocking, blur) |
| Speed | Slower (CPU-intensive) | Faster (hardware-accelerated) |
| Use Case | Medical images, CAD, archives | Photos, videos, music, web |
| Reversible? | Yes | No |
Example: Bank Cheque Scanning
- Lossy (JPEG): Blurs text → invalid for processing.
- Lossless (PNG): Preserves all details → used by Nabil Bank’s cheque readers.
7. Hardware Acceleration
Modern devices use GPUs/TPUs to speed up compression:
- NVIDIA NVENC: Encodes 4K60 H.264 in real-time (used in YouTube Live).
- Google TPUs: Run VP9 for YouTube videos (faster than CPUs).
- Apple A-series chips: Hardware-accelerate HEVC (used in iPhone videos).
Exam Tip
Compare JPEG vs. MPEG:
- JPEG: Single-frame, spatial compression (DCT).
- MPEG: Multi-frame, temporal + spatial (I/P/B frames).
- Exam trick: Always mention perceptual models (e.g., "MPEG exploits temporal redundancy between frames").
Trade-offs are key:
- Why lossy for 4K video? → "Bandwidth constraints (e.g., NTC’s 10 Mbps) require 90%+ compression; lossless would need 100 Mbps."
- Why lossless for medical images? → "No data loss allowed (e.g., tumor detection)."
Worked examples:
- For RLE, use a traffic light sequence or fax text.
- For JPEG, trace an 8×8 DCT block with quantization.
- For MPEG, draw an I/P/B frame timeline.
Real-world ties:
- eSewa/Khalti: "JPEG thumbnails for fast loading."
- Ncell/Pathao: "MPEG-4 for low-bandwidth streaming."
- Daraz: "HEVC for high-quality product videos."
Avoid common mistakes:
- ❌ "JPEG is lossless." → It’s lossy!
- ❌ "MPEG stores every frame." → It uses I/P/B frames.
- ❌ Ignoring perceptual models → Always mention "human vision limits."
Quick Revision Table
| Technique | Type | Key Idea | Example Use Case |
|---|---|---|---|
| RLE | Lossless | Replace repeated data | Fax machines, BMP files |
| Huffman | Lossless | Short codes for frequent symbols | PNG, ZIP files |
| JPEG | Lossy | DCT + quantization | Photos, web images |
| MPEG | Lossy | I/P/B frames + motion compensation | Videos, YouTube, Ncell calls |
| WebP | Lossy/Less | Combines JPEG + PNG | Google’s web images |
Based on the TU BCA syllabus for Multimedia System (CACS457), unit 4.
Discussion
Loading…