CACS457 Multimedia System

Multimedia SystemUnit 410 min read

Image & Video Compression: Techniques, Trade-offs & Real-world Use

Unit 4 of Multimedia System explores how image and video compression reduces file sizes while preserving quality, covering lossless (RLE, Huffman) and lossy (JPEG, MPEG) methods, perceptual models, and hardware acceleration. Includes comparisons, trade-offs, and Nepalese/global applications like eSewa’s video thumbnail

TAKEAWAYS:

  • Lossy vs. lossless: Lossy compression (JPEG, MPEG) sacrifices minor data for huge size gains (e.g., 90% reduction), while lossless (RLE, Huffman) preserves 100% data but offers modest savings (e.g., 50%).
  • Perceptual models: Humans ignore high-frequency details (e.g., fine textures in grass) and redundant colors—JPEG exploits this via DCT and quantization.
  • Video compression: MPEG uses inter-frame prediction (I/P/B frames) to store only changes between frames, cutting bandwidth by 90%+ for 4K streams.
  • Trade-offs: Higher compression → worse quality → more artifacts (blockiness, blur). Real-world examples: eSewa’s thumbnails use JPEG (lossy) for fast loading; Ncell’s video calls use MPEG (lossy) for smooth streaming.
  • Hardware matters: GPUs/TPUs accelerate compression (e.g., YouTube’s VP9 codec runs on Google’s TPUs).
  • Standards matter: JPEG for photos, MPEG-4 for videos, WebP for web—each balances speed, quality, and compatibility.

1. Why Compress Multimedia?

Multimedia files (images, videos) are huge:

  • A 1-minute 4K video at 60 fps = ~1.5 GB uncompressed.
  • A single 8MP photo = ~20 MB raw (16-bit color depth).
  • Problem: Networks (e.g., NTC’s 10 Mbps), storage (e.g., Daraz’s servers), and devices (e.g., Pathao’s phones) can’t handle this.

Solution: Compression reduces size while keeping perceived quality intact. Key metrics:

  • Compression ratio = Uncompressed size / Compressed size.
  • PSNR (Peak Signal-to-Noise Ratio): Measures quality loss (higher = better).
  • Bitrate: Bits per second (e.g., 10 Mbps for 4K Netflix).

2. Lossless Compression: No Data Lost

Definition: Reversibly reduces redundancy without losing any data. Used for medical images, CAD files, or archival photos.

A. Run-Length Encoding (RLE)

How it works:

  1. Replace consecutive identical values with a count + value pair.
  2. Example: "aaaaabbbbcccdde" → (5,a)(4,b)(3,c)(2,d)(1,e).

Worked Example: Traffic Light Sequence Assume a 10-second traffic light cycle recorded as RGB values every second: Red(255,0,0), Red(255,0,0), Green(0,255,0), ..., RLE compresses repeated Red frames into (2,Red) → 80% smaller.

Limitations:

  • Only works for simple patterns (e.g., fax machines, BMP files).
  • Not effective for complex images (e.g., photos).
graph LR
    A["Uncompressed: R R R G G B B"] -->|"RLE"| B["(3,R)(2,G)(2,B)"]
    B -->|"Decompress"| A

B. Huffman Coding

How it works:

  1. Assign shorter codes to frequent symbols, longer to rare ones.
  2. Example: For "aaaaabbbbcccdde":
    • a (5x) → 0, b (4x) → 10, c (3x) → 110, d (2x) → 1110, e (1x) → 1111.

Advantages:

  • Works for any data (text, audio, images).
  • Optimal for lossless compression (theoretical limit: entropy).

Disadvantage:

  • Slow for real-time apps (e.g., video calls).

Real-world use:

  • PNG images (lossless web format) use Huffman + LZ77.
  • ZIP files combine Huffman with LZ77.

3. Lossy Compression: Sacrifice Quality for Size

Definition: Permanently discards perceptually irrelevant data. Used for photos (JPEG), videos (MPEG), music (MP3).

A. JPEG for Images

How it works (8-step process):

  1. Color space conversion: RGB → YCbCr (luma + chroma).
    • Humans see luma (Y) better than color (CbCr), so chroma is subsampled (e.g., 4:2:0).
  2. Downsampling: Reduce resolution (e.g., 8MP → 2MP).
  3. DCT (Discrete Cosine Transform): Split image into 8×8 blocks and convert to frequency domain.
    • High frequencies = textures/details (e.g., grass blades). . Quantization: Discard high-frequency coefficients (e.g., keep only top 10%).
  4. Zig-zag scan: Reorder coefficients for Huffman coding.
  5. Entropy coding: Huffman/RLE on quantized data.
  6. Huffman coding: Assign short codes to frequent coefficients.

Artifacts:

  • Blocking: Visible 8×8 grid at high compression.
  • Blurring: Loss of high-frequency details.

Worked Example: eSewa Thumbnail

  • Original: 5MB PNG (lossless).
  • JPEG at 80% quality: 500 KB (90% smaller).
  • Trade-off: Slight blur but loads instantly on mobile.
flowchart TD
    A["RGB Image"] --> B["YCbCr\n(Subsample CbCr)"]
    B --> C["8×8 DCT Blocks"]
    C --> D["Quantize\n(Discard high freq)"]
    D --> E["Zig-zag Scan"]
    E --> F["Huffman\nEncode"]
    F --> G["JPEG File"]

B. MPEG for Video

How it works:

  1. Temporal compression: Store only changes between frames (unlike JPEG, which treats each frame independently).
  2. Frame types:
    • I-frame: Full frame (like JPEG), used as reference.
    • P-frame: Stores differences from previous I/P frame (predictive).
    • B-frame: Stores differences from past/future frames (bidirectional).
  3. Motion compensation: Track object movement (e.g., a car in traffic) and store only the delta.
  4. DCT + Quantization: Same as JPEG for each frame.

Example: Daraz Streaming

  • Uncompressed: 1080p60 = 1.5 Gbps.
  • MPEG-4 (H.264): 5 Mbps (300× smaller).
  • Why? Uses I/P/B frames + motion compensation.
sequenceDiagram
    participant I as I-Frame
    participant P as P-Frame
    participant B as B-Frame
    I->>P: "Store changes from me"
    P->>B: "Store changes from I/P"
    B->>P: "Bidirectional\n(uses future frame)"

4. Perceptual Models: Why Lossy Works

Humans can’t perceive all details:

Feature Perceptual Limit Exploited by...
Color resolution 8–10 bits per channel (24-bit RGB) JPEG’s 4:2:0 subsampling
Spatial resolution ~1 arcminute (30 cycles/degree) DCT’s high-frequency discard
Temporal resolution ~24 fps (film standard) MPEG’s motion compensation
Loudness 120 dB dynamic range (but we hear 0–120) MP3’s psychoacoustics

Example: Kathmandu Traffic Video

  • Uncompressed: 4K @ 30 fps = 10 Gbps.
  • MPEG-4 (H.265): 2 Mbps (5000× smaller).
  • Why? Motion compensation ignores tiny car movements; DCT blurs distant buildings.

5. Real-World Applications in Nepal

Company/Product Compression Technique Use Case Why It Matters
eSewa JPEG (lossy) User profile photos, receipt thumbnails Faster loading on low-bandwidth networks
Ncell Video Call H.264 (MPEG-4) Real-time 720p streaming Reduces data usage by 95%
Daraz Live Stream H.265 (HEVC) Product demos 50% smaller than H.264 at same quality
NEPSE Stock Charts PNG (lossless) + WebP High-frequency trading graphs Preserves all data for analysis
WhatsApp Status VP9 (Google’s codec) 1080p videos Balances quality and upload speed

Worked Example: Pathao Driver’s Video Feed

  • Problem: Driver’s phone must stream 360p @ 15 fps to the app in real-time.
  • Solution: H.264 with GOP (Group of Pictures) every 1 second (I-frame + 2 P-frames).
  • Result: 500 Kbps instead of 10 Mbps → works on 2G networks.

6. Trade-offs: Lossy vs. Lossless

Criteria Lossless (e.g., PNG, ZIP) Lossy (e.g., JPEG, MPEG)
Compression 2:1 to 5:1 10:1 to 100:1
Quality 100% original Artifacts (blocking, blur)
Speed Slower (CPU-intensive) Faster (hardware-accelerated)
Use Case Medical images, CAD, archives Photos, videos, music, web
Reversible? Yes No

Example: Bank Cheque Scanning

  • Lossy (JPEG): Blurs text → invalid for processing.
  • Lossless (PNG): Preserves all details → used by Nabil Bank’s cheque readers.

7. Hardware Acceleration

Modern devices use GPUs/TPUs to speed up compression:

  • NVIDIA NVENC: Encodes 4K60 H.264 in real-time (used in YouTube Live).
  • Google TPUs: Run VP9 for YouTube videos (faster than CPUs).
  • Apple A-series chips: Hardware-accelerate HEVC (used in iPhone videos).

Exam Tip

  1. Compare JPEG vs. MPEG:

    • JPEG: Single-frame, spatial compression (DCT).
    • MPEG: Multi-frame, temporal + spatial (I/P/B frames).
    • Exam trick: Always mention perceptual models (e.g., "MPEG exploits temporal redundancy between frames").
  2. Trade-offs are key:

    • Why lossy for 4K video? → "Bandwidth constraints (e.g., NTC’s 10 Mbps) require 90%+ compression; lossless would need 100 Mbps."
    • Why lossless for medical images? → "No data loss allowed (e.g., tumor detection)."
  3. Worked examples:

    • For RLE, use a traffic light sequence or fax text.
    • For JPEG, trace an 8×8 DCT block with quantization.
    • For MPEG, draw an I/P/B frame timeline.
  4. Real-world ties:

    • eSewa/Khalti: "JPEG thumbnails for fast loading."
    • Ncell/Pathao: "MPEG-4 for low-bandwidth streaming."
    • Daraz: "HEVC for high-quality product videos."
  5. Avoid common mistakes:

    • ❌ "JPEG is lossless." → It’s lossy!
    • ❌ "MPEG stores every frame." → It uses I/P/B frames.
    • ❌ Ignoring perceptual models → Always mention "human vision limits."

Quick Revision Table

Technique Type Key Idea Example Use Case
RLE Lossless Replace repeated data Fax machines, BMP files
Huffman Lossless Short codes for frequent symbols PNG, ZIP files
JPEG Lossy DCT + quantization Photos, web images
MPEG Lossy I/P/B frames + motion compensation Videos, YouTube, Ncell calls
WebP Lossy/Less Combines JPEG + PNG Google’s web images

Based on the TU BCA syllabus for Multimedia System (CACS457), unit 4.

Discussion

Loading…