CACS457 Multimedia System

Multimedia SystemUnit 311 min read

Audio Fundamentals, Compression & Real-World Applications

Unit 3 of Multimedia System explores how audio signals are digitized, compressed (lossy vs. lossless), and optimized for storage/transmission, with comparisons to real-world apps like WhatsApp voice notes and Ncell music streaming.

TAKEAWAYS

  • Audio compression reduces file size by discarding redundant or imperceptible data (lossy) or using mathematical patterns (lossless).
  • Sampling rate (e.g., 44.1 kHz) and bit depth (e.g., 16-bit) determine audio quality—higher values mean better fidelity but larger files.
  • MP3 (lossy) and FLAC (lossless) are industry standards for music, while AAC dominates streaming (Apple Music, Spotify).
  • MIDI synthesizes sound via instructions (small files) but lacks the realism of recorded audio.
  • Perceptual coding exploits human hearing limitations (e.g., masking) to improve compression efficiency.
  • Real-world trade-offs: WhatsApp uses Opus (low-latency, adaptive bitrate) for calls, while Ncell’s music app prioritizes AAC for balance between quality and data savings.

1. Audio Fundamentals: From Analog to Digital

Audio is a continuous analog signal (sound waves) that must be converted to digital form for computers to process. This involves three key steps:

1.1 Analog-to-Digital Conversion (ADC)

  • Sampling: Capturing the amplitude of the analog signal at discrete intervals (measured in Hz).
    • Example: A 44.1 kHz sampling rate means 44,100 samples per second.
    • Nyquist Theorem: Sampling rate must be ≥ 2 × highest frequency in the signal to avoid aliasing.
Analog WaveContinuous soundwave (e.g., microphoneSampling (44.1 kHz)44,100 discreteamplitude measurementsQuantizationAmplitude roundedto binary values (e.g.Digital FileEncoded as binary(e.g., WAV, MP3)
Step-by-step ADC process with real-world sampling rate example
  • Quantization: Assigning numerical values to each sample (e.g., 16-bit = 65,536 possible levels).
  • Encoding: Storing samples as binary data (e.g., WAV, AIFF).

1.2 Key Audio Parameters

Parameter Definition Example Values Impact on Quality/File Size
Sampling Rate Samples per second (Hz) 44.1 kHz, 48 kHz Higher = better fidelity
Bit Depth Bits per sample (resolution) 16-bit, 24-bit Higher = less noise
Channels Mono vs. Stereo vs. Surround 1 (Mono), 2 (Stereo) More channels = larger files
44.1 kHz (CD quality)48 kHz (video sync)Sampling Rate (Hz)16-bit (standard)24-bit (professional)Bit Depth (bits)MonoStereoSurround (5.1+)ChannelsAudio Parameters
Hierarchy of audio quality parameters with common values

Worked Example: CD Audio vs. MP3

  • A CD-quality audio file:
    • Sampling rate: 44.1 kHz
    • Bit depth: 16-bit
    • Channels: 2 (Stereo)
    • Uncompressed size: ~10 MB per minute (700 KB/sec).
  • An MP3 (compressed) version of the same audio might be ~1 MB/minute (100 KB/sec).

2. Audio Compression Techniques

Compression reduces file size by removing redundancy or irrelevant data. Two main types:

2.1 Lossless Compression

  • No data loss; original audio can be perfectly reconstructed.
  • Methods:
    • Run-Length Encoding (RLE): Replaces repeated values with a count. Example: "aaaaabbbbcccdde" → (5a,4b,3c,2d,e) → Compression ratio = 12/10 = 1.2 (20% reduction).
    • Dictionary Methods: Stores repeated patterns once (e.g., FLAC, ALAC).
    • Entropy Coding: Assigns shorter codes to frequent values (e.g., Huffman coding).
  • Use Cases: Archival audio, professional editing (FLAC, WMA Lossless).

2.2 Lossy Compression

  • Removes imperceptible or redundant data; irreversible but achieves high compression ratios.
  • Key Techniques:
    • Perceptual Coding: Exploits how humans hear:
      • Masking: Loud sounds hide quiet ones (e.g., a drum drowns out a hi-hat).
      • Frequency Irrelevance: Humans don’t hear ultra-high/low frequencies well.
    • ** Psychoacoustic Models**: Analyzes audio to discard inaudible components.
    • Hybrid Approaches: Combines time-domain and frequency-domain analysis (e.g., MP3, AAC).
06.2512.518.7525Critical Band 1 (20–100 Hz)10Critical Band 2 (100–200 Hz)15Critical Band 3 (200–500 Hz)25Masked Frequencies0
MP3 masking effect: Loud bass (Critical Band 3) masks high frequencies (0 dB inaudible)
Lossy Format Compression Ratio Use Case Example Apps
MP3 10:1 to 12:1 Music streaming Spotify, Gaana
AAC 7:1 to 10:1 Apple Music, YouTube iTunes, Netflix
Opus 2:1 to 10:1 Voice/video calls WhatsApp, Zoom
Vorbis 5:1 to 8:1 Open-source audio YouTube (WebM)

Worked Example: MP3 Compression in Ncell Music

  • Original WAV: 10 MB/minute (44.1 kHz, 16-bit, stereo).
  • MP3 (128 kbps): ~1.6 MB/minute (80% smaller).
  • Trade-off: Slight loss in high-frequency details (inaudible to most listeners).

3. Specialized Audio Formats

3.1 MIDI (Musical Instrument Digital Interface)

  • Not audio data; stores instructions (notes, timing, instruments) for synthesizers.
  • Advantages:
    • Tiny file sizes (e.g., a 5-minute song = ~50 KB).
    • Editable (change instruments, tempo without re-recording).
  • Disadvantages:
    • No "real" audio; relies on a sound font/synthesizer.
    • Poor for complex sounds (e.g., vocals, live recordings).
  • Use Cases: Video game soundtracks, ringtone editors, karaoke apps.

Comparison: MIDI vs. Digital Audio for Video Games

Feature MIDI Digital Audio (WAV/MP3)
File Size ~50 KB for 5 min ~10 MB (WAV), ~1 MB (MP3)
Quality Synthetic, limited instruments High-fidelity, realistic
Editability Fully editable (notes, BPM) Fixed recording
Latency Low (instant playback) Higher (file loading)
Best For Background music, chiptune Voiceovers, sound effects

Real-World Example: Pathao’s In-App Music

  • Uses AAC (lossy) for background music in the driver app to:
    • Save mobile data (critical for low-bandwidth users).
    • Balance quality and file size (~1 MB per song).

4. Real-World Applications of Audio Compression

4.1 WhatsApp Voice Messages (Opus Codec)

  • Why Opus?
    • Adaptive bitrate: Adjusts quality based on network speed.
    • Low latency: Optimized for real-time calls (~200 ms delay).
    • Compression: Reduces a 1-minute voice note from ~10 MB (WAV) to ~500 KB.
  • How It Works:
    1. Audio sampled at 48 kHz.
    2. Opus applies perceptual coding to discard inaudible frequencies.
    3. Encrypted and sent over WhatsApp’s servers.

4.2 Ncell Music Streaming (AAC)

  • Challenge: Nepal’s average internet speed (~10 Mbps) struggles with high-bitrate audio.
  • Solution:
    • Uses AAC at 128–192 kbps (balance of quality and data).
    • Adaptive streaming: Switches between 64 kbps (slow network) and 320 kbps (fast network).
  • Result: A 3-minute song downloads in ~1–2 MB vs. ~30 MB uncompressed.

4.3 eSewa’s IVR System (MIDI + Text-to-Speech)

  • Problem: Storing pre-recorded voice prompts for transactions (e.g., "Enter your PIN") wastes space.
  • Solution:
    • Uses MIDI-like synthesized speech for static prompts.
    • Dynamic prompts (e.g., balance updates) use low-bitrate MP3 snippets.
  • Savings: Reduces storage by 90% compared to WAV files.
User InputDTMF tones (e.g.,'1' for balance)MIDI ProcessingConverts todigital commandsTTS EngineGenerates speech(e.g., 'Your balance iOutputPlayed via IVRsystem
eSewa IVR workflow combining MIDI and TTS

5. Why Compression Matters for Video and Animation

5.1 Data Volume in Multimedia

  • 1 minute of uncompressed video:
    • 4K @ 60 FPS: ~1.5 GB (requires 150 Mbps bandwidth).
    • Audio (stereo, 48 kHz): ~10 MB/minute.
  • Compressed (H.264 + AAC):
    • ~50 MB for the same video (97% smaller).

5.2 Compression Techniques in Video

Technique Lossy/Lossless How It Works Example Formats
Intraframe Lossy Compresses each frame independently JPEG, PNG
Interframe Lossy Exploits similarity between frames MPEG, H.264
Motion Compensation Lossy Tracks moving objects across frames MP4, WebM
Discrete Cosine Transform (DCT) Lossy Converts spatial data to frequency domain JPEG, MP3

Worked Example: Daraz’s Product Videos

  • Problem: Sellers upload 1-minute videos (1080p, 30 FPS) → ~500 MB uncompressed.
  • Solution: Daraz’s platform compresses to ~20 MB using:
    • H.264 (interframe compression).
    • AAC audio (128 kbps).
  • Result: Faster uploads, lower bandwidth for viewers.

Exam Tip

  1. Lossy vs. Lossless:

    • Lossy (MP3, AAC) = high compression, irreversible, used for music/streaming.
    • Lossless (FLAC, WAV) = no quality loss, used for archiving.
    • Exam trick: Always mention perceptual coding for lossy formats.
  2. Sampling Rate/Bit Depth:

    • CD quality = 44.1 kHz, 16-bit.
    • Broadcast radio = 48 kHz, 24-bit.
    • Higher values = better quality but larger files.
  3. Real-World Scenarios:

    • WhatsApp: Opus codec for calls.
    • Ncell Music: AAC for streaming.
    • eSewa IVR: MIDI/text-to-speech for prompts.
    • Tip: Relate numerical examples to these apps (e.g., "AAC at 128 kbps reduces a song to 1 MB").
  4. Trade-offs:

    • MIDI vs. Digital Audio: MIDI saves space but lacks realism.
    • JPEG vs. MPEG: MPEG uses interframe compression for video; JPEG is for images.
  5. Compression Ratio Calculation:

    • Given: Original size = 100 KB, Compressed = 10 KB.
    • Ratio = Original/Compressed = 10:1.
  6. Diagrams:

    • Always draw sampling process, psychoacoustic model, or MPEG frame types if asked for steps.
    • Label every component (e.g., "I-frame," "P-frame," "Quantization step").

Final Note: Audio compression is about balancing quality, file size, and real-world constraints. Master the trade-offs (e.g., MIDI vs. WAV, MP3 vs. FLAC) and real-world examples (WhatsApp, Ncell, eSewa) to score full marks!

Based on the TU BCA syllabus for Multimedia System (CACS457), unit 3.

Discussion

Loading…