Multimedia SystemUnit 311 min read
Audio Fundamentals, Compression & Real-World Applications
Unit 3 of Multimedia System explores how audio signals are digitized, compressed (lossy vs. lossless), and optimized for storage/transmission, with comparisons to real-world apps like WhatsApp voice notes and Ncell music streaming.
TAKEAWAYS
- Audio compression reduces file size by discarding redundant or imperceptible data (lossy) or using mathematical patterns (lossless).
- Sampling rate (e.g., 44.1 kHz) and bit depth (e.g., 16-bit) determine audio quality—higher values mean better fidelity but larger files.
- MP3 (lossy) and FLAC (lossless) are industry standards for music, while AAC dominates streaming (Apple Music, Spotify).
- MIDI synthesizes sound via instructions (small files) but lacks the realism of recorded audio.
- Perceptual coding exploits human hearing limitations (e.g., masking) to improve compression efficiency.
- Real-world trade-offs: WhatsApp uses Opus (low-latency, adaptive bitrate) for calls, while Ncell’s music app prioritizes AAC for balance between quality and data savings.
1. Audio Fundamentals: From Analog to Digital
Audio is a continuous analog signal (sound waves) that must be converted to digital form for computers to process. This involves three key steps:
1.1 Analog-to-Digital Conversion (ADC)
- Sampling: Capturing the amplitude of the analog signal at discrete intervals (measured in Hz).
- Example: A 44.1 kHz sampling rate means 44,100 samples per second.
- Nyquist Theorem: Sampling rate must be ≥ 2 × highest frequency in the signal to avoid aliasing.
- Quantization: Assigning numerical values to each sample (e.g., 16-bit = 65,536 possible levels).
- Encoding: Storing samples as binary data (e.g., WAV, AIFF).
1.2 Key Audio Parameters
| Parameter | Definition | Example Values | Impact on Quality/File Size |
|---|---|---|---|
| Sampling Rate | Samples per second (Hz) | 44.1 kHz, 48 kHz | Higher = better fidelity |
| Bit Depth | Bits per sample (resolution) | 16-bit, 24-bit | Higher = less noise |
| Channels | Mono vs. Stereo vs. Surround | 1 (Mono), 2 (Stereo) | More channels = larger files |
Worked Example: CD Audio vs. MP3
- A CD-quality audio file:
- Sampling rate: 44.1 kHz
- Bit depth: 16-bit
- Channels: 2 (Stereo)
- Uncompressed size: ~10 MB per minute (700 KB/sec).
- An MP3 (compressed) version of the same audio might be ~1 MB/minute (100 KB/sec).
2. Audio Compression Techniques
Compression reduces file size by removing redundancy or irrelevant data. Two main types:
2.1 Lossless Compression
- No data loss; original audio can be perfectly reconstructed.
- Methods:
- Run-Length Encoding (RLE): Replaces repeated values with a count.
Example:
"aaaaabbbbcccdde"→(5a,4b,3c,2d,e)→ Compression ratio = 12/10 = 1.2 (20% reduction). - Dictionary Methods: Stores repeated patterns once (e.g., FLAC, ALAC).
- Entropy Coding: Assigns shorter codes to frequent values (e.g., Huffman coding).
- Run-Length Encoding (RLE): Replaces repeated values with a count.
Example:
- Use Cases: Archival audio, professional editing (FLAC, WMA Lossless).
2.2 Lossy Compression
- Removes imperceptible or redundant data; irreversible but achieves high compression ratios.
- Key Techniques:
- Perceptual Coding: Exploits how humans hear:
- Masking: Loud sounds hide quiet ones (e.g., a drum drowns out a hi-hat).
- Frequency Irrelevance: Humans don’t hear ultra-high/low frequencies well.
- ** Psychoacoustic Models**: Analyzes audio to discard inaudible components.
- Hybrid Approaches: Combines time-domain and frequency-domain analysis (e.g., MP3, AAC).
- Perceptual Coding: Exploits how humans hear:
| Lossy Format | Compression Ratio | Use Case | Example Apps |
|---|---|---|---|
| MP3 | 10:1 to 12:1 | Music streaming | Spotify, Gaana |
| AAC | 7:1 to 10:1 | Apple Music, YouTube | iTunes, Netflix |
| Opus | 2:1 to 10:1 | Voice/video calls | WhatsApp, Zoom |
| Vorbis | 5:1 to 8:1 | Open-source audio | YouTube (WebM) |
Worked Example: MP3 Compression in Ncell Music
- Original WAV: 10 MB/minute (44.1 kHz, 16-bit, stereo).
- MP3 (128 kbps): ~1.6 MB/minute (80% smaller).
- Trade-off: Slight loss in high-frequency details (inaudible to most listeners).
3. Specialized Audio Formats
3.1 MIDI (Musical Instrument Digital Interface)
- Not audio data; stores instructions (notes, timing, instruments) for synthesizers.
- Advantages:
- Tiny file sizes (e.g., a 5-minute song = ~50 KB).
- Editable (change instruments, tempo without re-recording).
- Disadvantages:
- No "real" audio; relies on a sound font/synthesizer.
- Poor for complex sounds (e.g., vocals, live recordings).
- Use Cases: Video game soundtracks, ringtone editors, karaoke apps.
Comparison: MIDI vs. Digital Audio for Video Games
| Feature | MIDI | Digital Audio (WAV/MP3) |
|---|---|---|
| File Size | ~50 KB for 5 min | ~10 MB (WAV), ~1 MB (MP3) |
| Quality | Synthetic, limited instruments | High-fidelity, realistic |
| Editability | Fully editable (notes, BPM) | Fixed recording |
| Latency | Low (instant playback) | Higher (file loading) |
| Best For | Background music, chiptune | Voiceovers, sound effects |
Real-World Example: Pathao’s In-App Music
- Uses AAC (lossy) for background music in the driver app to:
- Save mobile data (critical for low-bandwidth users).
- Balance quality and file size (~1 MB per song).
4. Real-World Applications of Audio Compression
4.1 WhatsApp Voice Messages (Opus Codec)
- Why Opus?
- Adaptive bitrate: Adjusts quality based on network speed.
- Low latency: Optimized for real-time calls (~200 ms delay).
- Compression: Reduces a 1-minute voice note from ~10 MB (WAV) to ~500 KB.
- How It Works:
- Audio sampled at 48 kHz.
- Opus applies perceptual coding to discard inaudible frequencies.
- Encrypted and sent over WhatsApp’s servers.
4.2 Ncell Music Streaming (AAC)
- Challenge: Nepal’s average internet speed (~10 Mbps) struggles with high-bitrate audio.
- Solution:
- Uses AAC at 128–192 kbps (balance of quality and data).
- Adaptive streaming: Switches between 64 kbps (slow network) and 320 kbps (fast network).
- Result: A 3-minute song downloads in ~1–2 MB vs. ~30 MB uncompressed.
4.3 eSewa’s IVR System (MIDI + Text-to-Speech)
- Problem: Storing pre-recorded voice prompts for transactions (e.g., "Enter your PIN") wastes space.
- Solution:
- Uses MIDI-like synthesized speech for static prompts.
- Dynamic prompts (e.g., balance updates) use low-bitrate MP3 snippets.
- Savings: Reduces storage by 90% compared to WAV files.
5. Why Compression Matters for Video and Animation
5.1 Data Volume in Multimedia
- 1 minute of uncompressed video:
- 4K @ 60 FPS: ~1.5 GB (requires 150 Mbps bandwidth).
- Audio (stereo, 48 kHz): ~10 MB/minute.
- Compressed (H.264 + AAC):
- ~50 MB for the same video (97% smaller).
5.2 Compression Techniques in Video
| Technique | Lossy/Lossless | How It Works | Example Formats |
|---|---|---|---|
| Intraframe | Lossy | Compresses each frame independently | JPEG, PNG |
| Interframe | Lossy | Exploits similarity between frames | MPEG, H.264 |
| Motion Compensation | Lossy | Tracks moving objects across frames | MP4, WebM |
| Discrete Cosine Transform (DCT) | Lossy | Converts spatial data to frequency domain | JPEG, MP3 |
Worked Example: Daraz’s Product Videos
- Problem: Sellers upload 1-minute videos (1080p, 30 FPS) → ~500 MB uncompressed.
- Solution: Daraz’s platform compresses to ~20 MB using:
- H.264 (interframe compression).
- AAC audio (128 kbps).
- Result: Faster uploads, lower bandwidth for viewers.
Exam Tip
Lossy vs. Lossless:
- Lossy (MP3, AAC) = high compression, irreversible, used for music/streaming.
- Lossless (FLAC, WAV) = no quality loss, used for archiving.
- Exam trick: Always mention perceptual coding for lossy formats.
Sampling Rate/Bit Depth:
- CD quality = 44.1 kHz, 16-bit.
- Broadcast radio = 48 kHz, 24-bit.
- Higher values = better quality but larger files.
Real-World Scenarios:
- WhatsApp: Opus codec for calls.
- Ncell Music: AAC for streaming.
- eSewa IVR: MIDI/text-to-speech for prompts.
- Tip: Relate numerical examples to these apps (e.g., "AAC at 128 kbps reduces a song to 1 MB").
Trade-offs:
- MIDI vs. Digital Audio: MIDI saves space but lacks realism.
- JPEG vs. MPEG: MPEG uses interframe compression for video; JPEG is for images.
Compression Ratio Calculation:
- Given: Original size = 100 KB, Compressed = 10 KB.
- Ratio = Original/Compressed = 10:1.
Diagrams:
- Always draw sampling process, psychoacoustic model, or MPEG frame types if asked for steps.
- Label every component (e.g., "I-frame," "P-frame," "Quantization step").
Final Note: Audio compression is about balancing quality, file size, and real-world constraints. Master the trade-offs (e.g., MIDI vs. WAV, MP3 vs. FLAC) and real-world examples (WhatsApp, Ncell, eSewa) to score full marks!
Based on the TU BCA syllabus for Multimedia System (CACS457), unit 3.
Discussion
Loading…