Multimedia ComputingUnit 210 min read
Text, Sound & Audio Systems: Encoding, Formats & Applications
Unit 2 of Multimedia Computing explores how text, sound, and audio are digitized, stored, and processed—covering ASCII/Unicode, sampling rates, compression (MP3, WAV), synthesis (FM, additive), and real-world applications in apps like eSewa’s voice authentication or Pathao’s audio feedback.
TAKEAWAYS:
- Text is encoded as binary (ASCII/Unicode) to represent characters, with Unicode supporting global scripts (Devanagari, Cyrillic).
- Sound waves are digitized via sampling (amplitude × frequency), measured in bits per sample (bps) and samples per second (Hz).
- Audio formats (WAV, MP3, AAC) balance compression (lossy vs. lossless) and file size—MP3 cuts high-frequency data to save space.
- Synthesis methods (FM, additive, subtractive) create sounds programmatically, used in games (e.g., 8-bit chiptunes) and virtual instruments.
- Real-world systems (eSewa’s voice OTP, WhatsApp voice messages, YouTube audio streams) rely on these principles for efficiency and clarity.
- Exam focus: Compare formats (table), calculate file sizes from sampling rates, and explain how compression works (e.g., MP3’s psychoacoustic model).
1. Text Representation: From Characters to Binary
Text is the foundation of multimedia, enabling user interfaces, subtitles, and metadata. Computers store text as binary numbers using encoding schemes.
1.1 Character Encoding: ASCII vs. Unicode
ASCII (American Standard Code for Information Interchange):
- Uses 7 bits (128 characters: A-Z, 0-9, punctuation, control codes).
- Limited to English; cannot represent Devanagari, Nepali, or emojis.
- Example:
'A'=01000001in binary.
Unicode:
- Supports all languages (including Nepali, Chinese, Arabic) and symbols (emojis, mathematical notation).
- Uses 16 bits (UTF-16) or 32 bits (UTF-32) per character.
- Example:
'न'(Devanagari "na") =U+0928in hexadecimal.
classDiagram
class ASCII {
+7 bits
+128 characters
+Limited to English
}
class Unicode {
+16/32 bits
+128,000+ characters
+Supports Nepali, emojis
}
ASCII -->|"Subset of"| Unicode1.2 Text in Multimedia
- Applications:
- eSewa app: Uses Unicode to display Nepali transaction details.
- WhatsApp messages: Relies on UTF-8 (variable-width encoding) for global language support.
- Subtitles in videos: Stored as
.srtor.vttfiles (Unicode text with timestamps).
Worked Example: Convert the Nepali word "नमस्ते" to Unicode hexadecimal.
- Breakdown:
न= U+0928,म= U+093E,स= U+0938,त= U+093E,े= U+0947. - Answer:
0928 093E 0938 093E 0947.
Shows how Nepali characters map to Unicode code points. (Image: User:Hintha, CC BY-SA 4.0, via Wikimedia Commons)
2. Sound and Audio: The Science of Digital Sound
Sound is a mechanical wave (vibrations in air) that we perceive as audio. To digitize sound, we sample its amplitude over time.
2.1 The Physics of Sound Waves
- Key properties:
- Amplitude: Loudness (measured in decibels, dB).
- Frequency: Pitch (measured in Hertz, Hz; human hearing: 20Hz–20kHz).
- Wavelength: Distance between wave peaks.
graph LR
A["Sound Wave"] --> B["Amplitude\n(Loudness)"]
A --> C["Frequency\n(Pitch)"]
A --> D["Wavelength\n(λ)"]
Shows amplitude, frequency, and wavelength of a sine wave. (Image: Kulayada, CC BY-SA 4.0, via Wikimedia Commons)
2.2 Digital Audio: Sampling and Quantization
To store sound digitally:
- Sampling: Measure amplitude at fixed intervals (e.g., 44,100 times/sec for CD quality).
- Quantization: Convert each sample to a binary number (e.g., 16-bit = 65,536 possible values).
Sampling Rate (Hz):
- CD Quality: 44.1 kHz (Nyquist theorem: sample at ≥ 2× highest frequency).
- Telephony: 8 kHz (sufficient for human speech).
- High-Fidelity: 96 kHz (used in studios).
Bit Depth (bps):
- 8-bit: 256 levels (low quality, e.g., old games).
- 16-bit: 65,536 levels (CD standard).
- 24-bit: 16.7 million levels (professional audio).
2.3 Audio File Formats: Trade-offs Between Quality and Size
| Format | Type | Compression | Use Case | Example |
|---|---|---|---|---|
| WAV | Uncompressed | None | High-quality editing | Audio CDs |
| MP3 | Lossy | Psychoacoustic | Music streaming | YouTube, Spotify |
| AAC | Lossy | Advanced | Mobile devices | iTunes, WhatsApp calls |
| FLAC | Lossless | Entropy coding | Archival, high-fidelity | Lossless music libraries |
| OGG | Lossy/Lossless | Vorbis | Open-source projects | Linux audio |
How MP3 Compression Works:
- Psychoacoustic Model: Removes frequencies humans can’t hear (e.g., >16kHz for most people).
- Bitrate: Lower bitrate (e.g., 128 kbps) = smaller file but less quality.
- Example: A 5-minute 44.1kHz, 16-bit WAV file = 52.9 MB. Same audio as MP3 at 128 kbps = ~6.4 MB.
3. Audio Synthesis: Creating Sound Programmatically
Instead of recording sound, synthesis generates it using algorithms. Used in games, virtual instruments, and sound effects.
3.1 Common Synthesis Methods
| Method | How It Works | Example Use Case |
|---|---|---|
| FM Synthesis | Modulates one waveform with another | Yamaha DX7 synthesizers |
| Additive | Sums multiple sine waves | Orchestral libraries (e.g., Vienna Symphonic Library) |
| Subtractive | Starts with noise, filters out frequencies | Electric guitar amps |
| Granular | Chops sound into tiny grains | Glitch-hop music |
Worked Example: FM Synthesis (Simplified)
- Carrier Wave: 440 Hz (A4 note).
- Modulator Wave: 100 Hz (sine wave).
- Result: A "metallic" tone (used in 8-bit games like Mario).
graph TD
A["Carrier: 440Hz"] --> B["Modulator: 100Hz"]
B -->|"FM Modulation"| C["Output: Complex Waveform"]4. Real-World Applications
4.1 eSewa: Voice Authentication
- How it works:
- User speaks a random phrase (e.g., "Your OTP is 1234").
- System converts speech to digital audio (16-bit, 8kHz sampling).
- Feature extraction (MFCC) identifies unique voice patterns.
- Matches against stored voiceprint (Unicode text prompts in Nepali).
- Why it matters: Uses audio sampling + compression to securely verify identity without passwords.
4.2 Pathao: Audio Feedback for Drivers
- How it works:
- Pathao’s app sends text-to-speech (TTS) instructions to drivers (e.g., "Turn left in 500m").
- TTS converts Unicode text to 16kHz, 8-bit audio (small file size for mobile).
- Driver hears real-time navigation via AAC compression (low bandwidth).
- Why it matters: Balances text encoding (Unicode) and audio compression (AAC) for clarity and speed.
4.3 YouTube: Adaptive Bitrate Streaming
- How it works:
- Uploaded video/audio is encoded in multiple bitrates (e.g., 64kbps, 128kbps, 320kbps).
- YouTube’s server detects user’s internet speed and delivers the highest possible quality.
- Uses AAC for audio (compressed but high-quality).
- Why it matters: Dynamic sampling rates ensure smooth playback on slow connections.
5. Exam Tips
Text Encoding:
- Know the difference between ASCII (7-bit) and Unicode (16/32-bit).
- Exam question: "Convert ‘Hello’ to binary using ASCII." → Use table lookup.
- Trick: Nepali exams often test Unicode for Devanagari (e.g., "नमस्ते").
Audio Calculations:
- File size formula:
Size (bytes) = (Sampling Rate × Bit Depth × Number of Channels × Duration) / 8Example: 1 minute of stereo (2 channels), 44.1kHz, 16-bit WAV:(44100 × 16 × 2 × 60) / 8 = 10,584,000 bytes ≈ 10.1 MB. - Compression ratio: Always compare original vs. compressed size.
- File size formula:
Synthesis vs. Recording:
- Synthesis = Programmatic (e.g., FM for game sounds).
- Recording = Digitizing real sound (e.g., WAV files).
- Exam question: "Why use FM synthesis in a mobile game?" → Answer: Smaller file size, no need for large audio files.
Format Comparisons:
- Lossless (FLAC, WAV): No quality loss, large files.
- Lossy (MP3, AAC): Smaller files, some quality loss.
- Exam question: "Which format would you use for archiving a music collection?" → FLAC (lossless).
Real-World Scenarios:
- Bank loan interest calculations → Not directly relevant, but audio compression is used in IVR systems (e.g., NMB Bank’s automated calls).
- Traffic routes → Not relevant, but text-to-speech (TTS) is used in GPS apps like Pathao or Google Maps for navigation.
6. Common Pitfalls
- Mixing up sampling rate and bit depth:
- ❌ "Higher sampling rate = better quality" (only if bit depth is fixed).
- ✅ Both matter: 44.1kHz + 8-bit = poor quality; 44.1kHz + 24-bit = high fidelity.
- Assuming all compression is lossless:
- MP3 is lossy (discards inaudible frequencies), while FLAC is lossless.
- Ignoring Unicode in Nepali context:
- Always use UTF-8 for Nepali text in apps (e.g., eSewa, Daraz).
7. Practice Questions (Self-Check)
- Convert the Nepali word "धन्यवाद" to Unicode hexadecimal.
- Calculate the file size of a 3-minute, mono, 22.05kHz, 16-bit WAV file.
- Why does WhatsApp use AAC instead of MP3 for voice messages?
- Draw a sound wave with:
- Amplitude = 0.5V (peak)
- Frequency = 440Hz
- Label the wavelength and period.
- Compare FM synthesis and additive synthesis in a table (how they work, use cases).
8. Further Reading
- Books:
- The Audio Programming Book (Will Pirkle) – Covers synthesis in detail.
- Multimedia Systems (Ralf Steinmetz) – Covers encoding and compression.
- Tools to Try:
- Audacity (edit audio, see sampling in action).
- Unicode Table (https://unicode-table.com/) – Explore Nepali characters.
- FM Synthesis Demo (https://www.jsfxr.com/) – Generate game sounds.
Based on the TU BIT syllabus for Multimedia Computing (BIT356), unit 2.
Discussion
Loading…