Multimedia ComputingUnit 210 min read

Text, Sound & Audio Systems: Encoding, Formats & Applications

Unit 2 of Multimedia Computing explores how text, sound, and audio are digitized, stored, and processed—covering ASCII/Unicode, sampling rates, compression (MP3, WAV), synthesis (FM, additive), and real-world applications in apps like eSewa’s voice authentication or Pathao’s audio feedback.

TAKEAWAYS:

  • Text is encoded as binary (ASCII/Unicode) to represent characters, with Unicode supporting global scripts (Devanagari, Cyrillic).
  • Sound waves are digitized via sampling (amplitude × frequency), measured in bits per sample (bps) and samples per second (Hz).
  • Audio formats (WAV, MP3, AAC) balance compression (lossy vs. lossless) and file size—MP3 cuts high-frequency data to save space.
  • Synthesis methods (FM, additive, subtractive) create sounds programmatically, used in games (e.g., 8-bit chiptunes) and virtual instruments.
  • Real-world systems (eSewa’s voice OTP, WhatsApp voice messages, YouTube audio streams) rely on these principles for efficiency and clarity.
  • Exam focus: Compare formats (table), calculate file sizes from sampling rates, and explain how compression works (e.g., MP3’s psychoacoustic model).

1. Text Representation: From Characters to Binary

Text is the foundation of multimedia, enabling user interfaces, subtitles, and metadata. Computers store text as binary numbers using encoding schemes.

1.1 Character Encoding: ASCII vs. Unicode

  • ASCII (American Standard Code for Information Interchange):

    • Uses 7 bits (128 characters: A-Z, 0-9, punctuation, control codes).
    • Limited to English; cannot represent Devanagari, Nepali, or emojis.
    • Example: 'A' = 01000001 in binary.
  • Unicode:

    • Supports all languages (including Nepali, Chinese, Arabic) and symbols (emojis, mathematical notation).
    • Uses 16 bits (UTF-16) or 32 bits (UTF-32) per character.
    • Example: 'न' (Devanagari "na") = U+0928 in hexadecimal.
classDiagram
    class ASCII {
        +7 bits
        +128 characters
        +Limited to English
    }
    class Unicode {
        +16/32 bits
        +128,000+ characters
        +Supports Nepali, emojis
    }
    ASCII -->|"Subset of"| Unicode

1.2 Text in Multimedia

  • Applications:
    • eSewa app: Uses Unicode to display Nepali transaction details.
    • WhatsApp messages: Relies on UTF-8 (variable-width encoding) for global language support.
    • Subtitles in videos: Stored as .srt or .vtt files (Unicode text with timestamps).

Worked Example: Convert the Nepali word "नमस्ते" to Unicode hexadecimal.

  • Breakdown: न = U+0928, म = U+093E, स = U+0938, त = U+093E, े = U+0947.
  • Answer: 0928 093E 0938 093E 0947.

unicode character map labelled diagram**Shows how Nepali characters map to Unicode code points. (Image: User:Hintha, CC BY-SA 4.0, via Wikimedia Commons)


2. Sound and Audio: The Science of Digital Sound

Sound is a mechanical wave (vibrations in air) that we perceive as audio. To digitize sound, we sample its amplitude over time.

2.1 The Physics of Sound Waves

  • Key properties:
    • Amplitude: Loudness (measured in decibels, dB).
    • Frequency: Pitch (measured in Hertz, Hz; human hearing: 20Hz–20kHz).
    • Wavelength: Distance between wave peaks.
graph LR
    A["Sound Wave"] --> B["Amplitude\n(Loudness)"]
    A --> C["Frequency\n(Pitch)"]
    A --> D["Wavelength\n(λ)"]

sound wave labelled diagram**Shows amplitude, frequency, and wavelength of a sine wave. (Image: Kulayada, CC BY-SA 4.0, via Wikimedia Commons)

2.2 Digital Audio: Sampling and Quantization

To store sound digitally:

  1. Sampling: Measure amplitude at fixed intervals (e.g., 44,100 times/sec for CD quality).
  2. Quantization: Convert each sample to a binary number (e.g., 16-bit = 65,536 possible values).
  • Sampling Rate (Hz):

    • CD Quality: 44.1 kHz (Nyquist theorem: sample at ≥ 2× highest frequency).
    • Telephony: 8 kHz (sufficient for human speech).
    • High-Fidelity: 96 kHz (used in studios).
  • Bit Depth (bps):

    • 8-bit: 256 levels (low quality, e.g., old games).
    • 16-bit: 65,536 levels (CD standard).
    • 24-bit: 16.7 million levels (professional audio).

2.3 Audio File Formats: Trade-offs Between Quality and Size

Format Type Compression Use Case Example
WAV Uncompressed None High-quality editing Audio CDs
MP3 Lossy Psychoacoustic Music streaming YouTube, Spotify
AAC Lossy Advanced Mobile devices iTunes, WhatsApp calls
FLAC Lossless Entropy coding Archival, high-fidelity Lossless music libraries
OGG Lossy/Lossless Vorbis Open-source projects Linux audio

How MP3 Compression Works:

  • Psychoacoustic Model: Removes frequencies humans can’t hear (e.g., >16kHz for most people).
  • Bitrate: Lower bitrate (e.g., 128 kbps) = smaller file but less quality.
  • Example: A 5-minute 44.1kHz, 16-bit WAV file = 52.9 MB. Same audio as MP3 at 128 kbps = ~6.4 MB.


3. Audio Synthesis: Creating Sound Programmatically

Instead of recording sound, synthesis generates it using algorithms. Used in games, virtual instruments, and sound effects.

3.1 Common Synthesis Methods

Method How It Works Example Use Case
FM Synthesis Modulates one waveform with another Yamaha DX7 synthesizers
Additive Sums multiple sine waves Orchestral libraries (e.g., Vienna Symphonic Library)
Subtractive Starts with noise, filters out frequencies Electric guitar amps
Granular Chops sound into tiny grains Glitch-hop music

Worked Example: FM Synthesis (Simplified)

  • Carrier Wave: 440 Hz (A4 note).
  • Modulator Wave: 100 Hz (sine wave).
  • Result: A "metallic" tone (used in 8-bit games like Mario).
graph TD
    A["Carrier: 440Hz"] --> B["Modulator: 100Hz"]
    B -->|"FM Modulation"| C["Output: Complex Waveform"]


4. Real-World Applications

4.1 eSewa: Voice Authentication

  • How it works:
    1. User speaks a random phrase (e.g., "Your OTP is 1234").
    2. System converts speech to digital audio (16-bit, 8kHz sampling).
    3. Feature extraction (MFCC) identifies unique voice patterns.
    4. Matches against stored voiceprint (Unicode text prompts in Nepali).
  • Why it matters: Uses audio sampling + compression to securely verify identity without passwords.

4.2 Pathao: Audio Feedback for Drivers

  • How it works:
    1. Pathao’s app sends text-to-speech (TTS) instructions to drivers (e.g., "Turn left in 500m").
    2. TTS converts Unicode text to 16kHz, 8-bit audio (small file size for mobile).
    3. Driver hears real-time navigation via AAC compression (low bandwidth).
  • Why it matters: Balances text encoding (Unicode) and audio compression (AAC) for clarity and speed.

4.3 YouTube: Adaptive Bitrate Streaming

  • How it works:
    1. Uploaded video/audio is encoded in multiple bitrates (e.g., 64kbps, 128kbps, 320kbps).
    2. YouTube’s server detects user’s internet speed and delivers the highest possible quality.
    3. Uses AAC for audio (compressed but high-quality).
  • Why it matters: Dynamic sampling rates ensure smooth playback on slow connections.


5. Exam Tips

  1. Text Encoding:

    • Know the difference between ASCII (7-bit) and Unicode (16/32-bit).
    • Exam question: "Convert ‘Hello’ to binary using ASCII." → Use table lookup.
    • Trick: Nepali exams often test Unicode for Devanagari (e.g., "नमस्ते").
  2. Audio Calculations:

    • File size formula: Size (bytes) = (Sampling Rate × Bit Depth × Number of Channels × Duration) / 8 Example: 1 minute of stereo (2 channels), 44.1kHz, 16-bit WAV: (44100 × 16 × 2 × 60) / 8 = 10,584,000 bytes ≈ 10.1 MB.
    • Compression ratio: Always compare original vs. compressed size.
  3. Synthesis vs. Recording:

    • Synthesis = Programmatic (e.g., FM for game sounds).
    • Recording = Digitizing real sound (e.g., WAV files).
    • Exam question: "Why use FM synthesis in a mobile game?" → Answer: Smaller file size, no need for large audio files.
  4. Format Comparisons:

    • Lossless (FLAC, WAV): No quality loss, large files.
    • Lossy (MP3, AAC): Smaller files, some quality loss.
    • Exam question: "Which format would you use for archiving a music collection?" → FLAC (lossless).
  5. Real-World Scenarios:

    • Bank loan interest calculations → Not directly relevant, but audio compression is used in IVR systems (e.g., NMB Bank’s automated calls).
    • Traffic routes → Not relevant, but text-to-speech (TTS) is used in GPS apps like Pathao or Google Maps for navigation.

6. Common Pitfalls

  • Mixing up sampling rate and bit depth:
    • ❌ "Higher sampling rate = better quality" (only if bit depth is fixed).
    • ✅ Both matter: 44.1kHz + 8-bit = poor quality; 44.1kHz + 24-bit = high fidelity.
  • Assuming all compression is lossless:
    • MP3 is lossy (discards inaudible frequencies), while FLAC is lossless.
  • Ignoring Unicode in Nepali context:
    • Always use UTF-8 for Nepali text in apps (e.g., eSewa, Daraz).

7. Practice Questions (Self-Check)

  1. Convert the Nepali word "धन्यवाद" to Unicode hexadecimal.
  2. Calculate the file size of a 3-minute, mono, 22.05kHz, 16-bit WAV file.
  3. Why does WhatsApp use AAC instead of MP3 for voice messages?
  4. Draw a sound wave with:
    • Amplitude = 0.5V (peak)
    • Frequency = 440Hz
    • Label the wavelength and period.
  5. Compare FM synthesis and additive synthesis in a table (how they work, use cases).

8. Further Reading

  • Books:
    • The Audio Programming Book (Will Pirkle) – Covers synthesis in detail.
    • Multimedia Systems (Ralf Steinmetz) – Covers encoding and compression.
  • Tools to Try:

Based on the TU BIT syllabus for Multimedia Computing (BIT356), unit 2.

Discussion

Loading…