BIT151 Microprocessor and Computer Architecture

Microprocessor and Computer ArchitectureUnit 1110 min read

Special Topics in Microprocessor & Architecture: RISC, CISC, Embedded Systems, Cache, Multiprocessing

Unit 11 of Microprocessor and Computer Architecture covers advanced topics like RISC vs CISC design philosophies, embedded system architectures, cache memory hierarchies, multiprocessing models, and real-world applications in Nepalese tech (eSewa, Ncell) and global systems (Google, WhatsApp). This note includes visual

Key Concepts & Visual Breakdown

1. RISC vs CISC: The CPU Design War

Definitions & Core Differences

classDiagram
    class RISC {
        + Simple instructions
        + Fixed-length opcodes
        + Load-store architecture
        + Pipelining support
        + Few addressing modes
    }
    class CISC {
        + Complex instructions (e.g., MUL, DIV in one cycle)
        + Variable-length opcodes
        + Memory-memory operations
        + Microcode control
        + Many addressing modes
    }
    RISC --> "Designed for" ARM Cortex
    CISC --> "Designed for" Intel x86

How It Works:

  • RISC (Reduced Instruction Set Computing):

    • Fewer, simpler instructions (e.g., ARM in smartphones).
    • Relies on compiler optimizations (e.g., loop unrolling) and hardware pipelining for speed.
    • Example: ADD R1, R2, R3 (all operands in registers).
    • Real-world use: Smartphones (Apple A-series, Qualcomm Snapdragon), Raspberry Pi.
  • CISC (Complex Instruction Set Computing):

    • Single instruction handles complex tasks (e.g., REP MOVSB in x86 for block memory copy).
    • Uses microcode (firmware-level instructions) to emulate complex ops.
    • Example: MUL AX, BX (multiplies two 16-bit numbers in one instruction).
    • Real-world use: PCs (Intel Core i7), servers (AMD EPYC).

Performance Trade-offs

Feature RISC CISC
Instruction Count High (10–20 per task) Low (1–5 per task)
Clock Speed Higher (simpler decoding) Lower (complex decoding)
Power Efficiency Better (mobile devices) Worse (desktops/servers)
Hardware Cost Cheaper (simpler ALU) Expensive (microcode ROM)
Example Chips ARM Cortex-M4, MIPS Intel 8086, x86-64

Worked Example: Adding Two Numbers

  • RISC (ARM):

    LDR  R1, [R0]   ; Load from memory to R1
    LDR  R2, [R4]   ; Load from memory to R2
    ADD  R3, R1, R2 ; Add registers
    STR  R3, [R8]   ; Store result
    

    Cycles: 4 (1 per instruction, pipelined).

  • CISC (x86):

    MOV  EAX, [R0]  ; Load
    ADD  EAX, [R4]  ; Add (memory-memory)
    MOV  [R8], EAX  ; Store
    

    Cycles: 3 (but ADD may take multiple cycles for memory ops).


Label key components: ALU, pipeline stages, register file, cache.


2. Embedded Systems: Tiny Computers, Big Impact

Architecture Overview

Key Features:

  • Harvard Architecture: Separate data/program buses (unlike von Neumann).
  • Real-time OS (RTOS): Used in eSewa payment terminals (FreeRTOS).
  • Peripheral Integration: ADC for sensor data (e.g., NTC’s smart meters).

Nepalese Applications

Device Microcontroller Role
eSewa POS Machine STM32 (ARM Cortex-M) Handles card swipes, encrypts data
Ncell IoT Router ESP32 (Xtensa LX6) Manages cellular connections
Pathao Bike Lock ATmega328P (AVR) GPS tracking + keyless entry

Worked Example: Traffic Light Controller (8051)

// Pseudocode for 8051 (used in Kathmandu traffic signals)
void main() {
    P1 = 0x01; // Green for East-West
    delay(30); // 30 seconds
    P1 = 0x02; // Yellow
    delay(5);
    P1 = 0x04; // Red for East-West, Green for North-South
    delay(25);
}

Why 8051?

  • Low cost (~$0.50 per chip).
  • Built-in timers for delays.
  • Real-world tie: Kathmandu’s traffic lights use similar MCUs with solar power.

Label: GPIO pins, UART TX/RX, ADC inputs, power pins.


3. Cache Memory: The Speed vs. Size Trade-off

Hierarchy & Mapping Techniques

mindmap
  root((Cache Hierarchy))
    L1["L1 Cache (KB range)"]
      --> "Split: Data (D$) + Instruction (I$)"
      --> "Associativity: 2-way/4-way set-associative"
    L2["L2 Cache (MB range)"]
      --> "Unified or split"
      --> "Shared between cores (multicore)"
    L3["L3 Cache (Multi-core shared)"]
      --> "Largest, slowest (e.g., 8MB in Intel i7)"
    Mapping["Mapping Policies"]
      --> "Direct-Mapped: Fast but conflicts"
      --> "Fully Associative: Flexible but slow"
      --> "Set-Associative: Balance (e.g., 4-way)"

How It Works:

  • Hit Rate: Probability data is in cache. Formula:
  • Miss Penalty: Time to fetch from next level. Example: L1 miss → 10 cycles; L2 miss → 50 cycles.

Worked Example: Cache Hit/Miss Calculation

Scenario: A program accesses memory addresses 0x100, 0x104, 0x108 in a direct-mapped 4KB L1 cache with 16-byte blocks.

  1. Block Offset: 4 bits (16 bytes = ).
  2. Index: bits.
  3. Tag: Remaining bits (32–4–8 = 20 bits).
Address Binary (Hex) Tag Index Offset Cache Block Hit/Miss
0x100 0001 0000 0000 0000 0001 00000000 0000 Block 0 Hit
0x104 0001 0000 0000 0010 0001 00000000 0010 Block 0 Hit
0x200 0010 0000 0000 0000 0010 00000000 0000 Block 0 Miss

Average Memory Access Time (AMAT): For 1ns hit time, 10ns miss penalty, 90% hit rate:


Label: Core, L1 I$/D$, L2, L3, main memory, and data flow arrows.


4. Multiprocessing: Parallelism in Action

Models & Synchronization

stateDiagram-v2
    [*] --> Waiting
    Waiting --> Running : "CPU assigned"
    Running --> Blocked : "I/O or lock"
    Running --> Finished : "Termination"
    Blocked --> Waiting : "Resource available"
    Finished --> [*]

Key Concepts:

  • SMP (Symmetric Multiprocessing): All cores share memory (e.g., Google servers).
  • MPP (Massively Parallel): Distributed memory (e.g., supercomputers).
  • Race Condition: When two threads access shared data without synchronization. Fix: Mutex locks, semaphores.

Worked Example: Bank Transaction (Race Condition)

Problem: Two customers deposit $100 each into an account with initial balance $0.

// Thread 1: Customer A
balance = read_memory(0x1000)  // Reads $0
balance = balance + 100        // $100
write_memory(0x1000, balance)  // Writes $100

// Thread 2: Customer B (same steps)

Outcome: Final balance = $100 (should be $200). Race condition!

Solution: Mutex Lock

lock.acquire()
balance = read_memory(0x1000)  // $0
balance = balance + 100        // $100
write_memory(0x1000, balance)  // $100
lock.release()

// Customer B waits until lock is free

Label: 4 cores, L2/L3 cache, integrated GPU, and memory controller.


In the Real World

  1. eSewa’s Payment Terminals

    • Topic: Embedded Systems + RISC Architecture
    • How: Uses an ARM Cortex-M4 MCU to:
      • Read NFC cards (via GPIO/UART).
      • Encrypt transactions (AES hardware accelerator).
      • Communicate with eSewa servers (Wi-Fi module).
    • Why RISC? Low power, real-time response for payments.
  2. Ncell’s 4G Base Stations

    • Topic: Multiprocessing + Cache Hierarchy
    • How:
      • Dual-core ARM processors handle:
        • Core 1: User data routing (L1 cache for low latency).
        • Core 2: Backhaul to NTC (L2 cache for bulk transfers).
      • SMP model ensures calls don’t drop during peak hours (e.g., 2077 BS).
  3. WhatsApp’s End-to-End Encryption

    • Topic: RISC vs CISC in Mobile Chips
    • How:
      • Smartphones (RISC): ARM chips (e.g., Snapdragon 888) use NEON SIMD for fast AES encryption.
      • Servers (CISC): Intel Xeon CPUs handle millions of messages via multithreading (hyper-threading).
  4. NEPSE’s Stock Trading System

    • Topic: Cache Memory + Multiprocessing
    • How:
      • L3 cache in trading servers reduces latency for high-frequency trades.
      • MPP model: Separate nodes for order matching, user auth, and reporting.

Exam Tip

What Examiners Look For

  1. Diagrams > Text:

    • Draw block diagrams for embedded systems, state diagrams for multiprocessing, and cache mapping tables.
    • Example: For RISC/CISC, show instruction count vs. clock cycles in a graph.
  2. Real-world Mapping:

    • eSewa → Embedded Systems (MCU + RTOS).
    • Ncell → Multiprocessing (SMP for base stations).
    • Bank loans → Cache hits/misses (e.g., "Why does a bank’s server need L3 cache?").
  3. Common Pitfalls:

    • Confusing RISC/CISC: Remember RISC = "simple but many instructions," CISC = "few but complex."
    • Cache calculations: Always show hit/miss tables with binary addresses.
    • Race conditions: Describe before/after with mutex locks.
  4. Short-Answer Tricks:

    • Define: "Cache associativity" → "Number of places a block can reside in a set."
    • Compare: "RISC vs CISC" → Use a 2-column table (see above).
    • Calculate: "AMAT" → Plug in hit rate and miss penalty.

Sample Exam Questions & Answers

Q1: Explain why embedded systems use Harvard architecture. Draw a block diagram. A:

Answer:

  • Separate buses prevent von Neumann bottleneck (memory contention).
  • Faster execution: Program and data can be fetched simultaneously.
  • Example: eSewa’s payment terminal reads card data (Data Bus) while fetching instructions (Program Bus).

Q2: Calculate the number of sets in a 64KB, 4-way set-associative cache with 64-byte blocks. A:

  1. Total blocks: blocks.
  2. Blocks per set: 4 (4-way).
  3. Number of sets: sets.

Q3: How does WhatsApp use RISC architecture to encrypt messages faster? A:

  • ARM NEON: SIMD (Single Instruction Multiple Data) unit in RISC chips processes 4 data items per instruction.
  • Example: AES encryption on Snapdragon 888 uses NEON to encrypt 4 bytes at once (vs. 1 byte in scalar mode).
  • Result: 4x speedup for end-to-end encryption.

Final Note:

  • Memorize: RISC/CISC features, cache mapping types, and embedded system components.
  • Practice: Draw diagrams for every concept. Examiners reward visual answers.
  • Apply: Relate to Nepalese tech (eSewa, Ncell) and global examples (Google, WhatsApp).

Based on the TU BIT syllabus for Microprocessor and Computer Architecture (BIT151), unit 11.

Discussion

Loading…