Microprocessor and Computer ArchitectureUnit 1110 min read
Special Topics in Microprocessor & Architecture: RISC, CISC, Embedded Systems, Cache, Multiprocessing
Unit 11 of Microprocessor and Computer Architecture covers advanced topics like RISC vs CISC design philosophies, embedded system architectures, cache memory hierarchies, multiprocessing models, and real-world applications in Nepalese tech (eSewa, Ncell) and global systems (Google, WhatsApp). This note includes visual
Key Concepts & Visual Breakdown
1. RISC vs CISC: The CPU Design War
Definitions & Core Differences
classDiagram
class RISC {
+ Simple instructions
+ Fixed-length opcodes
+ Load-store architecture
+ Pipelining support
+ Few addressing modes
}
class CISC {
+ Complex instructions (e.g., MUL, DIV in one cycle)
+ Variable-length opcodes
+ Memory-memory operations
+ Microcode control
+ Many addressing modes
}
RISC --> "Designed for" ARM Cortex
CISC --> "Designed for" Intel x86How It Works:
RISC (Reduced Instruction Set Computing):
- Fewer, simpler instructions (e.g., ARM in smartphones).
- Relies on compiler optimizations (e.g., loop unrolling) and hardware pipelining for speed.
- Example:
ADD R1, R2, R3(all operands in registers). - Real-world use: Smartphones (Apple A-series, Qualcomm Snapdragon), Raspberry Pi.
CISC (Complex Instruction Set Computing):
- Single instruction handles complex tasks (e.g.,
REP MOVSBin x86 for block memory copy). - Uses microcode (firmware-level instructions) to emulate complex ops.
- Example:
MUL AX, BX(multiplies two 16-bit numbers in one instruction). - Real-world use: PCs (Intel Core i7), servers (AMD EPYC).
- Single instruction handles complex tasks (e.g.,
Performance Trade-offs
| Feature | RISC | CISC |
|---|---|---|
| Instruction Count | High (10–20 per task) | Low (1–5 per task) |
| Clock Speed | Higher (simpler decoding) | Lower (complex decoding) |
| Power Efficiency | Better (mobile devices) | Worse (desktops/servers) |
| Hardware Cost | Cheaper (simpler ALU) | Expensive (microcode ROM) |
| Example Chips | ARM Cortex-M4, MIPS | Intel 8086, x86-64 |
Worked Example: Adding Two Numbers
RISC (ARM):
LDR R1, [R0] ; Load from memory to R1 LDR R2, [R4] ; Load from memory to R2 ADD R3, R1, R2 ; Add registers STR R3, [R8] ; Store resultCycles: 4 (1 per instruction, pipelined).
CISC (x86):
MOV EAX, [R0] ; Load ADD EAX, [R4] ; Add (memory-memory) MOV [R8], EAX ; StoreCycles: 3 (but
ADDmay take multiple cycles for memory ops).
Label key components: ALU, pipeline stages, register file, cache.
2. Embedded Systems: Tiny Computers, Big Impact
Architecture Overview
Key Features:
- Harvard Architecture: Separate data/program buses (unlike von Neumann).
- Real-time OS (RTOS): Used in eSewa payment terminals (FreeRTOS).
- Peripheral Integration: ADC for sensor data (e.g., NTC’s smart meters).
Nepalese Applications
| Device | Microcontroller | Role |
|---|---|---|
| eSewa POS Machine | STM32 (ARM Cortex-M) | Handles card swipes, encrypts data |
| Ncell IoT Router | ESP32 (Xtensa LX6) | Manages cellular connections |
| Pathao Bike Lock | ATmega328P (AVR) | GPS tracking + keyless entry |
Worked Example: Traffic Light Controller (8051)
// Pseudocode for 8051 (used in Kathmandu traffic signals)
void main() {
P1 = 0x01; // Green for East-West
delay(30); // 30 seconds
P1 = 0x02; // Yellow
delay(5);
P1 = 0x04; // Red for East-West, Green for North-South
delay(25);
}
Why 8051?
- Low cost (~$0.50 per chip).
- Built-in timers for delays.
- Real-world tie: Kathmandu’s traffic lights use similar MCUs with solar power.
Label: GPIO pins, UART TX/RX, ADC inputs, power pins.
3. Cache Memory: The Speed vs. Size Trade-off
Hierarchy & Mapping Techniques
mindmap
root((Cache Hierarchy))
L1["L1 Cache (KB range)"]
--> "Split: Data (D$) + Instruction (I$)"
--> "Associativity: 2-way/4-way set-associative"
L2["L2 Cache (MB range)"]
--> "Unified or split"
--> "Shared between cores (multicore)"
L3["L3 Cache (Multi-core shared)"]
--> "Largest, slowest (e.g., 8MB in Intel i7)"
Mapping["Mapping Policies"]
--> "Direct-Mapped: Fast but conflicts"
--> "Fully Associative: Flexible but slow"
--> "Set-Associative: Balance (e.g., 4-way)"How It Works:
- Hit Rate: Probability data is in cache. Formula:
- Miss Penalty: Time to fetch from next level. Example: L1 miss → 10 cycles; L2 miss → 50 cycles.
Worked Example: Cache Hit/Miss Calculation
Scenario: A program accesses memory addresses 0x100, 0x104, 0x108 in a direct-mapped 4KB L1 cache with 16-byte blocks.
- Block Offset: 4 bits (16 bytes = ).
- Index: bits.
- Tag: Remaining bits (32–4–8 = 20 bits).
| Address | Binary (Hex) | Tag | Index | Offset | Cache Block | Hit/Miss |
|---|---|---|---|---|---|---|
| 0x100 | 0001 0000 0000 0000 | 0001 | 00000000 | 0000 | Block 0 | Hit |
| 0x104 | 0001 0000 0000 0010 | 0001 | 00000000 | 0010 | Block 0 | Hit |
| 0x200 | 0010 0000 0000 0000 | 0010 | 00000000 | 0000 | Block 0 | Miss |
Average Memory Access Time (AMAT): For 1ns hit time, 10ns miss penalty, 90% hit rate:
Label: Core, L1 I$/D$, L2, L3, main memory, and data flow arrows.
4. Multiprocessing: Parallelism in Action
Models & Synchronization
stateDiagram-v2
[*] --> Waiting
Waiting --> Running : "CPU assigned"
Running --> Blocked : "I/O or lock"
Running --> Finished : "Termination"
Blocked --> Waiting : "Resource available"
Finished --> [*]Key Concepts:
- SMP (Symmetric Multiprocessing): All cores share memory (e.g., Google servers).
- MPP (Massively Parallel): Distributed memory (e.g., supercomputers).
- Race Condition: When two threads access shared data without synchronization. Fix: Mutex locks, semaphores.
Worked Example: Bank Transaction (Race Condition)
Problem: Two customers deposit $100 each into an account with initial balance $0.
// Thread 1: Customer A
balance = read_memory(0x1000) // Reads $0
balance = balance + 100 // $100
write_memory(0x1000, balance) // Writes $100
// Thread 2: Customer B (same steps)
Outcome: Final balance = $100 (should be $200). Race condition!
Solution: Mutex Lock
lock.acquire()
balance = read_memory(0x1000) // $0
balance = balance + 100 // $100
write_memory(0x1000, balance) // $100
lock.release()
// Customer B waits until lock is free
Label: 4 cores, L2/L3 cache, integrated GPU, and memory controller.
In the Real World
eSewa’s Payment Terminals
- Topic: Embedded Systems + RISC Architecture
- How: Uses an ARM Cortex-M4 MCU to:
- Read NFC cards (via GPIO/UART).
- Encrypt transactions (AES hardware accelerator).
- Communicate with eSewa servers (Wi-Fi module).
- Why RISC? Low power, real-time response for payments.
Ncell’s 4G Base Stations
- Topic: Multiprocessing + Cache Hierarchy
- How:
- Dual-core ARM processors handle:
- Core 1: User data routing (L1 cache for low latency).
- Core 2: Backhaul to NTC (L2 cache for bulk transfers).
- SMP model ensures calls don’t drop during peak hours (e.g., 2077 BS).
- Dual-core ARM processors handle:
WhatsApp’s End-to-End Encryption
- Topic: RISC vs CISC in Mobile Chips
- How:
- Smartphones (RISC): ARM chips (e.g., Snapdragon 888) use NEON SIMD for fast AES encryption.
- Servers (CISC): Intel Xeon CPUs handle millions of messages via multithreading (hyper-threading).
NEPSE’s Stock Trading System
- Topic: Cache Memory + Multiprocessing
- How:
- L3 cache in trading servers reduces latency for high-frequency trades.
- MPP model: Separate nodes for order matching, user auth, and reporting.
Exam Tip
What Examiners Look For
Diagrams > Text:
- Draw block diagrams for embedded systems, state diagrams for multiprocessing, and cache mapping tables.
- Example: For RISC/CISC, show instruction count vs. clock cycles in a graph.
Real-world Mapping:
- eSewa → Embedded Systems (MCU + RTOS).
- Ncell → Multiprocessing (SMP for base stations).
- Bank loans → Cache hits/misses (e.g., "Why does a bank’s server need L3 cache?").
Common Pitfalls:
- Confusing RISC/CISC: Remember RISC = "simple but many instructions," CISC = "few but complex."
- Cache calculations: Always show hit/miss tables with binary addresses.
- Race conditions: Describe before/after with mutex locks.
Short-Answer Tricks:
- Define: "Cache associativity" → "Number of places a block can reside in a set."
- Compare: "RISC vs CISC" → Use a 2-column table (see above).
- Calculate: "AMAT" → Plug in hit rate and miss penalty.
Sample Exam Questions & Answers
Q1: Explain why embedded systems use Harvard architecture. Draw a block diagram. A:
Answer:
- Separate buses prevent von Neumann bottleneck (memory contention).
- Faster execution: Program and data can be fetched simultaneously.
- Example: eSewa’s payment terminal reads card data (Data Bus) while fetching instructions (Program Bus).
Q2: Calculate the number of sets in a 64KB, 4-way set-associative cache with 64-byte blocks. A:
- Total blocks: blocks.
- Blocks per set: 4 (4-way).
- Number of sets: sets.
Q3: How does WhatsApp use RISC architecture to encrypt messages faster? A:
- ARM NEON: SIMD (Single Instruction Multiple Data) unit in RISC chips processes 4 data items per instruction.
- Example: AES encryption on Snapdragon 888 uses NEON to encrypt 4 bytes at once (vs. 1 byte in scalar mode).
- Result: 4x speedup for end-to-end encryption.
Final Note:
- Memorize: RISC/CISC features, cache mapping types, and embedded system components.
- Practice: Draw diagrams for every concept. Examiners reward visual answers.
- Apply: Relate to Nepalese tech (eSewa, Ncell) and global examples (Google, WhatsApp).
Based on the TU BIT syllabus for Microprocessor and Computer Architecture (BIT151), unit 11.
Discussion
Loading…