CACS155 Microprocessor and Computer Architecture

Microprocessor and Computer ArchitectureUnit 1219 min read

Review & Practical Applications of Microprocessors & Architecture

Unit 12 of Microprocessor and Computer Architecture reviews all key concepts (8085 architecture, instruction cycles, memory hierarchy, pipelining, RISC/CISC) through practical applications, real-world examples, and exam-focused problem-solving techniques.

Core Concepts Reviewed in Unit 12

This unit consolidates all previous units by:

  1. Mapping theory to real systems (how microprocessors power apps like eSewa or Ncell).
  2. Comparing architectures (8085 vs. modern CPUs, RISC vs. CISC).
  3. Tracing execution (step-by-step instruction cycles, pipelining bottlenecks).
  4. Design trade-offs (hardwired vs. microprogrammed control, cache vs. RAM).
  5. Practical programming (assembly code for real tasks like sensor data processing).

1. Real-World Applications of Microprocessor Concepts

Microprocessors and architecture principles are invisible but critical in everyday systems. Here’s how they work behind the scenes:

A. eSewa and Khalti: Secure Payment Processing

  • Concept: Instruction cycles, pipelining, and memory hierarchy When you pay a bill via eSewa, your phone’s microprocessor executes thousands of instructions in parallel (pipelining) to:
    1. Encrypt your card details (using cryptographic instructions in the CPU).
    2. Fetch transaction data from RAM (L1/L2 cache → main memory).
    3. Send the request to eSewa’s server (network stack handled by the CPU’s DMA controller).
  • Why it matters:
    • Pipelining speeds up transaction validation (e.g., checking your balance).
    • Cache memory reduces latency when fetching user profiles.
    • Hardwired control units in modern CPUs execute security checks faster than microprogrammed ones.

B. Pathao/Ncell Ride Booking: Real-Time Scheduling

  • Concept: Interrupts, priority queues, and parallel processing When you book a ride on Pathao:
    1. Your phone’s ARM Cortex-A processor (RISC architecture) sends a GPS request via Wi-Fi (handled by the CPU’s DMA controller).
    2. The Pathao server’s multi-core CPU uses interrupts to prioritize your request over background tasks (e.g., driver updates).
    3. The server’s memory hierarchy (L3 cache → SSD) stores driver locations for fast access.
  • Worked Example: Suppose Pathao’s server has a 4-core CPU with hyper-threading. If 1000 users request rides simultaneously:
    • Without pipelining: Each request would take ~500 cycles (serial processing).
    • With pipelining + multi-core: The same task takes ~125 cycles per core (4× speedup).
    • Real-world impact: Fewer "driver not found" errors during peak hours.

C. Daraz/NTC: Order Fulfillment and Network Routing

  • Concept: Bus architectures, memory-mapped I/O, and protocol stacks When you order from Daraz:
    1. Your 8086-compatible CPU (or modern x86) sends an HTTP request via the PCIe bus to your Wi-Fi card.
    2. The request travels through routers (like NTC’s backbone network), where each hop uses memory-mapped I/O to forward packets.
    3. Daraz’s database (likely SQL + Redis cache) uses indexed memory access to fetch your order status.
  • Visual: Network Packet Forwarding
    sequenceDiagram
      participant You as "Your PC (x86 CPU)"
      participant Router1 as "NTC Router (MIPS CPU)"
      participant Daraz as "Daraz Server (ARM CPU)"
      You->>Router1: HTTP Request (TCP/IP stack)
      Router1->>Daraz: Forwarded Packet (Memory-mapped I/O)
      Daraz-->>Router1: Response (L1 Cache → RAM)
      Router1-->>You: Acknowledged (Interrupt-driven)

D. Banks (NMB, Global IME): Loan Interest Calculation

  • Concept: Fixed-point arithmetic, ALU operations, and microprogramming When a bank calculates monthly loan installments, the microprocessor performs:
    1. Fixed-point multiplication (for interest rates like 8.5%).
    2. ALU operations to compute (principal + interest) / term.
    3. Microprogrammed control (in older systems) to handle edge cases (e.g., partial payments).
  • Worked Example: For a ₹500,000 loan at 8.5% for 5 years:
    • Monthly interest rate = .
    • 8085 assembly snippet (simplified):
      MVI A, 500000/256    ; Load principal (high byte)
      MVI B, 85            ; Load interest rate (8.5%)
      CALL MULTIPLY        ; Microprogrammed routine
      DCR C                ; Decrement for monthly rate
      
    • Real-world impact: Banks use floating-point units (FPUs) in modern CPUs to avoid rounding errors.

2. Comparing Architectures: 8085 vs. Modern CPUs

Feature 8085 Microprocessor (1976) Modern x86/ARM (2020s)
Architecture CISC (Complex Instruction Set) Mostly RISC (x86 has CISC legacy)
Word Size 8-bit (16-bit with external bus) 32/64-bit (x86-64, ARMv8)
Clock Speed 3 MHz 2–5 GHz (with turbo boost)
Pipelining No (sequential execution) 12–16 stage pipelines (e.g., Intel Core)
Cache None L1 (32–64 KB), L2 (256 KB–1 MB), L3 (shared)
Control Unit Hardwired (for most instructions) Hybrid (hardwired + microprogrammed for complex ops)
Memory Access 64 KB addressable (2^16) 64-bit: 16 EB (2^64) addressable
Real-World Use Embedded systems, old PCs Smartphones (ARM), PCs (x86), servers

Key Takeaway:

  • The 8085’s simplicity made it easy to program but slow for modern tasks.
  • Modern CPUs use pipelining, superscalar execution, and out-of-order processing to handle billions of instructions per second.

3. Instruction Execution: Tracing a Real Example

Let’s trace how an 8085 microprocessor executes:

MVI A, 05H    ; Move immediate value 05H to accumulator
ADD B        ; Add register B to accumulator
STA 2050H    ; Store result at memory location 2050H

Step-by-Step Execution (Instruction Cycle, Machine Cycle, T-States)

Step Action T-States Machine Cycles
Fetch MVI A, 05H Fetch opcode MVI from memory (address PC) 4 1 (Opcode Fetch)
Fetch operand 05H from next memory location 3 1 (Operand Fetch)
Execute MVI Load 05H into accumulator (A) 7 1 (Memory Write)
Fetch ADD B Fetch ADD opcode 4 1 (Opcode Fetch)
Execute ADD Add B to A (ALU operation) 4 1 (Memory Read)
Fetch STA 2050H Fetch STA opcode 4 1 (Opcode Fetch)
Execute STA Store A at 2050H (address calculation + memory write) 13 3 (Memory Read, Address Calculation, Memory Write)
Total 43 8

Visual: 8085 Instruction Cycle

stateDiagram-v2
    [*] --> FetchOpcode: T1-T4
    FetchOpcode --> Decode: T5-T6
    Decode --> FetchOperand: T7-T9 (if needed)
    FetchOperand --> Execute: T10-T13
    Execute --> [*]

Real-World Tie-In:

  • This is how eSewa’s server processes your payment request in microseconds, but scaled up with multi-core CPUs and SIMD instructions (Single Instruction Multiple Data).

4. Memory Hierarchy: Why Your Phone Doesn’t Freeze

Modern systems use a memory hierarchy to balance speed and cost. Here’s how it works in a smartphone (e.g., Samsung Galaxy):

Level Type Size Speed (ns) Cost per Bit Used For
L1 Cache SRAM 32–64 KB 0.5–1 High Frequently used instructions/data
L2 Cache SRAM 256 KB–1 MB 2–4 Medium Medium-access data
L3 Cache SRAM (shared) 1–8 MB 10–20 Low Multi-core communication
RAM DRAM 4–8 GB 50–100 Very Low Active apps, OS
Storage eMMC/NVMe SSD 64 GB–1 TB 10,000+ Very Low Apps, photos, OS

Example: Loading a WhatsApp Message

  1. Your phone’s ARM Cortex CPU checks L1 cache for the message.
  2. If not found (cache miss), it fetches from L2 cache (still fast).
  3. If still missing, it loads from RAM (slower but cheaper).
  4. If the message is in storage, the CPU uses DMA to transfer it to RAM without stalling.

Visual: Memory Access Latency

pie
    title Memory Access Time Breakdown
    "L1 Cache Hit" : 0.5
    "L2 Cache Hit" : 2
    "RAM Access" : 50
    "Storage Access" : 10000

5. Pipelining: How Modern CPUs Do More Work

Problem with 8085:

  • No pipelining → CPU stalls while waiting for memory/data.
  • Example: Fetching an instruction takes 4 T-states, but the ALU is idle.

Solution: Pipelining (Modern CPUs) Divide instruction execution into 5 stages:

  1. Fetch (get opcode from memory)
  2. Decode (determine operation)
  3. Execute (ALU operation)
  4. Memory Access (load/store)
  5. Writeback (update registers)

Visual: 5-Stage Pipeline

gantt
    title 5-Stage Pipeline Execution
    dateFormat  YYYY-MM-DD
    section Instruction 1
    Fetch :a1, 2023-01-01, 2d
    Decode :a2, 2023-01-02, 2d
    Execute :a3, 2023-01-03, 2d
    Memory :a4, 2023-01-04, 2d
    Writeback :a5, 2023-01-05, 2d
    section Instruction 2
    Fetch :b1, 2023-01-02, 2d
    Decode :b2, 2023-01-03, 2d
    Execute :b3, 2023-01-04, 2d
    Memory :b4, 2023-01-05, 2d
    Writeback :b5, 2023-01-06, 2d

Real-World Impact:

  • Without pipelining: 1 instruction per 43 T-states (8085).
  • With pipelining: 1 instruction per 1 T-state (theoretical max, but real CPUs achieve ~3–5 instructions/cycle).

Example: YouTube Video Playback

  • Your phone’s ARM CPU uses pipelining to:
    1. Fetch video frames from storage (DMA).
    2. Decode frames (NEON SIMD instructions).
    3. Render to screen (GPU offloading).
  • Result: Smooth 60 FPS playback even on a mid-range phone.

6. RISC vs. CISC: Why Your Laptop Uses Both

Feature RISC (ARM, MIPS) CISC (x86, 8085)
Instruction Set Simple, fixed-length (e.g., ADD R1, R2) Complex, variable-length (e.g., MUL AX, BX)
Pipelining Optimized for pipelines Harder to pipeline (variable cycles)
Memory Access Load/store architecture (no memory ops in ALU) ALU can access memory directly
Clock Speed Higher (simpler instructions) Lower (complex decoding)
Power Efficiency Better (mobile devices) Worse (desktops/servers)
Example CPUs ARM Cortex, MIPS Intel Core, AMD Ryzen

Why x86 (CISC) Still Dominates Desktops:

  • Backward compatibility (old software still runs).
  • Microcode emulates RISC-like efficiency.
  • Complex instructions reduce code size (e.g., REP MOVSB moves blocks of memory in one instruction).

Example: Compiling C Code

  • A RISC compiler (ARM) generates simpler, pipelined code.
  • A CISC compiler (x86) may use complex instructions but still relies on pipelining for speed.

7. Control Unit Design: Hardwired vs. Microprogrammed

Feature Hardwired Control Unit Microprogrammed Control Unit
Implementation Direct logic gates Control store (ROM) + sequencer
Speed Faster (direct paths) Slower (fetch microinstructions)
Flexibility Hard to modify Easy to update (change microcode)
Complexity High (custom logic for each instruction) Lower (standardized microinstructions)
Used In High-performance CPUs (e.g., 8085) Complex CPUs (e.g., VAX, some x86)

Example: 8085 vs. Modern CPUs

  • 8085: Uses a hardwired control unit for most instructions (fast but inflexible).
  • Modern CPUs: Use a hybrid approach:
    • Simple instructions (e.g., ADD) → hardwired.
    • Complex instructions (e.g., FPU operations) → microprogrammed.

Visual: Control Unit Block Diagram


8. Practical Applications: Putting It All Together

A. Designing a Traffic Light Controller (8085-Based)

Requirements:

  • 3 traffic lights (red, yellow, green).
  • Timings: Green (30s), Yellow (5s), Red (25s).
  • Use 8085’s timer and I/O ports.

Solution:

  1. Hardware:
    • 8085 CPU + 8255 PPI (Parallel Port Interface) for lights.
    • 8253 Timer for delays.
  2. Software (Pseudocode):
    START: MVI A, 00000001b  ; Green for North-South
           OUT 80H           ; Send to PPI
           CALL DELAY_30S    ; 30s delay using 8253
           MVI A, 00000010b  ; Yellow for North-South
           OUT 80H
           CALL DELAY_5S
           MVI A, 00000000b  ; Red for North-South (East-West green)
           OUT 80H
           CALL DELAY_25S
           JMP START
    
  3. Real-World Use:
    • Kathmandu’s smart traffic lights use ARM-based microcontrollers with similar logic but add sensor inputs (for adaptive timing).

B. Optimizing a Database Query (SQL + CPU Caching)

Scenario:

  • A bank’s NMB server runs a query:
    SELECT * FROM accounts WHERE balance > 1000000;
    

How the CPU Helps:

  1. Indexed Memory Access:
    • The CPU’s MMU (Memory Management Unit) uses B-tree indexes to jump directly to high-balance accounts (avoiding full table scan).
  2. Cache Optimization:
    • Frequently accessed account records stay in L3 cache.
  3. SIMD Instructions:
    • Modern CPUs use AVX-512 to compare balances in parallel.

Performance Gain:

  • Without optimization: 100 ms (full scan).
  • With optimization: 5 ms (cached + indexed).

9. Common Pitfalls and Exam Tips

A. What Examiners Love to Test

  1. Instruction Cycle Traces:
    • Always show T-states and machine cycles for 8085 questions.
    • Example: For ADD B, explain:
      • Opcode fetch (4 T-states).
      • Execute (4 T-states).
      • Total: 8 T-states (but 8085 takes 7 for ADD—watch for exceptions!).
  2. Memory Hierarchy Calculations:
    • Questions may ask: "If L1 cache hit rate is 90% and L2 is 80%, what’s average access time?"
    • Formula:
  3. Pipelining Hazards:
    • Examiners ask about data hazards, control hazards, and structural hazards.
    • Example: ADD R1, R2; SUB R1, R3 has a data hazard (R1 is read before written).
  4. RISC vs. CISC Trade-offs:
    • Compare power efficiency (RISC wins for phones) vs. code density (CISC wins for legacy systems).
  5. Control Unit Design:
    • Know when to use hardwired (speed) vs. microprogrammed (flexibility).

B. Model Answer Structure for Long Questions

Question: "Explain the instruction execution cycle of the 8085 microprocessor with an example. How does pipelining improve performance?"

Model Answer:

  1. Introduction (1 mark): "The 8085 microprocessor executes instructions in cycles: instruction cycle, machine cycle, and T-states. Pipelining, used in modern CPUs, overlaps these stages to improve throughput."

  2. Instruction Cycle Breakdown (4 marks):

    • Fetch: Opcode from memory (4 T-states).
    • Decode: Determine operation (3 T-states for MVI).
    • Execute: Perform ALU operation (4 T-states for ADD).
    • Example Trace: Use MVI A, 05H; ADD B (show T-states as above).
  3. Pipelining Explanation (3 marks):

    • "In pipelining, stages overlap: while one instruction is in Execute, the next is being Fetched. This reduces idle time from 43 T-states per instruction (8085) to ~1 T-state per stage (modern CPUs)."
    • Draw a 5-stage pipeline diagram (as above).
  4. Real-World Link (1 mark): "eSewa’s servers use pipelined CPUs to process 10,000+ transactions/sec, while an 8085 would take minutes for the same workload."

C. Short-Answer Tips

  • Memory Hierarchy: Always mention SRAM vs. DRAM, cache levels, and latency trade-offs.
  • RISC/CISC: Compare instruction complexity, power use, and examples (ARM vs. x86).
  • Control Unit: Say hardwired is fast but rigid; microprogrammed is flexible but slow.

10. Summary Table: Key Concepts at a Glance

Topic Key Idea Real-World Example
8085 Instruction Cycle Opcode fetch → decode → execute (T-states, machine cycles) eSewa payment processing
Pipelining Overlapping instruction stages for parallelism YouTube video decoding
Memory Hierarchy L1 > L2 > RAM > Storage (speed vs. cost) WhatsApp message loading
RISC vs. CISC RISC: simple, pipelined; CISC: complex, legacy support ARM (phones) vs. x86 (PCs)
Control Unit Hardwired (fast) vs. microprogrammed (flexible) 8085 (hardwired) vs. modern hybrid CPUs
Interrupts Priority-based task switching Pathao ride request handling

Exam Tip

How to Score Full Marks in Unit 12

  1. For theoretical questions:

    • Define the concept (e.g., "Pipelining is a technique to overlap instruction execution stages to improve throughput.").
    • Draw a diagram (e.g., 5-stage pipeline, memory hierarchy).
    • Give a real-world example (e.g., "Modern CPUs use pipelining to decode and execute billions of instructions per second in apps like Daraz.").
  2. For numerical problems:

    • Show all steps (e.g., T-state calculations for 8085).
    • Use formulas (e.g., average memory access time).
    • Assume reasonable values if missing (e.g., "Assume L1 cache hit time = 1 ns").
  3. For comparisons (RISC vs. CISC, hardwired vs. microprogrammed):

    • Use a table (as above) with 3–4 clear points.
    • Link to real systems (e.g., "ARM’s RISC design saves battery in smartphones").
  4. For assembly/practical questions:

    • Write pseudocode first, then map to 8085 instructions.
    • Explain hardware interactions (e.g., "The 8255 PPI is used to control traffic lights").
  5. Common Mistakes to Avoid:

    • Forgetting T-states: Always count them for 8085 questions.
    • Ignoring hazards: In pipelining, mention data hazards if relevant.
    • Overlooking real-world ties: Examiners reward 1–2 marks for practical examples.

memory hierarchy diagramA layered model showing L1 cache, L2 cache, RAM, and storage with latency and size labels. (Image: ComputerMemoryHierarchy.png: User:Danlash at en.wikipedia.or, Public domain, via Wikimedia Commons)

Based on the TU BCA syllabus for Microprocessor and Computer Architecture (CACS155), unit 12.

Discussion

Loading…