BIT151 Microprocessor and Computer Architecture

Microprocessor and Computer ArchitectureUnit 712 min read

Memory Hierarchy & Organization: Levels, Speed, and Trade-offs

Unit 7 of Microprocessor and Computer Architecture: explores how computers organize memory into a multi-level hierarchy (registers → cache → RAM → disk) to balance speed, cost, and capacity, explains memory mapping, addressing, and access times, and contrasts volatile vs. non-volatile storage with real-world examples l

TAKEAWAYS:

  • Memory hierarchy prioritizes speed over cost by using smaller, faster layers (registers, cache) for hot data and larger, slower layers (disk) for cold data.
  • Cache memory (L1, L2, L3) reduces CPU stalls by storing frequently accessed instructions/data, but requires associative mapping or set-associative schemes.
  • Virtual memory extends RAM using disk (swap space) but introduces page faults, increasing latency.
  • Memory-mapped I/O simplifies hardware access by treating peripherals as memory locations, unlike isolated I/O’s dedicated ports.
  • DRAM vs. SRAM: DRAM is cheaper but slower (used in RAM), while SRAM is faster (used in cache) but consumes more power.
  • Memory bandwidth (GB/s) and latency (ns) are critical for performance; pipelining and prefetching mitigate bottlenecks.

1. Memory Hierarchy: Why Multiple Layers?

Computers use a multi-level memory hierarchy to balance speed, cost, and capacity. The closer memory is to the CPU, the faster but more expensive it is. The goal is to minimize average access time while keeping costs low.

CPU RegistersL1 CacheL2 CacheL3 CacheRAMSSDHDDfaster, costlier
Speed vs. cost trade-off in memory hierarchy (logarithmic scale implied)

The Memory Hierarchy (From Fastest to Slowest)

figure: Memory Hierarchy
| Level       | Type          | Size (bits) | Speed (ns) | Cost (per bit) | Example Use Case                     |
|-------------|---------------|-------------|------------|----------------|-------------------------------------|
| Registers   | SRAM          | 32–64       | 0.1–1      | $$$$$$         | CPU registers (AX, BX, PC)         |
| L1 Cache    | SRAM          | KB–MB       | 1–5        | $$$$           | Frequently used instructions/data   |
| L2 Cache    | SRAM          | MB–16MB     | 5–20       | $$$            | Branch prediction, loop variables   |
| L3 Cache    | SRAM          | 16MB–100MB  | 20–50      | $$             | Shared across CPU cores            |
| Main RAM    | DRAM          | GB–TB       | 50–150     | $              | Running programs, OS data           |
| SSD         | Flash         | TB–PB       | 100K–1M    | $              | Boot OS, apps, user files           |
| HDD         | Magnetic      | PB+         | 10M–20M    | $              | Backups, large datasets             |

Why this hierarchy?

  • Locality Principle: Programs access the same data repeatedly (temporal locality) or nearby data (spatial locality).
  • Trade-off: Faster memory is smaller and expensive; slower memory is larger and cheap.

computer memory hierarchy diagramA layered cake of CPU registers (top), L1/L2/L3 cache, RAM, SSD, and HDD (bottom). (Image: ComputerMemoryHierarchy.png: User:Danlash at en.wikipedia.or, Public domain, via Wikimedia Commons)


2. Cache Memory: How It Works

Cache memory is a small, fast SRAM layer between the CPU and RAM. It stores copies of frequently used data to reduce access time.

Cache Organization

  1. Direct-Mapped Cache:
    • Each memory block maps to one cache line.
    • Simple but can cause conflicts (thrashing).
0[object Object][object Object]1—2[object Object]3[object Object]
Direct-mapped cache: 4 sets, 1-way associative (Block A and D conflict in Set 0)

Example: If two blocks hash to the same line, the LRU (Least Recently Used) policy evicts the older one.

  1. Associative Cache:
    • Any memory block can go to any cache line.
    • No conflicts but slower (requires comparison).
[object Object][object Object][object Object]Memory Block XCache Line 0Cache Line 1Cache Line 2
Associative cache: Full search for any block (no hashing)
  1. Set-Associative Cache (Most common):
    • Cache divided into sets (e.g., 4-way set-associative).
    • Each memory block maps to a set, not a single line.
0[object Object][object Object]1[object Object]
4-way set-associative cache: 2 sets, 4 lines each (Block A/B in Set 0)

Cache Hit/Miss

  • Hit: Data is in cache → fast access (~1–5 ns).
  • Miss: Data not in cache → RAM access (~50–150 ns).
  • Miss Penalty: The time lost when cache misses occur.

Worked Example: Cache Hit/Miss Calculation

  • Cache size: 16 KB, block size: 32 bytes.
  • Access time: RAM = 100 ns, Cache = 5 ns.
  • Hit rate: 90% (90% of accesses are cache hits).
  • Average access time (AAT):

In the Real World:

  • Ncell’s App Caching: When you open WhatsApp repeatedly, Ncell’s L3 cache stores the app’s binary to speed up launches.
  • Google Chrome’s Disk Cache: Frequently visited websites (e.g., Daraz) are cached in SSD to reduce load times.

3. Main Memory (RAM): DRAM vs. SRAM

RAM is volatile (loses data when powered off) and used for active programs.

DRAM (Dynamic RAM)

  • Cheaper, larger (used in main memory).
  • Requires refreshing (every ~8 ms).
  • Slower (~50–150 ns access time).
  • IMAGE: "DRAM chip close-up" | A DRAM chip showing memory cells with capacitors.

SRAM (Static RAM)

  • Faster, no refresh needed (used in cache).
  • More power-hungry (6 transistors per bit vs. 1 in DRAM).
  • IMAGE: "SRAM vs DRAM comparison" | A side-by-side of SRAM (top) and DRAM (bottom) cell structures.

Memory Addressing

  • Linear Addressing: Each memory location has a unique address (e.g., 32-bit address bus → 4 GB RAM).
  • Segmented Addressing: Divides memory into segments (e.g., code, data, stack).
  • Paged Addressing: Uses pages (fixed-size blocks) for virtual memory.

Example: 32-bit Address Bus

  • Address range: .
  • If word size = 32 bits, then number of words = 4 GB / 4 bytes = 1 GB words.

4. Virtual Memory: Extending RAM with Disk

Virtual memory simulates more RAM than physically available by using disk (swap space).

How It Works

  1. Page Fault: CPU requests a page not in RAM → OS loads it from disk.
  2. Page Replacement: If RAM is full, the OS evicts a page (using algorithms like LRU, FIFO).
  3. Thrashing: Too many page faults → CPU spends more time waiting for disk than executing.

Page Replacement Algorithms

Algorithm Description Example
FIFO Replace oldest page Not optimal for locality
LRU Replace least recently used Best for temporal locality
Optimal Replace page not used for longest time Impossible to implement
Clock Circular buffer with "used" bit Used in modern OS

Worked Example: Page Fault Calculation

  • RAM size: 4 pages, Disk access time: 10 ms, CPU time per page: 1 ms.
  • Page reference string: [1, 2, 3, 4, 1, 2, 5, 1, 2, 3, 4, 5]
  • Algorithm: LRU
  • Page faults:
    • Load 1, 2, 3, 4 → 4 faults.
    • Replace 4 with 5 → 1 fault.
    • Replace 3 with 1 → 1 fault.
    • Total faults = 6.
  • Total time:

In the Real World:

  • NEPSE’s Market Data: When traders open NEPSE’s trading app, the OS pages in only the necessary market data into RAM, keeping other data on SSD to save space.
  • Pathao’s Ride Booking: When you book a ride, Pathao’s backend caches frequently used driver locations in L3 cache to reduce latency.

5. Secondary Storage: SSD vs. HDD

Secondary storage is non-volatile (retains data when powered off).

Feature SSD (Solid State Drive) HDD (Hard Disk Drive)
Technology Flash memory (NAND) Magnetic platters
Speed 100K–1M IOPS 50–200 IOPS
Latency ~20–100 µs ~5–10 ms
Reliability No moving parts → more durable Moving parts → prone to failure
Cost $ (per GB) $ (per GB)
Use Case Boot OS, apps, databases Backups, large files

In the Real World:

  • eSewa’s Transaction Logs: eSewa stores millions of transactions on SSDs for fast retrieval during audits.
  • Daraz’s Order Database: Daraz uses HDDs for archived orders (cost-effective) and SSDs for active orders (fast access).

6. Memory-Mapped I/O vs. Isolated I/O

I/O devices (keyboard, disk) interact with the CPU via memory or dedicated ports.

08162431Address[31:2]30 bitsDeviceSelect1 bitsRegister Select1 bits
Memory-mapped I/O address decoding (simplified 32-bit example)
Feature Memory-Mapped I/O Isolated I/O
Address Space Part of main memory Dedicated I/O ports
Access Method MOV AX, [0x200] IN AX, 0x200
Simplicity Easier (uses MOV) More complex (requires IN/OUT)
Speed Faster (no port mapping) Slower (extra cycle)
Use Case Modern systems (PCIe) Legacy systems (8085)

Example: 8085 Microprocessor

  • Isolated I/O: Uses IN/OUT instructions.
    IN  AL, 0x20  ; Read from port 0x20
    OUT 0x20, AL  ; Write to port 0x20
    
  • Memory-Mapped I/O: Treats I/O as memory.
    MOV AL, [0x200]  ; Read from address 0x200
    MOV [0x200], AL  ; Write to address 0x200
    

In the Real World:

  • NTC’s Router Configuration: NTC’s routers use memory-mapped I/O to read/send network packets efficiently.
  • Khalti’s Payment Gateway: When you pay via Khalti, the payment request is sent via memory-mapped I/O to the payment server.

7. DMA (Direct Memory Access)

DMA allows peripherals (disk, network card) to transfer data directly to/from RAM without CPU intervention.

How DMA Works

  1. CPU sends DMA request to peripheral.
  2. DMA controller takes over the bus.
  3. Data transfers directly to RAM.
  4. CPU resumes when transfer is done.

Advantages:

  • Reduces CPU load (no need to handle every byte).
  • Faster I/O (e.g., disk reads).

Disadvantages:

  • Complexity (requires DMA controller).
  • Bus contention (DMA may slow down CPU).

In the Real World:

  • Ncell’s 5G Base Station: Uses DMA to transfer large amounts of user data (e.g., video calls) without bogging down the CPU.
  • YouTube Video Downloads: When you download a video, the network card uses DMA to write data directly to your SSD.

Exam Tip: How to Score Full Marks

  1. Define Key Terms Clearly

    • Always start with definitions (e.g., "Cache is a small, fast SRAM layer...").
    • Use figures (e.g., memory hierarchy table, cache mapping diagrams).
  2. Compare and Contrast

    • For DRAM vs. SRAM, memory-mapped vs. isolated I/O, or SSD vs. HDD, use a table with 3–4 columns (Feature, DRAM, SRAM, Comparison).
  3. Solve Numerical Problems

    • Practice cache hit/miss calculations and page fault analysis.
    • Show step-by-step working (e.g., "Hit rate = 80%, Miss rate = 20%").
  4. Real-World Applications

    • Link concepts to Nepali apps (e.g., "Khalti uses memory-mapped I/O for fast transactions").
    • Mention performance trade-offs (e.g., "SSDs are faster but more expensive than HDDs").
  5. Diagrams > Words

    • Always draw:
      • Memory hierarchy layers.
      • Cache mapping (direct, associative, set-associative).
      • DMA data transfer flow.
      • Page replacement algorithms (LRU, FIFO).
  6. Avoid Common Mistakes

    • ❌ Saying "Cache is faster than RAM" → Correct: "Cache is smaller and faster than RAM."
    • ❌ Mixing up DRAM and SRAM → Always specify which is used where.
    • ❌ Forgetting volatility → RAM is volatile; SSD/HDD are non-volatile.

Final Note: This unit is heavily exam-focused—expect 10–15 marks for definitions, 5–10 marks for comparisons, and 5–10 marks for numerical problems. Practice drawing diagrams and solving cache/page fault examples to master it.

Based on the TU BIT syllabus for Microprocessor and Computer Architecture (BIT151), unit 7.

Discussion

Loading…