Elective Computer Architecture

Computer ArchitectureUnit 1214 min read

Modern Processor Design & Case Studies: Cores, GPUs, RISC-V, ARM, x86, and Real-World Chips

Unit 12 of Computer Architecture explores how modern processors (CPUs, GPUs, NPUs) are designed—from multi-core architectures and out-of-order execution to specialized accelerators like TPUs and NPUs. It covers real-world case studies (Intel Core i9, Apple M-series, NVIDIA GPUs) and emerging trends (RISC-V, heterogeneo

TAKEAWAYS:

  • Multi-core vs. multi-processor: Understand how SMT (Simultaneous Multithreading) in Intel’s Hyper-Threading or ARM’s SMT differs from true multi-core, and why NUMA (Non-Uniform Memory Access) matters in servers like those used by NEPSE for stock trading.
  • Specialized accelerators: Learn how GPUs (NVIDIA A100) handle parallel tasks (e.g., Pathao’s route optimization) and NPUs (Apple’s Neural Engine) power on-device AI (e.g., Khalti’s fraud detection).
  • Pipeline hazards and solutions: Trace how branch prediction (used in Google’s Tensor Processing Units) and register renaming (Intel’s i9) resolve stalls in real pipelines.
  • Case studies: Compare ARM Cortex-A78 (in smartphones) vs. x86-64 (in Daraz’s servers) using Flynn’s Taxonomy and Amdahl’s Law.
  • RISC-V revolution: See why RISC-V is disrupting embedded systems (e.g., NTC’s IoT sensors) and how its open ISA enables custom cores.
  • Thermal and power limits: Calculate Dynamic Voltage and Frequency Scaling (DVFS) for a quad-core processor (e.g., cooling a laptop used for online exams).

1. Evolution of Processor Design: From Single-Core to Heterogeneous Systems

Modern processors are no longer just about raw clock speed. The shift toward multi-core, specialized accelerators, and energy efficiency defines today’s chips. Key trends:

  • Moore’s Law slowdown: Transistor scaling hit physical limits (~7nm), so designers added more cores (e.g., Intel’s 24-core i9-14900K) or accelerators (e.g., NVIDIA’s Hopper GPU for AI).
  • Heterogeneous computing: Combining CPUs (general-purpose), GPUs (parallel tasks), NPUs (AI), and DPUs (data processing). Example: A Daraz order server uses:
    • CPU for OS tasks,
    • GPU for real-time inventory updates,
    • NPU for recommendation algorithms.
CPU (General-Purpose)OS tasksGPU (Parallel)Graphics/AINPU (AI)Matrix opsDPU (I/O)Network/storage
Heterogeneous computing in a Daraz order server: task distribution across specialized cores.

2. Multi-Core and Multi-Threading Architectures

A. Single-Core vs. Multi-Core vs. Multi-Processor

Feature Single-Core Multi-Core (SMP) Multi-Processor (MPP)
Definition One ALU, one core Multiple cores on one die Multiple dies/sockets
Example Early Pentium Intel i5/i7/i9 Supercomputer (Ncell’s 5G)
Memory Access Uniform (UMA) UMA or NUMA NUMA (remote memory slower)
Scalability Limited Good for threads Best for distributed tasks
Use Case Embedded (RISC-V) Laptops, servers Cloud (Google, AWS)
Single CPU, multiple coresMultiple CPUs, shared memoryMultiple CPUs, distributed memorySingle-CoreMulti-CoreMulti-Processor
Evolution from single-core to multi-processor systems (simplified).

B. Simultaneous Multithreading (SMT)

  • How it works: A single core executes two threads simultaneously by interleaving instructions (e.g., Intel’s Hyper-Threading).
  • Advantage: Hides memory latency (e.g., while Thread 1 waits for data, Thread 2 runs).
  • Disadvantage: False sharing (threads modify same cache line) and resource contention (e.g., two threads competing for ALU).
  • Real-world example:
    • eSewa’s payment server: Uses SMT to handle 10,000+ concurrent transactions without stalls.
    • Worked example: If a core has 4 hardware threads and a task stalls for 100 cycles, the other 3 threads keep the core busy.
Cycle 0Core IdleCycle 10Thread 1dispatched (Stalls at Cycle 21SMT switches toThread 2 (Stalls at CyCycle 31SMT switches toThread 3 (Completes)Cycle 40Core Idle (Nexttask)
SMT in eSewa’s payment server: 4 hardware threads mask a 100-cycle stall.

C. NUMA (Non-Uniform Memory Access)

  • Problem: In multi-socket systems (e.g., a NEPSE trading server), memory is not equally fast.
    • Local memory: Accessed in ~50ns (same socket).
    • Remote memory: Accessed in ~200ns (other socket).
  • Solution: NUMA-aware OS (Linux, Windows) schedules threads to minimize remote access.
  • Example: A Daraz warehouse management system uses NUMA to:
    1. Place inventory threads on cores near their memory.
    2. Use NUMA balancing to migrate threads dynamically.

3. Specialized Accelerators: GPUs, NPUs, and DPUs

A. GPUs: From Graphics to AI

  • Key idea: SIMD (Single Instruction, Multiple Data) – thousands of cores execute the same operation on different data.
  • Example: NVIDIA A100 GPU in a Pathao server:
    • Task: Calculate 10,000 shortest paths for drivers.
    • GPU advantage: Processes 100x faster than a CPU.
    • CUDA cores: Handle floating-point math for route optimization.
CUDA Core 1,CUDA Core 2,CUDA Core 3,CUDA Core 40Thread Block 1,Thread Block 2,Thread Block 3,Thread Block 41SIMD Warp (32 threads),SIMD Warp (32 threads),SIMD Warp (32 2
NVIDIA A100 GPU core array: 6144 CUDA cores (simplified 4x4 grid) execute parallel threads for Pathao’s route optimization.

B. NPUs: AI at the Edge

  • Example: Apple A17 Pro Neural Engine (in iPhone 15 Pro):
    • Task: Run on-device fraud detection for Khalti payments.
    • How it works:
      1. INT8 quantization: Reduces AI model size (e.g., 1GB → 128MB).
      2. Matrix math: 16x16 SIMD units for fast inference.
    • Power savings: Uses 1/10th the energy of a CPU for the same task.

C. DPUs: Offloading I/O

  • Example: NVIDIA BlueField DPU in a Ncell 5G base station:
    • Task: Handle 10Gbps network traffic without CPU load.
    • Features:
      • SmartNIC: Processes packets (firewall, encryption) in hardware.
      • ARM cores: Run lightweight OS (e.g., Linux) for control.

4. Pipeline and Superscalar Techniques in Modern CPUs

A. Out-of-Order Execution (OoOE)

  • Problem: Instructions may not execute in order due to data hazards (e.g., ADD R1, R2, R3 depends on MUL R2, R4, R5).
  • Solution: Reorder Buffer (ROB) and Register Renaming (Intel’s i9).
  • Example: In a bank’s loan approval system:
    • Code:
      LOAD R1, [credit_score]  ; Hazard: depends on next instruction
      ADD R2, R1, 100           ; Stalls if LOAD not ready
      STORE R2, [approved]
      
    • OoOE fix: The CPU executes ADD speculatively while LOAD is pending.

B. Branch Prediction

  • Problem: Branches (IF, LOOP) cause pipeline stalls if mispredicted.
  • Solutions:
    1. Static prediction: Assume branches are not taken (simple but inaccurate).
    2. Dynamic prediction: Use Branch Target Buffer (BTB) (Intel’s i9 has a 4KB BTB).
  • Example: Google’s Tensor Processing Unit (TPU):
    • Uses perfect branch prediction for AI loops (e.g., matrix multiplication).

C. Pipeline Hazards and Solutions

Hazard Type Cause Solution Example (Intel i9)
Structural Resource conflict (e.g., two ALUs) Duplicate resources (e.g., 2 FPUs) i9 has 2x FPUs for SIMD
Data Read-after-write (RAW) Forwarding (bypass) ADD gets result from MUL early
Control Branch misprediction Delayed branching + BTB i9’s 256-entry BTB

5. Case Studies: Real-World Processors

A. Intel Core i9-14900K (Desktop)

  • Architecture: Hybrid (P-cores + E-cores)
    • P-cores (8): High-performance, OoOE, 32KB L1, 1MB L2.
    • E-cores (16): Efficiency, in-order, 128KB L2.
  • Use case: NEPSE trading terminal (low-latency order matching).
  • Key feature: Thread Director (OS schedules threads to P/E-cores dynamically).
Performance Cores (P-cores)High single-thread perfEfficiency Cores (E-cores)Energy-efficient multithreading
Intel’s hybrid architecture: P-cores vs. E-cores in the i9-14900K.

B. Apple M3 (ARM-based)

  • Architecture: Unified cores (no P/E split) with high IPC (Instructions Per Cycle).
  • Use case: Khalti’s fraud detection (low power, high efficiency).
  • Key feature: Neural Engine (16-core NPU for on-device ML).

C. NVIDIA Grace-Hopper (Supercomputer)

  • Architecture: CPU (Grace) + GPU (Hopper) in one package.
  • Use case: Weather forecasting (NTC’s climate models).
  • Key feature: NVLink (900GB/s bandwidth between CPU/GPU).

6. RISC-V: The Open ISA Revolution

A. Why RISC-V?

  • Problems with x86/ARM:
    • x86: Complex, proprietary (Intel/AMD).
    • ARM: Licensing fees (~$50M for high-end cores).
  • RISC-V advantages:
    • Open ISA: Free to use/modify.
    • Customizable: Add extensions (e.g., Zicsr for security, Zfinx for FPU).
    • Low power: Used in NTC’s IoT sensors and Raspberry Pi.

B. RISC-V Extensions

Extension Purpose Example Use Case
M Integer multiply/divide Financial calculations (NEPSE)
F Single-precision FPU Graphics (Pathao maps)
A Atomic operations Thread-safe databases
V Vector (SIMD) AI inference (Khalti)

C. Example: RISC-V in a Smart Meter

  • Hardware: SiFive FE310 (32-bit RISC-V).
  • Tasks:
    1. Read voltage/current sensors.
    2. Run SHA-256 (for secure NTC billing).
    3. Send data via LoRaWAN.
  • Code snippet (Assembly):
    # Read ADC (voltage)
    csrr a0, mcycle       # Get cycle count
    lw a1, 0x10000000     # Read ADC value
    add a2, a1, a0        # Combine for checksum
    

7. Thermal and Power Management

A. Dynamic Voltage and Frequency Scaling (DVFS)

  • How it works: Adjusts voltage (V) and clock speed (f) to save power.
    • P = C * V² * f (Power = Capacitance × Voltage² × Frequency).
  • Example: Laptop cooling during exams:
    • Idle: 1.2V @ 800MHz (low power).
    • Stress test: 1.4V @ 4.5GHz (high performance).

B. Thermal Throttling

  • Problem: If a core hits 100°C, performance drops.
  • Solutions:
    1. Throttle clock speed (e.g., i9 caps at 5.8GHz if hot).
    2. Park cores (e.g., Linux’s cpufreq governor).
  • Real-world: Ncell’s 5G base station shuts down non-critical cores during peak heat.

8. Flynn’s Taxonomy Revisited

Modern processors often combine categories. Examples:

Category Definition Example Processor Real-World Use Case
SISD Single instruction, single data Early Pentium (single-core) Legacy ATM machines
SIMD Single instruction, multiple data GPU (NVIDIA A100) Pathao’s route optimization
MISD Multiple instructions, single data Rare (e.g., pipeline stages) Not common
MIMD Multiple instructions, multiple data Multi-core (Intel i9) NEPSE trading system
SIMT SIMD + threading (GPU model) NVIDIA CUDA cores AI training (Google TPU)

Worked example:

  • Task: Render a 3D map for Pathao drivers.
  • Processor: NVIDIA RTX 4090 (SIMT).
  • Breakdown:
    • 1 instruction: "Render polygon."
    • 4096 threads: Each handles a pixel.
    • Result: 100x faster than a CPU.

In the Real World

  1. eSewa’s Payment Gateway:

    • Idea: NUMA architecture in their servers ensures low-latency access to transaction logs.
    • How: Threads handling Khalti payments are scheduled on cores local to their memory.
  2. Pathao’s Route Optimization:

    • Idea: GPU acceleration (SIMD) for calculating 10,000 shortest paths in real-time.
    • How: NVIDIA’s CUDA cores process each path in parallel, reducing response time from 5s → 50ms.
  3. Ncell’s 5G Base Stations:

    • Idea: Heterogeneous computing (CPU + DPU) offloads network tasks.
    • How: NVIDIA BlueField DPU handles firewall rules and encryption without burdening the CPU.
  4. Khalti’s Fraud Detection:

    • Idea: NPU (Neural Engine) runs on-device AI models.
    • How: Apple’s A17 Pro NPU detects fraudulent transactions in <10ms using INT8 quantization.
  5. NEPSE’s Trading System:

    • Idea: Out-of-order execution (OoOE) and NUMA for low-latency order matching.
    • How: Intel’s i9-14900K processes 10,000 orders/sec with minimal stalls.

Exam Tip

  1. Case studies are key: Always relate answers to real processors (Intel i9, ARM Cortex, RISC-V, NVIDIA GPUs). Example:

    • "Explain pipeline hazards" → "In Intel’s i9, structural hazards are avoided by duplicating ALUs, while data hazards use forwarding."
  2. Compare architectures:

    • Multi-core vs. multi-processor: Use a table (as above) and cite NUMA vs. UMA.
    • SMT vs. multi-core: Explain Hyper-Threading (Intel) vs. true multi-core (ARM).
  3. Math is critical:

    • Speedup: For a 4-stage pipeline, max speedup = 4 (if no hazards).
    • Power: Use P = C * V² * f to compare DVFS states.
  4. Draw diagrams:

    • Pipeline stages: Label IF, ID, EX, MEM, WB.
    • NUMA system: Show local vs. remote memory latency.
    • GPU core array: Draw CUDA cores + SIMD warps.
  5. RISC-V is hot:

    • Expect questions on custom extensions (e.g., "How would you add a cryptography extension to RISC-V?").
    • Mention SiFive, Alibaba’s XuanTie, and NTC’s IoT sensors.
  6. Common pitfalls:

    • ❌ "SMT = multi-core" → No! SMT is one core, multiple threads.
    • ❌ "All hazards are resolved by forwarding" → No! Control hazards need branch prediction.
    • ❌ "GPUs are only for graphics" → No! They’re for AI, databases, and parallel math.

Based on the PU BE Computer (PU) syllabus for Computer Architecture, unit 12.

Discussion

Loading…