CSC213 Computer Architecture

Computer ArchitectureUnit 511 min read

CISC vs. RISC: Architectures, Designs, and Performance Trade-offs

Unit 5 of Computer Architecture explores the core differences between Complex Instruction Set Computer (CISC) and Reduced Instruction Set Computer (RISC) architectures, their design philosophies, performance implications, and real-world applications in modern processors and systems. This note covers instruction set com

TAKEAWAYS:

  • CISC prioritizes complex, multi-cycle instructions (e.g., MUL, DIV) handled by microcode, while RISC uses simple, single-cycle instructions with hardware acceleration.
  • RISC achieves speed via pipelining, fixed-length instructions, and large register sets, while CISC relies on decoders and microprogrammed control units.
  • Register windows in RISC (e.g., SPARC) enable efficient procedure calls by overlapping local/global registers, reducing memory access.
  • Performance depends on context: CISC excels in legacy code, RISC in mobile/embedded systems (e.g., ARM in smartphones).
  • Modern CPUs (e.g., Intel’s x86-64, Apple’s M-series) blend CISC/RISC via decoding complex instructions into RISC-like micro-ops.
  • Exam focus: Compare architectures, explain pipelining gains, and link to real hardware (e.g., Intel Core vs. Raspberry Pi CPU).


1. Core Definitions: CISC vs. RISC

What Are They?

  • CISC (Complex Instruction Set Computer):

    • Design Philosophy: Fewer, complex instructions (e.g., REP MOVSB for block memory copy, FILD for floating-point load).
    • Hardware: Microprogrammed control unit decodes instructions into micro-ops.
    • Example: Intel x86 (Pentium, Core i7), AMD processors.
    • Key Feature: Variable-length instructions (1–15 bytes) with memory-to-memory operations (e.g., ADD [mem1], [mem2]).
  • RISC (Reduced Instruction Set Computer):

    • Design Philosophy: Simple, single-cycle instructions (e.g., ADD R1, R2, R3—only registers).
    • Hardware: Hardwired control unit, fixed-length instructions (e.g., 32-bit ARM), load-store architecture (no memory ops except LOAD/STORE).
    • Example: ARM (Apple M1, Qualcomm Snapdragon), MIPS, RISC-V.
    • Key Feature: Register-heavy design (e.g., ARM’s 31 general-purpose registers) to minimize memory access.

Visual: Instruction Set Complexity

mindmap
  root((CISC vs. RISC))
    CISC
      "Complex Instructions"
      "Microprogrammed Control"
      "Variable-Length (1-15 bytes)"
      "Memory-to-Memory Ops"
      "Example: Intel x86"
    RISC
      "Simple Instructions"
      "Hardwired Control"
      "Fixed-Length (e.g., 32-bit)"
      "Load-Store Architecture"
      "Example: ARM, MIPS"

2. Key Differences: A Comparison Table

Feature CISC RISC
Instruction Set Complex (e.g., MUL, DIV) Simple (e.g., ADD, SUB)
Instruction Length Variable (1–15 bytes) Fixed (e.g., 32/64-bit)
Addressing Modes Many (e.g., base+index) Few (e.g., register+offset)
Control Unit Microprogrammed Hardwired
Pipelining Limited (complex ops stall) Optimized (simple ops flow)
Registers Few (8–16) Many (32–128)
Memory Access Frequent (memory ops) Minimal (load-store only)
Performance Good for legacy code High throughput (mobile/embedded)
Examples Intel Core i9, AMD Ryzen Apple M1, Raspberry Pi 4

3. How They Work: Deep Dive

A. CISC: Microprogrammed Control

  1. Instruction Decoding:

    • A complex instruction (e.g., DIV) is broken into micro-instructions by a microprogram (firmware).
    • Example: DIV might take 20–30 micro-ops (fetch operands, perform division, store result).
    • IMAGE: microprogrammed control unit diagram | Block diagram of a microprogrammed control unit showing ROM, micro-instruction register, and control signals.
  2. Performance Trade-off:

    • Pros: Fewer instructions needed for complex tasks (e.g., REP MOVSB copies 1000 bytes in one instruction).
    • Cons: Longer execution time for complex ops (e.g., DIV stalls the pipeline).

B. RISC: Hardwired Control and Pipelining

  1. Fixed-Length Instructions:

    • All instructions are 32/64 bits (e.g., ARM’s ADD R0, R1, R2).
    • No memory-to-memory ops: Data moves via registers only (LOAD/STORE).
  2. Pipelining:

    • 5-stage pipeline (common in RISC):
      1. Fetch: Instruction from memory.
      2. Decode: Determine opcode/operands.
      3. Execute: ALU operation.
      4. Memory Access: LOAD/STORE.
      5. Writeback: Result to register.
    • Example: ARM Cortex-A76 completes 1 instruction per clock cycle (vs. CISC’s multi-cycle ops).
  3. Register Windows (Advanced RISC):

    • Problem: Procedure calls require saving/restoring registers (slow).
    • Solution: Overlapping register windows (e.g., SPARC’s 8 windows × 8 registers each).
    • How it works:
      • Global registers (shared across calls).
      • Local registers (for current procedure).
      • Input/output registers (for caller/callee).
    • Visual: Register Window Overlap
      stateDiagram-v2
        state "Global" as G
        state "Local" as L
        state "Input" as I
        state "Output" as O
        G --> L : "Procedure Call"
        L --> O : "Parameter Passing"
        O --> G : "Return"

4. Performance Analysis: Why RISC Wins in Modern Systems

A. Pipelining Gains

  • CISC Limitation: Complex instructions stall the pipeline (e.g., DIV takes 20 cycles).
  • RISC Advantage: Simple instructions flow smoothly through stages.
    • Example: ARM’s ADD takes 1 cycle; Intel’s ADD may take 1–3 cycles (depending on operands).

B. Real-World Example: Mobile Processors

  • Apple M1 (RISC-based ARM):
    • Uses 5-stage pipeline + out-of-order execution.
    • Result: 3.5× faster than Intel Core i7 in single-threaded tasks (e.g., compiling code).
  • Intel Core i9 (CISC):
    • Uses hybrid approach: Decodes x86 (CISC) into RISC-like micro-ops (e.g., ADD → ADD R1, R2, R3).
    • Result: Strong in multi-threaded tasks (e.g., video editing).

C. Worked Example: Loop Performance

Scenario: Copy 1000 integers from array1 to array2.

  • CISC (x86):
    REP MOVSB   ; Copies 1000 bytes in ~10 cycles (optimized by CPU).
    
  • RISC (ARM):
    LOOP:
      LDR R1, [array1], #4  ; Load + increment pointer
      STR R1, [array2], #4  ; Store + increment pointer
      SUBS R0, R0, #1       ; Decrement counter
      BNE LOOP              ; Branch if not zero
    
    • Cycles: ~4000 (1000 × 4 cycles per iteration).
    • Optimization: ARM’s loop unrolling or NEON SIMD can reduce this to ~500 cycles.

5. Real-World Applications

## In the Real World

  1. eSewa (Nepal):

    • RISC in Action: eSewa’s backend servers use ARM-based AWS Graviton processors (RISC) for cost-efficient, high-throughput transaction processing.
    • Why RISC? Low power consumption (critical for data centers in Kathmandu’s heat).
  2. Pathao (Ride-Hailing App):

    • Mobile CPUs: Pathao’s Android app runs on Qualcomm Snapdragon (ARM RISC).
    • Key Feature: Register windows in Snapdragon’s DSP (Digital Signal Processor) optimize GPS/location calculations.
  3. NTC’s Network Routers:

    • CISC in Routers: Older Cisco routers use x86 (CISC) for complex routing protocols (e.g., BGP).
    • RISC Shift: Newer models (e.g., Cisco’s Silicon One) use RISC-V for faster packet processing.
  4. Nepal Rastra Bank’s Core Banking:

    • Hybrid Approach: Banks use Intel Xeon (CISC) for legacy systems but ARM-based servers (RISC) for new microservices (e.g., mobile banking APIs).
  5. YouTube (Global):

    • Video Encoding: Uses ARM-based Google TPUs (RISC) for real-time video transcoding (e.g., converting 4K to 1080p).
    • Why? TPUs have custom RISC pipelines optimized for matrix math (used in video compression).

6. Modern CPUs: Blurring the Lines

A. Intel’s "RISC-ification" of x86

  • Problem: x86 (CISC) was slow for pipelining.
  • Solution: Decoding complex instructions into RISC-like micro-ops.
    • Example: MOV [mem], [mem] (CISC) → LOAD R1, [mem1]; STORE [mem2], R1 (RISC).
  • Result: Intel’s Skylake microarchitecture achieves 3–4 instructions per cycle.

B. Apple M1: RISC with CISC Tricks

  • ARM Core (RISC):
    • 8-core CPU with register renaming and out-of-order execution.
  • CISC-Like Features:
    • Supports x86 emulation (via Rosetta 2) for legacy apps.
    • Uses complex instructions (e.g., VMLA for vector math) but implements them efficiently.

Visual: Hybrid Architecture

flowchart TD
  A["CISC Instruction<br/>(e.g., MOVSB)"] --> B["Micro-op Decoder"]
  B --> C["RISC-like Micro-ops<br/>(ADD, LOAD, STORE)"]
  C --> D["5-Stage Pipeline<br/>(Fetch, Decode, Execute, Mem, WB)"]
  D --> E["Out-of-Order<br/>Execution Unit"]
  E --> F["Retirement<br/>(Commit to Arch. State)"]

7. Advantages and Disadvantages

Architecture Advantages Disadvantages
CISC - Fewer instructions for complex tasks. - Slow for pipelining.
- Backward compatibility (x86). - Higher power consumption.
- Good for legacy software. - Complex hardware (microprogrammed).
RISC - Faster execution (pipelining). - More instructions for complex tasks.
- Lower power (ideal for mobile). - Less backward compatibility.
- Simpler hardware (hardwired). - Requires more registers.

8. Exam Tip: How to Score Full Marks

  1. Compare CISC vs. RISC:

    • Use the table above as a template. Mention at least 4 differences (e.g., instruction length, control unit, pipelining).
    • Example Answer:

      "CISC uses variable-length instructions (e.g., 1–15 bytes) with memory-to-memory operations, while RISC employs fixed-length instructions (e.g., 32-bit) and a load-store architecture. CISC relies on microprogrammed control units, whereas RISC uses hardwired logic for faster decoding."

  2. Explain Pipelining:

    • Draw the 5-stage pipeline and explain hazards (e.g., data dependency stalls).
    • Example:

      "In RISC, pipelining allows one instruction per clock cycle. For example, ARM’s ADD R0, R1, R2 completes in 5 stages: Fetch (1), Decode (2), Execute (3), Memory (4), Writeback (5). CISC stalls here due to multi-cycle instructions like DIV."

  3. Register Windows:

    • Define overlapping windows and explain parameter passing.
    • Example:

      "SPARC’s register windows reduce context-switching overhead. When a function calls another, the caller’s output registers become the callee’s input registers, eliminating the need to save/restore registers to memory."

  4. Real-World Link:

    • Always tie theory to hardware (e.g., Intel vs. ARM, mobile vs. desktop).
    • Example:

      "The Raspberry Pi 4 uses ARM’s RISC architecture for its low power consumption (3W vs. 65W for an Intel i7), making it ideal for embedded systems like home automation controllers."

  5. Avoid Common Mistakes:

    • ❌ "RISC is always faster than CISC." → Context matters: CISC excels in legacy code.
    • ❌ "CISC has more registers." → False: RISC has more registers (e.g., ARM’s 31 vs. x86’s 16).
    • ✅ Focus on trade-offs: "RISC sacrifices instruction complexity for speed, while CISC prioritizes flexibility for legacy systems."

Based on the TU BSc CSIT syllabus for Computer Architecture (CSC213), unit 5.

Discussion

Loading…