Elective Computer Hardware Design

Computer Hardware DesignUnit 410 min read

The Processor: ALU, Control Unit, Registers & Pipeline

Unit 4 of Computer Hardware Design explores the core of the CPU—the Arithmetic Logic Unit (ALU), control unit, registers, instruction execution, and pipelining—with real-world examples from Nepali tech (e.g., Ncell billing systems, eSewa transactions) and visual breakdowns of how processors fetch, decode, and execute i

TAKEAWAYS:

  • The ALU performs arithmetic/logic operations (e.g., ADD, AND, NOT) using combinational circuits (full adders, multiplexers) and is controlled by the control unit.
  • Registers (PC, IR, MAR, MDR, ACC) act as temporary storage for instructions/data during execution, with their roles tied to the fetch-decode-execute cycle.
  • The control unit (hardwired or microprogrammed) generates timing and control signals using hardware (decoders, flip-flops) or microinstructions (stored in control memory).
  • Pipelining overlaps instruction stages (IF, ID, EX, MEM, WB) to boost throughput, but introduces hazards (data, control, structural) that require stalls or forwarding.
  • Real-world processors (e.g., Intel Core i7, ARM Cortex in smartphones) use superscalar pipelines (multiple execution units) and branch prediction to optimize performance.
  • Exam focus: Trace the fetch-decode-execute cycle, draw ALU circuits, explain pipeline hazards, and compare hardwired vs. microprogrammed control units.

1. The Arithmetic Logic Unit (ALU): The Brain of the CPU

The ALU is the heart of the processor, performing arithmetic (addition, subtraction, multiplication) and logical (AND, OR, NOT, XOR) operations. It is a combinational circuit built from smaller blocks like full adders, multiplexers (MUX), and decoders.

How the ALU Works

  1. Inputs:
    • Two operands (A and B) from registers or memory.
    • A control signal (e.g., ADD, SUB, AND) to select the operation.
  2. Operation Selection:
    • A decoder or priority encoder interprets the control signal.
    • A multiplexer routes the correct operation (e.g., A + B for ADD, A AND B for AND).
  3. Output:
    • Result (SUM or RESULT) and flags (e.g., Zero, Carry, Overflow).

ALU Circuit Example: 4-bit ALU

Worked Example: ALU Performing A AND B

Input A Input B Control Signal Operation Output (A AND B)
1010 1100 AND Bitwise AND 1000
0101 0011 ADD Addition 0110

Real-World Tie-In:

  • eSewa transactions: When you pay a bill, the ALU in the server’s CPU calculates the total amount (arithmetic) and checks if the transaction is valid (logical operations like IF (balance >= amount) THEN proceed).
  • Ncell billing: The ALU computes roaming charges by adding base rates and surcharges, then flags overflow if the total exceeds limits.

2. Registers: The CPU’s Temporary Memory

Registers are small, ultra-fast storage units inside the CPU (typically 32–256 bits) that hold:

  • Instructions (e.g., LOAD, STORE).
  • Data (operands, results).
  • Addresses (memory locations).

Key Registers in a Typical CPU

Register Abbreviation Role
Program Counter PC Holds the address of the next instruction to fetch.
Instruction Register IR Stores the current instruction being executed.
Memory Address Register MAR Holds the memory address for data/ instruction transfer.
Memory Data Register MDR Temporarily holds data read from/written to memory.
Accumulator ACC Stores intermediate results of ALU operations.
General Purpose R0-R15 Used for arithmetic/logic operations (e.g., R1 = R2 + R3).
General Purpose RegistersR0-R15Special Purpose RegistersPC, MAR, MDR, IR, SPFloating Point RegistersFP0-FP7
Hierarchy of CPU registers by function

Register Transfer Example: LOAD R1, [1000]

  1. MAR ← 1000 (address to load from).
  2. Memory → MDR (data at address 1000 is fetched).
  3. MDR → R1 (data stored in register R1).

Visual: Register Transfer Diagram

flowchart LR
    A["Memory[1000]"] -->|"Data"| B["MDR"]
    B -->|"Data"| C["R1"]
    D["PC"] -->|"Address"| E["MAR"]
    E -->|"Address"| A

Real-World Tie-In:

  • Pathao driver app: When you request a ride, the CPU’s registers store:
    • PC: Address of the next instruction (e.g., "update driver location").
    • MDR: Temporary data like your pickup location.
    • ACC: Intermediate calculations (e.g., fare = distance × rate).

3. The Control Unit: Orchestrating Execution

The control unit (CU) manages the fetch-decode-execute cycle by generating timing and control signals. It can be:

  1. Hardwired Control: Fixed logic circuits (faster, simpler).
  2. Microprogrammed Control: Uses a control memory to store microinstructions (more flexible).

Fetch-Decode-Execute Cycle

stateDiagram-v2
    [*] --> Fetch: PC → MAR → Memory → MDR → IR; PC++
    Fetch --> Decode: IR → Control Unit (decodes opcode)
    Decode --> Execute: ALU/Registers perform operation
    Execute --> [*]

Hardwired vs. Microprogrammed Control

Feature Hardwired Control Microprogrammed Control
Speed Faster (direct logic) Slower (fetches microinstructions)
Flexibility Inflexible (fixed logic) Flexible (can be reprogrammed)
Complexity Simpler design More complex (needs control memory)
Example Early CPUs (e.g., Intel 8086) Modern RISC processors (ARM)

Worked Example: Executing ADD R1, R2

  1. Fetch:
    • PC → MAR → Memory → MDR (ADD R1, R2).
    • MDR → IR; PC++.
  2. Decode:
    • Control unit decodes ADD → generates signals for ALU.
  3. Execute:
    • R1 → ALU A, R2 → ALU B.
    • ALU performs A + B → result → R1.

Real-World Tie-In:

  • NEPSE stock trading: When you buy shares, the CPU’s control unit:
    • Fetches your order (BUY NEPSE:100).
    • Decodes it to trigger ALU operations (e.g., deduct 100 × price from your balance).
    • Executes the transaction and updates registers with new stock counts.

4. Pipelining: Overlapping Execution for Speed

Pipelining divides instruction execution into 5 stages (IF, ID, EX, MEM, WB) and overlaps them to increase throughput.

Pipeline Stages

Stage Abbreviation Task
Fetch IF PC → MAR → Memory → IR; PC++
Decode ID Decode IR; read registers
Execute EX ALU operation (e.g., ADD, AND)
Memory MEM Access memory (load/store)
Writeback WB Write result to register file

Pipeline Hazards

Hazard Type Cause Solution
Data Hazard Instruction depends on previous result (e.g., ADD R1, R2; SUB R3, R1). Stall pipeline or use forwarding.
Control Hazard Branch/jump changes PC (e.g., BEQ). Branch prediction or delay slots.
Structural Hazard Two instructions need the same resource (e.g., ALU busy). Add more execution units (superscalar).

Visual: 5-Stage Pipeline Timing Diagram

Worked Example: Data Hazard in ADD R1, R2; SUB R3, R1

  1. Cycle 1: ADD in EX stage (result not yet in R1).
  2. Cycle 2: SUB tries to read R1 (data hazard).
    • Solution: Stall SUB until ADD completes in WB.

Real-World Tie-In:

  • Daraz order processing: When you place multiple orders in quick succession, the CPU’s pipeline:
    • Fetches order 1 → decodes → executes (deducts money).
    • Fetches order 2 → stalls if order 1’s payment isn’t confirmed (data hazard).
    • Uses forwarding to pass order 1’s result directly to order 2’s execution.

5. Advanced Topics: Superscalar and Out-of-Order Execution

Modern CPUs (e.g., Intel Core i7, ARM Cortex) use:

  • Superscalar: Multiple execution units (e.g., 2 ALUs, 1 FPU) to execute multiple instructions per cycle (IPC).
  • Out-of-Order Execution: Reorders instructions to avoid stalls (e.g., executes LOAD before ADD if data is ready).

In the Real World

  1. eSewa Transactions:

    • ALU: Calculates total bill amount (arithmetic) and checks validity (logical IF).
    • Registers: Store user ID (R1), transaction amount (R2), and response code (R3).
    • Pipeline: Processes multiple payments in parallel (superscalar).
  2. Ncell Billing System:

    • Control Unit: Decodes CHARGE:100 to trigger ALU for deduction.
    • Hazards: If two calls overlap, the pipeline stalls until the first call’s data is written back.
  3. Pathao Driver App:

    • ALU: Computes fare = distance × rate.
    • Pipelining: Handles multiple ride requests by overlapping fetch/decode stages.

Exam Tip

  1. Fetch-Decode-Execute Cycle:

    • Draw the state diagram and label each step (PC, MAR, MDR, IR).
    • Common mistake: Forgetting PC++ after fetching.
  2. ALU Circuits:

    • Must-know: Draw a 4-bit ALU with MUX and full adder blocks.
    • Exam question: "Design an ALU to support ADD, SUB, AND." → Use a 3:8 decoder for control signals.
  3. Pipeline Hazards:

    • Spot the hazard: Given LOAD R1, [100]; ADD R2, R1, R3, identify data hazard and propose stall/forwarding.
    • Superscalar bonus: Mention "modern CPUs use dynamic scheduling to hide latency."
  4. Control Unit Comparison:

    • Table question: Compare hardwired vs. microprogrammed (speed, flexibility, example CPUs).
  5. Real-World Application:

    • Link to Nepali tech: "How does the ALU in an eSewa server compute transaction fees?" → Arithmetic operations + flag checks.

Final Note: Focus on visuals (circuits, timing diagrams, state machines) and step-by-step traces (fetch-decode-execute). The exam tests both theory (e.g., pipeline stages) and application (e.g., "Explain how a data hazard occurs in a bank’s loan interest calculation").

Based on the TU BSc CSIT syllabus for Computer Hardware Design, unit 4.

Discussion

Loading…