Computer Hardware DesignUnit 410 min read
The Processor: ALU, Control Unit, Registers & Pipeline
Unit 4 of Computer Hardware Design explores the core of the CPU—the Arithmetic Logic Unit (ALU), control unit, registers, instruction execution, and pipelining—with real-world examples from Nepali tech (e.g., Ncell billing systems, eSewa transactions) and visual breakdowns of how processors fetch, decode, and execute i
TAKEAWAYS:
- The ALU performs arithmetic/logic operations (e.g.,
ADD,AND,NOT) using combinational circuits (full adders, multiplexers) and is controlled by the control unit. - Registers (PC, IR, MAR, MDR, ACC) act as temporary storage for instructions/data during execution, with their roles tied to the fetch-decode-execute cycle.
- The control unit (hardwired or microprogrammed) generates timing and control signals using hardware (decoders, flip-flops) or microinstructions (stored in control memory).
- Pipelining overlaps instruction stages (IF, ID, EX, MEM, WB) to boost throughput, but introduces hazards (data, control, structural) that require stalls or forwarding.
- Real-world processors (e.g., Intel Core i7, ARM Cortex in smartphones) use superscalar pipelines (multiple execution units) and branch prediction to optimize performance.
- Exam focus: Trace the fetch-decode-execute cycle, draw ALU circuits, explain pipeline hazards, and compare hardwired vs. microprogrammed control units.
1. The Arithmetic Logic Unit (ALU): The Brain of the CPU
The ALU is the heart of the processor, performing arithmetic (addition, subtraction, multiplication) and logical (AND, OR, NOT, XOR) operations. It is a combinational circuit built from smaller blocks like full adders, multiplexers (MUX), and decoders.
How the ALU Works
- Inputs:
- Two operands (
AandB) from registers or memory. - A control signal (e.g.,
ADD,SUB,AND) to select the operation.
- Two operands (
- Operation Selection:
- A decoder or priority encoder interprets the control signal.
- A multiplexer routes the correct operation (e.g.,
A + BforADD,A AND BforAND).
- Output:
- Result (
SUMorRESULT) and flags (e.g.,Zero,Carry,Overflow).
- Result (
ALU Circuit Example: 4-bit ALU
Worked Example: ALU Performing A AND B
| Input A | Input B | Control Signal | Operation | Output (A AND B) |
|---|---|---|---|---|
| 1010 | 1100 | AND |
Bitwise AND | 1000 |
| 0101 | 0011 | ADD |
Addition | 0110 |
Real-World Tie-In:
- eSewa transactions: When you pay a bill, the ALU in the server’s CPU calculates the total amount (arithmetic) and checks if the transaction is valid (logical operations like
IF (balance >= amount) THEN proceed). - Ncell billing: The ALU computes roaming charges by adding base rates and surcharges, then flags overflow if the total exceeds limits.
2. Registers: The CPU’s Temporary Memory
Registers are small, ultra-fast storage units inside the CPU (typically 32–256 bits) that hold:
- Instructions (e.g.,
LOAD,STORE). - Data (operands, results).
- Addresses (memory locations).
Key Registers in a Typical CPU
| Register | Abbreviation | Role |
|---|---|---|
| Program Counter | PC | Holds the address of the next instruction to fetch. |
| Instruction Register | IR | Stores the current instruction being executed. |
| Memory Address Register | MAR | Holds the memory address for data/ instruction transfer. |
| Memory Data Register | MDR | Temporarily holds data read from/written to memory. |
| Accumulator | ACC | Stores intermediate results of ALU operations. |
| General Purpose | R0-R15 | Used for arithmetic/logic operations (e.g., R1 = R2 + R3). |
Register Transfer Example: LOAD R1, [1000]
- MAR ←
1000(address to load from). - Memory → MDR (data at address
1000is fetched). - MDR → R1 (data stored in register
R1).
Visual: Register Transfer Diagram
flowchart LR
A["Memory[1000]"] -->|"Data"| B["MDR"]
B -->|"Data"| C["R1"]
D["PC"] -->|"Address"| E["MAR"]
E -->|"Address"| AReal-World Tie-In:
- Pathao driver app: When you request a ride, the CPU’s registers store:
- PC: Address of the next instruction (e.g., "update driver location").
- MDR: Temporary data like your pickup location.
- ACC: Intermediate calculations (e.g., fare = distance × rate).
3. The Control Unit: Orchestrating Execution
The control unit (CU) manages the fetch-decode-execute cycle by generating timing and control signals. It can be:
- Hardwired Control: Fixed logic circuits (faster, simpler).
- Microprogrammed Control: Uses a control memory to store microinstructions (more flexible).
Fetch-Decode-Execute Cycle
stateDiagram-v2
[*] --> Fetch: PC → MAR → Memory → MDR → IR; PC++
Fetch --> Decode: IR → Control Unit (decodes opcode)
Decode --> Execute: ALU/Registers perform operation
Execute --> [*]Hardwired vs. Microprogrammed Control
| Feature | Hardwired Control | Microprogrammed Control |
|---|---|---|
| Speed | Faster (direct logic) | Slower (fetches microinstructions) |
| Flexibility | Inflexible (fixed logic) | Flexible (can be reprogrammed) |
| Complexity | Simpler design | More complex (needs control memory) |
| Example | Early CPUs (e.g., Intel 8086) | Modern RISC processors (ARM) |
Worked Example: Executing ADD R1, R2
- Fetch:
- PC → MAR → Memory → MDR (
ADD R1, R2). - MDR → IR; PC++.
- PC → MAR → Memory → MDR (
- Decode:
- Control unit decodes
ADD→ generates signals for ALU.
- Control unit decodes
- Execute:
- R1 → ALU A, R2 → ALU B.
- ALU performs
A + B→ result → R1.
Real-World Tie-In:
- NEPSE stock trading: When you buy shares, the CPU’s control unit:
- Fetches your order (
BUY NEPSE:100). - Decodes it to trigger ALU operations (e.g., deduct
100 × pricefrom your balance). - Executes the transaction and updates registers with new stock counts.
- Fetches your order (
4. Pipelining: Overlapping Execution for Speed
Pipelining divides instruction execution into 5 stages (IF, ID, EX, MEM, WB) and overlaps them to increase throughput.
Pipeline Stages
| Stage | Abbreviation | Task |
|---|---|---|
| Fetch | IF | PC → MAR → Memory → IR; PC++ |
| Decode | ID | Decode IR; read registers |
| Execute | EX | ALU operation (e.g., ADD, AND) |
| Memory | MEM | Access memory (load/store) |
| Writeback | WB | Write result to register file |
Pipeline Hazards
| Hazard Type | Cause | Solution |
|---|---|---|
| Data Hazard | Instruction depends on previous result (e.g., ADD R1, R2; SUB R3, R1). |
Stall pipeline or use forwarding. |
| Control Hazard | Branch/jump changes PC (e.g., BEQ). |
Branch prediction or delay slots. |
| Structural Hazard | Two instructions need the same resource (e.g., ALU busy). | Add more execution units (superscalar). |
Visual: 5-Stage Pipeline Timing Diagram
Worked Example: Data Hazard in ADD R1, R2; SUB R3, R1
- Cycle 1:
ADDin EX stage (result not yet in R1). - Cycle 2:
SUBtries to read R1 (data hazard).- Solution: Stall
SUBuntilADDcompletes in WB.
- Solution: Stall
Real-World Tie-In:
- Daraz order processing: When you place multiple orders in quick succession, the CPU’s pipeline:
- Fetches order 1 → decodes → executes (deducts money).
- Fetches order 2 → stalls if order 1’s payment isn’t confirmed (data hazard).
- Uses forwarding to pass order 1’s result directly to order 2’s execution.
5. Advanced Topics: Superscalar and Out-of-Order Execution
Modern CPUs (e.g., Intel Core i7, ARM Cortex) use:
- Superscalar: Multiple execution units (e.g., 2 ALUs, 1 FPU) to execute multiple instructions per cycle (IPC).
- Out-of-Order Execution: Reorders instructions to avoid stalls (e.g., executes
LOADbeforeADDif data is ready).
In the Real World
eSewa Transactions:
- ALU: Calculates total bill amount (arithmetic) and checks validity (logical
IF). - Registers: Store user ID (R1), transaction amount (R2), and response code (R3).
- Pipeline: Processes multiple payments in parallel (superscalar).
- ALU: Calculates total bill amount (arithmetic) and checks validity (logical
Ncell Billing System:
- Control Unit: Decodes
CHARGE:100to trigger ALU for deduction. - Hazards: If two calls overlap, the pipeline stalls until the first call’s data is written back.
- Control Unit: Decodes
Pathao Driver App:
- ALU: Computes fare =
distance × rate. - Pipelining: Handles multiple ride requests by overlapping fetch/decode stages.
- ALU: Computes fare =
Exam Tip
Fetch-Decode-Execute Cycle:
- Draw the state diagram and label each step (PC, MAR, MDR, IR).
- Common mistake: Forgetting
PC++after fetching.
ALU Circuits:
- Must-know: Draw a 4-bit ALU with MUX and full adder blocks.
- Exam question: "Design an ALU to support
ADD,SUB,AND." → Use a 3:8 decoder for control signals.
Pipeline Hazards:
- Spot the hazard: Given
LOAD R1, [100]; ADD R2, R1, R3, identify data hazard and propose stall/forwarding. - Superscalar bonus: Mention "modern CPUs use dynamic scheduling to hide latency."
- Spot the hazard: Given
Control Unit Comparison:
- Table question: Compare hardwired vs. microprogrammed (speed, flexibility, example CPUs).
Real-World Application:
- Link to Nepali tech: "How does the ALU in an eSewa server compute transaction fees?" → Arithmetic operations + flag checks.
Final Note: Focus on visuals (circuits, timing diagrams, state machines) and step-by-step traces (fetch-decode-execute). The exam tests both theory (e.g., pipeline stages) and application (e.g., "Explain how a data hazard occurs in a bank’s loan interest calculation").
Based on the TU BSc CSIT syllabus for Computer Hardware Design, unit 4.
Discussion
Loading…