Embedded Systems ProgrammingUnit 68 min read
ARM Assembly: Code Optimization & Efficiency
Unit 6 of Embedded Systems Programming: This unit teaches ARM assembly language syntax, optimization techniques (register usage, pipelining, branch prediction), and how to write efficient low-level code for resource-constrained embedded systems.
TAKEAWAYS:
- ARM assembly uses registers (R0–R15) and thumb mode for compact code, critical for microcontrollers like STM32 or Raspberry Pi Pico.
- Optimization reduces power consumption and execution time by leveraging ARM-specific instructions (e.g.,
ADD,LDR,B). - Pipelining and branch prediction improve throughput in embedded firmware.
- Memory-mapped I/O and bit manipulation are essential for interfacing with hardware (e.g., GPIO, timers).
- Debugging in ARM assembly requires tools like Keil MDK or GNU Arm Embedded Toolchain.
- Real-world impact: Pathao’s route optimization (Dijkstra’s algorithm in assembly) or NTC’s traffic light control (state machines) rely on low-latency assembly code.
1. ARM Assembly Basics: Registers and Addressing Modes
ARM processors use 31 general-purpose registers (R0–R14) and 5 special-purpose registers (R15–PC, CPSR, SPSR). Unlike x86, ARM has no separate stack pointer and frame pointer in Thumb mode.
Key Registers
R0–R12: General-purpose (e.g., R0–R3 for function arguments) R13 (SP): Stack pointer R14 (LR): Link register (returns to caller) R15 (PC): Program counter CPSR: Condition flags (Z, N, C, V)
Addressing Modes
ARM supports 10 addressing modes, but embedded code often uses:
- Register addressing:
LDR R1, [R0](load from address in R0) - Immediate addressing:
MOV R1, #5(loads constant 5) - Offset addressing:
LDR R1, [R0, #4](loads from R0 + 4)
Example: Loading a value
sequenceDiagram
participant R0
participant Memory
R0->>Memory: [R0] = 0x2000
R0->>Memory: LDR R1, [R0] // Loads value at address 0x2000 into R1Trace Table:
| Step | R0 | R1 | Memory[0x2000] |
|---|---|---|---|
| 1 | 0x2000 | - | 0x42 |
| 2 | 0x2000 | 0x42 | 0x42 |
2. ARM Instruction Set: Data Transfer and Arithmetic
Data Transfer Instructions
LDR(Load Register):LDR R1, =0x1234(loads literal)STR(Store Register):STR R1, [R0](stores R1 to address in R0)PUSH/POP: Manages stack (e.g.,PUSH {R4, LR}saves registers)
Arithmetic Instructions
ADD:ADD R1, R2, R3(R1 = R2 + R3)SUB:SUB R1, R2, R3(R1 = R2 - R3)MUL:MUL R1, R2, R3(R1 = R2 × R3)
Example: Summing two numbers
MOV R1, #5
MOV R2, #3
ADD R3, R1, R2 ; R3 = 8
Trace:
| R1 | R2 | R3 |
|---|---|---|
| 5 | 3 | 8 |
3. Branch Instructions and Control Flow
Branches (B, BL, BX) change program flow. Thumb mode uses 16-bit branches (B is 24-bit in ARM mode).
Branch Types
| Instruction | Description | Example |
|---|---|---|
B |
Unconditional branch | B label |
BL |
Branch with link (saves LR) | BL func |
BX |
Branch with exchange (jump to) | BX R0 (jump to address in R0) |
Example: Conditional branch
sequenceDiagram
participant CPSR
participant PC
CPSR->>PC: Z flag set (R1 == 0)
PC->>PC: BEQ label // Branch if equal (Z=1)Code:
MOV R1, #0
CMP R1, #0
BEQ zero_flag ; Branch if R1 == 0
Trace:
| R1 | Z flag | Action |
|---|---|---|
| 0 | 1 | Branch to zero_flag |
4. Optimizing ARM Assembly Code
Optimization reduces code size and execution time, critical for embedded systems.
Optimization Techniques
| Technique | ARM Example | Benefit |
|---|---|---|
| Register reuse | ADD R1, R2, R3 (avoid MOV) |
Fewer instructions |
| Pipelining | Overlap instruction execution | Higher throughput |
| Branch prediction | B instructions optimized |
Reduces pipeline stalls |
| Bit manipulation | TST R1, #1 (check bit 0) |
Faster than division |
Example: Optimized Loop
LOOP:
LDR R1, [R0], #4 ; Load and increment R0 by 4
ADD R1, R1, R2 ; Add R2 to R1
CMP R0, #end_addr
BNE LOOP ; Branch if not equal
Trace Table:
| R0 | R1 | Action |
|---|---|---|
| 0x2000 | 0x42 | Load 0x42, R0 → 0x2004 |
| 0x2004 | 0x46 | Add R2, R0 → 0x2008 |
| ... | ... | ... |
5. Memory-Mapped I/O and Hardware Interaction
Embedded systems (e.g., STM32 microcontrollers) use memory-mapped I/O, where peripherals (GPIO, UART) are accessed like memory.
Example: Writing to GPIO (STM32)
MOV R0, #0x40021000 ; GPIOA base address
MOV R1, #0x01 ; Set bit 0 (LED)
STR R1, [R0] ; Write to GPIO output
Trace:
| R0 (GPIOA) | R1 (Value) | Action |
|---|---|---|
| 0x40021000 | 0x01 | LED turns ON |
6. Debugging ARM Assembly
Debugging tools:
- Keil MDK: GUI-based debugger for ARM Cortex-M.
- GNU Arm Embedded Toolchain: Command-line debugger (
arm-none-eabi-gdb). - Logic Analyzers: Hardware tools to inspect signals.
Example: Debugging a Loop
In the Real World
Pathao’s Route Optimization
- Uses ARM assembly to implement Dijkstra’s algorithm for real-time route calculation in microcontrollers.
- Why? Assembly reduces latency in GPS data processing.
NTC’s Traffic Light Control
- Embedded systems (e.g., STM32) use state machines in assembly to manage traffic signals.
- Why? Predictable timing and low power consumption.
Ncell’s Signal Processing
- ARM Cortex-M processors in baseband chips use optimized assembly for modulation/demodulation (e.g., LTE signals).
- Why? Faster than C for critical signal processing.
Worked Example: NEPSE’s Stock Price Update Suppose NEPSE’s server updates stock prices every 10ms. An optimized assembly loop (instead of C) reduces execution time from 5ms → 2ms, improving responsiveness.
LOOP:
LDR R1, [R0] ; Load stock price
ADD R1, R1, #1 ; Increment price
STR R1, [R0] ; Store back
SUB R2, R2, #1 ; Decrement counter
CMP R2, #0
BNE LOOP
Trace:
| R0 (Price) | R1 (New Price) | R2 (Counter) |
|---|---|---|
| 100 | 101 | 99 |
| 101 | 102 | 98 |
| ... | ... | ... |
Comparison: ARM vs. Thumb Mode
| Feature | ARM Mode (32-bit) | Thumb Mode (16-bit) |
|---|---|---|
| Instruction Size | 32-bit | 16-bit |
| Code Density | Lower | Higher |
| Execution Speed | Faster | Slower |
| Use Case | General-purpose | Microcontrollers |
Exam Tip
- Focus on optimization: Examiners test knowledge of register reuse, pipelining, and branch prediction.
- Hardware interaction: Always include memory-mapped I/O examples (e.g., GPIO, timers).
- Debugging: Know how to use Keil MDK or GNU tools for assembly debugging.
- Real-world mapping: Relate ARM assembly to Pathao’s routing or NTC’s traffic control in answers.
- Code tracing: Always show step-by-step register/memory changes in your answers.
Key Formula to Remember: For branch prediction, the misprediction penalty is often 2–4 cycles in ARM. Optimize branches to minimize this cost.
Based on the TU BSc CSIT syllabus for Embedded Systems Programming, unit 6.
Discussion
Loading…