Elective Embedded Systems Programming

Embedded Systems ProgrammingUnit 68 min read

ARM Assembly: Code Optimization & Efficiency

Unit 6 of Embedded Systems Programming: This unit teaches ARM assembly language syntax, optimization techniques (register usage, pipelining, branch prediction), and how to write efficient low-level code for resource-constrained embedded systems.

TAKEAWAYS:

  • ARM assembly uses registers (R0–R15) and thumb mode for compact code, critical for microcontrollers like STM32 or Raspberry Pi Pico.
  • Optimization reduces power consumption and execution time by leveraging ARM-specific instructions (e.g., ADD, LDR, B).
  • Pipelining and branch prediction improve throughput in embedded firmware.
  • Memory-mapped I/O and bit manipulation are essential for interfacing with hardware (e.g., GPIO, timers).
  • Debugging in ARM assembly requires tools like Keil MDK or GNU Arm Embedded Toolchain.
  • Real-world impact: Pathao’s route optimization (Dijkstra’s algorithm in assembly) or NTC’s traffic light control (state machines) rely on low-latency assembly code.

1. ARM Assembly Basics: Registers and Addressing Modes

ARM processors use 31 general-purpose registers (R0–R14) and 5 special-purpose registers (R15–PC, CPSR, SPSR). Unlike x86, ARM has no separate stack pointer and frame pointer in Thumb mode.

Key Registers

R0–R12: General-purpose (e.g., R0–R3 for function arguments) R13 (SP): Stack pointer R14 (LR): Link register (returns to caller) R15 (PC): Program counter CPSR: Condition flags (Z, N, C, V)


Addressing Modes

ARM supports 10 addressing modes, but embedded code often uses:

  • Register addressing: LDR R1, [R0] (load from address in R0)
  • Immediate addressing: MOV R1, #5 (loads constant 5)
  • Offset addressing: LDR R1, [R0, #4] (loads from R0 + 4)

Example: Loading a value

sequenceDiagram
    participant R0
    participant Memory
    R0->>Memory: [R0] = 0x2000
    R0->>Memory: LDR R1, [R0]  // Loads value at address 0x2000 into R1

Trace Table:

Step R0 R1 Memory[0x2000]
1 0x2000 - 0x42
2 0x2000 0x42 0x42

2. ARM Instruction Set: Data Transfer and Arithmetic

Data Transfer Instructions

  • LDR (Load Register): LDR R1, =0x1234 (loads literal)
  • STR (Store Register): STR R1, [R0] (stores R1 to address in R0)
  • PUSH/POP: Manages stack (e.g., PUSH {R4, LR} saves registers)

Arithmetic Instructions

  • ADD: ADD R1, R2, R3 (R1 = R2 + R3)
  • SUB: SUB R1, R2, R3 (R1 = R2 - R3)
  • MUL: MUL R1, R2, R3 (R1 = R2 × R3)

Example: Summing two numbers

MOV R1, #5
MOV R2, #3
ADD R3, R1, R2  ; R3 = 8

Trace:

R1 R2 R3
5 3 8

3. Branch Instructions and Control Flow

Branches (B, BL, BX) change program flow. Thumb mode uses 16-bit branches (B is 24-bit in ARM mode).

Branch Types

Instruction Description Example
B Unconditional branch B label
BL Branch with link (saves LR) BL func
BX Branch with exchange (jump to) BX R0 (jump to address in R0)

Example: Conditional branch

sequenceDiagram
    participant CPSR
    participant PC
    CPSR->>PC: Z flag set (R1 == 0)
    PC->>PC: BEQ label  // Branch if equal (Z=1)

Code:

MOV R1, #0
CMP R1, #0
BEQ zero_flag  ; Branch if R1 == 0

Trace:

R1 Z flag Action
0 1 Branch to zero_flag

4. Optimizing ARM Assembly Code

Optimization reduces code size and execution time, critical for embedded systems.

Optimization Techniques

Technique ARM Example Benefit
Register reuse ADD R1, R2, R3 (avoid MOV) Fewer instructions
Pipelining Overlap instruction execution Higher throughput
Branch prediction B instructions optimized Reduces pipeline stalls
Bit manipulation TST R1, #1 (check bit 0) Faster than division

Example: Optimized Loop

LOOP:
    LDR R1, [R0], #4  ; Load and increment R0 by 4
    ADD R1, R1, R2    ; Add R2 to R1
    CMP R0, #end_addr
    BNE LOOP         ; Branch if not equal

Trace Table:

R0 R1 Action
0x2000 0x42 Load 0x42, R0 → 0x2004
0x2004 0x46 Add R2, R0 → 0x2008
... ... ...

5. Memory-Mapped I/O and Hardware Interaction

Embedded systems (e.g., STM32 microcontrollers) use memory-mapped I/O, where peripherals (GPIO, UART) are accessed like memory.

Example: Writing to GPIO (STM32)

MOV R0, #0x40021000  ; GPIOA base address
MOV R1, #0x01        ; Set bit 0 (LED)
STR R1, [R0]         ; Write to GPIO output

Trace:

R0 (GPIOA) R1 (Value) Action
0x40021000 0x01 LED turns ON

6. Debugging ARM Assembly

Debugging tools:

  • Keil MDK: GUI-based debugger for ARM Cortex-M.
  • GNU Arm Embedded Toolchain: Command-line debugger (arm-none-eabi-gdb).
  • Logic Analyzers: Hardware tools to inspect signals.

Example: Debugging a Loop

startLDR R1, [R0]BEQ exit_loop (R1 == 0)BNE continue (R1 != 0)SUB R1, #1; B loop_startloop_startcheck_zerodecrementexit_loop
Debugging loop: Flow of control with conditional branch

In the Real World

  1. Pathao’s Route Optimization

    • Uses ARM assembly to implement Dijkstra’s algorithm for real-time route calculation in microcontrollers.
    • Why? Assembly reduces latency in GPS data processing.
  2. NTC’s Traffic Light Control

    • Embedded systems (e.g., STM32) use state machines in assembly to manage traffic signals.
    • Why? Predictable timing and low power consumption.
  3. Ncell’s Signal Processing

    • ARM Cortex-M processors in baseband chips use optimized assembly for modulation/demodulation (e.g., LTE signals).
    • Why? Faster than C for critical signal processing.

Worked Example: NEPSE’s Stock Price Update Suppose NEPSE’s server updates stock prices every 10ms. An optimized assembly loop (instead of C) reduces execution time from 5ms → 2ms, improving responsiveness.

LOOP:
    LDR R1, [R0]      ; Load stock price
    ADD R1, R1, #1    ; Increment price
    STR R1, [R0]      ; Store back
    SUB R2, R2, #1    ; Decrement counter
    CMP R2, #0
    BNE LOOP

Trace:

R0 (Price) R1 (New Price) R2 (Counter)
100 101 99
101 102 98
... ... ...

Comparison: ARM vs. Thumb Mode

Feature ARM Mode (32-bit) Thumb Mode (16-bit)
Instruction Size 32-bit 16-bit
Code Density Lower Higher
Execution Speed Faster Slower
Use Case General-purpose Microcontrollers
01234ARM (32-bit)4Thumb (16-bit)2Instruction size (bytes)
Instruction encoding size comparison (ARM vs. Thumb)

Exam Tip

  • Focus on optimization: Examiners test knowledge of register reuse, pipelining, and branch prediction.
  • Hardware interaction: Always include memory-mapped I/O examples (e.g., GPIO, timers).
  • Debugging: Know how to use Keil MDK or GNU tools for assembly debugging.
  • Real-world mapping: Relate ARM assembly to Pathao’s routing or NTC’s traffic control in answers.
  • Code tracing: Always show step-by-step register/memory changes in your answers.

Key Formula to Remember: For branch prediction, the misprediction penalty is often 2–4 cycles in ARM. Optimize branches to minimize this cost.

Based on the TU BSc CSIT syllabus for Embedded Systems Programming, unit 6.

Discussion

Loading…