Embedded Systems ProgrammingUnit 56 min read
Efficient C for ARM: Optimizations & Constraints
Unit 5 of Embedded Systems Programming covers memory-efficient C coding for ARM, register usage, bit manipulation, inline assembly, and compiler optimizations, with real-world constraints like limited RAM/ROM and power budgets.
TAKEAWAYS:
- ARM’s Harvard architecture (separate code/data buses) demands careful pointer usage and memory alignment for speed.
- Bitwise operations (
&,|,<<) replace loops to save cycles in resource-constrained systems like Ncell’s IoT sensors. - Inline assembly (
__asm) bridges C and ARM assembly for critical sections (e.g., Khalti’s payment gateways use it for cryptographic hashing). - Compiler flags (
-O2,-mcpu=cortex-m4) trade code size for speed—critical for eSewa’s embedded payment terminals. - Stack vs. heap: ARM’s small stack (often <1KB) forces global/static variables for persistent data (e.g., NTC’s smart meters).
- Endianness (ARM’s little-endian default) affects multi-byte data handling in NEPSE’s stock-trading servers.
1. ARM-Specific C Constraints
ARM microcontrollers (e.g., Cortex-M, Cortex-A) differ from x86 in:
- Memory hierarchy: Flash (code), SRAM (data), and peripherals (GPIO, UART) are separate.
- Registers: 13 general-purpose registers (R0–R12, LR, PC, SP, xPSR) vs. x86’s 8.
- Endianness: Little-endian by default (LSB first). Big-endian mode is rare but used in network protocols.
classDiagram
class ARM_Memory {
+Flash: Code storage (ROM)
+SRAM: Data/Stack (RAM)
+Peripherals: GPIO, UART, etc.
}
class ARM_Registers {
+R0-R12: General-purpose
+LR: Link Register
+PC: Program Counter
+SP: Stack Pointer
+xPSR: Status Register
}
ARM_Memory --> ARM_Registers : "Accessed via registers"Why it matters:
- Misaligned accesses (e.g.,
uint32_t* ptr = (uint32_t*)0x20000001;) cause hard faults on ARM. - Stack overflow crashes the system if SP grows beyond SRAM limits (common in Pathao’s bike-tracking firmware).
2. Memory Optimization Techniques
A. Data Types and Alignment
ARM prefers 4-byte alignment for performance. Use:
// Correct (4-byte aligned)
uint32_t aligned_var __attribute__((aligned(4)));
// Wrong (misaligned)
uint32_t* misaligned_ptr = (uint32_t*)0x20000001; // Hard fault!
Visual: Alignment Impact
B. Stack vs. Heap
- Stack: Fast but limited (e.g., 1KB in Cortex-M0). Use for short-lived data.
- Heap: Slower (malloc/free overhead) but dynamic. Avoid in ISRs (Interrupt Service Routines).
Example: NTC Smart Meter Firmware
// Bad: Stack overflow risk
void read_meter() {
uint16_t buffer[1024]; // 2KB on stack → CRASH!
}
// Good: Static array (heap-free)
static uint16_t buffer[1024]; // Allocated at startup
3. Bitwise Operations for Speed
Replace loops with bitwise ops. Example: Khalti’s payment validation checks flags in a single instruction.
// Slow: Loop
for (int i = 0; i < 8; i++) {
if (flags & (1 << i)) { /* ... */ }
}
// Fast: Bitwise
if (flags & 0xAA) { /* Check bits 1,3,5,7 */ }
Visual: Bitmasking in Action
4. Inline Assembly for Critical Code
Use __asm for performance-critical sections (e.g., Daraz’s order-fulfillment sensors).
// C: Slow division
uint32_t div_by_10(uint32_t x) { return x / 10; }
// ARM Assembly: Fast (10 cycles vs. 30+ in C)
uint32_t div_by_10_fast(uint32_t x) {
__asm__("UDIV %0, %1, #10" : "=r"(x) : "r"(x));
return x;
}
Trace: Register Usage
| Step | C Code | ARM Assembly | Registers |
|---|---|---|---|
| 1 | x = 1234 |
MOV R0, #1234 |
R0 = 1234 |
| 2 | x / 10 |
UDIV R0, R0, #10 |
R0 = 123 (result) |
5. Compiler Optimizations
Enable with -O2 or -Os (size vs. speed tradeoff). Key flags:
| Flag | Effect |
|---|---|
-mcpu=cortex-m4 |
Targets Cortex-M4 instructions. |
-mthumb |
Uses Thumb-2 (16/32-bit mixed) mode. |
-ffunction-sections |
Removes unused functions (saves flash). |
Example: eSewa’s Payment Terminal
arm-none-eabi-gcc -O2 -mcpu=cortex-m4 -mthumb main.c -o firmware.elf
6. Power-Efficient Coding
- Sleep modes: Use
__WFI()(Wait For Interrupt) to save power in NEPSE’s stock-exchange servers. - Clock gating: Disable unused peripherals (e.g., UART when idle).
// Enter low-power mode
__asm__("WFI"); // Waits for interrupt (saves ~90% power)
In the Real World
Khalti’s Payment Gateways
- Idea: Bitwise operations validate transaction flags in <100ns (vs. 1µs in C loops).
- How:
if (tx_flags & 0x03) { /* Approve */ }checks 2 bits for fraud detection.
Ncell’s IoT Sensors
- Idea: Inline assembly handles ADC (Analog-to-Digital Conversion) in fixed 50µs (vs. 200µs in C).
- How:
__asm__("ADC R0, R1");reads temperature sensors directly.
eSewa’s Embedded Terminals
- Idea: Stack-allocated buffers (max 512B) prevent crashes in high-transaction queues.
- How:
static uint8_t buffer[512];ensures no heap fragmentation.
Exam Tip
- Memory alignment: Always align data to 4/8 bytes. Misalignment = hard fault.
- Bitwise ops: Prefer
<<,>>,&over loops. Example:x |= (1 << 3)sets bit 3. - Compiler flags:
-O2optimizes speed;-Ossaves flash. Know the tradeoffs. - Stack vs. heap: ARM’s tiny stack (often <1KB) means global/static variables are safer.
- Endianness: ARM is little-endian by default. Network protocols (e.g., TCP) may require byte swaps.
Worked Example: Daraz Order Queue
// Simulates Daraz’s order fulfillment (FIFO queue)
typedef struct {
uint32_t order_id;
uint8_t status; // 0=pending, 1=shipped
} Order;
#define MAX_ORDERS 100
Order queue[MAX_ORDERS];
uint8_t head = 0, tail = 0;
// Add order (bitwise status check)
void add_order(uint32_t id) {
queue[tail].order_id = id;
queue[tail].status = 0; // Pending
tail = (tail + 1) % MAX_ORDERS;
if (tail == head) { /* Queue full */ }
}
// Check status (bitwise)
uint8_t is_shipped(uint32_t id) {
for (int i = 0; i < MAX_ORDERS; i++) {
if (queue[i].order_id == id) {
return queue[i].status & 0x01; // Bit 0 = shipped?
}
}
return 0;
}
Trace: Queue Operations
Based on the TU BSc CSIT syllabus for Embedded Systems Programming, unit 5.
Discussion
Loading…