Operating SystemUnit 1011 min read
Distributed Systems: Architecture, RPC, Security & Synchronization
Unit 10 of Operating System explores distributed systems—how multiple independent computers collaborate as a single system, covering transparency principles, communication methods (RPC, message passing), clock synchronization, security challenges, and real-world applications like cloud services and e-payment systems.
TAKEAWAYS:
- Distributed systems hide complexity (e.g., location, migration, replication) to appear as a single OS, using transparency layers like RPC and message passing.
- Remote Procedure Call (RPC) lets programs call functions on remote machines as if they were local, but requires marshalling (data serialization) and stub layers.
- Clock synchronization (e.g., Berkeley algorithm) is critical for distributed transactions (e.g., bank transfers) but faces drift and causality issues.
- Security threats (e.g., replay attacks, denial-of-service) demand authentication (passwords, biometrics) and authorization (access control lists).
- Consistency models (e.g., strong vs. eventual) trade off between correctness and performance—CAP theorem proves you can’t have all three (Consistency, Availability, Partition tolerance).
- Real-world use: eSewa uses distributed transactions for secure payments, while YouTube’s CDN relies on geographically replicated servers for low latency.
1. What is a Distributed System?
A distributed system is a collection of independent computers that appear to users as a single coherent system. Unlike centralized systems (e.g., a single mainframe), distributed systems:
- Share resources (e.g., files, printers, databases) across networks.
- Hide complexity (e.g., location of data, failures) via transparency.
- Enable scalability (e.g., Google’s search index spans thousands of servers).
Transparency Principles
Distributed systems use 8 transparency layers to mask heterogeneity:
- Location transparency: Users access resources without knowing their physical location (e.g., streaming a YouTube video from any server).
- Migration transparency: Processes can move between machines without user notice (e.g., cloud load balancing).
- Replication transparency: Multiple copies of data exist, but users see a single version (e.g., Daraz’s inventory across warehouses).
2. Communication in Distributed Systems
A. Message Passing
- How it works: Processes exchange messages (packets) via sockets or queues.
- Example: WhatsApp sends encrypted messages between phones over the internet.
- Pros: Simple, flexible.
- Cons: Requires manual error handling (e.g., retries for lost packets).
B. Remote Procedure Call (RPC)
RPC lets a program call a function on a remote machine as if it were local. Steps:
- Client stub marshals (serializes) arguments into a message.
- Message travels over the network (e.g., HTTP, TCP).
- Server stub unmarshals the message and invokes the function.
- Result is marshalled back to the client.
sequenceDiagram
Client->>Client Stub: Call func(args)
Client Stub->>Network: Marshal(args) → Send
Network-->>Server Stub: Receive
Server Stub->>Server: Unmarshal(args), Call func
Server-->>Server Stub: Return result
Server Stub->>Network: Marshal(result) → Send
Network-->>Client Stub: Receive
Client Stub-->>Client: Unmarshal(result)Real-world example:
- eSewa’s payment API: When you pay a bill, your app calls
eSewa.processPayment()via RPC, but the call happens on eSewa’s servers. - Fields in an RPC message:
+---------------------+ | RPC Header | +---------------------+ | Function Name | +---------------------+ | Arguments (Marshalled) | +---------------------+ | Checksum | +---------------------+
3. Clock Synchronization
Distributed systems need consistent time for:
- Ordering events (e.g., "Did transaction A happen before B?").
- Leases (e.g., "Is this file lock still valid?").
Challenges
- Clock drift: Machines’ clocks drift apart over time.
- Network latency: Messages take unpredictable time to deliver.
Solutions
| Algorithm | How It Works | Example Use Case |
|---|---|---|
| Berkeley | Master-slave: Master adjusts slaves. | NTP (Network Time Protocol). |
| Lamport Logical | Uses message timestamps (not real time). | Distributed databases (e.g., Cassandra). |
| Chronon | Hybrid: Combines physical + logical clocks. | Kubernetes cluster scheduling. |
Worked Example: Bank Transfer
- Problem: Two branches of Nabil Bank must agree on the order of transfers to avoid double-spending.
- Solution: Use Lamport timestamps to log transfers:
If timestamps are inconsistent, the system rolls back or retries.Transfer A (Branch 1 → Branch 2, timestamp 100) Transfer B (Branch 2 → Branch 1, timestamp 101)
4. Security in Distributed Systems
Threats
- Replay attacks: An attacker resends old messages (e.g., replaying a "transfer 1000 NPR" command).
- Denial-of-Service (DoS): Overloading a server (e.g., DDoS on Ncell’s website).
- Man-in-the-middle: Eavesdropping on unencrypted traffic (e.g., MITM on public Wi-Fi).
Solutions
| Mechanism | Description | Example |
|---|---|---|
| Authentication | Prove identity (e.g., passwords, OTPs). | eSewa’s SMS-based login. |
| Authorization | Check permissions (e.g., ACLs). | Google Drive’s "Viewer" role. |
| Encryption | Secure data in transit (e.g., TLS). | HTTPS on Daraz’s checkout page. |
| Digital Signatures | Verify message integrity. | Bitcoin transactions. |
How HTTPS secures WhatsApp messages (Image: Fleshgrinder and The People from The Tango! Desktop Project., Public domain, via Wikimedia Commons)
5. Consistency Models
Distributed systems trade off consistency, availability, and partition tolerance (CAP theorem).
| Model | Definition | Pros | Cons | Example Use Case |
|---|---|---|---|---|
| Strong | All nodes see the same data at the same time. | No stale reads. | High latency. | Bank accounts. |
| Eventual | Data will converge over time. | High availability. | Temporary inconsistencies. | Social media feeds. |
| Causal | Preserves cause-effect ordering. | Intuitive for users. | Complex to implement. | Chat apps (WhatsApp). |
Worked Example: NEPSE Stock Prices
- Problem: Multiple servers track stock prices. If Server A shows 1000 NPR and Server B shows 999 NPR, traders get confused.
- Solution: Use strong consistency for critical data (e.g., trades) and eventual consistency for non-critical data (e.g., news feeds).
6. Distributed vs. Centralized Systems
| Feature | Distributed System | Centralized System |
|---|---|---|
| Scalability | Scales horizontally (add more machines). | Scales vertically (bigger server). |
| Fault Tolerance | Survives node failures. | Single point of failure. |
| Cost | High initial setup (networking). | Lower cost for small scale. |
| Performance | Latency varies by location. | Low latency (local access). |
| Example | Google Cloud, eSewa. | Old mainframe banks. |
Real-world tie-in:
- Pathao’s ride-hailing: Uses a distributed system to match drivers and riders globally. If one server fails, others take over.
- NTC’s billing system: Centralized until recently; now uses distributed databases to handle peak traffic during festivals.
7. Deadlocks in Distributed Systems
Deadlocks can occur if processes hold resources while waiting for others. Example:
- Process P1 holds Resource A and waits for Resource B.
- Process P2 holds Resource B and waits for Resource A.
Prevention:
- Timeouts: Abort if a resource isn’t released within T seconds.
- Deadlock detection: Periodically check for cycles in the wait-for graph.
Worked Example: Daraz Order Queue
- Scenario: Two warehouses (A and B) hold parts of an order.
- Warehouse A ships Item 1 but waits for Item 2 from Warehouse B.
- Warehouse B ships Item 2 but waits for Item 1 from Warehouse A.
- Solution: Use a central coordinator (e.g., Kafka) to serialize order processing.
Exam Tip
- Diagrams are worth 50% of marks: Draw RPC sequences, clock synchronization algorithms, and distributed system layers clearly.
- CAP theorem is a favorite: Always explain trade-offs (e.g., "For NEPSE, we choose AP over CP to avoid downtime during Diwali").
- Real-world links: Relate examples to eSewa (RPC), Ncell (clock sync), or Daraz (consistency).
- Banker’s algorithm: Even though it’s Unit 9, it’s often tested with distributed deadlocks. Practice need matrix calculations.
- Security: Memorize authentication vs. authorization and TLS handshake steps.
Based on the TU BCA syllabus for Operating System (CACS251), unit 10.
Discussion
Loading…