CACS251 Operating System

Operating SystemUnit 1011 min read

Distributed Systems: Architecture, RPC, Security & Synchronization

Unit 10 of Operating System explores distributed systems—how multiple independent computers collaborate as a single system, covering transparency principles, communication methods (RPC, message passing), clock synchronization, security challenges, and real-world applications like cloud services and e-payment systems.

TAKEAWAYS:

  • Distributed systems hide complexity (e.g., location, migration, replication) to appear as a single OS, using transparency layers like RPC and message passing.
  • Remote Procedure Call (RPC) lets programs call functions on remote machines as if they were local, but requires marshalling (data serialization) and stub layers.
  • Clock synchronization (e.g., Berkeley algorithm) is critical for distributed transactions (e.g., bank transfers) but faces drift and causality issues.
  • Security threats (e.g., replay attacks, denial-of-service) demand authentication (passwords, biometrics) and authorization (access control lists).
  • Consistency models (e.g., strong vs. eventual) trade off between correctness and performance—CAP theorem proves you can’t have all three (Consistency, Availability, Partition tolerance).
  • Real-world use: eSewa uses distributed transactions for secure payments, while YouTube’s CDN relies on geographically replicated servers for low latency.

1. What is a Distributed System?

A distributed system is a collection of independent computers that appear to users as a single coherent system. Unlike centralized systems (e.g., a single mainframe), distributed systems:

  • Share resources (e.g., files, printers, databases) across networks.
  • Hide complexity (e.g., location of data, failures) via transparency.
  • Enable scalability (e.g., Google’s search index spans thousands of servers).

Transparency Principles

Distributed systems use 8 transparency layers to mask heterogeneity:

Access Transparency (hide differences in data representationLocation Transparency (hide where services are located)Migration Transparency (allow processes to move without userReplication Transparency (hide multiple copies of data)Persistence Transparency (hide volatility of storage)Concurrency Transparency (hide interleaved operations)Transparency Layers
Hierarchical breakdown of the 8 transparency layers in distributed systems
  • Location transparency: Users access resources without knowing their physical location (e.g., streaming a YouTube video from any server).
  • Migration transparency: Processes can move between machines without user notice (e.g., cloud load balancing).
  • Replication transparency: Multiple copies of data exist, but users see a single version (e.g., Daraz’s inventory across warehouses).

2. Communication in Distributed Systems

Function callMarshalled messageNetwork transportUnmarshalled callReturn valueMarshalled resultNetwork transportUnmarshalled resultClientClient StubNetworkServer StubServer
Physical network topology for RPC communication (simplified)

A. Message Passing

  • How it works: Processes exchange messages (packets) via sockets or queues.
  • Example: WhatsApp sends encrypted messages between phones over the internet.
  • Pros: Simple, flexible.
  • Cons: Requires manual error handling (e.g., retries for lost packets).

B. Remote Procedure Call (RPC)

RPC lets a program call a function on a remote machine as if it were local. Steps:

  1. Client stub marshals (serializes) arguments into a message.
  2. Message travels over the network (e.g., HTTP, TCP).
  3. Server stub unmarshals the message and invokes the function.
  4. Result is marshalled back to the client.
sequenceDiagram
    Client->>Client Stub: Call func(args)
    Client Stub->>Network: Marshal(args) → Send
    Network-->>Server Stub: Receive
    Server Stub->>Server: Unmarshal(args), Call func
    Server-->>Server Stub: Return result
    Server Stub->>Network: Marshal(result) → Send
    Network-->>Client Stub: Receive
    Client Stub-->>Client: Unmarshal(result)

Real-world example:

  • eSewa’s payment API: When you pay a bill, your app calls eSewa.processPayment() via RPC, but the call happens on eSewa’s servers.
  • Fields in an RPC message:
    +---------------------+
    | RPC Header          |
    +---------------------+
    | Function Name       |
    +---------------------+
    | Arguments (Marshalled) |
    +---------------------+
    | Checksum            |
    +---------------------+
    

3. Clock Synchronization

Distributed systems need consistent time for:

  • Ordering events (e.g., "Did transaction A happen before B?").
  • Leases (e.g., "Is this file lock still valid?").

Challenges

  • Clock drift: Machines’ clocks drift apart over time.
  • Network latency: Messages take unpredictable time to deliver.

Solutions

Algorithm How It Works Example Use Case
Berkeley Master-slave: Master adjusts slaves. NTP (Network Time Protocol).
Lamport Logical Uses message timestamps (not real time). Distributed databases (e.g., Cassandra).
Chronon Hybrid: Combines physical + logical clocks. Kubernetes cluster scheduling.

Worked Example: Bank Transfer

  • Problem: Two branches of Nabil Bank must agree on the order of transfers to avoid double-spending.
  • Solution: Use Lamport timestamps to log transfers:
    Transfer A (Branch 1 → Branch 2, timestamp 100)
    Transfer B (Branch 2 → Branch 1, timestamp 101)
    
    If timestamps are inconsistent, the system rolls back or retries.

4. Security in Distributed Systems

Threats

  • Replay attacks: An attacker resends old messages (e.g., replaying a "transfer 1000 NPR" command).
  • Denial-of-Service (DoS): Overloading a server (e.g., DDoS on Ncell’s website).
  • Man-in-the-middle: Eavesdropping on unencrypted traffic (e.g., MITM on public Wi-Fi).

Solutions

Mechanism Description Example
Authentication Prove identity (e.g., passwords, OTPs). eSewa’s SMS-based login.
Authorization Check permissions (e.g., ACLs). Google Drive’s "Viewer" role.
Encryption Secure data in transit (e.g., TLS). HTTPS on Daraz’s checkout page.
Digital Signatures Verify message integrity. Bitcoin transactions.

TLS handshake diagram**How HTTPS secures WhatsApp messages (Image: Fleshgrinder and The People from The Tango! Desktop Project., Public domain, via Wikimedia Commons)


5. Consistency Models

Distributed systems trade off consistency, availability, and partition tolerance (CAP theorem).

T1Write A=1 (Node 1)T2Read A (Node 2) →returns 1T3Write A=2 (Node 1)T4Read A (Node 2) →returns 2 (Strong ConsT5Read A (Node 2) →might return 1 (Eventu
Consistency models visualized through write-read operations
Model Definition Pros Cons Example Use Case
Strong All nodes see the same data at the same time. No stale reads. High latency. Bank accounts.
Eventual Data will converge over time. High availability. Temporary inconsistencies. Social media feeds.
Causal Preserves cause-effect ordering. Intuitive for users. Complex to implement. Chat apps (WhatsApp).

Worked Example: NEPSE Stock Prices

  • Problem: Multiple servers track stock prices. If Server A shows 1000 NPR and Server B shows 999 NPR, traders get confused.
  • Solution: Use strong consistency for critical data (e.g., trades) and eventual consistency for non-critical data (e.g., news feeds).

6. Distributed vs. Centralized Systems

Feature Distributed System Centralized System
Scalability Scales horizontally (add more machines). Scales vertically (bigger server).
Fault Tolerance Survives node failures. Single point of failure.
Cost High initial setup (networking). Lower cost for small scale.
Performance Latency varies by location. Low latency (local access).
Example Google Cloud, eSewa. Old mainframe banks.

Real-world tie-in:

  • Pathao’s ride-hailing: Uses a distributed system to match drivers and riders globally. If one server fails, others take over.
  • NTC’s billing system: Centralized until recently; now uses distributed databases to handle peak traffic during festivals.

7. Deadlocks in Distributed Systems

Deadlocks can occur if processes hold resources while waiting for others. Example:

  • Process P1 holds Resource A and waits for Resource B.
  • Process P2 holds Resource B and waits for Resource A.

Prevention:

  • Timeouts: Abort if a resource isn’t released within T seconds.
  • Deadlock detection: Periodically check for cycles in the wait-for graph.

Worked Example: Daraz Order Queue

  • Scenario: Two warehouses (A and B) hold parts of an order.
    • Warehouse A ships Item 1 but waits for Item 2 from Warehouse B.
    • Warehouse B ships Item 2 but waits for Item 1 from Warehouse A.
  • Solution: Use a central coordinator (e.g., Kafka) to serialize order processing.

Exam Tip

  1. Diagrams are worth 50% of marks: Draw RPC sequences, clock synchronization algorithms, and distributed system layers clearly.
  2. CAP theorem is a favorite: Always explain trade-offs (e.g., "For NEPSE, we choose AP over CP to avoid downtime during Diwali").
  3. Real-world links: Relate examples to eSewa (RPC), Ncell (clock sync), or Daraz (consistency).
  4. Banker’s algorithm: Even though it’s Unit 9, it’s often tested with distributed deadlocks. Practice need matrix calculations.
  5. Security: Memorize authentication vs. authorization and TLS handshake steps.

Based on the TU BCA syllabus for Operating System (CACS251), unit 10.

Discussion

Loading…