CACS352 Distributed System

Distributed SystemUnit 116 min read

Distributed Systems: Definitions, Architectures & Core Concepts

Unit 1 of Distributed System: Explores foundational concepts like definitions, key characteristics, communication models, and architectural styles—with real-world examples from eSewa, Daraz, and NTC to illustrate how distributed systems work in practice.

TAKEAWAYS:

  • A distributed system is a collection of independent computers that appear as a single coherent system to users, with transparency (e.g., location, migration, replication) as its core principle.
  • Communication in distributed systems is either synchronous (request-reply) or asynchronous (event-driven), with message-passing as the primary mechanism.
  • Architectures range from client-server (eSewa’s transaction processing) to peer-to-peer (BitTorrent-like file sharing) and hybrid (Daraz’s order management).
  • Challenges like scalability, fault tolerance, and security are addressed via replication, consistency models, and encryption (e.g., NTC’s network redundancy).
  • Identifiers, names, and addresses are distinct: identifiers are unique (e.g., a user’s eSewa ID), names are human-readable (e.g., eSewa.com), and addresses are locators (e.g., IP).
  • Security in distributed systems involves authentication (OAuth for apps), authorization (role-based access in banks), and confidentiality (encrypted transactions).

1. What is a Distributed System?

A distributed system is a collection of independent computers connected over a network that appear to users as a single coherent system. Unlike centralized systems, distributed systems distribute resources, data, and processing across multiple nodes to improve scalability, reliability, and performance.

Key Definitions

  • Distributed System: A system where components located on networked computers communicate and coordinate their actions by passing messages.
  • Transparency: The ability of a distributed system to hide its distributed nature from users. Types include:
    • Access transparency: Users access data without knowing its location.
    • Location transparency: Users interact with services without specifying their location.
    • Migration transparency: Processes can move between nodes without disruption.
    • Replication transparency: Multiple copies of data exist, but users access a single logical copy.
    • Concurrency transparency: Users see a consistent state despite concurrent operations.
    • Failure transparency: The system recovers from failures automatically.

Why Distributed Systems?

Distributed systems are used to:

  • Increase scalability (e.g., NTC’s network handles millions of users).
  • Improve reliability (e.g., banks use replication to prevent data loss).
  • Enhance performance (e.g., Daraz’s servers distribute order processing).
  • Enable global access (e.g., WhatsApp’s end-to-end encryption across continents).

2. Characteristics of Distributed Systems

Distributed systems exhibit the following key traits:

Characteristic Description Example
Concurrency Multiple processes execute simultaneously. Pathao’s ride allocation system handles multiple requests at once.
Scalability System performance improves with added resources. Ncell’s network scales with more base stations.
Fault Tolerance System continues operating despite failures. NEPSE’s trading system recovers from server crashes.
Transparency Users interact with the system without awareness of its distributed nature. eSewa’s payment system hides transaction routing.
Heterogeneity System components may use different hardware/software. A bank’s mainframe servers and cloud-based analytics.
Openness System can integrate with other systems. Google Maps API used by Pathao for navigation.
Dynamic Configuration Nodes can join/leave the system without downtime. Peer-to-peer file sharing (e.g., BitTorrent).

3. Communication in Distributed Systems

Communication is the backbone of distributed systems. It can be classified based on synchrony and message-passing:

Types of Communication

  1. Synchronous Communication:

    • Request-Reply Model: A client sends a request and waits for a response.
      sequenceDiagram
        participant Client
        participant Server
        Client->>Server: Request (e.g., "Transfer Rs. 1000")
        Server-->>Client: Response (e.g., "Transaction successful")
    • Example: When you check your eSewa balance, your device sends a request to eSewa’s server and waits for the response.
  2. Asynchronous Communication:

    • Event-Driven Model: Messages are sent without waiting for a response.
      sequenceDiagram
        participant Sensor
        participant Server
        Sensor->>Server: Event (e.g., "Temperature exceeds limit")
        Server--x Server: Process event (no immediate reply)
    • Example: NTC’s network monitors for traffic congestion and routes data asynchronously.
  3. Message-Passing:

    • Messages are exchanged between processes using protocols (e.g., HTTP, gRPC).
    • Direct Communication: Process A sends a message to process B directly.
    • Indirect Communication: Messages are sent to a mailbox (e.g., a queue or mailbox in middleware).

Communication Models

Model Description Example
Client-Server Clients request services from servers. eSewa’s mobile app (client) interacts with eSewa’s backend servers.
Peer-to-Peer (P2P) Nodes act as both clients and servers. BitTorrent for file sharing (users share files directly).
Hybrid Combines client-server and P2P models. Daraz’s order system: users (clients) interact with servers, but some data is replicated across peers.

4. Architectural Styles in Distributed Systems

Distributed systems can be organized using different architectural styles:

1. Client-Server Architecture

  • Structure:
  • Advantages:
    • Centralized control and management.
    • Scalable for large user bases (e.g., NTC’s network).
  • Disadvantages:
    • Single point of failure (server downtime affects all clients).
    • Bottleneck at the server.
  • Example: eSewa’s transaction processing
    • Clients (users) send payment requests to eSewa’s centralized servers.
    • Servers process transactions and return responses.

2. Peer-to-Peer (P2P) Architecture

  • Structure:
  • Advantages:
    • No single point of failure.
    • Decentralized control (e.g., file sharing).
  • Disadvantages:
    • Harder to manage and secure.
    • Scalability issues with too many peers.
  • Example: BitTorrent-like file sharing
    • Users (peers) share files directly without a central server.

3. Hybrid Architecture

  • Structure: Combines client-server and P2P models.
flowchart TD
    A["Client"] -->|"Request"| B["Server"]
    B -->|"Data"| C["Peer"]
    C -->|"Share"| D["Other Peers"]
    C -->|"Request"| B
    B -->|"Response"| A
  • Example: Daraz’s order management
    • Users (clients) interact with Daraz’s centralized servers for orders.
    • Some data (e.g., inventory) is replicated across multiple servers (P2P-like replication).

5. Identifiers, Names, and Addresses

In distributed systems, identifiers, names, and addresses serve different purposes:

Term Definition Example
Identifier A unique label for an entity (e.g., user, process). eSewa user ID: user_12345.
Name A human-readable identifier (e.g., eSewa.com). Domain name for eSewa’s website.
Address A locator (e.g., IP address) where the entity resides. eSewa’s server IP: 192.168.1.100.

Mapping Between Names and Addresses

  • Name Resolution: Converts names (e.g., eSewa.com) to addresses (e.g., IP).
    sequenceDiagram
        participant User
        participant DNS
        participant Server
        User->>DNS: "Resolve eSewa.com"
        DNS-->>User: "192.168.1.100"
        User->>Server: Connect to 192.168.1.100
  • Example: When you type eSewa.com in your browser, your device queries a DNS server to find the IP address of eSewa’s servers.

6. Security in Distributed Systems

Security is critical in distributed systems due to their open, heterogeneous, and dynamic nature. Key security challenges and solutions:

Security Threats

Threat Description Mitigation Strategy
Unauthorized Access Attackers gain access to resources without permission. Authentication: Use OAuth, biometrics, or passwords.
Data Tampering Attackers modify data in transit or storage. Encryption: Use TLS (e.g., HTTPS for eSewa transactions).
Denial-of-Service (DoS) Attackers overwhelm the system to cause outages. Rate Limiting: Restrict request rates (e.g., NTC’s network).
Man-in-the-Middle (MITM) Attackers intercept and alter messages. Encryption: Use symmetric/asymmetric encryption (e.g., WhatsApp).
Replay Attacks Attackers repeat valid data transmissions. Timestamps & Nonces: Add unique tokens to messages.

Security Management Techniques

  1. Authentication: Verify the identity of users/processes.
    • Example: eSewa’s login system uses passwords and 2FA (Two-Factor Authentication).
  2. Authorization: Control access to resources.
    • Example: Banking systems use role-based access control (RBAC).
  3. Confidentiality: Protect data from unauthorized access.
    • Example: Encrypted transactions in eSewa.
  4. Integrity: Ensure data is not altered.
    • Example: Digital signatures for NEPSE’s trading data.
  5. Availability: Ensure services are always accessible.
    • Example: Redundant servers in NTC’s network.

7. Challenges of Distributed Systems

Despite their advantages, distributed systems face several challenges:

ApplicationLatencyTransportReliabilityNetworkScalabilityData LinkConsistencyPhysicalFault Tolerance
Key challenges in distributed systems across OSI layers.
  1. Transparency vs. Complexity:
    • While transparency hides complexity, achieving it requires sophisticated design (e.g., replication, load balancing).
  2. Scalability:
    • Adding more nodes can lead to network congestion or coordination overhead.
    • Example: Ncell’s network must scale with increasing mobile users.
  3. Fault Tolerance:
    • Nodes may fail, requiring replication and checkpointing.
    • Example: Banks use distributed commit protocols to ensure transactions are consistent across nodes.
  4. Security:
    • Open networks are vulnerable to attacks (e.g., MITM, DoS).
    • Example: eSewa uses TLS encryption to secure transactions.
  5. Consistency:
    • Ensuring all nodes have the same view of data is challenging.
    • Example: CAP Theorem trade-offs in distributed databases (e.g., NEPSE’s trading systems).

8. Real-World Examples

1. eSewa: Client-Server Architecture with Security

  • Idea Used: Client-server architecture with synchronous communication and security (encryption, authentication).
  • How It Works:
    • Users (clients) send payment requests to eSewa’s servers.
    • Servers process transactions and return responses.
    • Security: TLS encryption ensures data confidentiality and integrity.

2. Daraz: Hybrid Architecture for Order Management

  • Idea Used: Hybrid architecture (client-server + P2P-like replication).
  • How It Works:
    • Users (clients) interact with Daraz’s centralized servers for orders.
    • Inventory data is replicated across multiple servers to ensure availability and scalability.

3. NTC: Fault Tolerance in Networking

  • Idea Used: Fault tolerance via replication and redundancy.
  • How It Works:
    • NTC’s network has multiple redundant paths and servers.
    • If one node fails, traffic is rerouted automatically to maintain availability.

9. Worked Example: Distributed Commit in Banking

Scenario: A bank (e.g., NMB) wants to transfer Rs. 10,000 from Account A to Account B across two branches (Branch 1 and Branch 2).

Two-Phase Commit (2PC) Protocol

  1. Prepare Phase:
    • Branch 1 sends a prepare request to Branch 2.
      sequenceDiagram
        participant Branch1
        participant Branch2
        Branch1->>Branch2: Prepare (Transfer Rs. 10,000)
        Branch2-->>Branch1: Acknowledge (Ready to commit)
    • Both branches check if they can commit (e.g., sufficient funds in Account A).
  2. Commit Phase:
    • Branch 1 sends a commit request to Branch 2.
      sequenceDiagram
        participant Branch1
        participant Branch2
        Branch1->>Branch2: Commit
        Branch2-->>Branch1: Acknowledge (Transaction committed)
    • Both branches update their databases.
    • If Branch 2 fails to commit, it sends a rollback request.

Why 2PC?

  • Ensures atomicity (all-or-nothing transaction).
  • Used in distributed databases (e.g., NEPSE’s trading systems).

10. Exam Tip

  • Focus on Definitions: TU/PU exams often ask for precise definitions (e.g., "What is a distributed system?").
  • Diagrams are Key: Draw sequence diagrams for communication models (e.g., 2PC, client-server) and flowcharts for architectures.
  • Real-World Tie-Ins: Relate concepts to eSewa, Daraz, NTC, or banks. For example:
    • "eSewa uses client-server architecture with TLS encryption for security."
    • "NEPSE’s trading system uses distributed commit to ensure transaction consistency."
  • Compare Architectures: Always compare client-server vs. P2P vs. hybrid in your answers.
  • Security is High-Weight: Expect questions on authentication, encryption, and DoS mitigation.
  • Transparency Types: Memorize the 6 types of transparency (access, location, migration, etc.) and give examples for each.

Visual Summary

1. Distributed System Layers

A distributed system can be visualized as a stack of layers:

  1. Application Layer: User interfaces (e.g., eSewa app).
  2. Middleware Layer: RPC, ORBs, messaging systems.
  3. Transport Layer: TCP, UDP.
  4. Network Layer: IP, routing.
  5. Hardware Layer: Servers, networks, storage.

2. Client-Server vs. Peer-to-Peer

mindmap
  root((Distributed Architectures))
    Client-Server
      Clients
      Servers
      Example: eSewa
    Peer-to-Peer
      Equal Nodes
      Decentralized
      Example: BitTorrent
    Hybrid
      Client-Server + P2P
      Example: Daraz

3. Name Resolution Process


IMAGE: "distributed system layers architecture" | A layered diagram of a distributed system showing application, middleware, transport, network, and hardware layers.

IMAGE: "client server architecture diagram" | A diagram showing clients interacting with centralized servers.

IMAGE: "peer to peer network topology" | A decentralized network where nodes act as both clients and servers.

IMAGE: "router hardware rack" | A real picture of a router rack used in distributed networks like NTC.

Based on the TU BCA syllabus for Distributed System (CACS352), unit 1.

Discussion

Loading…