CACS352 Distributed System

Distributed SystemUnit 1411 min read

Distributed OS vs. Network OS: Architectures, Workflows & Tradeoffs

Unit 14 of Distributed System: Explores how distributed operating systems (DOS) coordinate decentralized resources vs. network operating systems (NOS) that manage centralized services, comparing architectures, communication models, and real-world deployments like cloud clusters and enterprise networks.

TAKEAWAYS:

  • Distributed OSs share resources across nodes as a unified system (e.g., Google’s Borg), while network OSs centralize control (e.g., Windows Server managing clients).
  • Synchronization in DOS requires consensus algorithms (e.g., Paxos) to avoid conflicts, whereas NOS relies on hierarchical commands (e.g., Active Directory).
  • Fault tolerance in DOS uses redundancy (e.g., Apache Kafka’s brokers), while NOS isolates failures via failover nodes (e.g., NTC’s core routers).
  • Security in DOS depends on cryptographic peer-to-peer (P2P) trust (e.g., IPFS), but NOS enforces perimeter firewalls (e.g., bank VPNs).
  • Scalability in DOS is horizontal (e.g., AWS EC2 auto-scaling), while NOS scales vertically (e.g., Ncell’s centralized base stations).
  • Real-world tie: eSewa’s microservices (DOS) vs. NTC’s core network (NOS) illustrate how both models handle high transaction loads differently.

1. Definitions and Core Concepts

Distributed Operating System (DOS)

A DOS is an OS that manages a collection of independent computers as a single coherent system, where resources (CPU, storage, memory) are shared dynamically across nodes. Unlike traditional OSs, DOS does not rely on a central authority—instead, it uses distributed algorithms (e.g., consensus protocols) to coordinate tasks.

Key Idea: In a DOS, no single node has full control; instead, decentralized decision-making ensures resilience. Example: Google’s Borg (now Kubernetes) schedules tasks across thousands of machines without a central boss.

Network Operating System (NOS)

An NOS is an OS that manages networked resources (e.g., printers, file servers) from a centralized controller (e.g., a domain controller in Windows Server). While it coordinates network services, it does not share physical resources like a DOS does.

Key Idea: NOS centralizes management (e.g., Active Directory for user authentication) but does not distribute processing like DOS. Example: NTC’s core network OS routes traffic via centralized routers.


FIGURE 1: DOS vs. NOS Architecture

graph TD
    subgraph Distributed OS
        A["Node 1"] -->|"Shared"| B["Node 2"]
        B -->|"Decentralized"| C["Node 3"]
        C -->|"Consensus"| A
    end
    subgraph Network OS
        D["Central Controller"] -->|"Commands"| E["Client 1"]
        D -->|"Commands"| F["Client 2"]
        E -->|"Requests"| D
        F -->|"Requests"| D
    end

Caption: Left: DOS nodes communicate peer-to-peer; Right: NOS clients rely on a central controller.


2. How They Work: Workflows and Communication

Distributed OS Workflow

  1. Resource Discovery: Nodes advertise their capabilities (e.g., free CPU cycles) via distributed hash tables (DHTs) or gossip protocols.
  2. Task Scheduling: A master-slave or peer-to-peer approach assigns tasks. Example: Apache Hadoop splits big data jobs into chunks and distributes them.
  3. Synchronization: Uses locking mechanisms (e.g., Lamport clocks) or consensus algorithms (e.g., Paxos) to avoid race conditions.
  4. Fault Handling: If a node fails, replication (e.g., Kafka’s brokers) ensures data availability.

Worked Example: Google’s Borg (now Kubernetes) Scheduling

  • Problem: Schedule 10,000 containers across 100,000 machines with minimal latency.
  • Solution:
    1. Advertisement: Each machine reports its load (CPU/memory) to a global scheduler.
    2. Allocation: The scheduler uses bin-packing algorithms to place containers on the least busy node.
    3. Consensus: If two nodes report conflicting states, Paxos resolves the discrepancy.
    4. Fault Tolerance: If a node crashes, replica containers on other nodes take over.

Network OS Workflow

  1. Centralized Control: A domain controller (e.g., Windows Server) authenticates users and enforces policies.
  2. Client-Server Model: Clients (e.g., laptops) request services (e.g., file access) from the server.
  3. Resource Isolation: Each client gets a virtualized slice of resources (e.g., a shared printer queue).
  4. Security: Firewalls and VPNs protect the central server from attacks.

Worked Example: NTC’s Core Network OS

  • Problem: Route 10M+ calls daily across Nepal’s telecom infrastructure.
  • Solution:
    1. Centralized Routing: A core router (e.g., Cisco ASR) decides the best path for each call using OSPF protocols.
    2. Client Requests: Mobile phones (clients) send SIP messages to the router.
    3. Policy Enforcement: The router applies QoS rules (e.g., prioritize VoIP over data).
    4. Failover: If the core router fails, a standby router takes over in <100ms.

FIGURE 2: DOS Task Scheduling vs. NOS Command Flow

sequenceDiagram
    participant Node1 as Node 1
    participant Node2 as Node 2
    participant Master as Master Scheduler
    Node1->>Master: "I have 80% CPU free"
    Node2->>Master: "I have 60% CPU free"
    Master->>Node1: "Assign Task X"
    Master->>Node2: "Assign Task Y"
    Note right of Master: Distributed OS (decentralized)
sequenceDiagram
    participant Client as Client PC
    participant Server as Domain Controller
    Client->>Server: "Authenticate User"
    Server-->>Client: "Grant Access"
    Client->>Server: "Request File"
    Server-->>Client: "Send File"
    Note right of Server: Network OS (centralized)

3. Key Differences: Comparison Table

Feature Distributed OS (DOS) Network OS (NOS)
Control Decentralized (no single point) Centralized (single authority)
Resource Sharing Physical resources (CPU, storage) shared dynamically Virtualized resources (e.g., shared printers)
Scalability Horizontal (add more nodes) Vertical (upgrade central server)
Fault Tolerance Replication (e.g., Kafka brokers) Failover (standby servers)
Communication Peer-to-peer (e.g., gossip protocols) Client-server (e.g., HTTP, RPC)
Security Model Cryptographic (e.g., IPFS) Perimeter (firewalls, VPNs)
Example Google Borg, Apache Hadoop Windows Server, NTC Core Network

4. Real-World Applications

In the Real World

  1. eSewa’s Microservices (DOS)

    • Idea Used: Decentralized task scheduling (like Borg).
    • How: eSewa’s payment processing is split across microservices (e.g., fraud detection, transaction logging) running on Kubernetes clusters. If one service fails, others compensate.
    • Why DOS? High availability: If a region’s server crashes, transactions route to another cluster.
  2. Ncell’s Core Network (NOS)

    • Idea Used: Centralized routing control (like NTC’s OS).
    • How: Ncell’s core routers (e.g., Ericsson) use MPLS protocols to route calls/data via a centralized path. If a local switch fails, the core re-routes traffic instantly.
    • Why NOS? Predictable latency: Centralized control ensures all calls follow the same QoS rules.
  3. Daraz’s Order Queue (Hybrid Approach)

    • Idea Used: DOS for scalability + NOS for inventory control.
    • How:
      • DOS: Order processing is distributed across regional data centers (e.g., Kathmandu, Pokhara) to handle spikes.
      • NOS: Inventory management is centralized (e.g., a MySQL database in Singapore) to avoid double-selling.
    • Why Hybrid? Balances speed (DOS) and accuracy (NOS).


5. Advantages and Disadvantages

Distributed OS

Advantages Disadvantages
High Availability: No single point of failure. Complexity: Hard to debug (distributed logs).
Scalability: Add nodes without redesign. Latency: Peer-to-peer communication can be slow.
Resilience: Faults are localized. Consensus Overhead: Algorithms like Paxos slow down decisions.
Flexibility: Resources can be pooled dynamically. Security Risks: More attack surfaces (e.g., Sybil attacks).

Network OS

Advantages Disadvantages
Simplicity: Easier to manage (centralized). Bottleneck: Single point of failure (server).
Predictability: Deterministic performance. Scalability Limits: Vertical scaling is expensive.
Security: Firewalls protect the core. Downtime: Central server failure halts all services.
Policy Enforcement: Uniform rules across clients. Latency: Remote clients may experience delays.

6. When to Use Which?

Use DOS When... Use NOS When...
You need ultra-high availability (e.g., cloud services). You need strict control (e.g., enterprise file servers).
Horizontal scaling is critical (e.g., social media backends). Centralized logging/auditing is required (e.g., banks).
Decentralized trust is needed (e.g., blockchain). Low-latency responses are critical (e.g., VoIP).
Resource pooling is desired (e.g., Hadoop clusters). Simple deployment is preferred (e.g., small offices).

FIGURE 3: DOS vs. NOS Tradeoff Spectrum

mindmap
  root((DOS vs. NOS Tradeoffs))
    Decentralization
      High Availability
      Scalability
    Centralization
      Control
      Predictability
    Use Cases
      DOS: Cloud Services
      NOS: Enterprise Networks

7. Exam Tip

  • Focus on Contrasts: Exams test how DOS and NOS differ in control, scalability, and fault tolerance. Use the comparison table above to structure answers.
  • Real-World Tie: Always link concepts to Nepalese examples (e.g., eSewa for DOS, NTC for NOS). Example:

    "Unlike NTC’s centralized routing (NOS), eSewa’s payment system uses a DOS to distribute transactions across servers in Kathmandu and Pokhara, ensuring no single failure halts payments."

  • Diagrams: Draw sequence diagrams for workflows (e.g., task scheduling in DOS vs. client-server in NOS).
  • Common Pitfalls:
    • ❌ Saying "DOS has a central server" (it doesn’t).
    • ❌ Confusing distributed OS with distributed computing (e.g., MapReduce).
    • ❌ Forgetting to mention consensus algorithms (Paxos, Raft) in DOS explanations.

Based on the TU BCA syllabus for Distributed System (CACS352), unit 14.

Discussion

Loading…