Distributed SystemUnit 1411 min read
Distributed OS vs. Network OS: Architectures, Workflows & Tradeoffs
Unit 14 of Distributed System: Explores how distributed operating systems (DOS) coordinate decentralized resources vs. network operating systems (NOS) that manage centralized services, comparing architectures, communication models, and real-world deployments like cloud clusters and enterprise networks.
TAKEAWAYS:
- Distributed OSs share resources across nodes as a unified system (e.g., Google’s Borg), while network OSs centralize control (e.g., Windows Server managing clients).
- Synchronization in DOS requires consensus algorithms (e.g., Paxos) to avoid conflicts, whereas NOS relies on hierarchical commands (e.g., Active Directory).
- Fault tolerance in DOS uses redundancy (e.g., Apache Kafka’s brokers), while NOS isolates failures via failover nodes (e.g., NTC’s core routers).
- Security in DOS depends on cryptographic peer-to-peer (P2P) trust (e.g., IPFS), but NOS enforces perimeter firewalls (e.g., bank VPNs).
- Scalability in DOS is horizontal (e.g., AWS EC2 auto-scaling), while NOS scales vertically (e.g., Ncell’s centralized base stations).
- Real-world tie: eSewa’s microservices (DOS) vs. NTC’s core network (NOS) illustrate how both models handle high transaction loads differently.
1. Definitions and Core Concepts
Distributed Operating System (DOS)
A DOS is an OS that manages a collection of independent computers as a single coherent system, where resources (CPU, storage, memory) are shared dynamically across nodes. Unlike traditional OSs, DOS does not rely on a central authority—instead, it uses distributed algorithms (e.g., consensus protocols) to coordinate tasks.
Key Idea: In a DOS, no single node has full control; instead, decentralized decision-making ensures resilience. Example: Google’s Borg (now Kubernetes) schedules tasks across thousands of machines without a central boss.
Network Operating System (NOS)
An NOS is an OS that manages networked resources (e.g., printers, file servers) from a centralized controller (e.g., a domain controller in Windows Server). While it coordinates network services, it does not share physical resources like a DOS does.
Key Idea: NOS centralizes management (e.g., Active Directory for user authentication) but does not distribute processing like DOS. Example: NTC’s core network OS routes traffic via centralized routers.
FIGURE 1: DOS vs. NOS Architecture
graph TD
subgraph Distributed OS
A["Node 1"] -->|"Shared"| B["Node 2"]
B -->|"Decentralized"| C["Node 3"]
C -->|"Consensus"| A
end
subgraph Network OS
D["Central Controller"] -->|"Commands"| E["Client 1"]
D -->|"Commands"| F["Client 2"]
E -->|"Requests"| D
F -->|"Requests"| D
endCaption: Left: DOS nodes communicate peer-to-peer; Right: NOS clients rely on a central controller.
2. How They Work: Workflows and Communication
Distributed OS Workflow
- Resource Discovery: Nodes advertise their capabilities (e.g., free CPU cycles) via distributed hash tables (DHTs) or gossip protocols.
- Task Scheduling: A master-slave or peer-to-peer approach assigns tasks. Example: Apache Hadoop splits big data jobs into chunks and distributes them.
- Synchronization: Uses locking mechanisms (e.g., Lamport clocks) or consensus algorithms (e.g., Paxos) to avoid race conditions.
- Fault Handling: If a node fails, replication (e.g., Kafka’s brokers) ensures data availability.
Worked Example: Google’s Borg (now Kubernetes) Scheduling
- Problem: Schedule 10,000 containers across 100,000 machines with minimal latency.
- Solution:
- Advertisement: Each machine reports its load (CPU/memory) to a global scheduler.
- Allocation: The scheduler uses bin-packing algorithms to place containers on the least busy node.
- Consensus: If two nodes report conflicting states, Paxos resolves the discrepancy.
- Fault Tolerance: If a node crashes, replica containers on other nodes take over.
Network OS Workflow
- Centralized Control: A domain controller (e.g., Windows Server) authenticates users and enforces policies.
- Client-Server Model: Clients (e.g., laptops) request services (e.g., file access) from the server.
- Resource Isolation: Each client gets a virtualized slice of resources (e.g., a shared printer queue).
- Security: Firewalls and VPNs protect the central server from attacks.
Worked Example: NTC’s Core Network OS
- Problem: Route 10M+ calls daily across Nepal’s telecom infrastructure.
- Solution:
- Centralized Routing: A core router (e.g., Cisco ASR) decides the best path for each call using OSPF protocols.
- Client Requests: Mobile phones (clients) send SIP messages to the router.
- Policy Enforcement: The router applies QoS rules (e.g., prioritize VoIP over data).
- Failover: If the core router fails, a standby router takes over in <100ms.
FIGURE 2: DOS Task Scheduling vs. NOS Command Flow
sequenceDiagram
participant Node1 as Node 1
participant Node2 as Node 2
participant Master as Master Scheduler
Node1->>Master: "I have 80% CPU free"
Node2->>Master: "I have 60% CPU free"
Master->>Node1: "Assign Task X"
Master->>Node2: "Assign Task Y"
Note right of Master: Distributed OS (decentralized)sequenceDiagram
participant Client as Client PC
participant Server as Domain Controller
Client->>Server: "Authenticate User"
Server-->>Client: "Grant Access"
Client->>Server: "Request File"
Server-->>Client: "Send File"
Note right of Server: Network OS (centralized)3. Key Differences: Comparison Table
| Feature | Distributed OS (DOS) | Network OS (NOS) |
|---|---|---|
| Control | Decentralized (no single point) | Centralized (single authority) |
| Resource Sharing | Physical resources (CPU, storage) shared dynamically | Virtualized resources (e.g., shared printers) |
| Scalability | Horizontal (add more nodes) | Vertical (upgrade central server) |
| Fault Tolerance | Replication (e.g., Kafka brokers) | Failover (standby servers) |
| Communication | Peer-to-peer (e.g., gossip protocols) | Client-server (e.g., HTTP, RPC) |
| Security Model | Cryptographic (e.g., IPFS) | Perimeter (firewalls, VPNs) |
| Example | Google Borg, Apache Hadoop | Windows Server, NTC Core Network |
4. Real-World Applications
In the Real World
eSewa’s Microservices (DOS)
- Idea Used: Decentralized task scheduling (like Borg).
- How: eSewa’s payment processing is split across microservices (e.g., fraud detection, transaction logging) running on Kubernetes clusters. If one service fails, others compensate.
- Why DOS? High availability: If a region’s server crashes, transactions route to another cluster.
Ncell’s Core Network (NOS)
- Idea Used: Centralized routing control (like NTC’s OS).
- How: Ncell’s core routers (e.g., Ericsson) use MPLS protocols to route calls/data via a centralized path. If a local switch fails, the core re-routes traffic instantly.
- Why NOS? Predictable latency: Centralized control ensures all calls follow the same QoS rules.
Daraz’s Order Queue (Hybrid Approach)
- Idea Used: DOS for scalability + NOS for inventory control.
- How:
- DOS: Order processing is distributed across regional data centers (e.g., Kathmandu, Pokhara) to handle spikes.
- NOS: Inventory management is centralized (e.g., a MySQL database in Singapore) to avoid double-selling.
- Why Hybrid? Balances speed (DOS) and accuracy (NOS).
5. Advantages and Disadvantages
Distributed OS
| Advantages | Disadvantages |
|---|---|
| High Availability: No single point of failure. | Complexity: Hard to debug (distributed logs). |
| Scalability: Add nodes without redesign. | Latency: Peer-to-peer communication can be slow. |
| Resilience: Faults are localized. | Consensus Overhead: Algorithms like Paxos slow down decisions. |
| Flexibility: Resources can be pooled dynamically. | Security Risks: More attack surfaces (e.g., Sybil attacks). |
Network OS
| Advantages | Disadvantages |
|---|---|
| Simplicity: Easier to manage (centralized). | Bottleneck: Single point of failure (server). |
| Predictability: Deterministic performance. | Scalability Limits: Vertical scaling is expensive. |
| Security: Firewalls protect the core. | Downtime: Central server failure halts all services. |
| Policy Enforcement: Uniform rules across clients. | Latency: Remote clients may experience delays. |
6. When to Use Which?
| Use DOS When... | Use NOS When... |
|---|---|
| You need ultra-high availability (e.g., cloud services). | You need strict control (e.g., enterprise file servers). |
| Horizontal scaling is critical (e.g., social media backends). | Centralized logging/auditing is required (e.g., banks). |
| Decentralized trust is needed (e.g., blockchain). | Low-latency responses are critical (e.g., VoIP). |
| Resource pooling is desired (e.g., Hadoop clusters). | Simple deployment is preferred (e.g., small offices). |
FIGURE 3: DOS vs. NOS Tradeoff Spectrum
mindmap
root((DOS vs. NOS Tradeoffs))
Decentralization
High Availability
Scalability
Centralization
Control
Predictability
Use Cases
DOS: Cloud Services
NOS: Enterprise Networks7. Exam Tip
- Focus on Contrasts: Exams test how DOS and NOS differ in control, scalability, and fault tolerance. Use the comparison table above to structure answers.
- Real-World Tie: Always link concepts to Nepalese examples (e.g., eSewa for DOS, NTC for NOS). Example:
"Unlike NTC’s centralized routing (NOS), eSewa’s payment system uses a DOS to distribute transactions across servers in Kathmandu and Pokhara, ensuring no single failure halts payments."
- Diagrams: Draw sequence diagrams for workflows (e.g., task scheduling in DOS vs. client-server in NOS).
- Common Pitfalls:
- ❌ Saying "DOS has a central server" (it doesn’t).
- ❌ Confusing distributed OS with distributed computing (e.g., MapReduce).
- ❌ Forgetting to mention consensus algorithms (Paxos, Raft) in DOS explanations.
Based on the TU BCA syllabus for Distributed System (CACS352), unit 14.
Discussion
Loading…