BIT451 Network and System Administration

Network and System AdministrationUnit 1010 min read

Network Monitoring & Management: Tools, Metrics, Bandwidth, SDN

Unit 10 of Network and System Administration covers proactive monitoring (SNMP, Nagios, Zabbix), bandwidth management (QoS, traffic shaping), multicast protocols (IGMP, PIM), and Software-Defined Networking (SDN) architectures. Learn how to measure performance, detect faults, and optimize real-world networks like Ncell

TAKEAWAYS:

  • Monitoring ≠ Management: Tools like SNMP collect data (CPU, latency), while Zabbix/Nagios trigger alerts and automate fixes.
  • Bandwidth is a shared resource: QoS (DSCP, WRR) prioritizes voice/video over bulk transfers—critical for Pathao’s ride-hailing app.
  • Multicast saves bandwidth: IGMP/PIM let one stream (e.g., NTC’s live news) reach thousands without duplicating packets.
  • SDN decouples control from data: OpenFlow lets NEPSE’s trading platform dynamically reroute orders during peak hours.
  • Protocols have trade-offs: SNMPv3 is secure but complex; NetFlow is accurate but resource-heavy.
  • Real-world impact: A 1% latency drop in Daraz’s CDN can boost sales by 5–10%—monitoring directly drives revenue.

Core Concepts: What Is Network Monitoring?

Network monitoring is the continuous collection, analysis, and visualization of data from network devices (routers, switches, servers) to ensure:

  • Performance (latency, throughput)
  • Availability (uptime, failures)
  • Security (unauthorized access, DDoS)
[object Object][object Object][object Object][object Object]Ncell 5G Base StationNTC RouterSwitchServerWorkstation
Latency monitoring path: Ncell 5G base station → NTC server (real-world example with labeled link speeds)

How It Works: The Monitoring Cycle

stateDiagram-v2
    [*] --> Collect: SNMP/NetFlow pulls metrics
    Collect --> Analyze: Compare vs. baselines
    Analyze --> Alert: Threshold breached?
    Alert --> Act: Auto-remediate or notify admin
    Act --> [*]

Key Metrics Monitored

Category Metrics Tools
Performance Latency, jitter, packet loss, throughput Ping, MTR, iPerf
Availability Uptime, MTTR (Mean Time to Repair) Nagios, Zabbix
Security Failed logins, port scans, malware Wireshark, SIEM (e.g., Splunk)
Bandwidth Utilization, bottlenecks PRTG, SolarWinds
024.9849.9574.9399.9Latency (ms)15Throughput (Mbps)85Packet Loss (%)0.5Uptime (%)99.9
Example baseline metrics for a corporate network (Nepal Telecom)

Tools and Protocols: How Data Is Collected

sequenceDiagram
    participant Manager as SNMP Manager (Zabbix)
    participant Agent as SNMP Agent (Cisco Router)
    participant MIB as MIB Database
    Manager->>Agent: GET-REQUEST (ifInOctets)
    Agent->>MIB: Query OID 1.3.6.1.2.1.2.2.1.10.1
    MIB-->>Agent: 12,000 bytes
    Agent-->>Manager: GET-RESPONSE (12,000 bytes)
    Manager->>Agent: SET-REQUEST (threshold=10,000)
    Agent->>MIB: Update threshold
    MIB-->>Agent: ACK
    Agent-->>Manager: SET-RESPONSE (OK)
    Note right of MIB: MIB OID: 1.3.6.1.2.1.2.2.1.10.1 (ifInOctets)
    Note right of Manager: Real-world: Zabbix querying Cisco router via SNMPv2

1. SNMP (Simple Network Management Protocol)

  • Purpose: Standard for querying/managing network devices (routers, switches).
  • How it works:
    • Manager (e.g., Zabbix) sends GET/SET requests to agents (SNMP daemons) on devices.
    • Agents respond with MIB (Management Information Base) data (e.g., ifInOctets for traffic).
  • Versions:
    • SNMPv1/v2c: Unencrypted (insecure).
    • SNMPv3: Encrypted (authentication + privacy).

2. NetFlow/sFlow

  • Purpose: Track per-flow traffic (source/destination IP, port, bytes/packets).
  • Use case: Identify top talkers (e.g., a Daraz server flooding logs) or DDoS attacks.
  • Difference from SNMP:
    SNMP NetFlow
    Device-centric (CPU, memory) Flow-centric (traffic patterns)
    Polling-based Continuous sampling (low overhead)
    Example: "Switch CPU at 90%" "Top 5 flows: 192.168.1.100 → YouTube"

3. Open-Source Tools

  • Nagios: Alerts on thresholds (e.g., "Disk > 90%").
  • Zabbix: Agentless monitoring + auto-remediation (e.g., restart a crashed service).
  • PRTG: Bandwidth monitoring with historical trends.

Example: Ncell monitors 5G base stations using SNMP to detect overheating and auto-throttle traffic during peak hours.


Bandwidth Management: QoS and Traffic Shaping

VoIP (EF)Video (AF41)Data (BE)Torrent (Dropped)FRONTREARoutin
WRR queue example: Pathao’s QoS prioritizes voice (EF) over bulk transfers (BE)

Why Manage Bandwidth?

  • Congestion: Too many devices (e.g., Kathmandu traffic during Dashain) → packet loss.
  • Prioritization: Voice (Pathao calls) > Video (YouTube) > Bulk (eSewa file uploads).
  • Cost: ISPs charge by bandwidth usage (e.g., NTC’s enterprise clients).

Techniques

  1. Quality of Service (QoS)

    • DSCP (Differentiated Services Code Point): Marks packets (e.g., EF=46 for VoIP).
    • WRR (Weighted Round Robin): Allocates bandwidth (e.g., 60% to video, 30% to data).
    • CBQ (Class-Based Queuing): Hierarchical queues (e.g., "Gold" for NEPSE traders).
  2. Traffic Shaping

    • Token Bucket: Limits burst traffic (e.g., Daraz’s CDN caps sudden spikes).
    • Policing: Drops excess traffic (e.g., block torrenting on Ncell’s 4G).

Worked Example: Pathao’s Ride-Hailing QoS

  • Problem: Drivers complain of laggy GPS during peak hours (10 AM–6 PM).
  • Solution:
    • Classify traffic: GPS updates (UDP port 5000) → EF (Expedited Forwarding).
    • Police other apps: Limit background syncs to 10% bandwidth.
    • Result: 95% of GPS packets arrive <50ms; driver complaints drop by 80%.

Multicast: Efficient One-to-Many Communication

How It Works

  • Unicast: Sender → Each receiver (wastes bandwidth).
  • Multicast: Sender → Single group address (e.g., IGMPv3).
  • Protocols:
    • IGMP (Internet Group Management Protocol): Hosts join/leave groups.
    • PIM (Protocol Independent Multicast): Routers forward multicast traffic.

Real-World Use Cases

  1. NTC’s Live News: One stream to millions of TVs without duplicating packets.
  2. Stock Tickers (NEPSE): Real-time price updates to brokers.
  3. IPTV (Kantipur TV): Single feed to subscribers.

Example Trace:

  1. NTC’s server sends a multicast stream to group 239.1.1.1.
  2. Routers use PIM to forward only to subscribers (not the entire internet).
  3. Your TV’s set-top box joins the group via IGMP.

Software-Defined Networking (SDN): Centralized Control

[object Object][object Object][object Object][object Object]SDN ControllerSwitch 1Switch 2Host AHost B
SDN control plane (controller) managing multiple switches via OpenFlow protocol

What Is SDN?

  • Traditional networks: Control plane (routing decisions) is distributed in each device.
  • SDN: Control plane is centralized (SDN controller) + decoupled from data plane.
Application (e.g.,OpenDaylight)SDN ControllerSouthbound API (OpenFlow)Data Plane (Switches/Devices)Centralized control → Decoupled data plane
SDN architecture: Centralized control plane (controller) decoupled from data plane (switches)

How It Helps

  • Dynamic routing: NEPSE’s trading platform reroutes orders during peak hours.
  • Security: Whitelist/blacklist apps (e.g., block crypto mining on Ncell’s network).
  • Cost: Reduce hardware (e.g., Google’s SDN saves 30% on capex).

Example: Khalti uses SDN to prioritize payment confirmations over email traffic during Diwali sales.


Network Management: Beyond Monitoring

1. Fault Management

  • Tools: Nagios, Zabbix (auto-restart services).
  • Example: If a Daraz server crashes, Zabbix pings it every 5 seconds and restarts Apache.

2. Configuration Management

  • Tools: Ansible, Puppet (push configs to 100+ switches).
  • Example: NTC deploys new firewall rules to all routers via Ansible.

3. Performance Tuning

  • Techniques:
    • Baselining: Compare current metrics to historical data.
    • Load testing: Simulate 10,000 users on eSewa’s website.

In the Real World

  1. eSewa’s Payment Gateway

    • Monitoring: Uses Zabbix to track transaction latency (alerts if >200ms).
    • QoS: Prioritizes HTTPS traffic (port 443) over FTP uploads.
    • SDN: Dynamically scales servers during Dashain (5x normal load).
  2. Ncell’s 5G Network

    • Multicast: Broadcasts emergency alerts to all phones in a zone.
    • Bandwidth Management: Throttles Netflix during peak hours to prevent congestion.
  3. NEPSE’s Trading System

    • SDN: Reroutes orders to the nearest data center during market open.
    • NetFlow: Detects insider trading by spotting unusual flow patterns.

Exam Tip

How This Unit Is Tested:

  1. Definitions: Know SNMP vs. NetFlow vs. QoS (1–2 marks each).
  2. Scenarios: Given a latency issue, pick the right tool (e.g., "Use MTR to diagnose Pathao’s GPS lag").
  3. Diagrams: Draw a QoS queue or SDN architecture (label all layers).
  4. Calculations:
    • Bandwidth utilization = (Used Bandwidth / Total Bandwidth) × 100.
    • Example: A 100 Mbps link carries 70 Mbps → 70% utilization.
  5. Pros/Cons: Compare multicast vs. unicast (e.g., "Multicast saves bandwidth but requires IGMP support").

Common Pitfalls:

  • Confusing monitoring (data collection) with management (actions).
  • Forgetting SNMPv3 is secure but SNMPv1 is not.
  • Ignoring real-world constraints (e.g., "Ncell can’t use multicast for all traffic due to legacy devices").

Based on the TU BIT syllabus for Network and System Administration (BIT451), unit 10.

Discussion

Loading…