IT271 Networking and System Administration

Networking and System AdministrationUnit 913 min read

Network Monitoring & Backup: Tools, Strategies & Recovery

Unit 9 of Networking and System Administration covers proactive network monitoring (SNMP, syslog, Nagios), backup strategies (full/incremental/differential), disaster recovery plans, and real-world implementations in Nepali IT infrastructure like NTC’s network oversight and bank transaction logs.

TAKEAWAYS:

  • Network monitoring uses SNMP, syslog, and Nagios to track device health, traffic, and errors in real time—critical for troubleshooting outages like NTC’s fiber cuts.
  • Backup types (full, incremental, differential) differ in storage needs and recovery speed; banks use differential backups for daily transaction logs to balance speed and space.
  • RAID levels (0–6) trade redundancy for performance—RAID 1 mirrors data (used in Daraz’s servers) while RAID 5 stripes with parity (common in small offices).
  • Disaster recovery plans prioritize RTO (recovery time) and RPO (data loss tolerance); Ncell’s backup cell towers activate within 30 minutes (RTO) to minimize downtime.
  • Log analysis (via grep, awk, or ELK stack) uncovers security breaches—e.g., Pathao’s fraud detection flags unusual API calls in real-time logs.
  • Backup verification (e.g., restore tests) is as critical as the backup itself; NEPSE’s daily market data backups are validated nightly to ensure no corruption.

1. Network Monitoring: Watching the Digital Nervous System

Network monitoring is the proactive observation of network devices (routers, switches, servers) and traffic to detect anomalies, performance bottlenecks, or security threats before they disrupt services. Think of it as a doctor’s ECG for your network: continuous, data-driven, and alerting you to irregularities.

sequenceDiagram
    participant Router as NTC Router
    participant SNMP as SNMP Agent
    participant Nagios as Nagios Server
    participant Admin as Network Admin

    Router->>SNMP: CPU=98% (SNMP Trap)
    SNMP->>Nagios: Forward to Nagios
    Nagios->>Admin: Alert: "Router123 CPU Critical"
    Admin->>Router: SSH Command: "restart service"
    Router-->>Admin: Service Restarted
    Router-->>SNMP: CPU=45% (SNMP Update)
    SNMP-->>Nagios: Update Nagios Dashboard
SNMP trap handling workflow for NTC’s router monitoring system.
ApplicationDataTransportSegmentNetworkPacketData LinkFramePhysicalBits
OSI Model layers where SNMP (Network layer) and NetFlow (Transport/Network) operate for monitoring.

Key Components

Component Role Example Tools/Protocols
SNMP (Simple Network Management Protocol) Polls devices for stats (CPU, memory, interface errors) via GET/SET messages. snmpwalk, Zabbix, PRTG
Syslog Centralized logging of events (e.g., "Router X dropped 500 packets"). rsyslog, Graylog, Splunk
Nagios Monitors services (e.g., HTTP, SSH) and triggers alerts if thresholds breach. Nagios Core, Icinga
NetFlow/sFlow Tracks who is using bandwidth (e.g., a Daraz server under DDoS attack). SolarWinds, ManageEngine

How It Works:

  1. Agent (on devices) collects metrics (e.g., CPU usage).
  2. Manager (e.g., Nagios) polls agents via SNMP or reads syslog files.
  3. Alerts trigger if metrics exceed thresholds (e.g., "Switch Y’s temperature > 60°C").
sequenceDiagram
    participant Device as Router/Switch
    participant Agent as SNMP Agent
    participant Manager as Nagios Server
    participant Admin as Network Admin

    Device->>Agent: CPU=95% (SNMP Trap)
    Agent->>Manager: Forward trap to Nagios
    Manager->>Admin: Alert: "CPU Critical on Router123"
    Admin->>Device: SSH in, restart service

Real-World Example: NTC’s Network Oversight

  • Problem: Nepal’s fiber backbone (Nepal Fiber Company) faces frequent cuts due to landslides.
  • Solution: NTC uses SNMP + Nagios to monitor:
    • Link status (e.g., "Fiber segment Kathmandu-Pokhara down").
    • Traffic spikes (e.g., sudden 50% bandwidth drop during elections).
  • Action: Automated alerts trigger backup routes via BGP (Border Gateway Protocol) rerouting.

2. Backup Strategies: The 3-2-1 Rule in Action

The 3-2-1 backup rule is a golden standard:

  • 3 copies of data (original + 2 backups).
  • 2 different media (e.g., disk + tape).
  • 1 offsite (e.g., cloud or remote server).
AWS S3 (Offsite)2TBFull Backup (Weekly)Local RAID 5500GB (Mon)Local RAID 5500GB (Tue)Local RAID 5500GB (Wed)Differential Backups (Daily)Daraz Backup Strategy
Backup Type How It Works Pros Cons Example Use Case
Full Backup Copies all data every time. Fastest restore. Storage-intensive. Weekly backups of NEPSE’s daily trades.
Incremental Backs up only changes since last backup. Minimal storage. Slowest restore (needs all backups). Daily logs for Pathao’s ride data.
Differential Backs up changes since last full. Faster restore than incremental. Storage grows over time. Bank transaction backups (daily diffs).

Worked Example: Daraz’s Order Processing Backup

  • Scenario: Daraz processes 10,000 orders/day. A server crash risks losing unsent orders.
  • Strategy:
    • Full backup: Sunday night (2TB).
    • Differential backups: Mon–Sat nights (each ~500GB).
  • Restore Time:
    • If crash on Wednesday: Restore Sunday full + Wednesday diff (2.5TB) in 2 hours.
  • Offsite: Cloud (AWS S3) for disaster recovery.

3. RAID: Balancing Speed and Redundancy

RAID (Redundant Array of Independent Disks) combines multiple disks to improve performance or fault tolerance. Critical for servers handling high traffic (e.g., eSewa’s payment gateway).

RAID Level Description Redundancy? Performance Use Case
RAID 0 Striping (no redundancy). ❌ No ⚡ Very fast Temporary file storage.
RAID 1 Mirroring (exact copy). ✅ Yes 🐢 Slow Critical databases (banks).
RAID 5 Striping + parity (1 disk lost). ✅ Yes ⚡ Fast Web servers (Daraz).
RAID 6 Striping + dual parity (2 disks lost) ✅ Yes 🐢 Slow Long-term archives.

Worked Example: Ncell’s Call Log Backup

  • Problem: Ncell’s VoLTE system generates 1TB/day of call logs. A disk failure could lose billing records.
  • Solution: RAID 6 across 4 disks:
    • Striping: Distributes logs across disks for faster reads/writes.
    • Dual parity: Survives 2 disk failures (e.g., if Disk1 and Disk3 fail).
  • Result: No downtime during disk replacements.

4. Disaster Recovery: RTO vs. RPO

  • RTO (Recovery Time Objective): How fast must systems be back online?
    • Example: Ncell’s RTO = 30 minutes (backup cell towers activate automatically).
  • RPO (Recovery Point Objective): How much data loss is acceptable?
    • Example: eSewa’s RPO = 15 minutes (transaction logs backed every 15 mins).
stateDiagram-v2
    [*] --> Active
    Active --> Backup
    Backup --> Restore
    Restore --> Active
    
    state Active {
      [*] --> Normal
      Normal --> Failure
      Failure --> Alert
      Alert --> Backup
    }
    
    state Backup {
      [*] --> Cloud
      Cloud --> Local
      Local --> Restore
    }
    
    state Restore {
      [*] --> Partial
      Partial --> Full
      Full --> [*]
    }
    
    note right of Failure
      RTO = 30 mins (Ncell)
      RPO = 15 mins (eSewa)
    end
State transitions for Ncell’s disaster recovery with RTO/RPO constraints.
Scenario RTO RPO Strategy
Bank ATM failure 1 hour 5 minutes RAID 1 + daily incremental backups.
NTC fiber cut 2 hours 1 hour BGP rerouting + cloud backups.
Daraz server crash 30 minutes 10 minutes RAID 5 + hourly snapshots.

Disaster Recovery Plan (DRP) Steps:

  1. Identify critical systems (e.g., NEPSE’s trading platform).
  2. Classify data (e.g., real-time trades vs. historical reports).
  3. Test backups (e.g., restore a test server weekly).
  4. Document procedures (e.g., "If primary DC fails, failover to secondary DC in Chitwan").
Failover Trigger["Primary DC Down"]
    --> Check["Is Secondary DC Online?"]
        --> Yes["Failover to Secondary"]
        --> No["Activate Cloud DR Site"]
Failover Trigger --> Alert["Notify Team via Slack"]

5. Log Analysis: Hunting for Needles in Haystacks

Logs are digital breadcrumbs left by every transaction, error, or security event. Tools like grep, awk, or ELK Stack (Elasticsearch, Logstash, Kibana) turn raw logs into actionable insights.

Example: Pathao’s Fraud Detection

  • Log Sample:
    2023-10-15 14:30:23, API Call: /payments/process, UserID: 12345, Amount: $500, Status: FAILED
    2023-10-15 14:30:24, API Call: /payments/process, UserID: 12345, Amount: $500, Status: SUCCESS (from IP: 192.168.1.100)
    
  • Anomaly Detected:
    • Same user, same amount, but IP changed from 192.168.1.50 to 192.168.1.100 (likely a hacked account).
  • Action: Pathao’s system flags this for manual review and blocks the new IP.

Log Analysis Commands:

# Find failed SSH attempts in /var/log/auth.log
grep "Failed password" /var/log/auth.log | awk '{print $11}' | sort | uniq -c

# Count HTTP 500 errors in Apache logs
grep "500" /var/log/apache2/error.log | wc -l
LogSources["Apache/Nginx/Syslog"]
    --> Logstash["Parse & Filter Logs"]
        --> Elasticsearch["Index Logs"]
            --> Kibana["Visualize & Alert"]

6. Backup Verification: The Forgotten Step

80% of backups fail when tested. Always verify backups with:

  1. Integrity checks (e.g., md5sum on Linux).
  2. Restore tests (e.g., restore a backup to a test VM).
  3. Automated alerts (e.g., Nagios checks backup completion).

Example: NEPSE’s Market Data Backup

  • Process:
    1. Daily full backup of trade data to RAID 6 array.
    2. Nightly script runs restore to a test server.
    3. If restore fails, email alert to admins.
  • Result: Confirmed 100% recovery in 2022 after a disk failure.

In the Real World

  1. eSewa’s Payment Gateway

    • Idea Used: RAID 10 (mirrored striping) for transaction logs.
    • Why? Combines RAID 1’s redundancy with RAID 0’s speed to handle 50,000 transactions/minute during festivals like Dashain.
    • Backup: Incremental backups every 5 minutes to AWS Glacier (offsite).
  2. NTC’s Fiber Monitoring

    • Idea Used: SNMP + Nagios for real-time link status.
    • Why? Detects fiber cuts (e.g., 2021 Koshi landslide) and auto-reroutes traffic via BGP within 2 minutes.
    • Logs: Syslog centralizes alerts to NTC’s SOC (Security Operations Center).
  3. Pathao’s Ride Data

    • Idea Used: Differential backups + ELK Stack.
    • Why? Daily differential backups (500GB) of ride data, analyzed via Kibana to:
      • Detect fraud (sudden IP changes).
      • Optimize driver routes (traffic pattern logs).

Exam Tip

  1. Monitoring Questions:

    • Expect SNMP GET/SET or syslog format questions. Memorize:
      • SNMP community strings (public/private).
      • Syslog facility levels (e.g., auth, kern).
    • Trace Example: Draw a sequence diagram for Nagios checking a router’s CPU via SNMP.
  2. Backup Strategies:

    • Compare full vs. incremental vs. differential in a table (as above).
    • Worked Example: Given a scenario (e.g., "A bank with 5TB data"), calculate:
      • Storage needed for weekly full + daily incremental backups.
      • Restore time if 3 days of data are lost.
  3. RAID:

    • Must Know: RAID 0 (speed), RAID 1 (mirroring), RAID 5 (parity).
    • Exam Trick: If asked "Which RAID survives 2 disk failures?" → RAID 6.
  4. Disaster Recovery:

    • Define RTO/RPO and give real-world examples (e.g., Ncell’s 30-minute RTO).
    • Flowchart: Draw a failover process for a given scenario.
  5. Logs:

    • Practice: Given a log snippet, identify:
      • The source (e.g., Apache, SSH).
      • The anomaly (e.g., repeated failed logins).
    • Commands: Know grep, awk, and journalctl basics.

Visual Cheat Sheet for Exam:

mindmap
  root((Network Monitoring & Backup))
    Monitoring
      SNMP["GET/SET, Community Strings"]
      Syslog["Facilities: auth, kern, daemon"]
      Nagios["Active/Passive Checks"]
    Backup
      Types["Full/Incremental/Differential"]
      RAID["0-6, Redundancy vs. Speed"]
    DRP
      RTO/RPO["Time vs. Data Loss"]
      Plan["Failover Steps"]
    Logs
      Analysis["grep/awk, ELK Stack"]
      Verification["md5sum, Restore Tests"]

In the real world

  • NTC’s Fiber Monitoring: Uses SNMP + Nagios to track real-time fiber cuts (e.g., Kathmandu-Pokhara link) and triggers BGP rerouting within 2 hours (RTO) to restore service, ensuring minimal data loss (RPO = 1 hour).
  • eSewa’s Payment Gateway: Employs RAID 1 for critical transaction databases and hourly incremental backups to cloud storage (AWS) to meet RTO = 10 minutes and RPO = 5 minutes during peak hours.
  • Pathao’s Log Analysis: Deploys ELK Stack (Elasticsearch, Logstash, Kibana) to analyze 100GB/day of ride logs, flagging anomalies like fraudulent API calls (e.g., fake ride requests) in real-time using grep and awk scripts.

Based on the TU BITM syllabus for Networking and System Administration (IT271), unit 9.

Discussion

Loading…