IT271 Networking and System Administration

Networking and System AdministrationUnit 915 min read

Network Monitoring & Backup: Tools, Strategies & Recovery

Unit 9 of Networking and System Administration covers proactive network monitoring (SNMP, syslog, Nagios), backup strategies (full/incremental/differential), disaster recovery plans, and real-world tools like rsync, tar, and cloud backups. Learn how to diagnose failures, automate backups, and restore systems—with Nepal

TAKEAWAYS:

  • Monitoring vs. Backup: SNMP and syslog detect issues (e.g., server CPU spikes), while backups prevent data loss (e.g., restoring a crashed Daraz database).
  • Backup Types: Full backups copy everything; incremental backups only copy changes—choose based on RPO (Recovery Point Objective).
  • RAID Levels: RAID 1 mirrors data (faster reads, no redundancy loss), while RAID 5 stripes with parity (slower writes, fault-tolerant).
  • Disaster Recovery: The 3-2-1 rule (3 copies, 2 media types, 1 offsite) protects against ransomware (e.g., Ncell’s backup servers in Chitwan).
  • Tools in Action: cron + rsync automates daily backups; Nagios alerts admins to Pathao’s payment-gateway failures before users notice.
  • Exam Focus: Trace a backup/restore workflow (e.g., "How would you recover a deleted NEPSE trading file?"), and compare monitoring tools (Zabbix vs. PRTG).

1. Why Monitor and Backup?

Networks and systems fail. 93% of companies that lose data for 10+ days file for bankruptcy (University of Texas study). Monitoring prevents outages; backups recover from them.

Real-World Failures (and How Monitoring Helps)

Company/App Failure Scenario Monitoring Tool Used Outcome
NTC (Nepal) 2023 Kathmandu blackout (fiber cut) SNMP + Zabbix Identified the cut in 12 minutes; rerouted traffic via Pokhara.
eSewa Database corruption (unnoticed for 3 hours) Nagios + MySQL logs Auto-alerted admins; restored from 4-hour incremental backup.
Daraz (Nepal) Order queue crash (Black Friday) PRTG + custom scripts Detected stuck processes; rolled back to clean snapshot.

A real data center rack with redundant power (UPS) and cooling—critical for monitoring tools like Nagios to run 24/7.


2. Network Monitoring: Tools and Techniques

Monitoring collects metrics (CPU, bandwidth, errors) to predict failures before users complain.

A. Key Monitoring Protocols

Protocol Purpose Example Use Case Port
SNMP Queries devices (routers, switches) for stats (e.g., "How many packets dropped?"). Ncell monitors base stations for signal drops. 161/162
syslog Logs events (e.g., "User X failed login"). Banks log failed ATM transactions to detect fraud. 514
ICMP Ping tests connectivity (e.g., "Is Google reachable?"). Pathao checks if rider apps can hit their servers. 1/0

How SNMP Works (Simplified):

sequenceDiagram
    participant Manager as SNMP Manager (e.g., Zabbix)
    participant Agent as SNMP Agent (e.g., Router)
    Manager->>Agent: GET-REQUEST (e.g., "Interface errors?")
    Agent-->>Manager: GET-RESPONSE (e.g., "42 dropped packets")
    Note over Manager: Triggers alert if >50 errors/min
Tool Best For Nepali Example
Nagios Server uptime, service checks NTC uses it to monitor fiber routes.
Zabbix Custom dashboards (e.g., traffic heatmaps) Daraz tracks order-processing delays.
PRTG Bandwidth usage (e.g., "Is YouTube throttling?") Ncell caps data usage during peak hours.
Wireshark Packet-level debugging (e.g., "Why is my Khalti payment hanging?") Banks analyze TLS handshakes.

Worked Example: Monitoring a Kathmandu Traffic Route Problem: During Dashain, NTC’s fiber backbone between Thapathali and Balaju shows 90% utilization at 3 PM.

  1. Tool: PRTG monitors the link every 5 minutes.
  2. Alert: At 3:15 PM, PRTG detects >85% usage for 30+ minutes → triggers an email to NTC’s NOC.
  3. Action: NOC reroutes less critical traffic (e.g., video streaming) to a secondary link.
  4. Result: Usage drops to 60%; no outages.

A real PRTG dashboard overlaying NTC’s fiber map, highlighting the Thapathali-Balaju link in red (92% load).


3. Backup Strategies: Full, Incremental, Differential

Backups are classified by what they copy and how often.

A. Backup Types Compared

Type Copies Restore Time Storage Used Best For
Full All data every time Fastest Highest Critical data (e.g., NEPSE trading records)
Incremental Only changes since last backup Slowest Lowest Daily logs (e.g., eSewa transaction history)
Differential Changes since last full backup Medium Medium Weekly backups (e.g., Daraz product catalog)

Visual Workflow:

stateDiagram-v2
    [*] --> FullBackup: Day 1 (Full)
    FullBackup --> Incremental1: Day 2 (Incremental)
    Incremental1 --> Incremental2: Day 3 (Incremental)
    Incremental2 --> Differential1: Day 4 (Differential since Day 1)
    Differential1 --> Restore: "Restore: Full + Differential1"

B. Backup Tools and Commands

Tool/Command Purpose Example
tar Archive files (e.g., /var/www) tar -czvf backup.tar.gz /var/www
rsync Sync files to remote (e.g., cloud) rsync -avz /home/user/ backup@192.168.1.10:/backups/
cron Schedule backups (e.g., daily at 2 AM) 0 2 * * * /usr/bin/rsync ...
Veeam Enterprise backups (e.g., VMs) Used by Ncell for core network configs.

Worked Example: Backing Up eSewa’s Database Scenario: eSewa’s PostgreSQL database holds 100GB of transaction data. They need:

  • RPO: 1 hour (can lose up to 1 hour of data).
  • RTO: 4 hours (must restore within 4 hours).

Solution:

  1. Daily Full Backup at 2 AM (100GB) → stored on tape (offsite).
  2. Hourly Incremental backups (5GB/hour) → stored on cloud (AWS S3).
  3. Restore Process:
    • Latest full backup (2 AM) + incremental from 3 AM → 4-hour RTO.
    • If corruption is detected at 3:30 PM, restore from 2 AM full + 3 AM incremental.

A screenshot of an S3 bucket showing eSewa_db_full_20240501.tar and hourly incrementals like eSewa_db_incr_20240501_1500.tar.


4. RAID: Redundancy Without Backups

RAID (Redundant Array of Independent Disks) combines multiple disks for speed or fault tolerance. Critical for servers (e.g., NEPSE’s trading platform).

A. RAID Levels Explained

Level Description Fault Tolerance Performance Use Case
RAID 0 Striping (no redundancy) ❌ None ⚡ Fastest reads/writes Temporary storage (e.g., video editing)
RAID 1 Mirroring (exact copy) ✅ 1 disk failure 🏃 Moderate Critical data (e.g., bank logs)
RAID 5 Striping + parity (1 disk lost) ✅ 1 disk failure 🏃 Moderate Web servers (e.g., Daraz product DB)
RAID 10 RAID 1 + RAID 0 (mirrored stripes) ✅ 2 disk failures ⚡ Fastest High-availability (e.g., Ncell’s billing system)

Visual: RAID 1 vs. RAID 5

B. When to Use RAID vs. Backups

Scenario RAID Backup Why?
Single disk fails (e.g., laptop HDD) RAID 1 ✅ Yes RAID protects from hardware failure; backups protect from deletion/corruption.
Ransomware attack (e.g., Ncell’s servers) ❌ No ✅ Yes RAID doesn’t help if data is encrypted.
High I/O workload (e.g., Daraz’s order processing) RAID 10 ✅ Yes RAID 10 gives speed + redundancy.
A real photo of a server with 4 disks configured as RAID 10 (two mirrored pairs striped).

5. Disaster Recovery: Plans and Testing

A Disaster Recovery Plan (DRP) defines:

  1. RPO (Recovery Point Objective): Max data loss allowed (e.g., "1 hour of emails").
  2. RTO (Recovery Time Objective): Max downtime (e.g., "2 hours for NEPSE trading").
  3. Steps: Who does what? (e.g., "IT restores DB; PR team announces outage").

A. The 3-2-1 Backup Rule

  • 3 copies of data (e.g., primary DB + 2 backups).
  • 2 media types (e.g., disk + tape).
  • 1 offsite (e.g., cloud or Chitwan data center).

Example for a Nepali Bank:

  • Primary: Database on RAID 10 in Kathmandu.
  • Backup 1: Nightly incremental to AWS S3 (offsite).
  • Backup 2: Weekly full backup to tape (stored in Pokhara).

B. DRP Testing Methods

Method Description Example
Tabletop Exercise Discuss hypothetical scenarios (e.g., "Fire destroys server room"). NTC’s NOC team practices fiber-cut recovery.
Failover Test Force a switch to backup systems. eSewa simulates a primary DB crash.
Full Simulation Shut down primary; restore from backup. Daraz tests restoring order data from tape.

Worked Example: NTC’s Fiber Cut DRP

  1. RPO: 15 minutes (no more than 15 minutes of lost calls).
  2. RTO: 1 hour (restore service within 60 minutes).
  3. Steps:
    • 0:00–0:15: SNMP detects fiber cut; alerts NOC.
    • 0:15–0:30: NOC activates pre-configured backup route via satellite link.
    • 0:30–1:00: Restore voice traffic; reroute data via Pokhara node.
  4. Testing: Quarterly failover drills with satellite providers.

A map of Nepal’s fiber backbone showing the primary Kathmandu-Pokhara link and a dashed backup route via satellite.


6. Automating Backups with cron and Scripts

Manual backups fail. Automation ensures consistency.

A. cron Syntax

# Run at 2 AM daily, backup /var/www to remote server
0 2 * * * /usr/bin/rsync -avz /var/www/ user@backup-server:/backups/www/
  • 0 2 * * *: "At 2:00 AM, every day."
  • -avz: Archive mode (a), verbose (v), compress (z).

B. Script Example: Email Alerts on Backup Failure

#!/bin/bash
BACKUP_DIR="/backups/daraz_orders"
LOG_FILE="/var/log/backup.log"

# Run backup
rsync -avz /orders/ $BACKUP_DIR >> $LOG_FILE 2>&1

# Check exit status
if [ $? -ne 0 ]; then
    echo "Backup failed at $(date)" | mail -s "ALERT: Daraz Backup Failed" admin@daraz.com
fi

Why This Matters:

  • Daraz uses this to alert admins if the /orders/ directory isn’t backed up (e.g., due to disk full).
  • NEPSE uses similar scripts to verify trading data backups before market open.

7. Cloud vs. On-Premise Backups

Factor Cloud Backup (AWS S3, Google Drive) On-Premise (Tape, NAS)
Cost Pay-as-you-go (e.g., $0.023/GB/month) High upfront (e.g., tape drives)
Recovery Speed Slow (depends on internet) Fast (local storage)
Disaster Proof ✅ Yes (geographically distributed) ❌ No (localized risks)
Nepali Example eSewa uses AWS for offsite backups. NTC uses tapes for critical configs.

Worked Example: Khalti’s Cloud Backup

  • Primary: PostgreSQL DB on Kathmandu servers (RAID 10).
  • Backup: Hourly snapshots to AWS RDS (automated via pg_dump).
  • Restore: If Kathmandu floods, Khalti spins up a new DB in AWS Mumbai within 2 hours.

Exam Tip

How This Unit is Tested:

  1. Scenario-Based Questions (30%):

    • "NEPSE’s trading server crashes. The last full backup was yesterday at 2 AM, and incremental backups run every hour. If the crash happened at 4 PM, how would you restore?"
    • Answer: Full backup (2 AM) + incrementals from 3 AM and 4 AM.
  2. Tool Configuration (25%):

    • "Write a cron job to backup /var/log/ to a remote server daily at 3 AM."
    • Answer:
      0 3 * * * rsync -avz /var/log/ user@backup-server:/logs/
      
  3. RAID/Monitoring Comparisons (20%):

    • "Compare RAID 1 and RAID 5 for a Daraz order-processing server. Which would you choose and why?"
    • Answer:
      Criteria RAID 1 RAID 5
      Fault Tolerance 1 disk failure 1 disk failure
      Write Speed Slower (must write to 2 disks) Faster (parity calculated on-the-fly)
      Cost 2 disks for same space 3 disks for same space
      Best For Critical data (e.g., bank logs) High I/O (e.g., Daraz orders)
  4. DRP Diagrams (15%):

    • "Draw a state diagram for restoring a failed eSewa database using the 3-2-1 rule."
    • Answer:
      stateDiagram-v2
          [*] --> DetectFailure: "SNMP alerts DB down"
          DetectFailure --> CheckBackups: "Verify 3-2-1 rule"
          CheckBackups --> RestoreFull: "Restore latest full (2 AM)"
          RestoreFull --> ApplyIncrementals: "Apply 3 AM & 4 AM incrementals"
          ApplyIncrementals --> Verify: "Check data integrity"
          Verify --> [*]: "Service restored"
  5. Real-World Applications (10%):

    • "How does Pathao use monitoring to prevent rider app crashes during Diwali?"
    • Answer:
      • Tool: PRTG monitors API response times.
      • Alert: If >500ms latency for 5 minutes → auto-scales servers.
      • Backup: Daily snapshots of rider data to AWS.

Common Pitfalls:

  • Forgetting parity in RAID 5 (it’s what allows reconstruction).
  • Confusing incremental vs. differential backups (exam loves this!).
  • Ignoring offsite backups in DRP questions (always mention 3-2-1).

Final Checklist for Full Marks: ✅ Draw a state diagram for backup/restore. ✅ Compare RAID levels in a table. ✅ Write a cron job or rsync command. ✅ Explain one Nepali company’s monitoring/backup strategy (e.g., NTC’s fiber routes). ✅ Define RPO/RTO and give numbers (e.g., "NEPSE’s RPO is 15 minutes").

Based on the TU BIM syllabus for Networking and System Administration (IT271), unit 9.

Discussion

Loading…