Networking and System AdministrationUnit 915 min read
Network Monitoring & Backup: Tools, Strategies & Recovery
Unit 9 of Networking and System Administration covers proactive network monitoring (SNMP, syslog, Nagios), backup strategies (full/incremental/differential), disaster recovery plans, and real-world tools like rsync, tar, and cloud backups. Learn how to diagnose failures, automate backups, and restore systems—with Nepal
TAKEAWAYS:
- Monitoring vs. Backup: SNMP and syslog detect issues (e.g., server CPU spikes), while backups prevent data loss (e.g., restoring a crashed Daraz database).
- Backup Types: Full backups copy everything; incremental backups only copy changes—choose based on RPO (Recovery Point Objective).
- RAID Levels: RAID 1 mirrors data (faster reads, no redundancy loss), while RAID 5 stripes with parity (slower writes, fault-tolerant).
- Disaster Recovery: The 3-2-1 rule (3 copies, 2 media types, 1 offsite) protects against ransomware (e.g., Ncell’s backup servers in Chitwan).
- Tools in Action:
cron+rsyncautomates daily backups; Nagios alerts admins to Pathao’s payment-gateway failures before users notice. - Exam Focus: Trace a backup/restore workflow (e.g., "How would you recover a deleted NEPSE trading file?"), and compare monitoring tools (Zabbix vs. PRTG).
1. Why Monitor and Backup?
Networks and systems fail. 93% of companies that lose data for 10+ days file for bankruptcy (University of Texas study). Monitoring prevents outages; backups recover from them.
Real-World Failures (and How Monitoring Helps)
| Company/App | Failure Scenario | Monitoring Tool Used | Outcome |
|---|---|---|---|
| NTC (Nepal) | 2023 Kathmandu blackout (fiber cut) | SNMP + Zabbix | Identified the cut in 12 minutes; rerouted traffic via Pokhara. |
| eSewa | Database corruption (unnoticed for 3 hours) | Nagios + MySQL logs | Auto-alerted admins; restored from 4-hour incremental backup. |
| Daraz (Nepal) | Order queue crash (Black Friday) | PRTG + custom scripts | Detected stuck processes; rolled back to clean snapshot. |
A real data center rack with redundant power (UPS) and cooling—critical for monitoring tools like Nagios to run 24/7.
2. Network Monitoring: Tools and Techniques
Monitoring collects metrics (CPU, bandwidth, errors) to predict failures before users complain.
A. Key Monitoring Protocols
| Protocol | Purpose | Example Use Case | Port |
|---|---|---|---|
| SNMP | Queries devices (routers, switches) for stats (e.g., "How many packets dropped?"). | Ncell monitors base stations for signal drops. | 161/162 |
| syslog | Logs events (e.g., "User X failed login"). | Banks log failed ATM transactions to detect fraud. | 514 |
| ICMP | Ping tests connectivity (e.g., "Is Google reachable?"). | Pathao checks if rider apps can hit their servers. | 1/0 |
How SNMP Works (Simplified):
sequenceDiagram
participant Manager as SNMP Manager (e.g., Zabbix)
participant Agent as SNMP Agent (e.g., Router)
Manager->>Agent: GET-REQUEST (e.g., "Interface errors?")
Agent-->>Manager: GET-RESPONSE (e.g., "42 dropped packets")
Note over Manager: Triggers alert if >50 errors/minB. Popular Monitoring Tools
| Tool | Best For | Nepali Example |
|---|---|---|
| Nagios | Server uptime, service checks | NTC uses it to monitor fiber routes. |
| Zabbix | Custom dashboards (e.g., traffic heatmaps) | Daraz tracks order-processing delays. |
| PRTG | Bandwidth usage (e.g., "Is YouTube throttling?") | Ncell caps data usage during peak hours. |
| Wireshark | Packet-level debugging (e.g., "Why is my Khalti payment hanging?") | Banks analyze TLS handshakes. |
Worked Example: Monitoring a Kathmandu Traffic Route Problem: During Dashain, NTC’s fiber backbone between Thapathali and Balaju shows 90% utilization at 3 PM.
- Tool: PRTG monitors the link every 5 minutes.
- Alert: At 3:15 PM, PRTG detects >85% usage for 30+ minutes → triggers an email to NTC’s NOC.
- Action: NOC reroutes less critical traffic (e.g., video streaming) to a secondary link.
- Result: Usage drops to 60%; no outages.
A real PRTG dashboard overlaying NTC’s fiber map, highlighting the Thapathali-Balaju link in red (92% load).
3. Backup Strategies: Full, Incremental, Differential
Backups are classified by what they copy and how often.
A. Backup Types Compared
| Type | Copies | Restore Time | Storage Used | Best For |
|---|---|---|---|---|
| Full | All data every time | Fastest | Highest | Critical data (e.g., NEPSE trading records) |
| Incremental | Only changes since last backup | Slowest | Lowest | Daily logs (e.g., eSewa transaction history) |
| Differential | Changes since last full backup | Medium | Medium | Weekly backups (e.g., Daraz product catalog) |
Visual Workflow:
stateDiagram-v2
[*] --> FullBackup: Day 1 (Full)
FullBackup --> Incremental1: Day 2 (Incremental)
Incremental1 --> Incremental2: Day 3 (Incremental)
Incremental2 --> Differential1: Day 4 (Differential since Day 1)
Differential1 --> Restore: "Restore: Full + Differential1"B. Backup Tools and Commands
| Tool/Command | Purpose | Example |
|---|---|---|
tar |
Archive files (e.g., /var/www) |
tar -czvf backup.tar.gz /var/www |
rsync |
Sync files to remote (e.g., cloud) | rsync -avz /home/user/ backup@192.168.1.10:/backups/ |
cron |
Schedule backups (e.g., daily at 2 AM) | 0 2 * * * /usr/bin/rsync ... |
| Veeam | Enterprise backups (e.g., VMs) | Used by Ncell for core network configs. |
Worked Example: Backing Up eSewa’s Database Scenario: eSewa’s PostgreSQL database holds 100GB of transaction data. They need:
- RPO: 1 hour (can lose up to 1 hour of data).
- RTO: 4 hours (must restore within 4 hours).
Solution:
- Daily Full Backup at 2 AM (100GB) → stored on tape (offsite).
- Hourly Incremental backups (5GB/hour) → stored on cloud (AWS S3).
- Restore Process:
- Latest full backup (2 AM) + incremental from 3 AM → 4-hour RTO.
- If corruption is detected at 3:30 PM, restore from 2 AM full + 3 AM incremental.
A screenshot of an S3 bucket showing eSewa_db_full_20240501.tar and hourly incrementals like eSewa_db_incr_20240501_1500.tar.
4. RAID: Redundancy Without Backups
RAID (Redundant Array of Independent Disks) combines multiple disks for speed or fault tolerance. Critical for servers (e.g., NEPSE’s trading platform).
A. RAID Levels Explained
| Level | Description | Fault Tolerance | Performance | Use Case |
|---|---|---|---|---|
| RAID 0 | Striping (no redundancy) | ❌ None | ⚡ Fastest reads/writes | Temporary storage (e.g., video editing) |
| RAID 1 | Mirroring (exact copy) | ✅ 1 disk failure | 🏃 Moderate | Critical data (e.g., bank logs) |
| RAID 5 | Striping + parity (1 disk lost) | ✅ 1 disk failure | 🏃 Moderate | Web servers (e.g., Daraz product DB) |
| RAID 10 | RAID 1 + RAID 0 (mirrored stripes) | ✅ 2 disk failures | ⚡ Fastest | High-availability (e.g., Ncell’s billing system) |
Visual: RAID 1 vs. RAID 5
B. When to Use RAID vs. Backups
| Scenario | RAID | Backup | Why? |
|---|---|---|---|
| Single disk fails (e.g., laptop HDD) | RAID 1 | ✅ Yes | RAID protects from hardware failure; backups protect from deletion/corruption. |
| Ransomware attack (e.g., Ncell’s servers) | ❌ No | ✅ Yes | RAID doesn’t help if data is encrypted. |
| High I/O workload (e.g., Daraz’s order processing) | RAID 10 | ✅ Yes | RAID 10 gives speed + redundancy. |
5. Disaster Recovery: Plans and Testing
A Disaster Recovery Plan (DRP) defines:
- RPO (Recovery Point Objective): Max data loss allowed (e.g., "1 hour of emails").
- RTO (Recovery Time Objective): Max downtime (e.g., "2 hours for NEPSE trading").
- Steps: Who does what? (e.g., "IT restores DB; PR team announces outage").
A. The 3-2-1 Backup Rule
- 3 copies of data (e.g., primary DB + 2 backups).
- 2 media types (e.g., disk + tape).
- 1 offsite (e.g., cloud or Chitwan data center).
Example for a Nepali Bank:
- Primary: Database on RAID 10 in Kathmandu.
- Backup 1: Nightly incremental to AWS S3 (offsite).
- Backup 2: Weekly full backup to tape (stored in Pokhara).
B. DRP Testing Methods
| Method | Description | Example |
|---|---|---|
| Tabletop Exercise | Discuss hypothetical scenarios (e.g., "Fire destroys server room"). | NTC’s NOC team practices fiber-cut recovery. |
| Failover Test | Force a switch to backup systems. | eSewa simulates a primary DB crash. |
| Full Simulation | Shut down primary; restore from backup. | Daraz tests restoring order data from tape. |
Worked Example: NTC’s Fiber Cut DRP
- RPO: 15 minutes (no more than 15 minutes of lost calls).
- RTO: 1 hour (restore service within 60 minutes).
- Steps:
- 0:00–0:15: SNMP detects fiber cut; alerts NOC.
- 0:15–0:30: NOC activates pre-configured backup route via satellite link.
- 0:30–1:00: Restore voice traffic; reroute data via Pokhara node.
- Testing: Quarterly failover drills with satellite providers.
A map of Nepal’s fiber backbone showing the primary Kathmandu-Pokhara link and a dashed backup route via satellite.
6. Automating Backups with cron and Scripts
Manual backups fail. Automation ensures consistency.
A. cron Syntax
# Run at 2 AM daily, backup /var/www to remote server
0 2 * * * /usr/bin/rsync -avz /var/www/ user@backup-server:/backups/www/
0 2 * * *: "At 2:00 AM, every day."-avz: Archive mode (a), verbose (v), compress (z).
B. Script Example: Email Alerts on Backup Failure
#!/bin/bash
BACKUP_DIR="/backups/daraz_orders"
LOG_FILE="/var/log/backup.log"
# Run backup
rsync -avz /orders/ $BACKUP_DIR >> $LOG_FILE 2>&1
# Check exit status
if [ $? -ne 0 ]; then
echo "Backup failed at $(date)" | mail -s "ALERT: Daraz Backup Failed" admin@daraz.com
fi
Why This Matters:
- Daraz uses this to alert admins if the
/orders/directory isn’t backed up (e.g., due to disk full). - NEPSE uses similar scripts to verify trading data backups before market open.
7. Cloud vs. On-Premise Backups
| Factor | Cloud Backup (AWS S3, Google Drive) | On-Premise (Tape, NAS) |
|---|---|---|
| Cost | Pay-as-you-go (e.g., $0.023/GB/month) | High upfront (e.g., tape drives) |
| Recovery Speed | Slow (depends on internet) | Fast (local storage) |
| Disaster Proof | ✅ Yes (geographically distributed) | ❌ No (localized risks) |
| Nepali Example | eSewa uses AWS for offsite backups. | NTC uses tapes for critical configs. |
Worked Example: Khalti’s Cloud Backup
- Primary: PostgreSQL DB on Kathmandu servers (RAID 10).
- Backup: Hourly snapshots to AWS RDS (automated via
pg_dump). - Restore: If Kathmandu floods, Khalti spins up a new DB in AWS Mumbai within 2 hours.
Exam Tip
How This Unit is Tested:
Scenario-Based Questions (30%):
- "NEPSE’s trading server crashes. The last full backup was yesterday at 2 AM, and incremental backups run every hour. If the crash happened at 4 PM, how would you restore?"
- Answer: Full backup (2 AM) + incrementals from 3 AM and 4 AM.
Tool Configuration (25%):
- "Write a
cronjob to backup/var/log/to a remote server daily at 3 AM." - Answer:
0 3 * * * rsync -avz /var/log/ user@backup-server:/logs/
- "Write a
RAID/Monitoring Comparisons (20%):
- "Compare RAID 1 and RAID 5 for a Daraz order-processing server. Which would you choose and why?"
- Answer:
Criteria RAID 1 RAID 5 Fault Tolerance 1 disk failure 1 disk failure Write Speed Slower (must write to 2 disks) Faster (parity calculated on-the-fly) Cost 2 disks for same space 3 disks for same space Best For Critical data (e.g., bank logs) High I/O (e.g., Daraz orders)
DRP Diagrams (15%):
- "Draw a state diagram for restoring a failed eSewa database using the 3-2-1 rule."
- Answer:
stateDiagram-v2 [*] --> DetectFailure: "SNMP alerts DB down" DetectFailure --> CheckBackups: "Verify 3-2-1 rule" CheckBackups --> RestoreFull: "Restore latest full (2 AM)" RestoreFull --> ApplyIncrementals: "Apply 3 AM & 4 AM incrementals" ApplyIncrementals --> Verify: "Check data integrity" Verify --> [*]: "Service restored"
Real-World Applications (10%):
- "How does Pathao use monitoring to prevent rider app crashes during Diwali?"
- Answer:
- Tool: PRTG monitors API response times.
- Alert: If >500ms latency for 5 minutes → auto-scales servers.
- Backup: Daily snapshots of rider data to AWS.
Common Pitfalls:
- Forgetting parity in RAID 5 (it’s what allows reconstruction).
- Confusing incremental vs. differential backups (exam loves this!).
- Ignoring offsite backups in DRP questions (always mention 3-2-1).
Final Checklist for Full Marks:
✅ Draw a state diagram for backup/restore.
✅ Compare RAID levels in a table.
✅ Write a cron job or rsync command.
✅ Explain one Nepali company’s monitoring/backup strategy (e.g., NTC’s fiber routes).
✅ Define RPO/RTO and give numbers (e.g., "NEPSE’s RPO is 15 minutes").
Based on the TU BIM syllabus for Networking and System Administration (IT271), unit 9.
Discussion
Loading…