Networking and System AdministrationUnit 913 min read
Network Monitoring & Backup: Tools, Strategies & Recovery
Unit 9 of Networking and System Administration covers proactive network monitoring (SNMP, syslog, Nagios), backup strategies (full/incremental/differential), disaster recovery plans, and real-world implementations in Nepali IT infrastructure like NTC’s network oversight and bank transaction logs.
TAKEAWAYS:
- Network monitoring uses SNMP, syslog, and Nagios to track device health, traffic, and errors in real time—critical for troubleshooting outages like NTC’s fiber cuts.
- Backup types (full, incremental, differential) differ in storage needs and recovery speed; banks use differential backups for daily transaction logs to balance speed and space.
- RAID levels (0–6) trade redundancy for performance—RAID 1 mirrors data (used in Daraz’s servers) while RAID 5 stripes with parity (common in small offices).
- Disaster recovery plans prioritize RTO (recovery time) and RPO (data loss tolerance); Ncell’s backup cell towers activate within 30 minutes (RTO) to minimize downtime.
- Log analysis (via
grep,awk, or ELK stack) uncovers security breaches—e.g., Pathao’s fraud detection flags unusual API calls in real-time logs. - Backup verification (e.g.,
restoretests) is as critical as the backup itself; NEPSE’s daily market data backups are validated nightly to ensure no corruption.
1. Network Monitoring: Watching the Digital Nervous System
Network monitoring is the proactive observation of network devices (routers, switches, servers) and traffic to detect anomalies, performance bottlenecks, or security threats before they disrupt services. Think of it as a doctor’s ECG for your network: continuous, data-driven, and alerting you to irregularities.
sequenceDiagram
participant Router as NTC Router
participant SNMP as SNMP Agent
participant Nagios as Nagios Server
participant Admin as Network Admin
Router->>SNMP: CPU=98% (SNMP Trap)
SNMP->>Nagios: Forward to Nagios
Nagios->>Admin: Alert: "Router123 CPU Critical"
Admin->>Router: SSH Command: "restart service"
Router-->>Admin: Service Restarted
Router-->>SNMP: CPU=45% (SNMP Update)
SNMP-->>Nagios: Update Nagios DashboardSNMP trap handling workflow for NTC’s router monitoring system.Key Components
| Component | Role | Example Tools/Protocols |
|---|---|---|
| SNMP (Simple Network Management Protocol) | Polls devices for stats (CPU, memory, interface errors) via GET/SET messages. | snmpwalk, Zabbix, PRTG |
| Syslog | Centralized logging of events (e.g., "Router X dropped 500 packets"). | rsyslog, Graylog, Splunk |
| Nagios | Monitors services (e.g., HTTP, SSH) and triggers alerts if thresholds breach. | Nagios Core, Icinga |
| NetFlow/sFlow | Tracks who is using bandwidth (e.g., a Daraz server under DDoS attack). | SolarWinds, ManageEngine |
How It Works:
- Agent (on devices) collects metrics (e.g., CPU usage).
- Manager (e.g., Nagios) polls agents via SNMP or reads syslog files.
- Alerts trigger if metrics exceed thresholds (e.g., "Switch Y’s temperature > 60°C").
sequenceDiagram
participant Device as Router/Switch
participant Agent as SNMP Agent
participant Manager as Nagios Server
participant Admin as Network Admin
Device->>Agent: CPU=95% (SNMP Trap)
Agent->>Manager: Forward trap to Nagios
Manager->>Admin: Alert: "CPU Critical on Router123"
Admin->>Device: SSH in, restart serviceReal-World Example: NTC’s Network Oversight
- Problem: Nepal’s fiber backbone (Nepal Fiber Company) faces frequent cuts due to landslides.
- Solution: NTC uses SNMP + Nagios to monitor:
- Link status (e.g., "Fiber segment Kathmandu-Pokhara down").
- Traffic spikes (e.g., sudden 50% bandwidth drop during elections).
- Action: Automated alerts trigger backup routes via BGP (Border Gateway Protocol) rerouting.
2. Backup Strategies: The 3-2-1 Rule in Action
The 3-2-1 backup rule is a golden standard:
- 3 copies of data (original + 2 backups).
- 2 different media (e.g., disk + tape).
- 1 offsite (e.g., cloud or remote server).
| Backup Type | How It Works | Pros | Cons | Example Use Case |
|---|---|---|---|---|
| Full Backup | Copies all data every time. | Fastest restore. | Storage-intensive. | Weekly backups of NEPSE’s daily trades. |
| Incremental | Backs up only changes since last backup. | Minimal storage. | Slowest restore (needs all backups). | Daily logs for Pathao’s ride data. |
| Differential | Backs up changes since last full. | Faster restore than incremental. | Storage grows over time. | Bank transaction backups (daily diffs). |
Worked Example: Daraz’s Order Processing Backup
- Scenario: Daraz processes 10,000 orders/day. A server crash risks losing unsent orders.
- Strategy:
- Full backup: Sunday night (2TB).
- Differential backups: Mon–Sat nights (each ~500GB).
- Restore Time:
- If crash on Wednesday: Restore Sunday full + Wednesday diff (2.5TB) in 2 hours.
- Offsite: Cloud (AWS S3) for disaster recovery.
3. RAID: Balancing Speed and Redundancy
RAID (Redundant Array of Independent Disks) combines multiple disks to improve performance or fault tolerance. Critical for servers handling high traffic (e.g., eSewa’s payment gateway).
| RAID Level | Description | Redundancy? | Performance | Use Case |
|---|---|---|---|---|
| RAID 0 | Striping (no redundancy). | ❌ No | ⚡ Very fast | Temporary file storage. |
| RAID 1 | Mirroring (exact copy). | ✅ Yes | 🐢 Slow | Critical databases (banks). |
| RAID 5 | Striping + parity (1 disk lost). | ✅ Yes | ⚡ Fast | Web servers (Daraz). |
| RAID 6 | Striping + dual parity (2 disks lost) | ✅ Yes | 🐢 Slow | Long-term archives. |
Worked Example: Ncell’s Call Log Backup
- Problem: Ncell’s VoLTE system generates 1TB/day of call logs. A disk failure could lose billing records.
- Solution: RAID 6 across 4 disks:
- Striping: Distributes logs across disks for faster reads/writes.
- Dual parity: Survives 2 disk failures (e.g., if Disk1 and Disk3 fail).
- Result: No downtime during disk replacements.
4. Disaster Recovery: RTO vs. RPO
- RTO (Recovery Time Objective): How fast must systems be back online?
- Example: Ncell’s RTO = 30 minutes (backup cell towers activate automatically).
- RPO (Recovery Point Objective): How much data loss is acceptable?
- Example: eSewa’s RPO = 15 minutes (transaction logs backed every 15 mins).
stateDiagram-v2
[*] --> Active
Active --> Backup
Backup --> Restore
Restore --> Active
state Active {
[*] --> Normal
Normal --> Failure
Failure --> Alert
Alert --> Backup
}
state Backup {
[*] --> Cloud
Cloud --> Local
Local --> Restore
}
state Restore {
[*] --> Partial
Partial --> Full
Full --> [*]
}
note right of Failure
RTO = 30 mins (Ncell)
RPO = 15 mins (eSewa)
endState transitions for Ncell’s disaster recovery with RTO/RPO constraints.| Scenario | RTO | RPO | Strategy |
|---|---|---|---|
| Bank ATM failure | 1 hour | 5 minutes | RAID 1 + daily incremental backups. |
| NTC fiber cut | 2 hours | 1 hour | BGP rerouting + cloud backups. |
| Daraz server crash | 30 minutes | 10 minutes | RAID 5 + hourly snapshots. |
Disaster Recovery Plan (DRP) Steps:
- Identify critical systems (e.g., NEPSE’s trading platform).
- Classify data (e.g., real-time trades vs. historical reports).
- Test backups (e.g., restore a test server weekly).
- Document procedures (e.g., "If primary DC fails, failover to secondary DC in Chitwan").
Failover Trigger["Primary DC Down"]
--> Check["Is Secondary DC Online?"]
--> Yes["Failover to Secondary"]
--> No["Activate Cloud DR Site"]
Failover Trigger --> Alert["Notify Team via Slack"]
5. Log Analysis: Hunting for Needles in Haystacks
Logs are digital breadcrumbs left by every transaction, error, or security event. Tools like grep, awk, or ELK Stack (Elasticsearch, Logstash, Kibana) turn raw logs into actionable insights.
Example: Pathao’s Fraud Detection
- Log Sample:
2023-10-15 14:30:23, API Call: /payments/process, UserID: 12345, Amount: $500, Status: FAILED 2023-10-15 14:30:24, API Call: /payments/process, UserID: 12345, Amount: $500, Status: SUCCESS (from IP: 192.168.1.100) - Anomaly Detected:
- Same user, same amount, but IP changed from 192.168.1.50 to 192.168.1.100 (likely a hacked account).
- Action: Pathao’s system flags this for manual review and blocks the new IP.
Log Analysis Commands:
# Find failed SSH attempts in /var/log/auth.log
grep "Failed password" /var/log/auth.log | awk '{print $11}' | sort | uniq -c
# Count HTTP 500 errors in Apache logs
grep "500" /var/log/apache2/error.log | wc -l
LogSources["Apache/Nginx/Syslog"]
--> Logstash["Parse & Filter Logs"]
--> Elasticsearch["Index Logs"]
--> Kibana["Visualize & Alert"]
6. Backup Verification: The Forgotten Step
80% of backups fail when tested. Always verify backups with:
- Integrity checks (e.g.,
md5sumon Linux). - Restore tests (e.g., restore a backup to a test VM).
- Automated alerts (e.g., Nagios checks backup completion).
Example: NEPSE’s Market Data Backup
- Process:
- Daily full backup of trade data to RAID 6 array.
- Nightly script runs
restoreto a test server. - If restore fails, email alert to admins.
- Result: Confirmed 100% recovery in 2022 after a disk failure.
In the Real World
eSewa’s Payment Gateway
- Idea Used: RAID 10 (mirrored striping) for transaction logs.
- Why? Combines RAID 1’s redundancy with RAID 0’s speed to handle 50,000 transactions/minute during festivals like Dashain.
- Backup: Incremental backups every 5 minutes to AWS Glacier (offsite).
NTC’s Fiber Monitoring
- Idea Used: SNMP + Nagios for real-time link status.
- Why? Detects fiber cuts (e.g., 2021 Koshi landslide) and auto-reroutes traffic via BGP within 2 minutes.
- Logs: Syslog centralizes alerts to NTC’s SOC (Security Operations Center).
Pathao’s Ride Data
- Idea Used: Differential backups + ELK Stack.
- Why? Daily differential backups (500GB) of ride data, analyzed via Kibana to:
- Detect fraud (sudden IP changes).
- Optimize driver routes (traffic pattern logs).
Exam Tip
Monitoring Questions:
- Expect SNMP GET/SET or syslog format questions. Memorize:
- SNMP community strings (
public/private). - Syslog facility levels (e.g.,
auth,kern).
- SNMP community strings (
- Trace Example: Draw a sequence diagram for Nagios checking a router’s CPU via SNMP.
- Expect SNMP GET/SET or syslog format questions. Memorize:
Backup Strategies:
- Compare full vs. incremental vs. differential in a table (as above).
- Worked Example: Given a scenario (e.g., "A bank with 5TB data"), calculate:
- Storage needed for weekly full + daily incremental backups.
- Restore time if 3 days of data are lost.
RAID:
- Must Know: RAID 0 (speed), RAID 1 (mirroring), RAID 5 (parity).
- Exam Trick: If asked "Which RAID survives 2 disk failures?" → RAID 6.
Disaster Recovery:
- Define RTO/RPO and give real-world examples (e.g., Ncell’s 30-minute RTO).
- Flowchart: Draw a failover process for a given scenario.
Logs:
- Practice: Given a log snippet, identify:
- The source (e.g., Apache, SSH).
- The anomaly (e.g., repeated failed logins).
- Commands: Know
grep,awk, andjournalctlbasics.
- Practice: Given a log snippet, identify:
Visual Cheat Sheet for Exam:
mindmap
root((Network Monitoring & Backup))
Monitoring
SNMP["GET/SET, Community Strings"]
Syslog["Facilities: auth, kern, daemon"]
Nagios["Active/Passive Checks"]
Backup
Types["Full/Incremental/Differential"]
RAID["0-6, Redundancy vs. Speed"]
DRP
RTO/RPO["Time vs. Data Loss"]
Plan["Failover Steps"]
Logs
Analysis["grep/awk, ELK Stack"]
Verification["md5sum, Restore Tests"]In the real world
- NTC’s Fiber Monitoring: Uses SNMP + Nagios to track real-time fiber cuts (e.g., Kathmandu-Pokhara link) and triggers BGP rerouting within 2 hours (RTO) to restore service, ensuring minimal data loss (RPO = 1 hour).
- eSewa’s Payment Gateway: Employs RAID 1 for critical transaction databases and hourly incremental backups to cloud storage (AWS) to meet RTO = 10 minutes and RPO = 5 minutes during peak hours.
- Pathao’s Log Analysis: Deploys ELK Stack (Elasticsearch, Logstash, Kibana) to analyze 100GB/day of ride logs, flagging anomalies like fraudulent API calls (e.g., fake ride requests) in real-time using
grepandawkscripts.
Based on the TU BITM syllabus for Networking and System Administration (IT271), unit 9.
Discussion
Loading…