CACS406 Network Administration

Network AdministrationUnit 611 min read

Server Hardware Strategies & Management: Types, Procurement, Risk & Maintenance

Unit 6 of Network Administration covers server hardware selection, procurement strategies, risk management, and maintenance techniques for enterprise networks, with real-world examples from Nepali tech companies like Ncell and NEPSE.

TAKEAWAYS:

  • Server types (web, file, mail, database, proxy, etc.) are specialized for specific roles in enterprise networks, each requiring distinct hardware configurations.
  • Hardware strategies include redundancy (RAID, hot-swappable components), scalability (blade servers, modular designs), and energy efficiency (PUE, virtualization).
  • Procurement planning follows a structured approach: needs assessment → vendor evaluation → budgeting → risk mitigation (e.g., SLAs, warranties).
  • Risk management in server hardware involves physical threats (fire, theft), hardware failures (disk crashes, overheating), and supply chain risks (vendor reliability).
  • Maintenance includes preventive (regular cleaning, firmware updates) and corrective (RMA, replacement) actions, with documentation as a critical component.
  • Cost-benefit analysis balances upfront hardware costs against long-term TCO (Total Cost of Ownership), including energy, cooling, and downtime risks.

1. Types of Servers and Their Hardware Requirements

Servers are specialized computers designed to handle specific tasks in a network. Their hardware must align with their role to ensure performance, reliability, and security.

Common Server Types and Hardware Needs

WebServerDatabaseServerFileServerMailServerProxyServerServerType
Hierarchy of server types with key hardware requirements

Worked Example: NEPSE’s Trading Server Hardware

Nepal Stock Exchange (NEPSE) requires ultra-low latency and high availability for its trading platform. Their server hardware includes:

  • Dual Intel Xeon Platinum CPUs (for real-time order matching).
  • 1TB+ RAM (to handle thousands of concurrent trades).
  • RAID 10 storage (for transaction logs and audit trails).
  • Redundant power supplies (2x PSUs) and hot-swappable fans.
  • 10Gbps NICs (to connect to exchange participants).

Why?

  • Latency: Multi-core CPUs and NVMe SSDs reduce delay in order execution.
  • Reliability: RAID 10 ensures no data loss if a disk fails.
  • Scalability: Additional NICs allow future expansion to 40Gbps.

2. Hardware Strategies for Servers

Hardware strategies ensure servers meet performance, reliability, and cost-efficiency goals.

A. Redundancy and Fault Tolerance

Redundancy prevents single points of failure. Key techniques:

  • RAID Levels (for storage redundancy):

  • Hot-Swappable Components:

    • Power Supplies (PSUs): Replaceable without shutdown.
    • Fans: Prevent overheating (critical in data centers).
    • Hard Drives: RAID arrays allow drive replacement while running.
  • Uninterruptible Power Supply (UPS):

    • Provides backup power during outages (e.g., Ncell’s data centers use UPS to avoid service disruption during load shedding).

B. Scalability Strategies

Servers must grow with demand. Common approaches:

  1. Vertical Scaling (Scaling Up):
    • Adding more CPU cores, RAM, or faster storage to a single server.
    • Example: Upgrading a Daraz order-processing server from 32GB RAM to 128GB during Diwali sales.
  2. Horizontal Scaling (Scaling Out):
    • Adding more servers (e.g., load balancers distributing traffic across multiple web servers).
    • Example: Khalti’s payment gateway uses multiple servers to handle peak transactions during Dashain.
Server 1Server 2Load BalancerClient
Load balancing architecture for horizontal scalability
  1. Modular Servers (Blade Servers):
    • Definition: Servers with removable blades (e.g., HP BladeSystem).
    • Advantages:
      • Shared power/cooling reduces costs.
      • Easy to add/remove blades.
    • Disadvantage: Single point of failure if the chassis fails.

C. Energy Efficiency and Cooling

Data centers consume ~1-2% of global electricity. Strategies to reduce costs:

  • Power Usage Effectiveness (PUE):
    • Formula: PUE = Total Facility Power / IT Equipment Power
    • Goal: PUE < 1.2 (best-in-class).
  • Cooling Methods:
    • Air Cooling: Fans + cold aisles (used in most small data centers).
    • Liquid Cooling: Immersion cooling (used by Google’s data centers).
    • Free Cooling: Using outside air (e.g., NTC’s fiber optic nodes in cooler climates).
Server RoomCooling SystemPower SupplyMonitoringEnergy efficiency focus
Key components for energy-efficient server room design

3. Server Procurement Planning

Procurement is a structured process to select the right hardware while managing risks.

Step-by-Step Procurement Plan

A. Needs Assessment

  • Questions to Answer:
    • What is the server’s role? (e.g., eSewa’s authentication server needs high security).
    • What is the expected workload? (e.g., Pathao’s ride-matching server requires low latency).
    • What are compliance requirements? (e.g., PCI-DSS for banks).

B. Vendor Evaluation

Criteria Example Vendors Key Considerations
Hardware Quality Dell, HP, Lenovo MTBF (Mean Time Between Failures), warranties
Support Cisco, IBM 24/7 on-site support, SLAs
Price Local resellers Bulk discounts, hidden costs
Local Availability Nepal Data Center, Ncell Faster repairs, local expertise

C. Budgeting

  • Cost Components:
    • Upfront Costs: Server hardware, racks, UPS.
    • Operational Costs: Electricity, cooling, maintenance.
    • Hidden Costs: Downtime, data loss, upgrades.
  • Example Budget for a Small Enterprise:
    Item Cost (USD)
    Dell PowerEdge Server $3,500
    1TB NVMe Storage $800
    Rack Mounting $200
    UPS (1kVA) $500
    Total $4,000

D. Risk Management

Risk Mitigation Strategy
Hardware Failure RAID, redundant PSUs, UPS
Vendor Delays Multiple quotes, local backup vendors
Security Breaches Firewalls, EDR (Endpoint Detection & Response)
Power Outages UPS + diesel generator backup

Worked Example: Kathmandu Traffic Management System

  • Risk: Server failure during peak hours (e.g., Dashain) causes traffic gridlock.
  • Solution:
    • Redundant servers (active-passive setup).
    • Battery-backed UPS for 30-minute runtime.
    • Cloud backup of traffic data (AWS/Nepal Cloud).

4. Server Hardware Maintenance

Maintenance ensures longevity, performance, and security.

A. Preventive Maintenance

Task Frequency Tools Used
Dust Cleaning Monthly Compressed air, microfiber cloth
Firmware Updates Quarterly Dell EMC OpenManage, HP ILO
Temperature Checks Weekly Server monitoring tools (Zabbix, Nagios)
Backup Testing Bi-weekly Veeam, Acronis

B. Corrective Maintenance

  • RMA (Return Merchandise Authorization):
    • Process for replacing faulty hardware (e.g., dead motherboard in a Pathao server).
  • Hardware Replacement:
    • Example: Replacing a failed 10Gbps NIC in a Ncell core router.
  • Documentation:
    • Log all maintenance in a CMDB (Configuration Management Database).

In the Real World

  1. Ncell’s 4G/5G Core Network:

    • Hardware Strategy: Uses Cisco ASR 9000 routers with hot-swappable line cards for redundancy.
    • Why? Ensures 99.999% uptime (5 nines) for voice/data services.
    • Real Example: During the 2023 monsoon floods, Ncell’s redundant power systems kept the network running despite power outages in Kathmandu.
  2. eSewa’s Payment Gateway Servers:

    • Hardware: Dell PowerEdge R750 with RAID 10 storage and TLS 1.3 acceleration cards.
    • Why?
      • RAID 10 prevents data loss if a disk fails during high transaction volumes (e.g., Dashain).
      • TLS acceleration speeds up encrypted payments (critical for PCI-DSS compliance).
  3. Nepal Stock Exchange (NEPSE) Trading Servers:

    • Hardware: Custom-built servers with FPGA accelerators for ultra-low-latency trading.
    • Why?
      • FPGAs reduce order execution time from 10ms to <1ms.
      • Redundant NICs ensure no network bottleneck during high-frequency trading.

Exam Tip

This unit is heavily tested on:

  1. Server Types & Hardware Matching (5-mark questions):

    • Example: "Which server type requires the highest RAM and why?" → Database server (for caching query results).
    • Key Point: Memorize the hardware requirements for each server type (CPU, RAM, storage, NIC).
  2. Procurement Planning (4+4-mark questions):

    • Structure Your Answer:
      1. Needs Assessment (1 mark) → Define server role and workload.
      2. Vendor Selection (2 marks) → Compare Dell vs. HP vs. local vendors.
      3. Budgeting (1 mark) → Breakdown of costs (hardware + operational).
      4. Risk Management (2 marks) → Redundancy, SLAs, backup plans.
    • Example Question: "Prepare a procurement plan for a file server handling 1TB of student records for TU."
      • Answer Tip: Mention RAID 6 for storage, UPS for power, and local vendor (Nepal Data Center) for faster support.
  3. Hardware Strategies (5-mark questions):

    • Common Pitfalls:
      • ❌ Saying "RAID 0 is best for reliability" → Wrong! RAID 0 has no redundancy.
      • ❌ Ignoring cooling in data center design → Always mention PUE or hot/cold aisles.
    • Example Question: "What are the essential hardware strategies for a web server hosting a Nepali news website?"
      • Answer:
        • Redundancy: RAID 1 for OS, RAID 10 for user uploads.
        • Scalability: Load balancer + multiple web servers.
        • Cooling: Air conditioning with PUE < 1.5.
        • Backup: Daily cloud backups (AWS S3).

Final Checklist for Full Marks: ✅ Define key terms (RAID, PUE, hot-swappable). ✅ Compare at least two hardware strategies (e.g., RAID 5 vs. RAID 10). ✅ Apply to a real-world scenario (e.g., Ncell, eSewa, NEPSE). ✅ Use visuals (tables, diagrams, mermaid flows) to explain complex ideas. ✅ Structure answers clearly (e.g., procurement plan steps).

Based on the TU BCA syllabus for Network Administration (CACS406), unit 6.

Discussion

Loading…