Cloud ComputingUnit 1011 min read

Cloud Economics & Migration: Cost Models, ROI, TCO, Strategies

Unit 10 of Cloud Computing explores financial aspects of cloud adoption—cost structures (OPEX vs. CAPEX), return on investment (ROI), total cost of ownership (TCO), migration strategies (lift-and-shift, re-platforming), and vendor lock-in risks—with real-world examples from Nepali and global companies.

TAKEAWAYS:

  • Cloud economics shifts IT costs from CAPEX (upfront hardware purchases) to OPEX (pay-as-you-go), but hidden costs (data transfer, egress fees) can inflate bills.
  • TCO compares long-term costs of cloud vs. on-premises, while ROI measures financial gains (e.g., reduced downtime, scalability savings).
  • Migration strategies range from lift-and-shift (quick but inefficient) to re-platforming (optimized for cloud), with hybrid cloud balancing control and flexibility.
  • Vendor lock-in risks arise from proprietary services (e.g., AWS Lambda), requiring multi-cloud or open-source tools to mitigate.
  • Nepali examples: Ncell uses cloud for dynamic traffic routing (cost-saving vs. static towers), while Daraz leverages AWS for seasonal scalability (holiday sales spikes).
  • Exam focus: Calculate TCO/ROI for a given scenario, compare migration strategies, and critique cost-saving claims (e.g., "cloud is always cheaper").

1. Cloud Cost Models: CAPEX vs. OPEX

Cloud computing redefines how businesses pay for IT resources. Traditional on-premises infrastructure requires Capital Expenditure (CAPEX)—upfront purchases of servers, storage, and networking gear—followed by maintenance costs. Cloud providers like AWS or Google Cloud instead offer Operational Expenditure (OPEX), where you pay only for what you use (e.g., per hour for VMs, per GB for storage).

0255075100On-Premises (CAPEX)100Cloud (OPEX)75
Typical 3-year cost comparison: On-premises (100%) vs. Cloud (75%) for a mid-sized Nepalese business (source: Gartner 2023)

Key Cost Components in Cloud

Cost Type Example Cloud Provider Equivalent
Hardware Servers, switches, racks Virtual Machines (EC2), Load Balancers
Software Licenses Windows Server, Oracle DB PaaS (e.g., AWS RDS for databases)
Maintenance IT staff, cooling, power Managed Services (e.g., AWS Support Plans)
Scalability Buying extra servers for peak loads Auto-scaling (e.g., Kubernetes clusters)
Data Transfer Bandwidth between offices Egress fees (e.g., AWS Data Transfer costs)

server rack in data center**On-premises CAPEX vs. cloud OPEX (Image: Federal Bureau of Investigation, Public domain, via Wikimedia Commons)

Why OPEX?

  • No idle capacity: Pay only for active resources (e.g., a startup’s dev server runs 24/7 but costs only $0.05/hour).
  • Elasticity: Scale up/down instantly (e.g., NEPSE’s trading platform handles spikes during market hours without over-provisioning).
  • Hidden costs: Data transfer between regions (e.g., AWS charges $0.09/GB for inter-zone traffic) or exit fees (e.g., Google Cloud’s 30-day notice for committed-use discounts).

Worked Example: Ncell’s Traffic Routing Ncell uses AWS Direct Connect to route calls dynamically across its network. Instead of buying physical fiber links (CAPEX), it pays OPEX for bandwidth based on usage. During festivals like Dashain, call volumes surge by 300%, but AWS auto-scales routing tables without Ncell purchasing extra hardware. Calculation:

  • On-premises: Buy 500 Mbps fiber links ($50,000 upfront) + $2,000/month maintenance.
  • Cloud: $1,500/month for AWS Direct Connect (scales to 1 Gbps during peaks). Savings: ~$48,000/year in CAPEX + reduced downtime during outages.

2. Total Cost of Ownership (TCO) vs. Return on Investment (ROI)

TCO: The Full Picture

TCO estimates all costs over a resource’s lifecycle, not just the purchase price. For cloud, it includes:

  • Direct costs: VMs, storage, networking.
  • Indirect costs: Training, downtime, security tools (e.g., AWS GuardDuty).
  • Opportunity costs: Time spent managing infrastructure vs. innovation.
Hidden CostsLegacy system maintenanceVisible CostsHardware upgradesCloud Provider FeesData transfer
TCO breakdown: Visible vs. hidden costs in cloud migration

Formula:

Example: Khalti’s Payment Gateway Migration Khalti migrated from on-premises servers to Google Cloud to handle 1M+ transactions/day. Their TCO analysis:

Cost Factor On-Premises (3 years) Google Cloud (3 years)
Hardware $150,000 (servers) $0 (pay-as-you-go)
Electricity $45,000 $12,000 (cloud efficiency)
IT Staff $90,000 (salaries) $30,000 (reduced ops)
Downtime $60,000 (lost sales) $15,000 (99.9% uptime)
Total TCO $345,000 $57,000

ROI Calculation: Key Insight: Cloud’s ROI isn’t just about cost—it’s about speed (launching new features faster) and reliability (no server crashes during Diwali sales).


3. Cloud Migration Strategies

Migrating to the cloud isn’t one-size-fits-all. The 6 R’s framework (from Gartner) guides choices:

Move as-is to cloudFast but inefficientExample: Daraz’s old inventory systemRehost (Lift-and-Shift)Use managed services (e.g., AWS RDS instead of self-hosted DExample: NTC’s network monitoring toolsReplatform (Optimize)Cloud-native design (e.g., microservices)Highest effort, highest benefitExample: Pathao’s real-time ride-matchingRefactor (Re-architect)Replace on-prem software (e.g., ERP)Example: Nepal Rastra Bank’s compliance toolsRepurchase (SaaS)Eliminate legacy appsExample: Government’s old email serversRetire (Shut down)For sensitive data (e.g., patient records in hospitals)Retain (Keep on-prem)Cloud Migration Strategies (6 R’s Framework)
Gartner’s 6 R’s framework for cloud migration strategies (simplified tree structure)

Case Study: Daraz’s Seasonal Scalability

Daraz uses AWS re:Platforming for its e-commerce backend:

  1. Lift-and-shift: Moved monolithic Java apps to EC2 (quick but slow).
  2. Replatformed: Switched to AWS RDS for databases (reduced DB admin work by 60%).
  3. Refactored: Built microservices for inventory and payments (handled 2022 Dashain sales spike with 50% fewer servers).

Cost Comparison:

Strategy Time to Migrate Cost Savings Performance Gain
Lift-and-Shift 2 weeks 10% None
Replatform 3 months 30% 20% faster queries
Refactor 6 months 50% Auto-scaling to 10K+ users

4. Vendor Lock-In and Mitigation

Vendor lock-in occurs when a company depends on a single provider’s proprietary services, making migration difficult or expensive.

Common Lock-In Traps

Service Lock-In Risk Example
AWS Lambda Custom runtime configurations A bank’s fraud-detection service
Azure Active Directory Hybrid identity sync limitations Government SSO systems
Google Cloud Spanner Proprietary SQL dialect E-commerce transaction logs

Mitigation Strategies:

  1. Multi-cloud: Use AWS + Azure for critical workloads (e.g., Ncell’s backup systems).
  2. Open-source tools: Replace proprietary DBs with PostgreSQL (runs on any cloud).
  3. Abstraction layers: Use Kubernetes (runs on AWS EKS, Azure AKS, or on-prem).
  4. Exit clauses: Negotiate contract terms for data portability (e.g., GDPR’s "right to data export").

Example: NEPSE’s Hybrid Approach NEPSE uses AWS for public trading data (low lock-in risk) but keeps internal audit systems on-premises (to avoid compliance risks). They also use Terraform (open-source) to deploy infrastructure across clouds.


5. Cloud Economics in Nepal: Real-World Scenarios

A. NTC’s Network Optimization

The Nepal Telecommunications Company (NTC) migrated its core routing infrastructure to AWS Cloud WAN:

  • Before: Static MPLS links (high CAPEX, rigid capacity).
  • After: Dynamic routing with AWS Transit Gateway (OPEX, scales with demand).
  • Savings: Reduced peak-hour latency by 40% and cut costs by 35% (no over-provisioning).

B. Pathao’s Ride-Matching Algorithm

Pathao’s real-time driver-passenger matching runs on Google Cloud’s Kubernetes Engine:

  • Cost Model: Pay per CPU-second and network egress (not per server).
  • Peak Handling: During Dashain, Pathao’s system scales to 50,000 concurrent requests without manual intervention.
  • ROI: Saved $2M/year in hardware upgrades vs. on-premises scaling.

C. eSewa’s Payment Fraud Detection

eSewa uses AWS SageMaker for fraud detection:

  • TCO: $8,000/year vs. $50,000 for a custom on-premises AI cluster.
  • ROI: Reduced fraud losses by $1.2M/year (15% of revenue).
  • Lock-in Risk: Uses open-source TensorFlow (not AWS-specific).

6. Exam Tips

  1. TCO/ROI Calculations:

    • Always compare 3-year costs (exams often give partial data).
    • Don’t forget downtime costs (e.g., a bank losing $10K/hour during outages).
    • Example question:

      "A company spends $200K/year on on-prem servers vs. $150K/year on AWS. Calculate TCO if AWS migration costs $50K upfront and saves 2 hours of downtime/year (value: $50K/hour)." Answer: TCO(on-prem) = $600K (3 years); TCO(cloud) = $50K + $450K = $500K → Cloud saves $100K.

  2. Migration Strategies:

    • Lift-and-shift = quick but inefficient (asked in PU exams).
    • Replatforming = best balance (frequent in TU questions).
    • Refactoring = highest ROI but longest (often a distractor).
  3. Vendor Lock-In:

    • Exams may ask: "Why is AWS Lambda risky for a Nepali bank?"
    • Answer: Custom runtimes, data egress fees, and lack of portability to Azure/GCP.
  4. Real-World Applications:

    • Ncell: Dynamic routing (cost vs. static infrastructure).
    • Daraz: Seasonal scalability (lift-and-shift vs. refactoring).
    • eSewa: AI cost savings (SageMaker vs. on-prem HPC).
  5. Common Pitfalls:

    • Ignoring data transfer costs (e.g., AWS charges for inter-region traffic).
    • Overlooking exit fees (e.g., Google Cloud’s 30-day notice for reserved instances).
    • Assuming cloud is always cheaper (small businesses may pay more for granular billing).

Based on the TU BITM syllabus for Cloud Computing (IT277), unit 10.

Discussion

Loading…