Cloud ComputingUnit 613 min read

Cloud Storage: Types, Models, Technologies & Real-World Use

Unit 6 of Cloud Computing explores cloud storage architectures, data persistence models (object, block, file), storage-as-a-service (STaaS), scalability techniques, and real-world implementations like AWS S3, Google Cloud Storage, and NEPSE’s data lakes. Covers encryption, redundancy, and cost optimization with worked

Cloud Storage Fundamentals

Cloud storage is the on-demand, scalable, and elastic storage of data on remote servers managed by cloud providers. Unlike traditional storage (e.g., local hard drives or NAS), cloud storage abstracts physical infrastructure, offering accessibility, durability, and automatic scaling. Key characteristics:

  • Multi-tenancy: Shared infrastructure with logical isolation.
  • Pay-as-you-go: Costs tied to usage (storage volume, bandwidth, operations).
  • Geographic distribution: Data replicated across regions/zones for resilience.

Why Cloud Storage Matters

mindmap
  root((Cloud Storage Advantages))
    Cost
      No upfront hardware investment
      Pay per GB/month
    Scalability
      Auto-scaling for seasonal spikes (e.g., Daraz’s Diwali sales)
    Durability
      99.999999999% (11 9’s) for critical data (e.g., Ncell’s customer records)
    Accessibility
      Global low-latency access via CDNs
    Management
      Automated backups, snapshots, lifecycle policies

Storage Service Models

Cloud storage is delivered via three primary models, each with distinct use cases and trade-offs.

1. Object Storage

Definition: Stores unstructured data (e.g., images, videos, logs) as objects with metadata and a unique identifier (e.g., user_123/profile.jpg). No hierarchical folders—just a flat namespace.

How It Works

erDiagram
    BUCKET ||--o{ OBJECT : "contains"
    OBJECT {
        string key
        string metadata
        byte[] data
        timestamp created_at
    }
    BUCKET {
        string name
        string region
        string access_policy
    }
  • Example: AWS S3, Google Cloud Storage, Backblaze B2.
  • Use Cases:
    • Media hosting (YouTube, Netflix).
    • Backup archives (eSewa’s transaction logs).
    • Static website content (Daraz’s product images).

Worked Example: eSewa’s Transaction Logs

eSewa stores 10 million daily transactions in S3 with:

  • Key: txn_<timestamp>_<user_id>.json
  • Metadata: {"amount": 5000, "status": "completed", "service": "electricity"}
  • Lifecycle Rule: Move to Glacier Deep Archive after 30 days (cost: $0.00099/GB/month).

Cost Calculation:

  • 10M transactions × avg 1KB = 10TB/month.
  • S3 Standard: $0.023/GB → $230/month.
  • Glacier: $0.00099/GB → $9.90/month (96% savings).

2. Block Storage

Definition: Stores structured data (e.g., databases, VM disks) as fixed-size blocks (typically 4KB–1MB). Mounted as a volume to a VM or container.

How It Works

  • Example: AWS EBS, Azure Disk Storage, Google Persistent Disk.
  • Use Cases:
    • Databases (MySQL on NEPSE’s trading platform).
    • Boot volumes for VMs (Pathao’s backend servers).
    • Enterprise apps (ERP systems in Nepali banks).

Comparison: Object vs. Block Storage

Feature Object Storage Block Storage
Data Type Unstructured (files, blobs) Structured (blocks, disks)
Access Pattern HTTP/REST API (GET/PUT) Low-latency, high IOPS
Use Case Backups, media, logs Databases, VMs, OS disks
Scalability Horizontal (add buckets) Vertical (resize volume)
Cost $0.02–$0.05/GB/month $0.10–$0.50/GB/month

3. File Storage

Definition: Provides shared file systems (e.g., NFS, SMB) accessible via standard protocols. Supports hierarchical directories and concurrent access.

How It Works

sequenceDiagram
    participant User
    participant Client as "NFS Client"
    participant Server as "Cloud File Storage (e.g., AWS EFS)"
    User->>Client: Open file.txt
    Client->>Server: NFS Request (read)
    Server-->>Client: File data
    Client-->>User: Display content
  • Example: AWS EFS, Azure Files, Google Filestore.
  • Use Cases:
    • Collaborative editing (Google Docs-like apps).
    • Development environments (shared code repos for Nepali startups).
    • Media workflows (video editing teams).

Real-World Example: Kathmandu Traffic Management The Kathmandu Metropolitan City uses AWS EFS to share real-time traffic camera feeds across 50+ monitoring stations. Files are structured as:

/traffic/
  ├── 2023-10/
  │   ├── camera_001_10-00-00.mp4
  │   └── metadata.json
  └── rules/
      └── speed_limits.csv

Storage Architectures

Cloud providers use distributed architectures to ensure durability, availability, and performance.

1. Storage Classes

Providers offer tiered storage based on access frequency and cost. Example: AWS S3 classes:

mindmap
  root((AWS S3 Storage Classes))
    Standard
      Frequent access, low latency
      $0.023/GB
    Intelligent-Tiering
      Auto-moves data between tiers
      $0.0225/GB + monitoring fee
    Infrequent Access (IA)
      $0.0125/GB + retrieval fee
    Glacier
      Archive (3–5 hrs retrieval)
      $0.0036/GB
    Glacier Deep Archive
      Compliance (12+ hrs retrieval)
      $0.00099/GB

Worked Example: NEPSE’s Historical Data NEPSE stores 20 years of stock market data (50TB) in S3 Glacier Deep Archive:

  • Cost: 50TB × $0.00099/GB/month = $49.50/month (vs. $1,150 in Standard).
  • Retrieval: Query via AWS Athena (serverless SQL) when analysts need old trades.

2. Redundancy and Replication

Cloud storage ensures durability via:

  • Geographic Replication: Data copied across Availability Zones (AZs) or Regions.
  • Erasure Coding: Splits data into fragments + parity bits (e.g., 10+2 erasure coding = 10% overhead for 2x fault tolerance).
  • Versioning: Retains multiple versions of objects (e.g., file.txt?v=1, file.txt?v=2).

Example: Google Cloud Storage uses dual-copy replication within a region and multi-regional replication for global access.


3. Caching Layers

To reduce latency, cloud storage uses:

  • CDNs (Content Delivery Networks): Edge caches for static content (e.g., Daraz’s product images).
  • In-Memory Caches: Redis/Memcached for frequently accessed data (e.g., Ncell’s subscriber lookup tables).
sequenceDiagram
    participant User
    participant CDN as "Cloudflare (Edge Cache)"
    participant Origin as "AWS S3 (Origin)"
    User->>CDN: Request image.jpg
    CDN-->>User: Serve from cache (90% hits)
    User->>CDN: Request new_image.jpg
    CDN->>Origin: Fetch missing file
    Origin-->>CDN: Return file
    CDN-->>User: Serve + cache

Cloud Storage Technologies

1. Distributed File Systems

  • HDFS (Hadoop Distributed File System): Used by Google Cloud Storage and AWS S3 for large-scale analytics.
  • Ceph: Open-source, used by OpenStack Swift.
  • Google File System (GFS): Inspired Colossus (Google’s storage backbone).

2. Object Storage Backends

  • Ceph RADOS: Used by OpenStack Swift.
  • MinIO: Open-source S3-compatible storage (used by Nepali startups for cost savings).
  • Azure Blob Storage: Uses Azure Data Lake Storage Gen2 for hierarchical namespace.

3. Hybrid and Edge Storage

  • AWS Outposts: Extends S3/EBS to on-premises data centers (used by Nepali banks for compliance).
  • Google Distributed Cloud: Runs Anthos on-premises with cloud storage integration.
  • Edge Caching: Cloudflare Workers cache data at 300+ locations globally.

Security in Cloud Storage

1. Data Encryption

  • At Rest: AES-256 (e.g., AWS S3 Server-Side Encryption).
  • In Transit: TLS 1.2+ (e.g., HTTPS for S3 API).
  • Customer-Managed Keys: AWS KMS, Google Cloud KMS, Azure Key Vault.

2. Access Control

  • IAM Policies: Fine-grained permissions (e.g., s3:GetObject for read-only access).
  • Bucket Policies: Restrict access by IP/region (e.g., block public access to NEPSE’s trading data).
  • VPC Endpoints: Private network access to S3 (avoids public internet exposure).

3. Compliance

  • GDPR: Right to erasure (e.g., Khalti must delete user data on request).
  • HIPAA: Encryption for healthcare data (e.g., CIAA hospitals using Azure Blob Storage).
  • PCI DSS: Tokenization for payment data (e.g., eSewa’s credit card storage).

Performance Optimization

1. Latency Reduction

  • Region Selection: Store data closer to users (e.g., ap-south-1 for Nepal vs. us-east-1).
  • Transfer Acceleration: AWS S3 Transfer Acceleration (uses CloudFront edge locations).

2. Bandwidth Management

  • Compression: Gzip/Brotli for text data (e.g., JSON logs).
  • Chunked Uploads: Parallel uploads for large files (e.g., 5GB video to Google Cloud Storage).

3. Lifecycle Management

Automate data transitions:

  1. Hot → Cool: Move infrequently accessed data to S3 IA.
  2. Cool → Archive: Transition to Glacier after 90 days.
  3. Archive → Delete: Expire old logs after 7 years (compliance).
stateDiagram-v2
    [*] --> Hot
    Hot --> Cool: After 30 days of inactivity
    Cool --> Glacier: After 90 days of inactivity
    Glacier --> DeepArchive: After 1 year
    DeepArchive --> [*]: After 7 years (retention policy)

In the Real World

  1. eSewa’s Transaction Processing

    • Technology: AWS S3 (object storage) + DynamoDB (metadata).
    • How It Works:
      • Each transaction is stored as a JSON object in S3 with a key like txn_20231001_123456.json.
      • Lifecycle Rule: Moves to Glacier after 30 days.
      • Security: Encrypted with AWS KMS and restricted via IAM roles.
    • Impact: Handles 50,000 transactions/hour during festivals (e.g., Dashain).
  2. Daraz’s Inventory Management

    • Technology: AWS EFS (file storage) + S3 (product images).
    • How It Works:
      • Product catalog stored in EFS (shared across 100+ servers).
      • High-res images in S3 with CloudFront CDN for global delivery.
      • Cost Optimization: Uses S3 Intelligent-Tiering to auto-move old images to cheaper storage.
    • Impact: Reduces image load time from 2s → 200ms via CDN.
  3. NEPSE’s Trading Platform

    • Technology: Google Cloud Storage (for historical data) + Firestore (real-time trades).
    • How It Works:
      • Real-time trades: Stored in Firestore (NoSQL database).
      • Historical data (20+ years): Archived in Coldline Storage ($0.003/GB).
      • Disaster Recovery: Data replicated across asia-southeast1 (Singapore) and europe-west1 (Belgium).
    • Impact: Ensures 99.999% uptime during market crashes.

Exam Tip

What Examiners Look For

  1. Definitions: Clearly distinguish object, block, and file storage (1 mark each).
  2. Worked Examples: Show cost calculations (e.g., S3 vs. Glacier) or lifecycle policies (2 marks).
  3. Diagrams: Draw storage class transitions (state diagram) or erasure coding (1 mark).
  4. Real-World Links: Relate to Nepali companies (eSewa, Daraz, NEPSE) for application-based questions (3 marks).
  5. Security: Mention encryption (AES-256), IAM policies, and compliance (GDPR/HIPAA) (2 marks).

Common Pitfalls

  • Confusing Block vs. Object: Block storage is for databases/VMs; object is for files/media.
  • Ignoring Costs: Always calculate GB/month × rate for storage classes.
  • Overlooking Redundancy: Assume multi-AZ replication unless stated otherwise.
  • Vague Answers: Instead of "cloud storage is scalable," say:

    "AWS S3 auto-scales to petabytes by distributing data across 100+ servers in a region, ensuring <99.99% availability even during traffic spikes like Daraz’s Diwali sales."

Sample Exam Question & Answer

Question: "Explain how Ncell could use cloud storage to reduce costs for storing 1PB of customer call logs while ensuring compliance with Nepali data laws. Include a storage class recommendation and security measures."

Model Answer: Ncell can optimize storage costs and compliance using AWS S3 with the following strategy:

  1. Storage Class:

    • Hot Storage (S3 Standard): For recent logs (≤30 days) accessed frequently for fraud detection.
      • Cost: 100TB × $0.023/GB/month = $2,300/month.
    • Cool Storage (S3 IA): For logs 30–90 days old (infrequent access).
      • Cost: 500TB × $0.0125/GB/month = $625/month.
    • Archive (Glacier Deep Archive): For logs >90 days old (compliance retention).
      • Cost: 400TB × $0.00099/GB/month = $39.60/month.
    • Total Cost: $2,964.60/month (vs. $23,000 in on-prem HDDs).
  2. Security Measures:

    • Encryption: Enable S3 Server-Side Encryption (SSE-S3) with AES-256 for all buckets.
    • Access Control:
      • Restrict IAM roles to least privilege (e.g., s3:GetObject for analytics teams).
      • Use VPC Endpoints to avoid public internet exposure.
    • Compliance:
      • Enable S3 Object Lock to prevent deletion for 7 years (as per Nepali telecom regulations).
      • Store logs in Asia-Pacific (Mumbai) region to comply with Nepali data localization laws.
  3. Performance:

    • Use S3 Intelligent-Tiering to auto-move logs between classes.
    • For real-time fraud detection, cache hot logs in Amazon ElastiCache (Redis).

Visual Aid:

stateDiagram-v2
    [*] --> Hot: "≤30 days (S3 Standard)"
    Hot --> Cool: "After 30 days (S3 IA)"
    Cool --> Archive: "After 90 days (Glacier)"
    Archive --> [*]: "After 7 years (retention policy)"

Based on the TU BIM syllabus for Cloud Computing (IT277), unit 6.

Discussion

Loading…