Cloud ComputingUnit 615 min read

Cloud Storage: Types, Models, Security & Real-World Use

Unit 6 of Cloud Computing explores cloud storage architectures, data redundancy, storage-as-a-service models, security best practices, and real-world applications like eSewa’s transaction logs and Ncell’s customer data. Learn how data is stored, accessed, and protected in cloud environments, with comparisons to traditi

Cloud Storage Fundamentals

What is Cloud Storage?

Cloud storage is a service model that allows users to store, manage, and retrieve data over the internet using third-party cloud providers instead of local storage devices (like hard drives or SSDs). It leverages distributed storage systems, data replication, and scalable infrastructure to ensure availability, durability, and accessibility.

Key Characteristics:

  • On-demand scalability: Storage capacity can be increased or decreased dynamically.
  • Multi-tenancy: Multiple users or organizations share the same physical storage infrastructure.
  • Geographic distribution: Data is replicated across multiple data centers for redundancy.
  • Pay-as-you-go pricing: Users pay only for the storage they use.

How Cloud Storage Works

Cloud storage relies on a three-tier architecture:

  1. Client Tier: Users or applications that request storage services (e.g., uploading a file to Google Drive).
  2. Storage Tier: Physical or virtual storage systems (e.g., hard drives, SSDs, or object storage like Amazon S3).
  3. Interface Tier: APIs, web portals, or command-line tools that interact with the storage system.

Data Storage Models:

Cloud storage typically uses one of these models:

  1. Block Storage: Data is split into fixed-size blocks (e.g., 4KB) and stored independently. Used for databases and boot volumes.
    • Example: AWS EBS (Elastic Block Store).
  2. File Storage: Hierarchical file systems (folders, subfolders) with metadata (e.g., permissions, timestamps).
    • Example: Google Drive, Dropbox.
  3. Object Storage: Data is stored as objects with unique identifiers (keys) and metadata. Scalable and used for unstructured data.
    • Example: Amazon S3, Azure Blob Storage.

Cloud Storage Service Models

Cloud storage is often delivered as part of broader cloud service models (IaaS, PaaS, SaaS), but it can also be offered as a standalone service. Here’s how it fits into the cloud service hierarchy:

Service Model Cloud Storage Role Example Providers
IaaS Block or file storage as a virtual resource. AWS EBS, Azure Disk Storage
PaaS Managed storage for applications (e.g., databases). Google Cloud SQL, Azure SQL Database
SaaS Built-in storage for end-user applications. Google Workspace, Microsoft 365
Standalone Dedicated storage service for backups or archives. Backblaze B2, Wasabi Hot Storage

In the Real World

  1. eSewa’s Transaction Logs: eSewa uses object storage (like AWS S3) to store transaction records securely. Each transaction is stored as an immutable object with metadata (timestamp, user ID, amount). This ensures compliance with Nepal’s financial regulations and allows quick retrieval for audits.

    • Why object storage? Scalability (millions of transactions/day) and durability (data never lost due to replication).
  2. Ncell’s Customer Data: Ncell’s CRM system relies on block storage (AWS EBS) for real-time customer data (e.g., call logs, billing records). The data is replicated across multiple Availability Zones in AWS to prevent downtime during peak usage (e.g., during festivals like Dashain).

    • Why block storage? Low-latency access for database operations (e.g., updating a customer’s plan).
  3. Daraz’s Product Catalog: Daraz uses Google Cloud Storage (GCS) to host product images, descriptions, and inventory data. During sales events (e.g., 11.11), the system scales storage dynamically to handle spikes in uploads (e.g., new product listings).

    • Why object storage? Cost-efficiency for static assets (images, videos) and easy CDN integration.

Worked Example: NTC’s Network Log Storage

Scenario: The Nepal Telecommunications Corporation (NTC) needs to store 1TB of network logs daily for compliance. They choose Azure Blob Storage (object storage) with the following configuration:

  • Storage Tier: Hot storage (frequently accessed logs) + Cool storage (archived logs older than 30 days).
  • Redundancy: Geo-redundant storage (GRS) to replicate data across Nepal and Singapore.
  • Lifecycle Policy: Automatically move logs to Cool storage after 30 days, then archive to tape after 1 year.

Cost Calculation:

  • Hot storage: $0.02/GB/month → 1TB/month = $20/month.
  • Cool storage: $0.01/GB/month → 900GB/month (after 30 days) = $9/month.
  • Data transfer: $0.05/GB for cross-region replication → 1TB/month = $50/month.
  • Total: ~$79/month (scalable for future growth).

Cloud Storage Architectures

Cloud storage systems use distributed architectures to ensure high availability and fault tolerance. Here’s a breakdown:

1. Distributed File Systems

These systems spread data across multiple servers while presenting a unified file system to users. Examples:

  • Google File System (GFS): Used by Google for large-scale data storage (e.g., YouTube videos).
  • Hadoop Distributed File System (HDFS): Used for big data analytics (e.g., Facebook’s data processing).
Read/WriteReplicatedReplicatedReplicatedClientNameNode (Metadata)DataNode 1DataNode 2DataNode 3Block 1
HDFS architecture: NameNode manages metadata while DataNodes store replicated blocks (e.g., YouTube’s video chunks).

How it works:

  • The NameNode tracks metadata (file locations, permissions).
  • DataNodes store actual data blocks (typically 128MB–1GB each).
  • Data is replicated 3x by default for fault tolerance.

2. Object Storage Architectures

Object storage systems store data as objects (key-value pairs) with metadata. Key components:

  • Object: Data + metadata (e.g., user_123_profile.jpg with size=1MB, upload_date=2024-05-01).
  • Bucket: A container for objects (analogous to a folder).
  • API Gateway: Handles HTTP requests (e.g., PUT, GET, DELETE).
PUT/GET ObjectTracksClientAPI GatewayMetadata DBStorage Node 1Storage Node 2Bucket: user_uploadsObject 1 (img1.jpg)Object 2 (img2.jpg)
Object storage workflow: API Gateway routes requests to storage nodes (e.g., AWS S3).

Example: Amazon S3

  • Consistency Model: Strong consistency for PUT/DELETE, eventual consistency for GET after a PUT.
  • Storage Classes:
    • S3 Standard: High durability (99.999999999%), low latency.
    • S3 IA (Infrequent Access): Cheaper for rarely accessed data.
    • S3 Glacier: Ultra-cheap archival storage (retrieval takes hours).

3. Hybrid Storage Architectures

Combine on-premises storage with cloud storage for performance and cost optimization. Example:

  • Nepal Rastra Bank’s Data Center:
    • On-premises: High-performance SSDs for real-time transaction processing.
    • Cloud (AWS): Cold storage for archived financial records (e.g., 10-year-old loan data).
    • Sync: Data is replicated nightly using AWS Storage Gateway.
graph LR
    A["On-Premises\n(SSD Arrays)"] -->|"Nightly Sync"| B["AWS Storage Gateway"]
    B --> C["AWS S3\n(Cold Storage)"]
    D["Cloud App\n(e.g., Analytics)"] -->|"Direct Access"| C

Cloud Storage Security

Security is critical in cloud storage due to the shared responsibility model (provider secures infrastructure; user secures data). Key mechanisms:

1. Data Encryption

  • At Rest: Data encrypted on storage devices (e.g., AES-256).
    • Example: Google Cloud encrypts data by default using customer-supplied keys (CMEK).
  • In Transit: TLS/SSL for data moving between client and storage.
  • Client-Side: Users encrypt data before uploading (e.g., using AWS KMS).
Data at RestAES-256Data in TransitTLS 1.3Data in UseHomomorphic Encryption
Encryption types: Protecting data in storage, transit, and processing.

2. Access Control

  • IAM (Identity and Access Management): Role-based access (e.g., storage-admin, read-only).
  • Bucket Policies: Define who can access a bucket (e.g., only user@example.com).
  • Temporary Credentials: Short-lived tokens (e.g., AWS STS).

3. Compliance and Auditing

  • Standards: HIPAA (healthcare), GDPR (EU data), PCI-DSS (payments).
  • Logging: Track access (e.g., AWS CloudTrail logs all S3 API calls).
  • Immutable Storage: Prevent deletion/modification (e.g., AWS S3 Object Lock).

Worked Example: eSewa’s Secure Transaction Storage

Requirements:

  • Store transaction data with immutability (for audits).
  • Encrypt data at rest and in transit.
  • Restrict access to only eSewa’s compliance team.
Transaction RequestEncryptStoreLogUser DeviceeSewa AppAWS KMSS3 Bucket (Encrypted)Database (Tokenized)
eSewa’s security: KMS encrypts data before S3 storage; tokens replace raw card numbers.

Solution:

  1. Storage: AWS S3 with Object Lock (WORM mode) to prevent deletions.
  2. Encryption:
    • At rest: AWS KMS with customer-managed keys.
    • In transit: TLS 1.3 for all API calls.
  3. Access Control:
    • IAM role eSewa-Compliance-ReadOnly with least-privilege permissions.
    • Bucket policy to allow only esewa.com domain access.
  4. Audit Trail: Enable AWS CloudTrail to log all PutObject and GetObject events.

Cost:

  • S3 Standard: $0.023/GB → 500GB/month = $11.50.
  • KMS: $1/month for key management.
  • CloudTrail: $0.10/GB logged → 10GB/month = $1.
  • Total: ~$13.50/month (scalable).

Cloud Storage Performance Optimization

Performance depends on latency, throughput, and consistency. Techniques to optimize:

Technique Use Case Example
Caching Reduce latency for frequently accessed data. Amazon CloudFront (CDN caching).
Sharding Distribute load across multiple storage nodes. MongoDB’s sharded clusters.
Erasure Coding Reduce storage overhead for replication. Ceph (uses 6+3 erasure coding).
Tiered Storage Move cold data to cheaper storage. AWS S3 Lifecycle Policies.

Worked Example: Pathao’s Ride Data Storage

Scenario: Pathao needs to store 10TB of ride data daily (GPS logs, driver info) with low latency for real-time analytics. Solution:

  1. Hot Storage: AWS S3 Standard for recent data (<30 days).
  2. Cold Storage: S3 Glacier Deep Archive for older data (>1 year).
  3. Caching: Amazon ElastiCache (Redis) for frequently queried driver locations.
  4. Sharding: Data partitioned by ride_id across 10 S3 buckets.

Performance Metrics:

  • Read Latency: 10–50ms (cached data), 100–300ms (S3).
  • Write Throughput: 10GB/s (using S3 Transfer Acceleration).
  • Cost: ~$200/month (scalable for 1M daily rides).

Cloud Storage vs. Traditional Storage

Compare cloud storage to on-premises storage (e.g., NAS, SAN) using key metrics:

Metric Cloud Storage Traditional Storage
Scalability Infinite (pay-as-you-go). Limited by hardware capacity.
Maintenance Fully managed by provider. User responsible for hardware/software.
Cost Pay for what you use (no upfront CAPEX). High upfront cost (servers, racks).
Availability 99.99%–99.999% SLA (multi-region). Depends on local infrastructure.
Disaster Recovery Built-in (geo-replication). Requires manual setup (e.g., offsite backups).
Accessibility Global access via internet. Limited to on-premises or VPN.
Security Shared responsibility model. Fully user’s responsibility.

Cloud Storage Providers Comparison

Major providers offer similar services but with different strengths:

Provider Key Offerings Best For Pricing Model
AWS S3, EBS, EFS, Glacier. Enterprises needing global reach. Pay-per-GB + data transfer fees.
Azure Blob Storage, Azure Files, Cool Storage. Microsoft ecosystem (e.g., .NET apps). Tiered pricing (Hot/Cool/Archive).
Google Cloud Cloud Storage, Persistent Disk, Coldline. AI/ML workloads (integrates with BigQuery). Nearline/Coldline for archival.
IBM Cloud Cloud Object Storage, File Storage. Hybrid cloud (on-prem + cloud). Predictable pricing for reserved capacity.
Oracle Cloud Object Storage, Block Volume. Oracle database workloads. Flexible pricing (hourly/daily).

Worked Example: Kathmandu Traffic Route Optimization

Scenario: The Kathmandu Metropolitan City (KMC) wants to store 5TB of traffic camera data monthly for AI-based congestion analysis. Solution:

  • Provider: Google Cloud Storage (GCS) for cost-efficiency.
  • Storage Classes:
    • Hot Storage: First 30 days (for real-time analysis).
    • Coldline: Next 6 months (historical data).
    • Archive: Data older than 1 year.
  • Cost:
    • Hot: 5TB × $0.02/GB = $100/month.
    • Coldline: 4TB × $0.01/GB = $40/month.
    • Archive: 1TB × $0.004/GB = $4/month.
    • Total: ~$144/month (vs. $500/month for on-premises NAS).

Exam Tip

For the Cloud Storage unit in TU exams, focus on these high-yield areas:

  1. Definitions and Models:
    • Differentiate block, file, and object storage with examples (e.g., EBS vs. S3).
    • Explain distributed file systems (GFS, HDFS) and their components (NameNode, DataNode).
  2. Architectures:
    • Draw and label a cloud storage architecture diagram (client → API → storage nodes).
    • Compare hybrid storage (on-prem + cloud) with pure cloud.
  3. Security:
    • List 3 encryption methods (at rest, in transit, client-side) and where each is used.
    • Describe IAM roles and bucket policies with a real-world example (e.g., eSewa’s access control).
  4. Performance Optimization:
    • Explain caching, sharding, and tiered storage with a worked example (e.g., Pathao’s ride data).
  5. Cost Calculations:
    • Practice pricing scenarios (e.g., "Calculate the monthly cost for storing 10TB in AWS S3 Standard").
  6. Real-World Applications:
    • Relate concepts to Nepali companies (e.g., NTC’s logs, Daraz’s product catalog) or global platforms (e.g., YouTube’s GFS).
  7. Diagrams:
    • Be ready to draw:
      • A distributed file system (NameNode + DataNodes).
      • An object storage workflow (client → API → bucket → object).
      • A hybrid storage setup (on-prem + cloud).

Common Pitfalls:

  • Confusing SaaS storage (e.g., Google Drive) with IaaS storage (e.g., AWS EBS).
  • Forgetting to mention replication or geo-redundancy in high-availability scenarios.
  • Overlooking cost factors (e.g., data transfer fees, storage class differences).

Based on the TU BITM syllabus for Cloud Computing (IT277), unit 6.

Discussion

Loading…