Cloud ComputingUnit 613 min read
Cloud Storage: Types, Models, Technologies & Real-World Use
Unit 6 of Cloud Computing explores cloud storage architectures, data persistence models (object, block, file), storage-as-a-service (STaaS), scalability techniques, and real-world implementations like AWS S3, Google Cloud Storage, and NEPSE’s data lakes. Covers encryption, redundancy, and cost optimization with worked
Cloud Storage Fundamentals
Cloud storage is the on-demand, scalable, and elastic storage of data on remote servers managed by cloud providers. Unlike traditional storage (e.g., local hard drives or NAS), cloud storage abstracts physical infrastructure, offering accessibility, durability, and automatic scaling. Key characteristics:
- Multi-tenancy: Shared infrastructure with logical isolation.
- Pay-as-you-go: Costs tied to usage (storage volume, bandwidth, operations).
- Geographic distribution: Data replicated across regions/zones for resilience.
Why Cloud Storage Matters
mindmap
root((Cloud Storage Advantages))
Cost
No upfront hardware investment
Pay per GB/month
Scalability
Auto-scaling for seasonal spikes (e.g., Daraz’s Diwali sales)
Durability
99.999999999% (11 9’s) for critical data (e.g., Ncell’s customer records)
Accessibility
Global low-latency access via CDNs
Management
Automated backups, snapshots, lifecycle policiesStorage Service Models
Cloud storage is delivered via three primary models, each with distinct use cases and trade-offs.
1. Object Storage
Definition: Stores unstructured data (e.g., images, videos, logs) as objects with metadata and a unique identifier (e.g., user_123/profile.jpg). No hierarchical folders—just a flat namespace.
How It Works
erDiagram
BUCKET ||--o{ OBJECT : "contains"
OBJECT {
string key
string metadata
byte[] data
timestamp created_at
}
BUCKET {
string name
string region
string access_policy
}- Example: AWS S3, Google Cloud Storage, Backblaze B2.
- Use Cases:
- Media hosting (YouTube, Netflix).
- Backup archives (eSewa’s transaction logs).
- Static website content (Daraz’s product images).
Worked Example: eSewa’s Transaction Logs
eSewa stores 10 million daily transactions in S3 with:
- Key:
txn_<timestamp>_<user_id>.json - Metadata:
{"amount": 5000, "status": "completed", "service": "electricity"} - Lifecycle Rule: Move to Glacier Deep Archive after 30 days (cost: $0.00099/GB/month).
Cost Calculation:
- 10M transactions × avg 1KB = 10TB/month.
- S3 Standard: $0.023/GB → $230/month.
- Glacier: $0.00099/GB → $9.90/month (96% savings).
2. Block Storage
Definition: Stores structured data (e.g., databases, VM disks) as fixed-size blocks (typically 4KB–1MB). Mounted as a volume to a VM or container.
How It Works
- Example: AWS EBS, Azure Disk Storage, Google Persistent Disk.
- Use Cases:
- Databases (MySQL on NEPSE’s trading platform).
- Boot volumes for VMs (Pathao’s backend servers).
- Enterprise apps (ERP systems in Nepali banks).
Comparison: Object vs. Block Storage
| Feature | Object Storage | Block Storage |
|---|---|---|
| Data Type | Unstructured (files, blobs) | Structured (blocks, disks) |
| Access Pattern | HTTP/REST API (GET/PUT) | Low-latency, high IOPS |
| Use Case | Backups, media, logs | Databases, VMs, OS disks |
| Scalability | Horizontal (add buckets) | Vertical (resize volume) |
| Cost | $0.02–$0.05/GB/month | $0.10–$0.50/GB/month |
3. File Storage
Definition: Provides shared file systems (e.g., NFS, SMB) accessible via standard protocols. Supports hierarchical directories and concurrent access.
How It Works
sequenceDiagram
participant User
participant Client as "NFS Client"
participant Server as "Cloud File Storage (e.g., AWS EFS)"
User->>Client: Open file.txt
Client->>Server: NFS Request (read)
Server-->>Client: File data
Client-->>User: Display content- Example: AWS EFS, Azure Files, Google Filestore.
- Use Cases:
- Collaborative editing (Google Docs-like apps).
- Development environments (shared code repos for Nepali startups).
- Media workflows (video editing teams).
Real-World Example: Kathmandu Traffic Management The Kathmandu Metropolitan City uses AWS EFS to share real-time traffic camera feeds across 50+ monitoring stations. Files are structured as:
/traffic/
├── 2023-10/
│ ├── camera_001_10-00-00.mp4
│ └── metadata.json
└── rules/
└── speed_limits.csv
Storage Architectures
Cloud providers use distributed architectures to ensure durability, availability, and performance.
1. Storage Classes
Providers offer tiered storage based on access frequency and cost. Example: AWS S3 classes:
mindmap
root((AWS S3 Storage Classes))
Standard
Frequent access, low latency
$0.023/GB
Intelligent-Tiering
Auto-moves data between tiers
$0.0225/GB + monitoring fee
Infrequent Access (IA)
$0.0125/GB + retrieval fee
Glacier
Archive (3–5 hrs retrieval)
$0.0036/GB
Glacier Deep Archive
Compliance (12+ hrs retrieval)
$0.00099/GBWorked Example: NEPSE’s Historical Data NEPSE stores 20 years of stock market data (50TB) in S3 Glacier Deep Archive:
- Cost: 50TB × $0.00099/GB/month = $49.50/month (vs. $1,150 in Standard).
- Retrieval: Query via AWS Athena (serverless SQL) when analysts need old trades.
2. Redundancy and Replication
Cloud storage ensures durability via:
- Geographic Replication: Data copied across Availability Zones (AZs) or Regions.
- Erasure Coding: Splits data into fragments + parity bits (e.g., 10+2 erasure coding = 10% overhead for 2x fault tolerance).
- Versioning: Retains multiple versions of objects (e.g.,
file.txt?v=1,file.txt?v=2).
Example: Google Cloud Storage uses dual-copy replication within a region and multi-regional replication for global access.
3. Caching Layers
To reduce latency, cloud storage uses:
- CDNs (Content Delivery Networks): Edge caches for static content (e.g., Daraz’s product images).
- In-Memory Caches: Redis/Memcached for frequently accessed data (e.g., Ncell’s subscriber lookup tables).
sequenceDiagram
participant User
participant CDN as "Cloudflare (Edge Cache)"
participant Origin as "AWS S3 (Origin)"
User->>CDN: Request image.jpg
CDN-->>User: Serve from cache (90% hits)
User->>CDN: Request new_image.jpg
CDN->>Origin: Fetch missing file
Origin-->>CDN: Return file
CDN-->>User: Serve + cacheCloud Storage Technologies
1. Distributed File Systems
- HDFS (Hadoop Distributed File System): Used by Google Cloud Storage and AWS S3 for large-scale analytics.
- Ceph: Open-source, used by OpenStack Swift.
- Google File System (GFS): Inspired Colossus (Google’s storage backbone).
2. Object Storage Backends
- Ceph RADOS: Used by OpenStack Swift.
- MinIO: Open-source S3-compatible storage (used by Nepali startups for cost savings).
- Azure Blob Storage: Uses Azure Data Lake Storage Gen2 for hierarchical namespace.
3. Hybrid and Edge Storage
- AWS Outposts: Extends S3/EBS to on-premises data centers (used by Nepali banks for compliance).
- Google Distributed Cloud: Runs Anthos on-premises with cloud storage integration.
- Edge Caching: Cloudflare Workers cache data at 300+ locations globally.
Security in Cloud Storage
1. Data Encryption
- At Rest: AES-256 (e.g., AWS S3 Server-Side Encryption).
- In Transit: TLS 1.2+ (e.g., HTTPS for S3 API).
- Customer-Managed Keys: AWS KMS, Google Cloud KMS, Azure Key Vault.
2. Access Control
- IAM Policies: Fine-grained permissions (e.g.,
s3:GetObjectfor read-only access). - Bucket Policies: Restrict access by IP/region (e.g., block public access to NEPSE’s trading data).
- VPC Endpoints: Private network access to S3 (avoids public internet exposure).
3. Compliance
- GDPR: Right to erasure (e.g., Khalti must delete user data on request).
- HIPAA: Encryption for healthcare data (e.g., CIAA hospitals using Azure Blob Storage).
- PCI DSS: Tokenization for payment data (e.g., eSewa’s credit card storage).
Performance Optimization
1. Latency Reduction
- Region Selection: Store data closer to users (e.g.,
ap-south-1for Nepal vs.us-east-1). - Transfer Acceleration: AWS S3 Transfer Acceleration (uses CloudFront edge locations).
2. Bandwidth Management
- Compression: Gzip/Brotli for text data (e.g., JSON logs).
- Chunked Uploads: Parallel uploads for large files (e.g., 5GB video to Google Cloud Storage).
3. Lifecycle Management
Automate data transitions:
- Hot → Cool: Move infrequently accessed data to S3 IA.
- Cool → Archive: Transition to Glacier after 90 days.
- Archive → Delete: Expire old logs after 7 years (compliance).
stateDiagram-v2
[*] --> Hot
Hot --> Cool: After 30 days of inactivity
Cool --> Glacier: After 90 days of inactivity
Glacier --> DeepArchive: After 1 year
DeepArchive --> [*]: After 7 years (retention policy)In the Real World
eSewa’s Transaction Processing
- Technology: AWS S3 (object storage) + DynamoDB (metadata).
- How It Works:
- Each transaction is stored as a JSON object in S3 with a key like
txn_20231001_123456.json. - Lifecycle Rule: Moves to Glacier after 30 days.
- Security: Encrypted with AWS KMS and restricted via IAM roles.
- Each transaction is stored as a JSON object in S3 with a key like
- Impact: Handles 50,000 transactions/hour during festivals (e.g., Dashain).
Daraz’s Inventory Management
- Technology: AWS EFS (file storage) + S3 (product images).
- How It Works:
- Product catalog stored in EFS (shared across 100+ servers).
- High-res images in S3 with CloudFront CDN for global delivery.
- Cost Optimization: Uses S3 Intelligent-Tiering to auto-move old images to cheaper storage.
- Impact: Reduces image load time from 2s → 200ms via CDN.
NEPSE’s Trading Platform
- Technology: Google Cloud Storage (for historical data) + Firestore (real-time trades).
- How It Works:
- Real-time trades: Stored in Firestore (NoSQL database).
- Historical data (20+ years): Archived in Coldline Storage ($0.003/GB).
- Disaster Recovery: Data replicated across asia-southeast1 (Singapore) and europe-west1 (Belgium).
- Impact: Ensures 99.999% uptime during market crashes.
Exam Tip
What Examiners Look For
- Definitions: Clearly distinguish object, block, and file storage (1 mark each).
- Worked Examples: Show cost calculations (e.g., S3 vs. Glacier) or lifecycle policies (2 marks).
- Diagrams: Draw storage class transitions (state diagram) or erasure coding (1 mark).
- Real-World Links: Relate to Nepali companies (eSewa, Daraz, NEPSE) for application-based questions (3 marks).
- Security: Mention encryption (AES-256), IAM policies, and compliance (GDPR/HIPAA) (2 marks).
Common Pitfalls
- Confusing Block vs. Object: Block storage is for databases/VMs; object is for files/media.
- Ignoring Costs: Always calculate GB/month × rate for storage classes.
- Overlooking Redundancy: Assume multi-AZ replication unless stated otherwise.
- Vague Answers: Instead of "cloud storage is scalable," say:
"AWS S3 auto-scales to petabytes by distributing data across 100+ servers in a region, ensuring <99.99% availability even during traffic spikes like Daraz’s Diwali sales."
Sample Exam Question & Answer
Question: "Explain how Ncell could use cloud storage to reduce costs for storing 1PB of customer call logs while ensuring compliance with Nepali data laws. Include a storage class recommendation and security measures."
Model Answer: Ncell can optimize storage costs and compliance using AWS S3 with the following strategy:
Storage Class:
- Hot Storage (S3 Standard): For recent logs (≤30 days) accessed frequently for fraud detection.
- Cost: 100TB × $0.023/GB/month = $2,300/month.
- Cool Storage (S3 IA): For logs 30–90 days old (infrequent access).
- Cost: 500TB × $0.0125/GB/month = $625/month.
- Archive (Glacier Deep Archive): For logs >90 days old (compliance retention).
- Cost: 400TB × $0.00099/GB/month = $39.60/month.
- Total Cost: $2,964.60/month (vs. $23,000 in on-prem HDDs).
- Hot Storage (S3 Standard): For recent logs (≤30 days) accessed frequently for fraud detection.
Security Measures:
- Encryption: Enable S3 Server-Side Encryption (SSE-S3) with AES-256 for all buckets.
- Access Control:
- Restrict IAM roles to least privilege (e.g.,
s3:GetObjectfor analytics teams). - Use VPC Endpoints to avoid public internet exposure.
- Restrict IAM roles to least privilege (e.g.,
- Compliance:
- Enable S3 Object Lock to prevent deletion for 7 years (as per Nepali telecom regulations).
- Store logs in Asia-Pacific (Mumbai) region to comply with Nepali data localization laws.
Performance:
- Use S3 Intelligent-Tiering to auto-move logs between classes.
- For real-time fraud detection, cache hot logs in Amazon ElastiCache (Redis).
Visual Aid:
stateDiagram-v2
[*] --> Hot: "≤30 days (S3 Standard)"
Hot --> Cool: "After 30 days (S3 IA)"
Cool --> Archive: "After 90 days (Glacier)"
Archive --> [*]: "After 7 years (retention policy)"Based on the TU BIM syllabus for Cloud Computing (IT277), unit 6.
Discussion
Loading…