Cloud ComputingUnit 615 min read
Cloud Storage: Types, Models, Security & Real-World Use
Unit 6 of Cloud Computing explores cloud storage architectures, data redundancy, storage-as-a-service models, security best practices, and real-world applications like eSewa’s transaction logs and Ncell’s customer data. Learn how data is stored, accessed, and protected in cloud environments, with comparisons to traditi
Cloud Storage Fundamentals
What is Cloud Storage?
Cloud storage is a service model that allows users to store, manage, and retrieve data over the internet using third-party cloud providers instead of local storage devices (like hard drives or SSDs). It leverages distributed storage systems, data replication, and scalable infrastructure to ensure availability, durability, and accessibility.
Key Characteristics:
- On-demand scalability: Storage capacity can be increased or decreased dynamically.
- Multi-tenancy: Multiple users or organizations share the same physical storage infrastructure.
- Geographic distribution: Data is replicated across multiple data centers for redundancy.
- Pay-as-you-go pricing: Users pay only for the storage they use.
How Cloud Storage Works
Cloud storage relies on a three-tier architecture:
- Client Tier: Users or applications that request storage services (e.g., uploading a file to Google Drive).
- Storage Tier: Physical or virtual storage systems (e.g., hard drives, SSDs, or object storage like Amazon S3).
- Interface Tier: APIs, web portals, or command-line tools that interact with the storage system.
Data Storage Models:
Cloud storage typically uses one of these models:
- Block Storage: Data is split into fixed-size blocks (e.g., 4KB) and stored independently. Used for databases and boot volumes.
- Example: AWS EBS (Elastic Block Store).
- File Storage: Hierarchical file systems (folders, subfolders) with metadata (e.g., permissions, timestamps).
- Example: Google Drive, Dropbox.
- Object Storage: Data is stored as objects with unique identifiers (keys) and metadata. Scalable and used for unstructured data.
- Example: Amazon S3, Azure Blob Storage.
Cloud Storage Service Models
Cloud storage is often delivered as part of broader cloud service models (IaaS, PaaS, SaaS), but it can also be offered as a standalone service. Here’s how it fits into the cloud service hierarchy:
| Service Model | Cloud Storage Role | Example Providers |
|---|---|---|
| IaaS | Block or file storage as a virtual resource. | AWS EBS, Azure Disk Storage |
| PaaS | Managed storage for applications (e.g., databases). | Google Cloud SQL, Azure SQL Database |
| SaaS | Built-in storage for end-user applications. | Google Workspace, Microsoft 365 |
| Standalone | Dedicated storage service for backups or archives. | Backblaze B2, Wasabi Hot Storage |
In the Real World
eSewa’s Transaction Logs: eSewa uses object storage (like AWS S3) to store transaction records securely. Each transaction is stored as an immutable object with metadata (timestamp, user ID, amount). This ensures compliance with Nepal’s financial regulations and allows quick retrieval for audits.
- Why object storage? Scalability (millions of transactions/day) and durability (data never lost due to replication).
Ncell’s Customer Data: Ncell’s CRM system relies on block storage (AWS EBS) for real-time customer data (e.g., call logs, billing records). The data is replicated across multiple Availability Zones in AWS to prevent downtime during peak usage (e.g., during festivals like Dashain).
- Why block storage? Low-latency access for database operations (e.g., updating a customer’s plan).
Daraz’s Product Catalog: Daraz uses Google Cloud Storage (GCS) to host product images, descriptions, and inventory data. During sales events (e.g., 11.11), the system scales storage dynamically to handle spikes in uploads (e.g., new product listings).
- Why object storage? Cost-efficiency for static assets (images, videos) and easy CDN integration.
Worked Example: NTC’s Network Log Storage
Scenario: The Nepal Telecommunications Corporation (NTC) needs to store 1TB of network logs daily for compliance. They choose Azure Blob Storage (object storage) with the following configuration:
- Storage Tier: Hot storage (frequently accessed logs) + Cool storage (archived logs older than 30 days).
- Redundancy: Geo-redundant storage (GRS) to replicate data across Nepal and Singapore.
- Lifecycle Policy: Automatically move logs to Cool storage after 30 days, then archive to tape after 1 year.
Cost Calculation:
- Hot storage: $0.02/GB/month → 1TB/month = $20/month.
- Cool storage: $0.01/GB/month → 900GB/month (after 30 days) = $9/month.
- Data transfer: $0.05/GB for cross-region replication → 1TB/month = $50/month.
- Total: ~$79/month (scalable for future growth).
Cloud Storage Architectures
Cloud storage systems use distributed architectures to ensure high availability and fault tolerance. Here’s a breakdown:
1. Distributed File Systems
These systems spread data across multiple servers while presenting a unified file system to users. Examples:
- Google File System (GFS): Used by Google for large-scale data storage (e.g., YouTube videos).
- Hadoop Distributed File System (HDFS): Used for big data analytics (e.g., Facebook’s data processing).
How it works:
- The NameNode tracks metadata (file locations, permissions).
- DataNodes store actual data blocks (typically 128MB–1GB each).
- Data is replicated 3x by default for fault tolerance.
2. Object Storage Architectures
Object storage systems store data as objects (key-value pairs) with metadata. Key components:
- Object: Data + metadata (e.g.,
user_123_profile.jpgwithsize=1MB,upload_date=2024-05-01). - Bucket: A container for objects (analogous to a folder).
- API Gateway: Handles HTTP requests (e.g.,
PUT,GET,DELETE).
Example: Amazon S3
- Consistency Model: Strong consistency for
PUT/DELETE, eventual consistency forGETafter aPUT. - Storage Classes:
- S3 Standard: High durability (99.999999999%), low latency.
- S3 IA (Infrequent Access): Cheaper for rarely accessed data.
- S3 Glacier: Ultra-cheap archival storage (retrieval takes hours).
3. Hybrid Storage Architectures
Combine on-premises storage with cloud storage for performance and cost optimization. Example:
- Nepal Rastra Bank’s Data Center:
- On-premises: High-performance SSDs for real-time transaction processing.
- Cloud (AWS): Cold storage for archived financial records (e.g., 10-year-old loan data).
- Sync: Data is replicated nightly using AWS Storage Gateway.
graph LR
A["On-Premises\n(SSD Arrays)"] -->|"Nightly Sync"| B["AWS Storage Gateway"]
B --> C["AWS S3\n(Cold Storage)"]
D["Cloud App\n(e.g., Analytics)"] -->|"Direct Access"| CCloud Storage Security
Security is critical in cloud storage due to the shared responsibility model (provider secures infrastructure; user secures data). Key mechanisms:
1. Data Encryption
- At Rest: Data encrypted on storage devices (e.g., AES-256).
- Example: Google Cloud encrypts data by default using customer-supplied keys (CMEK).
- In Transit: TLS/SSL for data moving between client and storage.
- Client-Side: Users encrypt data before uploading (e.g., using AWS KMS).
2. Access Control
- IAM (Identity and Access Management): Role-based access (e.g.,
storage-admin,read-only). - Bucket Policies: Define who can access a bucket (e.g., only
user@example.com). - Temporary Credentials: Short-lived tokens (e.g., AWS STS).
3. Compliance and Auditing
- Standards: HIPAA (healthcare), GDPR (EU data), PCI-DSS (payments).
- Logging: Track access (e.g., AWS CloudTrail logs all S3 API calls).
- Immutable Storage: Prevent deletion/modification (e.g., AWS S3 Object Lock).
Worked Example: eSewa’s Secure Transaction Storage
Requirements:
- Store transaction data with immutability (for audits).
- Encrypt data at rest and in transit.
- Restrict access to only eSewa’s compliance team.
Solution:
- Storage: AWS S3 with Object Lock (WORM mode) to prevent deletions.
- Encryption:
- At rest: AWS KMS with customer-managed keys.
- In transit: TLS 1.3 for all API calls.
- Access Control:
- IAM role
eSewa-Compliance-ReadOnlywith least-privilege permissions. - Bucket policy to allow only
esewa.comdomain access.
- IAM role
- Audit Trail: Enable AWS CloudTrail to log all
PutObjectandGetObjectevents.
Cost:
- S3 Standard: $0.023/GB → 500GB/month = $11.50.
- KMS: $1/month for key management.
- CloudTrail: $0.10/GB logged → 10GB/month = $1.
- Total: ~$13.50/month (scalable).
Cloud Storage Performance Optimization
Performance depends on latency, throughput, and consistency. Techniques to optimize:
| Technique | Use Case | Example |
|---|---|---|
| Caching | Reduce latency for frequently accessed data. | Amazon CloudFront (CDN caching). |
| Sharding | Distribute load across multiple storage nodes. | MongoDB’s sharded clusters. |
| Erasure Coding | Reduce storage overhead for replication. | Ceph (uses 6+3 erasure coding). |
| Tiered Storage | Move cold data to cheaper storage. | AWS S3 Lifecycle Policies. |
Worked Example: Pathao’s Ride Data Storage
Scenario: Pathao needs to store 10TB of ride data daily (GPS logs, driver info) with low latency for real-time analytics. Solution:
- Hot Storage: AWS S3 Standard for recent data (<30 days).
- Cold Storage: S3 Glacier Deep Archive for older data (>1 year).
- Caching: Amazon ElastiCache (Redis) for frequently queried driver locations.
- Sharding: Data partitioned by
ride_idacross 10 S3 buckets.
Performance Metrics:
- Read Latency: 10–50ms (cached data), 100–300ms (S3).
- Write Throughput: 10GB/s (using S3 Transfer Acceleration).
- Cost: ~$200/month (scalable for 1M daily rides).
Cloud Storage vs. Traditional Storage
Compare cloud storage to on-premises storage (e.g., NAS, SAN) using key metrics:
| Metric | Cloud Storage | Traditional Storage |
|---|---|---|
| Scalability | Infinite (pay-as-you-go). | Limited by hardware capacity. |
| Maintenance | Fully managed by provider. | User responsible for hardware/software. |
| Cost | Pay for what you use (no upfront CAPEX). | High upfront cost (servers, racks). |
| Availability | 99.99%–99.999% SLA (multi-region). | Depends on local infrastructure. |
| Disaster Recovery | Built-in (geo-replication). | Requires manual setup (e.g., offsite backups). |
| Accessibility | Global access via internet. | Limited to on-premises or VPN. |
| Security | Shared responsibility model. | Fully user’s responsibility. |
Cloud Storage Providers Comparison
Major providers offer similar services but with different strengths:
| Provider | Key Offerings | Best For | Pricing Model |
|---|---|---|---|
| AWS | S3, EBS, EFS, Glacier. | Enterprises needing global reach. | Pay-per-GB + data transfer fees. |
| Azure | Blob Storage, Azure Files, Cool Storage. | Microsoft ecosystem (e.g., .NET apps). | Tiered pricing (Hot/Cool/Archive). |
| Google Cloud | Cloud Storage, Persistent Disk, Coldline. | AI/ML workloads (integrates with BigQuery). | Nearline/Coldline for archival. |
| IBM Cloud | Cloud Object Storage, File Storage. | Hybrid cloud (on-prem + cloud). | Predictable pricing for reserved capacity. |
| Oracle Cloud | Object Storage, Block Volume. | Oracle database workloads. | Flexible pricing (hourly/daily). |
Worked Example: Kathmandu Traffic Route Optimization
Scenario: The Kathmandu Metropolitan City (KMC) wants to store 5TB of traffic camera data monthly for AI-based congestion analysis. Solution:
- Provider: Google Cloud Storage (GCS) for cost-efficiency.
- Storage Classes:
- Hot Storage: First 30 days (for real-time analysis).
- Coldline: Next 6 months (historical data).
- Archive: Data older than 1 year.
- Cost:
- Hot: 5TB × $0.02/GB = $100/month.
- Coldline: 4TB × $0.01/GB = $40/month.
- Archive: 1TB × $0.004/GB = $4/month.
- Total: ~$144/month (vs. $500/month for on-premises NAS).
Exam Tip
For the Cloud Storage unit in TU exams, focus on these high-yield areas:
- Definitions and Models:
- Differentiate block, file, and object storage with examples (e.g., EBS vs. S3).
- Explain distributed file systems (GFS, HDFS) and their components (NameNode, DataNode).
- Architectures:
- Draw and label a cloud storage architecture diagram (client → API → storage nodes).
- Compare hybrid storage (on-prem + cloud) with pure cloud.
- Security:
- List 3 encryption methods (at rest, in transit, client-side) and where each is used.
- Describe IAM roles and bucket policies with a real-world example (e.g., eSewa’s access control).
- Performance Optimization:
- Explain caching, sharding, and tiered storage with a worked example (e.g., Pathao’s ride data).
- Cost Calculations:
- Practice pricing scenarios (e.g., "Calculate the monthly cost for storing 10TB in AWS S3 Standard").
- Real-World Applications:
- Relate concepts to Nepali companies (e.g., NTC’s logs, Daraz’s product catalog) or global platforms (e.g., YouTube’s GFS).
- Diagrams:
- Be ready to draw:
- A distributed file system (NameNode + DataNodes).
- An object storage workflow (client → API → bucket → object).
- A hybrid storage setup (on-prem + cloud).
- Be ready to draw:
Common Pitfalls:
- Confusing SaaS storage (e.g., Google Drive) with IaaS storage (e.g., AWS EBS).
- Forgetting to mention replication or geo-redundancy in high-availability scenarios.
- Overlooking cost factors (e.g., data transfer fees, storage class differences).
Based on the TU BITM syllabus for Cloud Computing (IT277), unit 6.
Discussion
Loading…