Big Data and AnalyticsUnit 19 min read
Big Data: Definitions, Characteristics, Sources & Challenges
Unit 1 of Big Data and Analytics introduces core concepts: what big data is, its 5Vs (volume, velocity, variety, veracity, value), sources (structured/unstructured), challenges (storage, processing, privacy), and real-world applications in Nepalese and global tech ecosystems.
What is Big Data?
Big data refers to extremely large and complex datasets that traditional data-processing tools cannot handle efficiently. It is not just about size—it’s about extracting meaningful insights from data that grows exponentially over time.
Why is it called "Big"?
The term is defined by the 5Vs framework (introduced by IBM and Gartner), which categorizes its key attributes:
| Attribute | Definition | Example |
|---|---|---|
| Volume | Massive scale of data (petabytes to exabytes). | Ncell processes 100+ TB/day of call logs, SMS, and internet traffic. |
| Velocity | Speed at which data is generated and processed. | Pathao receives 10,000+ ride requests per minute during peak hours. |
| Variety | Diverse data types (structured, semi-structured, unstructured). | eSewa handles transaction logs (structured), user reviews (text), and payment images (unstructured). |
| Veracity | Data accuracy, reliability, and quality. | NEPSE stock data must be 99.9% accurate to avoid trading errors. |
| Value | Potential business insights from data. | Daraz uses customer browsing data to predict demand and optimize inventory. |
Sources of Big Data
Data originates from three primary sources:
1. Structured Data
- Definition: Organized, tabular data (e.g., databases, spreadsheets).
- Examples:
- Bank transaction records (e.g., Nabil Bank loan applications).
- NTC’s monthly internet usage reports.
- Format: SQL databases, Excel sheets.
- Visual:
flowchart TD A["Structured Data"] --> B["SQL Databases\n(e.g., MySQL, PostgreSQL)"] A --> C["Spreadsheets\n(e.g., Excel, Google Sheets)"] A --> D["Relational Tables\n(e.g., Customer ID, Transaction Date)"]
2. Semi-Structured Data
- Definition: Data with no fixed schema but some organizational tags (e.g., JSON, XML).
- Examples:
- WhatsApp message logs (timestamp, sender, media type).
- Khalti transaction metadata (user ID, amount, status).
- Format: JSON, XML, NoSQL databases.
- Visual:
flowchart TD A["Semi-Structured Data"] --> B["JSON\n(e.g., {\"user\": \"ABC123\", \"amount\": 500})"] A --> C["XML\n(e.g., <transaction><date>2024-05-20</date>)</transaction>"] A --> D["NoSQL\n(e.g., MongoDB collections)"]
3. Unstructured Data
- Definition: Data without predefined format (e.g., text, images, videos).
- Examples:
- YouTube comments and video content.
- Facebook posts, memes, and live streams.
- Nepali news websites (e.g., Kantipur, Republica) articles and reader comments.
- Format: Text, images, audio, video.
- Visual:
Challenges in Handling Big Data
Despite its value, big data presents technical and ethical hurdles:
Technical Challenges
| Challenge | Description | Example |
|---|---|---|
| Storage | Requires scalable infrastructure (e.g., cloud storage). | Google uses 100+ petabytes of storage for YouTube and Gmail. |
| Processing | Needs high-performance tools (e.g., Hadoop, Spark). | Ncell processes 100M+ call records daily using distributed systems. |
| Analysis | Complex algorithms needed for pattern recognition. | Daraz uses machine learning to detect fraudulent orders. |
| Privacy & Security | Risk of data breaches and misuse. | Khalti encrypts transactions but faces phishing attacks targeting user credentials. |
Ethical Challenges
- Data Privacy: Laws like Nepal’s Data Privacy Act (2018) regulate how companies (e.g., eSewa) handle user data.
- Bias in Algorithms: AI models trained on biased data (e.g., NEPSE stock predictions) can reinforce inequalities.
- Misuse: Governments or corporations may exploit data for surveillance (e.g., NTC’s internet throttling during exams).
Real-World Applications in Nepal & Globally
1. eSewa & Digital Payments
- Idea Used: Velocity & Variety
- How?
- Processes 50,000+ transactions per minute during festivals (e.g., Dashain, Tihar).
- Handles structured (transaction IDs, amounts) and unstructured (user complaints in chat) data.
- Visual:
2. Pathao & Ride-Sharing Optimization
- Idea Used: Velocity & Value
- How?
- Uses real-time GPS data (velocity) to match drivers and riders.
- Predicts peak demand (value) to deploy more drivers in Kathmandu traffic hotspots (e.g., Thapathali, Lakshmi Marg).
- Visual:
flowchart TD A["Pathao App"] --> B["User Request\n(Location, Time)"] B --> C["Real-Time GPS\n(Google Maps API)"] C --> D["Driver Matching\n(Algorithm)"] D --> E["Route Optimization\n(Avoiding Traffic Jams)"]
3. NEPSE & Stock Market Predictions
- Idea Used: Veracity & Value
- How?
- Analyzes historical stock data (structured) and news sentiment (unstructured) to predict trends.
- Example: If Kantipur publishes a positive article about a company, NEPSE’s algorithms may increase its predicted stock value.
- Visual:
Worked Example: Calculating Data Growth in NTC
Scenario: NTC’s internet users grew from 5M (2018) to 20M (2024). If data usage per user increased from 5GB/month to 20GB/month, calculate the total data volume in 2024 and classify it using the 5Vs.
Solution:
- Total Users (2024): 20M
- Data per User (2024): 20GB/month
- Total Monthly Data:
- 5Vs Classification:
- Volume: 400TB/month (high volume).
- Velocity: Data grows ~30% annually (fast velocity).
- Variety: Includes web browsing, videos, VoIP calls (diverse types).
- Veracity: Must ensure no packet loss in NTC’s backbone network.
- Value: Helps NTC optimize bandwidth allocation and predict peak hours.
Big Data vs. Traditional Data Processing
| Feature | Big Data | Traditional Data |
|---|---|---|
| Scale | Petabytes/exabytes (e.g., Google’s 20M+ servers). | Gigabytes/terabytes (e.g., Excel spreadsheets). |
| Tools | Hadoop, Spark, NoSQL. | SQL databases, Excel, R. |
| Processing Speed | Real-time (e.g., Pathao’s ride matching). | Batch processing (e.g., monthly bank reports). |
| Use Case | Fraud detection, personalized ads, traffic prediction. | Inventory management, simple analytics. |
| Cost | High (cloud infrastructure). | Low (on-premise servers). |
Exam Tip
- Define Big Data using the 5Vs—examiners love this framework.
- Compare structured vs. unstructured data with Nepali examples (e.g., NEPSE vs. Facebook comments).
- Explain challenges with real-world fixes:
- Challenge: Ncell’s call data is too large.
- Solution: Use Hadoop for distributed storage.
- Link theory to applications:
- If asked about velocity, discuss eSewa’s transaction speed.
- If asked about veracity, discuss NEPSE’s data accuracy needs.
- Avoid vague answers—always tie concepts to Nepali tech companies (e.g., Khalti, Daraz, NTC).
Key Formula to Remember: (Higher numerator + lower denominator = more valuable insights.)
Based on the TU BITM syllabus for Big Data and Analytics (IT278), unit 1.
Discussion
Loading…