IT278 Big Data and Analytics
Big Data and Analytics notes
10 chapter notes, in syllabus order. Each starts with the key points.
Unit 1
Big Data: Definitions, Characteristics, Sources & ChallengesUnit 1 of Big Data and Analytics introduces core concepts: what big data is, its 5Vs (volume, velocity, variety, veracity, value), sources (structured/unstructured), challenges (storage, processing, privacy), and real-world applications in Nepalese and global tech ecosystems.9 min readUnit 2
Big Data Architecture: Layers, Models, and ComponentsUnit 2 of Big Data and Analytics explores the foundational architecture of big data systems, covering the Lambda architecture, Kappa architecture, data lakes, data warehouses, and distributed storage/compute frameworks. It explains how these components interact to process, store, and analyze massive datasets efficientl7 min readUnit 3
Hadoop & HDFS: Architecture, HDFS Design, Data Storage & Fault ToleranceUnit 3 of Big Data and Analytics explores Hadoop’s ecosystem, HDFS architecture (NameNode, DataNode, replication), data storage mechanics, and fault tolerance mechanisms like replication and rack awareness, with real-world applications in Nepal’s eSewa and Ncell systems.11 min readUnit 4
MapReduce: Model, Phases, Use Cases & OptimizationUnit 4 of Big Data and Analytics explores MapReduce—its core architecture, the map and reduce phases, data partitioning, combiners, and real-world deployments (e.g., Hadoop). Learn how it processes petabytes of data across clusters, with comparisons to Spark and worked examples like log analysis or Ncell call-detail re9 min readUnit 5
NoSQL Databases: Types, Models, Use Cases & ComparisonUnit 5 of Big Data and Analytics explores NoSQL databases—how they differ from SQL, their four core data models (document, key-value, column-family, graph), real-world applications in Nepalese and global tech, and when to use them over traditional relational databases.12 min readUnit 6
Apache Spark: Architecture, Components & Big Data ProcessingUnit 6 of Big Data and Analytics covers Apache Spark’s architecture, core components (Spark Core, Spark SQL, MLlib, GraphX, and Spark Streaming), its in-memory processing model, and how it outperforms Hadoop MapReduce for iterative and real-time analytics. Includes real-world use cases, performance comparisons, and han8 min readUnit 7
Data Ingestion & Stream Processing: Pipelines, Tools & Real-Time AnalyticsUnit 7 of Big Data and Analytics explores how to ingest massive data streams (batch vs. real-time), key tools like Kafka and Flume, stream processing frameworks (Spark Streaming, Flink), and their applications in fraud detection, IoT, and social media analytics—with Nepalese examples like Ncell’s network monitoring and5 min readUnit 8
Big Data Analytics: Techniques, Tools & ApplicationsUnit 8 of Big Data and Analytics explores core techniques for extracting insights from massive datasets, including descriptive, predictive, and prescriptive analytics, along with tools like regression, clustering, classification, and association rules. It covers real-world applications in Nepal (e.g., NTC’s network opt11 min readUnit 9
Machine Learning on Big Data: Algorithms, Scalability & ApplicationsUnit 9 of Big Data and Analytics explores how traditional machine learning techniques are adapted for big data—scalable algorithms (e.g., stochastic gradient descent, ensemble methods), distributed training frameworks (Spark MLlib), and real-world use cases in fraud detection, recommendation systems, and predictive ana9 min readUnit 10
Big Data Visualization & Applications: Tools, Techniques & Real-World ImpactUnit 10 of Big Data and Analytics explores how to transform raw big data into actionable insights through visualization techniques, tools, and real-world applications—covering dashboards, geospatial analytics, and industry use cases from eSewa to global tech giants.8 min read