Big Data and Analytics notes

10 chapter notes, in syllabus order. Each starts with the key points.

Unit 1

Big Data: Definitions, Characteristics, Sources & ChallengesUnit 1 of Big Data and Analytics introduces core concepts: what big data is, its 5Vs (volume, velocity, variety, veracity, value), sources (structured/unstructured), challenges (storage, processing, privacy), and real-world applications in Nepalese and global tech ecosystems.9 min read

Unit 2

Big Data Architecture: Layers, Models, and ComponentsUnit 2 of Big Data and Analytics explores the foundational architecture of big data systems, covering the Lambda architecture, Kappa architecture, data lakes, data warehouses, and distributed storage/compute frameworks. It explains how these components interact to process, store, and analyze massive datasets efficientl7 min read

Unit 3

Hadoop & HDFS: Architecture, HDFS Design, Data Storage & Fault ToleranceUnit 3 of Big Data and Analytics explores Hadoop’s ecosystem, HDFS architecture (NameNode, DataNode, replication), data storage mechanics, and fault tolerance mechanisms like replication and rack awareness, with real-world applications in Nepal’s eSewa and Ncell systems.11 min read

Unit 4

MapReduce: Model, Phases, Use Cases & OptimizationUnit 4 of Big Data and Analytics explores MapReduce—its core architecture, the map and reduce phases, data partitioning, combiners, and real-world deployments (e.g., Hadoop). Learn how it processes petabytes of data across clusters, with comparisons to Spark and worked examples like log analysis or Ncell call-detail re9 min read

Unit 5

NoSQL Databases: Types, Models, Use Cases & ComparisonUnit 5 of Big Data and Analytics explores NoSQL databases—how they differ from SQL, their four core data models (document, key-value, column-family, graph), real-world applications in Nepalese and global tech, and when to use them over traditional relational databases.12 min read

Unit 6

Apache Spark: Architecture, Components & Big Data ProcessingUnit 6 of Big Data and Analytics covers Apache Spark’s architecture, core components (Spark Core, Spark SQL, MLlib, GraphX, and Spark Streaming), its in-memory processing model, and how it outperforms Hadoop MapReduce for iterative and real-time analytics. Includes real-world use cases, performance comparisons, and han8 min read

Unit 7

Data Ingestion & Stream Processing: Pipelines, Tools & Real-Time AnalyticsUnit 7 of Big Data and Analytics explores how to ingest massive data streams (batch vs. real-time), key tools like Kafka and Flume, stream processing frameworks (Spark Streaming, Flink), and their applications in fraud detection, IoT, and social media analytics—with Nepalese examples like Ncell’s network monitoring and5 min read

Unit 8

Big Data Analytics: Techniques, Tools & ApplicationsUnit 8 of Big Data and Analytics explores core techniques for extracting insights from massive datasets, including descriptive, predictive, and prescriptive analytics, along with tools like regression, clustering, classification, and association rules. It covers real-world applications in Nepal (e.g., NTC’s network opt11 min read

Unit 9

Machine Learning on Big Data: Algorithms, Scalability & ApplicationsUnit 9 of Big Data and Analytics explores how traditional machine learning techniques are adapted for big data—scalable algorithms (e.g., stochastic gradient descent, ensemble methods), distributed training frameworks (Spark MLlib), and real-world use cases in fraud detection, recommendation systems, and predictive ana9 min read

Unit 10

Big Data Visualization & Applications: Tools, Techniques & Real-World ImpactUnit 10 of Big Data and Analytics explores how to transform raw big data into actionable insights through visualization techniques, tools, and real-world applications—covering dashboards, geospatial analytics, and industry use cases from eSewa to global tech giants.8 min read