IT278 Big Data and Analytics
Big Data and Analytics notes
9 chapter notes, in syllabus order. Each starts with the key points.
Unit 1
Big Data Basics: Definitions, Characteristics, and Real-World ImpactUnit 1 of Big Data and Analytics introduces the core concepts of big data—its 5Vs (Volume, Velocity, Variety, Veracity, and Value), sources, challenges, and applications in modern industries. This note covers definitions, real-world examples, and how big data transforms decision-making in Nepal and globally.13 min readUnit 2
Big Data Architecture: Layers, Models, and Real-World SystemsUnit 2 of Big Data and Analytics explores the foundational architecture of big data systems, covering the Lambda architecture, Kappa architecture, data lakes vs. data warehouses, distributed storage, and processing layers (batch vs. real-time). It explains how companies like Google, Facebook, and Ncell design scalable 9 min readUnit 3
Hadoop & HDFS: Architecture, HDFS Design, Data Flow & Fault ToleranceUnit 3 of Big Data and Analytics explores Hadoop’s distributed ecosystem and HDFS (Hadoop Distributed File System), covering its architecture, components (NameNode, DataNode), replication, data flow, and real-world applications like eSewa’s transaction logs and Ncell’s call detail records.11 min readUnit 4
MapReduce: Model, Phases, Use Cases & OptimizationUnit 4 of Big Data and Analytics explores MapReduce—a scalable programming model for processing vast datasets across distributed clusters. This note covers its architecture, phases (map, shuffle, reduce), real-world applications (e.g., Google search indexing), and optimizations like combiners and partitioning, with vis7 min readUnit 5
NoSQL Databases: Types, Models, and Use CasesUnit 5 of Big Data and Analytics explores NoSQL databases—how they differ from SQL, their four core data models (document, key-value, column-family, graph), real-world applications in Nepal (e.g., eSewa’s transaction logs, Daraz’s inventory scaling), and trade-offs in scalability, flexibility, and consistency. Includes12 min readUnit 6
Apache Spark: Architecture, Components, and ApplicationsUnit 6 of Big Data and Analytics explores Apache Spark, its architecture, core components (RDDs, DataFrames, Spark SQL), execution model, and real-world applications in analytics, machine learning, and stream processing. This note covers how Spark differs from Hadoop MapReduce, its distributed computing model, and hand8 min readUnit 7
Data Ingestion & Stream Processing: Pipelines, Tools & Real-Time AnalyticsUnit 7 of Big Data and Analytics explores how raw data is collected, processed in real-time, and stored for analytics. Learn about batch vs. stream processing, ingestion tools (Kafka, Flume), stream processing frameworks (Spark Streaming, Flink), and real-world applications like fraud detection in eSewa or traffic rout15 min readUnit 8
Big Data Analytics: Techniques, Tools & Real-World ImpactUnit 8 of Big Data and Analytics explores core techniques for extracting insights from massive datasets—descriptive, predictive, prescriptive analytics, and their tools (SQL, R, Python, Tableau). Covers real-world applications in Nepal (e.g., Ncell’s churn prediction, Daraz’s demand forecasting) and global platforms (G9 min readUnit 9
Machine Learning on Big Data: Algorithms, Scalability & ApplicationsUnit 9 of Big Data and Analytics explores how machine learning (ML) techniques are adapted for big data—scalable algorithms, distributed training, feature engineering for massive datasets, and real-world deployments in analytics pipelines. Covers supervised/unsupervised deep learning, model optimization, and ethical co9 min readUnit 10
Big Data Visualization and Applications · note coming