Data Science and AnalyticsUnit 110 min read
Data Science Basics: Definitions, Scope & Applications
Unit 1 of Data Science and Analytics introduces core concepts—what data science is, its evolution, key components (data, tools, techniques), and real-world applications in industries like finance, healthcare, and e-commerce. Learn how it differs from traditional statistics and business intelligence, and explore its rol
What is Data Science?
Data Science is an interdisciplinary field that combines statistics, programming, domain expertise, and machine learning to extract meaningful insights from structured and unstructured data. It bridges the gap between raw data and actionable knowledge, enabling data-driven decision-making.
Key Definitions
- Data: Raw facts and figures (e.g., numbers, text, images) collected for analysis.
- Information: Processed data that answers specific questions (e.g., "Customer churn rate is 15%").
- Insight: Strategic understanding derived from data (e.g., "Offer discounts to high-churn customers").
- Analytics: Techniques (descriptive, predictive, prescriptive) used to analyze data.
How Data Science Works
Evolution of Data Science
Data Science has evolved through three major phases:
Descriptive Analytics (1960s–1990s)
- Focus: Summarizing historical data (e.g., sales reports).
- Tools: Excel, SQL, basic statistics.
- Example: A bank analyzing past loan defaults to report trends.
Predictive Analytics (2000s–Present)
- Focus: Forecasting future trends using machine learning.
- Tools: Python (Scikit-learn), R, TensorFlow.
- Example: Ncell predicting customer churn based on usage patterns.
Prescriptive Analytics (Emerging)
- Focus: Recommending optimal actions.
- Tools: Optimization algorithms, reinforcement learning.
- Example: Daraz suggesting inventory restock levels to minimize losses.
Core Components of Data Science
| Component | Description | Example Tools/Techniques |
|---|---|---|
| Data Collection | Gathering data from sources (databases, APIs, sensors). | Python (Requests), SQL, Web Scraping |
| Data Cleaning | Handling missing values, outliers, and inconsistencies. | Pandas, OpenRefine |
| Exploratory Data Analysis (EDA) | Visualizing patterns (graphs, charts) and summarizing statistics. | Matplotlib, Seaborn, Tableau |
| Machine Learning | Building models to predict or classify data. | Scikit-learn, TensorFlow, PyTorch |
| Big Data Tools | Processing large-scale data (Hadoop, Spark). | Hadoop, Spark, Kafka |
| Domain Knowledge | Understanding the business context (e.g., finance, healthcare). | Industry reports, expert interviews |
Data Science vs. Related Fields
| Field | Focus | Key Difference from Data Science |
|---|---|---|
| Statistics | Hypothesis testing, probability, sampling. | Data Science applies statistics + programming + ML. |
| Business Intelligence (BI) | Dashboards, reports for business decisions. | BI is reactive; Data Science is predictive/prescriptive. |
| Machine Learning (ML) | Algorithms to learn from data. | Data Science includes data collection, cleaning, and storytelling. |
| Data Engineering | Building pipelines to store/process data. | Data Science uses engineered data to build models. |
Applications in Nepal and Globally
## In the real world
eSewa (Nepal)
- Idea Used: Predictive Analytics
- How: eSewa uses transaction data to predict fraudulent activities (e.g., unusual login locations) and flags them for review. Machine learning models analyze spending patterns to detect anomalies in real time.
Pathao (Ride-Hailing App)
- Idea Used: Clustering & Optimization
- How: Pathao’s algorithm clusters high-demand areas (e.g., Thapathali during rush hour) and optimizes driver routes using real-time traffic data from NTC. This reduces wait times and fuel costs.
Nepal Rastra Bank (NRB) Loan Approvals
- Idea Used: Classification Models
- How: NRB uses logistic regression and decision trees to classify loan applicants as "high-risk" or "low-risk" based on credit history, income, and collateral. This reduces default rates by 20%.
YouTube (Global)
- Idea Used: Collaborative Filtering (Recommender Systems)
- How: YouTube’s algorithm clusters users by watch history and recommends videos based on similar users’ preferences. This keeps users engaged for 70% of watch time via personalized content.
Google Maps (Global)
- Idea Used: Geospatial Data + Machine Learning
- How: Google Maps uses real-time traffic data from NTC (Nepal) and global sources to predict congestion. It reroutes users dynamically, saving 20–30 minutes in Kathmandu traffic during peak hours.
Daraz (E-Commerce)
- Idea Used: Association Rule Mining (Market Basket Analysis)
- How: Daraz analyzes purchase histories to suggest "Frequently Bought Together" items (e.g., "Customers who bought a phone also bought a case"). This increases cross-selling by 15%.
Worked Example: Predicting Customer Churn for Ncell
Scenario: Ncell wants to reduce customer churn (subscriptions canceled). They have data on:
- Call duration, SMS usage, data consumption.
- Customer complaints, payment history.
- Demographic details (age, location).
Step-by-Step Approach
Data Collection
- Extract data from Ncell’s CRM system (e.g., 10,000 customers over 6 months).
- IMAGE: CRM database schema labelled diagram | How customer data is structured in a telecom database.
Data Cleaning
- Handle missing values (e.g., 5% of call duration records are empty → impute with median).
- Remove outliers (e.g., a customer using 1TB data/month when average is 5GB).
Exploratory Data Analysis (EDA)
- Plot churn rate by region (e.g., Pokhara has 25% churn vs. Kathmandu’s 12%).
- IMAGE: customer churn rate by region (bar chart) | Highlighting regional differences.
Model Building
- Use Logistic Regression to predict churn (binary: 1 = churned, 0 = retained).
- Features:
call_duration,complaints,payment_delay,data_usage. - Equation:
- Train-test split: 80% training, 20% testing.
Evaluation
- Metrics: Accuracy (82%), Precision (85%), Recall (78%).
- Confusion Matrix:
| | Predicted No Churn | Predicted Churn | |---------------|--------------------|-----------------| | Actual No Churn| 1,200 | 150 | | Actual Churn | 300 | 400 |
Actionable Insight
- Recommendation: Target Pokhara customers with loyalty discounts and improve network coverage in high-churn areas.
- Impact: Reduces churn by 18% in 6 months.
Tools and Technologies
| Category | Tools | Use Case |
|---|---|---|
| Programming | Python, R, SQL | Data cleaning, modeling |
| Visualization | Tableau, Power BI, Matplotlib | Dashboards, reports |
| Big Data | Hadoop, Spark, Apache Kafka | Processing large datasets (e.g., NTC traffic) |
| Machine Learning | Scikit-learn, TensorFlow, PyTorch | Predictive models |
| Cloud Platforms | AWS, Google Cloud, Azure | Scalable data storage (e.g., Daraz inventory) |
Challenges in Data Science
Data Quality Issues
- Problem: Missing/inconsistent data (e.g., 30% of NEPSE stock records have null values).
- Solution: Use imputation (mean/median) or flag records for manual review.
Bias in Models
- Problem: A loan approval model trained on historical data may favor urban over rural applicants.
- Solution: Audit datasets for bias and use techniques like fairness-aware ML.
Scalability
- Problem: A model working on 100K records may fail on 10M records (e.g., Pathao’s surge pricing).
- Solution: Use distributed computing (Spark) and optimize algorithms.
Ethical Concerns
- Problem: Privacy risks (e.g., Khalti sharing transaction data without consent).
- Solution: Anonymize data and comply with Nepal’s Data Privacy Act (2018).
## Exam Tip
This unit is conceptual but often tested with short-answer and scenario-based questions. Focus on:
- Definitions: Differentiate data science from statistics/BI (expect 2–3 marks).
- Applications: Link concepts to real-world examples (e.g., "How does YouTube use clustering?").
- Workflows: Describe the data science pipeline (collection → cleaning → modeling → deployment).
- Tools: Name one tool for each component (e.g., "Pandas for cleaning," "Tableau for visualization").
- Challenges: Explain one ethical or technical challenge (e.g., bias in loan approvals).
Common Exam Questions:
- "Explain the difference between descriptive and predictive analytics with an example from eSewa."
- "Draw a flowchart of the data science workflow and label each step."
- "How would you reduce churn for Ncell using data science? Outline steps."
Key Formula to Remember: For logistic regression (classification):
Based on the PU BE Computer (PU) syllabus for Data Science and Analytics (CMP422), unit 1.
Discussion
Loading…