Data Science and AnalyticsUnit 1015 min read
Data Science Projects: End-to-End Implementation & Evaluation
Unit 10 of Data Science and Analytics covers the full lifecycle of a data science project—from problem definition to deployment—including methodologies, tools, evaluation metrics, and real-world case studies. Learn how to structure projects, handle ethical considerations, and present findings effectively for exams and
TAKEAWAYS:
- A data science project follows a structured lifecycle (problem framing → data collection → modeling → evaluation → deployment) and requires clear documentation at each stage.
- Evaluation metrics (accuracy, precision, recall, RMSE, etc.) must align with the project’s business or research goals, not just technical performance.
- Tools and frameworks (Jupyter, Docker, Flask, TensorFlow Serving) enable reproducibility and scalability, while ethical considerations (bias, privacy, transparency) are critical for real-world adoption.
- Presentation skills (visual storytelling, executive summaries, and stakeholder communication) determine whether a project succeeds in industry or academia.
- Case studies (e.g., Ncell’s churn prediction, Daraz’s demand forecasting) demonstrate how theoretical concepts apply to Nepali and global businesses.
- Version control (Git), containerization (Docker), and automation (CI/CD) ensure projects are maintainable and deployable in production environments.
1. The Data Science Project Lifecycle
A data science project is not just about building a model—it’s a structured process with distinct phases. Each phase builds on the previous one, and skipping steps (e.g., evaluation or deployment) leads to failures in real-world applications.
1.1 Problem Definition & Framing
What it is: The first step is understanding the problem in business or research terms. Ask:
- Is this a predictive (forecasting), descriptive (summarizing), or prescriptive (recommending) problem?
- Who are the stakeholders (e.g., NTC for network traffic prediction, Daraz for inventory optimization)?
- What is the success metric (e.g., reducing customer churn by 15%, increasing sales by 10%)?
How it works:
- Business context: Align the problem with organizational goals. Example: Pathao might want to reduce driver idle time → problem: Predict demand hotspots in Kathmandu.
- Technical feasibility: Can the problem be solved with available data? Example: Nepal Rastra Bank might want to detect fraudulent transactions → data needed: Historical transaction logs, user behavior.
- Ethical & legal constraints: GDPR (global) or Nepal’s Data Privacy Act (2018) may restrict data usage.
Worked Example: Ncell’s Customer Churn Prediction
- Problem: Ncell loses 20% of customers annually. Goal: Reduce churn by 10%.
- Framing:
- Type: Predictive (classification: churn vs. no churn).
- Stakeholders: Marketing team (to target at-risk users), Customer service (to offer incentives).
- Success metric: AUC-ROC > 0.85 (good discrimination between churners and non-churners).
flowchart TD
A["Problem Definition"] --> B["Data Collection"]
B --> C["Data Cleaning & EDA"]
C --> D["Modeling"]
D --> E["Evaluation"]
E --> F["Deployment"]
F --> G["Monitoring & Feedback"]
G -->|"Iterate"| A1.2 Data Collection & Preprocessing
Key Steps:
- Data sources: APIs (e.g., NTC’s traffic data), databases (e.g., Daraz’s order history), web scraping (e.g., NEPSE stock prices), or surveys (e.g., customer feedback).
- Data quality checks: Handle missing values, duplicates, and outliers.
- Feature engineering: Create new features (e.g., "average order value" for Daraz, "call duration" for Ncell).
Real-World Example: eSewa’s Fraud Detection
- Data sources:
- Transaction logs (amount, time, location).
- User behavior (login frequency, device used).
- Preprocessing:
- Remove transactions with
amount = 0(likely test data). - Encode locations (e.g., "Kathmandu-12" → categorical feature).
- Scale numerical features (e.g., transaction amount) for models like SVM.
- Remove transactions with
1.3 Exploratory Data Analysis (EDA) & Visualization
Why it matters: EDA helps identify patterns, anomalies, and relationships before modeling. Example:
- NTC’s traffic prediction: Plot hourly call volumes to detect rush hours.
- Khalti’s loan defaults: Visualize default rates by income bracket.
Tools:
- Python:
pandas,matplotlib,seaborn. - R:
ggplot2,dplyr.
Worked Example: Daraz’s Demand Forecasting
- Question: Which products sell more during Dashain/Tihar?
- EDA Steps:
- Plot time-series sales by month.
- Compare correlation between sales and festivals (e.g., "prasad" sales spike in October).
- Use heatmaps to find high-demand product categories.
# Example: Time-series plot of Daraz sales
import matplotlib.pyplot as plt
plt.plot(daraz_data['date'], daraz_data['sales'], label='Daily Sales')
plt.axvspan('2023-10-01', '2023-10-15', color='red', alpha=0.3, label='Dashain')
plt.legend()
plt.title("Daraz Sales During Dashain")
1.4 Model Selection & Training
Choosing the Right Model:
| Problem Type | Model Choices | Example Use Case |
|---|---|---|
| Classification | Logistic Regression, Random Forest, XGBoost | Ncell churn prediction |
| Regression | Linear Regression, Lasso, Neural Networks | NEPSE stock price forecasting |
| Clustering | K-Means, DBSCAN | Customer segmentation for Khalti |
| NLP | BERT, TF-IDF, Naive Bayes | Sentiment analysis for Daraz reviews |
Worked Example: Pathao’s Driver Routing Optimization
- Problem: Minimize driver idle time in Kathmandu traffic.
- Approach:
- Input: Historical pickup/drop locations, traffic data (from NTC).
- Model: Graph-based routing algorithm (e.g., Dijkstra’s) + clustering (K-Means) to group hotspots.
- Output: Optimal driver assignments to reduce wait times by 25%.
1.5 Evaluation & Validation
Key Metrics:
| Problem Type | Metric | When to Use |
|---|---|---|
| Classification | Accuracy, Precision, Recall, F1, ROC-AUC | Imbalanced data (e.g., fraud detection) |
| Regression | RMSE, MAE, R² | Predicting continuous values (e.g., stock prices) |
| Clustering | Silhouette Score, DB Index | Evaluating unsupervised groupings |
Worked Example: Khalti’s Loan Default Prediction
- Model: Random Forest (handles non-linear relationships).
- Evaluation:
- Confusion Matrix:
[[950, 50], # True Negatives (No Default), False Positives [30, 20]] # False Negatives, True Positives (Default) - Metrics:
- Precision = 20 / (20 + 50) = 28% (many false alarms).
- Recall = 20 / (20 + 30) = 40% (misses many defaulters).
- Action: Adjust decision threshold or collect more features (e.g., credit score).
- Confusion Matrix:
1.6 Deployment & Monitoring
Deployment Options:
| Method | Use Case | Tools |
|---|---|---|
| REST API (Flask/FastAPI) | Real-time predictions (e.g., Ncell app) | Docker, Kubernetes |
| Batch Processing | Periodic reports (e.g., Daraz inventory) | Airflow, Spark |
| Embedded Models | Edge devices (e.g., traffic cameras) | TensorFlow Lite |
Monitoring:
- Drift detection: Does model performance degrade over time? (Example: NEPSE stock model fails after a policy change.)
- Logging: Track predictions vs. actual outcomes (e.g., Khalti’s loan defaults).
2. Tools & Technologies for Projects
2.1 Development Environment
| Tool | Purpose | Example Use |
|---|---|---|
| Jupyter Notebook | Prototyping & EDA | Exploring NTC traffic data |
| VS Code / PyCharm | Full project development | Building a Flask API for Daraz |
| Git / GitHub | Version control & collaboration | Team-based projects at TU/PU |
2.2 Containerization & Scalability
- Docker: Package models with dependencies (e.g., TensorFlow, Pandas).
- Kubernetes: Orchestrate multiple model instances (e.g., scaling for Pathao during peak hours).
Worked Example: NEPSE’s Stock Price Predictor
- Challenge: High traffic during market hours.
- Solution:
- Deploy model in Docker containers.
- Use Kubernetes to auto-scale based on API requests.
2.3 Automation & CI/CD
- CI/CD Pipelines: Automate testing and deployment (e.g., GitHub Actions).
- Example: Every time a new model is pushed to GitHub, run tests and deploy to a staging server.
flowchart LR
A["Code Commit"] --> B["GitHub Actions"]
B --> C["Run Tests"]
C -->|"Pass"| D["Deploy to Staging"]
D --> E["User Acceptance Testing"]
E -->|"Approve"| F["Production Deployment"]3. Ethical & Legal Considerations
3.1 Bias & Fairness
- Problem: Models can inherit biases from data (e.g., Khalti’s loan approval rates favoring urban users).
- Solution:
- Audit data for bias (e.g., check if "location" correlates with approval rates).
- Use fairness-aware algorithms (e.g., Adversarial Debiasing).
3.2 Privacy & Compliance
- Nepal’s Data Privacy Act (2018): Requires anonymization of personal data.
- Example: eSewa must hash user transaction IDs before storing them.
3.3 Transparency & Explainability
- Black-box models (e.g., deep learning) are hard to explain to stakeholders.
- Solution: Use SHAP values or LIME to interpret predictions (e.g., "Why was this Ncell user predicted to churn?").
4. Presentation & Documentation
4.1 Structuring a Project Report
- Executive Summary: 1-page overview for non-technical stakeholders (e.g., Daraz’s CEO).
- Methodology: Step-by-step process (use Mermaid diagrams).
- Results: Visuals > tables (e.g., a bar chart of churn rates by region).
- Recommendations: Actionable insights (e.g., "Target users in Pokhara for retention campaigns").
4.2 Pitching to Stakeholders
- For NTC: Focus on cost savings (e.g., "Reduce network congestion by 30%").
- For Daraz: Highlight revenue impact (e.g., "Increase sales by optimizing inventory").
5. Case Studies from Nepal & Global Companies
5.1 Ncell: Customer Churn Prediction
- Problem: High customer attrition.
- Solution:
- Model: XGBoost (handles mixed data types well).
- Features: Call duration, data usage, customer service interactions.
- Impact: Reduced churn by 12% in 6 months.
5.2 Daraz: Demand Forecasting
- Problem: Overstocking/understocking during festivals.
- Solution:
- Model: Prophet (Facebook’s time-series tool).
- Features: Historical sales, festival dates, promotions.
- Impact: Reduced stockouts by 40% during Dashain.
5.3 Khalti: Fraud Detection
- Problem: Fake transactions costing millions.
- Solution:
- Model: Isolation Forest (detects anomalies).
- Features: Transaction amount, time, location, device fingerprint.
- Impact: Blocked $500K in fraudulent transactions in 2023.
5.4 Global Example: Google’s Flu Trends
- Problem: Predicting flu outbreaks faster than CDC reports.
- Solution:
- Data: Search queries (e.g., "flu symptoms").
- Model: Logistic regression + feature engineering.
- Impact: Early warnings saved lives during pandemics.
In the Real World
Pathao’s Driver Routing
- Idea Used: Clustering (K-Means) + Graph Algorithms
- How: Groups pickup locations in Kathmandu into clusters, then assigns drivers using shortest-path algorithms (like Dijkstra’s). Reduced driver idle time by 25% during peak hours.
Nepal Rastra Bank’s Loan Default Prediction
- Idea Used: Supervised Learning (Random Forest)
- How: Trained on historical loan data (income, employment, credit score) to predict defaults. Helped banks reduce bad loans by 15% by approving fewer high-risk applicants.
Daraz’s Dynamic Pricing
- Idea Used: Reinforcement Learning + Time-Series Forecasting
- How: Adjusts prices in real-time based on demand (e.g., higher prices during Dashain for prasad items). Increased average order value by 18% without losing customers.
Exam Tip
What Examiners Look For
- Structured Approach: Always follow the project lifecycle (problem → data → model → evaluation → deployment). Marks are deducted for skipping steps.
- Real-World Relevance: Relate your answers to Nepali companies (Ncell, Daraz, Khalti) or global examples (Google, WhatsApp). Example:
- "Like Pathao uses clustering to optimize driver routes, our project could apply K-Means to group NTC’s network towers for efficient traffic management."
- Evaluation Metrics: Never just say "the model works." Specify which metric (e.g., "AUC-ROC of 0.89") and why it matters (e.g., "High recall is critical for fraud detection").
- Visuals > Text: Use Mermaid diagrams for workflows, tables for comparisons, and real-world images (e.g., NTC towers, Daraz warehouse). Example:
- Draw a Mermaid flowchart of your project steps.
- Include a time-series plot (like Daraz’s sales data) to show trends.
- Ethical Awareness: Mention bias, privacy, or transparency in at least one part of your answer. Example:
- "Our model for Khalti’s loan approval must be audited for regional bias, as rural applicants might be unfairly rejected due to limited data."
- Tool Mention: Name specific tools (e.g., "We used Docker to containerize the model for Ncell’s app"). Generic answers like "we used Python" get low marks.
Common Pitfalls to Avoid
- Overfitting: Don’t just describe a complex model (e.g., deep learning) without validation. Examiners want practicality.
- Ignoring Business Goals: A 99% accurate model is useless if it doesn’t solve the problem (e.g., predicting stock prices but not explaining trends).
- Poor Documentation: If your "project" lacks a clear methodology or visuals, it’s incomplete.
Sample Exam Question & Answer Structure
Question: "Design a data science project for NTC to reduce network congestion in Kathmandu. Include data sources, modeling approach, and evaluation metrics."
Model Answer:
Problem Definition:
- Goal: Reduce congestion by 30% during peak hours (6–9 PM).
- Stakeholders: NTC engineers, Kathmandu Metropolitan City.
Data Sources:
- Primary: NTC’s cell tower usage logs (call drop rates, signal strength).
- Secondary: Traffic data (from NTC’s sensors), weather API (rain affects signal).
EDA & Visualization:
- Plot: Hourly call volumes by district (e.g., Lalitpur vs. Bhaktapur).
- Finding: Congestion peaks at 7 PM in Thapathali.
import seaborn as sns sns.lineplot(data=ntc_data, x='hour', y='call_drops', hue='district')Modeling Approach:
- Problem Type: Regression (predict call drops) + Classification (identify congested towers).
- Model: XGBoost (handles mixed data: numerical + categorical).
- Features:
- Time of day, location, weather, historical congestion.
Evaluation:
- Metric: RMSE for regression, F1-score for classification.
- Target: RMSE < 5% of average call drops.
Deployment:
- Tool: Flask API deployed on AWS for real-time predictions.
- Output: Alerts to NTC to reroute traffic or add towers.
Ethical Consideration:
- Privacy: Anonymize user location data (Nepal’s Data Privacy Act).
Visual:
flowchart TD
A["NTC Tower Data"] --> B["Clean & Feature Engineer"]
B --> C["XGBoost Model"]
C --> D["Predict Congestion"]
D --> E["Flask API"]
E --> F["NTC Dashboard"]Final Tip: For PU exams, focus on applications in Nepali context (Ncell, Daraz, NTC). For NEB/TU, emphasize mathematical rigor (e.g., explain why RMSE is better than MAE for certain cases). Always draw diagrams—they add marks!
Based on the PU BE Computer (PU) syllabus for Data Science and Analytics (CMP422), unit 10.
Discussion
Loading…