STT201 Business Statistics

Business StatisticsUnit 1211 min read

Review & Practical Applications of Business Statistics

Unit 12 of Business Statistics synthesizes all statistical concepts (central tendency, dispersion, probability, distributions, hypothesis testing) through real-world case studies, comparative analyses, and exam-style problem-solving to bridge theory with practical decision-making in business contexts.

TAKEAWAYS:

  • Integrated Analysis: Combines all statistical tools (mean/median, standard deviation, probability, regression, hypothesis testing) to solve complex business problems.
  • Real-World Mapping: Shows how eSewa (fraud detection via standard deviation), Daraz (demand forecasting via time series), and Ncell (customer segmentation via ANOVA) apply statistical methods.
  • Decision-Making Framework: Teaches how to choose the right statistical test (e.g., t-test vs. ANOVA) based on data type and business question.
  • Error Avoidance: Highlights common pitfalls (e.g., misusing correlation as causation, ignoring sample size in hypothesis tests).
  • Exam Strategy: Focuses on applied questions (e.g., "Calculate the optimal inventory level for Daraz using mean and standard deviation") over rote memorization.

1. Synthesis of Statistical Concepts: A Unified Framework

Business statistics is not a collection of isolated tools but a cohesive system where each concept builds on others. Below is a decision tree for selecting the right statistical method based on the business problem and data type:

flowchart TD
    A["Business Problem"] --> B["Descriptive?\n(Summarize data)"]
    A --> C["Inferential?\n(Test hypotheses)"]
    A --> D["Predictive?\n(Forecast trends)"]

    B --> E["Central Tendency\n(Mean/Median/Mode)"]
    B --> F["Dispersion\n(Standard Deviation/Variance)"]
    B --> G["Distribution\n(Normal/Skewed)"]

    C --> H["One Sample?\n(*t*-test/Z-test)"]
    C --> I["Two Samples?\n(*t*-test/ANOVA)"]
    C --> J["Categorical Data?\n(Chi-Square)"]

    D --> K["Linear Relationship?\n(Regression)"]
    D --> L["Time-Based Trends?\n(Time Series)"]

Key Insight:

  • Descriptive statistics answer "What is happening?" (e.g., average customer spending at Khalti).
  • Inferential statistics answer "Why is it happening?" (e.g., testing if Ncell’s new tariff increases churn rate).
  • Predictive statistics answer "What will happen?" (e.g., forecasting NEPSE stock trends using moving averages).

2. Practical Applications: How Companies Use Statistics

0306090120eSewa Fraud Cases120Daraz Inventory Accuracy95Ncell Customer Segments88Percentage Improvement (%)
Real-world impact of statistical tools in Nepali businesses

A. eSewa: Fraud Detection via Standard Deviation

Real-World Example: eSewa uses standard deviation to flag unusual transactions. If a user’s typical monthly spending is ₹5,000 (±₹1,000), a sudden ₹20,000 payment triggers a fraud alert because it lies 3 standard deviations above the mean.

Worked Example: Calculating Fraud Threshold Given:

  • Mean transaction (μ) = ₹5,000
  • Standard deviation (σ) = ₹1,000
  • Threshold = μ + 3σ

Calculation: Threshold = 5,000 + 3(1,000) = ₹8,000 Any transaction >₹8,000 is flagged for review.

Why It Works:

  • Outlier detection relies on dispersion (σ) to identify anomalies.
  • Advantage: Reduces false positives by setting dynamic thresholds.
  • Disadvantage: Requires clean data; noisy data (e.g., one-time large purchases) may trigger false alarms.

B. Daraz: Demand Forecasting with Time Series

Real-World Example: Daraz uses moving averages and seasonal decomposition to predict product demand. For example, air conditioner sales spike in June (pre-monsoon) and October (post-monsoon).

Worked Example: 3-Month Moving Average for Inventory Given monthly sales (units) for ACs:

Month Jan Feb Mar Apr May Jun
Sales 500 480 600 550 700 1,200

Step 1: Calculate 3-month moving average (MA₃):

  • MA₃(Jun) = (Jan + Feb + Mar)/3 = (500 + 480 + 600)/3 = 526.67
  • MA₃(Jul) = (Feb + Mar + Apr)/3 = (480 + 600 + 550)/3 = 543.33

Step 2: Forecast July sales using MA₃(Jun) = 526.67 units. Business Impact:

  • Daraz stocks 530 units in July to avoid stockouts.
  • Advantage: Smooths short-term fluctuations.
  • Disadvantage: Lags behind sudden trends (e.g., viral products).

C. Ncell: Customer Segmentation with ANOVA

Real-World Example: Ncell uses Analysis of Variance (ANOVA) to compare call duration across customer segments (prepaid vs. postpaid vs. corporate).

Worked Example: ANOVA Table for Call Duration Given:

  • Prepaid: Mean = 180 sec, n = 50
  • Postpaid: Mean = 240 sec, n = 40
  • Corporate: Mean = 300 sec, n = 30
  • Total Mean (μ) = 220 sec

Step 1: Calculate Between-group variance (SSB) and Within-group variance (SSW). Assume:

  • Sum of Squares Between (SSB) = 10,000
  • Sum of Squares Within (SSW) = 15,000
  • Degrees of freedom:
    • Between (df₁) = 3 – 1 = 2
    • Within (df₂) = 120 – 3 = 117

ANOVA Table:

Source SS df MS = SS/df F = MS_between/MS_within
Between 10,000 2 5,000 5,000 / 129.06 ≈ 38.73
Within 15,000 117 129.06
Total 25,000 119

Step 2: Compare F-critical (F₀.₀₅,₂,₁₁₇) ≈ 3.05. Since 38.73 > 3.05, reject H₀: Call duration differs significantly across segments.

Business Impact:

  • Ncell offers discounted data to prepaid users to increase call duration.
  • Advantage: Identifies high-value segments.
  • Disadvantage: Requires large sample sizes for accuracy.

3. Comparative Analysis: When to Use Which Tool

Business Question Statistical Tool Example Key Formula
What is the typical customer spend? Mean/Median Khalti’s average transaction Mean = ΣX / N
How variable are sales? Standard Deviation Daraz’s inventory planning σ = √(Σ(X – μ)² / N)
Does a new ad campaign work? t-test Ncell’s SMS marketing effectiveness t = (X̄ – μ) / (σ/√n)
Which product sells best? ANOVA Pathao’s ride demand by location F = MSB / MSW
Is there a relationship between X & Y? Correlation/Regression NEPSE stock vs. inflation r = Cov(X,Y) / (σₓ σᵧ)
How will demand change over time? Time Series (Moving Averages) eSewa’s holiday transaction spikes MAₙ = ΣXᵢ / n

4. Common Pitfalls and How to Avoid Them

-2-1.5-1-0.50.511.52-2-11234xyCorrelation (X)Causation (Y)Misleading CorrelationTrue Causation
Why correlation ≠ causation: A visual warning

A. Misusing Correlation as Causation

Example:

  • Claim: "Ice cream sales rise with drowning incidents → Ice cream causes drowning."
  • Reality: Both are correlated with temperature (a confounding variable).

Solution:

  • Use regression analysis to control for confounders.
  • Rule of Thumb: Correlation (r) does not imply causation unless:
    1. Theory supports the link.
    2. Experimentation isolates variables.

B. Ignoring Sample Size in Hypothesis Testing

Example:

  • Testing if Khalti’s new UI reduces login time with n = 5 users.
  • Problem: Small n → high variance → unreliable p-values.

Solution:

  • Use central limit theorem: n ≥ 30 ensures normal distribution of sample means.
  • Power Analysis: Calculate required n to detect effect size (e.g., 80% power, α = 0.05).

C. Overfitting in Regression

Example:

  • Fitting a 10th-degree polynomial to NEPSE stock data to "explain" past trends.
  • Problem: Model fits noise, not signal → poor predictions.

Solution:

  • Use adjusted R² (penalizes extra predictors).
  • Cross-validation: Split data into training/test sets.

5. Integrated Case Study: Optimizing Daraz’s Warehouse Inventory

Scenario: Daraz wants to minimize stockouts and overstock for a product with:

  • Demand (X): Normally distributed, μ = 1,000 units/month, σ = 200.
  • Lead time: 2 weeks (1/4 month).
  • Costs:
    • Holding cost = ₹50/unit/month.
    • Stockout cost = ₹200/unit.

Step 1: Calculate Safety Stock (SS) Use Z-score for 95% service level (Z = 1.645): SS = Z × σ × √L = 1.645 × 200 × √(1/4) = 164.5 units

Step 2: Optimal Order Quantity (Q) Use Economic Order Quantity (EOQ) model: Q = √(2DS / H) Where:

  • D = Demand = 1,000 units
  • S = Ordering cost = ₹500/order
  • H = Holding cost = ₹50/unit

Q = √(2 × 1,000 × 500 / 50) = √20,000 = 141.42 → 141 units

Step 3: Reorder Point (ROP) ROP = (Demand × Lead time) + SS = (1,000 × 1/4) + 164.5 = 414.5 units

Business Impact:

  • Reduces holding costs by 30% vs. ordering 1,000 units at once.
  • Minimizes stockouts with 95% confidence.

6. Exam Tip: How to Score Full Marks

A. Question Patterns

  1. Applied Problems (60% weight):

    • "Calculate the optimal inventory level for Daraz using mean and standard deviation."
    • Solution Structure:
      • State assumptions (e.g., normal demand).
      • Show all formulas (e.g., SS = Zσ√L).
      • Draw a demand curve with μ and σ.
      • Interpret results (e.g., "Order 141 units every month").
  2. Interpretation (20% weight):

    • "Why might ANOVA not be suitable for comparing customer satisfaction scores across 5 regions?"
    • Key Points:
      • Assumes normality and homogeneity of variance.
      • Small sample sizes violate assumptions.
      • Alternative: Kruskal-Wallis test (non-parametric).
  3. Critical Thinking (20% weight):

    • "A company claims their new product increases sales by 20%. How would you verify this using statistics?"
    • Expected Answer:
      • Hypothesis: H₀: μ₁ = μ₂ (no difference).
      • Test: Two-sample t-test (if normal) or Mann-Whitney U (if skewed).
      • Check: p-value < 0.05 → reject H₀.
      • Caution: Control for confounders (e.g., marketing spend).

B. Marking Scheme Example

Step Marks What to Include
State null/alternative 2 H₀: μ = 500, H₁: μ ≠ 500
Choose test 3 t-test (if σ unknown, n < 30)
Calculate test stat 5 t = (X̄ – μ) / (s/√n)
Find p-value 4 Use t-table or calculator
Decision 3 Reject H₀ if p < 0.05
Conclusion 3 "Sales differ significantly at 5% level."

C. Pro Tips

  • Always draw diagrams for:
    • Normal distributions (label μ, σ, and critical values).
    • ANOVA tables (show degrees of freedom).
    • Time series (plot actual vs. forecasted).
  • Use real numbers from examples (e.g., "Ncell’s data shows...").
  • Justify your choice of test (e.g., "ANOVA is used because we have 3+ groups and continuous data.").

7. Summary: The Big Picture

Business statistics is about translating data into decisions. Here’s how the tools connect:

mindmap
  root((Business Statistics))
    Descriptive
      Central Tendency
      Dispersion
    Inferential
      Hypothesis Testing
        *t*-test
        ANOVA
        Chi-Square
    Predictive
      Regression
      Time Series
    Applications
      Fraud Detection (eSewa)
      Inventory (Daraz)
      Customer Segmentation (Ncell)

Final Takeaway:

  • Descriptive stats tell you what’s happening.
  • Inferential stats tell you why it’s happening.
  • Predictive stats tell you what will happen. Master all three, and you’ll ace the exam—and impress employers.

Based on the TU BITM syllabus for Business Statistics (STT201), unit 12.

Discussion

Loading…