STT201 Business Statistics

Business StatisticsUnit 112 min read

Statistics Basics: Data, Types, Collection & Uses

Unit 1 of Business Statistics introduces core concepts like data types, statistical variables, data collection methods, and ethical considerations—essential for analyzing real-world business problems using quantitative tools.

TAKEAWAYS:

  • Statistics is the science of collecting, analyzing, and interpreting data to make informed business decisions.
  • Data can be classified as primary (original) or secondary (existing), and qualitative (descriptive) or quantitative (numerical).
  • Variables are categorized as nominal, ordinal, interval, or ratio, each requiring different statistical treatments.
  • Sampling techniques (random, stratified, systematic) ensure representative data for accurate analysis.
  • Ethical data collection prioritizes privacy, accuracy, and transparency in business research.

1. Introduction to Statistics

Statistics is the backbone of decision-making in business. It helps in collecting, organizing, analyzing, and interpreting data to derive meaningful insights. Business statistics is particularly useful in:

  • Market research (e.g., customer preferences, demand forecasting).
  • Financial analysis (e.g., risk assessment, investment decisions).
  • Operational efficiency (e.g., supply chain optimization, quality control).

Key Definitions

  • Population: The entire group being studied (e.g., all customers of a bank).
  • Sample: A subset of the population used for analysis (e.g., 500 customers surveyed out of 10,000).
  • Parameter: A numerical value describing a population (e.g., mean height of all Nepali adults).
  • Statistic: A numerical value describing a sample (e.g., mean height of 1,000 surveyed Nepali adults).

Why Study Statistics?

  • Reduces uncertainty in decision-making.
  • Identifies trends (e.g., sales growth, customer behavior).
  • Optimizes resources (e.g., inventory management, advertising spend).

2. Types of Data

Data can be classified based on source and nature:

UPrimarySecondarySurveys, Experiments, Direct ObservationsGovernment Reports, Company Records, NEPSE DataPrimary Data, Secondary Data
Overlap: Data that can be both Primary and Secondary (e.g., Company Records if collected first-hand but reused)
UPrimarySecondarySurveys, Experiments, Direct ObservationsGovernment Reports, Company Records, NEPSE DataPrimary Data, Secondary Data
Primary vs. Secondary Data Sources

A. Based on Source

Type Definition Example
Primary Data Collected firsthand for a specific purpose. Surveys, experiments, direct observations.
Secondary Data Existing data collected for other purposes. Government reports, company records, NEPSE data.

B. Based on Nature

Type Definition Example
Qualitative (Categorical) Descriptive, non-numerical data. Customer feedback ("satisfied," "dissatisfied").
Quantitative (Numerical) Numerical data that can be measured. Sales revenue (Rs. 50,000), employee count (200).

3. Types of Variables

Variables are the characteristics being measured. They are classified as:

UNominal (Gender: Male, Female)Ordinal (Satisfaction: Poor, Average, Good)Interval (Temperature: 20°C, 30°C)Categories only (no order)Order matters (no equal intervals)Equal intervals (no true zero)Ratio (true zero: Height, Weight)
Measurement Scales Hierarchy (Ratio > Interval > Ordinal > Nominal)

A. Classification of Variables

Type Definition Example Measurement Scale
Nominal Categories with no order. Gender (Male, Female), Brands (Nike, Adidas). Categories only.
Ordinal Categories with a meaningful order. Customer satisfaction (Poor, Average, Good). Order matters.
Interval Numerical data with equal intervals but no true zero. Temperature (°C), IQ scores. Differences matter.
Ratio Numerical data with a true zero. Height (165 cm), Weight (60 kg), Sales (Rs. 10,000). Ratios matter (e.g., 2x height).

Visual: Measurement Scales


4. Data Collection Methods

Data collection is the first step in statistical analysis. Methods include:

017.53552.570Surveys70Experiments40Observations30Published Sources60Internal Records50Frequency of Use (%)
Primary vs. Secondary Data Collection Methods (Nepali Businesses, 2023)

A. Primary Data Collection

  1. Surveys/Questionnaires
    • Directly ask respondents for information.
    • Example: E-Sewa surveys customers on service satisfaction.
  2. Experiments
    • Manipulate variables to observe effects.
    • Example: Testing two ad campaigns to see which drives more sales.
  3. Observations
    • Record behavior without intervention.
    • Example: Counting foot traffic in a Daraz store.

B. Secondary Data Collection

  1. Published Sources
    • Government reports, company annual reports.
    • Example: NEPSE data for stock market analysis.
  2. Internal Records
    • Company databases, sales reports.
    • Example: Ncell’s customer call logs for network performance.

Advantages and Disadvantages

Method Advantages Disadvantages
Primary Data Tailored to specific needs, up-to-date. Time-consuming, expensive.
Secondary Data Quick, cost-effective. May be outdated or irrelevant.

5. Sampling Techniques

Sampling ensures that data is representative of the population. Common techniques:

UPopulationSampleUnsampled (N-100)Representative SubsetSample (n=100)
Population vs. Sample (n=100, N=1000)

A. Probability Sampling

  1. Simple Random Sampling
    • Every member has an equal chance of selection.
    • Example: Randomly selecting 100 customers from a list of 10,000.
  2. Stratified Sampling
    • Population divided into subgroups (strata), then sampled.
    • Example: Surveying students from different faculties (Management, Engineering) proportionally.
  3. Systematic Sampling
    • Select every k-th member from a list.
    • Example: Surveying every 10th customer entering a Pathao store.

B. Non-Probability Sampling

  1. Convenience Sampling
    • Select easily accessible members.
    • Example: Surveying students in a TU classroom.
  2. Purposive Sampling
    • Select members based on specific criteria.
    • Example: Interviewing top 10 Daraz sellers for case studies.

Visual: Sampling Techniques

graph TD
    A["Sampling Techniques"] --> B["Probability Sampling"]
    A --> C["Non-Probability Sampling"]
    B --> B1["Simple Random"]
    B --> B2["Stratified"]
    B --> B3["Systematic"]
    C --> C1["Convenience"]
    C --> C2["Purposive"]
    B1 --> B1a["Example: Lottery"]
    B2 --> B2a["Example: Divide by Gender"]
    B3 --> B3a["Example: Every 10th Person"]
    C1 --> C1a["Example: Convenient Students"]
    C2 --> C2a["Example: Expert Judgment"]

6. Ethical Considerations in Data Collection

Ethics ensures data is collected fairly, accurately, and respectfully:

  • Informed Consent: Participants must know how their data will be used.
  • Confidentiality: Personal data should be anonymized.
  • Accuracy: Avoid bias in questions or sampling.
  • Transparency: Clearly state the purpose of data collection.

Example: E-Sewa’s Ethical Practices

  • Uses opt-in consent for customer surveys.
  • Anonymizes transaction data for security.
  • Avoids leading questions in feedback forms.

7. Worked Examples

Example 1: Classifying Variables

Data: Customer ratings for a Khalti app (1-5 stars).

  • Variable Type: Ordinal (ordered categories with no equal intervals).
  • Why? Ratings are ordered (1 < 2 < 3), but the difference between 1 and 2 isn’t numerically equal to 2 and 3.

Example 2: Sampling in Real Life

Scenario: NTC wants to survey customer satisfaction for its broadband service.

  • Population: All 5 million broadband users.
  • Sample: 500 users selected via stratified sampling (dividing by urban/rural areas).
  • Why Stratified? Ensures representation from all regions.

Example 3: Primary vs. Secondary Data

Scenario Data Type Example
NEPSE analyzing stock trends Secondary Historical stock prices from NEPSE database.
Daraz conducting a survey on delivery times Primary Directly asking customers via email.

8. In the Real World

  1. eSewa’s Data Collection

    • Uses primary data (customer transactions) to analyze spending patterns.
    • Applies stratified sampling to segment users by age, location, and transaction frequency.
    • Ethical Note: Ensures GDPR-compliant data handling for user privacy.
  2. Pathao’s Ride Demand Prediction

    • Collects secondary data (historical ride requests, weather data) to predict peak hours.
    • Uses ratio variables (distance, fare) to optimize pricing algorithms.
    • Real-World Impact: Reduces wait times by 30% during rush hours.
  3. NEPSE’s Index Calculation

    • Relies on secondary data (company stock prices, trading volumes).
    • Uses interval variables (stock prices) to compute daily index changes.
    • Example Calculation:
      • If NEPSE index rises from 1,200 to 1,250, the percentage change is:

9. Exam Tip

  • Focus on definitions: Know the difference between population vs. sample, parameter vs. statistic, and variable types.
  • Practice classification: Given a dataset, classify variables (nominal/ordinal/interval/ratio) and identify data types (primary/secondary).
  • Sampling questions: Expect questions on when to use stratified vs. random sampling.
  • Ethics: Always justify why a method is ethical (e.g., "stratified sampling ensures representation").
  • Real-world tie-ins: Relate examples to Nepali businesses (e.g., Daraz’s inventory management, Ncell’s network data).

10. Common Mistakes to Avoid

  • Confusing nominal and ordinal data: Remember, ordinal has a meaningful order.
  • Ignoring sampling bias: Always justify why a sample is representative.
  • Overlooking ethics: Never collect data without consent or anonymization.
  • Misapplying scales: Ratio data allows ratios (e.g., "twice as tall"), but interval does not (e.g., 20°C is not "twice as hot" as 10°C).

11. Practice Questions

  1. Classify the following variables:

    • Blood type (A, B, AB, O) → Nominal
    • Customer loyalty tier (Bronze, Silver, Gold) → Ordinal
    • Temperature in °C → Interval
    • Annual salary → Ratio
  2. A bank wants to survey customer satisfaction. Suggest a probability sampling method and justify why.

  3. Differentiate between primary and secondary data with examples from Nepali businesses.


12. Summary Table

Concept Definition Example Key Use in Business
Primary Data Collected firsthand. E-Sewa surveys. Customized insights.
Secondary Data Existing data. NEPSE reports. Quick trend analysis.
Nominal Variable Categories only. Brands (Nike, Adidas). Market segmentation.
Ratio Variable True zero, ratios allowed. Sales revenue (Rs. 50,000). Financial forecasting.
Stratified Sampling Subgroups sampled proportionally. Surveying TU students by faculty. Representative results.

13. Final Visual: Data Collection Flowchart


Exam Alert: Always label axes, justify sampling methods, and tie examples to real businesses (e.g., Daraz, NEPSE, Ncell). Use tables and diagrams to organize answers clearly.

Based on the TU BBM syllabus for Business Statistics (STT201), unit 1.

Discussion

Loading…