StatisticsUnit 915 min read

Practical Applications of Statistics in Tourism & Business – Case Studies

Unit 9 of Statistics: explores how statistical tools are applied to real‑world tourism and business problems in Nepal, with step‑by‑step case studies, visual analyses, and exam‑focused tips.

Key points

  • Statistical techniques turn raw tourism data into actionable insights for marketing, pricing and forecasting.
  • Correlation and regression reveal demand drivers such as seasonality, exchange rates and festival calendars.
  • Time‑series decomposition helps tourism operators separate trend, seasonal and irregular components for better capacity planning.
  • Sampling and presentation methods (frequency tables, box‑plots) enable reliable surveys of tourist satisfaction.
  • Real‑world case studies illustrate how e‑payment platforms, airlines and travel agencies use these methods daily.

1. Why Statistics Matter in Tourism & Business

Tourism operators collect massive amounts of data: daily arrivals, hotel occupancy, average spend, online bookings, and customer feedback. Statistics converts this “noise” into information that supports:

Decision Area Typical Statistical Tool Example in Nepal
Pricing strategy Regression analysis Determining ticket price elasticity for Nepal Airlines
Capacity planning Time‑series forecasting Forecasting peak season hotel rooms for Pokhara
Market segmentation Cluster analysis (not covered in syllabus but built on measures of central tendency) Grouping domestic vs. foreign tourists
Service quality improvement Box‑and‑whisker plots of satisfaction scores Analyzing TripAdvisor ratings for Annapurna treks

Understanding these tools is essential for any tourism manager or business analyst.


2. Case Study 1 – Forecasting Monthly Tourist Arrivals

2.1 Data Collection

A regional tourism board recorded monthly international arrivals (in thousands) for the past 3 years.

Month Year 1 Year 2 Year 3
Jan 12.5 13.0 13.4
Feb 14.2 14.8 15.1
Mar 18.9 19.5 20.2
Apr 22.1 22.8 23.5
May 30.4 31.2 32.0
Jun 45.6 46.8 48.1
Jul 58.9 60.3 61.7
Aug 55.2 56.5 57.9
Sep 40.1 41.0 42.3
Oct 28.7 29.4 30.1
Nov 18.3 18.9 19.5
Dec 13.8 14.2 14.7

2.2 Time‑Series Decomposition

We decompose the series into Trend (T), Seasonal (S), and Irregular (I) components using the Classical Additive Model:

24681012102030405060Year 1 (Actual)Year 2 (Actual)Year 3 (Actual)
Actual tourist arrivals (thousands) vs. Trend (CMA) over 3 years
2468101220406080100120140xTrendSeasonalActual
Decomposition of tourist arrivals into Trend and Seasonal components

Step 1 – Centered Moving Average (CMA) for Trend

A 12‑month centered moving average smooths the data. The CMA values (in thousands) are:

Month CMA
Jan 21.2
Feb 22.0
Mar 24.1
Apr 27.3
May 33.5
Jun 44.0
Jul 55.0
Aug 55.0
Sep 44.0
Oct 33.5
Nov 27.3
Dec 24.1

Step 2 – Seasonal Indices

Seasonal index for month m = (average of (Y / CMA) for that month) ÷ overall average of ratios.

Calculated indices (rounded):

Month Seasonal Index
Jan 0.62
Feb 0.71
Mar 0.79
Apr 0.87
May 0.96
Jun 1.04
Jul 1.12
Aug 1.06
Sep 0.92
Oct 0.84
Nov 0.71
Dec 0.62

Step 3 – Deseasonalized Data

. For example, for July Year 3:

2.3 Forecast for Next Year (2025)

Assume the trend continues linearly. Fit a simple linear regression to the CMA values (time = 1…36). The estimated trend equation:

For month m of the next year (t = 37 + m‑1), compute and re‑apply the seasonal index.

Example – Forecast for July 2025

  • t = 43 → (thousands)
  • Seasonal index for July = 1.12
  • Forecast thousand arrivals.

2.4 Interpretation

  • Trend shows a steady increase of ~0.95 k per month, reflecting growing international interest.
  • Seasonal peaks (July‑August) coincide with the monsoon trekking season and festivals like Rato Machhindranath.
  • The forecast helps hotels in Pokhara allocate extra staff and negotiate bulk linen contracts before the high‑season surge.

3. Case Study 2 – Determining Factors of Tourist Expenditure

3.1 Objective

Identify which variables most strongly influence average daily spend (NRS) of foreign tourists visiting Kathmandu.

3.2 Variables Collected (sample = 120 tourists)

Variable Symbol Description
Age (years) Tourist’s age
Length of stay (days) Number of nights
Exchange rate (NRS/USD) Rate at time of visit
Festival period (0/1) 1 if visit overlaps a major festival
Daily spend (NRS) Dependent variable

3.3 Correlation Matrix

Key observations

UStayFestivalHigh Spend, Medium SpendFestival Period, Non-Festival PeriodYoung, Old
Venn diagram showing overlap between Stay, Festival, and Age influencing Spend
  • Strongest positive correlation: Stay ↔ Spend (r = 0.62) – longer stays raise daily spend.
  • Moderate positive correlation with Festival period (r = 0.28) – festivals boost spending.
  • Age and exchange rate show weak relationships.

3.4 Simple Linear Regression (Stay → Spend)

Model:

0123456789102 Days (NRS 1,200)4 Days (NRS 2,100)6 Days (NRS 3,100)
Number line showing daily spend increase per extra night stayed (NRS 850/day)
1234567891010002000300040005000xyy = 200 + 500x(2, 1200)(3, 1700)(4, 2100)(5, 2600)(6, 3100)
Scatter plot with regression line: Spend (NPR) vs Days Stayed

Using least‑squares:

Interpretation – A tourist staying one extra night is expected to increase average daily spend by NRS 850.

3.5 Multiple Regression (All predictors)

  • – 71 % of variance explained.
  • Significant predictors (p < 0.05): Stay, Festival, Exchange rate. Age is not significant.

3.6 Business Application

A travel agency (e.g., Pathao Tours) can use the model to price bundled packages. If a package includes a 5‑day stay during Dashain, the expected daily spend rises by:

Thus, the agency can set a premium of NRS 6,400 per tourist for the festival bundle, ensuring profitability while matching market willingness.


4. Sampling Techniques for Tourist Satisfaction Surveys

4.1 Types of Sampling

Technique How it works When to use in tourism
Simple Random Sampling (SRS) Every tourist has equal chance of selection Small‑scale hotel surveys
Systematic Sampling Choose every k‑th respondent after a random start Queue at a popular attraction
Stratified Sampling Divide population into strata (e.g., domestic vs. foreign) and sample each National tourism board wants representation across source markets
Cluster Sampling Select whole groups (e.g., hotels) then survey all guests in chosen clusters Cost‑effective for remote trekking lodges

4.2 Worked Example – Stratified Sample for Kathmandu Visitor Survey

  • Population: 150,000 foreign tourists per year.
  • Strata: Europe (30 %), Asia (45 %), America (25 %).
  • Desired sample size: n = 600.
Domestic Tourists (45%)Foreign Tourists (Asia) (30%)Foreign Tourists (Europe) (15%)Foreign Tourists (Americas) (10%)
Stratified sampling distribution for Kathmandu visitor survey (n=120)

Randomly select tourists from each stratum using passport data at the airport. This ensures the final satisfaction index reflects the true mix of source markets.

Europe (30%)Asia (45%)America (25%)

4.3 Presentation of Survey Results

  • Frequency distribution of satisfaction scores (1‑5).
  • Ogive (cumulative frequency) to identify the percentile at which 80 % of tourists are “satisfied” (score ≥ 4).
  • Box‑and‑whisker plot to compare satisfaction across regions.

Interpretation – Asian tourists show the widest spread (more variability), suggesting targeted service improvements.


5. Comparative Summary of Statistical Tools for Tourism Decision‑Making


6. Real‑World Applications

6.1 eSewa & Khalti – Transaction Volume Forecasting

Both e‑payment platforms experience seasonal spikes during festivals (Dashain, Tihar). They apply time‑series forecasting similar to the tourist arrivals example to allocate server capacity and schedule promotional discounts.

6.2 Daraz – Demand‑Driven Inventory

Daraz uses regression analysis to predict product demand based on variables such as search volume, price discounts, and holiday periods. The same regression logic helps a tour operator forecast bookings for a new trekking package.

6.3 Ncell – Customer Churn Prediction

Ncell analyses call‑detail records with correlation coefficients to identify factors (e.g., low data usage, high bill amount) that correlate with churn. Tourism firms can adopt the same approach to predict which repeat visitors are likely to stop using a travel agency’s services.


7. In the real world

  • Daraz’s “Flash Sale” algorithm uses a time‑series model to predict the optimal discount level that maximizes sales without eroding profit margins. The model incorporates seasonal indices similar to those calculated for tourist arrivals.
  • NTC’s network traffic monitoring applies standard deviation to detect abnormal spikes that may indicate network attacks; tourism websites (e.g., Booking.com Nepal) use the same technique to spot sudden surges in booking requests during a festival, prompting auto‑scaling of cloud resources.
  • A Nepali bank’s loan‑interest calculator employs a regression equation where the dependent variable is the interest rate and predictors include credit score, loan amount, and prevailing market rate—mirroring the multiple regression used to estimate tourist daily spend.

8. Worked Example Tied to a Real Situation

Scenario – A boutique hotel in Pokhara wants to set a seasonal room rate for the monsoon trekking season (July‑August). Historical data shows:

Month Avg. Occupancy (%) Avg. Daily Rate (NRS)
Jun 68 5,200
Jul 92 6,800
Aug 88 6,500
Sep 55 4,800

Step 1 – Compute Coefficient of Variation (CV) for each month:

Assume SD of daily rate for July = 800 NRS, for August = 750 NRS.

Step 2 – Decide Pricing Strategy

  • Lower CV indicates more stable demand; the hotel can raise rates with less risk of unsold rooms.
  • Since July’s CV is slightly higher, set a modest increase of NRS 200 over the historical average, yielding NRS 7,000.
  • For August, apply a larger increase of NRS 300 (NRS 6,800) because demand is still high but variability is marginally lower.

Result – Projected revenue for the two months:

The hotel can compare these figures with the baseline (no price change) to justify the new rates.


9. Advantages & Disadvantages of Key Techniques

Technique Advantages Disadvantages
Mean & Standard Deviation Simple, intuitive, good for symmetric data Sensitive to extreme values; may mislead for skewed tourism spend
Coefficient of Variation Allows comparison across different scales (e.g., hotel vs. airline) Requires positive mean; meaningless if mean ≈ 0
Correlation Quick insight into linear relationships Does not imply causation; ignores non‑linear patterns
Regression Provides predictive equations, quantifies impact of each factor Assumes linearity, requires large sample, multicollinearity can distort results
Time‑Series Decomposition Separates trend/seasonality, improves forecasts Needs long historical series; irregular shocks can be hard to model
Stratified Sampling Guarantees representation of key sub‑populations More planning effort; requires reliable strata information

In the real world

  • eSewa & Khalti use time-series forecasting to predict transaction volumes during festivals like Dashain and Tihar, ensuring sufficient server capacity and cash reserves. For example, during Dashain, transaction volumes spike by 30–40% compared to regular days, requiring dynamic load balancing.

  • Daraz applies multiple regression analysis to forecast demand for products like trekking gear and festival-related items. For instance, during the monsoon season, demand for raincoats and hiking boots increases by 25–30%, allowing Daraz to optimize inventory levels and reduce stockouts.

  • Ncell leverages customer churn prediction models (using logistic regression) to identify subscribers likely to switch operators. For example, if a customer’s call drop rate exceeds 5% and data usage drops by 20% in a month, Ncell proactively offers promotions to retain them.

Exam tip

  • Know the formulas for mean, median, mode, range, quartile deviation, standard deviation, CV, correlation coefficient , simple linear regression (slope & intercept), and the additive time‑series model. Memorise the steps; the exam often asks for a complete calculation (e.g., “Compute the seasonal indices for the given data”).
  • Practice interpreting graphs: be able to read a line chart of tourist arrivals, identify peaks, and explain what they represent (festival, weather, etc.).
  • Remember the hierarchy: descriptive → dispersion → relationship → forecasting. Many short‑answer questions ask you to state why a particular tool is chosen for a tourism problem.
  • Use the “5‑point checklist” when tackling a case study: (1) Identify the decision problem, (2) Choose the appropriate statistical tool, (3) Show calculations step‑by‑step, (4) Interpret the numeric result in tourism terms, (5) State a practical recommendation.
  • Time management: allocate ~5 min for data cleaning, ~10 min for calculations, ~5 min for interpretation. Write the final answer in clear bullet form; examiners reward neat, structured presentation.

Based on the TU BTTM syllabus for Statistics (STT301), unit 9.

Discussion

Loading…