Basic StatisticsUnit 18 min read
Introduction to Statistics: Definitions, Types, and IT Applications
Unit 1 of Basic Statistics covers the foundational concepts of statistics, including its definition, types (descriptive vs. inferential), role in IT, and key terminologies like population, sample, parameter, and statistic. This note explains how statistics is applied in real-world scenarios like eSewa, Daraz, and Ncell
TAKEAWAYS:
- Statistics is the science of collecting, analyzing, interpreting, and presenting data to make informed decisions.
- Descriptive statistics summarizes data (e.g., mean, median), while inferential statistics uses samples to draw conclusions about populations.
- Population refers to the entire group being studied, while a sample is a subset used for analysis.
- Parameters describe population characteristics, while statistics describe sample characteristics.
- Statistics plays a critical role in IT for data-driven decision-making, fraud detection, and system optimization.
- Visual tools like tables, graphs, and charts help communicate statistical insights effectively.
What is Statistics?
Statistics is the science of collecting, organizing, analyzing, interpreting, and presenting data to uncover patterns, trends, and insights. It helps in making informed decisions based on empirical evidence rather than guesswork.
Key Definitions:
- Data: Raw facts or figures (e.g., marks of students, sales records).
- Population: The entire group being studied (e.g., all students in TU).
- Sample: A subset of the population (e.g., 100 students from TU).
- Parameter: A numerical value describing a population (e.g., average height of all Nepali adults).
- Statistic: A numerical value describing a sample (e.g., average height of 500 surveyed Nepali adults).
Types of Statistics
Statistics is broadly divided into two types:
1. Descriptive Statistics
Summarizes and describes data using tables, graphs, and summary measures (e.g., mean, median, mode).
- Purpose: Organize and present data in a meaningful way.
- Example: Calculating the average marks of students in a class.
2. Inferential Statistics
Uses sample data to make predictions or inferences about a population.
- Purpose: Draw conclusions, test hypotheses, and make generalizations.
- Example: Predicting the average salary of IT professionals in Nepal based on a sample survey.
Role of Statistics in Information Technology (IT)
Statistics is essential in IT for:
- Data Analysis: Companies like Daraz use statistics to analyze customer purchase patterns and optimize inventory.
- Fraud Detection: Ncell and banks use statistical models to detect unusual transaction patterns (e.g., sudden large withdrawals).
- System Optimization: eSewa uses statistical algorithms to predict peak usage times and allocate server resources efficiently.
- Quality Assurance: Software companies use statistical sampling to test software quality before release.
- Machine Learning: Algorithms in WhatsApp (e.g., spam detection) rely on statistical techniques like regression and classification.
Real-World Example: Daraz’s Order Queue
Daraz uses statistical sampling to estimate delivery times. If 80% of orders in Kathmandu are delivered within 2 days, Daraz can set realistic delivery expectations for customers.
flowchart TD
A["Customer Places Order"] --> B["Order Enters Queue"]
B --> C["Statistical Sampling: Estimate Delivery Time"]
C --> D["Assign Delivery Slot"]
D --> E["Update Customer: 'Delivered in 2 days'"]Parameters vs. Statistics
| Parameter | Statistic |
|---|---|
| Describes population | Describes sample |
| Example: Mean height of all Nepali adults | Example: Mean height of 500 surveyed Nepali adults |
| Fixed value (if population is known) | Varies with sample selection |
| Notated with Greek letters (e.g., μ for mean) | Notated with Roman letters (e.g., x̄ for sample mean) |
Worked Example: Population vs. Sample
Scenario: NTC wants to estimate the average monthly internet usage of Nepali households.
- Population: All 8 million households in Nepal.
- Sample: 500 households surveyed in Kathmandu, Pokhara, and Biratnagar.
Calculations:
Population Mean (μ):
- If NTC surveys every household, the average usage is the population mean.
- Example: hours/month (hypothetical).
Sample Mean (x̄):
- From the 500 households, the average usage is calculated as:
- Suppose the sum is 22,500 hours:
- Here, , but in practice, samples may vary slightly.
In the Real World
- eSewa and Khalti:
- Application: These apps use descriptive statistics to summarize transaction data (e.g., average transaction amount, peak usage hours).
- How: They analyze millions of transactions to identify trends like "most transactions occur between 6–9 PM."
Daraz and NEPSE:
- Application: Inferential statistics predicts stock prices (NEPSE) and sales trends (Daraz).
- How: Daraz uses sample data from past sales to forecast demand for Diwali season, while NEPSE analysts use historical stock data to predict market trends.
Ncell and Banks:
- Application: Statistical sampling detects fraud.
- How: Ncell flags unusual call patterns (e.g., sudden international calls from a local number) by comparing a user’s data against population norms.
Exam Tip
Definitions Matter:
- Always distinguish between population/parameter and sample/statistic. Examiners often test this with numerical examples.
- Example question: "If the average salary of all IT professionals in Nepal is ₹50,000, what is the parameter?" Answer: The parameter is μ = ₹50,000 (population mean).
Real-World Applications:
- Link concepts to Nepali companies (e.g., Daraz’s inventory, Ncell’s fraud detection). Examiners love context!
- Example: "How does eSewa use statistics?" → Descriptive stats for transaction summaries; inferential stats for fraud detection.
Visuals in Exams:
- If asked to "distinguish between descriptive and inferential statistics," draw a table (like above) or a flowchart showing how data moves from sample to population.
- Example:
flowchart LR A["Sample Data"] --> B["Descriptive Stats\n(Summarize)"] B --> C["Inferential Stats\n(Predict Population)"] C --> D["Decision Making"]
Common Pitfalls:
- Mixing parameters and statistics: Always use Greek letters (μ, σ) for population and Roman letters (x̄, s) for samples.
- Ignoring units: In examples, always specify units (e.g., "marks out of 100," "hours/month").
Practice Questions (Based on Past Exams)
Describe the role of statistics in IT (5 marks):
- Use the Daraz/Ncell/eSewa examples from this note.
- Mention data analysis, fraud detection, and optimization.
Distinguish between descriptive and inferential statistics (5 marks):
- Use the table and flowchart from this note.
- Add an example: "Descriptive: Average marks of 50 students; Inferential: Predicting TU pass rate for all students."
Short notes:
- Parameter vs. Statistic: Use the definition table and NTC internet usage example.
- Use of Box and Whisker Plot: (Save for Unit 2, but hint: It’s a descriptive tool to show dispersion and outliers in data.)
Summary
- Statistics is the backbone of data-driven decision-making in IT and business.
- Descriptive stats = summarize; Inferential stats = predict.
- Population = whole; Sample = part.
- Parameters (μ, σ) describe populations; statistics (x̄, s) describe samples.
- Real-world tie: Every app (eSewa, Daraz) and service (NTC, Ncell) uses these concepts daily.
Final Tip: For exams, always relate answers to Nepali examples—examiners reward local context!
Based on the TU BIT syllabus for Basic Statistics (STA154), unit 1.
Discussion
Loading…