Information RetrievalUnit 110 min read
Info Retrieval: Role, Architecture & Human Needs
Unit 1 of Information Retrieval explores why information is vital to human life, how retrieval systems function, and the core architecture behind modern search engines—from basic definitions to real-world applications in Nepalese tech ecosystems.
Key points
- Information retrieval bridges human needs and digital systems by transforming unstructured data into actionable knowledge.
- The **architecture** of IR systems follows a pipeline: *crawling → indexing → querying → ranking → presentation*.
- **Human-centric IR** prioritizes relevance, accessibility, and context over raw data volume (e.g., eSewa’s search vs. Google’s).
- **Challenges** include ambiguity (e.g., "bank" as financial vs. river), multilingualism (Nepali vs. English), and scalability (NEPSE vs. YouTube).
- **Real-world tie-ins**: Pathao’s ride-finding uses IR to match queries like "go to Thapathali" to driver locations; Daraz’s product search relies on indexing.
- **Evaluation** hinges on metrics like precision/recall, but user satisfaction (e.g., Khalti’s transaction search) often trumps pure metrics.
1. Why Information Matters: The Human Need
Information retrieval (IR) isn’t just about searching—it’s about survival, decision-making, and problem-solving. Humans rely on information to:
- Act: A farmer checks weather forecasts (NTC’s Khabar app) before planting.
- Learn: Students use ePustakalaya to find textbooks for TU exams.
- Connect: WhatsApp’s search uses IR to find old messages or contacts.
Visual: How humans process information
mindmap
root((Human Information Needs))
Needs["1. Survival (e.g., NTC alerts)"]
Needs["2. Decision (e.g., Daraz product reviews)"]
Needs["3. Learning (e.g., eSewa guides)"]
Needs["4. Social (e.g., Pathao ride history)"]
Needs["5. Entertainment (e.g., YouTube recommendations)"]Real-world example:
- eSewa’s "Bill Payment" search: When you type "Ncell bill," the system must:
- Disambiguate "Ncell" (telecom vs. user name).
- Match to the correct service provider (Ncell Ltd.).
- Return the exact payment link—not a generic "mobile recharge" page. This fails if the IR system lacks Nepal-specific telecom entity indexing.
2. What Is Information Retrieval?
Definition:
Information Retrieval (IR) is the process of obtaining relevant information from a collection of unstructured or semi-structured data in response to a user’s query.
Key Components:
| Component | Role | Nepal Example |
|---|---|---|
| Data Collection | Gathers raw data (text, images, etc.). | Daraz’s product catalogs. |
| Preprocessing | Cleans/structures data (e.g., removing stopwords like "the"). | Khalti’s transaction logs. |
| Indexing | Creates a searchable map (e.g., inverted index). | Google’s "docID → keyword" mapping. |
| Query Processing | Interprets user input (e.g., "best laptop under Rs. 50k"). | Pathao’s "go to Patan" location query. |
| Ranking | Orders results by relevance (e.g., TF-IDF, PageRank). | NEPSE’s stock search prioritizing volume. |
| Presentation | Displays results (e.g., snippets, images). | YouTube’s video thumbnails. |
3. The IR Pipeline: How It Works
flowchart TD A["User Query\n(e.g., 'TU BSc CS syllabus 2080')"] --> B["Query Analysis\n(Split into terms: TU, BSc, CS, syllabus, 2080)"] B --> C["Index Lookup\n(Finds docs with these terms)"] C --> D["Ranking\n(TF-IDF, BM25, or learned-to-rank models)"] D --> E["Result Presentation\n(Snippets, images, ads)"] E --> F["User Feedback\n(Clicks, dwell time)\n→ Improves future queries"]
Worked Example: Daraz Order Tracking
- Query: User searches "order #DZ123456789".
- Preprocessing: System ignores "order" (stopword) and focuses on
#DZ123456789. - Index Lookup: Checks inverted index for
DZ123456789→ matches to a single order record. - Ranking: Since only one match exists, it’s ranked #1 (no competition).
- Presentation: Shows order status, items, and delivery estimate.
- Failure case: If Daraz’s index is outdated, the query returns "No results."
4. Challenges in IR (With Nepalese Context)
| Challenge | Cause | Nepal Example | Solution in IR |
|---|---|---|---|
| Ambiguity | Words have multiple meanings. | "Bank" → NABIL vs. river bank. | Word sense disambiguation (WSD) models. |
| Multilingualism | Nepali vs. English queries. | "मेरो खातामा पैसा छ?" vs. "balance check." | Bilingual indexing (e.g., eSewa’s Nepali-English support). |
| Noisy Data | Typos, slang, or incomplete queries. | "laptop 50k" vs. "laptop under 50 thousand." | Query expansion (synonyms, spell-check). |
| Scalability | Millions of documents (e.g., NEPSE data). | Searching 1000+ stocks in real-time. | Distributed indexing (e.g., Elasticsearch). |
| Cultural Bias | Western-trained models may fail. | "Best hotel in Kathmandu" → returns Thamel hotels only. | Localized training data (e.g., Pathao’s Kathmandu traffic patterns). |
5. IR Architecture: From Crawlers to Results
classDiagram
class Crawler {
+fetch(URLs)
+parse(HTML/text)
}
class Indexer {
+buildInvertedIndex()
+tokenize(text)
}
class QueryProcessor {
+parseQuery()
+expandQuery()
}
class Ranker {
+scoreDocs()
+applyTFIDF()
}
class UserInterface {
+displayResults()
+handleFeedback()
}
Crawler --> Indexer : "feeds parsed data"
Indexer --> QueryProcessor : "provides index"
QueryProcessor --> Ranker : "passes query terms"
Ranker --> UserInterface : "returns ranked list"Real-world tie-in: NTC’s "Service Request" Portal
- Crawler: NTC’s bot scans its website for updates on electricity outages.
- Indexer: Stores keywords like "load shedding," "ward number," and "schedule."
- Query: User types "load shedding schedule today Kathmandu."
- Ranking: Results prioritize:
- Official NTC announcements (high authority).
- Ward-specific schedules (high relevance).
- Presentation: Shows a map of affected areas (geospatial IR).
6. Evaluation Metrics: How Good Is Your Search?
IR systems are judged by:
- Precision: % of returned results that are relevant.
- Example: If 10 results appear for "TU BSc CS syllabus" and 8 are correct, precision = 80%.
- Recall: % of relevant results actually returned.
- Example: If 100 syllabi exist but only 8 are shown, recall = 8%.
- F1-Score: Harmonic mean of precision/recall (balances both).
- User-Centric Metrics:
- Click-Through Rate (CTR): % of users who click a result (e.g., Khalti’s transaction search).
- Dwell Time: How long users stay on a result page (e.g., Daraz product pages).
Comparison Table:
| Metric | High Value Means | Nepal Example |
|---|---|---|
| Precision | Fewer irrelevant results. | eSewa shows only valid bill payment options. |
| Recall | Fewer missed relevant results. | NEPSE returns all stocks matching "bank." |
| CTR | Users trust the search. | Pathao’s "go to" queries have 90%+ CTR. |
| Dwell Time | Results satisfy the user. | YouTube videos with high watch time. |
7. IR in Nepal’s Tech Ecosystem
| Company/App | IR Technique Used | Example Query | Why It Matters |
|---|---|---|---|
| eSewa | Keyword matching + entity recognition | "Ncell Rs. 500 recharge" | Avoids scams by linking to official Ncell. |
| Khalti | Transaction log indexing | "Payment to ABC Bank on 2080-05-15" | Fraud detection via user history. |
| Daraz | Product attribute indexing | "laptop under 50k with SSD" | Filters by price, brand, and specs. |
| Pathao | Geospatial + real-time indexing | "go to Thapathali now" | Matches to nearest driver + traffic data. |
| NEPSE | Stock metadata + volume ranking | "top 10 stocks in banking sector" | Prioritizes liquidity and relevance. |
| YouTube | Multimedia indexing (audio, text) | "how to make roti" | Transcribes videos for search. |
Worked Example: Khalti’s Payment Search
- Query: User searches "paid to Nepal Oil Corp on 2080-05-20."
- Preprocessing:
- Removes "paid to" (stopwords).
- Recognizes "Nepal Oil Corp" as an entity (not just keywords).
- Parses date format.
- Index Lookup: Checks transaction logs for:
- Merchant: Nepal Oil Corp.
- Date: 2080-05-20.
- Amount: Any (user may not remember).
- Ranking: Returns the most recent match (highest relevance).
- Presentation: Shows transaction ID, amount, and merchant details.
- Failure: If the date is stored as "2080/05/20" but user types "20-05-2080," the system must handle date ambiguity.
Exam Tip
What examiners test for this unit:
- Definitions: Be ready to explain IR, precision, recall, and ranking algorithms (TF-IDF/BM25) in one sentence each.
- Architecture: Draw the IR pipeline (crawler → indexer → query processor → ranker → UI) and label each component.
- Challenges: Link ambiguity, multilingualism, and scalability to Nepalese examples (e.g., "How would you improve Daraz’s product search for Nepali users?").
- Metrics: Calculate precision/recall given a confusion matrix. Example:
- Given: 100 docs, 10 relevant, system returns 5 (all relevant).
- Precision = 5/5 = 100%; Recall = 5/10 = 50%.
- Given: 100 docs, 10 relevant, system returns 5 (all relevant).
- Real-world applications: Expect 2–3 marks on how IR is used in apps like eSewa, Pathao, or NEPSE. Use the table above as a reference.
- Shortcomings: Critique a system (e.g., "Why might Khalti’s search fail for rural users?" → Answer: Limited mobile data + Nepali script OCR issues).
Common pitfalls:
- Forgetting to mention user feedback loops (e.g., clicks improving rankings).
- Ignoring Nepal-specific challenges (e.g., low internet bandwidth affects multimedia IR).
- Overcomplicating answers with advanced terms (e.g., "BERT") when basic TF-IDF suffices.
Final Visual Summary
mindmap
root((Information Retrieval in Nepal))
Basics["1. Definitions: IR = finding info from data"]
Pipeline["2. Steps: Crawl → Index → Query → Rank → Present"]
Challenges["3. Issues: Ambiguity, Nepali-English, Scale"]
Metrics["4. Evaluation: Precision, Recall, CTR"]
Examples["5. Apps: eSewa, Khalti, Daraz, Pathao, NEPSE"]
Exam["6. Focus: Architecture, metrics, real-world fixes"]Based on the TU BSc CSIT syllabus for Information Retrieval (CSC413), unit 1.
Discussion
Loading…