CSC413 Information Retrieval

Information RetrievalUnit 110 min read

Info Retrieval: Role, Architecture & Human Needs

Unit 1 of Information Retrieval explores why information is vital to human life, how retrieval systems function, and the core architecture behind modern search engines—from basic definitions to real-world applications in Nepalese tech ecosystems.

Key points

  • Information retrieval bridges human needs and digital systems by transforming unstructured data into actionable knowledge.
  • The **architecture** of IR systems follows a pipeline: *crawling → indexing → querying → ranking → presentation*.
  • **Human-centric IR** prioritizes relevance, accessibility, and context over raw data volume (e.g., eSewa’s search vs. Google’s).
  • **Challenges** include ambiguity (e.g., "bank" as financial vs. river), multilingualism (Nepali vs. English), and scalability (NEPSE vs. YouTube).
  • **Real-world tie-ins**: Pathao’s ride-finding uses IR to match queries like "go to Thapathali" to driver locations; Daraz’s product search relies on indexing.
  • **Evaluation** hinges on metrics like precision/recall, but user satisfaction (e.g., Khalti’s transaction search) often trumps pure metrics.

1. Why Information Matters: The Human Need

Information retrieval (IR) isn’t just about searching—it’s about survival, decision-making, and problem-solving. Humans rely on information to:

  • Act: A farmer checks weather forecasts (NTC’s Khabar app) before planting.
  • Learn: Students use ePustakalaya to find textbooks for TU exams.
  • Connect: WhatsApp’s search uses IR to find old messages or contacts.

Visual: How humans process information

mindmap
  root((Human Information Needs))
    Needs["1. Survival (e.g., NTC alerts)"]
    Needs["2. Decision (e.g., Daraz product reviews)"]
    Needs["3. Learning (e.g., eSewa guides)"]
    Needs["4. Social (e.g., Pathao ride history)"]
    Needs["5. Entertainment (e.g., YouTube recommendations)"]

Real-world example:

  • eSewa’s "Bill Payment" search: When you type "Ncell bill," the system must:
    1. Disambiguate "Ncell" (telecom vs. user name).
    2. Match to the correct service provider (Ncell Ltd.).
    3. Return the exact payment link—not a generic "mobile recharge" page. This fails if the IR system lacks Nepal-specific telecom entity indexing.

2. What Is Information Retrieval?

Definition:

Information Retrieval (IR) is the process of obtaining relevant information from a collection of unstructured or semi-structured data in response to a user’s query.

Key Components:

Component Role Nepal Example
Data Collection Gathers raw data (text, images, etc.). Daraz’s product catalogs.
Preprocessing Cleans/structures data (e.g., removing stopwords like "the"). Khalti’s transaction logs.
Indexing Creates a searchable map (e.g., inverted index). Google’s "docID → keyword" mapping.
Query Processing Interprets user input (e.g., "best laptop under Rs. 50k"). Pathao’s "go to Patan" location query.
Ranking Orders results by relevance (e.g., TF-IDF, PageRank). NEPSE’s stock search prioritizing volume.
Presentation Displays results (e.g., snippets, images). YouTube’s video thumbnails.

3. The IR Pipeline: How It Works

flowchart TD
  A["User Query\n(e.g., 'TU BSc CS syllabus 2080')"] --> B["Query Analysis\n(Split into terms: TU, BSc, CS, syllabus, 2080)"]
  B --> C["Index Lookup\n(Finds docs with these terms)"]
  C --> D["Ranking\n(TF-IDF, BM25, or learned-to-rank models)"]
  D --> E["Result Presentation\n(Snippets, images, ads)"]
  E --> F["User Feedback\n(Clicks, dwell time)\n→ Improves future queries"]

Worked Example: Daraz Order Tracking

  1. Query: User searches "order #DZ123456789".
  2. Preprocessing: System ignores "order" (stopword) and focuses on #DZ123456789.
  3. Index Lookup: Checks inverted index for DZ123456789 → matches to a single order record.
  4. Ranking: Since only one match exists, it’s ranked #1 (no competition).
  5. Presentation: Shows order status, items, and delivery estimate.
    • Failure case: If Daraz’s index is outdated, the query returns "No results."

4. Challenges in IR (With Nepalese Context)

Challenge Cause Nepal Example Solution in IR
Ambiguity Words have multiple meanings. "Bank" → NABIL vs. river bank. Word sense disambiguation (WSD) models.
Multilingualism Nepali vs. English queries. "मेरो खातामा पैसा छ?" vs. "balance check." Bilingual indexing (e.g., eSewa’s Nepali-English support).
Noisy Data Typos, slang, or incomplete queries. "laptop 50k" vs. "laptop under 50 thousand." Query expansion (synonyms, spell-check).
Scalability Millions of documents (e.g., NEPSE data). Searching 1000+ stocks in real-time. Distributed indexing (e.g., Elasticsearch).
Cultural Bias Western-trained models may fail. "Best hotel in Kathmandu" → returns Thamel hotels only. Localized training data (e.g., Pathao’s Kathmandu traffic patterns).

5. IR Architecture: From Crawlers to Results

classDiagram
  class Crawler {
    +fetch(URLs)
    +parse(HTML/text)
  }
  class Indexer {
    +buildInvertedIndex()
    +tokenize(text)
  }
  class QueryProcessor {
    +parseQuery()
    +expandQuery()
  }
  class Ranker {
    +scoreDocs()
    +applyTFIDF()
  }
  class UserInterface {
    +displayResults()
    +handleFeedback()
  }
  Crawler --> Indexer : "feeds parsed data"
  Indexer --> QueryProcessor : "provides index"
  QueryProcessor --> Ranker : "passes query terms"
  Ranker --> UserInterface : "returns ranked list"

Real-world tie-in: NTC’s "Service Request" Portal

  1. Crawler: NTC’s bot scans its website for updates on electricity outages.
  2. Indexer: Stores keywords like "load shedding," "ward number," and "schedule."
  3. Query: User types "load shedding schedule today Kathmandu."
  4. Ranking: Results prioritize:
    • Official NTC announcements (high authority).
    • Ward-specific schedules (high relevance).
  5. Presentation: Shows a map of affected areas (geospatial IR).

IR systems are judged by:

  1. Precision: % of returned results that are relevant.
    • Example: If 10 results appear for "TU BSc CS syllabus" and 8 are correct, precision = 80%.
  2. Recall: % of relevant results actually returned.
    • Example: If 100 syllabi exist but only 8 are shown, recall = 8%.
  3. F1-Score: Harmonic mean of precision/recall (balances both).
  4. User-Centric Metrics:
    • Click-Through Rate (CTR): % of users who click a result (e.g., Khalti’s transaction search).
    • Dwell Time: How long users stay on a result page (e.g., Daraz product pages).

Comparison Table:

Metric High Value Means Nepal Example
Precision Fewer irrelevant results. eSewa shows only valid bill payment options.
Recall Fewer missed relevant results. NEPSE returns all stocks matching "bank."
CTR Users trust the search. Pathao’s "go to" queries have 90%+ CTR.
Dwell Time Results satisfy the user. YouTube videos with high watch time.

7. IR in Nepal’s Tech Ecosystem

Company/App IR Technique Used Example Query Why It Matters
eSewa Keyword matching + entity recognition "Ncell Rs. 500 recharge" Avoids scams by linking to official Ncell.
Khalti Transaction log indexing "Payment to ABC Bank on 2080-05-15" Fraud detection via user history.
Daraz Product attribute indexing "laptop under 50k with SSD" Filters by price, brand, and specs.
Pathao Geospatial + real-time indexing "go to Thapathali now" Matches to nearest driver + traffic data.
NEPSE Stock metadata + volume ranking "top 10 stocks in banking sector" Prioritizes liquidity and relevance.
YouTube Multimedia indexing (audio, text) "how to make roti" Transcribes videos for search.

Worked Example: Khalti’s Payment Search

  1. Query: User searches "paid to Nepal Oil Corp on 2080-05-20."
  2. Preprocessing:
    • Removes "paid to" (stopwords).
    • Recognizes "Nepal Oil Corp" as an entity (not just keywords).
    • Parses date format.
  3. Index Lookup: Checks transaction logs for:
    • Merchant: Nepal Oil Corp.
    • Date: 2080-05-20.
    • Amount: Any (user may not remember).
  4. Ranking: Returns the most recent match (highest relevance).
  5. Presentation: Shows transaction ID, amount, and merchant details.
    • Failure: If the date is stored as "2080/05/20" but user types "20-05-2080," the system must handle date ambiguity.

Exam Tip

What examiners test for this unit:

  1. Definitions: Be ready to explain IR, precision, recall, and ranking algorithms (TF-IDF/BM25) in one sentence each.
  2. Architecture: Draw the IR pipeline (crawler → indexer → query processor → ranker → UI) and label each component.
  3. Challenges: Link ambiguity, multilingualism, and scalability to Nepalese examples (e.g., "How would you improve Daraz’s product search for Nepali users?").
  4. Metrics: Calculate precision/recall given a confusion matrix. Example:
    • Given: 100 docs, 10 relevant, system returns 5 (all relevant).
      • Precision = 5/5 = 100%; Recall = 5/10 = 50%.
  5. Real-world applications: Expect 2–3 marks on how IR is used in apps like eSewa, Pathao, or NEPSE. Use the table above as a reference.
  6. Shortcomings: Critique a system (e.g., "Why might Khalti’s search fail for rural users?" → Answer: Limited mobile data + Nepali script OCR issues).

Common pitfalls:

  • Forgetting to mention user feedback loops (e.g., clicks improving rankings).
  • Ignoring Nepal-specific challenges (e.g., low internet bandwidth affects multimedia IR).
  • Overcomplicating answers with advanced terms (e.g., "BERT") when basic TF-IDF suffices.

Final Visual Summary

mindmap
  root((Information Retrieval in Nepal))
    Basics["1. Definitions: IR = finding info from data"]
    Pipeline["2. Steps: Crawl → Index → Query → Rank → Present"]
    Challenges["3. Issues: Ambiguity, Nepali-English, Scale"]
    Metrics["4. Evaluation: Precision, Recall, CTR"]
    Examples["5. Apps: eSewa, Khalti, Daraz, Pathao, NEPSE"]
    Exam["6. Focus: Architecture, metrics, real-world fixes"]

Based on the TU BSc CSIT syllabus for Information Retrieval (CSC413), unit 1.

Discussion

Loading…