Information RetrievalUnit 614 min read
Advanced Retrieval: Semantics, Learning, and Beyond
Unit 6 of Information Retrieval explores cutting-edge techniques like semantic search, machine learning in retrieval, and hybrid systems, with real-world applications in eSewa, Ncell, and global platforms like Google. Learn how these methods improve accuracy, handle ambiguity, and adapt to user behavior—critical for mo
Core Concepts & Techniques
Advanced retrieval techniques go beyond keyword matching to understand meaning, context, and user intent. This unit covers:
- Semantic Search: Moving from exact matches to understanding meaning (e.g., synonyms, word sense disambiguation).
- Machine Learning in IR: Using algorithms (e.g., neural networks, clustering) to rank results dynamically.
- Hybrid Retrieval: Combining traditional and modern methods (e.g., keyword + semantic).
- Personalization & Context: Adapting results to user history, location, or device.
- Multimodal Retrieval: Searching across text, images, audio, or video (e.g., reverse image search).
- Real-Time & Streaming Retrieval: Handling live data (e.g., news, social media).
1. Semantic Search: Beyond Keywords
Why Keyword Search Fails
Traditional search (e.g., Boolean retrieval) relies on exact term matches. Problems:
- Synonyms: "Car" vs. "Automobile" → Missed matches.
- Polysemy: "Java" (programming vs. coffee) → Wrong results.
- Context: "Bank" (finance vs. river) → Ambiguity.
How Semantic Search Works
Uses knowledge graphs, word embeddings, or ontologies to map relationships between terms. Example:
- Query: "Find cheap hotels near Kathmandu airport"
- Keyword search: Matches "hotel" + "Kathmandu" + "airport" (but misses synonyms like "lodging" or "TIA").
- Semantic search: Understands "airport" = "Tribhuvan International Airport (TIA)" and expands to nearby areas.
Key Techniques
| Technique | How It Works | Example Use Case |
|---|---|---|
| Word Embeddings | Converts words to vectors (e.g., Word2Vec, GloVe) to find semantic similarity. | Google’s "King - Man + Woman ≈ Queen" logic. |
| Knowledge Graphs | Links entities (e.g., "Nepal" → "Kathmandu" → "Durbar Square"). | Google’s search graph for "Ncell plans". |
| Ontologies | Structured hierarchies (e.g., "Animal" → "Mammal" → "Dog"). | Medical search for symptoms/diseases. |
Worked Example: eSewa’s Semantic Search
Scenario: A user searches "electricity bill payment for my house in Lalitpur".
- Keyword Search: Finds pages with "electricity," "bill," "payment," "Lalitpur" (but may miss "house" or "Nepal Electricity Authority").
- Semantic Search:
- Expands "house" → "residential connection."
- Links "electricity bill" → "NEA" (Nepal Electricity Authority).
- Uses location data to filter Lalitpur-specific results.
- Result: Direct link to NEA’s e-payment portal for residential users.
Mermaid Diagram: Semantic Expansion Process
flowchart TD
A["User Query: 'electricity bill payment Lalitpur'"] --> B["Keyword Extraction"]
B --> C["Semantic Expansion"]
C --> D1["house" → "residential connection"]
C --> D2["electricity bill" → "NEA"]
C --> D3["Lalitpur" → "Nepal Electricity Authority (NEA) zone"]
D1 & D2 & D3 --> E["Rank & Filter"]
E --> F["Top Result: NEA e-payment portal"]2. Machine Learning in Information Retrieval
Why Use ML?
- Adaptability: Learns from user behavior (e.g., clicks, dwell time).
- Personalization: Recommends based on history (e.g., YouTube, Netflix).
- Scalability: Handles billions of queries (e.g., Google’s RankBrain).
Key ML Models
| Model/Algorithm | Role in IR | Example Platform |
|---|---|---|
| TF-IDF + SVM | Classifies documents by relevance. | Early email spam filters. |
| Neural Networks | Deep learning for semantic matching (e.g., BERT, Transformers). | Google’s "BERT" for natural language queries. |
| Clustering (K-Means) | Groups similar documents (e.g., news articles by topic). | Flipboard’s content aggregation. |
| Reinforcement Learning | Optimizes rankings based on user feedback (e.g., clicks). | Facebook’s News Feed algorithm. |
Worked Example: Ncell’s Recommendation System
Scenario: Ncell wants to recommend a new plan to a user who frequently uses data at night.
- Data Collected:
- Night-time data usage (10 PM–4 AM): 5 GB/month.
- Current plan: 3 GB/day (expensive for night use).
- ML Approach:
- Clustering: Groups users with similar night-time usage patterns.
- Collaborative Filtering: Recommends plans popular among similar users (e.g., "Night Unlimited" plan).
- Ranking: Uses past clicks to prioritize the "Night Unlimited" plan over others.
Mermaid Diagram: Ncell’s Recommendation Pipeline
3. Hybrid Retrieval Systems
What’s the Problem?
- Keyword-only: Misses context (e.g., "Python" as a snake vs. programming).
- Semantic-only: Too slow for large-scale searches (e.g., Daraz’s product catalog).
Solution: Combine Methods
| Component | Example Techniques | Why Combine? |
|---|---|---|
| Retrieval | BM25 (keyword) + Word2Vec (semantic) | Fast filtering + meaning-aware ranking. |
| Ranking | TF-IDF + Neural scoring | Balances precision and recall. |
| Re-ranking | User feedback + collaborative signals | Personalizes final results. |
Worked Example: Daraz’s Search
Scenario: User searches "iPhone 13 case for women in Nepal".
- Keyword Retrieval (BM25): Finds pages with "iPhone 13," "case," "women."
- Semantic Expansion:
- "Women" → "pink," "floral," "minimalist."
- "Nepal" → filters to Nepali sellers (e.g., Daraz Nepal inventory).
- Hybrid Ranking:
- Combines keyword matches (price, brand) with semantic relevance (design preferences).
- Result: Top 3 cases with high ratings + matching descriptions.
4. Personalization & Context-Aware Search
How It Works
- User Profile: Past queries, clicks, location (e.g., "Kathmandu traffic" vs. "Pokhara traffic").
- Context Signals:
- Device: Mobile vs. desktop (e.g., simplified results for slow connections).
- Time: "Best restaurants in Thamel at 8 PM" (vs. 2 AM).
- Location: "ATMs near my current GPS" (e.g., Khalti’s "Near Me" feature).
Techniques
| Technique | Example Application |
|---|---|
| Session-Based Filtering | Pathao’s ride recommendations based on recent trips. |
| Location-Aware Ranking | Google Maps’ "Nearby" results. |
| Device Adaptation | Mobile-friendly search results on slow networks. |
Worked Example: Khalti’s "Pay Near Me"
Scenario: User opens Khalti to pay a bill but is unsure of the merchant’s Khalti ID.
- Location Context: Khalti uses GPS to show nearby merchants (e.g., "NTC office, Thapathali").
- Personalization:
- If the user frequently pays NTC bills, Khalti pre-filters NTC-related merchants.
- Hybrid Search:
- Keyword: "NTC" + "Thapathali."
- Semantic: Expands "bill payment" → "electricity," "telecom," "tax."
5. Multimodal Retrieval
What It Is
Searching across multiple data types:
- Text + Images (e.g., reverse image search).
- Audio + Text (e.g., Shazam for songs).
- Video + Transcripts (e.g., YouTube’s search).
Key Applications
| Modality Combination | Example Platform | Use Case |
|---|---|---|
| Text + Image | Google Lens, Pinterest | Search for "red dress" → finds images + product pages. |
| Audio + Text | Shazam, Spotify | Identify a song → shows lyrics, artist. |
| Video + Text | YouTube, TikTok | Search "how to cook dal bhat" → videos + recipes. |
Worked Example: Google Lens in Nepal
Scenario: User takes a photo of a rare Nepali flower (e.g., Rhododendron) while hiking.
- Image Processing: Google Lens extracts visual features.
- Text Search: Matches against a database of Nepali flora.
- Hybrid Result:
- Name: "Nepali Rhododendron (Lali Gurans)."
- Related text: Wikipedia page, local nurseries (e.g., Godavari Nursery, Kathmandu).
- Bonus: Shows similar flowers in the region.
6. Real-Time & Streaming Retrieval
Why It Matters
- News: Users want updates on "Nepal earthquake 2023" in real time.
- Social Media: Twitter/Pathao feeds update every second.
- Stocks: NEPSE share prices change continuously.
Techniques
| Technique | How It Works | Example Use Case |
|---|---|---|
| Stream Processing | Processes data as it arrives (e.g., Apache Kafka). | Live sports scores on Pathao. |
| Incremental Indexing | Updates indexes without full rebuilds. | Google News’ real-time article ranking. |
| Event-Based Triggers | Alerts for specific events (e.g., "Ncell outage in Lalitpur"). | NTC’s service status updates. |
Worked Example: NEPSE Live Share Prices
Scenario: User searches "NEPSE top gainers today" at 3 PM.
- Streaming Data: NEPSE’s API pushes real-time price updates.
- Incremental Ranking:
- Compares current prices with 9 AM opening prices.
- Filters stocks with >5% gain.
- Result: Top 5 stocks (e.g., "Global IME Bank," "Nabil Bank") with live price graphs.
## In the Real World
eSewa’s Semantic Search
- Idea Used: Semantic expansion + knowledge graphs.
- How: When you search "electricity bill payment," eSewa doesn’t just match keywords—it understands you mean "NEA bill" and shows the correct payment portal, even if you type "my house’s power bill."
Ncell’s Hybrid Recommendation System
- Idea Used: Collaborative filtering + reinforcement learning.
- How: If you always use data at night, Ncell’s app recommends the "Night Unlimited" plan before you even think to search for it, based on your usage patterns.
Google Lens in Nepal
- Idea Used: Multimodal retrieval (image + text).
- How: Point your phone at a rare Nepali plant (e.g., Dendrobium orchid), and Google Lens will show you its name, where it grows, and even local sellers on Daraz or Sano Commerce.
Pathao’s Ride Recommendations
- Idea Used: Context-aware personalization.
- How: If you usually take rides to Thamel at night, Pathao will suggest drivers near your current location and pre-select the "Night Surge" pricing option if demand is high.
NTC’s Service Status Updates
- Idea Used: Real-time streaming retrieval.
- How: During a power outage, NTC’s website doesn’t wait for you to refresh—it pushes live updates via web sockets, so you see "Outage in Lalitpur: Estimated restore time: 2 hours" instantly.
## Exam Tip
What Examiners Look For
Definitions with Examples:
- Don’t just say "semantic search uses word embeddings." Explain how Word2Vec would turn "car" and "automobile" into similar vectors.
- Exam Pitfall: Vague answers like "ML improves search" → Fix: "TF-IDF + SVM classifies spam emails by training on labeled datasets (e.g., Gmail’s ‘Promotions’ folder)."
Comparisons:
- Always compare keyword vs. semantic vs. hybrid retrieval in a table (as shown above). Examiners love structured comparisons.
Real-World Applications:
- Must link to Nepali platforms (eSewa, Ncell, Daraz) or global ones (Google, YouTube).
- Example Question: "How does Pathao use context-aware search?" Answer: "Pathao combines location (GPS), time (rush hour), and user history (frequent destinations) to rank drivers. If you often go to Thamel at 8 PM, it prioritizes drivers near Thamel and shows surge pricing warnings."
Diagrams:
- Draw (or describe) a hybrid retrieval pipeline or semantic expansion graph. Even if you can’t draw, describe it step-by-step with arrows: "Query → BM25 (keyword) → Semantic Layer (Word2Vec) → Neural Ranking → Final Results."
Math Lite:
- For TF-IDF or cosine similarity, show a mini-worked example:
- Query: "best laptop under 50000"
- Document: "Dell Inspiron 15, 8GB RAM, 50000 NPR"
- TF-IDF Score: Calculate weights for "laptop," "Dell," "50000" (even if simplified).
- For TF-IDF or cosine similarity, show a mini-worked example:
Shortcuts for Full Marks:
- Use bullet lists for pros/cons (e.g., "Semantic search: ✅ Handles synonyms ❌ Slower than keyword").
- Bold key terms in your answer (e.g., "Knowledge graphs link entities like ‘Nepal’ → ‘Kathmandu’ → ‘Durbar Square’").
Common Mistakes to Avoid
- ❌ "Advanced retrieval is just better search." → Fix: "It’s semantic understanding + ML + multimodal data (e.g., Google Lens)."
- ❌ Ignoring Nepali examples. Always tie to eSewa, Ncell, Daraz, or NEPSE.
- ❌ Describing how to build a search engine instead of explaining techniques (e.g., don’t write about crawlers—focus on ranking or personalization).
Based on the TU BSc CSIT syllabus for Information Retrieval (CSC413), unit 6.
Discussion
Loading…