CSC413 Information Retrieval

Information RetrievalUnit 711 min read

Question Answering Systems: QA Pipelines, Models & Real-World Apps

Unit 7 of Information Retrieval explores how QA systems process natural language queries to extract precise answers, covering pipelines, retrieval vs. generation models, and evaluation metrics—with real-world examples from eSewa, Ncell, and Google.

TAKEAWAYS:

  • QA systems bridge the gap between unstructured text and direct answers by combining retrieval (finding relevant documents) and generation (extracting answers).
  • Key components include question classification, candidate answer retrieval, and answer validation—often using TF-IDF, BM25, or neural models like BERT.
  • Open-domain QA (e.g., Google’s search snippets) relies on pre-indexed corpora, while closed-domain QA (e.g., eSewa’s FAQs) uses domain-specific datasets.
  • Evaluation metrics like Exact Match (EM) and F1 score measure precision/recall trade-offs in answer correctness.
  • Real-world applications span customer support (Pathao’s chatbots), financial queries (Ncell’s billing), and legal research (Nepal Law Commission databases).
  • Challenges include ambiguity, context dependency, and scalability—addressed via hybrid models (e.g., retrieval + fine-tuning).

1. What Are Question Answering (QA) Systems?

QA systems automate the process of answering natural language questions by extracting precise information from unstructured text. Unlike traditional search engines (which return documents), QA systems directly return answers—e.g., a date, entity, or short text span.

How QA Systems Work: The Pipeline

Final AnswerPost-ProcessingAnswer ValidationCandidate Answer RetrievalQuestion ClassificationUser Question
Step-by-step QA pipeline hierarchy (top-down)
  • Step 1: Question Classification Categorize the question to determine the answer type (e.g., who, when, how). Example:

    • "When was Tribhuvan University founded?" → Time-based question.
    • "What is the capital of Nepal?" → Entity-based question.
  • Step 2: Candidate Answer Retrieval Use information retrieval (IR) techniques (e.g., BM25, TF-IDF) to fetch relevant documents or passages. For example:

    • For "What is the interest rate for Ncell loans?", the system retrieves Ncell’s official loan policy documents.
  • Step 3: Answer Validation Rank candidate answers using:

    • Lexical matching (exact word overlap).
    • Semantic similarity (e.g., BERT embeddings).
    • Contextual cues (e.g., negation, temporal references).
  • Step 4: Post-Processing Refine answers for readability (e.g., rephrasing, unit conversion). Example:

    • Raw answer: "The university was established in 1959."
    • Processed: "Tribhuvan University was founded in 1959."

2. Types of QA Systems

Type Description Example IR Technique Used
Open-Domain QA Answers questions from general knowledge (e.g., Wikipedia, web). Google’s "People Also Ask" snippets. BM25, Dense Retrieval (DPR).
Closed-Domain QA Answers questions from a specific dataset (e.g., FAQs, legal texts). eSewa’s chatbot for bill payments. TF-IDF, Keyword Matching.
Conversational QA Maintains context across multiple turns (e.g., chatbots). Pathao’s customer support bot. Memory Networks, Transformers.
Multi-Hop QA Requires reasoning across multiple documents (e.g., "Who wrote X after Y?"). Ncell’s troubleshooting guides. Graph Neural Networks (GNNs).

3. Key Techniques in QA Systems

A. Retrieval-Based QA

  • How it works: Uses IR models (e.g., BM25) to fetch relevant passages, then extracts answers via rule-based matching or span selection.

  • Example: For "What is the deadline for TU exam form submission?", the system:

    1. Retrieves TU’s official exam notice.
    2. Extracts the date span: "The deadline is 2024-05-15."

    Worked Example: TU Exam Deadline

    Document (TU Notice):
    "Students must submit exam forms by **May 15, 2024**, at 11:59 PM."
    
    Query: "When is the last date for TU exam form submission?"
    
    • BM25 Score: High for the sentence containing "deadline" + "May 15".
    • Answer Span: "May 15, 2024".

B. Generation-Based QA

  • How it works: Uses seq2seq models (e.g., BERT, T5) to generate answers from scratch, even if the exact phrase isn’t in the corpus.
  • Example: Google’s LaMDA generates conversational responses like:
    • User: "How do I reset my Daraz password?"
    • QA System: "Go to Daraz’s login page, click ‘Forgot Password,’ and enter your email."

C. Hybrid QA

  • Combines retrieval + generation for robustness. Example:
    1. Retrieve top-5 passages for "What is the fine for late NTC bill payment?".
    2. Use BERT to re-rank and generate a concise answer: "A fine of 10% of the unpaid amount is applied after the due date."

4. Real-World Applications in Nepal

२०७८eSewa integratesQA for bill payment qu२०७९Ncell launchestroubleshooting chatbo२०८०NEPSE adopts QAfor stock query automa
Nepal’s QA adoption timeline (2078–2080)

Example 1: eSewa’s Bill Payment QA

  • Problem: Users ask "How to pay Ncell bill via eSewa?"
  • QA Pipeline:
    1. Classification: Procedural question → triggers step-by-step guide.
    2. Retrieval: Fetches eSewa’s Ncell payment FAQ.
    3. Generation: Rephrases steps for clarity:

      "1. Open eSewa app. 2. Select ‘Pay Bills.’ 3. Choose ‘Ncell.’ 4. Enter phone number and amount."

Example 2: Ncell’s Troubleshooting Chatbot

  • Problem: "My Ncell 4G is not working."
  • Multi-Hop QA:
    1. Retrieves documents: Ncell network outages, device settings, SIM issues.
    2. Ranks answers by relevance (e.g., "Check if your SIM is inserted properly").
    3. If unresolved, escalates to human agent.

Example 3: NEPSE Stock Query System

  • Problem: "What was the closing price of NTC on 2023-12-01?"
  • Closed-Domain QA:
    • Retrieves NEPSE’s historical data CSV.
    • Extracts: "NTC closed at Rs. 125.50."

5. Evaluation Metrics for QA Systems

Metric Definition When to Use
Exact Match (EM) % of answers exactly matching the ground truth. Factoid QA (e.g., dates, names).
F1 Score Harmonic mean of precision/recall for answer spans. Extractive QA (e.g., SQuAD datasets).
BLEU Measures n-gram overlap with reference answers (used in generation QA). Open-ended answers (e.g., chatbots).
MRR (Mean Reciprocal Rank) Average rank of the first correct answer in a list. Retrieval-based QA.
021.2542.563.7585Exact Match (EM)78F1 Score85BLEU Score62
Average performance metrics for Nepali QA systems (sample dataset)

Worked Example: Evaluating a QA System for TU Exam Dates

  • Ground Truth Answer: "The TU exam starts on June 1, 2024."
  • System Answer: "Exams begin 1st June 2024."
    • EM: 0 (format mismatch).
    • F1: 0.95 (partial match on date).

6. Challenges and Solutions

Challenge Cause Solution
Ambiguity "What is the capital of Nepal?" (Kathmandu vs. historical capitals). Use contextual embeddings (BERT) or disambiguation prompts.
Lack of Training Data Low-resource languages (e.g., Nepali QA). Few-shot learning or data augmentation.
Scalability Slow retrieval for large corpora. Approximate Nearest Neighbors (ANN) or distributed indexing.
Hallucinations Generation models invent false answers. Retrieval augmentation (hybrid models).

7. Advanced Techniques

A. Text Shingling for Answer Extraction

  • What it does: Splits text into overlapping shingles (e.g., 3-gram sequences) to find exact matches.
  • Example: For the query "TU’s first chancellor", shingles might be:
    • "first chancellor of TU" (from a historical document).
  • Use Case: eSewa’s legal QA for contract clauses.

B. Rocchio Algorithm for QA Classification

  • How it works: Adjusts query vectors based on positive/negative feedback.
  • Example: Classify "process scheduling" as Operating System vs. Automata:
    1. Compute centroids for each class from training data.
    2. Update query vector:
    3. Assign to the class with the highest cosine similarity.

Worked Example: Classifying "process scheduling"

Class Relevance Feedback Updated Query Vector Similarity Score
Operating System High Closer to OS centroid 0.85
Automata Low Far from Automata centroid 0.30

Answer: Classified as Operating System.


8. Question Answering vs. Recommendation Systems

Feature Question Answering (QA) Recommendation Systems
Goal Extract precise answers from text. Predict user preferences (e.g., Daraz products).
Input Natural language questions. User behavior (clicks, ratings).
Output Text spans, entities, or facts. Item rankings (e.g., "Top 5 laptops").
Key Techniques TF-IDF, BERT, Span Extraction. Collaborative Filtering, Deep Learning.
Example in Nepal Ncell’s chatbot for bill queries. Daraz’s "Customers Also Bought" section.

Exam Tip

  1. For definitions:

    • QA systems retrieve + generate answers; retrieval-based uses IR, generation-based uses seq2seq.
    • Text shingling = overlapping n-grams for exact match retrieval.
  2. For algorithms:

    • Rocchio: Update query vectors using , , weights.
    • BM25: Prefer documents with high-term frequency but low collection frequency.
  3. For applications:

    • Open-domain: Google, Wikipedia.
    • Closed-domain: eSewa, Ncell FAQs.
    • Multi-hop: NTC’s troubleshooting guides.
  4. For evaluation:

    • EM = exact match rate (strict).
    • F1 = balance between precision/recall (lenient).
  5. Common pitfalls:

    • Don’t confuse QA (answer extraction) with recommendation (ranking).
    • Ambiguity is a key challenge—mention context or disambiguation techniques.

Visual Summary

Based on the TU BSc CSIT syllabus for Information Retrieval (CSC413), unit 7.

Discussion

Loading…