Information RetrievalUnit 711 min read
Question Answering Systems: QA Pipelines, Models & Real-World Apps
Unit 7 of Information Retrieval explores how QA systems process natural language queries to extract precise answers, covering pipelines, retrieval vs. generation models, and evaluation metrics—with real-world examples from eSewa, Ncell, and Google.
TAKEAWAYS:
- QA systems bridge the gap between unstructured text and direct answers by combining retrieval (finding relevant documents) and generation (extracting answers).
- Key components include question classification, candidate answer retrieval, and answer validation—often using TF-IDF, BM25, or neural models like BERT.
- Open-domain QA (e.g., Google’s search snippets) relies on pre-indexed corpora, while closed-domain QA (e.g., eSewa’s FAQs) uses domain-specific datasets.
- Evaluation metrics like Exact Match (EM) and F1 score measure precision/recall trade-offs in answer correctness.
- Real-world applications span customer support (Pathao’s chatbots), financial queries (Ncell’s billing), and legal research (Nepal Law Commission databases).
- Challenges include ambiguity, context dependency, and scalability—addressed via hybrid models (e.g., retrieval + fine-tuning).
1. What Are Question Answering (QA) Systems?
QA systems automate the process of answering natural language questions by extracting precise information from unstructured text. Unlike traditional search engines (which return documents), QA systems directly return answers—e.g., a date, entity, or short text span.
How QA Systems Work: The Pipeline
Step 1: Question Classification Categorize the question to determine the answer type (e.g., who, when, how). Example:
- "When was Tribhuvan University founded?" → Time-based question.
- "What is the capital of Nepal?" → Entity-based question.
Step 2: Candidate Answer Retrieval Use information retrieval (IR) techniques (e.g., BM25, TF-IDF) to fetch relevant documents or passages. For example:
- For "What is the interest rate for Ncell loans?", the system retrieves Ncell’s official loan policy documents.
Step 3: Answer Validation Rank candidate answers using:
- Lexical matching (exact word overlap).
- Semantic similarity (e.g., BERT embeddings).
- Contextual cues (e.g., negation, temporal references).
Step 4: Post-Processing Refine answers for readability (e.g., rephrasing, unit conversion). Example:
- Raw answer: "The university was established in 1959."
- Processed: "Tribhuvan University was founded in 1959."
2. Types of QA Systems
| Type | Description | Example | IR Technique Used |
|---|---|---|---|
| Open-Domain QA | Answers questions from general knowledge (e.g., Wikipedia, web). | Google’s "People Also Ask" snippets. | BM25, Dense Retrieval (DPR). |
| Closed-Domain QA | Answers questions from a specific dataset (e.g., FAQs, legal texts). | eSewa’s chatbot for bill payments. | TF-IDF, Keyword Matching. |
| Conversational QA | Maintains context across multiple turns (e.g., chatbots). | Pathao’s customer support bot. | Memory Networks, Transformers. |
| Multi-Hop QA | Requires reasoning across multiple documents (e.g., "Who wrote X after Y?"). | Ncell’s troubleshooting guides. | Graph Neural Networks (GNNs). |
3. Key Techniques in QA Systems
A. Retrieval-Based QA
How it works: Uses IR models (e.g., BM25) to fetch relevant passages, then extracts answers via rule-based matching or span selection.
Example: For "What is the deadline for TU exam form submission?", the system:
- Retrieves TU’s official exam notice.
- Extracts the date span: "The deadline is 2024-05-15."
Worked Example: TU Exam Deadline
Document (TU Notice): "Students must submit exam forms by **May 15, 2024**, at 11:59 PM." Query: "When is the last date for TU exam form submission?"- BM25 Score: High for the sentence containing "deadline" + "May 15".
- Answer Span: "May 15, 2024".
B. Generation-Based QA
- How it works: Uses seq2seq models (e.g., BERT, T5) to generate answers from scratch, even if the exact phrase isn’t in the corpus.
- Example: Google’s LaMDA generates conversational responses like:
- User: "How do I reset my Daraz password?"
- QA System: "Go to Daraz’s login page, click ‘Forgot Password,’ and enter your email."
C. Hybrid QA
- Combines retrieval + generation for robustness. Example:
- Retrieve top-5 passages for "What is the fine for late NTC bill payment?".
- Use BERT to re-rank and generate a concise answer: "A fine of 10% of the unpaid amount is applied after the due date."
4. Real-World Applications in Nepal
Example 1: eSewa’s Bill Payment QA
- Problem: Users ask "How to pay Ncell bill via eSewa?"
- QA Pipeline:
- Classification: Procedural question → triggers step-by-step guide.
- Retrieval: Fetches eSewa’s Ncell payment FAQ.
- Generation: Rephrases steps for clarity:
"1. Open eSewa app. 2. Select ‘Pay Bills.’ 3. Choose ‘Ncell.’ 4. Enter phone number and amount."
Example 2: Ncell’s Troubleshooting Chatbot
- Problem: "My Ncell 4G is not working."
- Multi-Hop QA:
- Retrieves documents: Ncell network outages, device settings, SIM issues.
- Ranks answers by relevance (e.g., "Check if your SIM is inserted properly").
- If unresolved, escalates to human agent.
Example 3: NEPSE Stock Query System
- Problem: "What was the closing price of NTC on 2023-12-01?"
- Closed-Domain QA:
- Retrieves NEPSE’s historical data CSV.
- Extracts: "NTC closed at Rs. 125.50."
5. Evaluation Metrics for QA Systems
| Metric | Definition | When to Use |
|---|---|---|
| Exact Match (EM) | % of answers exactly matching the ground truth. | Factoid QA (e.g., dates, names). |
| F1 Score | Harmonic mean of precision/recall for answer spans. | Extractive QA (e.g., SQuAD datasets). |
| BLEU | Measures n-gram overlap with reference answers (used in generation QA). | Open-ended answers (e.g., chatbots). |
| MRR (Mean Reciprocal Rank) | Average rank of the first correct answer in a list. | Retrieval-based QA. |
Worked Example: Evaluating a QA System for TU Exam Dates
- Ground Truth Answer: "The TU exam starts on June 1, 2024."
- System Answer: "Exams begin 1st June 2024."
- EM: 0 (format mismatch).
- F1: 0.95 (partial match on date).
6. Challenges and Solutions
| Challenge | Cause | Solution |
|---|---|---|
| Ambiguity | "What is the capital of Nepal?" (Kathmandu vs. historical capitals). | Use contextual embeddings (BERT) or disambiguation prompts. |
| Lack of Training Data | Low-resource languages (e.g., Nepali QA). | Few-shot learning or data augmentation. |
| Scalability | Slow retrieval for large corpora. | Approximate Nearest Neighbors (ANN) or distributed indexing. |
| Hallucinations | Generation models invent false answers. | Retrieval augmentation (hybrid models). |
7. Advanced Techniques
A. Text Shingling for Answer Extraction
- What it does: Splits text into overlapping shingles (e.g., 3-gram sequences) to find exact matches.
- Example: For the query "TU’s first chancellor", shingles might be:
- "first chancellor of TU" (from a historical document).
- Use Case: eSewa’s legal QA for contract clauses.
B. Rocchio Algorithm for QA Classification
- How it works: Adjusts query vectors based on positive/negative feedback.
- Example: Classify "process scheduling" as Operating System vs. Automata:
- Compute centroids for each class from training data.
- Update query vector:
- Assign to the class with the highest cosine similarity.
Worked Example: Classifying "process scheduling"
| Class | Relevance Feedback | Updated Query Vector | Similarity Score |
|---|---|---|---|
| Operating System | High | Closer to OS centroid | 0.85 |
| Automata | Low | Far from Automata centroid | 0.30 |
Answer: Classified as Operating System.
8. Question Answering vs. Recommendation Systems
| Feature | Question Answering (QA) | Recommendation Systems |
|---|---|---|
| Goal | Extract precise answers from text. | Predict user preferences (e.g., Daraz products). |
| Input | Natural language questions. | User behavior (clicks, ratings). |
| Output | Text spans, entities, or facts. | Item rankings (e.g., "Top 5 laptops"). |
| Key Techniques | TF-IDF, BERT, Span Extraction. | Collaborative Filtering, Deep Learning. |
| Example in Nepal | Ncell’s chatbot for bill queries. | Daraz’s "Customers Also Bought" section. |
Exam Tip
For definitions:
- QA systems retrieve + generate answers; retrieval-based uses IR, generation-based uses seq2seq.
- Text shingling = overlapping n-grams for exact match retrieval.
For algorithms:
- Rocchio: Update query vectors using , , weights.
- BM25: Prefer documents with high-term frequency but low collection frequency.
For applications:
- Open-domain: Google, Wikipedia.
- Closed-domain: eSewa, Ncell FAQs.
- Multi-hop: NTC’s troubleshooting guides.
For evaluation:
- EM = exact match rate (strict).
- F1 = balance between precision/recall (lenient).
Common pitfalls:
- Don’t confuse QA (answer extraction) with recommendation (ranking).
- Ambiguity is a key challenge—mention context or disambiguation techniques.
Visual Summary
Based on the TU BSc CSIT syllabus for Information Retrieval (CSC413), unit 7.
Discussion
Loading…