Knowledge EngineeringUnit 113 min read
Knowledge Engineering: Foundations, Representation & Real-World Systems
Unit 1 of Knowledge Engineering introduces the core concepts of knowledge engineering—its definition, goals, and the critical role of knowledge representation. This note covers the evolution from traditional AI to knowledge-based systems, key representation techniques (propositional logic, predicate logic, ontologies),
TAKEAWAYS:
- Knowledge engineering bridges AI and domain expertise to build systems that reason (not just compute), using representations like logic, ontologies, and probabilistic models.
- Knowledge representation (KRep) is the "language" for machines—propositional logic is rigid (true/false), predicate logic adds variables (e.g.,
Parent(X,Y)), and ontologies define shared vocabularies (e.g., NTC’s network topology). - Ontologies (e.g., OWL) are the backbone of semantic web apps like eSewa’s service catalog or Daraz’s product hierarchies, enabling machines to understand relationships (e.g.,
Order → Payment → Delivery). - Knowledge acquisition is the bottleneck: 80% of a system’s success depends on gathering accurate, structured data (e.g., Pathao’s driver-location rules vs. real-time traffic).
- Applications span healthcare (diagnosis rules), finance (loan eligibility), and social media (sentiment analysis via ontologies of emotions).
- Exam focus: Compare logic types, critique ontology limitations, and link theory to real systems (e.g., "How would you represent Kathmandu’s traffic rules as a predicate logic rule?").
1. What Is Knowledge Engineering?
Knowledge engineering (KE) is the art and science of designing systems that encode human expertise to solve complex problems. Unlike traditional AI (which relies on data patterns), KE focuses on explicit knowledge representation—structuring facts, rules, and relationships so machines can reason (not just predict).
Why Does It Matter?
- Human knowledge is unstructured: A doctor’s diagnosis isn’t just data; it’s rules like "If fever + rash → suspect dengue" or "If X-ray shows fracture → immobilize joint".
- Machines need "common sense": Google Maps doesn’t just plot routes—it understands "traffic jam → reroute" (a rule-based inference).
- Domain-specific systems: eSewa’s bill payment workflow isn’t possible without encoding rules like "User → Select Service → Verify Payment → Confirm".
Visual: The Knowledge Engineering Pipeline
flowchart LR
A["Domain Expertise\n(e.g., NTC engineer)"] --> B["Knowledge Acquisition\n(Interviews, documents, data)"]
B --> C["Knowledge Representation\n(Logics, ontologies, rules)"]
C --> D["Knowledge Base\n(e.g., Daraz’s product ontology)"]
D --> E["Inference Engine\n(e.g., eSewa’s payment validator)"]
E --> F["Decision/Output\n(e.g., ‘Payment failed: insufficient balance’)"]Key Idea: KE turns explicit knowledge (rules, facts) into actionable systems.
2. Knowledge Representation: The "Language" for Machines
Knowledge representation (KRep) is how we encode information so computers can process it. Poor representation = garbage output. Good representation = systems like:
- Ncell’s customer service chatbot: Uses rules like "If ‘battery drain’ → suggest ‘update software’".
- NEPSE’s stock alerts: "If share price < 100 → trigger ‘buy’ signal" (propositional logic).
Types of Knowledge Representation
| Type | Example | Strengths | Weaknesses | Real-World Use |
|---|---|---|---|---|
| Propositional Logic | IF (Rainy AND Cold) THEN (Carry Umbrella) |
Simple, fast for binary decisions | No variables (e.g., can’t say "Carry Umbrella for any person") | Traffic light timings (NTC) |
| Predicate Logic | Parent(X,Y) ∧ Adult(X) → CanVote(X) |
Handles variables, flexible rules | Computationally expensive for large datasets | Pathao’s driver assignment rules |
| Ontologies | Class: Vehicle → Subclass: Bike → Property: max_speed=60 |
Shared vocabularies, hierarchical | Requires manual curation (e.g., updating Daraz’s product categories) | eSewa’s service ontology |
| Frames/Slots | Person: [name: "Ramesh", age: 30, occupation: "Engineer"] |
Structured data, easy to extend | Rigid schema (hard to modify) | Bank loan applicant profiles |
| Semantic Networks | Doctor → knows → Disease → treats → Patient |
Visual, intuitive relationships | Scalability issues for large graphs | Healthcare diagnosis systems |
Worked Example: Representing Kathmandu Traffic Rules
Problem: How would you encode the rule "Vehicles >2 wheels must stop at red lights" in:
- Propositional logic?
- Predicate logic?
Solution:
Propositional Logic (too rigid):
IF (Light = Red AND VehicleType = Car) THEN (Stop)Fails for: Bikes, buses, or if the light turns yellow.
Predicate Logic (flexible):
Vehicle(X) ∧ Wheels(X, Y) ∧ Y > 1 ∧ Light(Location, Red) → MustStop(X)Works for: Any vehicle with >1 wheel (cars, buses) at any location.
Visual: Predicate Logic in Action
graph TD
A["Vehicle(X)"] --> B["Wheels(X, 4)"]
B --> C["Light(Thamel, Red)"]
C --> D["MustStop(X)"]Real-World Tie-In: NTC’s traffic management system uses predicate-like rules to flag violations (e.g., Vehicle(X) ∧ Speed(X, >80) → Fine(X)).
3. Ontologies: The "Dictionary" for Machines
An ontology is a formal, shared specification of a domain’s concepts, relationships, and constraints. Think of it as a thesaurus + grammar rules for machines.
Why Ontologies?
- eSewa’s service catalog: Defines
Service → Subclass: BillPayment → Property: due_date. - Daraz’s product search: Uses
Product → has → Category → has → Subcategoryto recommend items. - NEPSE’s stock data: Links
Company → owns → Shares → has → Price.
Ontology Languages: OWL vs. RDF
| Feature | OWL (Web Ontology Language) | RDF (Resource Description Framework) |
|---|---|---|
| Purpose | Defines classes and relationships | Describes resources and their properties |
| Example | Class: Animal → Subclass: Dog → Property: barks |
Dog → hasProperty → barks → true |
| Use Case | Complex hierarchies (e.g., healthcare) | Simple data linking (e.g., social media profiles) |
| Tool Support | Protégé, TopBraid | SPARQL queries, GraphDB |
Visual: OWL Ontology Graph (NEPSE Stock Data)
graph TD
A["Company\n(e.g., NMB Bank)"] --> B["Owns\n(relationship)"]
B --> C["Shares\n(Class)"]
C --> D["Has\n(Property)"]
D --> E["Price\n(Datatype: 120.50)"]
C --> F["Has\n(Property)"]
F --> G["Dividend\n(Datatype: 5.00)"]Real-World Example: NEPSE’s official ontology (hypothetical) would link companies to shares, dividends, and market trends.
4. Knowledge Acquisition: The Hardest Part
Definition: The process of extracting knowledge from experts, documents, or data to build a knowledge base.
Methods
- Interviews: Ask domain experts (e.g., a cardiologist for a diagnosis system).
- Document Analysis: Mine textbooks, manuals (e.g., NTC’s traffic rules).
- Observation: Watch experts work (e.g., Pathao drivers handling rush-hour traffic).
- Machine Learning: Auto-extract patterns (e.g., WhatsApp’s spam detector learns from labeled messages).
Challenges
- Bottleneck: Experts are busy (e.g., a doctor may not have time to document all rules).
- Ambiguity: Natural language is vague (e.g., "usually works" vs. "always fails").
- Scalability: Manual encoding is slow (e.g., updating Daraz’s 100,000+ product categories).
Worked Example: Acquiring Knowledge for a Loan Approval System Scenario: A bank wants to automate loan approvals. How would you gather rules?
Steps:
- Interview a loan officer:
- "What’s the minimum salary for a 5-lakh loan?" →
Salary(X) ≥ 50,000 → ApproveLoan(X, 500,000). - "What if the applicant has a default?" →
DefaultHistory(X, true) → RejectLoan(X).
- "What’s the minimum salary for a 5-lakh loan?" →
- Analyze past data:
- "80% of loans >10 lakh are approved if CIBIL > 700" →
LoanAmount(X, >1000000) ∧ CIBIL(X, >700) → 0.8 Probability(Approved).
- "80% of loans >10 lakh are approved if CIBIL > 700" →
- Validate with stakeholders:
- "Does this cover all cases?" (e.g., co-applicants, collateral).
Visual: Knowledge Acquisition Workflow
flowchart LR
A["Domain Expert\n(e.g., Bank Officer)"] --> B["Interviews\n(Structured questions)"]
B --> C["Documents\n(Loan policies, past data)"]
C --> D["Observation\n(Watch approval process)"]
D --> E["Knowledge Base\n(Rules, exceptions)"]
E --> F["Prototype\n(Test with sample cases)"]
F --> G["Refine\n(Fix errors, add rules)"]5. Applications of Knowledge Engineering
A. Healthcare: Diagnosis Systems
- Example: A system that encodes rules like:
Symptom(Fever) ∧ Symptom(Rash) ∧ Duration(>3 days) → Suspect(Dengue) - Real-World: Nepal’s Health Management Information System (HMIS) uses KE to flag outbreaks (e.g., "If 5+ cases in a week → Alert District Hospital").
B. Finance: Fraud Detection
- Example: Rules like:
Transaction(Amount, >100,000) ∧ Location(India) ∧ Time(3 AM) → Flag(Fraud) - Real-World: Nabil Bank’s KE system detects unusual transactions (e.g., a student suddenly transferring 5 lakh).
C. E-Commerce: Product Recommendations
- Example: Ontology links:
User(Bought: Laptop) → Suggest(Accessories: Mouse, Bag) - Real-World: Daraz’s "Frequently Bought Together" uses KE to infer relationships.
D. Social Media: Sentiment Analysis
- Example: Ontology of emotions:
Word("awful") → Sentiment(Negative) → Score(-2) Word("great") → Sentiment(Positive) → Score(+1) - Real-World: Pathao’s customer feedback analyzer uses KE to classify reviews.
## In the Real World
eSewa’s Bill Payment Workflow
- Idea Used: Predicate logic + ontologies
- How: eSewa encodes rules like:
The ontology definesUser(LoggedIn) ∧ SelectService(Electricity) ∧ VerifyPayment(Success) → ConfirmTransactionService → Subclasses (Electricity, Water, Telephone)andPayment → Methods (Khalti, eBanking).
Ncell’s Customer Service Chatbot
- Idea Used: Propositional logic for FAQs + NLP for open-ended questions
- How: For "My phone won’t charge":
- Check if the rule
BatteryDead ∧ ChargingPortDamaged → ReplacePortmatches. - If not, use NLP to extract symptoms (e.g., "light flashes but no power" → suggest cleaning port).
- Check if the rule
NEPSE’s Stock Alerts
- Idea Used: Predicate logic + probabilistic reasoning
- How: A rule like:
But with uncertainty: "If CIBIL(CompanyX) < 600 → Reduce Alert Confidence by 30%".SharePrice(CompanyX, <100) ∧ Volume(>1000) ∧ Trend(Down3Days) → Alert("Buy Opportunity")
## Exam Tip
Compare and Contrast: Always link theory to real systems. For example:
- "Propositional logic is like a traffic light (only 3 states: red/green/yellow), while predicate logic is like a GPS (handles variables like ‘destination’)."
- "Ontologies are like a restaurant menu (shared structure), but frames are like a recipe card (fixed slots)."
Worked Examples Are Gold:
- If asked to "represent X as a predicate logic rule", always:
- Identify entities (e.g.,
Person,Loan). - Define relationships (e.g.,
AppliesFor,Approves). - Add constraints (e.g.,
Income(X) > 50,000).
- Identify entities (e.g.,
- Example Question: "How would you represent ‘A student gets a scholarship if their GPA > 3.5 and family income < 50,000’?"
Answer:
Student(X) ∧ GPA(X, >3.5) ∧ Income(FamilyOf(X), <50000) → EligibleForScholarship(X)
- If asked to "represent X as a predicate logic rule", always:
Ontology Questions:
- Expect questions like "How would you model a university ontology?"
- Answer Structure:
- Classes:
Person → Student, Faculty. - Properties:
Student → has → GPA, EnrolledIn. - Relationships:
Faculty → Teaches → Course. - Constraints:
GPA ∈ [0, 4].
- Classes:
Applications:
- Always tie to Nepali contexts. For example:
- "How could KE improve NTC’s traffic management?" → Use predicate logic for
Vehicle(X) ∧ Speed(X, >80) → Fine(X). - "How does Daraz use ontologies?" → Product hierarchies (
Electronics → Laptops → Subclass: Gaming).
- "How could KE improve NTC’s traffic management?" → Use predicate logic for
- Always tie to Nepali contexts. For example:
Avoid Common Mistakes:
- ❌ "Ontologies are just databases." → No! Ontologies define meaning (e.g.,
Dogis-aAnimal). - ❌ "Knowledge acquisition is easy." → No! It’s the hardest part (80% of project time).
- ❌ "Predicate logic is always better." → No! Use propositional logic for simple, fast decisions (e.g., traffic lights).
- ❌ "Ontologies are just databases." → No! Ontologies define meaning (e.g.,
Final Visual: Knowledge Engineering in a Nepali App
mindmap
root((Knowledge Engineering in Nepal))
eSewa
Predicate Logic: "User → Service → Payment → Confirm"
Ontology: "Service → Subclasses (Electricity, Water)"
Ncell Chatbot
Propositional Rules: "Symptom → Suggest Fix"
NLP: "Extract keywords from user input"
Daraz
Ontology: "Product → Category → Subcategory"
Recommendation Engine: "If bought X → suggest Y"
NTC
Predicate Rules: "Vehicle → Speed → Fine"
Traffic Simulation: "Model intersections as state space"Based on the TU BCA syllabus for Knowledge Engineering (CACS458), unit 1.
Discussion
Loading…