CACS458 Knowledge Engineering

Knowledge EngineeringUnit 113 min read

Knowledge Engineering: Foundations, Representation & Real-World Systems

Unit 1 of Knowledge Engineering introduces the core concepts of knowledge engineering—its definition, goals, and the critical role of knowledge representation. This note covers the evolution from traditional AI to knowledge-based systems, key representation techniques (propositional logic, predicate logic, ontologies),

TAKEAWAYS:

  • Knowledge engineering bridges AI and domain expertise to build systems that reason (not just compute), using representations like logic, ontologies, and probabilistic models.
  • Knowledge representation (KRep) is the "language" for machines—propositional logic is rigid (true/false), predicate logic adds variables (e.g., Parent(X,Y)), and ontologies define shared vocabularies (e.g., NTC’s network topology).
  • Ontologies (e.g., OWL) are the backbone of semantic web apps like eSewa’s service catalog or Daraz’s product hierarchies, enabling machines to understand relationships (e.g., Order → Payment → Delivery).
  • Knowledge acquisition is the bottleneck: 80% of a system’s success depends on gathering accurate, structured data (e.g., Pathao’s driver-location rules vs. real-time traffic).
  • Applications span healthcare (diagnosis rules), finance (loan eligibility), and social media (sentiment analysis via ontologies of emotions).
  • Exam focus: Compare logic types, critique ontology limitations, and link theory to real systems (e.g., "How would you represent Kathmandu’s traffic rules as a predicate logic rule?").

1. What Is Knowledge Engineering?

Knowledge engineering (KE) is the art and science of designing systems that encode human expertise to solve complex problems. Unlike traditional AI (which relies on data patterns), KE focuses on explicit knowledge representation—structuring facts, rules, and relationships so machines can reason (not just predict).

Why Does It Matter?

  • Human knowledge is unstructured: A doctor’s diagnosis isn’t just data; it’s rules like "If fever + rash → suspect dengue" or "If X-ray shows fracture → immobilize joint".
  • Machines need "common sense": Google Maps doesn’t just plot routes—it understands "traffic jam → reroute" (a rule-based inference).
  • Domain-specific systems: eSewa’s bill payment workflow isn’t possible without encoding rules like "User → Select Service → Verify Payment → Confirm".

Visual: The Knowledge Engineering Pipeline

flowchart LR
    A["Domain Expertise\n(e.g., NTC engineer)"] --> B["Knowledge Acquisition\n(Interviews, documents, data)"]
    B --> C["Knowledge Representation\n(Logics, ontologies, rules)"]
    C --> D["Knowledge Base\n(e.g., Daraz’s product ontology)"]
    D --> E["Inference Engine\n(e.g., eSewa’s payment validator)"]
    E --> F["Decision/Output\n(e.g., ‘Payment failed: insufficient balance’)"]

Key Idea: KE turns explicit knowledge (rules, facts) into actionable systems.


2. Knowledge Representation: The "Language" for Machines

Knowledge representation (KRep) is how we encode information so computers can process it. Poor representation = garbage output. Good representation = systems like:

  • Ncell’s customer service chatbot: Uses rules like "If ‘battery drain’ → suggest ‘update software’".
  • NEPSE’s stock alerts: "If share price < 100 → trigger ‘buy’ signal" (propositional logic).

Types of Knowledge Representation

Type Example Strengths Weaknesses Real-World Use
Propositional Logic IF (Rainy AND Cold) THEN (Carry Umbrella) Simple, fast for binary decisions No variables (e.g., can’t say "Carry Umbrella for any person") Traffic light timings (NTC)
Predicate Logic Parent(X,Y) ∧ Adult(X) → CanVote(X) Handles variables, flexible rules Computationally expensive for large datasets Pathao’s driver assignment rules
Ontologies Class: Vehicle → Subclass: Bike → Property: max_speed=60 Shared vocabularies, hierarchical Requires manual curation (e.g., updating Daraz’s product categories) eSewa’s service ontology
Frames/Slots Person: [name: "Ramesh", age: 30, occupation: "Engineer"] Structured data, easy to extend Rigid schema (hard to modify) Bank loan applicant profiles
Semantic Networks Doctor → knows → Disease → treats → Patient Visual, intuitive relationships Scalability issues for large graphs Healthcare diagnosis systems

Worked Example: Representing Kathmandu Traffic Rules

Problem: How would you encode the rule "Vehicles >2 wheels must stop at red lights" in:

  1. Propositional logic?
  2. Predicate logic?

Solution:

  1. Propositional Logic (too rigid):

    IF (Light = Red AND VehicleType = Car) THEN (Stop)
    

    Fails for: Bikes, buses, or if the light turns yellow.

  2. Predicate Logic (flexible):

    Vehicle(X) ∧ Wheels(X, Y) ∧ Y > 1 ∧ Light(Location, Red) → MustStop(X)
    

    Works for: Any vehicle with >1 wheel (cars, buses) at any location.

Visual: Predicate Logic in Action

graph TD
    A["Vehicle(X)"] --> B["Wheels(X, 4)"]
    B --> C["Light(Thamel, Red)"]
    C --> D["MustStop(X)"]

Real-World Tie-In: NTC’s traffic management system uses predicate-like rules to flag violations (e.g., Vehicle(X) ∧ Speed(X, >80) → Fine(X)).


3. Ontologies: The "Dictionary" for Machines

An ontology is a formal, shared specification of a domain’s concepts, relationships, and constraints. Think of it as a thesaurus + grammar rules for machines.

Why Ontologies?

  • eSewa’s service catalog: Defines Service → Subclass: BillPayment → Property: due_date.
  • Daraz’s product search: Uses Product → has → Category → has → Subcategory to recommend items.
  • NEPSE’s stock data: Links Company → owns → Shares → has → Price.

Ontology Languages: OWL vs. RDF

Feature OWL (Web Ontology Language) RDF (Resource Description Framework)
Purpose Defines classes and relationships Describes resources and their properties
Example Class: Animal → Subclass: Dog → Property: barks Dog → hasProperty → barks → true
Use Case Complex hierarchies (e.g., healthcare) Simple data linking (e.g., social media profiles)
Tool Support Protégé, TopBraid SPARQL queries, GraphDB

Visual: OWL Ontology Graph (NEPSE Stock Data)

graph TD
    A["Company\n(e.g., NMB Bank)"] --> B["Owns\n(relationship)"]
    B --> C["Shares\n(Class)"]
    C --> D["Has\n(Property)"]
    D --> E["Price\n(Datatype: 120.50)"]
    C --> F["Has\n(Property)"]
    F --> G["Dividend\n(Datatype: 5.00)"]

Real-World Example: NEPSE’s official ontology (hypothetical) would link companies to shares, dividends, and market trends.


4. Knowledge Acquisition: The Hardest Part

Definition: The process of extracting knowledge from experts, documents, or data to build a knowledge base.

Methods

  1. Interviews: Ask domain experts (e.g., a cardiologist for a diagnosis system).
  2. Document Analysis: Mine textbooks, manuals (e.g., NTC’s traffic rules).
  3. Observation: Watch experts work (e.g., Pathao drivers handling rush-hour traffic).
  4. Machine Learning: Auto-extract patterns (e.g., WhatsApp’s spam detector learns from labeled messages).

Challenges

  • Bottleneck: Experts are busy (e.g., a doctor may not have time to document all rules).
  • Ambiguity: Natural language is vague (e.g., "usually works" vs. "always fails").
  • Scalability: Manual encoding is slow (e.g., updating Daraz’s 100,000+ product categories).

Worked Example: Acquiring Knowledge for a Loan Approval System Scenario: A bank wants to automate loan approvals. How would you gather rules?

Steps:

  1. Interview a loan officer:
    • "What’s the minimum salary for a 5-lakh loan?" → Salary(X) ≥ 50,000 → ApproveLoan(X, 500,000).
    • "What if the applicant has a default?" → DefaultHistory(X, true) → RejectLoan(X).
  2. Analyze past data:
    • "80% of loans >10 lakh are approved if CIBIL > 700" → LoanAmount(X, >1000000) ∧ CIBIL(X, >700) → 0.8 Probability(Approved).
  3. Validate with stakeholders:
    • "Does this cover all cases?" (e.g., co-applicants, collateral).

Visual: Knowledge Acquisition Workflow

flowchart LR
    A["Domain Expert\n(e.g., Bank Officer)"] --> B["Interviews\n(Structured questions)"]
    B --> C["Documents\n(Loan policies, past data)"]
    C --> D["Observation\n(Watch approval process)"]
    D --> E["Knowledge Base\n(Rules, exceptions)"]
    E --> F["Prototype\n(Test with sample cases)"]
    F --> G["Refine\n(Fix errors, add rules)"]

5. Applications of Knowledge Engineering

A. Healthcare: Diagnosis Systems

  • Example: A system that encodes rules like:
    Symptom(Fever) ∧ Symptom(Rash) ∧ Duration(>3 days) → Suspect(Dengue)
    
  • Real-World: Nepal’s Health Management Information System (HMIS) uses KE to flag outbreaks (e.g., "If 5+ cases in a week → Alert District Hospital").

B. Finance: Fraud Detection

  • Example: Rules like:
    Transaction(Amount, >100,000) ∧ Location(India) ∧ Time(3 AM) → Flag(Fraud)
    
  • Real-World: Nabil Bank’s KE system detects unusual transactions (e.g., a student suddenly transferring 5 lakh).

C. E-Commerce: Product Recommendations

  • Example: Ontology links:
    User(Bought: Laptop) → Suggest(Accessories: Mouse, Bag)
    
  • Real-World: Daraz’s "Frequently Bought Together" uses KE to infer relationships.

D. Social Media: Sentiment Analysis

  • Example: Ontology of emotions:
    Word("awful") → Sentiment(Negative) → Score(-2)
    Word("great") → Sentiment(Positive) → Score(+1)
    
  • Real-World: Pathao’s customer feedback analyzer uses KE to classify reviews.

## In the Real World

  1. eSewa’s Bill Payment Workflow

    • Idea Used: Predicate logic + ontologies
    • How: eSewa encodes rules like:
      User(LoggedIn) ∧ SelectService(Electricity) ∧ VerifyPayment(Success) → ConfirmTransaction
      
      The ontology defines Service → Subclasses (Electricity, Water, Telephone) and Payment → Methods (Khalti, eBanking).
  2. Ncell’s Customer Service Chatbot

    • Idea Used: Propositional logic for FAQs + NLP for open-ended questions
    • How: For "My phone won’t charge":
      • Check if the rule BatteryDead ∧ ChargingPortDamaged → ReplacePort matches.
      • If not, use NLP to extract symptoms (e.g., "light flashes but no power" → suggest cleaning port).
  3. NEPSE’s Stock Alerts

    • Idea Used: Predicate logic + probabilistic reasoning
    • How: A rule like:
      SharePrice(CompanyX, <100) ∧ Volume(>1000) ∧ Trend(Down3Days) → Alert("Buy Opportunity")
      
      But with uncertainty: "If CIBIL(CompanyX) < 600 → Reduce Alert Confidence by 30%".

## Exam Tip

  1. Compare and Contrast: Always link theory to real systems. For example:

    • "Propositional logic is like a traffic light (only 3 states: red/green/yellow), while predicate logic is like a GPS (handles variables like ‘destination’)."
    • "Ontologies are like a restaurant menu (shared structure), but frames are like a recipe card (fixed slots)."
  2. Worked Examples Are Gold:

    • If asked to "represent X as a predicate logic rule", always:
      1. Identify entities (e.g., Person, Loan).
      2. Define relationships (e.g., AppliesFor, Approves).
      3. Add constraints (e.g., Income(X) > 50,000).
    • Example Question: "How would you represent ‘A student gets a scholarship if their GPA > 3.5 and family income < 50,000’?" Answer:
      Student(X) ∧ GPA(X, >3.5) ∧ Income(FamilyOf(X), <50000) → EligibleForScholarship(X)
      
  3. Ontology Questions:

    • Expect questions like "How would you model a university ontology?"
    • Answer Structure:
      1. Classes: Person → Student, Faculty.
      2. Properties: Student → has → GPA, EnrolledIn.
      3. Relationships: Faculty → Teaches → Course.
      4. Constraints: GPA ∈ [0, 4].
  4. Applications:

    • Always tie to Nepali contexts. For example:
      • "How could KE improve NTC’s traffic management?" → Use predicate logic for Vehicle(X) ∧ Speed(X, >80) → Fine(X).
      • "How does Daraz use ontologies?" → Product hierarchies (Electronics → Laptops → Subclass: Gaming).
  5. Avoid Common Mistakes:

    • ❌ "Ontologies are just databases." → No! Ontologies define meaning (e.g., Dog is-a Animal).
    • ❌ "Knowledge acquisition is easy." → No! It’s the hardest part (80% of project time).
    • ❌ "Predicate logic is always better." → No! Use propositional logic for simple, fast decisions (e.g., traffic lights).

Final Visual: Knowledge Engineering in a Nepali App

mindmap
  root((Knowledge Engineering in Nepal))
    eSewa
      Predicate Logic: "User → Service → Payment → Confirm"
      Ontology: "Service → Subclasses (Electricity, Water)"
    Ncell Chatbot
      Propositional Rules: "Symptom → Suggest Fix"
      NLP: "Extract keywords from user input"
    Daraz
      Ontology: "Product → Category → Subcategory"
      Recommendation Engine: "If bought X → suggest Y"
    NTC
      Predicate Rules: "Vehicle → Speed → Fine"
      Traffic Simulation: "Model intersections as state space"

Based on the TU BCA syllabus for Knowledge Engineering (CACS458), unit 1.

Discussion

Loading…