CSC410 Data Warehousing and Data Mining

Data Warehousing and Data Mining TU Board 2080 question paper

12 questionsSit this paper (timed)

Tribhuvan University

Bachelor of Science in Computer Science and Information Technology

Semester 7 · TU Board 2080

Course Title: Data Warehousing and Data Mining (CSC410)

Full Marks: 60Pass Marks: 24Time: 3 hours

Candidates are required to give their answers in their own words as far as practicable. The figures in the margin indicate full marks.

Group A

Attempt any TWO questions.(2 × 10 = 20)

  1. 1.

    State Apriori property. Find frequent item sets and association rules from the transaction database given below using Apriori algorithm. Assume min. support is 50% and min confidence is 75%.

    Transaction ID
    Items Purchased
    1
    Bread, Cheese, Egg, Juice
    2
    Bread, Cheese, Juice
    3
    Bread, Milk, Yogurt
    4
    Bread, Juice, Milk
    5
    Cheese, Juice, Milk

    10
  2. 2.

    How classification differs from regression. Train ID3 classifierusing the dataset given below. Then predict class label for the data [Age=Mid, Competition=Yes, Type=HW].

    [figure in the original paper]

    10
  3. 3.

    Why the concept of data mart is important? Discuss different data warehouse schema with examples.

    10

Group B

Attempt any EIGHT questions.(8 × 5 = 40)

  1. 4.

    How KDD differs from data mining? Explain various stages of KDD with suitable block diagram.

    5
  2. 5.

    Discuss different ways of smoothing noisy data along with suitable examples.

    5
  3. 6.

    How many cuboids are possible from 5-dimensional data? Discuss the concept of full cube and iceberg cube.

    5
  4. 7.

    How K-medoids clustering differs from K-means clustering? Divide the following data points into two clusters using kmedoids algorithm. Show computation up to 3 iterations. {(70,85), (65,80), (72,88), (75,90), (60,50), (64,55), (62,52), (63,58)}.

    5
  5. 8.

    Discuss working of DBSCAN algorithm.

    5
  6. 9.

    Which algorithm is used for training multi-layer perceptron? Discuss the algorithm in detail.

    5
  7. 10.

    Explain the OLAP operations with examples.

    5
  8. 11.

    Discuss the concept of multimedia data mining along with the concept of similarity search.

    5
  9. 12.

    Write down short notes on:

    1. Support Vector Machine
    2. Multi-dimensional Data Model
    5

— The End —