Tribhuvan University
Bachelor of Science in Computer Science and Information Technology
Semester 7 · TU Board 2079
Course Title: Data Warehousing and Data Mining (CSC410)
Full Marks: 60Pass Marks: 24Time: 3 hours
Candidates are required to give their answers in their own words as far as practicable. The figures in the margin indicate full marks.
Group A
Attempt any two question.(2 × 10 = 20)
- 1.10
Discuss any two drawbacks of Apriori algorithm. Find frequent item-sets and association rules from the transaction database given below using FP-growth algorithm. Assume minimum support is 50% and minimum confidence is 60%.
Transaction_ID
Items purchased
1
Sausage, peanut, Beer
2
peanut, Beer, Apple
3
Apple, Milk
4
Sausage, peanut, Apple
5
Sausage, peanut, Beer, Milk
6
Sausage, peanut, Beer, AppleAnswer comingAlso asked in 2080
- 2.10
When multilayer perceptron is better choice over other classification algorithms? Consider a multilayer feed-forward neural network given below. Let the learning rate be 0.5. Assume initial values of weights and biases as given in the table below. Train the network for the training tuples (1, 1, 0) and (0, 1, 1), where last number is target output. Show weight and bias updates by using back-propagation algorithm. Assume that sigmoid activation function is used in the network.
w13
w14
w23
w24
w35
w45
b3
b4
b5
0.5
0.2
-0.3
0.5
0.1
0.3
0.6
-0.4
0.8 - 3.10
Why OLAP operations are used? Discuss various OLAP operation with suitable example of each.
Group B
Attempt any eight questions.(8 × 5 = 40)
- 4.5
Suppose that we have 5 dimensional data. What will be total number of cuboids generated? If we consider each dimension has 5 levels, what will be the number of cuboids generated?
- 5.5
Discuss different types of attributes with suitable example of each.
- 6.5
Why data normalization is important in data mining? Explain min-max and Z-score normalization approach.
- 7.5
What are two categories of hierarchical clustering? Divide the following data points into two clusters using agglomerative clustering.
{ {(2,10), ((2,5), (8,4), (5,8), (7,5), (6,4)) - 8.5
Discuss the concept of K-means++ and Mini-batch K-means algorithm.
Answer comingAlso asked in Model question
- 9.5
What is confusion matrix? Discuss various classification measures along with their mathematical formulae.
- 10.5
What are application areas of graph mining? Explain the concept behind inductive logic programming with suitable demonstration.
- 11.5
Discuss the concept of text mining with its practical implications.
- 12.5
Write down short notes on:
- Data Mart
- Market Basket Analysis
— The End —