0% found this document useful (0 votes)

61 views21 pages

Naive Bayes and Decision Tree Classification

Decision Trees are popular in machine learning for their intuitive structure that mimics human decision-making, making them easy to understand. They operate by recursively splitting datasets based on attribute values until reaching leaf nodes, with techniques like Information Gain and Gini Index used for attribute selection. Random Forest, an ensemble method of multiple decision trees, enhances predictive accuracy and mitigates overfitting, making it suitable for both classification and regression tasks.

Uploaded by

priyajecay

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

61 views21 pages

Naive Bayes and Decision Tree Classification

Uploaded by

priyajecay

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 21

Why use Decision Trees?

There are various algorithms in Machine learning, so choosing the best algorithm for the given
dataset and problem is the main point to remember while creating a machine learning model.
Below are the two reasons for using the Decision tree:

o Decision Trees usually mimic human thinking ability while making a decision, so it is
easy to understand.
o The logic behind the decision tree can be easily understood because it shows a tree-like
structure.

Decision Tree Terminologies

 Root Node: Root node is from where the decision tree starts. It represents the entire
dataset, which further gets divided into two or more homogeneous sets.
 Leaf Node: Leaf nodes are the final output node, and the tree cannot be segregated
further after getting a leaf node.
 Splitting: Splitting is the process of dividing the decision node/root node into sub-nodes
according to the given conditions.
 Branch/Sub Tree: A tree formed by splitting the tree.
 Pruning: Pruning is the process of removing the unwanted branches from the tree.
 Parent/Child node: The root node of the tree is called the parent node, and other nodes
are called the child nodes.

How does the Decision Tree algorithm Work?

In a decision tree, for predicting the class of the given dataset, the algorithm starts from the root
node of the tree. This algorithm compares the values of root attribute with the record (real
dataset) attribute and, based on the comparison, follows the branch and jumps to the next node.

For the next node, the algorithm again compares the attribute value with the other sub-nodes
and move further. It continues the process until it reaches the leaf node of the tree. The complete
process can be better understood using the below algorithm:
o Step-1: Begin the tree with the root node, says S, which contains the complete dataset.
o Step-2: Find the best attribute in the dataset using Attribute Selection Measure
(ASM).
o Step-3: Divide the S into subsets that contains possible values for the best attributes.
o Step-4: Generate the decision tree node, which contains the best attribute.
o Step-5: Recursively make new decision trees using the subsets of the dataset created in
step -3. Continue this process until a stage is reached where you cannot further classify
the nodes and called the final node as a leaf node.

Example: Suppose there is a candidate who has a job offer and wants to decide whether he
should accept the offer or Not. So, to solve this problem, the decision tree starts with the root
node (Salary attribute by ASM). The root node splits further into the next decision node
(distance from the office) and one leaf node based on the corresponding labels. The next
decision node further gets split into one decision node (Cab facility) and one leaf node. Finally,
the decision node splits into two leaf nodes (Accepted offers and Declined offer). Consider the
below diagram:

Attribute Selection Measures

While implementing a Decision tree, the main issue arises that how to select the best attribute
for the root node and for sub-nodes. So, to solve such problems there is a technique which is
called as Attribute selection measure or ASM. By this measurement, we can easily select the
best attribute for the nodes of the tree. There are two popular techniques for ASM, which are:

o Information Gain
o Gini Index

1. Information Gain:
o Information gain is the measurement of changes in entropy after the segmentation of a
dataset based on an attribute.
o It calculates how much information a feature provides us about a class.
o According to the value of information gain, we split the node and build the decision tree.
o A decision tree algorithm always tries to maximize the value of information gain, and a
node/attribute having the highest information gain is split first. It can be calculated using
the below formula:

1. Information Gain= Entropy(S)- [(Weighted Avg) *Entropy(each feature)

Entropy: Entropy is a metric to measure the impurity in a given attribute. It specifies

randomness in data. Entropy can be calculated as:

Entropy(s)= -P(yes)log2 P(yes)- P(no) log2 P(no)

Where,

o S= Total number of samples

o P(yes)= probability of yes
o P(no)= probability of no

2. Gini Index:
o Gini index is a measure of impurity or purity used while creating a decision tree in the
CART(Classification and Regression Tree) algorithm.
o An attribute with the low Gini index should be preferred as compared to the high Gini
index.
o It only creates binary splits, and the CART algorithm uses the Gini index to create
binary splits.
o Gini index can be calculated using the below formula:

Gini Index= 1- ∑jPj2

Pruning: Getting an Optimal Decision tree

Pruning is a process of deleting the unnecessary nodes from a tree in order to get the optimal
decision tree.

A too-large tree increases the risk of overfitting, and a small tree may not capture all the
important features of the dataset. Therefore, a technique that decreases the size of the learning
tree without reducing accuracy is known as Pruning. There are mainly two types of
tree pruning technology used:

o Cost Complexity Pruning

o Reduced Error Pruning.

Advantages of the Decision Tree

o It is simple to understand as it follows the same process which a human follow while
making any decision in real-life.
o It can be very useful for solving decision-related problems.
o It helps to think about all the possible outcomes for a problem.
o There is less requirement of data cleaning compared to other algorithms.

Disadvantages of the Decision Tree

o The decision tree contains lots of layers, which makes it complex.
o It may have an overfitting issue, which can be resolved using the Random Forest
algorithm.
o For more class labels, the computational complexity of the decision tree may increase.

Random Forest Algorithm

Random Forest is a popular machine learning algorithm that belongs to the supervised learning
technique. It can be used for both Classification and Regression problems in ML. It is based on
the concept of ensemble learning, which is a process of combining multiple classifiers to solve
a complex problem and to improve the performance of the model.

As the name suggests, "Random Forest is a classifier that contains a number of decision trees
on various subsets of the given dataset and takes the average to improve the predictive
accuracy of that dataset." Instead of relying on one decision tree, the random forest takes the
prediction from each tree and based on the majority votes of predictions, and it predicts the final
output.

The greater number of trees in the forest leads to higher accuracy and prevents the
problem of overfitting.

The below diagram explains the working of the Random Forest algorithm:

Note:
the Decision
To better
Treeunderstand
Algorithm.the Random Forest Algorithm, you should have knowledge of
Assumptions for Random Forest

Since the random forest combines multiple trees to predict the class of the dataset, it is possible
that some decision trees may predict the correct output, while others may not. But together, all
the trees predict the correct output. Therefore, below are two assumptions for a better Random
forest classifier:

o There should be some actual values in the feature variable of the dataset so that the
classifier can predict accurate results rather than a guessed result.
o The predictions from each tree must have very low correlations.

Why use Random Forest?

Below are some points that explain why we should use the Random Forest algorithm:

<="" li="">
o It takes less training time as compared to other algorithms.
o It predicts output with high accuracy, even for the large dataset it runs efficiently.
o It can also maintain accuracy when a large proportion of data is missing.

How does Random Forest algorithm work?

Random Forest works in two-phase first is to create the random forest by combining N decision
tree, and second is to make predictions for each tree created in the first phase.

The Working process can be explained in the below steps and diagram:

Step-1: Select random K data points from the training set.

Step-2: Build the decision trees associated with the selected data points (Subsets).

Step-3: Choose the number N for decision trees that you want to build.

Step-4: Repeat Step 1 & 2.

Step-5: For new data points, find the predictions of each decision tree, and assign the new data
points to the category that wins the majority votes.

The working of the algorithm can be better understood by the below example:

Example: Suppose there is a dataset that contains multiple fruit images. So, this dataset is given
to the Random forest classifier. The dataset is divided into subsets and given to each decision
tree. During the training phase, each decision tree produces a prediction result, and when a new
data point occurs, then based on the majority of results, the Random Forest classifier predicts
the final decision. Consider the below image:
Applications of Random Forest

There are mainly four sectors where Random forest mostly used:

1. Banking: Banking sector mostly uses this algorithm for the identification of loan risk.
2. Medicine: With the help of this algorithm, disease trends and risks of the disease can be
identified.
3. Land Use: We can identify the areas of similar land use by this algorithm.
4. Marketing: Marketing trends can be identified using this algorithm.

Advantages of Random Forest

o Random Forest is capable of performing both Classification and Regression tasks.
o It is capable of handling large datasets with high dimensionality.
o It enhances the accuracy of the model and prevents the overfitting issue.

Disadvantages of Random Forest

o Although random forest can be used for both classification and regression tasks, it is not
more suitable for Regression tasks.
Naiver Bayes
Bayesian Decision Theory

Bayesian framework assumes that we always have a prior distribution for

everything.
– The prior may be very vague.
– When we see some data, we combine our prior
distribution with a likelihood termto get a posterior
distribution.
– The likelihood term takes into account how
probable the observed data is giventhe parameters
of the model.
• It favors parameter settings that make the data likely.
• It fights the prior
• With enough data the likelihood terms always win.
Given database:
Example:
Losses and Risks:
Discriminant Functions

Classification can also be seen as implementing a set of discriminant functions,

gi(x), i = 1, . . K, such that we

Building model Using Naiver Bayes

Bayesian networks
A Bayesian network, Bayes network, belief network, Bayes(ian)
model or probabilistic directed acyclic graphical model is a probabilistic graphical model (a
type of statistical model) that represents a set of variables and their conditional
dependencies via a directed acyclic graph (DAG). For example, a Bayesian network could
represent the probabilistic relationships between diseases and symptoms. Given symptoms, the
network can be used to compute the probabilities of the presence of various diseases.
Bayesian Net Example:
Consider the following Bayesian network:

Thus, the independence expressed in this Bayesian net are thatA and B
are (absolutely) independent.
C is independent of B given A.
D is independent of C given A and B.
E is independent of A, B, and D given C.

Suppose that the net further records the following probabilities:

Some sample computations:

Prob(D=T):

P(D=T) =

P(D=T,A=T,B=T) + P(D=T,A=T,B=F) + P(D=T,A=F,B=T) + P(D=T,A=F,B=F) =

P(D=T|A=T,B=T) P(A=T,B=T) + P(D=T|A=T,B=F) P(A=T,B=F) +P(D=T|A=F,B=T)
P(A=F,B=T) + P(D=T|A=F,B=F) P(A=F,B=F) =
(since A and B are independent absolutely)

P(D=T|A=T,B=T) P(A=T) P(B=T) + P(D=T|A=T,B=F) P(A=T) P(B=F) +P(D=T|A=F,B=T)

P(A=F) P(B=T) + P(D=T|A=F,B=F) P(A=F) P(B=F) =

0.70.30.6 + 0.80.30.4 + 0.10.70.6 + 0.20.70.4 = 0.32

Prob(A=T|C=T):
P(A=T|C=T) = P(C=T|A=T)P(A=T) / P(C=T).

Now, P(C=T) = P(C=T,A=T) + P(C=T,A=F) =

P(C=T|A=T)P(A=T) + P(C=T|A=F)P(A=F) =
0.8*0.3+ 0.4*0.7 = 0.52

So P(C=T|A=T)P(A=T) / P(C=T) = 0.8*0.3/0.52= 0.46.

Association rule
Association rule mining is explained using the Apriori Algorithm.
Apriori Algorithm:
Confidence:

Parametric Methods

Parametric Estimation

 X = { xt }t where xt ~ p (x)
 Parametric estimation:

Assume a form for p (x | θ) and estimate θ, its sufficient statistics,using X

e.g., N ( μ, σ2) where θ = { μ, σ2}

Maximum Likelihood Estimation:

Likelihood of θ given the sample X

l (θ|X) = p (X |θ) = ∏t p (xt|θ)

Log likelihood

L(θ|X) = log l (θ|X) = ∑t log p (xt|θ)

Maximum likelihood estimator (MLE)

θ* = argmaxθ L(θ|X)

Examples: Bernoulli/Multinomial:

Gaussian (Normal) Distribution:

Bias and Variance:

Classification
Regression
Linear Regression:

Other Error Measures:

Decision Tree Classification Algorithm
No ratings yet
Decision Tree Classification Algorithm
30 pages
2.unit 2
No ratings yet
2.unit 2
23 pages
Decision Tree & Random ForestNotes
No ratings yet
Decision Tree & Random ForestNotes
11 pages
Lecture-4 Unit 2
No ratings yet
Lecture-4 Unit 2
73 pages
Unit 4
No ratings yet
Unit 4
33 pages
DS Unit - 4
No ratings yet
DS Unit - 4
76 pages
Decision Tree Classification Algorithm
No ratings yet
Decision Tree Classification Algorithm
4 pages
NOTES
No ratings yet
NOTES
18 pages
Unit 3
No ratings yet
Unit 3
25 pages
Chapter 4classification and Prediction
No ratings yet
Chapter 4classification and Prediction
19 pages
Decision Tree and Random Forest
No ratings yet
Decision Tree and Random Forest
41 pages
Tree
No ratings yet
Tree
7 pages
Decision Tree
No ratings yet
Decision Tree
11 pages
Lecture Note #5 - PEC-CS701E
No ratings yet
Lecture Note #5 - PEC-CS701E
16 pages
Decision Tree Algorithm in Machine Learning
No ratings yet
Decision Tree Algorithm in Machine Learning
17 pages
Decision Tree (Autosaved)
No ratings yet
Decision Tree (Autosaved)
14 pages
Decision Tree Classification Algorithm
No ratings yet
Decision Tree Classification Algorithm
11 pages
Deciosn Tree
No ratings yet
Deciosn Tree
5 pages
Decision Tree Classification Algorithm
No ratings yet
Decision Tree Classification Algorithm
14 pages
Decsion Tree
No ratings yet
Decsion Tree
6 pages
Decision Tree Algorithm
No ratings yet
Decision Tree Algorithm
5 pages
Decision Tree Algorithm Guide
No ratings yet
Decision Tree Algorithm Guide
10 pages
DS Tech M 3 1
No ratings yet
DS Tech M 3 1
13 pages
Decision Trees
No ratings yet
Decision Trees
3 pages
CSL0777 L25
No ratings yet
CSL0777 L25
39 pages
Decision Tree Classification Algorithm
No ratings yet
Decision Tree Classification Algorithm
14 pages
Main Algorithms Used in Machine Learning Lecture Notes
No ratings yet
Main Algorithms Used in Machine Learning Lecture Notes
26 pages
Decision Tree & Random Forest
No ratings yet
Decision Tree & Random Forest
34 pages
Lecture Notes 3
No ratings yet
Lecture Notes 3
11 pages
Lecture 7.1 - Decision Tree Classification
No ratings yet
Lecture 7.1 - Decision Tree Classification
15 pages
UNIT 2 - Groups (Decision Tree)
No ratings yet
UNIT 2 - Groups (Decision Tree)
20 pages
Ch5 Data Science
No ratings yet
Ch5 Data Science
60 pages
Lab 2
No ratings yet
Lab 2
3 pages
Decision Tree and Random Forest Overview
No ratings yet
Decision Tree and Random Forest Overview
15 pages
Supervised Learning Algorithm DT
No ratings yet
Supervised Learning Algorithm DT
15 pages
Decision Tree
No ratings yet
Decision Tree
7 pages
Unit 3.2 Decision Tree Algorithm Wit Examples
No ratings yet
Unit 3.2 Decision Tree Algorithm Wit Examples
85 pages
1822 B.E Cse Batchno 149
No ratings yet
1822 B.E Cse Batchno 149
66 pages
13.decision Tree
No ratings yet
13.decision Tree
29 pages
What Is Decision Tree
No ratings yet
What Is Decision Tree
35 pages
Decision Trees for Data Enthusiasts
No ratings yet
Decision Trees for Data Enthusiasts
52 pages
Decision Trees and Probabilistic Models
No ratings yet
Decision Trees and Probabilistic Models
32 pages
Classification 4
No ratings yet
Classification 4
16 pages
Decisiontree
No ratings yet
Decisiontree
6 pages
08 Decision - Tree
No ratings yet
08 Decision - Tree
9 pages
Decision Tree
No ratings yet
Decision Tree
16 pages
Decision Tree
No ratings yet
Decision Tree
24 pages
Decision Trees - A Complete Introduction With Examples - by Shubham Koli - Medium
No ratings yet
Decision Trees - A Complete Introduction With Examples - by Shubham Koli - Medium
22 pages
Unit-3 Introduction To Machine Learning Algorithms
No ratings yet
Unit-3 Introduction To Machine Learning Algorithms
18 pages
2179 Unit 3
No ratings yet
2179 Unit 3
29 pages
FMLanswerkey-IT 2
No ratings yet
FMLanswerkey-IT 2
11 pages
ML Unit 3 Qa
No ratings yet
ML Unit 3 Qa
26 pages
Unit-4 (1) .Docx ML
No ratings yet
Unit-4 (1) .Docx ML
42 pages
Unit 1 ML (DT)
No ratings yet
Unit 1 ML (DT)
24 pages
DMDW 04
No ratings yet
DMDW 04
10 pages
Decision Tree
0% (1)
Decision Tree
16 pages
Cours #4-Decision Tree
No ratings yet
Cours #4-Decision Tree
18 pages
Decision Tree
No ratings yet
Decision Tree
5 pages
An Introduction To Random Forest Algorithm For Beginners
No ratings yet
An Introduction To Random Forest Algorithm For Beginners
16 pages
Algorithm For Estimating The Degree of Polymerization of Paper Insulation Impregnated With Inhibited Insulating Oil
No ratings yet
Algorithm For Estimating The Degree of Polymerization of Paper Insulation Impregnated With Inhibited Insulating Oil
6 pages
Associative Classifier Simplified
No ratings yet
Associative Classifier Simplified
8 pages
Machine Learning for Heart Failure Prediction
No ratings yet
Machine Learning for Heart Failure Prediction
15 pages
Group 12
No ratings yet
Group 12
54 pages
Thesis
No ratings yet
Thesis
49 pages
Module 5
No ratings yet
Module 5
16 pages
Classification Unit-4
No ratings yet
Classification Unit-4
19 pages
9 Lecture AI 09
No ratings yet
9 Lecture AI 09
57 pages
Understanding Decision Trees and ID3
No ratings yet
Understanding Decision Trees and ID3
4 pages
Dwdm-Unit-3 R20
No ratings yet
Dwdm-Unit-3 R20
11 pages
C2 W4 Lab 01 Decision Trees
No ratings yet
C2 W4 Lab 01 Decision Trees
6 pages
Assignment 2 Solution
No ratings yet
Assignment 2 Solution
4 pages
Classification and Prediction Techniques
No ratings yet
Classification and Prediction Techniques
41 pages
AIML 4th SEM ML Lab Manual
No ratings yet
AIML 4th SEM ML Lab Manual
33 pages
AI Unit 4 NEW
No ratings yet
AI Unit 4 NEW
60 pages
2019 SCC361 Questions
No ratings yet
2019 SCC361 Questions
6 pages
Unit 3
No ratings yet
Unit 3
27 pages
DWDM Notes 1
No ratings yet
DWDM Notes 1
62 pages
UNIT 4 Supervised Learning
No ratings yet
UNIT 4 Supervised Learning
38 pages
Crime Data Mediante Machine Learning
No ratings yet
Crime Data Mediante Machine Learning
6 pages
A Study of Suspicious E-Mail Detection Techniques
No ratings yet
A Study of Suspicious E-Mail Detection Techniques
8 pages
Naive Bayes and Decision Tree Classification
No ratings yet
Naive Bayes and Decision Tree Classification
21 pages
AD3461 ML Lab Manual
No ratings yet
AD3461 ML Lab Manual
32 pages
Decision Tree
No ratings yet
Decision Tree
36 pages
ML CLASS 6 Decision Tree Algorithm
No ratings yet
ML CLASS 6 Decision Tree Algorithm
21 pages
ML Unit2
No ratings yet
ML Unit2
38 pages
Credit Card Fraud Detection Using Naive Bayesian and C4.5 Decision
No ratings yet
Credit Card Fraud Detection Using Naive Bayesian and C4.5 Decision
5 pages
Machine Learning Data Classification Guide
No ratings yet
Machine Learning Data Classification Guide
156 pages
Gemini v. ChatGPT v. Mistral
No ratings yet
Gemini v. ChatGPT v. Mistral
23 pages
Introduction To Machine Learning 9
No ratings yet
Introduction To Machine Learning 9
3 pages

Naive Bayes and Decision Tree Classification

Uploaded by

Naive Bayes and Decision Tree Classification

Uploaded by

Why use Decision Trees?

Decision Tree Terminologies

How does the Decision Tree algorithm Work?

Attribute Selection Measures

1. Information Gain= Entropy(S)- [(Weighted Avg) *Entropy(each feature)

Entropy: Entropy is a metric to measure the impurity in a given attribute. It specifies

Entropy(s)= -P(yes)log2 P(yes)- P(no) log2 P(no)

o S= Total number of samples

Gini Index= 1- ∑jPj2

o Cost Complexity Pruning

Advantages of the Decision Tree

Disadvantages of the Decision Tree

Random Forest Algorithm

Why use Random Forest?

How does Random Forest algorithm work?

Step-1: Select random K data points from the training set.

Step-4: Repeat Step 1 & 2.

Advantages of Random Forest

Disadvantages of Random Forest

Bayesian framework assumes that we always have a prior distribution for

Classification can also be seen as implementing a set of discriminant functions,

Building model Using Naiver Bayes

Suppose that the net further records the following probabilities:

Some sample computations:

P(D=T,A=T,B=T) + P(D=T,A=T,B=F) + P(D=T,A=F,B=T) + P(D=T,A=F,B=F) =

P(D=T|A=T,B=T) P(A=T) P(B=T) + P(D=T|A=T,B=F) P(A=T) P(B=F) +P(D=T|A=F,B=T)

0.7*0.3*0.6 + 0.8*0.3*0.4 + 0.1*0.7*0.6 + 0.2*0.7*0.4 = 0.32

Now, P(C=T) = P(C=T,A=T) + P(C=T,A=F) =

So P(C=T|A=T)P(A=T) / P(C=T) = 0.8*0.3/0.52= 0.46.

Assume a form for p (x | θ) and estimate θ, its sufficient statistics,using X

e.g., N ( μ, σ2) where θ = { μ, σ2}

Maximum Likelihood Estimation:

Likelihood of θ given the sample X

L(θ|X) = log l (θ|X) = ∑t log p (xt|θ)

Maximum likelihood estimator (MLE)

Gaussian (Normal) Distribution:

Other Error Measures:

You might also like

0.70.30.6 + 0.80.30.4 + 0.10.70.6 + 0.20.70.4 = 0.32