BTCE | 5th Sem
AiML SubjectUnit 3

(5th sem) AIML Chapter 3: Complete Concepts Guide

Chapter 3: Decision Tree Learning and Artificial Neural Networks -> Generated and Prepared By Thiruselvan (ThiruXD)

13.1 Introduction to Decision Tree Learning

Definition: Decision Tree Learning is a supervised learning technique that constructs a tree-structured model for classification and prediction by recursively partitioning training data based on attribute values.

Key Concepts:

  • Machine Learning Classification: The process of assigning data instances to predefined classes by learning a mapping between input attributes and output labels.
  • Decision Tree: A tree-structured model consisting of nodes and branches representing decision rules.
  • Supervised Learning: Learning from labeled training examples where the correct output is known.

How It Works:

  1. Start with the entire dataset at the root node.
  2. Select the best attribute to split the data based on a splitting criterion.
  3. Partition the data into subsets based on attribute values.
  4. Repeat recursively for each subset until a stopping condition is met.
  5. Leaf nodes provide the final classification.

Advantages of Decision Trees:

AdvantageDescription
Easy InterpretationProduces human-readable rules
Minimal Data PreparationRequires little preprocessing
Handles Mixed DataWorks with numerical and categorical attributes
Efficient LearningComputationally efficient for many applications
Decision SupportUseful for business and medical decisions

Classification Process:

Training Data → Learning Algorithm → Classification Model → New Instance → Prediction

13.2 Decision Tree Components

Four Structural Components:

ComponentDescriptionExample
Root NodeTopmost node; represents entire dataset; contains most informative attributeWeather
Internal NodesIntermediate decision points; test on an attributeHumidity, Wind
BranchesConnect nodes; represent outcomes of attribute testsSunny, Rainy
Leaf NodesTerminal nodes; provide final classificationPlay = Yes/No

Structure:

                [Root Node: Weather]
               /         |          \
          Sunny       Rainy        Cloudy
            |           |            |
    [Internal: Humidity] [Leaf: Play=No] [Internal: Wind]
         /      \                        /      \
      High     Low                   Strong    Weak
       |        |                       |         |
   [Leaf: No] [Leaf: Yes]          [Leaf: No] [Leaf: Yes]

13.3 Applications of Decision Trees

DomainApplication
ClassificationDisease diagnosis, spam detection, credit approval
Decision SupportBusiness planning, risk assessment, financial forecasting
Data MiningCustomer behavior analysis, fraud detection, sales forecasting

13.4 Characteristics of Decision Tree Learning

CharacteristicDescription
InterpretabilityClearly shows which attributes are important and why decisions are made
SimplicityEasy to construct, visualize, and explain
Predictive CapabilityAchieves high accuracy through recursive partitioning and nonlinear modeling

UNIT 14: DECISION TREE REPRESENTATION

14.1 Decision Tree Representation

Tree Structure:

  • Hierarchical arrangement of nodes and branches
  • Starting from root node, data instances travel through branches based on attribute values
  • Ending at leaf nodes that provide final classification

Attribute-Based Splitting:

  • Process of dividing a dataset into subsets based on attribute values
  • Objective: Create subsets more homogeneous than the original dataset
  • Attribute selection uses measures like Entropy and Information Gain

Splitting Process:

Dataset → Weather? → Sunny / Rainy / Cloudy
                     ↓
              Subset1 / Subset2 / Subset3

14.2 Appropriate Problems for Decision Tree Learning

CharacteristicDescriptionExample
Attribute-Value InstancesFixed collections of attribute-value pairsOutlook, Temperature, Humidity, Wind
Discrete Target FunctionsTarget variable consists of discrete categoriesYes/No, Pass/Fail, Approved/Rejected
Noisy Training DataCan tolerate moderate errors and inconsistenciesMedical records with minor inaccuracies

14.3 Entropy

Definition: Entropy measures the impurity, uncertainty, or disorder in a dataset.

Mathematical Formula:

Entropy(S)=−∑i=1cpilog⁡2(pi)Entropy(S) = -\sum_{i=1}^{c} p_i \log_2(p_i)

Where:

  • pip_i = proportion of instances belonging to class ii
  • cc = number of classes

Interpretation:

Entropy ValueMeaning
0Pure dataset (all instances same class)
0.5Partially mixed
1.0Maximum disorder (equal distribution in binary)

Example Calculation:

  • Dataset: 9 Positive, 5 Negative (Total = 14)
  • P(Positive) = 9/14, P(Negative) = 5/14
  • Entropy = −(9/14)log₂(9/14) − (5/14)log₂(5/14) ≈ 0.94

14.4 Information Gain

Definition: Information Gain measures the reduction in entropy achieved after splitting the dataset on a particular attribute.

Mathematical Formula:

Gain(S,A)=Entropy(S)−∑v∈Values(A)∣Sv∣∣S∣Entropy(Sv)Gain(S, A) = Entropy(S) - \sum_{v \in Values(A)} \frac{|S_v|}{|S|} Entropy(S_v)

Where:

  • SS = dataset
  • AA = attribute
  • SvS_v = subset of SS where attribute A=vA = v

Attribute Selection Process:

  1. Calculate dataset entropy
  2. Split dataset using an attribute
  3. Calculate entropy of each subset
  4. Compute weighted average entropy
  5. Subtract from original entropy
  6. Select attribute with highest Information Gain

Example:

AttributeInformation Gain
Outlook0.246
Temperature0.029
Humidity0.151
Wind0.048

UNIT 15: ID3 ALGORITHM

15.1 Introduction to ID3 Algorithm

History: Developed by Ross Quinlan in 1986. One of the earliest and most influential decision tree learning algorithms.

Purpose:

  • Construct decision tree by repeatedly selecting attribute with highest Information Gain
  • Minimize classification errors
  • Reduce uncertainty
  • Generate understandable rules
  • Build compact decision trees

15.2 Working of ID3 Algorithm

Selection of Root Node:

  1. Compute entropy for each attribute
  2. Calculate Information Gain
  3. Choose attribute with highest gain as root

Recursive Partitioning:

  • After root selection, dataset partitioned into subsets
  • Same process repeated recursively for each subset
  • Each subset treated as new dataset

Termination Conditions:

  • All examples belong to same class
  • No attributes remain
  • Dataset becomes empty

Tree Construction: Top-down approach


15.3 Steps of ID3 Algorithm

StepDescription
Step 1Calculate entropy of training dataset
Step 2Calculate Information Gain for every attribute
Step 3Select attribute with maximum Information Gain
Step 4Create node using this attribute
Step 5Repeat recursively until tree is complete

Algorithm Flow:

Start → Calculate Entropy → Calculate Information Gain → Select Best Attribute
→ Create Node → Partition Data → Recurse → Stop

15.4 Illustrative Example

Training Dataset:

OutlookHumidityPlay Tennis
SunnyHighNo
SunnyNormalYes
RainyHighYes
RainyNormalYes

Information Gain Values:

  • Outlook = 0.45
  • Humidity = 0.18

Result: Outlook becomes root node (highest gain)

Classification Example:

  • New instance: Outlook = Sunny, Humidity = Normal
  • Path: Outlook → Sunny → Humidity → Normal → Play = Yes

15.5 Advantages and Limitations of ID3

AdvantagesLimitations
Simple and easy to understandOverfitting
Fast learningHandles categorical attributes only
Handles multiple attributesSensitive to noise
Generates human-readable rulesBias toward multi-valued attributes

UNIT 16: INTRODUCTION TO ARTIFICIAL NEURAL NETWORKS

16.1 Introduction to Artificial Neural Networks

Definition: ANNs are computational systems inspired by the structure and functioning of the human brain, consisting of interconnected artificial neurons that work together to solve complex problems.

Inspiration from Biological Neurons:

  • Human brain contains ~86 billion neurons connected through trillions of synapses
  • Biological neuron components: Dendrites, Cell Body (Soma), Axon, Synapses
  • Dendrites receive signals; cell body processes; axon transmits; synapses connect

History of Neural Networks:

YearMilestone
1943McCulloch-Pitts Model (first mathematical neuron)
1958Perceptron (Frank Rosenblatt)
1969Minsky & Papert identify limitations
1986Backpropagation (Rumelhart, Hinton, Williams)
2000+Deep Learning Era
PresentModern AI Systems

Characteristics of Neural Networks:

CharacteristicDescription
Learning AbilityLearn from examples and improve through training
GeneralizationCorrectly classify unseen data
AdaptationAdapt to changing environments
Fault ToleranceContinue functioning if some neurons fail
Parallel ProcessingMultiple neurons process simultaneously
Nonlinear ModelingModel complex nonlinear relationships

16.2 Biological versus Artificial Neurons

Biological Neuron Structure:

ComponentFunction
DendritesReceive signals from neighboring neurons
Cell BodyProcesses incoming information
AxonCarries output signals away
SynapsesConnections for communication; strength = synaptic weights

Artificial Neuron Model:

  • Receives inputs
  • Multiplies each input by a weight
  • Sums weighted inputs
  • Applies activation function
  • Produces output

Mathematical Formula:

Net=∑i=1nwixi+bNet = \sum_{i=1}^{n} w_i x_i + b y=f(Net)y = f(Net)

Comparison:

Biological NeuronArtificial Neuron
Dendrites receive signalsInputs receive data
Synapses determine signal strengthWeights determine importance
Cell body processes informationSummation unit processes inputs
Axon transmits outputOutput node generates result
Learns biologicallyLearns mathematically

16.3 Applications of Neural Networks

ApplicationDescriptionExamples
Pattern RecognitionIdentify regularities in dataHandwriting, face, fingerprint recognition
Speech RecognitionConvert spoken language to textVoice assistants, transcription
Image ProcessingAnalyze and interpret imagesObject detection, medical imaging
Medical DiagnosisAnalyze medical data for disease detectionCancer detection, heart disease prediction

UNIT 17: NEURAL NETWORK REPRESENTATION

17.1 Neural Network Architecture

Three Layer Types:

LayerDescriptionFunction
Input LayerFirst layer; receives external dataAccepts raw input; distributes to hidden layers
Hidden LayerBetween input and output; performs computationsExtracts features; learns patterns; nonlinear transformations
Output LayerFinal layer; produces resultProduces final prediction; classification

Architecture Diagram:

Input Layer    Hidden Layer    Output Layer
   x1 ──┐
        ├──→ h1 ──┐
   x2 ──┤        ├──→ y
        ├──→ h2 ──┤
   x3 ──┘        └──→

17.2 Components of Neural Networks

ComponentDescriptionRole
InputsData values supplied to networkRepresent features/attributes
WeightsNumerical values on connectionsDetermine importance of inputs
BiasAdditional parameter added to weighted sumShifts activation function; improves flexibility
Activation FunctionDetermines neuron outputIntroduces nonlinearity

Activation Functions:

FunctionFormulaRangeCharacteristics
Step1 if Net > Threshold, else 0{0,1}Simple, binary output
Sigmoidf(x) = 1/(1+e⁻ˣ)(0,1)Smooth, differentiable
ReLUf(x) = max(0,x)[0,∞)Efficient, avoids vanishing gradient
Tanhtanh(x)(−1,1)Zero-centered

Neuron Computation:

Inputs → Weights Applied → Summation + Bias → Activation Function → Output

17.3 Appropriate Problems for Neural Networks

Problem TypeDescriptionExamples
Nonlinear ClassificationComplex relationships not linearly separableFraud detection, disease diagnosis
Complex Pattern RecognitionIdentifying hidden patternsFace, speech, fingerprint recognition
Prediction ProblemsForecasting future valuesWeather, sales, stock prediction

17.4 Types of Neural Networks

TypeStructureAdvantagesLimitations
Single LayerInput → Output (no hidden layer)Easy to implement; fast trainingCannot solve nonlinear problems (XOR)
Multi-LayerInput → Hidden Layer(s) → OutputHigh learning capability; solves nonlinear problemsRequires more computation; longer training

UNIT 18: PERCEPTRON MODEL

18.1 Introduction to Perceptron

Definition: A perceptron is a computational model that simulates a biological neuron by receiving inputs, processing them through weighted connections, and generating an output based on an activation function.

Mathematical Formula:

Net=∑i=1nwixi+bNet = \sum_{i=1}^{n} w_i x_i + b

Historical Background:

YearContribution
1943McCulloch and Pitts proposed first artificial neuron model
1958Frank Rosenblatt introduced the Perceptron
1969Minsky and Papert identified limitations
1986Backpropagation revived neural network research
PresentDeep learning evolved from perceptron concepts

18.2 Structure of a Perceptron

Components:

ComponentDescription
InputsFeatures/attributes (x₁, x₂, …, xₙ)
WeightsImportance of each input (w₁, w₂, …, wₙ)
Summation FunctionCalculates weighted sum + bias
BiasShifts activation threshold
Activation FunctionDetermines output (Step function)
OutputBinary classification (0 or 1)

Structure Diagram:

x1 ──(w1)──┐
           │
x2 ──(w2)──┼──→ Σ ──→ Activation ──→ Output
           │
x3 ──(w3)──┘
           │
     Bias ─┘

Example Calculation:

  • x₁=1, x₂=0, x₃=1
  • w₁=0.5, w₂=0.3, w₃=0.8
  • Bias = 0.2
  • Net = (1×0.5) + (0×0.3) + (1×0.8) + 0.2 = 1.5
  • If threshold = 1, Output = 1

18.3 Perceptron Learning Rule

Weight Initialization: Small random values (e.g., w₁=0.2, w₂=0.4, w₃=0.1)

Error Calculation:

Error=Target−OutputError = Target - Output

Weight Update Mechanism:

wnew=wold+η(Target−Output)xw_{new} = w_{old} + \eta(Target - Output)x

Where:

  • η = Learning rate
  • x = Input value

Learning Rate Effects:

Learning RateEffect
Very SmallSlow learning
ModerateStable learning
Very LargeUnstable learning

Example Weight Update:

  • Input = 1, Weight = 0.4, Target = 1, Output = 0, η = 0.2
  • w_new = 0.4 + 0.2(1-0)(1) = 0.6

Learning Cycle:

Training Data → Prediction → Calculate Error → Update Weights → Repeat Until Error = 0

18.4 Applications of Perceptrons

ApplicationDescriptionExamples
Binary ClassificationTwo-class problemsPass/Fail, Yes/No, Spam/Not Spam
Pattern RecognitionSimple pattern identificationCharacter recognition, shape classification
Decision SupportAssisting decision-makingLoan approval, risk assessment

18.5 Limitations of Perceptrons

Linear Separability Constraint:

  • Perceptron can only classify linearly separable data
  • A dataset is linearly separable if a straight line can divide the classes

XOR Problem:

Input AInput BOutput
000
011
101
110
  • No single straight line can separate XOR output classes
  • Single-layer perceptron cannot solve XOR problem

Other Limitations:

  • Limited learning capability
  • Binary output restriction
  • No hidden layers
  • Poor performance on complex data

Solution: Multi-Layer Perceptrons (MLPs) and Backpropagation Algorithm


UNIT 19: BACKPROPAGATION ALGORITHM

19.1 Introduction to Backpropagation

Definition: Backpropagation is a supervised learning algorithm used for training Multi-Layer Perceptrons by adjusting connection weights based on errors produced during prediction.

Need for Backpropagation:

  • Perceptron learning rule works only for single-layer networks
  • Real-world problems require multiple hidden layers
  • Hidden layer weights cannot be directly determined
  • Backpropagation calculates each neuron’s contribution to overall error

Multi-Layer Learning:

  • Input Layer
  • One or More Hidden Layers
  • Output Layer
  • All layers learn simultaneously

Concept Flow:

Training Data → Forward Propagation → Predicted Output → Error Calculation
→ Backward Propagation → Weight Adjustment → Improved Prediction

19.2 Forward Propagation

Definition: First phase where input data passes through network layer by layer until output is generated.

Steps:

  1. Input Processing: Present input data to input layer
  2. Weighted Sum Calculation: Net = Σwᵢxᵢ + b
  3. Activation Function Application: Apply sigmoid, ReLU, etc.
  4. Output Generation: Produce final prediction

Forward Propagation Flow:

Input Data → Weighted Sum → Activation Function → Hidden Layer
→ Output Layer → Prediction

19.3 Error Computation

Definition: Evaluating prediction accuracy by computing error between predicted and actual output.

Error Function (Mean Squared Error):

E=12(Target−Output)2E = \frac{1}{2}(Target - Output)^2

Example:

  • Target = 1, Predicted = 0.8
  • E = ½(1−0.8)² = 0.02

Performance Measures:

  • Mean Squared Error (MSE)
  • Accuracy
  • Precision
  • Recall

Error Computation Process:

Target Output → Compare → Predicted Output → Error Calculation

19.4 Backward Propagation

Definition: Heart of backpropagation; error propagates backward through network to adjust weights.

Error Propagation:

  1. Output layer receives error first
  2. Error transmitted backward to hidden layers
  3. Each neuron receives error proportional to its contribution
  4. Network determines which connections need adjustment

Gradient Descent:

  • Optimization technique to minimize error
  • Moves weights in direction opposite to error gradient
  • Continuously moves toward minimum error point

Error Propagation Flow:

Output Error → Output Layer → Hidden Layer → Input Layer

19.5 Weight Update Process

Learning Rate (η):

ValueEffect
0.01Slow but stable
0.05Moderate
0.1Faster but less stable
Very LargeUnstable learning

Weight Adjustment Formula:

wnew=wold−η∂E∂ww_{new} = w_{old} - \eta \frac{\partial E}{\partial w}

Where:

  • w_old = Current weight
  • η = Learning rate
  • ∂E/∂w = Error gradient

Training Cycle:

Initialize Weights → Forward Propagation → Calculate Error
→ Backward Propagation → Update Weights → Repeat Training

19.6 Advantages and Applications

Advantages:

AdvantageDescription
Learns Complex RelationshipsModels highly nonlinear patterns
High Prediction AccuracyAccurate predictions for real-world problems
Adaptive LearningContinuously improves through training
Supports Multi-Layer NetworksEnables deep neural network learning
Automatic Feature LearningLearns hidden patterns automatically

Applications:

DomainExamples
Image RecognitionFace recognition, object detection, medical image analysis
ForecastingWeather, sales, stock market, demand forecasting
Intelligent SystemsAutonomous vehicles, voice assistants, recommendation systems

CHAPTER 3 SUMMARY TABLE

UnitTopicKey Concepts
13Introduction to Decision TreesClassification, components, applications, characteristics
14Decision Tree RepresentationTree structure, entropy, information gain, attribute splitting
15ID3 AlgorithmEntropy calculation, information gain, recursive partitioning
16Introduction to ANNBiological inspiration, history, characteristics, applications
17Neural Network RepresentationArchitecture, components, activation functions, types
18Perceptron ModelStructure, learning rule, applications, XOR limitation
19BackpropagationForward propagation, error computation, backward propagation, weight update

KEY FORMULAS REFERENCE

ConceptFormula
EntropyEntropy(S)=−∑i=1cpilog⁡2(pi)Entropy(S) = -\sum_{i=1}^{c} p_i \log_2(p_i)
Information Gain$Gain(S,A) = Entropy(S) - \sum_{v \in Values(A)} \frac{
Neuron Net InputNet=∑i=1nwixi+bNet = \sum_{i=1}^{n} w_i x_i + b
Perceptron Outputy=f(Net)y = f(Net)
Error (MSE)E=12(Target−Output)2E = \frac{1}{2}(Target - Output)^2
Weight Updatewnew=wold−η∂E∂ww_{new} = w_{old} - \eta \frac{\partial E}{\partial w}
Sigmoidf(x)=11+e−xf(x) = \frac{1}{1+e^{-x}}
ReLUf(x)=max⁡(0,x)f(x) = \max(0,x)
Tanhf(x)=tanh⁡(x)f(x) = \tanh(x)

On this page