(5th sem) AIML Chapter 4: Complete Concepts Guide
Chapter 4: Bayesian Learning and Clustering Techniques -> Generated and Prepared By Thiruselvan (ThiruXD)
20.1 Introduction to Bayesian Learning
Definition: Bayesian Learning is a statistical approach to machine learning that uses probability theory to represent and reason about uncertainty. Unlike traditional learning methods that produce a single hypothesis, Bayesian learning considers multiple hypotheses and assigns probabilities to them based on available evidence.
Foundation: Bayes’ Theorem provides a mathematical framework for updating beliefs when new information becomes available.
Basic Concept Flow:
Prior Knowledge → New Observations → Bayes Theorem → Updated Knowledge → PredictionProbabilistic Learning Framework Components:
| Component | Description |
|---|---|
| Hypothesis Space | The set of all possible hypotheses |
| Prior Probability | Initial belief about a hypothesis before observing data |
| Training Data | Evidence used for learning |
| Posterior Probability | Updated probability after observing data |
| Prediction Mechanism | Selection of the most probable hypothesis |
Framework Flow:
Hypothesis Space → Prior Probabilities → Training Data → Bayesian Updating
→ Posterior Probabilities → Prediction20.2 Characteristics of Bayesian Learning
1. Handling Uncertainty:
Bayesian learning manages uncertainty by assigning probabilities to multiple possible outcomes rather than making definite conclusions.
Example — Medical Diagnosis:
| Disease | Probability |
|---|---|
| Influenza | 0.60 |
| Pneumonia | 0.25 |
| COVID-19 | 0.15 |
As more symptoms become available, these probabilities are updated.
Advantages of Handling Uncertainty:
- Supports decision-making under incomplete information
- Provides confidence measures for predictions
- Reduces impact of noisy data
- Improves robustness in real-world applications
2. Prior and Posterior Knowledge:
| Type | Definition | Example |
|---|---|---|
| Prior Probability | Initial belief before observing evidence | P(Rain Tomorrow) = 0.40 |
| Posterior Probability | Updated belief after observing evidence | P(Rain |
Learning from Experience:
Initial Disease Probability = 40% → New Symptom Added → Updated Probability = 70%3. Continuous Learning: Bayesian systems continuously improve predictions by incorporating new observations incrementally.
20.3 Applications of Bayesian Learning
| Domain | Application | Description |
|---|---|---|
| Medical Diagnosis | Disease prediction | Calculate probability of diseases based on symptoms |
| Spam Filtering | Email classification | Learn probability of spam based on word patterns |
| Decision Support | Business decisions | Evaluate uncertain situations for managers |
| Financial Analysis | Risk assessment | Credit risk, fraud detection, stock prediction |
| Weather Forecasting | Rain prediction | Update probabilities based on atmospheric conditions |
| Robotics | Navigation | Handle uncertainty in sensor data |
| NLP | Text classification | Sentiment analysis, language translation |
Advantages of Bayesian Learning:
- Handles uncertainty effectively
- Incorporates prior knowledge
- Learns incrementally from new data
- Provides probabilistic predictions
- Works well with small datasets
- Supports decision-making under uncertainty
- Reduces overfitting in many applications
Limitations:
- Requires probability estimates
- Computationally expensive for large hypothesis spaces
- Performance depends on quality of prior knowledge
- Assumptions may not always hold in real-world scenarios
UNIT 21: BAYES THEOREM AND CONCEPT LEARNING
21.1 Bayes Theorem
Definition: Bayes Theorem describes the relationship between conditional probabilities and provides a method for calculating the probability of an event based on prior knowledge.
Mathematical Expression:
Components:
| Component | Symbol | Meaning |
|---|---|---|
| Posterior Probability | P(H | D) |
| Likelihood Probability | P(D | H) |
| Prior Probability | P(H) | Initial belief about hypothesis |
| Evidence Probability | P(D) | Overall probability of observing data |
Components Flow:
Bayes Theorem
├── Prior: P(H)
├── Likelihood: P(D|H)
├── Evidence: P(D)
└── Posterior: P(H|D)Probability Concepts:
| Concept | Definition | Example |
|---|---|---|
| Conditional Probability | P(A | B) — likelihood of A given B |
| Joint Probability | P(A∩B) — likelihood of both A and B | P(Rain and Thunderstorm) |
| Marginal Probability | P(A) — likelihood of A regardless of B | P(Rain) |
21.2 Bayes Theorem in Machine Learning
Posterior Probability: Machine learning algorithms use posterior probability to make predictions. The hypothesis with the highest posterior probability is selected.
Bayesian Learning Process:
Training Data → Prior Probabilities → Bayes Theorem → Posterior Probabilities → PredictionLikelihood Function: Measures how well a hypothesis explains the observed data.
Example — Email Classification: Words like “Free”, “Winner”, “Prize” strongly support hypothesis that email is spam.
Importance of Likelihood:
- Measures compatibility between data and hypothesis
- Helps distinguish competing hypotheses
- Forms basis of Bayesian classification
Bayesian Learning Framework Steps:
- Define hypotheses
- Assign prior probabilities
- Collect training data
- Compute likelihood values
- Calculate posterior probabilities
- Select most probable hypothesis
21.3 Bayes Theorem and Concept Learning
Concept Learning Overview: Concept learning is the process of learning a target concept from training examples. Bayesian learning provides a probabilistic framework for this.
Hypothesis Evaluation Example:
| Hypothesis | Posterior Probability |
|---|---|
| H1 | 0.10 |
| H2 | 0.25 |
| H3 | 0.55 |
| H4 | 0.10 |
Result: H3 selected (highest posterior probability)
Learning from Examples:
Training Examples → Probability Update → Hypothesis Evaluation → Concept LearnedAdvantages of Bayesian Concept Learning:
- Handles noisy training data
- Supports uncertain environments
- Considers multiple hypotheses
- Produces probabilistic predictions
- Provides better generalization
21.4 Applications of Bayes Theorem
| Application | Description | Examples |
|---|---|---|
| Classification | Assign instances to classes | Spam detection, medical diagnosis, sentiment analysis |
| Prediction | Forecast future values | Weather, disease, financial forecasting |
| Medical Diagnosis | Identify diseases | Influenza, pneumonia, COVID-19 |
| Spam Filtering | Classify emails | Spam / Not Spam |
| Decision Support | Assist decision-making | Business analytics, risk assessment |
Bayesian Classification Flow:
Input Data → Bayes Theorem → Probability Calculation → Class AssignmentAdvantages of Bayes Theorem:
- Provides mathematical framework for uncertainty
- Combines prior knowledge with new evidence
- Produces probabilistic predictions
- Works effectively with incomplete information
- Supports incremental learning
- Handles noisy datasets
- Forms basis for many ML algorithms
Limitations:
- Requires accurate probability estimates
- Performance depends on prior knowledge
- Computationally expensive for large datasets
- Assumptions may not reflect real-world conditions
UNIT 22: ML, LS ERROR HYPOTHESIS AND MDL PRINCIPLE
22.1 Maximum Likelihood (ML) Hypothesis
Definition: The Maximum Likelihood Hypothesis is the hypothesis that maximizes the probability of observing the training data.
Mathematical Formula:
Where:
- h_ML = Maximum Likelihood Hypothesis
- H = Hypothesis Space
- D = Training Data
- P(D|h) = Probability of observing data D given hypothesis h
Hypothesis Selection Example:
| Hypothesis | Likelihood |
|---|---|
| H1 | 0.30 |
| H2 | 0.75 |
| H3 | 0.45 |
Result: H2 selected (highest likelihood)
ML Estimation Example — Coin Toss:
- Number of Heads = 80
- Number of Tosses = 100
- P(Head) = 80/100 = 0.8
Applications of ML Estimation:
- Parameter estimation
- Classification models
- Regression analysis
- Bayesian learning
- Neural network training
22.2 Least Squared Error (LS) Hypothesis
Definition: The Least Squared Error Hypothesis minimizes the prediction error between actual and predicted values.
Error Function:
Where:
- y_i = Actual Value
- ŷ_i = Predicted Value
- E = Total Squared Error
Error Minimization Process:
Initial Model → Calculate Error → Adjust Parameters → Reduced Error → Better ModelImportance of Least Squared Error:
- Simple to compute
- Penalizes large errors
- Suitable for regression problems
- Widely used in machine learning
Example:
- Actual Output = 50
- Predicted Output = 45
- Error = 50 - 45 = 5
- Squared Error = 25
22.3 ML for Predicting Values
Definition: Maximum Likelihood estimation is used for prediction tasks by estimating model parameters that best explain observed data.
Prediction Model Flow:
Historical Data → ML Estimation → Model Building → Future PredictionParameter Estimation Example — Linear Model:
Values of a and b estimated using training data. Once estimated, model predicts future values.
Applications:
- House price prediction
- Weather forecasting
- Sales forecasting
- Demand prediction
Importance of ML Prediction:
- Improves forecasting accuracy
- Supports intelligent decision-making
- Learns patterns from historical data
- Provides data-driven predictions
22.4 Minimum Description Length (MDL) Principle
Definition: MDL is a model selection technique where the best hypothesis provides the shortest complete description of both the model and the training data.
Concept:
Training Data → Candidate Models → Description Length Evaluation → Select Minimum LengthModel Selection:
MDL selects model that minimizes:
MDL vs Overfitting/Underfitting:
| Condition | Training Accuracy | Testing Accuracy | Problem |
|---|---|---|---|
| Overfitting | High | Low | Model too complex |
| Underfitting | Low | Low | Model too simple |
| Ideal (MDL) | High | High | Balanced complexity |
Advantages of MDL:
- Prevents overfitting
- Encourages simpler models
- Improves generalization
- Supports efficient model selection
- Reduces unnecessary complexity
Applications:
- Decision Tree Learning
- Bayesian Learning
- Data Compression
- Pattern Recognition
- Knowledge Discovery
- Machine Learning Model Selection
UNIT 23: BAYESIAN CLASSIFIERS
23.1 Bayes Optimal Classifier
Definition: The Bayes Optimal Classifier is the theoretically optimal classifier that minimizes the probability of classification error by considering all possible hypotheses.
Mathematical Formula:
Where:
- v_BO = Bayes Optimal prediction
- v_j = Possible class value
- h_i = Hypothesis
- P(h_i|D) = Posterior probability of hypothesis
- P(v_j|h_i) = Probability that hypothesis predicts class v_j
Working Principle:
Training Data → Posterior Probabilities → Combine All Hypotheses → Final ClassExample — Medical Diagnosis:
| Hypothesis | Probability | Prediction |
|---|---|---|
| H1 | 0.40 | Disease |
| H2 | 0.30 | Disease |
| H3 | 0.20 | No Disease |
| H4 | 0.10 | Disease |
Combined Probability:
- Disease = 0.40 + 0.30 + 0.10 = 0.80
- No Disease = 0.20
- Final Prediction: Disease
Advantages:
- Minimum possible classification error
- Considers all hypotheses
- Strong theoretical foundation
- Provides optimal predictions
Limitations:
- Computationally expensive
- Difficult to evaluate large hypothesis spaces
- Requires posterior probabilities for all hypotheses
23.2 Gibbs Algorithm
Definition: The Gibbs Algorithm approximates the Bayes Optimal Classifier by randomly selecting one hypothesis according to its posterior probability.
Algorithm Procedure:
- Calculate posterior probabilities for all hypotheses
- Randomly select one hypothesis according to its probability
- Use selected hypothesis for classification
- Repeat process when required
Example:
| Hypothesis | Posterior Probability |
|---|---|
| H1 | 0.50 |
| H2 | 0.30 |
| H3 | 0.20 |
Algorithm randomly selects one hypothesis. If H1 selected, prediction = output of H1.
Approximation to Bayes Optimal:
- Expected error of Gibbs ≤ 2 × error of Bayes Optimal
- Practical alternative when exact Bayesian classification is infeasible
Advantages:
- Computationally efficient
- Easy implementation
- Suitable for large hypothesis spaces
Limitations:
- Random selection may lead to inconsistent predictions
- Accuracy lower than Bayes Optimal
23.3 Naïve Bayes Classifier
Definition: Naïve Bayes applies Bayes Theorem with the simplifying assumption that all attributes are conditionally independent given the class label.
Conditional Independence Assumption: Attributes are independent given the class.
Example — Email Classification: Attributes: Contains “Free”, Contains “Winner”, Contains “Lottery” Naïve Bayes assumes occurrence of one word does not influence another when class (Spam/Not Spam) is known.
Bayes Theorem for Classification:
Where:
- C = Class
- X = Feature vector
- P(C|X) = Posterior probability
- P(X|C) = Likelihood
- P(C) = Prior probability
Classification Procedure:
- Calculate prior probabilities of classes
- Calculate likelihood probabilities of attributes
- Apply Bayes Theorem
- Compute posterior probabilities
- Select class with maximum probability
Classification Flow:
Input Features → Calculate Priors → Calculate Likelihoods → Apply Bayes Rule
→ Posterior Probabilities → Final ClassExample: Email contains: Free, Winner, Offer
- P(Spam|Email) = 0.95
- P(Not Spam|Email) = 0.05
- Classification: Spam
Advantages:
- Simple and easy to implement
- Fast training and prediction
- Works well with large datasets
- Handles high-dimensional data effectively
- Requires less training data
Limitations:
- Independence assumption may not hold
- Performance decreases when attributes are highly correlated
- Sensitive to zero-frequency problems
23.4 Applications of Bayesian Classifiers
| Application | Description | Examples |
|---|---|---|
| Text Classification | Assign documents to categories | News categorization, topic classification |
| Spam Detection | Identify spam emails | Email filtering |
| Sentiment Analysis | Identify opinions in text | Product reviews, social media monitoring |
| Medical Diagnosis | Disease prediction | Symptom-based diagnosis |
| Fraud Detection | Identify fraudulent activities | Credit card fraud |
| Recommendation Systems | Suggest products | E-commerce, streaming services |
Text Classification Flow:
Document → Bayesian Classifier → Sports / TechnologySpam Detection Flow:
Email → Feature Extraction → Naïve Bayes Model → Spam / Not SpamSentiment Analysis Flow:
Customer Review → Bayesian Classifier → Positive / NegativeUNIT 24: BAYESIAN BELIEF NETWORKS AND EM ALGORITHM
24.1 Bayesian Belief Networks (BBN)
Definition: A Bayesian Belief Network is a graphical representation used to model uncertain knowledge and probabilistic relationships among variables. It combines probability theory and graph theory.
Structure:
- Nodes represent variables
- Directed edges represent probabilistic dependencies
- Each node has a Conditional Probability Table (CPT)
Example:
Rain → Wet Road → Traffic Jam- Rain influences whether road becomes wet
- Wet road increases probability of traffic congestion
Graphical Representation — DAG:
A
/ \
▼ ▼
B C
\ /
▼
D- A influences B and C
- B and C together influence D
- No cycles allowed
Conditional Dependencies:
Cloudy → Rain → Wet GrassP(WetGrass | Rain) = Probability grass is wet given it has rained.
24.2 Components of Bayesian Belief Networks
| Component | Description | Example |
|---|---|---|
| Nodes | Represent variables or events | Disease, Fever, Rain, Traffic |
| Directed Edges | Represent dependencies between variables | Smoking → Lung Cancer |
| Probability Tables (CPT) | Specify probability of node given parent states | P(Wet Road |
Conditional Probability Table Example:
| Rain | P(Wet Road = Yes) |
|---|---|
| Yes | 0.95 |
| No | 0.05 |
Interpretation:
- If it rains, probability of wet road is 95%
- If it does not rain, probability is only 5%
Advantages of BBN:
- Handling uncertainty
- Knowledge representation
- Probabilistic inference
- Decision support
- Learning capability
Applications:
- Medical diagnosis
- Risk analysis
- Fault detection
- Weather prediction
- Expert systems
24.3 Expectation Maximization (EM) Algorithm
Definition: EM is an iterative statistical technique for estimating unknown parameters when data contains missing values or hidden variables.
EM Framework:
Initial Parameters → E-Step → M-Step → Updated Parameters → Repeat Until ConvergenceE-Step (Expectation Step):
- Missing or hidden information is estimated
- Expected values calculated using current parameters
- Probabilities of hidden variables computed
Example: Estimate probability that each customer belongs to a specific cluster.
E-Step Flow:
Current Parameters → Estimate Hidden Data → Expected ValuesM-Step (Maximization Step):
- Parameters updated using expected values from E-step
- Likelihood of observed data maximized
M-Step Flow:
Expected Values → Update Parameters → Maximize LikelihoodEM Iteration Process:
Start → E-Step → M-Step → Convergence? → Yes: Stop / No: RepeatAdvantages:
- Handles missing data
- Robust parameter estimation
- Flexible framework
- Supports unsupervised learning
Limitations:
- May converge to local optima
- Slow convergence for large datasets
- Sensitive to initialization
24.4 Applications of EM Algorithm
| Application | Description | Examples |
|---|---|---|
| Clustering | Group similar data points | Customer segmentation, market analysis |
| Missing Data Analysis | Estimate missing values | Medical databases, survey analysis |
| Pattern Recognition | Identify hidden patterns | Face recognition, speech recognition |
Clustering Flow:
Dataset → EM Algorithm → ClustersMissing Data Analysis Flow:
Incomplete Data → EM Algorithm → Estimated ValuesPattern Recognition Flow:
Raw Data → EM Learning → Pattern DiscoveryUNIT 25: K-MEANS CLUSTERING
25.1 Introduction to Clustering
Definition: Clustering is an unsupervised learning technique that groups similar data objects into clusters such that objects within the same cluster are more similar to each other than to objects in different clusters.
Supervised vs Unsupervised Learning:
| Aspect | Supervised Learning | Unsupervised Learning |
|---|---|---|
| Labels | Known labels | No labels available |
| Goal | Predict output | Discover patterns |
| Example | Classification | Clustering |
Objectives of Clustering:
- Discover hidden patterns
- Organize large datasets
- Identify similarities among data points
- Support decision-making and data analysis
Characteristics of Good Clustering:
- High similarity within clusters
- Low similarity between clusters
- Compact cluster structure
- Meaningful separation of groups
Clustering Concept:
Dataset → Clustering → Cluster1 / Cluster2 / Cluster325.2 K-Means Clustering
Definition: K-Means partitions a dataset into K predefined clusters by finding K cluster centers (centroids) and assigning each data point to the nearest centroid.
Basic Principle:
- “Means” refers to average value of data points within a cluster
- Objective: Minimize distance between data points and assigned cluster centers
K-Means Process:
Dataset → Select K Centroids → Assign Data Points to Clusters
→ Recalculate Centroids → Repeat Until StableCluster Formation Example:
- K = 3
- Cluster 1 → High Spending Customers
- Cluster 2 → Medium Spending Customers
- Cluster 3 → Low Spending Customers
Cluster Formation Flow:
Customer Data → K-Means Algorithm → Cluster A / Cluster B / Cluster C25.3 Steps of K-Means Algorithm
Step 1: Initialization
- Select number of clusters (K)
- Randomly choose K initial centroids
Step 2: Assignment Step
-
Assign each data point to nearest centroid
-
Distance measured using Euclidean Distance:
Step 3: Centroid Update
-
Recalculate centroid of each cluster
-
New centroid = average of all points in cluster:
Step 4: Convergence
- Repeat assignment and update steps
- Stop when centroids no longer change significantly
Convergence Process:
Initialize → Assign Points → Update Centroids → Converged? → Yes: Stop / No: RepeatExample — Student Marks:
| Student | Marks |
|---|---|
| S1 | 35 |
| S2 | 40 |
| S3 | 42 |
| S4 | 75 |
| S5 | 78 |
| S6 | 80 |
K = 2:
- Cluster 1: 35, 40, 42 (low performers)
- Cluster 2: 75, 78, 80 (high performers)
25.4 Advantages and Limitations
Strengths of K-Means:
| Strength | Description |
|---|---|
| Simplicity | Easy to understand and implement |
| Fast Execution | Computationally efficient for large datasets |
| Scalability | Works effectively with large volumes of data |
| Easy Interpretation | Cluster centers provide meaningful summaries |
Weaknesses of K-Means:
| Weakness | Description |
|---|---|
| Choosing K | Selecting correct number of clusters is difficult |
| Sensitive to Initialization | Different initial centroids produce different results |
| Sensitive to Outliers | Extreme values affect cluster centers |
| Assumption of Spherical Clusters | Performs best with compact, well-separated clusters |
Advantages Flow:
K-Means → Simple / Fast / Scalable / InterpretableLimitations Flow:
K-Means → Outliers / Initial Centroids / K Value / Cluster Shape25.5 Applications of K-Means
| Application | Description | Benefits |
|---|---|---|
| Customer Segmentation | Divide customers into groups | Personalized marketing, customer retention |
| Image Segmentation | Divide image into regions | Medical imaging, object recognition |
| Data Analysis | Discover hidden patterns | Market research, social network analysis |
| Document Clustering | Group similar documents | News categorization |
| Recommendation Systems | Suggest products | E-commerce |
| Fraud Detection | Identify anomalies | Financial transactions |
Customer Segmentation Flow:
Customer Data → K-Means → Group1 / Group2 / Group3Image Segmentation Flow:
Image → K-Means → Segmented RegionsUNIT 26: HIERARCHICAL CLUSTERING AND CLUSTER VALIDATION
26.1 Hierarchical Clustering
Definition: Hierarchical Clustering builds a hierarchy of clusters that can be visualized as a tree-like structure called a dendrogram.
Types:
| Type | Approach | Description |
|---|---|---|
| Agglomerative | Bottom-up | Each point starts as own cluster; merge similar clusters |
| Divisive | Top-down | All points in one cluster; repeatedly split |
Agglomerative Clustering:
Step 1: A B C D (each separate)
Step 2: (A,B) C D (merge closest)
Step 3: AB C D
Step 4: ABC D
Step 5: ABCD (all merged)Agglomerative Process:
A B C D
| | | |
+----+ | |
AB | |
+-----+ |
ABC |
+-----+
ABCDDivisive Clustering:
Step 1: ABCD (all together)
Step 2: AB CD (split)
Step 3: A B C D (continue splitting)Divisive Process:
ABCD
/ \
AB CD
/ \ / \
A B C DAdvantages of Agglomerative:
- Simple to understand
- Produces complete cluster hierarchy
- No need to specify number of clusters initially
- Useful for exploratory data analysis
Disadvantages:
- Computationally expensive for large datasets
- Once clusters merged, cannot be separated
- Sensitive to noise and outliers
26.2 Dendrogram Representation
Definition: A dendrogram is a tree-like graphical representation showing how clusters are formed during hierarchical clustering.
Shows:
- Order of cluster formation
- Similarity levels among clusters
- Hierarchical relationships
Dendrogram Structure:
Distance
|
10 | --------
| | |
8 | ---- |
| | |
6 | ---- |
| | |
4 |-- |
|
+-------------------------
A B C DInterpretation:
- Horizontal cut through dendrogram produces clusters
- Example: Cut at distance level 6
- Cluster 1 → A, B
- Cluster 2 → C, D
Dendrogram Interpretation Flow:
Cluster Formation → Dendrogram → Select Cutting Level → Obtain Clusters26.3 Cluster Validity Measures
1. Silhouette Coefficient:
Measures how similar an object is to its own cluster compared to other clusters.
Range: -1 ≤ S ≤ 1
| Value | Meaning |
|---|---|
| Near 1 | Well-clustered |
| Near 0 | Overlapping clusters |
| Near -1 | Incorrect clustering |
Advantages:
- Easy interpretation
- Measures both cohesion and separation
2. Dunn Index:
Evaluates cluster quality by comparing:
- Minimum distance between clusters
- Maximum diameter within clusters
Characteristics:
- Larger values preferred
- Encourages compact clusters
- Promotes well-separated clusters
Dunn Index Flow:
Inter-Cluster Distance (Higher) + Intra-Cluster Distance (Lower) → Higher Dunn Index3. Davies-Bouldin Index:
Evaluates clustering quality by measuring cluster similarity.
| DB Index | Quality |
|---|---|
| Low | Better |
| High | Poor |
Advantages:
- Simple computation
- Widely used in clustering evaluation
26.4 Evaluation of Clustering Results
Cluster Quality Assessment:
A good clustering solution should exhibit:
| Characteristic | Description |
|---|---|
| High Intra-Cluster Similarity | Objects within a cluster should be highly similar |
| Low Inter-Cluster Similarity | Different clusters should be clearly separated |
Cluster Quality Flow:
Good Clustering → High Intra-Similarity / Low Inter-SimilarityComparison of Clustering Techniques:
| Feature | K-Means | Hierarchical |
|---|---|---|
| Number of Clusters | Required (K) | Not required |
| Scalability | High | Moderate |
| Dendrogram Support | No | Yes |
| Interpretability | Moderate | High |
| Computational Cost | Low | High |
Comparison Flow:
Clustering Methods → K-Means (Fast & Simple) / Hierarchical (Detailed Hierarchy)26.5 Applications of Hierarchical Clustering
| Application | Description | Examples |
|---|---|---|
| Bioinformatics | Analyze gene expression data | Gene sequencing, DNA analysis |
| Document Clustering | Group similar documents | Research papers, news articles |
| Market Research | Segment customers | Customer segmentation, product recommendation |
| Image Processing | Group similar images | Medical diagnosis |
| Social Network Analysis | Identify communities | Fraud detection |
Bioinformatics Flow:
Gene Data → Hierarchical Clustering → Gene GroupsDocument Clustering Flow:
Documents → Clustering → Topic GroupsMarket Research Flow:
Customer Data → Hierarchical Clustering → Customer SegmentsCHAPTER 4 SUMMARY TABLE
| Unit | Topic | Key Concepts |
|---|---|---|
| 20 | Introduction to Bayesian Learning | Probability, uncertainty, prior/posterior, applications |
| 21 | Bayes Theorem and Concept Learning | Bayes theorem, likelihood, posterior, classification |
| 22 | ML, LS Error Hypothesis, MDL | Maximum likelihood, least squared error, MDL principle |
| 23 | Bayesian Classifiers | Bayes optimal, Gibbs, Naïve Bayes, applications |
| 24 | Bayesian Belief Networks and EM | BBN, DAG, CPT, EM algorithm, applications |
| 25 | K-Means Clustering | Centroids, assignment, update, convergence, applications |
| 26 | Hierarchical Clustering | Agglomerative, divisive, dendrogram, validation |
KEY FORMULAS REFERENCE
| Concept | Formula |
|---|---|
| Bayes Theorem | $P(H |
| Maximum Likelihood | $h_{ML} = \arg\max_{h \in H} P(D |
| Least Squared Error | |
| MDL Principle | $\text{Minimize } L(h) + L(D |
| Bayes Optimal Classifier | $v_{BO} = \arg\max_{v_j} \sum_{h_i} P(v_j |
| Naïve Bayes | $P(C |
| Euclidean Distance | |
| Centroid | |
| Silhouette Coefficient |
CONCEPT RELATIONSHIPS MAP
BAYESIAN LEARNING
|
┌──────────────────┼──────────────────┐
▼ ▼ ▼
Bayes Theorem Bayesian Classifiers BBN & EM
| | |
┌────┴────┐ ┌────┴────┐ ┌────┴────┐
▼ ▼ ▼ ▼ ▼ ▼
Prior Posterior Naïve Bayes BBN EM
P(H) P(H|D) Bayes Optimal (DAG) (E/M)
| |
┌────┴────┐ ┌────┴────┐
▼ ▼ ▼ ▼
ML MDL Gibbs Applications
Hypothesis Principle Algorithm
|
┌────┴────┐
▼ ▼
LS Error Prediction
Hypothesis Models
CLUSTERING TECHNIQUES
|
┌──────────────────┼──────────────────┐
▼ ▼ ▼
K-Means Hierarchical Cluster
Clustering Clustering Validation
| | |
┌────┴────┐ ┌────┴────┐ ┌────┴────┐
▼ ▼ ▼ ▼ ▼ ▼
Centroids Assign Agglom Divisive Silhouette Dunn
Points erative Coefficient Index
| |
┌────┴────┐ ┌────┴────┐
▼ ▼ ▼ ▼
Update Converge Dendrogram Applications
Centroids(5th sem) AIML Chapter 3: Questions & Answers
Chapter 3: Decision Tree Learning and Artificial Neural Networks -> Generated and Prepared By Thiruselvan (ThiruXD)
(5th sem) AIML Chapter 4: Questions & Answers
Chapter 4: Bayesian Learning and Clustering Techniques -> Generated and Prepared By Thiruselvan (ThiruXD)