BTCE | 5th Sem
Data Science SubjectUnit 3

DS Unit 3: Complete Concept Guide

Unit III: Statistical Foundations and Exploratory Data Analysis -> Generated and Prepared By Thiruselvan (ThiruXD)

1. Introduction to Statistics

1.1 What is Statistics?

Statistics is the science of collecting, organizing, summarizing, analyzing, and interpreting data to make informed decisions. In Data Science, statistics provides the mathematical foundation for understanding data patterns, quantifying uncertainty, and validating hypotheses.

1.2 Branches of Statistics

BranchPurposeExample
Descriptive StatisticsSummarizes and describes the main features of a datasetMean sales per month, standard deviation of test scores
Inferential StatisticsDraws conclusions about a population from a sampleEstimating voter preference from a survey of 1,000 people

1.3 Key Terminology

  • Population: The entire group of interest (e.g., all customers of a company).
  • Sample: A subset of the population used for analysis.
  • Parameter: A numerical summary of a population (e.g., population mean μ).
  • Statistic: A numerical summary of a sample (e.g., sample mean x̄).
  • Variable: A characteristic being measured (e.g., age, income).
  • Observation: A single data point or record.

1.4 Types of Data

TypeDescriptionExamples
NominalCategories without orderGender, city, color
OrdinalCategories with orderEducation level, satisfaction rating
IntervalOrdered, equal intervals, no true zeroTemperature (°C), IQ scores
RatioOrdered, equal intervals, true zeroHeight, weight, income

2. Descriptive Statistics

Descriptive statistics summarize data using measures of central tendency and measures of dispersion.

2.1 Measures of Central Tendency

These measures identify the “center” or “typical value” of a dataset.

(a) Mean (Arithmetic Average)

  • Formula (Population): μ = (Σxᵢ) / N
  • Formula (Sample): x̄ = (Σxᵢ) / n
  • Use: Best for symmetric, interval/ratio data without outliers.
  • Limitation: Highly sensitive to outliers.

Example: Data: 10, 20, 30, 40, 50 Mean = (10+20+30+40+50)/5 = 30

(b) Median

  • The middle value when data is sorted.
  • If n is odd: middle value.
  • If n is even: average of the two middle values.
  • Use: Best for skewed data or data with outliers.

Example: Data: 10, 20, 30, 40, 1000 Median = 30 (robust to the outlier 1000)

(c) Mode

  • The most frequently occurring value.
  • A dataset can have no mode, one mode (unimodal), two modes (bimodal), or multiple modes.
  • Use: Best for categorical/nominal data.

Example: Data: 2, 3, 3, 5, 7, 7, 7, 9 Mode = 7

Comparison of Central Tendency Measures

MeasureSensitive to Outliers?Best For
MeanYesSymmetric numerical data
MedianNoSkewed data, outliers
ModeNoCategorical data

Skewness and Central Tendency

  • Symmetric distribution: Mean = Median = Mode
  • Right-skewed (positive): Mode < Median < Mean
  • Left-skewed (negative): Mean < Median < Mode

2.2 Measures of Dispersion

Dispersion measures how spread out the data is around the center.

(a) Range

  • Formula: Range = Maximum − Minimum
  • Pros: Simple to compute.
  • Cons: Only uses two values; sensitive to outliers.

Example: Data: 5, 10, 15, 20 → Range = 20 − 5 = 15

(b) Variance

  • Measures the average squared deviation from the mean.
  • Population Variance: σ² = Σ(xᵢ − μ)² / N
  • Sample Variance: s² = Σ(xᵢ − x̄)² / (n − 1)
  • Note: Dividing by (n−1) gives an unbiased estimator (Bessel’s correction).

(c) Standard Deviation (SD)

  • The square root of variance; expressed in the same units as the data.
  • Population SD: σ = √σ²
  • Sample SD: s = √s²
  • Interpretation: Larger SD → more spread; smaller SD → data clustered near mean.

Example: Data: 2, 4, 4, 4, 5, 5, 7, 9 Mean = 5 Variance (population) = [(−3)²+(−1)²+(−1)²+(−1)²+0²+0²+2²+4²]/8 = (9+1+1+1+0+0+4+16)/8 = 32/8 = 4 SD = √4 = 2

(d) Interquartile Range (IQR)

  • Formula: IQR = Q3 − Q1
  • Q1 = 25th percentile, Q3 = 75th percentile.
  • Use: Robust measure of spread; used in box plots and outlier detection.
  • Outlier rule: Value < Q1 − 1.5×IQR or > Q3 + 1.5×IQR → potential outlier.

(e) Coefficient of Variation (CV)

  • Formula: CV = (SD / Mean) × 100%
  • Use: Compares variability between datasets with different units or means.

Summary Table of Dispersion Measures

MeasureFormulaUnitsOutlier Sensitivity
RangeMax − MinOriginalHigh
VarianceΣ(x−μ)²/NSquaredHigh
Standard Deviation√VarianceOriginalHigh
IQRQ3 − Q1OriginalLow
CVSD/Mean × 100%Moderate

2.3 Five-Number Summary

A compact description of a dataset:

  1. Minimum
  2. Q1 (25th percentile)
  3. Median (Q2)
  4. Q3 (75th percentile)
  5. Maximum

Use: Foundation of the box plot.


3. Basics of Probability

3.1 Core Concepts

  • Experiment: A process with uncertain outcomes (e.g., tossing a coin).
  • Sample Space (S): All possible outcomes (e.g., {H, T}).
  • Event (E): A subset of the sample space (e.g., getting Heads).
  • Probability: P(E) = Number of favorable outcomes / Total outcomes
    • 0 ≤ P(E) ≤ 1
    • P(S) = 1, P(∅) = 0

3.2 Rules of Probability

RuleFormula
ComplementP(A’) = 1 − P(A)
Addition (mutually exclusive)P(A∪B) = P(A) + P(B)
Addition (general)P(A∪B) = P(A) + P(B) − P(A∩B)
Multiplication (independent)P(A∩B) = P(A) × P(B)
Conditional ProbabilityP(A
Bayes’ TheoremP(A

3.3 Random Variables

  • Discrete: Countable outcomes (e.g., number of heads in 10 tosses).
  • Continuous: Infinite outcomes in a range (e.g., height, time).

3.4 Probability Distributions

DistributionTypeUse Case
BernoulliDiscreteSingle trial (success/failure)
BinomialDiscreteNumber of successes in n trials
PoissonDiscreteCount of events in a fixed interval
Normal (Gaussian)ContinuousNatural phenomena, measurement errors
UniformContinuousEqual probability across range
ExponentialContinuousTime between events

3.5 Normal Distribution (Gaussian)

  • Bell-shaped, symmetric around mean μ.
  • Empirical Rule (68–95–99.7):
    • ~68% of data within μ ± 1σ
    • ~95% within μ ± 2σ
    • ~99.7% within μ ± 3σ
  • Standard Normal (Z): μ = 0, σ = 1; Z = (x − μ) / σ
  • Central Limit Theorem (CLT): The sampling distribution of the mean approaches normality as sample size increases (n ≥ 30), regardless of population distribution.

4. Correlation and Covariance

4.1 Covariance

  • Measures the direction of the linear relationship between two variables.
  • Formula (Sample): Cov(X, Y) = Σ[(xᵢ − x̄)(yᵢ − ȳ)] / (n − 1)
  • Interpretation:
    • Cov > 0 → X and Y tend to move together.
    • Cov < 0 → X and Y tend to move in opposite directions.
    • Cov ≈ 0 → No linear relationship.
  • Limitation: Magnitude depends on units → hard to interpret.

4.2 Correlation (Pearson’s r)

  • Standardized measure of linear relationship; unit-free.
  • Formula: r = Cov(X, Y) / (σₓ × σᵧ)
  • Range: −1 ≤ r ≤ +1
r ValueInterpretation
+1Perfect positive linear relationship
+0.7 to +0.9Strong positive
+0.3 to +0.7Moderate positive
0No linear relationship
−0.3 to −0.7Moderate negative
−0.7 to −0.9Strong negative
−1Perfect negative linear relationship

4.3 Spearman’s Rank Correlation

  • Measures monotonic (not necessarily linear) relationships.
  • Based on ranks rather than raw values.
  • Use: Non-linear relationships, ordinal data, or when outliers exist.

4.4 Important Notes

  • Correlation ≠ Causation.
  • Correlation only captures linear relationships (for Pearson).
  • Always visualize data (scatter plot) alongside correlation.

5. Exploratory Data Analysis (EDA)

5.1 What is EDA?

EDA is the process of exploring data visually and statistically to:

  • Understand data structure, patterns, and anomalies.
  • Form hypotheses for further analysis.
  • Guide feature selection and model building.

Coined by: John Tukey (1977)

5.2 Goals of EDA

  1. Maximize insight into a dataset.
  2. Uncover underlying structure.
  3. Extract important variables.
  4. Detect outliers and anomalies.
  5. Test underlying assumptions.
  6. Develop parsimonious models.

5.3 Steps in EDA

Step 1: Data Understanding

  • Check dimensions (rows × columns).
  • Inspect data types (numeric, categorical, datetime).
  • View first/last rows, summary statistics.

Step 2: Data Cleaning

  • Handle missing values (imputation or removal).
  • Fix inconsistent formats.
  • Remove duplicates.

Step 3: Univariate Analysis

  • Analyze one variable at a time.
  • Numeric: Histogram, box plot, mean, median, SD.
  • Categorical: Bar chart, frequency table, mode.

Step 4: Bivariate Analysis

  • Analyze relationships between two variables.
  • Numeric vs Numeric: Scatter plot, correlation.
  • Categorical vs Numeric: Box plot, group means.
  • Categorical vs Categorical: Contingency table, stacked bar chart.

Step 5: Multivariate Analysis

  • Analyze three or more variables simultaneously.
  • Techniques: Pair plots, correlation heatmaps, 3D scatter plots, PCA.

Step 6: Pattern Identification and Trend Analysis

  • Identify seasonality, cycles, trends in time-series data.
  • Detect clusters, correlations, and associations.
  • Use moving averages, decomposition, and aggregation.

5.4 Common EDA Techniques

TechniquePurpose
Summary statisticsQuick overview of distribution
HistogramShape of numeric distribution
Box plotSpread, median, outliers
Scatter plotRelationship between two numerics
Correlation matrixPairwise linear relationships
HeatmapVisualize correlation or missing data
Pair plotAll pairwise relationships
Bar chartCompare categories
Line chartTrends over time
Count plotFrequency of categories

5.5 Outlier Detection

Methods:

  1. IQR Method: Outliers < Q1 − 1.5×IQR or > Q3 + 1.5×IQR
  2. Z-score Method: |Z| > 3 → outlier (Z = (x − μ)/σ)
  3. Visual: Box plots, scatter plots
  4. Isolation Forest / DBSCAN: Machine learning-based methods

Handling Outliers:

  • Remove (if erroneous)
  • Cap/Winsorize (replace with percentile values)
  • Transform (log, square root)
  • Keep (if genuine and important)

5.6 Missing Data Handling

MethodDescription
Listwise deletionRemove rows with missing values
Mean/Median/Mode imputationReplace with central tendency
Forward/Backward fillUse adjacent values (time series)
KNN imputationUse similar records
Regression imputationPredict missing values
Multiple imputationGenerate several plausible values

5.7 Pattern Identification and Trend Analysis

  • Trend: Long-term increase or decrease.
  • Seasonality: Regular periodic fluctuations.
  • Cyclical: Irregular long-term oscillations.
  • Noise: Random variation.

Tools:

  • Moving averages
  • Exponential smoothing
  • Time-series decomposition (trend + seasonal + residual)
  • Autocorrelation (ACF/PACF)

6. EDA in Practice — Python Example

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns

# Load data
df = pd.read_csv('data.csv')

# 1. Basic info
print(df.shape)
print(df.info())
print(df.describe())

# 2. Missing values
print(df.isnull().sum())

# 3. Univariate - histogram
df['age'].hist(bins=20)
plt.title('Age Distribution')
plt.show()

# 4. Box plot
sns.boxplot(x=df['income'])
plt.show()

# 5. Bivariate - scatter
sns.scatterplot(x='age', y='income', data=df)
plt.show()

# 6. Correlation matrix
corr = df.corr(numeric_only=True)
sns.heatmap(corr, annot=True, cmap='coolwarm')
plt.show()

# 7. Pair plot
sns.pairplot(df[['age', 'income', 'score']])
plt.show()

# 8. Outlier detection via IQR
Q1 = df['income'].quantile(0.25)
Q3 = df['income'].quantile(0.75)
IQR = Q3 - Q1
outliers = df[(df['income'] < Q1 - 1.5*IQR) | (df['income'] > Q3 + 1.5*IQR)]
print(f"Outliers:{len(outliers)}")

7. Key Formulas — Quick Reference

ConceptFormula
Meanx̄ = Σxᵢ / n
MedianMiddle value of sorted data
ModeMost frequent value
Variance (sample)s² = Σ(xᵢ − x̄)² / (n − 1)
Standard Deviations = √s²
IQRQ3 − Q1
CovarianceCov(X,Y) = Σ[(xᵢ−x̄)(yᵢ−ȳ)] / (n−1)
Correlationr = Cov(X,Y) / (σₓσᵧ)
Z-scoreZ = (x − μ) / σ
CV(s / x̄) × 100%
ProbabilityP(E) = favorable / total
BayesP(A

8. Summary — Unit III at a Glance

STATISTICAL FOUNDATIONS & EDA
│
├── Introduction to Statistics
│   ├── Descriptive vs Inferential
│   ├── Population vs Sample
│   └── Types of Data (Nominal, Ordinal, Interval, Ratio)
│
├── Descriptive Statistics
│   ├── Central Tendency: Mean, Median, Mode
│   └── Dispersion: Range, Variance, SD, IQR, CV
│
├── Probability Basics
│   ├── Rules & Conditional Probability
│   ├── Bayes' Theorem
│   └── Distributions (Normal, Binomial, Poisson...)
│
├── Correlation & Covariance
│   ├── Covariance (direction)
│   ├── Pearson's r (strength + direction)
│   └── Spearman's rank (monotonic)
│
└── Exploratory Data Analysis
    ├── Univariate → Bivariate → Multivariate
    ├── Outlier detection (IQR, Z-score)
    ├── Missing data handling
    └── Pattern & trend analysis

9. Exam-Focused Points

  1. Difference between mean, median, mode — when to use each.
  2. Variance vs Standard Deviation — SD is in original units.
  3. Covariance vs Correlation — correlation is standardized.
  4. Correlation ≠ Causation — classic exam question.
  5. IQR method for outliers — memorize the 1.5×IQR rule.
  6. Empirical rule (68-95-99.7) — for normal distributions.
  7. Central Limit Theorem — n ≥ 30 rule of thumb.
  8. EDA steps — univariate → bivariate → multivariate.
  9. Skewness and mean/median relationship.
  10. When to use Spearman vs Pearson correlation.

On this page