DS Unit 3: Complete Concept Guide
Unit III: Statistical Foundations and Exploratory Data Analysis -> Generated and Prepared By Thiruselvan (ThiruXD)
1. Introduction to Statistics
1.1 What is Statistics?
Statistics is the science of collecting, organizing, summarizing, analyzing, and interpreting data to make informed decisions. In Data Science, statistics provides the mathematical foundation for understanding data patterns, quantifying uncertainty, and validating hypotheses.
1.2 Branches of Statistics
| Branch | Purpose | Example |
|---|---|---|
| Descriptive Statistics | Summarizes and describes the main features of a dataset | Mean sales per month, standard deviation of test scores |
| Inferential Statistics | Draws conclusions about a population from a sample | Estimating voter preference from a survey of 1,000 people |
1.3 Key Terminology
- Population: The entire group of interest (e.g., all customers of a company).
- Sample: A subset of the population used for analysis.
- Parameter: A numerical summary of a population (e.g., population mean μ).
- Statistic: A numerical summary of a sample (e.g., sample mean x̄).
- Variable: A characteristic being measured (e.g., age, income).
- Observation: A single data point or record.
1.4 Types of Data
| Type | Description | Examples |
|---|---|---|
| Nominal | Categories without order | Gender, city, color |
| Ordinal | Categories with order | Education level, satisfaction rating |
| Interval | Ordered, equal intervals, no true zero | Temperature (°C), IQ scores |
| Ratio | Ordered, equal intervals, true zero | Height, weight, income |
2. Descriptive Statistics
Descriptive statistics summarize data using measures of central tendency and measures of dispersion.
2.1 Measures of Central Tendency
These measures identify the “center” or “typical value” of a dataset.
(a) Mean (Arithmetic Average)
- Formula (Population): μ = (Σxᵢ) / N
- Formula (Sample): x̄ = (Σxᵢ) / n
- Use: Best for symmetric, interval/ratio data without outliers.
- Limitation: Highly sensitive to outliers.
Example: Data: 10, 20, 30, 40, 50 Mean = (10+20+30+40+50)/5 = 30
(b) Median
- The middle value when data is sorted.
- If n is odd: middle value.
- If n is even: average of the two middle values.
- Use: Best for skewed data or data with outliers.
Example: Data: 10, 20, 30, 40, 1000 Median = 30 (robust to the outlier 1000)
(c) Mode
- The most frequently occurring value.
- A dataset can have no mode, one mode (unimodal), two modes (bimodal), or multiple modes.
- Use: Best for categorical/nominal data.
Example: Data: 2, 3, 3, 5, 7, 7, 7, 9 Mode = 7
Comparison of Central Tendency Measures
| Measure | Sensitive to Outliers? | Best For |
|---|---|---|
| Mean | Yes | Symmetric numerical data |
| Median | No | Skewed data, outliers |
| Mode | No | Categorical data |
Skewness and Central Tendency
- Symmetric distribution: Mean = Median = Mode
- Right-skewed (positive): Mode < Median < Mean
- Left-skewed (negative): Mean < Median < Mode
2.2 Measures of Dispersion
Dispersion measures how spread out the data is around the center.
(a) Range
- Formula: Range = Maximum − Minimum
- Pros: Simple to compute.
- Cons: Only uses two values; sensitive to outliers.
Example: Data: 5, 10, 15, 20 → Range = 20 − 5 = 15
(b) Variance
- Measures the average squared deviation from the mean.
- Population Variance: σ² = Σ(xᵢ − μ)² / N
- Sample Variance: s² = Σ(xᵢ − x̄)² / (n − 1)
- Note: Dividing by (n−1) gives an unbiased estimator (Bessel’s correction).
(c) Standard Deviation (SD)
- The square root of variance; expressed in the same units as the data.
- Population SD: σ = √σ²
- Sample SD: s = √s²
- Interpretation: Larger SD → more spread; smaller SD → data clustered near mean.
Example: Data: 2, 4, 4, 4, 5, 5, 7, 9 Mean = 5 Variance (population) = [(−3)²+(−1)²+(−1)²+(−1)²+0²+0²+2²+4²]/8 = (9+1+1+1+0+0+4+16)/8 = 32/8 = 4 SD = √4 = 2
(d) Interquartile Range (IQR)
- Formula: IQR = Q3 − Q1
- Q1 = 25th percentile, Q3 = 75th percentile.
- Use: Robust measure of spread; used in box plots and outlier detection.
- Outlier rule: Value < Q1 − 1.5×IQR or > Q3 + 1.5×IQR → potential outlier.
(e) Coefficient of Variation (CV)
- Formula: CV = (SD / Mean) × 100%
- Use: Compares variability between datasets with different units or means.
Summary Table of Dispersion Measures
| Measure | Formula | Units | Outlier Sensitivity |
|---|---|---|---|
| Range | Max − Min | Original | High |
| Variance | Σ(x−μ)²/N | Squared | High |
| Standard Deviation | √Variance | Original | High |
| IQR | Q3 − Q1 | Original | Low |
| CV | SD/Mean × 100 | % | Moderate |
2.3 Five-Number Summary
A compact description of a dataset:
- Minimum
- Q1 (25th percentile)
- Median (Q2)
- Q3 (75th percentile)
- Maximum
Use: Foundation of the box plot.
3. Basics of Probability
3.1 Core Concepts
- Experiment: A process with uncertain outcomes (e.g., tossing a coin).
- Sample Space (S): All possible outcomes (e.g., {H, T}).
- Event (E): A subset of the sample space (e.g., getting Heads).
- Probability: P(E) = Number of favorable outcomes / Total outcomes
- 0 ≤ P(E) ≤ 1
- P(S) = 1, P(∅) = 0
3.2 Rules of Probability
| Rule | Formula |
|---|---|
| Complement | P(A’) = 1 − P(A) |
| Addition (mutually exclusive) | P(A∪B) = P(A) + P(B) |
| Addition (general) | P(A∪B) = P(A) + P(B) − P(A∩B) |
| Multiplication (independent) | P(A∩B) = P(A) × P(B) |
| Conditional Probability | P(A |
| Bayes’ Theorem | P(A |
3.3 Random Variables
- Discrete: Countable outcomes (e.g., number of heads in 10 tosses).
- Continuous: Infinite outcomes in a range (e.g., height, time).
3.4 Probability Distributions
| Distribution | Type | Use Case |
|---|---|---|
| Bernoulli | Discrete | Single trial (success/failure) |
| Binomial | Discrete | Number of successes in n trials |
| Poisson | Discrete | Count of events in a fixed interval |
| Normal (Gaussian) | Continuous | Natural phenomena, measurement errors |
| Uniform | Continuous | Equal probability across range |
| Exponential | Continuous | Time between events |
3.5 Normal Distribution (Gaussian)
- Bell-shaped, symmetric around mean μ.
- Empirical Rule (68–95–99.7):
- ~68% of data within μ ± 1σ
- ~95% within μ ± 2σ
- ~99.7% within μ ± 3σ
- Standard Normal (Z): μ = 0, σ = 1; Z = (x − μ) / σ
- Central Limit Theorem (CLT): The sampling distribution of the mean approaches normality as sample size increases (n ≥ 30), regardless of population distribution.
4. Correlation and Covariance
4.1 Covariance
- Measures the direction of the linear relationship between two variables.
- Formula (Sample): Cov(X, Y) = Σ[(xᵢ − x̄)(yᵢ − ȳ)] / (n − 1)
- Interpretation:
- Cov > 0 → X and Y tend to move together.
- Cov < 0 → X and Y tend to move in opposite directions.
- Cov ≈ 0 → No linear relationship.
- Limitation: Magnitude depends on units → hard to interpret.
4.2 Correlation (Pearson’s r)
- Standardized measure of linear relationship; unit-free.
- Formula: r = Cov(X, Y) / (σₓ × σᵧ)
- Range: −1 ≤ r ≤ +1
| r Value | Interpretation |
|---|---|
| +1 | Perfect positive linear relationship |
| +0.7 to +0.9 | Strong positive |
| +0.3 to +0.7 | Moderate positive |
| 0 | No linear relationship |
| −0.3 to −0.7 | Moderate negative |
| −0.7 to −0.9 | Strong negative |
| −1 | Perfect negative linear relationship |
4.3 Spearman’s Rank Correlation
- Measures monotonic (not necessarily linear) relationships.
- Based on ranks rather than raw values.
- Use: Non-linear relationships, ordinal data, or when outliers exist.
4.4 Important Notes
- Correlation ≠ Causation.
- Correlation only captures linear relationships (for Pearson).
- Always visualize data (scatter plot) alongside correlation.
5. Exploratory Data Analysis (EDA)
5.1 What is EDA?
EDA is the process of exploring data visually and statistically to:
- Understand data structure, patterns, and anomalies.
- Form hypotheses for further analysis.
- Guide feature selection and model building.
Coined by: John Tukey (1977)
5.2 Goals of EDA
- Maximize insight into a dataset.
- Uncover underlying structure.
- Extract important variables.
- Detect outliers and anomalies.
- Test underlying assumptions.
- Develop parsimonious models.
5.3 Steps in EDA
Step 1: Data Understanding
- Check dimensions (rows × columns).
- Inspect data types (numeric, categorical, datetime).
- View first/last rows, summary statistics.
Step 2: Data Cleaning
- Handle missing values (imputation or removal).
- Fix inconsistent formats.
- Remove duplicates.
Step 3: Univariate Analysis
- Analyze one variable at a time.
- Numeric: Histogram, box plot, mean, median, SD.
- Categorical: Bar chart, frequency table, mode.
Step 4: Bivariate Analysis
- Analyze relationships between two variables.
- Numeric vs Numeric: Scatter plot, correlation.
- Categorical vs Numeric: Box plot, group means.
- Categorical vs Categorical: Contingency table, stacked bar chart.
Step 5: Multivariate Analysis
- Analyze three or more variables simultaneously.
- Techniques: Pair plots, correlation heatmaps, 3D scatter plots, PCA.
Step 6: Pattern Identification and Trend Analysis
- Identify seasonality, cycles, trends in time-series data.
- Detect clusters, correlations, and associations.
- Use moving averages, decomposition, and aggregation.
5.4 Common EDA Techniques
| Technique | Purpose |
|---|---|
| Summary statistics | Quick overview of distribution |
| Histogram | Shape of numeric distribution |
| Box plot | Spread, median, outliers |
| Scatter plot | Relationship between two numerics |
| Correlation matrix | Pairwise linear relationships |
| Heatmap | Visualize correlation or missing data |
| Pair plot | All pairwise relationships |
| Bar chart | Compare categories |
| Line chart | Trends over time |
| Count plot | Frequency of categories |
5.5 Outlier Detection
Methods:
- IQR Method: Outliers < Q1 − 1.5×IQR or > Q3 + 1.5×IQR
- Z-score Method: |Z| > 3 → outlier (Z = (x − μ)/σ)
- Visual: Box plots, scatter plots
- Isolation Forest / DBSCAN: Machine learning-based methods
Handling Outliers:
- Remove (if erroneous)
- Cap/Winsorize (replace with percentile values)
- Transform (log, square root)
- Keep (if genuine and important)
5.6 Missing Data Handling
| Method | Description |
|---|---|
| Listwise deletion | Remove rows with missing values |
| Mean/Median/Mode imputation | Replace with central tendency |
| Forward/Backward fill | Use adjacent values (time series) |
| KNN imputation | Use similar records |
| Regression imputation | Predict missing values |
| Multiple imputation | Generate several plausible values |
5.7 Pattern Identification and Trend Analysis
- Trend: Long-term increase or decrease.
- Seasonality: Regular periodic fluctuations.
- Cyclical: Irregular long-term oscillations.
- Noise: Random variation.
Tools:
- Moving averages
- Exponential smoothing
- Time-series decomposition (trend + seasonal + residual)
- Autocorrelation (ACF/PACF)
6. EDA in Practice — Python Example
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
# Load data
df = pd.read_csv('data.csv')
# 1. Basic info
print(df.shape)
print(df.info())
print(df.describe())
# 2. Missing values
print(df.isnull().sum())
# 3. Univariate - histogram
df['age'].hist(bins=20)
plt.title('Age Distribution')
plt.show()
# 4. Box plot
sns.boxplot(x=df['income'])
plt.show()
# 5. Bivariate - scatter
sns.scatterplot(x='age', y='income', data=df)
plt.show()
# 6. Correlation matrix
corr = df.corr(numeric_only=True)
sns.heatmap(corr, annot=True, cmap='coolwarm')
plt.show()
# 7. Pair plot
sns.pairplot(df[['age', 'income', 'score']])
plt.show()
# 8. Outlier detection via IQR
Q1 = df['income'].quantile(0.25)
Q3 = df['income'].quantile(0.75)
IQR = Q3 - Q1
outliers = df[(df['income'] < Q1 - 1.5*IQR) | (df['income'] > Q3 + 1.5*IQR)]
print(f"Outliers:{len(outliers)}")7. Key Formulas — Quick Reference
| Concept | Formula |
|---|---|
| Mean | x̄ = Σxᵢ / n |
| Median | Middle value of sorted data |
| Mode | Most frequent value |
| Variance (sample) | s² = Σ(xᵢ − x̄)² / (n − 1) |
| Standard Deviation | s = √s² |
| IQR | Q3 − Q1 |
| Covariance | Cov(X,Y) = Σ[(xᵢ−x̄)(yᵢ−ȳ)] / (n−1) |
| Correlation | r = Cov(X,Y) / (σₓσᵧ) |
| Z-score | Z = (x − μ) / σ |
| CV | (s / x̄) × 100% |
| Probability | P(E) = favorable / total |
| Bayes | P(A |
8. Summary — Unit III at a Glance
STATISTICAL FOUNDATIONS & EDA
│
├── Introduction to Statistics
│ ├── Descriptive vs Inferential
│ ├── Population vs Sample
│ └── Types of Data (Nominal, Ordinal, Interval, Ratio)
│
├── Descriptive Statistics
│ ├── Central Tendency: Mean, Median, Mode
│ └── Dispersion: Range, Variance, SD, IQR, CV
│
├── Probability Basics
│ ├── Rules & Conditional Probability
│ ├── Bayes' Theorem
│ └── Distributions (Normal, Binomial, Poisson...)
│
├── Correlation & Covariance
│ ├── Covariance (direction)
│ ├── Pearson's r (strength + direction)
│ └── Spearman's rank (monotonic)
│
└── Exploratory Data Analysis
├── Univariate → Bivariate → Multivariate
├── Outlier detection (IQR, Z-score)
├── Missing data handling
└── Pattern & trend analysis9. Exam-Focused Points
- Difference between mean, median, mode — when to use each.
- Variance vs Standard Deviation — SD is in original units.
- Covariance vs Correlation — correlation is standardized.
- Correlation ≠ Causation — classic exam question.
- IQR method for outliers — memorize the 1.5×IQR rule.
- Empirical rule (68-95-99.7) — for normal distributions.
- Central Limit Theorem — n ≥ 30 rule of thumb.
- EDA steps — univariate → bivariate → multivariate.
- Skewness and mean/median relationship.
- When to use Spearman vs Pearson correlation.