DS Unit 3: Question with Answers
Unit III: Statistical Foundations and Exploratory Data Analysis -> Generated and Prepared By Thiruselvan (ThiruXD)
SECTION A: MULTIPLE CHOICE QUESTIONS (50 MCQs)
Descriptive Statistics & Central Tendency
Q1. Which measure of central tendency is most affected by outliers?
- Median
- Mode
- Mean
- Range
Answer: C) MeanExplanation: The mean uses every value in its calculation, so extreme values pull it significantly.
Q2. For a perfectly symmetric distribution:
- Mean > Median > Mode
- Mean = Median = Mode
- Mean < Median < Mode
- Mode > Mean > Median
Answer: B) Mean = Median = ModeExplanation: In a symmetric distribution, all three measures coincide at the center.
Q3. The mode of the dataset {2, 3, 3, 5, 7, 7, 7, 9} is:
- 3
- 5
- 7
- 9
Answer: C) 7Explanation: 7 appears three times, more than any other value.
Q4. For a right-skewed (positively skewed) distribution:
- Mean < Median < Mode
- Mode < Median < Mean
- Mean = Median = Mode
- Median < Mode < Mean
Answer: B) Mode < Median < MeanExplanation: The tail pulls the mean to the right, making it the largest.
Q5. The median of {10, 20, 30, 40, 50, 60} is:
- 30
- 35
- 40
- 45
Answer: B) 35Explanation: For even n, median = average of two middle values = (30+40)/2 = 35.
Q6. Which measure is best for categorical (nominal) data?
- Mean
- Median
- Mode
- Standard Deviation
Answer: C) ModeExplanation: Mode works on categories; mean/median require numerical ordering.
Q7. The arithmetic mean of the first 10 natural numbers is:
- 5
- 5.5
- 6
- 10
Answer: B) 5.5Explanation: Sum = 55, n = 10, Mean = 55/10 = 5.5.
Q8. If every value in a dataset is multiplied by 3, the mean:
- Stays the same
- Increases by 3
- Multiplies by 3
- Divides by 3
Answer: C) Multiplies by 3Explanation: Mean is a linear operator: Mean(3x) = 3 × Mean(x).
Measures of Dispersion
Q9. Variance is expressed in:
- Original units
- Squared units
- Percentage
- Logarithmic units
Answer: B) Squared unitsExplanation: Variance squares deviations, so units are squared. SD restores original units.
Q10. The standard deviation of {2, 4, 4, 4, 5, 5, 7, 9} (mean = 5) is:
- 1
- 2
- 3
- 4
Answer: B) 2Explanation: Variance = 32/8 = 4, SD = √4 = 2.
Q11. Which measure of dispersion is most robust to outliers?
- Range
- Variance
- Standard Deviation
- IQR
Answer: D) IQRExplanation: IQR uses quartiles and ignores extreme values.
Q12. If a dataset has variance 25, its standard deviation is:
- 5
- 25
- 625
- 12.5
Answer: A) 5Explanation: SD = √Variance = √25 = 5.
Q13. The IQR of a dataset with Q1 = 20 and Q3 = 50 is:
- 20
- 30
- 35
- 70
Answer: B) 30Explanation: IQR = Q3 − Q1 = 50 − 20 = 30.
Q14. Using the 1.5×IQR rule, an outlier is any value:
- Below Q1 or above Q3
- Below Q1 − IQR or above Q3 + IQR
- Below Q1 − 1.5×IQR or above Q3 + 1.5×IQR
- Below Q1 − 2×IQR or above Q3 + 2×IQR
Answer: C) Below Q1 − 1.5×IQR or above Q3 + 1.5×IQRExplanation: Tukey’s standard outlier detection rule.
Q15. The coefficient of variation is useful because it:
- Has no units
- Compares variability across datasets with different units
- Ignores the mean
- Is always less than 1
Answer: B) Compares variability across datasets with different unitsExplanation: CV = (SD/Mean)×100% is unit-free, enabling comparison.
Q16. Adding a constant to every value in a dataset changes the variance by:
- Adding the same constant
- Multiplying by the constant
- No change
- Squaring the constant
Answer: C) No changeExplanation: Variance measures spread; shifting all values equally doesn’t change spread.
Q17. Range is calculated as:
- Q3 − Q1
- Max − Min
- Mean − Median
- Variance²
Answer: B) Max − MinExplanation: Range is the simplest measure of dispersion.
Probability
Q18. The probability of the sample space S is:
- 0
- 0.5
- 1
- Undefined
Answer: C) 1Explanation: P(S) = 1 by definition; some outcome must occur.
Q19. If P(A) = 0.4, then P(A’) is:
- 0.4
- 0.5
- 0.6
- 1.4
Answer: C) 0.6Explanation: P(A’) = 1 − P(A) = 1 − 0.4 = 0.6.
Q20. For mutually exclusive events A and B:
- P(A∩B) = P(A)×P(B)
- P(A∩B) = 0
- P(A∪B) = 0
- P(A) = P(B)
Answer: B) P(A∩B) = 0Explanation: Mutually exclusive events cannot occur together.
Q21. Bayes’ theorem is used to:
- Calculate variance
- Update probabilities with new evidence
- Find the mode
- Compute IQR
Answer: B) Update probabilities with new evidenceExplanation: Bayes’ theorem: P(A|B) = P(B|A)P(A)/P(B).
Q22. In the Empirical Rule, approximately what percentage of data lies within μ ± 2σ?
- 68%
- 95%
- 99.7%
- 50%
**Answer: B) 95%**Explanation: The 68-95-99.7 rule for normal distributions.
Q23. The Central Limit Theorem states that sampling distribution of the mean approaches normality when:
- n < 10
- n ≥ 30
- Population is normal only
- Variance is zero
Answer: B) n ≥ 30Explanation: CLT holds for large samples (n ≥ 30) regardless of population shape.
Q24. A standard normal distribution has:
- μ = 1, σ = 0
- μ = 0, σ = 1
- μ = 0, σ = 0
- μ = 1, σ = 1
Answer: B) μ = 0, σ = 1Explanation: Standard normal Z has mean 0 and SD 1.
Q25. If P(A) = 0.5, P(B) = 0.4, and A, B are independent, P(A∩B) is:
- 0.9
- 0.2
- 0.1
- 0.5
Answer: B) 0.2Explanation: P(A∩B) = P(A)×P(B) = 0.5×0.4 = 0.2.
Q26. Which distribution models the number of successes in n independent trials?
- Normal
- Poisson
- Binomial
- Exponential
Answer: C) BinomialExplanation: Binomial = n independent Bernoulli trials.
Q27. The Poisson distribution models:
- Time between events
- Count of events in a fixed interval
- Height of people
- Coin toss outcomes
Answer: B) Count of events in a fixed intervalExplanation: Poisson is for rare event counts per unit time/space.
Correlation & Covariance
Q28. Covariance measures:
- Strength of relationship
- Direction of linear relationship
- Causation
- Probability
Answer: B) Direction of linear relationshipExplanation: Covariance sign indicates direction; magnitude is unit-dependent.
Q29. The range of Pearson’s correlation coefficient r is:
- 0 to 1
- −1 to 0
- −1 to +1
- −∞ to +∞
Answer: C) −1 to +1Explanation: r is standardized, bounded between −1 and +1.
Q30. If r = 0, it means:
- Perfect positive correlation
- Perfect negative correlation
- No linear relationship
- Strong causation
Answer: C) No linear relationshipExplanation: r = 0 means no linear relationship (may still be non-linear).
Q31. Correlation of +0.85 indicates:
- Weak positive relationship
- Strong positive relationship
- No relationship
- Strong negative relationship
Answer: B) Strong positive relationshipExplanation: |r| > 0.7 typically indicates strong correlation.
Q32. Which correlation is used for ordinal data or non-linear monotonic relationships?
- Pearson
- Spearman
- Covariance
- Z-score
Answer: B) SpearmanExplanation: Spearman’s rank correlation works on ranks, capturing monotonic relationships.
Q33. Correlation does not imply:
- Association
- Relationship
- Causation
- Linear trend
Answer: C) CausationExplanation: Classic statistical principle — correlation ≠ causation.
Q34. If Cov(X,Y) = 0, then:
- X and Y are independent
- X and Y have no linear relationship
- X and Y are perfectly correlated
- X = Y
Answer: B) X and Y have no linear relationshipExplanation: Zero covariance means no linear relationship; they may still be non-linearly dependent.
Q35. The formula for Pearson’s r is:
- Cov(X,Y) / (σₓ + σᵧ)
- Cov(X,Y) / (σₓ × σᵧ)
- Cov(X,Y) × σₓ × σᵧ
- σₓ × σᵧ / Cov(X,Y)
**Answer: B) Cov(X,Y) / (σₓ × σᵧ)**Explanation: r standardizes covariance by the product of standard deviations.
EDA
Q36. EDA stands for:
- Exploratory Data Analysis
- Estimated Data Analysis
- External Data Assessment
- Empirical Data Analytics
Answer: A) Exploratory Data AnalysisExplanation: Coined by John Tukey (1977).
Q37. Which plot is best for visualizing the distribution of a single numeric variable?
- Scatter plot
- Histogram
- Pie chart
- Heatmap
Answer: B) HistogramExplanation: Histograms show frequency distribution of a numeric variable.
Q38. A box plot displays all of the following EXCEPT:
- Median
- Quartiles
- Outliers
- Correlation
Answer: D) CorrelationExplanation: Box plots show five-number summary + outliers; not correlation.
Q39. Which EDA technique examines relationships between two numeric variables?
- Histogram
- Bar chart
- Scatter plot
- Pie chart
Answer: C) Scatter plotExplanation: Scatter plots reveal patterns, trends, and correlations between two numerics.
Q40. In EDA, the correct sequence is:
- Multivariate → Bivariate → Univariate
- Univariate → Bivariate → Multivariate
- Bivariate → Univariate → Multivariate
- Multivariate → Univariate → Bivariate
Answer: B) Univariate → Bivariate → MultivariateExplanation: Start simple (one variable) and progressively add complexity.
Q41. A Z-score of +3.5 indicates:
- The value is 3.5 units above the mean
- The value is 3.5 SDs above the mean
- The value is an outlier by IQR rule
- The value equals the mean
Answer: B) The value is 3.5 SDs above the meanExplanation: Z = (x − μ)/σ; |Z| > 3 typically flags outliers.
Q42. Which is NOT a method for handling missing data?
- Mean imputation
- KNN imputation
- Listwise deletion
- Correlation matrix
Answer: D) Correlation matrixExplanation: Correlation matrix is a visualization/analysis tool, not an imputation method.
Q43. A correlation heatmap is used to:
- Show missing values only
- Visualize pairwise correlations
- Plot time series
- Display categorical counts
Answer: B) Visualize pairwise correlationsExplanation: Heatmaps color-code correlation coefficients across variable pairs.
Q44. Which plot is best for detecting outliers in a single variable?
- Pie chart
- Box plot
- Stacked bar chart
- Line chart
Answer: B) Box plotExplanation: Box plots explicitly mark outliers beyond whiskers.
Q45. Seasonality in time-series data refers to:
- Random noise
- Regular periodic fluctuations
- Long-term increase
- One-time spike
Answer: B) Regular periodic fluctuationsExplanation: Seasonality repeats at fixed intervals (e.g., monthly, yearly).
Q46. Which measure is part of the five-number summary?
- Mean
- Mode
- Q1
- Variance
Answer: C) Q1Explanation: Five-number summary = Min, Q1, Median, Q3, Max.
Q47. Pair plots are most useful for:
- Single variable analysis
- Pairwise relationships in multivariate data
- Time-series forecasting
- Text analysis
Answer: B) Pairwise relationships in multivariate dataExplanation: Pair plots show scatter plots for all variable pairs plus distributions.
Q48. The best chart to show trend over time is:
- Pie chart
- Line chart
- Histogram
- Box plot
Answer: B) Line chartExplanation: Line charts connect data points chronologically to show trends.
Q49. Which is a robust (outlier-resistant) measure of central tendency?
- Mean
- Median
- Range
- Variance
Answer: B) MedianExplanation: Median depends only on position, not magnitude of extremes.
Q50. Skewness of a dataset can be detected using:
- Mean and Median comparison
- Only the mode
- Range alone
- Sample size
Answer: A) Mean and Median comparisonExplanation: If Mean > Median → right skew; Mean < Median → left skew.
SECTION B: THEORY QUESTIONS (20)
Q1. Define statistics. Differentiate between descriptive and inferential statistics with examples.
Answer: Statistics is the science of collecting, organizing, summarizing, analyzing, and interpreting data to support decision-making.
| Aspect | Descriptive Statistics | Inferential Statistics |
|---|---|---|
| Purpose | Summarize and describe data | Draw conclusions about population from sample |
| Tools | Mean, median, SD, charts | Hypothesis testing, confidence intervals, regression |
| Scope | Entire dataset | Sample → population |
| Example | Average marks of a class | Estimating national average from a sample survey |
Example: Computing the average salary of 100 employees is descriptive; using those 100 to estimate the average salary of all 10,000 employees in a company is inferential.
Q2. Explain the measures of central tendency. When should each be used?
Answer:
- Mean: Arithmetic average; best for symmetric numeric data without outliers. Formula: x̄ = Σxᵢ/n.
- Median: Middle value of sorted data; best for skewed data or when outliers exist. Robust to extremes.
- Mode: Most frequent value; best for categorical/nominal data. A dataset may be unimodal, bimodal, or multimodal.
When to use:
- Mean → symmetric distributions (e.g., heights)
- Median → income, house prices (skewed)
- Mode → favorite color, product preference
Relationship: Symmetric → Mean = Median = Mode; Right-skewed → Mode < Median < Mean; Left-skewed → Mean < Median < Mode.
Q3. Discuss measures of dispersion. Why is standard deviation preferred over variance?
Answer: Dispersion measures spread around the center:
- Range: Max − Min; simple but outlier-sensitive.
- Variance: Average squared deviation; σ² = Σ(x−μ)²/N.
- Standard Deviation: √Variance; in original units.
- IQR: Q3 − Q1; robust to outliers.
- CV: (SD/Mean)×100%; unit-free comparison.
Why SD over Variance:
- SD is in the same units as data, making interpretation intuitive.
- Variance is in squared units (e.g., squared rupees), which is meaningless in context.
- SD aligns with the normal distribution’s empirical rule (68-95-99.7).
Q4. Explain the Empirical Rule and Central Limit Theorem.
Answer:Empirical Rule (Normal Distribution):
- ~68% of data within μ ± 1σ
- ~95% within μ ± 2σ
- ~99.7% within μ ± 3σ
Central Limit Theorem (CLT): For a population with mean μ and finite variance σ², the sampling distribution of the sample mean x̄ approaches a normal distribution as sample size n increases, regardless of the population’s shape. Rule of thumb: n ≥ 30.
Significance:
- CLT enables inference (confidence intervals, hypothesis tests) without knowing population distribution.
- Foundation for many statistical procedures in data science.
Q5. Define covariance and correlation. How do they differ?
Answer:Covariance: Measures direction of linear relationship. Cov(X,Y) = Σ[(xᵢ−x̄)(yᵢ−ȳ)]/(n−1)
- Positive → variables move together
- Negative → variables move oppositely
- Magnitude depends on units → hard to interpret
Correlation (Pearson’s r): Standardized covariance. r = Cov(X,Y)/(σₓσᵧ), range −1 to +1.
| Aspect | Covariance | Correlation |
|---|---|---|
| Units | Product of units | Unit-free |
| Range | −∞ to +∞ | −1 to +1 |
| Interpretation | Direction only | Direction + strength |
| Sensitivity to scale | Yes | No |
Note: Correlation ≠ Causation.
Q6. What is Exploratory Data Analysis? Discuss its goals and steps.
Answer: EDA (coined by John Tukey, 1977) is the process of exploring data visually and statistically to understand structure, patterns, and anomalies before formal modeling.
Goals:
- Maximize insight into data
- Uncover underlying structure
- Extract important variables
- Detect outliers and anomalies
- Test assumptions
- Develop parsimonious models
Steps:
- Data Understanding: Dimensions, types, summary statistics
- Data Cleaning: Missing values, duplicates, inconsistencies
- Univariate Analysis: Histogram, box plot, mean/median
- Bivariate Analysis: Scatter plot, correlation, group comparisons
- Multivariate Analysis: Pair plots, heatmaps, PCA
- Pattern & Trend Analysis: Seasonality, cycles, trends
Q7. Explain various techniques for handling missing data.
Answer:
| Method | Description | Pros | Cons |
|---|---|---|---|
| Listwise deletion | Remove rows with missing values | Simple | Loses data |
| Mean/Median/Mode imputation | Replace with central value | Easy | Reduces variance |
| Forward/Backward fill | Use adjacent values | Good for time series | Ignores other variables |
| KNN imputation | Use similar records | Context-aware | Computationally heavy |
| Regression imputation | Predict from other variables | Uses relationships | Can overfit |
| Multiple imputation | Generate several plausible values | Statistically sound | Complex |
Choice depends on: Missingness mechanism (MCAR, MAR, MNAR), data size, and analysis goals.
Q8. What are outliers? Discuss methods for detecting and handling them.
Answer: Outliers are data points significantly different from the majority.
Detection Methods:
- IQR Method: Outlier if x < Q1 − 1.5×IQR or x > Q3 + 1.5×IQR
- Z-score: Outlier if |Z| > 3 where Z = (x−μ)/σ
- Visual: Box plots, scatter plots
- ML-based: Isolation Forest, DBSCAN
Handling Methods:
- Remove: If erroneous
- Cap/Winsorize: Replace with percentile values
- Transform: Log, square root
- Keep: If genuine and informative
- Separate analysis: Treat as a distinct group
Q9. Differentiate between Pearson and Spearman correlation.
Answer:
| Aspect | Pearson | Spearman |
|---|---|---|
| Measures | Linear relationship | Monotonic relationship |
| Data type | Interval/Ratio | Ordinal or Ranked |
| Formula | Cov(X,Y)/(σₓσᵧ) | 1 − 6Σd²/(n(n²−1)) |
| Outlier sensitivity | High | Low |
| Assumption | Linear, normal | Monotonic |
| Use case | Height vs weight | Ranking vs satisfaction |
Example: Pearson for temperature vs ice cream sales; Spearman for education rank vs income rank.
Q10. Explain the five-number summary and its role in box plots.
Answer: The five-number summary consists of:
- Minimum
- Q1 (25th percentile)
- Median (Q2)
- Q3 (75th percentile)
- Maximum
Box Plot Components:
- Box: Spans Q1 to Q3 (IQR)
- Line inside box: Median
- Whiskers: Extend to min/max within 1.5×IQR
- Points beyond whiskers: Outliers
Use: Quick visual comparison of distributions across groups; reveals skewness, spread, and outliers.
Q11. Describe the normal distribution and its properties.
Answer: The normal (Gaussian) distribution is a continuous probability distribution characterized by:
- Bell-shaped, symmetric curve
- Defined by μ (mean) and σ (SD)
- Mean = Median = Mode
- Empirical Rule: 68-95-99.7
- Total area under curve = 1
- Asymptotic: Tails approach but never touch x-axis
- Standard Normal: μ=0, σ=1; Z = (x−μ)/σ
Significance: Many natural phenomena follow normal distribution; CLT ensures sample means are normally distributed for large n.
Q12. What is skewness? How does it affect measures of central tendency?
Answer: Skewness measures asymmetry of a distribution.
- Symmetric: Mean = Median = Mode
- Right-skewed (positive): Tail extends right; Mode < Median < Mean
- Left-skewed (negative): Tail extends left; Mean < Median < Mode
Impact:
- In skewed data, mean is pulled toward the tail.
- Median remains robust and better represents the “typical” value.
- Example: Income distribution is right-skewed; median income is more representative than mean.
Formula (Pearson’s coefficient): Skewness = 3(Mean − Median)/SD
Q13. Discuss probability rules with examples.
Answer:
| Rule | Formula | Example |
|---|---|---|
| Complement | P(A’) = 1 − P(A) | P(not rain) = 1 − 0.3 = 0.7 |
| Addition (mutually exclusive) | P(A∪B) = P(A) + P(B) | P(1 or 2 on die) = 1/6 + 1/6 = 1/3 |
| Addition (general) | P(A∪B) = P(A)+P(B)−P(A∩B) | P(king or heart) = 4/52 + 13/52 − 1/52 = 16/52 |
| Multiplication (independent) | P(A∩B) = P(A)×P(B) | P(two heads) = 0.5×0.5 = 0.25 |
| Conditional | P(A | B) = P(A∩B)/P(B) |
| Bayes | P(A | B) = P(B |
Q14. Explain the data types in statistics with examples.
Answer:
| Type | Description | Examples | Operations |
|---|---|---|---|
| Nominal | Categories, no order | Gender, city, color | Mode, frequency |
| Ordinal | Ordered categories | Education level, ratings | Median, mode |
| Interval | Equal intervals, no true zero | Temperature (°C), IQ | Mean, SD |
| Ratio | Equal intervals, true zero | Height, weight, income | All operations |
Importance: Determines which statistical operations are valid. E.g., computing mean of nominal data is meaningless.
Q15. What is pattern identification and trend analysis in EDA?
Answer:Pattern identification involves discovering recurring structures in data:
- Trend: Long-term increase/decrease
- Seasonality: Regular periodic fluctuations (e.g., sales spike every December)
- Cyclical: Irregular long-term oscillations (e.g., economic cycles)
- Noise: Random variation
Trend Analysis Techniques:
- Moving Average: Smooths short-term fluctuations
- Exponential Smoothing: Weights recent observations more
- Time-Series Decomposition: Separates trend + seasonal + residual
- ACF/PACF: Autocorrelation for dependency detection
Application: Forecasting, anomaly detection, business planning.
Q16. Explain univariate, bivariate, and multivariate analysis.
Answer:
| Analysis | Variables | Techniques | Example |
|---|---|---|---|
| Univariate | One | Histogram, box plot, mean, SD | Distribution of ages |
| Bivariate | Two | Scatter plot, correlation, cross-tab | Age vs income |
| Multivariate | Three+ | Pair plot, heatmap, PCA, 3D scatter | Age, income, spending together |
Progression: Start univariate to understand each variable; move to bivariate to find relationships; then multivariate for complex interactions.
Q17. What is the coefficient of variation? Why is it useful?
Answer:CV = (Standard Deviation / Mean) × 100%
Usefulness:
- Unit-free: Compares variability across datasets with different units (e.g., weight in kg vs height in cm).
- Relative measure: Indicates spread as a percentage of the mean.
- Risk assessment: In finance, CV measures risk per unit of return.
Example: Stock A (mean return 10%, SD 2%, CV = 20%) vs Stock B (mean 15%, SD 4%, CV = 26.7%). Stock A has lower relative risk.
Limitation: Undefined when mean = 0.
Q18. Describe common data visualization techniques used in EDA.
Answer:
| Chart | Purpose | Best For |
|---|---|---|
| Histogram | Distribution of numeric variable | Shape, skewness |
| Box plot | Five-number summary, outliers | Comparing groups |
| Scatter plot | Relationship between two numerics | Correlation |
| Line chart | Trend over time | Time series |
| Bar chart | Compare categories | Categorical data |
| Pie chart | Proportions | Parts of whole (limited use) |
| Heatmap | Correlation/matrix visualization | Multivariate |
| Pair plot | All pairwise relationships | Quick multivariate overview |
Principles: Choose based on data type, question, and audience; avoid clutter; label axes; use color meaningfully.
Q19. Explain Bayes’ theorem with a real-world example.
Answer:Bayes’ Theorem: P(A|B) = [P(B|A) × P(A)] / P(B)
Example (Medical Test):
- Disease prevalence: P(D) = 0.01
- Test sensitivity: P(+|D) = 0.99
- False positive rate: P(+|no D) = 0.05
P(+) = P(+|D)P(D) + P(+|no D)P(no D) = 0.99×0.01 + 0.05×0.99 = 0.0099 + 0.0495 = 0.0594
P(D|+) = (0.99 × 0.01) / 0.0594 ≈ 0.167
Interpretation: Despite a positive test, only ~16.7% chance of actually having the disease — because the disease is rare. This illustrates the base rate fallacy.
Q20. Discuss the importance of EDA in the data science lifecycle.
Answer: EDA is critical because it:
- Reveals data quality issues: Missing values, duplicates, outliers
- Informs feature engineering: Which variables matter, transformations needed
- Guides model selection: Linear vs non-linear, distribution assumptions
- Prevents garbage-in-garbage-out: Ensures clean, understood data
- Generates hypotheses: Patterns suggest relationships to test
- Communicates insights: Visualizations aid stakeholder understanding
- Detects bias: Reveals sampling or measurement bias early
Position in lifecycle: After data acquisition/cleaning, before modeling. Iterative — may loop back to cleaning.
SECTION C: ANALYTICAL QUESTIONS (10)
Q1. Given the dataset: 12, 15, 18, 20, 22, 25, 30, 35, 40, 100. Calculate mean, median, mode, variance, SD, IQR, and identify outliers using the 1.5×IQR rule.
Answer:
Sorted: 12, 15, 18, 20, 22, 25, 30, 35, 40, 100
Mean: Sum = 317, n = 10 → x̄ = 31.7
Median: (22+25)/2 = 23.5
Mode: No repetition → No mode
Variance: Deviations from 31.7: −19.7, −16.7, −13.7, −11.7, −9.7, −6.7, −1.7, 3.3, 8.3, 68.3 Squares: 388.09, 278.89, 187.69, 136.89, 94.09, 44.89, 2.89, 10.89, 68.89, 4664.89 Sum = 5878.9 s² = 5878.9/9 = 653.21
SD: √653.21 = 25.56
IQR: Q1 = median of lower half (12,15,18,20,22) = 18 Q3 = median of upper half (25,30,35,40,100) = 35 IQR = 35 − 18 = 17
Outlier boundaries: Lower = 18 − 1.5×17 = −7.5 Upper = 35 + 1.5×17 = 60.5
Outliers: 100 (> 60.5) → 100 is an outlier.
Interpretation: The mean (31.7) is inflated by 100; median (23.5) better represents typical value. Right-skewed distribution.
Q2. Two datasets have the following statistics. Compare their variability using CV.
- Dataset A: Mean = 50, SD = 10
- Dataset B: Mean = 80, SD = 12
Answer:
CV(A) = (10/50) × 100% = 20%CV(B) = (12/80) × 100% = 15%
Conclusion: Although Dataset B has a larger absolute SD (12 > 10), Dataset A has higher relative variability (20% > 15%). CV reveals that Dataset A’s spread is larger relative to its mean.
Insight: Absolute SD can be misleading when means differ; CV provides fair comparison.
Q3. Compute Pearson’s correlation for X = {1, 2, 3, 4, 5} and Y = {2, 4, 5, 4, 5}. Interpret the result.
Answer:
x̄ = 3, ȳ = 4
| x | y | x−x̄ | y−ȳ | (x−x̄)(y−ȳ) | (x−x̄)² | (y−ȳ)² |
|---|---|---|---|---|---|---|
| 1 | 2 | −2 | −2 | 4 | 4 | 4 |
| 2 | 4 | −1 | 0 | 0 | 1 | 0 |
| 3 | 5 | 0 | 1 | 0 | 0 | 1 |
| 4 | 4 | 1 | 0 | 0 | 1 | 0 |
| 5 | 5 | 2 | 1 | 2 | 4 | 1 |
Σ(x−x̄)(y−ȳ) = 6 Σ(x−x̄)² = 10 Σ(y−ȳ)² = 6
r = 6 / √(10×6) = 6/√60 = 6/7.746 = 0.7746
Interpretation: r ≈ 0.77 → strong positive linear relationship. As X increases, Y tends to increase.
Q4. A dataset has Q1 = 25, Q3 = 55, Median = 40. Determine the shape of distribution and identify outlier boundaries.
Answer:
IQR = 55 − 25 = 30
Outlier boundaries:
- Lower = 25 − 1.5×30 = 25 − 45 = −20
- Upper = 55 + 1.5×30 = 55 + 45 = 100
Any value < −20 or > 100 is an outlier.
Shape:
- Distance Q1 to Median = 40 − 25 = 15
- Distance Median to Q3 = 55 − 40 = 15
Since distances are equal, the middle 50% is symmetric. However, overall shape depends on whiskers. If upper whisker is longer → right-skewed; if lower whisker is longer → left-skewed.
Given equal quartile spread: Likely symmetric around the median.
Q5. In a normal distribution with μ = 100 and σ = 15, what percentage of data lies between 85 and 115? What about between 70 and 130?
Answer:
Between 85 and 115: Z₁ = (85−100)/15 = −1 Z₂ = (115−100)/15 = +1 → Within μ ± 1σ → ~68%
Between 70 and 130: Z₁ = (70−100)/15 = −2 Z₂ = (130−100)/15 = +2 → Within μ ± 2σ → ~95%
Between 55 and 145: Z = ±3 → ~99.7%
Application: In IQ testing (μ=100, σ=15), ~68% score between 85–115; ~95% between 70–130.
Q6. A company’s monthly sales for 6 months: 120, 135, 150, 145, 160, 175. Compute the 3-month moving average and identify the trend.
Answer:
| Month | Sales | 3-Month MA |
|---|---|---|
| 1 | 120 | — |
| 2 | 135 | — |
| 3 | 150 | (120+135+150)/3 = 135 |
| 4 | 145 | (135+150+145)/3 = 143.33 |
| 5 | 160 | (150+145+160)/3 = 151.67 |
| 6 | 175 | (145+160+175)/3 = 160 |
Trend: Moving average increases from 135 → 143.33 → 151.67 → 160 → upward trend.
Insight: Smoothing removes month-to-month fluctuations, revealing consistent growth. Useful for forecasting.
Q7. Given P(A) = 0.3, P(B) = 0.5, P(A∩B) = 0.15. Are A and B independent? Find P(A∪B) and P(A|B).
Answer:
Independence check: P(A)×P(B) = 0.3×0.5 = 0.15 P(A∩B) = 0.15 Since P(A∩B) = P(A)×P(B) → A and B are independent.
P(A∪B) = P(A) + P(B) − P(A∩B) = 0.3 + 0.5 − 0.15 = 0.65
P(A|B) = P(A∩B)/P(B) = 0.15/0.5 = 0.3
Note: Since independent, P(A|B) = P(A) = 0.3, confirming independence.
Q8. A dataset has mean = 60, SD = 8. A value of 78 is observed. Is it an outlier using the Z-score method?
Answer:
Z-score = (x − μ)/σ = (78 − 60)/8 = 18/8 = 2.25
Decision:
- |Z| > 3 → definite outlier
- 2 < |Z| < 3 → potential outlier, investigate
- |Z| ≤ 2 → not an outlier
Conclusion: Z = 2.25 → Potential outlier (not extreme). It’s 2.25 SDs above the mean, which occurs in ~1.2% of cases in a normal distribution. Investigate context before removing.
Q9. Compare two investment options using CV:
- Stock X: Mean return = 12%, SD = 3%
- Stock Y: Mean return = 18%, SD = 6%
Which is better risk-adjusted?
Answer:
CV(X) = (3/12) × 100% = 25%CV(Y) = (6/18) × 100% = 33.33%
Interpretation:
- Stock Y has higher absolute return (18% > 12%) but also higher risk.
- Stock X has lower CV → less risk per unit of return.
- For risk-averse investors → Stock X is better risk-adjusted.
- For risk-seeking investors → Stock Y may be preferred for higher returns.
Insight: CV enables apples-to-apples comparison of investments with different scales.
Q10. Given the correlation matrix below, interpret the relationships:
| Age | Income | Spending | |
|---|---|---|---|
| Age | 1.00 | 0.65 | −0.20 |
| Income | 0.65 | 1.00 | 0.80 |
| Spending | −0.20 | 0.80 | 1.00 |
Answer:
Interpretations:
- Age vs Income (r = 0.65): Moderate positive — older people tend to earn more (career progression).
- Age vs Spending (r = −0.20): Weak negative — older people tend to spend slightly less.
- Income vs Spending (r = 0.80): Strong positive — higher income → higher spending.
- Diagonal (r = 1.00): Each variable perfectly correlated with itself.
Insights for Business:
- Income is the strongest predictor of spending.
- Age has opposing effects: positive on income, negative on spending.
- Multicollinearity concern: Income and Spending are highly correlated (0.80); may cause issues in regression if both used as predictors.
- Segmentation: Target high-income younger customers for premium products.
Caveat: Correlation ≠ causation. Age doesn’t cause income; confounding factors (experience, education) exist.