BTCE | 5th Sem
Data Science SubjectUnit 3

DS Unit 3: Question with Answers

Unit III: Statistical Foundations and Exploratory Data Analysis -> Generated and Prepared By Thiruselvan (ThiruXD)

SECTION A: MULTIPLE CHOICE QUESTIONS (50 MCQs)


Descriptive Statistics & Central Tendency

Q1. Which measure of central tendency is most affected by outliers?

  1. Median
  2. Mode
  3. Mean
  4. Range

Answer: C) MeanExplanation: The mean uses every value in its calculation, so extreme values pull it significantly.


Q2. For a perfectly symmetric distribution:

  1. Mean > Median > Mode
  2. Mean = Median = Mode
  3. Mean < Median < Mode
  4. Mode > Mean > Median

Answer: B) Mean = Median = ModeExplanation: In a symmetric distribution, all three measures coincide at the center.


Q3. The mode of the dataset {2, 3, 3, 5, 7, 7, 7, 9} is:

  1. 3
  2. 5
  3. 7
  4. 9

Answer: C) 7Explanation: 7 appears three times, more than any other value.


Q4. For a right-skewed (positively skewed) distribution:

  1. Mean < Median < Mode
  2. Mode < Median < Mean
  3. Mean = Median = Mode
  4. Median < Mode < Mean

Answer: B) Mode < Median < MeanExplanation: The tail pulls the mean to the right, making it the largest.


Q5. The median of {10, 20, 30, 40, 50, 60} is:

  1. 30
  2. 35
  3. 40
  4. 45

Answer: B) 35Explanation: For even n, median = average of two middle values = (30+40)/2 = 35.


Q6. Which measure is best for categorical (nominal) data?

  1. Mean
  2. Median
  3. Mode
  4. Standard Deviation

Answer: C) ModeExplanation: Mode works on categories; mean/median require numerical ordering.


Q7. The arithmetic mean of the first 10 natural numbers is:

  1. 5
  2. 5.5
  3. 6
  4. 10

Answer: B) 5.5Explanation: Sum = 55, n = 10, Mean = 55/10 = 5.5.


Q8. If every value in a dataset is multiplied by 3, the mean:

  1. Stays the same
  2. Increases by 3
  3. Multiplies by 3
  4. Divides by 3

Answer: C) Multiplies by 3Explanation: Mean is a linear operator: Mean(3x) = 3 × Mean(x).


Measures of Dispersion

Q9. Variance is expressed in:

  1. Original units
  2. Squared units
  3. Percentage
  4. Logarithmic units

Answer: B) Squared unitsExplanation: Variance squares deviations, so units are squared. SD restores original units.


Q10. The standard deviation of {2, 4, 4, 4, 5, 5, 7, 9} (mean = 5) is:

  1. 1
  2. 2
  3. 3
  4. 4

Answer: B) 2Explanation: Variance = 32/8 = 4, SD = √4 = 2.


Q11. Which measure of dispersion is most robust to outliers?

  1. Range
  2. Variance
  3. Standard Deviation
  4. IQR

Answer: D) IQRExplanation: IQR uses quartiles and ignores extreme values.


Q12. If a dataset has variance 25, its standard deviation is:

  1. 5
  2. 25
  3. 625
  4. 12.5

Answer: A) 5Explanation: SD = √Variance = √25 = 5.


Q13. The IQR of a dataset with Q1 = 20 and Q3 = 50 is:

  1. 20
  2. 30
  3. 35
  4. 70

Answer: B) 30Explanation: IQR = Q3 − Q1 = 50 − 20 = 30.


Q14. Using the 1.5×IQR rule, an outlier is any value:

  1. Below Q1 or above Q3
  2. Below Q1 − IQR or above Q3 + IQR
  3. Below Q1 − 1.5×IQR or above Q3 + 1.5×IQR
  4. Below Q1 − 2×IQR or above Q3 + 2×IQR

Answer: C) Below Q1 − 1.5×IQR or above Q3 + 1.5×IQRExplanation: Tukey’s standard outlier detection rule.


Q15. The coefficient of variation is useful because it:

  1. Has no units
  2. Compares variability across datasets with different units
  3. Ignores the mean
  4. Is always less than 1

Answer: B) Compares variability across datasets with different unitsExplanation: CV = (SD/Mean)×100% is unit-free, enabling comparison.


Q16. Adding a constant to every value in a dataset changes the variance by:

  1. Adding the same constant
  2. Multiplying by the constant
  3. No change
  4. Squaring the constant

Answer: C) No changeExplanation: Variance measures spread; shifting all values equally doesn’t change spread.


Q17. Range is calculated as:

  1. Q3 − Q1
  2. Max − Min
  3. Mean − Median
  4. Variance²

Answer: B) Max − MinExplanation: Range is the simplest measure of dispersion.


Probability

Q18. The probability of the sample space S is:

  1. 0
  2. 0.5
  3. 1
  4. Undefined

Answer: C) 1Explanation: P(S) = 1 by definition; some outcome must occur.


Q19. If P(A) = 0.4, then P(A’) is:

  1. 0.4
  2. 0.5
  3. 0.6
  4. 1.4

Answer: C) 0.6Explanation: P(A’) = 1 − P(A) = 1 − 0.4 = 0.6.


Q20. For mutually exclusive events A and B:

  1. P(A∩B) = P(A)×P(B)
  2. P(A∩B) = 0
  3. P(A∪B) = 0
  4. P(A) = P(B)

Answer: B) P(A∩B) = 0Explanation: Mutually exclusive events cannot occur together.


Q21. Bayes’ theorem is used to:

  1. Calculate variance
  2. Update probabilities with new evidence
  3. Find the mode
  4. Compute IQR

Answer: B) Update probabilities with new evidenceExplanation: Bayes’ theorem: P(A|B) = P(B|A)P(A)/P(B).


Q22. In the Empirical Rule, approximately what percentage of data lies within μ ± 2σ?

  1. 68%
  2. 95%
  3. 99.7%
  4. 50%

**Answer: B) 95%**Explanation: The 68-95-99.7 rule for normal distributions.


Q23. The Central Limit Theorem states that sampling distribution of the mean approaches normality when:

  1. n < 10
  2. n ≥ 30
  3. Population is normal only
  4. Variance is zero

Answer: B) n ≥ 30Explanation: CLT holds for large samples (n ≥ 30) regardless of population shape.


Q24. A standard normal distribution has:

  1. μ = 1, σ = 0
  2. μ = 0, σ = 1
  3. μ = 0, σ = 0
  4. μ = 1, σ = 1

Answer: B) μ = 0, σ = 1Explanation: Standard normal Z has mean 0 and SD 1.


Q25. If P(A) = 0.5, P(B) = 0.4, and A, B are independent, P(A∩B) is:

  1. 0.9
  2. 0.2
  3. 0.1
  4. 0.5

Answer: B) 0.2Explanation: P(A∩B) = P(A)×P(B) = 0.5×0.4 = 0.2.


Q26. Which distribution models the number of successes in n independent trials?

  1. Normal
  2. Poisson
  3. Binomial
  4. Exponential

Answer: C) BinomialExplanation: Binomial = n independent Bernoulli trials.


Q27. The Poisson distribution models:

  1. Time between events
  2. Count of events in a fixed interval
  3. Height of people
  4. Coin toss outcomes

Answer: B) Count of events in a fixed intervalExplanation: Poisson is for rare event counts per unit time/space.


Correlation & Covariance

Q28. Covariance measures:

  1. Strength of relationship
  2. Direction of linear relationship
  3. Causation
  4. Probability

Answer: B) Direction of linear relationshipExplanation: Covariance sign indicates direction; magnitude is unit-dependent.


Q29. The range of Pearson’s correlation coefficient r is:

  1. 0 to 1
  2. −1 to 0
  3. −1 to +1
  4. −∞ to +∞

Answer: C) −1 to +1Explanation: r is standardized, bounded between −1 and +1.


Q30. If r = 0, it means:

  1. Perfect positive correlation
  2. Perfect negative correlation
  3. No linear relationship
  4. Strong causation

Answer: C) No linear relationshipExplanation: r = 0 means no linear relationship (may still be non-linear).


Q31. Correlation of +0.85 indicates:

  1. Weak positive relationship
  2. Strong positive relationship
  3. No relationship
  4. Strong negative relationship

Answer: B) Strong positive relationshipExplanation: |r| > 0.7 typically indicates strong correlation.


Q32. Which correlation is used for ordinal data or non-linear monotonic relationships?

  1. Pearson
  2. Spearman
  3. Covariance
  4. Z-score

Answer: B) SpearmanExplanation: Spearman’s rank correlation works on ranks, capturing monotonic relationships.


Q33. Correlation does not imply:

  1. Association
  2. Relationship
  3. Causation
  4. Linear trend

Answer: C) CausationExplanation: Classic statistical principle — correlation ≠ causation.


Q34. If Cov(X,Y) = 0, then:

  1. X and Y are independent
  2. X and Y have no linear relationship
  3. X and Y are perfectly correlated
  4. X = Y

Answer: B) X and Y have no linear relationshipExplanation: Zero covariance means no linear relationship; they may still be non-linearly dependent.


Q35. The formula for Pearson’s r is:

  1. Cov(X,Y) / (σₓ + σᵧ)
  2. Cov(X,Y) / (σₓ × σᵧ)
  3. Cov(X,Y) × σₓ × σᵧ
  4. σₓ × σᵧ / Cov(X,Y)

**Answer: B) Cov(X,Y) / (σₓ × σᵧ)**Explanation: r standardizes covariance by the product of standard deviations.


EDA

Q36. EDA stands for:

  1. Exploratory Data Analysis
  2. Estimated Data Analysis
  3. External Data Assessment
  4. Empirical Data Analytics

Answer: A) Exploratory Data AnalysisExplanation: Coined by John Tukey (1977).


Q37. Which plot is best for visualizing the distribution of a single numeric variable?

  1. Scatter plot
  2. Histogram
  3. Pie chart
  4. Heatmap

Answer: B) HistogramExplanation: Histograms show frequency distribution of a numeric variable.


Q38. A box plot displays all of the following EXCEPT:

  1. Median
  2. Quartiles
  3. Outliers
  4. Correlation

Answer: D) CorrelationExplanation: Box plots show five-number summary + outliers; not correlation.


Q39. Which EDA technique examines relationships between two numeric variables?

  1. Histogram
  2. Bar chart
  3. Scatter plot
  4. Pie chart

Answer: C) Scatter plotExplanation: Scatter plots reveal patterns, trends, and correlations between two numerics.


Q40. In EDA, the correct sequence is:

  1. Multivariate → Bivariate → Univariate
  2. Univariate → Bivariate → Multivariate
  3. Bivariate → Univariate → Multivariate
  4. Multivariate → Univariate → Bivariate

Answer: B) Univariate → Bivariate → MultivariateExplanation: Start simple (one variable) and progressively add complexity.


Q41. A Z-score of +3.5 indicates:

  1. The value is 3.5 units above the mean
  2. The value is 3.5 SDs above the mean
  3. The value is an outlier by IQR rule
  4. The value equals the mean

Answer: B) The value is 3.5 SDs above the meanExplanation: Z = (x − μ)/σ; |Z| > 3 typically flags outliers.


Q42. Which is NOT a method for handling missing data?

  1. Mean imputation
  2. KNN imputation
  3. Listwise deletion
  4. Correlation matrix

Answer: D) Correlation matrixExplanation: Correlation matrix is a visualization/analysis tool, not an imputation method.


Q43. A correlation heatmap is used to:

  1. Show missing values only
  2. Visualize pairwise correlations
  3. Plot time series
  4. Display categorical counts

Answer: B) Visualize pairwise correlationsExplanation: Heatmaps color-code correlation coefficients across variable pairs.


Q44. Which plot is best for detecting outliers in a single variable?

  1. Pie chart
  2. Box plot
  3. Stacked bar chart
  4. Line chart

Answer: B) Box plotExplanation: Box plots explicitly mark outliers beyond whiskers.


Q45. Seasonality in time-series data refers to:

  1. Random noise
  2. Regular periodic fluctuations
  3. Long-term increase
  4. One-time spike

Answer: B) Regular periodic fluctuationsExplanation: Seasonality repeats at fixed intervals (e.g., monthly, yearly).


Q46. Which measure is part of the five-number summary?

  1. Mean
  2. Mode
  3. Q1
  4. Variance

Answer: C) Q1Explanation: Five-number summary = Min, Q1, Median, Q3, Max.


Q47. Pair plots are most useful for:

  1. Single variable analysis
  2. Pairwise relationships in multivariate data
  3. Time-series forecasting
  4. Text analysis

Answer: B) Pairwise relationships in multivariate dataExplanation: Pair plots show scatter plots for all variable pairs plus distributions.


Q48. The best chart to show trend over time is:

  1. Pie chart
  2. Line chart
  3. Histogram
  4. Box plot

Answer: B) Line chartExplanation: Line charts connect data points chronologically to show trends.


Q49. Which is a robust (outlier-resistant) measure of central tendency?

  1. Mean
  2. Median
  3. Range
  4. Variance

Answer: B) MedianExplanation: Median depends only on position, not magnitude of extremes.


Q50. Skewness of a dataset can be detected using:

  1. Mean and Median comparison
  2. Only the mode
  3. Range alone
  4. Sample size

Answer: A) Mean and Median comparisonExplanation: If Mean > Median → right skew; Mean < Median → left skew.


SECTION B: THEORY QUESTIONS (20)


Q1. Define statistics. Differentiate between descriptive and inferential statistics with examples.

Answer: Statistics is the science of collecting, organizing, summarizing, analyzing, and interpreting data to support decision-making.

AspectDescriptive StatisticsInferential Statistics
PurposeSummarize and describe dataDraw conclusions about population from sample
ToolsMean, median, SD, chartsHypothesis testing, confidence intervals, regression
ScopeEntire datasetSample → population
ExampleAverage marks of a classEstimating national average from a sample survey

Example: Computing the average salary of 100 employees is descriptive; using those 100 to estimate the average salary of all 10,000 employees in a company is inferential.


Q2. Explain the measures of central tendency. When should each be used?

Answer:

  • Mean: Arithmetic average; best for symmetric numeric data without outliers. Formula: x̄ = Σxᵢ/n.
  • Median: Middle value of sorted data; best for skewed data or when outliers exist. Robust to extremes.
  • Mode: Most frequent value; best for categorical/nominal data. A dataset may be unimodal, bimodal, or multimodal.

When to use:

  • Mean → symmetric distributions (e.g., heights)
  • Median → income, house prices (skewed)
  • Mode → favorite color, product preference

Relationship: Symmetric → Mean = Median = Mode; Right-skewed → Mode < Median < Mean; Left-skewed → Mean < Median < Mode.


Q3. Discuss measures of dispersion. Why is standard deviation preferred over variance?

Answer: Dispersion measures spread around the center:

  • Range: Max − Min; simple but outlier-sensitive.
  • Variance: Average squared deviation; σ² = Σ(x−μ)²/N.
  • Standard Deviation: √Variance; in original units.
  • IQR: Q3 − Q1; robust to outliers.
  • CV: (SD/Mean)×100%; unit-free comparison.

Why SD over Variance:

  1. SD is in the same units as data, making interpretation intuitive.
  2. Variance is in squared units (e.g., squared rupees), which is meaningless in context.
  3. SD aligns with the normal distribution’s empirical rule (68-95-99.7).

Q4. Explain the Empirical Rule and Central Limit Theorem.

Answer:Empirical Rule (Normal Distribution):

  • ~68% of data within μ ± 1σ
  • ~95% within μ ± 2σ
  • ~99.7% within μ ± 3σ

Central Limit Theorem (CLT): For a population with mean μ and finite variance σ², the sampling distribution of the sample mean x̄ approaches a normal distribution as sample size n increases, regardless of the population’s shape. Rule of thumb: n ≥ 30.

Significance:

  • CLT enables inference (confidence intervals, hypothesis tests) without knowing population distribution.
  • Foundation for many statistical procedures in data science.

Q5. Define covariance and correlation. How do they differ?

Answer:Covariance: Measures direction of linear relationship. Cov(X,Y) = Σ[(xᵢ−x̄)(yᵢ−ȳ)]/(n−1)

  • Positive → variables move together
  • Negative → variables move oppositely
  • Magnitude depends on units → hard to interpret

Correlation (Pearson’s r): Standardized covariance. r = Cov(X,Y)/(σₓσᵧ), range −1 to +1.

AspectCovarianceCorrelation
UnitsProduct of unitsUnit-free
Range−∞ to +∞−1 to +1
InterpretationDirection onlyDirection + strength
Sensitivity to scaleYesNo

Note: Correlation ≠ Causation.


Q6. What is Exploratory Data Analysis? Discuss its goals and steps.

Answer: EDA (coined by John Tukey, 1977) is the process of exploring data visually and statistically to understand structure, patterns, and anomalies before formal modeling.

Goals:

  1. Maximize insight into data
  2. Uncover underlying structure
  3. Extract important variables
  4. Detect outliers and anomalies
  5. Test assumptions
  6. Develop parsimonious models

Steps:

  1. Data Understanding: Dimensions, types, summary statistics
  2. Data Cleaning: Missing values, duplicates, inconsistencies
  3. Univariate Analysis: Histogram, box plot, mean/median
  4. Bivariate Analysis: Scatter plot, correlation, group comparisons
  5. Multivariate Analysis: Pair plots, heatmaps, PCA
  6. Pattern & Trend Analysis: Seasonality, cycles, trends

Q7. Explain various techniques for handling missing data.

Answer:

MethodDescriptionProsCons
Listwise deletionRemove rows with missing valuesSimpleLoses data
Mean/Median/Mode imputationReplace with central valueEasyReduces variance
Forward/Backward fillUse adjacent valuesGood for time seriesIgnores other variables
KNN imputationUse similar recordsContext-awareComputationally heavy
Regression imputationPredict from other variablesUses relationshipsCan overfit
Multiple imputationGenerate several plausible valuesStatistically soundComplex

Choice depends on: Missingness mechanism (MCAR, MAR, MNAR), data size, and analysis goals.


Q8. What are outliers? Discuss methods for detecting and handling them.

Answer: Outliers are data points significantly different from the majority.

Detection Methods:

  1. IQR Method: Outlier if x < Q1 − 1.5×IQR or x > Q3 + 1.5×IQR
  2. Z-score: Outlier if |Z| > 3 where Z = (x−μ)/σ
  3. Visual: Box plots, scatter plots
  4. ML-based: Isolation Forest, DBSCAN

Handling Methods:

  • Remove: If erroneous
  • Cap/Winsorize: Replace with percentile values
  • Transform: Log, square root
  • Keep: If genuine and informative
  • Separate analysis: Treat as a distinct group

Q9. Differentiate between Pearson and Spearman correlation.

Answer:

AspectPearsonSpearman
MeasuresLinear relationshipMonotonic relationship
Data typeInterval/RatioOrdinal or Ranked
FormulaCov(X,Y)/(σₓσᵧ)1 − 6Σd²/(n(n²−1))
Outlier sensitivityHighLow
AssumptionLinear, normalMonotonic
Use caseHeight vs weightRanking vs satisfaction

Example: Pearson for temperature vs ice cream sales; Spearman for education rank vs income rank.


Q10. Explain the five-number summary and its role in box plots.

Answer: The five-number summary consists of:

  1. Minimum
  2. Q1 (25th percentile)
  3. Median (Q2)
  4. Q3 (75th percentile)
  5. Maximum

Box Plot Components:

  • Box: Spans Q1 to Q3 (IQR)
  • Line inside box: Median
  • Whiskers: Extend to min/max within 1.5×IQR
  • Points beyond whiskers: Outliers

Use: Quick visual comparison of distributions across groups; reveals skewness, spread, and outliers.


Q11. Describe the normal distribution and its properties.

Answer: The normal (Gaussian) distribution is a continuous probability distribution characterized by:

  • Bell-shaped, symmetric curve
  • Defined by μ (mean) and σ (SD)
  • Mean = Median = Mode
  • Empirical Rule: 68-95-99.7
  • Total area under curve = 1
  • Asymptotic: Tails approach but never touch x-axis
  • Standard Normal: μ=0, σ=1; Z = (x−μ)/σ

Significance: Many natural phenomena follow normal distribution; CLT ensures sample means are normally distributed for large n.


Q12. What is skewness? How does it affect measures of central tendency?

Answer: Skewness measures asymmetry of a distribution.

  • Symmetric: Mean = Median = Mode
  • Right-skewed (positive): Tail extends right; Mode < Median < Mean
  • Left-skewed (negative): Tail extends left; Mean < Median < Mode

Impact:

  • In skewed data, mean is pulled toward the tail.
  • Median remains robust and better represents the “typical” value.
  • Example: Income distribution is right-skewed; median income is more representative than mean.

Formula (Pearson’s coefficient): Skewness = 3(Mean − Median)/SD


Q13. Discuss probability rules with examples.

Answer:

RuleFormulaExample
ComplementP(A’) = 1 − P(A)P(not rain) = 1 − 0.3 = 0.7
Addition (mutually exclusive)P(A∪B) = P(A) + P(B)P(1 or 2 on die) = 1/6 + 1/6 = 1/3
Addition (general)P(A∪B) = P(A)+P(B)−P(A∩B)P(king or heart) = 4/52 + 13/52 − 1/52 = 16/52
Multiplication (independent)P(A∩B) = P(A)×P(B)P(two heads) = 0.5×0.5 = 0.25
ConditionalP(AB) = P(A∩B)/P(B)
BayesP(AB) = P(B

Q14. Explain the data types in statistics with examples.

Answer:

TypeDescriptionExamplesOperations
NominalCategories, no orderGender, city, colorMode, frequency
OrdinalOrdered categoriesEducation level, ratingsMedian, mode
IntervalEqual intervals, no true zeroTemperature (°C), IQMean, SD
RatioEqual intervals, true zeroHeight, weight, incomeAll operations

Importance: Determines which statistical operations are valid. E.g., computing mean of nominal data is meaningless.


Q15. What is pattern identification and trend analysis in EDA?

Answer:Pattern identification involves discovering recurring structures in data:

  • Trend: Long-term increase/decrease
  • Seasonality: Regular periodic fluctuations (e.g., sales spike every December)
  • Cyclical: Irregular long-term oscillations (e.g., economic cycles)
  • Noise: Random variation

Trend Analysis Techniques:

  1. Moving Average: Smooths short-term fluctuations
  2. Exponential Smoothing: Weights recent observations more
  3. Time-Series Decomposition: Separates trend + seasonal + residual
  4. ACF/PACF: Autocorrelation for dependency detection

Application: Forecasting, anomaly detection, business planning.


Q16. Explain univariate, bivariate, and multivariate analysis.

Answer:

AnalysisVariablesTechniquesExample
UnivariateOneHistogram, box plot, mean, SDDistribution of ages
BivariateTwoScatter plot, correlation, cross-tabAge vs income
MultivariateThree+Pair plot, heatmap, PCA, 3D scatterAge, income, spending together

Progression: Start univariate to understand each variable; move to bivariate to find relationships; then multivariate for complex interactions.


Q17. What is the coefficient of variation? Why is it useful?

Answer:CV = (Standard Deviation / Mean) × 100%

Usefulness:

  1. Unit-free: Compares variability across datasets with different units (e.g., weight in kg vs height in cm).
  2. Relative measure: Indicates spread as a percentage of the mean.
  3. Risk assessment: In finance, CV measures risk per unit of return.

Example: Stock A (mean return 10%, SD 2%, CV = 20%) vs Stock B (mean 15%, SD 4%, CV = 26.7%). Stock A has lower relative risk.

Limitation: Undefined when mean = 0.


Q18. Describe common data visualization techniques used in EDA.

Answer:

ChartPurposeBest For
HistogramDistribution of numeric variableShape, skewness
Box plotFive-number summary, outliersComparing groups
Scatter plotRelationship between two numericsCorrelation
Line chartTrend over timeTime series
Bar chartCompare categoriesCategorical data
Pie chartProportionsParts of whole (limited use)
HeatmapCorrelation/matrix visualizationMultivariate
Pair plotAll pairwise relationshipsQuick multivariate overview

Principles: Choose based on data type, question, and audience; avoid clutter; label axes; use color meaningfully.


Q19. Explain Bayes’ theorem with a real-world example.

Answer:Bayes’ Theorem: P(A|B) = [P(B|A) × P(A)] / P(B)

Example (Medical Test):

  • Disease prevalence: P(D) = 0.01
  • Test sensitivity: P(+|D) = 0.99
  • False positive rate: P(+|no D) = 0.05

P(+) = P(+|D)P(D) + P(+|no D)P(no D) = 0.99×0.01 + 0.05×0.99 = 0.0099 + 0.0495 = 0.0594

P(D|+) = (0.99 × 0.01) / 0.0594 ≈ 0.167

Interpretation: Despite a positive test, only ~16.7% chance of actually having the disease — because the disease is rare. This illustrates the base rate fallacy.


Q20. Discuss the importance of EDA in the data science lifecycle.

Answer: EDA is critical because it:

  1. Reveals data quality issues: Missing values, duplicates, outliers
  2. Informs feature engineering: Which variables matter, transformations needed
  3. Guides model selection: Linear vs non-linear, distribution assumptions
  4. Prevents garbage-in-garbage-out: Ensures clean, understood data
  5. Generates hypotheses: Patterns suggest relationships to test
  6. Communicates insights: Visualizations aid stakeholder understanding
  7. Detects bias: Reveals sampling or measurement bias early

Position in lifecycle: After data acquisition/cleaning, before modeling. Iterative — may loop back to cleaning.


SECTION C: ANALYTICAL QUESTIONS (10)


Q1. Given the dataset: 12, 15, 18, 20, 22, 25, 30, 35, 40, 100. Calculate mean, median, mode, variance, SD, IQR, and identify outliers using the 1.5×IQR rule.

Answer:

Sorted: 12, 15, 18, 20, 22, 25, 30, 35, 40, 100

Mean: Sum = 317, n = 10 → x̄ = 31.7

Median: (22+25)/2 = 23.5

Mode: No repetition → No mode

Variance: Deviations from 31.7: −19.7, −16.7, −13.7, −11.7, −9.7, −6.7, −1.7, 3.3, 8.3, 68.3 Squares: 388.09, 278.89, 187.69, 136.89, 94.09, 44.89, 2.89, 10.89, 68.89, 4664.89 Sum = 5878.9 s² = 5878.9/9 = 653.21

SD: √653.21 = 25.56

IQR: Q1 = median of lower half (12,15,18,20,22) = 18 Q3 = median of upper half (25,30,35,40,100) = 35 IQR = 35 − 18 = 17

Outlier boundaries: Lower = 18 − 1.5×17 = −7.5 Upper = 35 + 1.5×17 = 60.5

Outliers: 100 (> 60.5) → 100 is an outlier.

Interpretation: The mean (31.7) is inflated by 100; median (23.5) better represents typical value. Right-skewed distribution.


Q2. Two datasets have the following statistics. Compare their variability using CV.

  • Dataset A: Mean = 50, SD = 10
  • Dataset B: Mean = 80, SD = 12

Answer:

CV(A) = (10/50) × 100% = 20%CV(B) = (12/80) × 100% = 15%

Conclusion: Although Dataset B has a larger absolute SD (12 > 10), Dataset A has higher relative variability (20% > 15%). CV reveals that Dataset A’s spread is larger relative to its mean.

Insight: Absolute SD can be misleading when means differ; CV provides fair comparison.


Q3. Compute Pearson’s correlation for X = {1, 2, 3, 4, 5} and Y = {2, 4, 5, 4, 5}. Interpret the result.

Answer:

x̄ = 3, ȳ = 4

xyx−x̄y−ȳ(x−x̄)(y−ȳ)(x−x̄)²(y−ȳ)²
12−2−2444
24−10010
3501001
4410010
5521241

Σ(x−x̄)(y−ȳ) = 6 Σ(x−x̄)² = 10 Σ(y−ȳ)² = 6

r = 6 / √(10×6) = 6/√60 = 6/7.746 = 0.7746

Interpretation: r ≈ 0.77 → strong positive linear relationship. As X increases, Y tends to increase.


Q4. A dataset has Q1 = 25, Q3 = 55, Median = 40. Determine the shape of distribution and identify outlier boundaries.

Answer:

IQR = 55 − 25 = 30

Outlier boundaries:

  • Lower = 25 − 1.5×30 = 25 − 45 = −20
  • Upper = 55 + 1.5×30 = 55 + 45 = 100

Any value < −20 or > 100 is an outlier.

Shape:

  • Distance Q1 to Median = 40 − 25 = 15
  • Distance Median to Q3 = 55 − 40 = 15

Since distances are equal, the middle 50% is symmetric. However, overall shape depends on whiskers. If upper whisker is longer → right-skewed; if lower whisker is longer → left-skewed.

Given equal quartile spread: Likely symmetric around the median.


Q5. In a normal distribution with μ = 100 and σ = 15, what percentage of data lies between 85 and 115? What about between 70 and 130?

Answer:

Between 85 and 115: Z₁ = (85−100)/15 = −1 Z₂ = (115−100)/15 = +1 → Within μ ± 1σ → ~68%

Between 70 and 130: Z₁ = (70−100)/15 = −2 Z₂ = (130−100)/15 = +2 → Within μ ± 2σ → ~95%

Between 55 and 145: Z = ±3 → ~99.7%

Application: In IQ testing (μ=100, σ=15), ~68% score between 85–115; ~95% between 70–130.


Q6. A company’s monthly sales for 6 months: 120, 135, 150, 145, 160, 175. Compute the 3-month moving average and identify the trend.

Answer:

MonthSales3-Month MA
1120—
2135—
3150(120+135+150)/3 = 135
4145(135+150+145)/3 = 143.33
5160(150+145+160)/3 = 151.67
6175(145+160+175)/3 = 160

Trend: Moving average increases from 135 → 143.33 → 151.67 → 160 → upward trend.

Insight: Smoothing removes month-to-month fluctuations, revealing consistent growth. Useful for forecasting.


Q7. Given P(A) = 0.3, P(B) = 0.5, P(A∩B) = 0.15. Are A and B independent? Find P(A∪B) and P(A|B).

Answer:

Independence check: P(A)×P(B) = 0.3×0.5 = 0.15 P(A∩B) = 0.15 Since P(A∩B) = P(A)×P(B) → A and B are independent.

P(A∪B) = P(A) + P(B) − P(A∩B) = 0.3 + 0.5 − 0.15 = 0.65

P(A|B) = P(A∩B)/P(B) = 0.15/0.5 = 0.3

Note: Since independent, P(A|B) = P(A) = 0.3, confirming independence.


Q8. A dataset has mean = 60, SD = 8. A value of 78 is observed. Is it an outlier using the Z-score method?

Answer:

Z-score = (x − μ)/σ = (78 − 60)/8 = 18/8 = 2.25

Decision:

  • |Z| > 3 → definite outlier
  • 2 < |Z| < 3 → potential outlier, investigate
  • |Z| ≤ 2 → not an outlier

Conclusion: Z = 2.25 → Potential outlier (not extreme). It’s 2.25 SDs above the mean, which occurs in ~1.2% of cases in a normal distribution. Investigate context before removing.


Q9. Compare two investment options using CV:

  • Stock X: Mean return = 12%, SD = 3%
  • Stock Y: Mean return = 18%, SD = 6%

Which is better risk-adjusted?

Answer:

CV(X) = (3/12) × 100% = 25%CV(Y) = (6/18) × 100% = 33.33%

Interpretation:

  • Stock Y has higher absolute return (18% > 12%) but also higher risk.
  • Stock X has lower CV → less risk per unit of return.
  • For risk-averse investors → Stock X is better risk-adjusted.
  • For risk-seeking investors → Stock Y may be preferred for higher returns.

Insight: CV enables apples-to-apples comparison of investments with different scales.


Q10. Given the correlation matrix below, interpret the relationships:

AgeIncomeSpending
Age1.000.65−0.20
Income0.651.000.80
Spending−0.200.801.00

Answer:

Interpretations:

  1. Age vs Income (r = 0.65): Moderate positive — older people tend to earn more (career progression).
  2. Age vs Spending (r = −0.20): Weak negative — older people tend to spend slightly less.
  3. Income vs Spending (r = 0.80): Strong positive — higher income → higher spending.
  4. Diagonal (r = 1.00): Each variable perfectly correlated with itself.

Insights for Business:

  • Income is the strongest predictor of spending.
  • Age has opposing effects: positive on income, negative on spending.
  • Multicollinearity concern: Income and Spending are highly correlated (0.80); may cause issues in regression if both used as predictors.
  • Segmentation: Target high-income younger customers for premium products.

Caveat: Correlation ≠ causation. Age doesn’t cause income; confounding factors (experience, education) exist.


On this page