DS Unit 3 & 4: Q/A Bank
Generated and Prepared By Thiruselvan (ThiruXD)
Q1. Differentiate descriptive statistics from Exploratory Data Analysis with suitable objectives and examples.
Answer:
| Aspect | Descriptive Statistics | Exploratory Data Analysis (EDA) |
|---|---|---|
| Definition | Summarizes and describes the main features of a dataset using numerical measures and basic charts | A broader, iterative process of exploring data visually and statistically to uncover patterns, anomalies, and relationships before modeling |
| Objective | To condense data into a few meaningful numbers (mean, median, SD) | To maximize insight, detect outliers, test assumptions, and generate hypotheses |
| Scope | Narrow — focused on summary measures | Broad — includes cleaning, visualization, transformation, and pattern discovery |
| Tools | Mean, median, mode, variance, SD, range | Histograms, box plots, scatter plots, correlation matrices, pair plots |
| Output | Summary tables and simple charts | Insights, hypotheses, and direction for further analysis |
| Nature | Confirmatory (describes what is) | Exploratory (discovers what might be) |
Examples:
- Descriptive Statistics: Calculating the average salary of 100 employees = ₹45,000; SD = ₹8,000.
- EDA: Plotting salary distribution, discovering it is right-skewed, identifying 3 outliers, and finding that salary correlates with years of experience (r = 0.78).
Relationship: Descriptive statistics is a subset/tool used within EDA. EDA uses descriptive measures plus visualization to explore data holistically.
Q2. For the observations 5, 6, 8, 9, and 13, calculate the mean, median, and mode. Comment briefly on the effect of the largest value.
Answer:
Data: 5, 6, 8, 9, 13 (already sorted)
Mean: x̄ = (5 + 6 + 8 + 9 + 13) / 5 = 41 / 5 = 8.2
Median: n = 5 (odd) → middle value = 3rd value = 8
Mode: No value repeats → No mode
Effect of the largest value (13):
- The mean (8.2) is pulled above the median (8) because 13 is much larger than the other values.
- Without 13, the mean of {5, 6, 8, 9} = 28/4 = 7.0.
- The largest value increases the mean by 1.2 units but does not affect the median (which remains 8).
- Conclusion: The mean is sensitive to extreme values; the median is robust. For skewed data, the median is a better representative value.
Q3. Calculate the population variance and standard deviation for the data 2, 3, and 5.
Answer:
Data: 2, 3, 5 N = 3
Mean: μ = (2 + 3 + 5) / 3 = 10 / 3 = 3.333
Deviations and squares:
| x | x − μ | (x − μ)² |
|---|---|---|
| 2 | −1.333 | 1.778 |
| 3 | −0.333 | 0.111 |
| 5 | 1.667 | 2.778 |
Σ(x − μ)² = 1.778 + 0.111 + 2.778 = 4.667
Population Variance: σ² = Σ(x − μ)² / N = 4.667 / 3 = 1.556
Population Standard Deviation: σ = √1.556 = 1.247
Answer: Variance = 1.556, Standard Deviation = 1.247
Q4. A bag contains 5 red, 3 blue, and 2 green balls. One ball is selected at random. Find the probability that it is red or blue.
Answer:
Total balls: 5 + 3 + 2 = 10
Events:
- P(Red) = 5/10
- P(Blue) = 3/10
- Red and Blue are mutually exclusive (a ball cannot be both)
Using Addition Rule (mutually exclusive): P(Red ∪ Blue) = P(Red) + P(Blue) = 5/10 + 3/10 = 8/10 = 0.8 = 80%
Answer: The probability that the selected ball is red or blue is 0.8 (or 80%).
Q5. The correlation between study time and examination score is 0.85. Interpret its direction and strength.
Answer:
r = 0.85
Direction: Positive — as study time increases, examination scores tend to increase.
Strength: Strong — |r| = 0.85 is well above 0.7, indicating a strong linear relationship.
Interpretation:
- There is a strong positive linear relationship between study time and exam scores.
- Students who study more tend to score higher.
- r² = 0.85² = 0.7225 → approximately 72.25% of the variation in exam scores can be explained by study time (coefficient of determination).
Caveat: Correlation ≠ Causation. Other factors (IQ, teaching quality, prior knowledge) may influence scores.
Q6. A company wants to present monthly sales for one year. Select an appropriate chart and justify your choice.
Answer:
Recommended Chart: Line Chart
Justification:
- Time-series data: Monthly sales over 12 months is continuous temporal data.
- Trend visibility: Line charts clearly show upward/downward trends, peaks, and dips.
- Seasonality detection: Reveals seasonal patterns (e.g., December spikes).
- Continuity: The connected line emphasizes the flow of time.
- Comparison: Multiple years can be overlaid for comparison.
Alternative: A column chart could also work if the focus is on comparing discrete monthly values rather than trends. However, a line chart is superior for showing the overall pattern and trend.
Enhancement: Add a moving average line to smooth fluctuations and highlight the underlying trend.
Q7. For a dataset, Q1 = 18 and Q3 = 30. Using the IQR rule, determine whether the value 55 is an outlier.
Answer:
IQR = Q3 − Q1 = 30 − 18 = 12
Outlier boundaries:
- Lower boundary = Q1 − 1.5 × IQR = 18 − 1.5(12) = 18 − 18 = 0
- Upper boundary = Q3 + 1.5 × IQR = 30 + 1.5(12) = 30 + 18 = 48
Decision: Since 55 > 48, the value 55 is an outlier.
Answer: Yes, 55 is an outlier (it exceeds the upper boundary of 48).
Q8. The monthly salaries of five employees are 25, 26, 27, 28, and 90 thousand rupees. Compute the mean and median, and select the better representative value.
Answer:
Data (in thousands): 25, 26, 27, 28, 90
Mean: x̄ = (25 + 26 + 27 + 28 + 90) / 5 = 196 / 5 = 39.2 thousand rupees
Median: n = 5 (odd) → middle value = 3rd value = 27 thousand rupees
Comparison:
| Measure | Value | Effect of 90 |
|---|---|---|
| Mean | 39.2 | Pulled upward significantly |
| Median | 27 | Unaffected |
Better Representative Value: Median (27 thousand rupees)
Reason: The salary 90 is an extreme outlier. The mean (39.2) is inflated and does not represent the typical salary. Four out of five employees earn between 25–28 thousand. The median (27) accurately reflects the typical employee’s salary. For skewed data with outliers, the median is the better measure of central tendency.
Q9. Two production lines have the same mean output, but Line A has a standard deviation of 2 units and Line B has 9 units. Compare their consistency.
Answer:
Given:
- Mean output (same for both lines)
- SD(A) = 2 units
- SD(B) = 9 units
Comparison:
| Aspect | Line A | Line B |
|---|---|---|
| Standard Deviation | 2 | 9 |
| Variability | Low | High |
| Consistency | High | Low |
| Reliability | More predictable | Less predictable |
Interpretation:
- Line A is more consistent because its standard deviation is smaller (2 < 9).
- Line A’s output deviates from the mean by only ~2 units on average.
- Line B’s output deviates by ~9 units, indicating high variability and less predictability.
- Since both have the same mean, Line A is preferable for stable production.
Coefficient of Variation (if mean = 50):
- CV(A) = (2/50) × 100 = 4%
- CV(B) = (9/50) × 100 = 18%
- Line A has much lower relative variability.
Conclusion: Line A is 4.5 times more consistent than Line B (9/2 = 4.5).
Q10. Two fair coins are tossed together. Calculate the probability of obtaining exactly one head.
Answer:
Sample Space (S): {HH, HT, TH, TT} → Total outcomes = 4
Favorable outcomes (exactly one head): {HT, TH} → 2 outcomes
Probability: P(exactly one head) = 2/4 = 1/2 = 0.5 = 50%
Alternative method (Binomial): P(X = 1) = C(2,1) × (0.5)¹ × (0.5)¹ = 2 × 0.5 × 0.5 = 0.5
Answer: The probability of obtaining exactly one head is 0.5 (or 50%).
Q11. Ice-cream sales and drowning incidents both rise during summer. Explain why their positive correlation does not establish causation. (2 Marks)
Answer:
Positive correlation: Ice-cream sales and drowning incidents both increase during summer → r > 0.
Why correlation ≠ causation:
- Confounding variable: Summer heat is the lurking variable that causes both:
- Hot weather → more ice-cream purchases
- Hot weather → more swimming → more drownings
- No direct mechanism: Eating ice cream does not cause drowning.
- Coincidental trend: Both variables respond to a third factor (temperature), creating a spurious correlation.
Conclusion: The correlation is spurious — driven by a common external factor (summer), not by a cause-effect relationship between ice-cream sales and drowning. To establish causation, controlled experiments or causal inference methods are required.
Q12. Design a student-performance dashboard by proposing two KPIs, two suitable charts, and one useful filter.
Answer:
Dashboard: Student Performance Dashboard
Two KPIs:
- Average CGPA: Overall mean CGPA across all students (e.g., 7.8/10) — with trend indicator (▲/▼ vs last semester).
- Pass Percentage: Percentage of students passing all subjects (e.g., 92%) — with target benchmark.
Two Suitable Charts:
- Bar Chart — Subject-wise Average Marks: Compares performance across subjects (Math, Physics, CS, etc.). Quickly identifies weak subjects.
- Line Chart — CGPA Trend Over Semesters: Shows how average CGPA changes across semesters (Sem 1 to Sem 8). Reveals improvement or decline.
One Useful Filter:
- Department/Branch Filter: Allows viewing performance for CSE, ECE, ME, etc. Enables department-specific analysis.
Additional Components (optional):
- Scatter plot: Attendance vs CGPA (relationship)
- Box plot: CGPA distribution by department
- Table: Top 10 and bottom 10 performers
Design Principles:
- KPIs at top-left (F-pattern)
- Consistent color scheme
- Drill-down from department → student level
Q13. A report uses a three-dimensional pie chart with ten categories and similar colors. Identify two problems and recommend an improved visualization.
Answer:
Problems Identified:
- 3D Pie Chart Distortion: The 3D effect distorts proportions. Slices in the foreground appear larger than those in the background, misleading viewers about actual values.
- Too Many Categories (10): Human eyes cannot accurately compare 10 slices. Small slices become indistinguishable, and labels overlap.
- Similar Colors: Ten similar colors make it impossible to distinguish categories, especially for colorblind viewers.
Recommended Improved Visualization:
Option 1: Horizontal Bar Chart
- Sort bars in descending order by value.
- Clear labels and values on each bar.
- Easy comparison of all 10 categories.
- No distortion.
Option 2: Treemap
- Shows hierarchical proportions with area encoding.
- Handles many categories better than pie charts.
- Can use a sequential color palette.
Option 3: Top 5 + “Others” Pie Chart
- Show only the top 5 categories; group the rest as “Others.”
- Use distinct, colorblind-friendly colors.
- Add a bar chart for the detailed breakdown.
Best Recommendation: Horizontal Bar Chart — because it accurately represents values, handles 10 categories, and enables easy comparison.
Q14. Describe the main steps of Exploratory Data Analysis for identifying missing values, outliers, distributions, relationships, and trends.
Answer:
Step-by-Step EDA Process:
Step 1: Data Understanding
- Check dimensions (rows × columns), data types, and first/last rows.
- Use
df.info(),df.describe(),df.head().
Step 2: Missing Values
- Identify:
df.isnull().sum() - Visualize: Heatmap of missingness (
sns.heatmap(df.isnull())) - Handle: Deletion, mean/median/mode imputation, KNN imputation, or forward fill.
Step 3: Outlier Detection
- IQR Method: Values < Q1 − 1.5×IQR or > Q3 + 1.5×IQR
- Z-score Method: |Z| > 3
- Visual: Box plots, scatter plots
- Handle: Remove, cap, transform, or keep based on context.
Step 4: Distribution Analysis
- Numeric: Histograms, density plots, box plots
- Categorical: Bar charts, count plots
- Check: Shape (symmetric/skewed), modality, spread.
Step 5: Relationship Analysis
- Numeric vs Numeric: Scatter plots, correlation matrix, heatmap
- Categorical vs Numeric: Box plots, group means
- Categorical vs Categorical: Contingency tables, stacked bars
Step 6: Trend Analysis
- Time series: Line charts, moving averages
- Identify: Trend, seasonality, cycles, noise
- Tools: Decomposition, ACF/PACF
Step 7: Pattern Identification
- Clusters, correlations, associations
- Dimensionality reduction (PCA) for multivariate patterns
Step 8: Documentation & Iteration
- Summarize findings
- Form hypotheses
- Iterate back to cleaning if needed
Q15. Find the population variance and standard deviation of 5, 5, 5, and 5, and interpret the result.
Answer:
Data: 5, 5, 5, 5 N = 4
Mean: μ = (5 + 5 + 5 + 5) / 4 = 20 / 4 = 5
Deviations: Each value: 5 − 5 = 0
Squared deviations: 0² + 0² + 0² + 0² = 0
Population Variance: σ² = 0 / 4 = 0
Population Standard Deviation: σ = √0 = 0
Interpretation:
- Variance and SD are both zero.
- This means there is no variability in the data — all values are identical.
- Every observation equals the mean (5).
- In real-world data, zero variance is rare and often indicates a constant or error.
Q16. The probability of rain tomorrow is 0.30. Find the probability of no rain and explain the rule used.
Answer:
Given: P(Rain) = 0.30
Rule Used: Complement Rule P(A’) = 1 − P(A)
Calculation: P(No Rain) = 1 − P(Rain) = 1 − 0.30 = 0.70
Explanation: The complement rule states that the probability of an event not occurring is 1 minus the probability of it occurring. Since either it rains or it doesn’t (exhaustive events), the probabilities must sum to 1.
Answer: P(No Rain) = 0.70 (or 70%)
Q17. A correlation matrix reports r(X,Y) = -0.82, r(X,Z) = 0.25, and r(Y,Z) = 0.10. Identify the strongest relationship and interpret it.
Answer:
Correlation values:
- r(X,Y) = −0.82
- r(X,Z) = 0.25
- r(Y,Z) = 0.10
Strongest Relationship: X and Y (|r| = 0.82, the largest absolute value)
Interpretation:
- Direction: Negative — as X increases, Y tends to decrease.
- Strength: Strong — |r| = 0.82 > 0.7.
- r² = 0.82² = 0.6724 → ~67% of the variation in Y can be explained by X (and vice versa).
Other relationships:
- X and Z: Weak positive (r = 0.25).
- Y and Z: Very weak positive (r = 0.10) — negligible linear relationship.
Conclusion: The X–Y relationship is the strongest and most significant. It shows a strong inverse linear association.
Caveat: Correlation ≠ Causation.
Q18. Quarterly revenue values are 120, 135, 128, 150, and 168 lakh rupees over five consecutive quarters. Identify the overall pattern and choose a suitable chart.
Answer:
Data: 120, 135, 128, 150, 168 (lakh rupees)
Overall Pattern:
- Q1 → Q2: 120 → 135 (increase)
- Q2 → Q3: 135 → 128 (slight decrease)
- Q3 → Q4: 128 → 150 (increase)
- Q4 → Q5: 150 → 168 (increase)
Pattern: Overall upward trend with a minor dip in Q3. Revenue is growing consistently.
Growth Rate:
- Overall growth = (168 − 120) / 120 × 100 = 40% over 5 quarters.
Suitable Chart: Line Chart
Justification:
- Time-series data (consecutive quarters).
- Line chart clearly shows the upward trend and the Q3 dip.
- Enables forecasting and trend analysis.
- Can add a trend line or moving average for clarity.
Alternative: Column chart for discrete quarterly comparison; but line chart is superior for trend visualization.
Q19. Class A has Q1 = 42 and Q3 = 58, while Class B has Q1 = 35 and Q3 = 65. Compare their variability using the interquartile range.
Answer:
Class A: IQR(A) = Q3 − Q1 = 58 − 42 = 16
Class B: IQR(B) = Q3 − Q1 = 65 − 35 = 30
Comparison:
| Class | Q1 | Q3 | IQR | Variability |
|---|---|---|---|---|
| A | 42 | 58 | 16 | Lower |
| B | 35 | 65 | 30 | Higher |
Interpretation:
- Class B has greater variability (IQR = 30) compared to Class A (IQR = 16).
- The middle 50% of Class A’s data spans 16 units; Class B’s spans 30 units.
- Class A’s scores are more consistent/concentrated around the median.
- Class B’s scores are more spread out.
Conclusion: Class A is more consistent; Class B has higher dispersion in the middle 50% of its data.
Q20. A dashboard uses many unrelated colors, truncated axes, missing labels, and no visual hierarchy. Recommend four changes based on effective visualization principles.
Answer:
Four Recommended Changes:
- Use a Consistent, Meaningful Color Palette:
- Replace unrelated colors with a purposeful scheme (sequential, diverging, or categorical).
- Limit to 5–7 colors.
- Use color to encode meaning, not decoration.
- Ensure colorblind-friendly palettes.
- Fix Truncated Axes:
- Start bar chart axes at zero to avoid exaggerating differences.
- For line charts, truncated axes may be acceptable if clearly labeled, but avoid misleading scales.
- Maintain proportional representation.
- Add Missing Labels:
- Label all axes with variable names and units.
- Add descriptive titles that state the key takeaway.
- Include legends for colors and symbols.
- Add data source and timestamp.
- Establish Visual Hierarchy:
- Place most important KPIs at top-left (F-pattern reading).
- Use size, color, and position to indicate importance.
- Group related metrics together.
- Use white space to separate sections.
- Apply Gestalt principles (proximity, similarity) for grouping.
Additional Recommendations:
- Remove chartjunk (unnecessary gridlines, 3D effects).
- Ensure consistent fonts and formatting.
- Add filters for interactivity.
- Test with target audience.
Q21. A health survey must show the distribution of patient ages and the relationship between age and income. Select one chart for each objective and justify both choices.
Answer:
Objective 1: Distribution of Patient AgesChart: Histogram
Justification:
- Age is a continuous numeric variable.
- Histogram bins ages into intervals (e.g., 0–10, 10–20) and shows frequency.
- Reveals shape (symmetric/skewed), modality, and spread.
- Ideal for understanding the demographic profile.
Objective 2: Relationship Between Age and IncomeChart: Scatter Plot
Justification:
- Both age and income are numeric variables.
- Scatter plot reveals correlation, trends, clusters, and outliers.
- Can add a regression line to show the trend.
- Color-coding by gender or region adds a third dimension.
Summary:
| Objective | Chart | Why |
|---|---|---|
| Distribution of ages | Histogram | Shows frequency distribution of continuous data |
| Age vs Income relationship | Scatter plot | Shows relationship between two numeric variables |
Additional: A box plot could show income distribution by age group; a heatmap could show correlation strength.
Q22. For two variables, covariance is -18 and correlation is -0.72. Interpret both values without claiming causation.
Answer:
Covariance = −18:
- Direction: Negative — the two variables tend to move in opposite directions. When one increases, the other tends to decrease.
- Magnitude: The value −18 depends on the units of the variables, so its absolute size is not directly interpretable without context.
- Covariance only indicates direction, not strength.
Correlation = −0.72:
- Direction: Negative — confirms the inverse relationship.
- Strength: Moderate to strong (|r| = 0.72 is close to 0.7 threshold).
- r² = 0.72² = 0.5184 → ~52% of the variation in one variable is associated with the other.
- Correlation is unit-free and standardized, making it more interpretable than covariance.
Combined Interpretation:
- There is a moderate-to-strong negative linear relationship between the two variables.
- As one variable increases, the other tends to decrease.
- The negative covariance (−18) and negative correlation (−0.72) are consistent.
Important Caveat:
- Correlation ≠ Causation. We cannot claim that one variable causes the other to change.
- A third confounding variable may influence both.
- The relationship is linear; non-linear relationships may exist but are not captured by r.
Q23. One card is drawn from a standard deck of 52 cards. Find the probability that it is a face card or a heart.
Answer:
Total cards: 52
Events:
- Face cards: Jack, Queen, King in each suit = 4 × 3 = 12 cards
- Hearts: 13 cards
- Face cards that are hearts: Jack, Queen, King of hearts = 3 cards (intersection)
Using General Addition Rule: P(Face ∪ Heart) = P(Face) + P(Heart) − P(Face ∩ Heart)
P(Face) = 12/52 P(Heart) = 13/52 P(Face ∩ Heart) = 3/52
P(Face ∪ Heart) = 12/52 + 13/52 − 3/52 = 22/52 = 11/26 ≈ 0.4231
Answer: The probability is 11/26 ≈ 0.423 (or 42.3%).
Q24. Datasets A = {20, 20, 20, 20} and B = {10, 15, 25, 30} have the same mean. Compare their dispersion using variance or standard deviation.
Answer:
Dataset A: 20, 20, 20, 20 Mean(A) = 80/4 = 20
Dataset B: 10, 15, 25, 30 Mean(B) = 80/4 = 20
Both have the same mean (20).
Dataset A — Variance: All deviations = 0 σ²(A) = 0/4 = 0 σ(A) = 0
Dataset B — Variance:
| x | x − 20 | (x − 20)² |
|---|---|---|
| 10 | −10 | 100 |
| 15 | −5 | 25 |
| 25 | 5 | 25 |
| 30 | 10 | 100 |
Σ(x − μ)² = 100 + 25 + 25 + 100 = 250 σ²(B) = 250/4 = 62.5 σ(B) = √62.5 = 7.91
Comparison:
| Dataset | Mean | Variance | SD |
|---|---|---|---|
| A | 20 | 0 | 0 |
| B | 20 | 62.5 | 7.91 |
Interpretation:
- Dataset A has zero dispersion — all values are identical.
- Dataset B has significant dispersion — values range from 10 to 30.
- Although both have the same mean, Dataset B is far more variable.
- Conclusion: Mean alone does not describe a dataset; dispersion measures are essential.
Q25. For the observations 10, 12, 12, 13, 18, and 25, calculate the mean, median, and mode, and comment on their differences.
Answer:
Data (sorted): 10, 12, 12, 13, 18, 25 n = 6
Mean: x̄ = (10 + 12 + 12 + 13 + 18 + 25) / 6 = 90 / 6 = 15
Median: n = 6 (even) → average of 3rd and 4th values = (12 + 13) / 2 = 12.5
Mode: 12 appears twice → Mode = 12
Comparison:
| Measure | Value |
|---|---|
| Mean | 15 |
| Median | 12.5 |
| Mode | 12 |
Comment on Differences:
- Mean > Median > Mode → The distribution is right-skewed (positively skewed).
- The value 25 is an extreme high value that pulls the mean upward.
- The median (12.5) is more representative of the typical value.
- The mode (12) shows the most common observation.
- Ordering: Mode (12) < Median (12.5) < Mean (15) confirms right skewness.
- Recommendation: For skewed data, use the median as the measure of central tendency.