BTCE | 5th Sem
Data Science SubjectUnit 4

DS Unit 4: Question with Answers

Unit IV: Data Visualization, Analytics, and Real-World Case Studies -> Generated and Prepared By Thiruselvan (ThiruXD)

SECTION A: MULTIPLE CHOICE QUESTIONS (50 MCQs)


Introduction to Data Visualization

Q1. Data visualization is best defined as:

  1. Storing data in databases
  2. Graphical representation of data and information
  3. Cleaning raw data
  4. Writing SQL queries

Answer: B) Graphical representation of data and information

Explanation: Visualization transforms data into visual contexts like charts and graphs.


Q2. The human brain processes images approximately how many times faster than text?

  1. 10×
  2. 100×
  3. 1,000×
  4. 60,000×

Answer: D) 60,000×

Explanation: Visual processing is dramatically faster than reading text.


Q3. Who is credited with inventing the bar, line, and pie charts in the 18th century?

  1. John Tukey
  2. William Playfair
  3. Edward Tufte
  4. Florence Nightingale

Answer: B) William Playfair

Explanation: Playfair pioneered these foundational chart types.


Q4. Florence Nightingale is famous for which visualization?

  1. Pie chart
  2. Coxcomb diagram
  3. Scatter plot
  4. Histogram

Answer: B) Coxcomb diagram

Explanation: She used polar area diagrams to show mortality causes in the Crimean War.


Q5. Which is NOT a primary purpose of data visualization?

  1. Pattern discovery
  2. Storytelling
  3. Data encryption
  4. Decision support

Answer: C) Data encryption

Explanation: Encryption is a security function, unrelated to visualization.


Principles of Effective Visualization

Q6. The data-ink ratio principle was proposed by:

  1. William Playfair
  2. Edward Tufte
  3. John Tukey
  4. Hans Rosling

Answer: B) Edward Tufte

Explanation: Tufte emphasized maximizing data-ink and minimizing chartjunk.


Q7. “Chartjunk” refers to:

  1. Useful data points
  2. Unnecessary visual elements that distract from data
  3. Missing data
  4. Outliers in a dataset

Answer: B) Unnecessary visual elements that distract from data

Explanation: Coined by Tufte; decorations that don’t convey information.


Q8. Which Gestalt principle states that objects close together are perceived as grouped?

  1. Similarity
  2. Proximity
  3. Closure
  4. Continuity

Answer: B) Proximity

Explanation: Spatial closeness creates perceived grouping.


Q9. A color scheme going from light blue to dark blue for temperature data is:

  1. Diverging
  2. Categorical
  3. Sequential
  4. Random

Answer: C) Sequential

Explanation: Sequential palettes show ordered progression from low to high.


Q10. Which color combination should be avoided for accessibility?

  1. Blue-orange
  2. Red-green
  3. Purple-yellow
  4. Black-white

Answer: B) Red-green

Explanation: Red-green colorblindness is the most common form.


Q11. Starting a bar chart’s y-axis at a value other than zero:

  1. Improves accuracy
  2. Exaggerates differences
  3. Is always recommended
  4. Has no effect

Answer: B) Exaggerates differences

Explanation: Truncated axes mislead viewers about relative magnitudes.


Q12. A diverging color scheme is best for:

  1. Categorical data
  2. Data with a meaningful midpoint
  3. Time series
  4. Hierarchical data

Answer: B) Data with a meaningful midpoint

Explanation: Diverging schemes show deviation from a center (e.g., profit/loss).


Q13. The Gestalt principle of “closure” means:

  1. Objects close together are grouped
  2. The brain fills in missing information
  3. Similar objects are related
  4. Elements in a line are continuous

Answer: B) The brain fills in missing information

Explanation: Viewers perceive complete shapes even with gaps.


Types of Charts and Graphs

Q14. Which chart is best for showing the frequency distribution of a continuous variable?

  1. Pie chart
  2. Histogram
  3. Bar chart
  4. Scatter plot

Answer: B) Histogram

Explanation: Histograms bin continuous data and show frequency.


Q15. The key difference between a histogram and a bar chart is:

  1. Histogram bars touch; bar chart bars are separate
  2. Bar chart bars touch; histogram bars are separate
  3. They are identical
  4. Histograms are only for categorical data

Answer: A) Histogram bars touch; bar chart bars are separate

Explanation: Histograms represent continuous intervals; bar charts represent discrete categories.


Q16. A box plot displays all of the following EXCEPT:

  1. Median
  2. Quartiles
  3. Outliers
  4. Mean

Answer: D) Mean

Explanation: Box plots show five-number summary; mean is not typically displayed.


Q17. Which chart is best for showing the relationship between two numeric variables?

  1. Pie chart
  2. Bar chart
  3. Scatter plot
  4. Histogram

Answer: C) Scatter plot

Explanation: Scatter plots reveal correlations and patterns between two numerics.


Q18. A bubble chart is a scatter plot with:

  1. A third variable represented by bubble size
  2. Only categorical data
  3. No axes
  4. Time on both axes

Answer: A) A third variable represented by bubble size

Explanation: Bubble size encodes a third dimension of data.


Q19. Which chart is most suitable for showing parts of a whole?

  1. Line chart
  2. Pie chart
  3. Scatter plot
  4. Histogram

Answer: B) Pie chart

Explanation: Pie charts show proportions of a total.


Q20. The main limitation of a pie chart is:

  1. It cannot show percentages
  2. It is hard to compare many slices
  3. It requires numeric data only
  4. It cannot be colored

Answer: B) It is hard to compare many slices

Explanation: Human eyes struggle to compare angles; limit to 2–5 categories.


Q21. Which chart is best for showing trends over time?

  1. Pie chart
  2. Line chart
  3. Box plot
  4. Treemap

Answer: B) Line chart

Explanation: Line charts connect data points chronologically to show trends.


Q22. A candlestick chart is primarily used in:

  1. Healthcare
  2. Finance
  3. Agriculture
  4. Social media

Answer: B) Finance

Explanation: Candlesticks show open, high, low, close prices of securities.


Q23. A choropleth map uses:

  1. Points sized by value
  2. Color-coded regions by value
  3. Flow lines between locations
  4. 3D bars

Answer: B) Color-coded regions by value

Explanation: Choropleth maps shade geographic regions based on data values.


Q24. A treemap is best for showing:

  1. Time series
  2. Hierarchical proportions
  3. Correlation
  4. Distribution

Answer: B) Hierarchical proportions

Explanation: Nested rectangles represent hierarchical data proportions.


Q25. Which chart shows stages in a process like a sales pipeline?

  1. Funnel chart
  2. Scatter plot
  3. Histogram
  4. Box plot

Answer: A) Funnel chart

Explanation: Funnels show progressive reduction through stages.


Q26. A violin plot combines:

  1. Bar chart and line chart
  2. Box plot and density plot
  3. Pie chart and scatter plot
  4. Histogram and heatmap

Answer: B) Box plot and density plot

Explanation: Violin plots show distribution shape plus summary statistics.


Q27. A sparkline is:

  1. A large detailed chart
  2. A tiny inline chart showing trend
  3. A 3D chart
  4. A geographic map

Answer: B) A tiny inline chart showing trend

Explanation: Sparklines are compact charts embedded in text or tables.


Q28. A Sankey diagram is used to show:

  1. Distribution
  2. Flow between categories
  3. Correlation
  4. Outliers

Answer: B) Flow between categories

Explanation: Sankey diagrams visualize transfers or flows with varying widths.


Q29. Which chart is best for comparing distributions across multiple groups?

  1. Pie chart
  2. Box plot
  3. Line chart
  4. Gauge

Answer: B) Box plot

Explanation: Side-by-side box plots efficiently compare distributions.


Q30. A heatmap is most commonly used to visualize:

  1. Time series
  2. Correlation matrices
  3. Single values
  4. Text data

Answer: B) Correlation matrices

Explanation: Heatmaps use color intensity to show matrix values like correlations.


Dashboards

Q31. A dashboard is:

  1. A single chart
  2. A visual display of key metrics on one screen
  3. A database
  4. A programming language

Answer: B) A visual display of key metrics on one screen

Explanation: Dashboards consolidate KPIs for quick monitoring.


Q32. Which type of dashboard is designed for executives to track long-term KPIs?

  1. Operational
  2. Strategic
  3. Analytical
  4. Tactical

Answer: B) Strategic

Explanation: Strategic dashboards focus on high-level, long-term metrics.


Q33. Real-time server uptime monitoring is an example of which dashboard type?

  1. Strategic
  2. Operational
  3. Analytical
  4. Historical

Answer: B) Operational

Explanation: Operational dashboards monitor real-time processes.


Q34. In dashboard design, the most important KPIs should be placed:

  1. Bottom-right
  2. Top-left
  3. Center
  4. Randomly

Answer: B) Top-left

Explanation: Users read in an F-pattern, starting top-left.


Q35. KPI stands for:

  1. Key Performance Indicator
  2. Known Process Integration
  3. Key Parameter Index
  4. Kernel Processing Interface

Answer: A) Key Performance Indicator

Explanation: KPIs are measurable values showing progress toward goals.


Q36. Which dashboard tool is developed by Microsoft?

  1. Tableau
  2. Power BI
  3. Qlik Sense
  4. Looker Studio

Answer: B) Power BI

Explanation: Power BI is Microsoft’s business analytics service.


Q37. Tableau is owned by:

  1. Microsoft
  2. Google
  3. Salesforce
  4. Oracle

Answer: C) Salesforce

Explanation: Salesforce acquired Tableau in 2019.


Q38. “Drill-down” in a dashboard means:

  1. Deleting data
  2. Navigating from summary to detailed data
  3. Exporting to Excel
  4. Changing colors

Answer: B) Navigating from summary to detailed data

Explanation: Drill-downs let users explore deeper levels of data.


Storytelling with Data

Q39. The three pillars of data storytelling are:

  1. Data, visuals, narrative
  2. Code, database, server
  3. Mean, median, mode
  4. Tables, charts, maps

Answer: A) Data, visuals, narrative

Explanation: Effective storytelling combines all three elements.


Q40. Pre-attentive attributes are:

  1. Slow cognitive processes
  2. Visual properties processed instantly by the brain
  3. Statistical measures
  4. Database features

Answer: B) Visual properties processed instantly by the brain

Explanation: Color, size, position, and shape are processed pre-attentively.


Q41. Which is NOT a pre-attentive attribute?

  1. Color
  2. Size
  3. Position
  4. Statistical significance

Answer: D) Statistical significance

Explanation: Statistical significance requires conscious analysis, not pre-attention.


Q42. A good chart title should:

  1. Be vague
  2. State the key takeaway
  3. Only describe axes
  4. Be omitted

Answer: B) State the key takeaway

Explanation: Titles should communicate the insight, not just describe the chart.


Q43. The recommended narrative structure for data storytelling is:

  1. Random order
  2. Setup → Conflict → Resolution
  3. Only data tables
  4. Alphabetical

Answer: B) Setup → Conflict → Resolution

Explanation: A narrative arc engages the audience and drives action.


Visualization Tools

Q44. Which Python library is built on Matplotlib and provides statistical visualizations?

  1. Plotly
  2. Seaborn
  3. Bokeh
  4. Altair

Answer: B) Seaborn

Explanation: Seaborn extends Matplotlib with statistical plotting functions.


Q45. Which Python library is best for interactive web-based charts?

  1. Matplotlib
  2. Plotly
  3. NumPy
  4. Pandas

Answer: B) Plotly

Explanation: Plotly creates interactive, web-ready visualizations.


Q46. Matplotlib is best described as:

  1. A high-level statistical library
  2. A low-level foundational plotting library
  3. A database
  4. A dashboard tool

Answer: B) A low-level foundational plotting library

Explanation: Matplotlib offers full control for custom static plots.


Q47. Which tool is free and cloud-based from Google?

  1. Tableau
  2. Power BI
  3. Looker Studio
  4. Qlik Sense

Answer: C) Looker Studio

Explanation: Google Looker Studio (formerly Data Studio) is free and cloud-based.


Q48. D3.js is primarily used for:

  1. Database queries
  2. Custom web-based visualizations
  3. Statistical modeling
  4. Data cleaning

Answer: B) Custom web-based visualizations

Explanation: D3.js is a JavaScript library for bespoke web visualizations.


Real-World Case Studies

Q49. In healthcare analytics, predicting 30-day hospital readmissions primarily uses:

  1. Time-series forecasting
  2. Classification models
  3. Clustering
  4. Association rules

Answer: B) Classification models

Explanation: Readmission prediction is a binary classification problem.


Q50. In agriculture analytics, NDVI is used to measure:

  1. Soil pH
  2. Vegetation health from satellite imagery
  3. Rainfall
  4. Temperature

Answer: B) Vegetation health from satellite imagery

Explanation: NDVI (Normalized Difference Vegetation Index) assesses plant health.


SECTION B: THEORY QUESTIONS (20)


Q1. Define data visualization. Explain its importance in data science.

Answer: Data visualization is the graphical representation of data and information using visual elements like charts, graphs, maps, and dashboards. It transforms raw data into a visual context to make patterns, trends, and outliers easier to detect.

Importance in Data Science:

  1. Faster comprehension: The brain processes images ~60,000× faster than text.
  2. Pattern discovery: Reveals trends and anomalies invisible in tables.
  3. Storytelling: Communicates insights to non-technical stakeholders.
  4. Decision support: Enables data-driven decisions.
  5. Memory retention: Visual information is retained longer.
  6. Error detection: Visualizations expose data quality issues.
  7. Exploratory analysis: Essential for EDA before modeling.

Q2. Discuss the principles of effective data visualization.

Answer:

  1. Clarity: Message immediately understandable.
  2. Accuracy: Truthful representation; no distorted axes.
  3. Simplicity: Remove chartjunk; maximize data-ink ratio (Tufte).
  4. Relevance: Show only data supporting the message.
  5. Consistency: Uniform colors, scales, labels.
  6. Context: Titles, labels, units, legends.
  7. Audience awareness: Design for the target audience.
  8. Gestalt principles: Proximity, similarity, closure, continuity.
  9. Color theory: Sequential, diverging, categorical palettes.
  10. Accessibility: Colorblind-friendly palettes.

Data-Ink Ratio: Maximize ink used for data; minimize non-data ink.


Q3. Explain the Gestalt principles and their application in visualization.

Answer:

PrincipleDescriptionApplication
ProximityClose objects perceived as groupedGroup related data points
SimilaritySimilar objects perceived as relatedSame color for same category
EnclosureObjects within a boundary groupedBoxes, backgrounds
ClosureBrain fills missing informationLine charts with gaps
ContinuityElements in a line perceived as continuousTrend lines
ConnectionConnected objects perceived as relatedNetwork diagrams

Application: These principles guide layout, grouping, and color choices to make charts intuitive and reduce cognitive load.


Q4. Compare and contrast different types of charts used in data visualization.

Answer:

ChartPurposeBest ForLimitation
Bar chartCompare categoriesFew categoriesNot for continuous
HistogramDistributionContinuous dataBin size sensitive
Pie chartParts of whole2–5 categoriesHard to compare slices
Scatter plotRelationshipTwo numericsOverplotting
Line chartTrend over timeTime seriesNot for categories
Box plotDistribution summaryComparing groupsHides modality
HeatmapMatrix visualizationCorrelationColor perception issues
TreemapHierarchical proportionsNested dataHard to read labels
FunnelProcess stagesSales pipelinesOnly sequential stages
Violin plotDistribution shapeGroup comparisonComplex for laypeople

Selection depends on: Data type, question, audience, and message.


Q5. Explain the difference between histogram and bar chart. When should each be used?

Answer:

AspectHistogramBar Chart
Data typeContinuousCategorical
BarsTouch each otherSeparate
X-axisNumeric intervals (bins)Categories
PurposeShow distributionCompare categories
OrderFixed by valueCan be reordered
GapNo gapsGaps between bars

When to use:

  • Histogram: Distribution of age, income, temperature.
  • Bar chart: Sales by product, count by city, survey responses.

Common mistake: Using a bar chart for continuous data or a histogram for categories.


Q6. Describe the components and interpretation of a box plot.

Answer: A box plot (box-and-whisker plot) displays the five-number summary:

Components:

  1. Minimum: Lowest value (within 1.5×IQR)
  2. Q1: 25th percentile
  3. Median (Q2): 50th percentile
  4. Q3: 75th percentile
  5. Maximum: Highest value (within 1.5×IQR)
  6. Box: Spans Q1 to Q3 (IQR)
  7. Whiskers: Extend to min/max within 1.5×IQR
  8. Outliers: Points beyond whiskers

Interpretation:

  • Box height: Spread of middle 50%
  • Median position: Skewness indicator
  • Whisker length: Tail behavior
  • Outlier points: Extreme values

Use: Comparing distributions across groups; detecting outliers; assessing skewness.


Q7. What is a dashboard? Discuss its types and design principles.

Answer: A dashboard is a visual display of key metrics and trends consolidated on a single screen, enabling quick monitoring and decision-making.

Types:

TypePurposeAudience
StrategicLong-term KPIsExecutives
OperationalReal-time monitoringManagers
AnalyticalDeep analysisAnalysts
TacticalDepartmental progressTeam leads

Design Principles:

  1. Know your audience.
  2. Prioritize metrics (top-left placement).
  3. Use appropriate charts.
  4. Maintain consistency.
  5. Enable interactivity (filters, drill-downs).
  6. Avoid clutter (5–9 key metrics).
  7. Provide context (targets, benchmarks).
  8. Ensure performance.

Components: KPI cards, charts, filters, tables, alerts, drill-downs.


Q8. Explain the concept of storytelling with data. What are its three pillars?

Answer: Data storytelling is the art of combining data, visuals, and narrative to communicate insights and drive action.

Three Pillars:

  1. Data: Accurate, relevant, sufficient.
  2. Visuals: Clear, appropriate charts.
  3. Narrative: Context, meaning, call to action.

Framework (Knaflic):

  1. Understand the context (audience, message, action).
  2. Choose an appropriate visual.
  3. Eliminate clutter.
  4. Focus attention (pre-attentive attributes).
  5. Think like a designer.
  6. Tell a story (Setup → Conflict → Resolution).

Best Practices:

  • Lead with the “so what?”
  • Use descriptive titles stating the takeaway.
  • Annotate charts.
  • Use color strategically.
  • One message per chart.
  • End with a call to action.

Q9. What are pre-attentive attributes? How are they used in visualization?

Answer: Pre-attentive attributes are visual properties processed instantly by the brain without conscious effort.

Types:

  • Color (hue, intensity)
  • Size (length, area, volume)
  • Position (2D location)
  • Orientation (angle, slope)
  • Shape (form)
  • Motion (animation)

Use in Visualization:

  1. Highlight key data: Use contrasting color for important points.
  2. Guide the eye: Position critical information strategically.
  3. Group data: Use similar colors for related categories.
  4. Show magnitude: Use size to encode values.
  5. Reduce cognitive load: Pre-attentive processing is automatic.

Example: In a scatter plot, coloring outliers red draws immediate attention.


Q10. Compare Tableau and Power BI as visualization tools.

Answer:

AspectTableauPower BI
VendorSalesforceMicrosoft
Ease of useDrag-and-dropDrag-and-drop + DAX
Data connectivityExtensiveStrong Microsoft integration
CostExpensiveAffordable/freemium
VisualizationsBest-in-classVery good
CommunityLargeLarge and growing
MobileYesYes
LanguageVizQLDAX, M
Best forEnterprise analyticsBusiness users, Office integration

Tableau strengths: Superior visual design, fast rendering, advanced analytics. Power BI strengths: Affordable, Office 365 integration, frequent updates, natural language queries.


Q11. Discuss the Python visualization libraries used in data science.

Answer:

LibraryTypeStrengthsUse Case
MatplotlibLow-levelFull control, publication-qualityCustom static plots
SeabornStatisticalBeautiful defaults, statistical plotsEDA, statistical graphics
PlotlyInteractiveHover, zoom, 3D, DashWeb dashboards
BokehInteractiveStreaming data, server appsReal-time dashboards
AltairDeclarativeSimple syntax, Vega-LiteQuick interactive charts

Example Code:

# Matplotlib
plt.hist(df['age'], bins=20)
plt.show()

# Seaborn
sns.scatterplot(x='age', y='income', hue='gender', data=df)

# Plotly
px.scatter(df, x='age', y='income', color='gender').show()

Selection: Matplotlib for control, Seaborn for EDA, Plotly for interactivity.


Q12. Explain the principles of color theory in data visualization.

Answer: Color theory guides effective color use in charts.

Types of Color Schemes:

  1. Sequential: Light to dark for ordered data (e.g., temperature, population density).
  2. Diverging: Two hues from neutral center for data with meaningful midpoint (e.g., profit/loss, correlation).
  3. Categorical: Distinct colors for unordered categories (e.g., product types, regions).

Best Practices:

  • Limit to 5–7 colors.
  • Use colorblind-friendly palettes (avoid red-green).
  • Use color to encode meaning, not decoration.
  • Maintain consistent color meaning across charts.
  • Consider cultural associations.
  • Ensure sufficient contrast.

Common Mistakes:

  • Rainbow palettes for sequential data.
  • Too many colors.
  • Using color without legend.
  • Ignoring accessibility.

Q13. Describe the different types of dashboards with examples.

Answer:

TypePurposeAudienceExample
StrategicLong-term KPIsExecutivesAnnual revenue growth, market share
OperationalReal-time monitoringManagersDaily sales, server uptime, inventory
AnalyticalDeep analysisAnalystsCustomer segmentation, cohort analysis
TacticalDepartmental progressTeam leadsCampaign performance, project milestones

Strategic: Quarterly business review dashboards. Operational: Live call center metrics. Analytical: Marketing attribution analysis. Tactical: Weekly sales team performance.

Key difference: Strategic = “Where are we going?”; Operational = “What’s happening now?”; Analytical = “Why did it happen?”; Tactical = “How are we doing this quarter?”


Q14. What are the common mistakes in data visualization? How can they be avoided?

Answer:

MistakeProblemFix
Truncated y-axisExaggerates differencesStart at zero for bar charts
3D pie chartsDistorts proportionsUse 2D or bar charts
Too many colorsOverwhelmingLimit to 5–7 colors
Missing labelsUninterpretableLabel axes, units, legends
Pie chart with many slicesHard to compareUse bar chart
Dual axes misleadingFalse correlationsUse carefully with labels
ChartjunkDistracts from dataMaximize data-ink ratio
Wrong chart typeMisrepresents dataMatch chart to data/question
Ignoring audienceInappropriate complexityDesign for user expertise
No contextUninterpretableAdd benchmarks, targets

General Rule: Clarity > decoration; accuracy > aesthetics.


Q15. Explain the role of visualization in exploratory data analysis (EDA).

Answer: Visualization is central to EDA:

  1. Distribution analysis: Histograms, box plots reveal shape, spread, outliers.
  2. Relationship discovery: Scatter plots, heatmaps show correlations.
  3. Pattern identification: Line charts reveal trends, seasonality.
  4. Outlier detection: Box plots, scatter plots expose anomalies.
  5. Missing data patterns: Heatmaps show missingness structure.
  6. Group comparisons: Box plots, violin plots compare distributions.
  7. Multivariate exploration: Pair plots, PCA visualizations.
  8. Hypothesis generation: Visual patterns suggest relationships to test.

Tools: Matplotlib, Seaborn, Plotly, Pandas plotting.

Workflow: Univariate → Bivariate → Multivariate visual exploration.


Q16. Discuss the healthcare case study of predicting hospital readmissions.

Answer:Problem: Hospitals face penalties for high 30-day readmission rates.

Data: Patient demographics, diagnosis codes, medication history, length of stay, discharge disposition.

Approach:

  1. Data Collection: EHR (Electronic Health Records).
  2. EDA: Identify patterns — diabetics readmitted more; elderly at higher risk.
  3. Feature Engineering: Comorbidity index, medication count, prior admissions.
  4. Modeling: Logistic regression, random forest, gradient boosting.
  5. Visualization: Risk score dashboards for care coordinators.

Outcome: 15% reduction in readmissions; $2M annual savings.

Visualizations Used:

  • Heatmap of readmission rates by department
  • Kaplan-Meier survival curves
  • Feature importance bar charts
  • Patient risk score dashboards

Key Insight: Visualization made risk scores actionable for non-technical care staff.


Q17. Explain the finance case study of credit card fraud detection.

Answer:Problem: Detect fraudulent transactions in real time.

Data: Transaction amount, merchant category, location, time, customer history.

Approach:

  1. EDA: Fraud rate ~0.2%; highly imbalanced dataset.
  2. Feature Engineering: Velocity (transactions/hour), geo-distance, unusual merchant.
  3. Modeling: Isolation Forest, autoencoders, XGBoost.
  4. Visualization: Real-time alert dashboard; confusion matrix; ROC curves.

Outcome: 90% fraud detection with <1% false positive rate.

Visualizations:

  • Scatter plot of transaction amount vs time (fraud highlighted)
  • ROC curve
  • Confusion matrix heatmap
  • Real-time alert feed

Challenge: Class imbalance; concept drift; real-time processing.


Q18. Discuss the social media analytics case study of sentiment analysis.

Answer:Problem: Monitor brand perception across Twitter, Facebook, Instagram.

Data: Tweets, posts, comments, hashtags.

Approach:

  1. Data Collection: APIs (Twitter API, Facebook Graph).
  2. NLP: Tokenization, sentiment scoring (VADER, BERT).
  3. EDA: Volume over time, sentiment distribution.
  4. Visualization: Sentiment timeline, word clouds, influencer network graphs.

Outcome: Early detection of PR crises; 40% faster response time.

Visualizations:

  • Sentiment timeline (stacked area chart)
  • Word cloud of frequent terms
  • Network graph of influencers
  • Geographic heatmap of mentions

Key Insight: Real-time visualization enabled rapid crisis response.


Q19. Describe the agriculture case study of precision farming.

Answer:Problem: Optimize crop yield while minimizing water and fertilizer.

Data: Soil moisture, temperature, humidity, NDVI (satellite vegetation index), weather forecasts.

Approach:

  1. Data Collection: IoT sensors, drones, satellite imagery.
  2. EDA: Correlate soil moisture with yield.
  3. Modeling: Random forest for yield prediction; RL for irrigation scheduling.
  4. Visualization: Farm maps, time-series of soil metrics, yield prediction dashboards.

Outcome: 20% water savings; 12% yield increase.

Visualizations:

  • NDVI maps (choropleth)
  • Soil moisture time series
  • Yield prediction heatmaps
  • Drone imagery overlays

Key Insight: Visualization made complex sensor data actionable for farmers.


Q20. Explain the cybersecurity case study of network intrusion detection.

Answer:Problem: Detect anomalous network traffic indicating attacks.

Data: Network logs (source IP, destination IP, port, protocol, bytes, duration).

Approach:

  1. EDA: Baseline normal traffic patterns.
  2. Feature Engineering: Packet rate, connection count, entropy.
  3. Modeling: Isolation Forest, autoencoders, LSTM.
  4. Visualization: SOC (Security Operations Center) dashboards; attack maps.

Outcome: 95% detection rate; 60% reduction in false alerts.

Visualizations:

  • Network topology graphs
  • Anomaly score timelines
  • Geo-IP attack maps
  • Severity heatmaps

Key Insight: Real-time dashboards enabled rapid threat response.


SECTION C: ANALYTICAL QUESTIONS (10)


Q1. A dataset has the following chart requirements. Recommend the most appropriate chart type for each and justify:

(a) Monthly sales over 3 years(b) Distribution of customer ages(c) Relationship between advertising spend and revenue(d) Market share of 4 product categories(e) Comparison of salary distributions across 5 departments

Answer:

RequirementBest ChartJustification
(a) Monthly sales over 3 yearsLine chartShows trend over time; continuous temporal data
(b) Distribution of customer agesHistogramContinuous numeric variable; shows frequency distribution
(c) Ad spend vs revenueScatter plotTwo numeric variables; reveals correlation
(d) Market share of 4 categoriesPie chartParts of a whole; ≤5 categories is acceptable
(e) Salary across 5 departmentsBox plotCompares distributions across groups; shows median, IQR, outliers

Additional considerations:

  • For (d), a bar chart would also work and enable easier comparison.
  • For (a), add a moving average line to highlight trend.
  • For (c), add a regression line and color by region if available.

Q2. Analyze the following chart and identify its flaws:

A bar chart showing “Average Customer Satisfaction” with y-axis starting at 3.5 (out of 5), where Department A = 4.0 and Department B = 4.5. The bars appear to show B as three times taller than A.

Answer:

Flaws Identified:

  1. Truncated y-axis: Starting at 3.5 instead of 0 exaggerates differences. Department B appears 3× taller but is only 12.5% higher (4.5 vs 4.0).
  2. Misleading visual proportion: The visual ratio (3:1) does not match the data ratio (1.125:1).
  3. Missing context: No sample sizes, no error bars, no time period.
  4. No baseline reference: What is the industry average? What is “good”?

Corrections:

  • Start y-axis at 0 (bars would appear nearly equal).
  • Or use a dot plot to show actual values without bar-length bias.
  • Add error bars for confidence intervals.
  • Include sample sizes and time period.

Lesson: Truncated axes are one of the most common and misleading visualization mistakes.


Q3. A hospital wants to build a dashboard for monitoring patient readmissions. Design the dashboard layout and specify components.

Answer:

Dashboard Type: Operational + Analytical (hybrid)

Layout (F-pattern, top-left priority):

┌─────────────────────────────────────────────────────────────┐
│  HOSPITAL READMISSION DASHBOARD          [Date Range Filter] │
├──────────────┬──────────────┬──────────────┬────────────────┤
│ 30-Day       │ High-Risk    │ Avg Length   │ Cost Impact    │
│ Readmission  │ Patients     │ of Stay      │ ($)            │
│ Rate: 12.3%  │ 47           │ 5.2 days     │ $1.2M          │
│ ▼ 2.1%       │ ▲ 5          │ ▼ 0.3        │ ▲ $80K         │
├──────────────┴──────────────┴──────────────┴────────────────┤
│                                                              │
│  Readmission Trend (Line Chart)    │  Readmission by Dept   │
│  12 months                         │  (Bar Chart)           │
│                                    │                        │
├────────────────────────────────────┼────────────────────────┤
│                                    │                        │
│  Risk Factor Importance            │  Patient Risk Scores   │
│  (Horizontal Bar Chart)            │  (Table with color)    │
│                                    │                        │
└────────────────────────────────────┴────────────────────────┘

Components:

  1. KPI Cards: Readmission rate, high-risk count, avg LOS, cost impact (with trend arrows).
  2. Line Chart: 12-month readmission trend.
  3. Bar Chart: Readmission rate by department.
  4. Horizontal Bar: Feature importance (comorbidity, age, prior admissions).
  5. Table: Patient list with risk scores (color-coded).
  6. Filters: Date range, department, diagnosis category.

Design Principles Applied:

  • Most important KPI (readmission rate) at top-left.
  • Consistent color scheme (red = high risk, green = low).
  • Drill-down from department to patient level.
  • Alerts for threshold breaches.

Q4. Given a correlation heatmap showing the following values, interpret the relationships and suggest business actions:

PriceQualityReviewsSales
Price1.000.720.15−0.35
Quality0.721.000.600.45
Reviews0.150.601.000.78
Sales−0.350.450.781.00

Answer:

Interpretations:

  1. Price vs Quality (r = 0.72): Strong positive — higher-priced products are perceived as higher quality.
  2. Price vs Reviews (r = 0.15): Weak positive — price barely affects review count.
  3. Price vs Sales (r = −0.35): Moderate negative — higher prices reduce sales.
  4. Quality vs Reviews (r = 0.60): Moderate positive — better quality → more reviews.
  5. Quality vs Sales (r = 0.45): Moderate positive — quality drives sales.
  6. Reviews vs Sales (r = 0.78): Strong positive — reviews strongly drive sales.

Business Actions:

  • Prioritize quality: It drives both reviews and sales.
  • Encourage reviews: Strongest predictor of sales; incentivize customer feedback.
  • Price sensitivity: Consider premium pricing only if quality justifies it.
  • Multicollinearity concern: Quality and Reviews are correlated (0.60); be cautious in regression.
  • Segmentation: Target quality-conscious buyers willing to pay premium.

Caveat: Correlation ≠ causation; run A/B tests to confirm.


Q5. A marketing team has a dashboard showing campaign performance. Analyze the following data and recommend optimizations:

CampaignImpressionsClicksCTRConversionsConv. RateCostRevenueROAS
A100,0002,0002.0%1005.0%$1,000$3,0003.0
B50,0001,5003.0%755.0%$800$2,2502.8
C200,0002,0001.0%402.0%$1,200$1,2001.0
D80,0003,2004.0%1605.0%$1,500$6,4004.3

Answer:

Analysis:

CampaignCTRConv. RateROASAssessment
A2.0%5.0%3.0Good
B3.0%5.0%2.8Good CTR, lower ROAS
C1.0%2.0%1.0Poor — break-even only
D4.0%5.0%4.3Best performer

Recommendations:

  1. Scale Campaign D: Highest ROAS (4.3) and CTR (4.0%). Increase budget.
  2. Optimize Campaign C: Lowest CTR (1.0%) and conversion (2.0%). Either pause or redesign creative/targeting.
  3. Investigate Campaign B: Good CTR but lower ROAS — landing page or pricing issue?
  4. Maintain Campaign A: Solid performer; test incremental improvements.
  5. Visualization: Use a bubble chart with CTR (x), Conv. Rate (y), bubble size = Revenue, color = ROAS.

Dashboard Components:

  • KPI cards: Total ROAS, Total Revenue, Best Campaign
  • Bar chart: ROAS by campaign
  • Scatter plot: CTR vs Conversion Rate
  • Table: Detailed metrics with conditional formatting

Q6. Design a data storytelling narrative for presenting a 15% decline in customer satisfaction scores to executives.

Answer:

Narrative Structure: Setup → Conflict → Resolution

1. Setup (Context):

  • “Customer satisfaction has been our North Star metric, averaging 4.2/5 for three years.”
  • Visual: Line chart showing stable satisfaction over 3 years.

2. Conflict (Problem):

  • “This quarter, satisfaction dropped to 3.6/5 — a 15% decline.”
  • Visual: Highlighted drop on the line chart with annotation.
  • “The decline is concentrated in our mobile app users (−22%).”
  • Visual: Bar chart comparing satisfaction by channel.

3. Root Cause (Analysis):

  • “Analysis reveals the drop coincides with the v3.0 app release.”
  • Visual: Scatter plot of app rating vs time, with release date marked.
  • “Top complaints: slow loading (45%), login issues (30%), missing features (25%).”
  • Visual: Horizontal bar chart of complaint categories.

4. Resolution (Action):

  • “We recommend: (1) rollback of the loading optimization, (2) hotfix for login, (3) feature roadmap review.”
  • Visual: Timeline with milestones.
  • “Projected recovery to 4.0/5 within 8 weeks.”
  • Visual: Forecast line with confidence interval.

Storytelling Principles Applied:

  • Lead with the “so what?” (15% decline).
  • One message per chart.
  • Annotations explain key points.
  • Ends with clear call to action.

Q7. A cybersecurity team needs a dashboard for real-time threat monitoring. Design the dashboard and specify visualizations.

Answer:

Dashboard Type: Operational (real-time)

Layout:

┌─────────────────────────────────────────────────────────────┐
│  SECURITY OPERATIONS CENTER (SOC)        [Live] [Last 24h]  │
├──────────────┬──────────────┬──────────────┬────────────────┤
│ Active       │ Threats      │ Blocked      │ System         │
│ Threats: 12  │ Today: 1,247 │ IPs: 89      │ Health: 98%    │
│ ▲ 3          │ ▲ 15%        │ ▲ 12         │ ●              │
├──────────────┴──────────────┴──────────────┴────────────────┤
│                                                              │
│  Threat Timeline (Real-time Line)  │  Attack Sources (Map)  │
│                                    │                        │
├────────────────────────────────────┼────────────────────────┤
│                                    │                        │
│  Top Attack Types (Bar)            │  Severity Heatmap      │
│                                    │  (Time × Type)         │
│                                    │                        │
├────────────────────────────────────┴────────────────────────┤
│  Recent Alerts (Table with color-coded severity)            │
└─────────────────────────────────────────────────────────────┘

Components:

  1. KPI Cards: Active threats, threats today, blocked IPs, system health.
  2. Real-time Line Chart: Threat count over last 24 hours.
  3. Geo-IP Map: Attack sources worldwide.
  4. Bar Chart: Top attack types (DDoS, SQL injection, phishing).
  5. Heatmap: Severity by time and attack type.
  6. Alert Table: Recent events with severity color coding.
  7. Filters: Time range, severity, attack type.

Design Principles:

  • Red/yellow/green for severity.
  • Auto-refresh every 30 seconds.
  • Drill-down from alert to details.
  • Sound/visual alerts for critical threats.

Q8. Analyze the following Python code and explain what each line visualizes. Then suggest improvements.

import matplotlib.pyplot as plt
import seaborn as sns
import pandas as pd

df = pd.read_csv('sales.csv')

plt.figure(figsize=(10, 6))
plt.hist(df['revenue'], bins=30, color='steelblue', edgecolor='black')
plt.xlabel('Revenue')
plt.ylabel('Frequency')
plt.title('Revenue Distribution')
plt.show()

sns.scatterplot(x='ad_spend', y='revenue', hue='region', data=df)
plt.title('Ad Spend vs Revenue by Region')
plt.show()

sns.heatmap(df.corr(), annot=True, cmap='coolwarm', fmt='.2f')
plt.title('Correlation Matrix')
plt.show()

Answer:

Line-by-Line Explanation:

  1. import matplotlib.pyplot as plt — Imports Matplotlib for static plotting.
  2. import seaborn as sns — Imports Seaborn for statistical visualizations.
  3. import pandas as pd — Imports Pandas for data manipulation.
  4. df = pd.read_csv('sales.csv') — Loads sales data.
  5. plt.figure(figsize=(10, 6)) — Creates a 10×6 inch figure.
  6. plt.hist(df['revenue'], bins=30, ...) — Plots histogram of revenue with 30 bins.
  7. plt.xlabel/ylabel/title — Adds labels and title.
  8. plt.show() — Displays the histogram.
  9. sns.scatterplot(...) — Scatter plot of ad spend vs revenue, colored by region.
  10. sns.heatmap(df.corr(), ...) — Correlation heatmap with annotations.

Improvements:

  1. Histogram: Add KDE curve (sns.histplot(..., kde=True)) to show density.
  2. Scatter plot: Add regression line (sns.lmplot) to show trend.
  3. Heatmap: Mask upper triangle to reduce redundancy (mask=np.triu(df.corr())).
  4. Consistency: Use a consistent color palette across all charts.
  5. Context: Add mean/median lines to histogram.
  6. Interactivity: Consider Plotly for interactive versions.
  7. Accessibility: Use colorblind-friendly palettes.
  8. Titles: Make titles state the takeaway, not just describe.

Improved Code:

sns.histplot(df['revenue'], bins=30, kde=True, color='steelblue')
plt.axvline(df['revenue'].mean(), color='red', linestyle='--', label='Mean')
plt.legend()

Q9. A retail company wants to understand customer purchase patterns. Given the following visualizations, interpret the insights and recommend actions:

(a) Heatmap shows high correlation (0.85) between “time on site” and “purchase amount”(b) Box plot shows median purchase amount varies by age group: 18-25 (45),26−35(45), 26-35 (80), 36-50 (120),51+(120), 51+ (90)(c) Line chart shows purchases peak on weekends and during December(d) Scatter plot shows no relationship between gender and purchase amount

Answer:

Interpretations:

(a) Time on site vs Purchase amount (r = 0.85):

  • Strong positive correlation — engaged visitors spend more.
  • Action: Improve site engagement (recommendations, reviews, easy navigation).

(b) Purchase by age group:

  • 36–50 spend most (120median);18–25spendleast(120 median); 18–25 spend least (45).
  • Action: Target 36–50 with premium products; create entry-level offerings for 18–25.

(c) Temporal patterns:

  • Weekend and December peaks.
  • Action: Schedule promotions for weekends; prepare inventory for December; run weekday flash sales to boost traffic.

(d) Gender vs Purchase:

  • No relationship — gender is not a useful segmentation variable for spend.
  • Action: Focus segmentation on age and behavior instead of gender.

Overall Recommendations:

  1. Segmentation: Age-based and behavior-based.
  2. Engagement: Invest in site experience (strongest predictor).
  3. Timing: Align marketing with weekend and holiday peaks.
  4. Personalization: Recommend products based on browsing time.
  5. Dashboard: Build a customer analytics dashboard with these four views.

Visualization Strategy:

  • Heatmap for correlation discovery.
  • Box plots for segment comparison.
  • Line charts for temporal trends.
  • Scatter plots for relationship testing.

Q10. Design a complete visualization strategy for a healthcare provider wanting to monitor patient outcomes. Include chart types, dashboard layout, and storytelling approach.

Answer:

Objective: Monitor patient outcomes across departments, identify improvement areas, and communicate to stakeholders.

Visualization Strategy:

1. Data Sources:

  • EHR (diagnoses, treatments, outcomes)
  • Patient satisfaction surveys
  • Readmission records
  • Length of stay data
  • Cost data

2. Chart Selection by Question:

QuestionChart Type
What is the overall outcome rate?KPI card
How has it changed over time?Line chart
Which departments perform best/worst?Bar chart
What is the distribution of outcomes?Box plot
What factors correlate with outcomes?Heatmap
How do patient segments differ?Grouped bar chart
Where are outliers?Scatter plot

3. Dashboard Layout:

┌─────────────────────────────────────────────────────────────┐
│  PATIENT OUTCOMES DASHBOARD              [Dept] [Period]    │
├──────────────┬──────────────┬──────────────┬────────────────┤
│ Outcome Rate │ Readmission  │ Avg LOS      │ Satisfaction   │
│ 92.3%        │ 11.2%        │ 4.8 days     │ 4.3/5          │
│ ▲ 1.2%       │ ▼ 0.8%       │ ▼ 0.2        │ ▲ 0.1          │
├──────────────┴──────────────┴──────────────┴────────────────┤
│                                                              │
│  Outcome Trend (Line)              │  By Department (Bar)   │
│                                    │                        │
├────────────────────────────────────┼────────────────────────┤
│                                    │                        │
│  Outcome Distribution (Box)        │  Correlation (Heatmap) │
│                                    │                        │
├────────────────────────────────────┴────────────────────────┤
│  Patient Detail (Table with drill-down)                     │
└─────────────────────────────────────────────────────────────┘

4. Storytelling Approach:

Setup: “Our overall outcome rate is 92.3%, up 1.2% from last quarter.”Conflict: “However, the Cardiology department lags at 87%, and readmissions for diabetic patients are 18%.”

Analysis: “Correlation analysis shows that follow-up appointment compliance (r = 0.72) and medication adherence (r = 0.68) are the strongest predictors of positive outcomes.”

Resolution: “We recommend: (1) automated follow-up scheduling, (2) pharmacist-led medication counseling, (3) targeted intervention for diabetic patients. Projected outcome rate: 94.5% in 6 months.”

5. Audience-Specific Views:

  • Executives: Strategic KPIs, trends, financial impact.
  • Department heads: Department-specific metrics, benchmarks.
  • Care coordinators: Patient-level risk scores, action items.

6. Design Principles:

  • Red/yellow/green for performance.
  • Drill-down from hospital → department → patient.
  • Consistent color meaning.
  • Annotations for key events.
  • Mobile-responsive for rounds.


On this page