DS Unit 2: Questions & Answers
Unit 2: Data Acquisition & Data Preparation -> Generated and Prepared By Thiruselvan (ThiruXD)
Part A: Multiple Choice Questions (MCQs)
1. Approximately what percentage of a data scientist’s time is typically spent on finding, cleaning and preparing data rather than modelling?
A) 20–30%
B) 40–50%
C) 70–80%
D) 90–95%
Answer: C
2. Which of the following is a primary source of data?
A) Government census records
B) Kaggle datasets
C) A survey designed and conducted by the data science team for the current problem
D) Company ERP sales history
Answer: C
3. JSON API responses and XML files are classic examples of:
A) Structured data
B) Semi-structured data
C) Unstructured data
D) Binary data
Answer: B
4. Which data quality dimension is violated when the same person appears twice with slightly different name spellings or city names?
A) Accuracy
B) Completeness
C) Uniqueness
D) Timeliness
Answer: C
5. In the IQR method for outlier detection, a value is flagged as an outlier if it lies outside:
A) [Q1 − 1.5×IQR, Q3 + 1.5×IQR]
B) [Mean − 2σ, Mean + 2σ]
C) [Min, Max]
D) [Q1, Q3]
Answer: A
6. Which imputation method is most robust when the data contains extreme outliers?
A) Mean imputation
B) Median imputation
C) Mode imputation
D) Forward fill
Answer: B
7. Min-Max normalisation transforms a feature into the range:
A) (−1, 1)
B) (0, 1)
C) Mean = 0, Std = 1
D) Any arbitrary range
Answer: B
Formula: ( x' = \dfrac{x - \min}{\max - \min} )
8. One-Hot Encoding is preferred over Label Encoding when the categorical variable is:
A) Ordinal (has a natural order)
B) Nominal (no natural order)
C) Continuous
D) Binary only
Answer: B
9. Which of the following is not a typical challenge in data integration?
A) Schema mismatch
B) Entity resolution
C) Feature selection using PCA
D) Conflicting values from different sources
Answer: C
10. The golden rule of data cleaning is:
A) Always overwrite the original raw data after cleaning
B) Clean a copy, log every change, and keep the process reproducible
C) Delete all rows that contain any missing value
D) Apply the same imputation method to every column
Answer: B
Part B: Theory / Short Answer Questions
1. Differentiate between Primary and Secondary sources of data with two advantages and two disadvantages of each.
Model Answer:
- Primary: Collected by the team specifically for the current problem (surveys, experiments, interviews).
- Advantages: Fresh, exact fit to the question, full control over quality and ownership.
- Disadvantages: Expensive, time-consuming, requires expertise.
- Secondary: Already collected by someone else (census, company databases, Kaggle).
- Advantages: Cheap and fast, large historical coverage.
- Disadvantages: May not perfectly match the research question; quality and bias often unknown.
2. Explain the six dimensions of data quality with a short example for each.
Model Answer:
- Accuracy – value correctly reflects reality (no typos).
- Completeness – no required fields are missing.
- Consistency – same fact does not contradict itself across sources.
- Timeliness – data is recent enough to be relevant.
- Validity – values obey domain rules (age > 0).
- Uniqueness – each real-world entity appears only once.
3. List four common reasons for missing values and five standard techniques to handle them.
Model Answer:
Reasons: data-entry errors, non-response, sensor/system failure, merging mismatches.
Techniques: deletion (if few & random), mean/median/mode imputation, forward/backward fill, predictive imputation (regression/KNN), flag as “Unknown” category.
4. What is the difference between Min-Max normalisation and Z-score standardisation? When would you prefer each?
Model Answer:
- Min-Max: scales to [0, 1]. Preferred when you need values in a fixed range or when the distribution is not Gaussian.
- Z-score: transforms to mean = 0, std = 1. Preferred for distance-based algorithms (KNN, clustering, SVM) and when features have different units/scales.
5. Briefly explain Label Encoding versus One-Hot Encoding and state when each should be used.
Model Answer:
- Label Encoding assigns a single integer to each category. Use only for ordinal variables (Low/Medium/High) because it implies an artificial order.
- One-Hot Encoding creates a binary column for each category. Safe default for nominal variables; avoids false ordinal relationships but increases dimensionality.
Part C: Long Answer / Essay-Type Questions (Tough)
1. “Garbage in, garbage out” is the fundamental principle of data preparation. Critically discuss the six data-quality dimensions and, using a realistic dirty customer dataset, demonstrate how each dimension can be violated and how a proper cleaning pipeline would restore quality.
Model Answer Framework:
- Define each of the six dimensions with precise meaning.
- Present a dirty table containing duplicates, inconsistent city names, invalid ages, missing values, etc.
- Map each problem to the violated dimension.
- Walk through the cleaning pipeline (handle missing → remove duplicates → fix errors → smooth noise/outliers → standardise formats).
- Show the cleaned table and explain residual risks (e.g., imputation bias).
- Conclude with the importance of logging and reproducibility.
2. Design a complete data-preparation workflow for the Swiggy late-delivery prediction problem. For each stage (collection, quality assessment, cleaning, integration, transformation, reduction) specify:
- Data sources and types
- Likely quality issues
- Concrete techniques you would apply
- Tools you would use
- Potential pitfalls
Model Answer Framework:
Follow the exact six-step scenario given in the module and expand each step with technical detail, tools (Pandas, Scikit-learn, SQL, APIs), and risk analysis (selection bias, concept drift, entity resolution failures, etc.).
3. Compare and contrast structured, semi-structured and unstructured data across schema, storage, query ease, share of real-world data, and preprocessing effort. Then explain why modern data-science projects almost always involve a mixture of all three types and how the preprocessing pipeline must adapt.
Model Answer Framework:
- Detailed comparison table.
- Real-world example (e-commerce: order tables + JSON API responses + customer reviews/images).
- Discuss the extra steps required for unstructured data (NLP, computer vision, feature extraction) before it can join structured data.
- Argue that the ability to handle all three types is a core skill that separates junior from senior practitioners.
4. A dataset contains both continuous numeric features (Age, Salary, Distance) and categorical features (City, Weather, Payment Method). Explain, with formulas and examples, the complete sequence of transformation steps you would apply before feeding the data to a K-Nearest Neighbours model. Justify every choice.
Model Answer Framework:
- Handle missing values first.
- Encode categorical variables (One-Hot for nominal, Label only if ordinal).
- Apply Min-Max or Z-score to numeric features (justify Z-score for distance-based KNN).
- Discuss why order of operations matters.
- Mention dimensionality concerns after One-Hot and possible need for feature selection/reduction.
- Show numerical worked examples for both normalisation formulas.