DS Unit 1 & 2 : Analytical Questions
Generated and Prepared By Thiruselvan (ThiruXD)
1. A sensor network in a smart factory frequently drops connection resulting in missing temperature readings. Formulate a strategy to handle these missing values without losing valuable time series trends.
Strategy:
- Diagnose the missingness pattern: Check if missing values are MCAR, MAR, or MNAR. In sensor networks, dropouts are often intermittent and temporary (MAR related to network status).
- Prefer time-aware methods over simple mean/median:
- Forward fill / Backward fill (carry last known value) for short gaps.
- Linear or spline interpolation for moderate gaps to preserve trend and smoothness.
- Seasonal decomposition + imputation or Kalman filtering if strong periodic patterns exist.
- For longer gaps: Use a predictive model (LSTM, Prophet, or KNN with time features) trained on surrounding sensors or historical patterns.
- Additional safeguards:
- Flag imputed values with a binary indicator column so the model knows they are estimated.
- Set a maximum gap length; beyond that, treat the segment as missing or use multi-sensor fusion.
- Never use global mean/median — it destroys local trends and seasonality.
This approach retains temporal continuity while minimising distortion of the underlying process.
2. You have discovered noisy data in patients’ blood pressure records where some readings are negative numbers. Formulate a method to identify and correct these anomalies.
Method:
- Detection:
- Domain rule: Systolic and diastolic BP cannot be ≤ 0 or unrealistically high (e.g., > 300 mmHg). Flag all values outside physiological range (typically 40–250 mmHg for systolic).
- Statistical methods: IQR rule (values outside Q1−1.5×IQR or Q3+1.5×IQR) and Z-score (|z| > 3).
- Visual inspection: Box plots and histograms.
- Correction strategy (in order of preference):
- Investigate root cause first (sensor error, data-entry mistake, unit conversion error).
- If few anomalies and neighbouring readings exist → interpolate or use forward/backward fill.
- If clear data-entry error (e.g., negative sign accidentally added) → correct the sign or replace with absolute value after verification.
- Otherwise → treat as missing and impute using median of similar patients (same age, gender, diagnosis) or predictive models.
- Cap extreme but still positive outliers at a clinically reasonable upper limit (winsorisation).
- Documentation: Log every change and create a quality flag so clinicians and models know which values were corrected.
Never simply delete rows containing negative BP — it may remove critically ill patients.
3. A company relies heavily on AI tools but lacks a fundamental data science strategy. Analyse the potential pitfalls of jumping straight into AI without a solid data science foundation.
Major pitfalls:
- Garbage-in, garbage-out: AI models amplify poor data quality (missing values, bias, inconsistency). Results look sophisticated but are unreliable.
- Wrong problem framing: Teams optimise the wrong metric or solve a non-business-critical problem.
- High project failure rate: Industry data shows 85–90% of AI/ML projects never reach production when data foundations are weak.
- Lack of reproducibility and governance: No data lineage, version control, or monitoring → models degrade silently (data/concept drift).
- Ethical and regulatory risk: Biased training data leads to discriminatory outcomes; no model cards or audit trails → legal exposure (GDPR, DPDP, EU AI Act).
- Wasted investment: Expensive GPUs and AutoML tools cannot compensate for dirty, siloed, or insufficient data.
- Talent and process gaps: Without proper data engineering, EDA, and evaluation discipline, the organisation cannot diagnose why models fail.
Conclusion: AI is a powerful layer built on top of solid data science (problem understanding → quality data → lifecycle → evaluation → monitoring). Skipping the foundation turns AI into expensive theatre.
4. A dataset contains ages ranging from 18 to 90 and incomes ranging from 20,000 USD to 1,50,000 USD. Analyse why passing this raw data to a distance-based machine learning algorithm will cause issues.
Core problem – Scale dominance:
Distance-based algorithms (KNN, K-Means, hierarchical clustering, SVM with RBF kernel, etc.) rely on Euclidean (or similar) distance.
Income values are orders of magnitude larger than age. A difference of $10,000 in income completely overwhelms a difference of 20 years in age. Consequently:
- The algorithm effectively ignores the Age feature.
- Neighbours or clusters are determined almost solely by income.
- Model performance and interpretability collapse.
Solution: Apply feature scaling before modelling:
- Z-score standardisation (preferred for most distance-based methods) → mean 0, variance 1.
- Or Min-Max normalisation to [0,1].
After scaling, both features contribute on equal footing.
5. Explain how the volume, velocity and variety of big data impact the data acquisition phase in a modern e-commerce application.
- Volume: Millions of daily transactions, clickstreams, and product views generate terabytes. Traditional single-server databases cannot store or query this scale → must use distributed storage (HDFS, S3, data lakes) and sampling or incremental loading strategies during acquisition.
- Velocity: Real-time clickstreams, inventory updates, and pricing changes arrive continuously. Batch collection (nightly dumps) is insufficient → need streaming acquisition via Kafka, Kinesis, or Flink, plus near-real-time pipelines.
- Variety: Structured (orders, payments), semi-structured (JSON product catalogues, API logs), and unstructured (customer reviews, images, chat transcripts). Acquisition systems must support multiple connectors, schema-on-read, and flexible ingestion rather than rigid ETL into a single relational schema.
Impact: Acquisition architecture must be scalable, multi-modal, and streaming-capable, otherwise downstream analysis is already compromised.
6. Evaluate the impact of using mean imputation vs k-nearest neighbours imputation on a dataset with 30% missing values in a critical healthcare feature.
Mean imputation:
- Pros: Simple, fast, preserves sample size.
- Cons (severe at 30% missing):
- Reduces variance → underestimates uncertainty.
- Distorts correlations and relationships with other variables.
- Introduces bias if data is not MCAR.
- In healthcare, can hide high-risk patients (e.g., mean BP masks hypertensive cases).
- Model performance and clinical reliability suffer significantly.
KNN imputation:
- Pros: Uses similarity to other patients → preserves local structure, correlations, and complex patterns better.
- Cons: Computationally expensive, sensitive to choice of k and distance metric, can still propagate noise if neighbours are poor.
- At 30% missingness it is far superior to mean, but still not perfect; may require multiple imputation or domain-informed models for critical features.
Recommendation: Avoid mean imputation for a critical healthcare feature at 30% missingness. Prefer KNN, iterative (MICE), or model-based imputation, plus sensitivity analysis and missingness indicators.
7. In an agricultural context, if a farmer wants to optimise crop yield, map out what data sources they would collect and how a data scientist should utilise them.
Data sources:
- Primary / IoT: Soil moisture, temperature, pH, nutrient sensors; weather stations; drone / satellite imagery.
- Secondary: Historical yield records, government soil maps, market prices, pest/disease databases.
- External: Hyper-local weather forecasts, satellite NDVI data, seed/fertiliser quality data.
How a data scientist utilises them:
- Integrate multi-modal data (time-series sensors + imagery + tabular history).
- Clean: handle sensor dropouts, calibrate readings, remove cloud-contaminated satellite images.
- Feature engineering: growing degree days, soil water balance, vegetation indices.
- Build predictive models (yield forecast) and prescriptive models (optimal irrigation/fertiliser schedules).
- Deploy recommendations via mobile app or dashboard with uncertainty estimates.
- Continuously monitor and retrain as new seasons arrive.
8. If a project fails during the evaluation phase because the model is inaccurate, analyse which previous lifecycle phases might be responsible and why.
Possible responsible phases:
- Business Understanding: Problem was poorly framed or success metrics were wrong → model optimises the wrong target.
- Data Collection: Insufficient volume, biased sampling, or missing critical variables → model lacks necessary information.
- Data Cleaning / Preprocessing: Unhandled missing values, noise, or leakage → model learns artefacts.
- Feature Engineering: Weak or irrelevant features; failure to capture interactions or domain knowledge.
- Modelling: Inappropriate algorithm choice, inadequate hyperparameter tuning, or data leakage between train/test.
- Even Evaluation design itself: Wrong metric (e.g., accuracy on imbalanced data) or contaminated test set.
Most common root causes lie in the early phases (problem framing and data quality), not just the modelling step.
9. During data collection from web scraping you notice significant HTML tags and irrelevant symbols in your text data. Outline the specific data cleaning techniques required.
Required techniques (in sequence):
- Remove HTML tags: Use BeautifulSoup, lxml, or regular expressions (
re.sub(r'<.*?>', '', text)). - Decode HTML entities: Convert
&, , etc. to proper characters. - Remove special symbols and punctuation (or selectively keep if needed for sentiment).
- Normalise whitespace: Collapse multiple spaces, tabs, newlines.
- Lower-casing (unless case is meaningful).
- Remove or replace URLs, email addresses, phone numbers.
- Handle non-ASCII / emoji (remove or convert to text descriptions).
- Optional advanced: Spelling correction, stop-word removal, lemmatisation/stemming — depending on downstream task.
- Validation: Check residual noise with sample inspection and character-frequency analysis.
Always keep the original raw scrape for auditability.
10. A company wants to predict housing prices. One feature is “neighbourhood” with 50 distinct text labels. Analyse the pros and cons of One-Hot Encoding for this feature.
Pros:
- No artificial ordinal relationship is imposed (neighbourhoods are nominal).
- Most algorithms can use the resulting binary features directly.
- Preserves all information about which neighbourhood a house belongs to.
Cons:
- Curse of dimensionality: 50 new binary columns → sparse matrix, increased memory and computation.
- Risk of multicollinearity (dummy-variable trap) if not handled (drop one category).
- Rare neighbourhoods create very sparse columns that may overfit or be dropped by regularisation.
- Tree-based models (Random Forest, XGBoost) can handle categorical features natively or with target/ordinal encoding more efficiently; One-Hot is often unnecessary and suboptimal for them.
- If new neighbourhoods appear in production, the encoding breaks unless a robust pipeline is built.
Practical recommendation:
- For linear models / neural nets → One-Hot (or embeddings).
- For tree-based models → prefer native categorical support, target encoding, or CatBoost encoding.
- Consider grouping rare neighbourhoods into an “Other” category to reduce dimensionality.