BTCE | 5th Sem
Data Science SubjectUnit 1

DS Unit 1: Questions & Answers

Unit 1: Foundations of Data Science -> Generated and Prepared By Thiruselvan (ThiruXD)

Part A: Multiple Choice Questions (MCQs)

1. Which of the following best explains why the 2006 release of Apache Hadoop is considered a turning point in the evolution of Data Science?

A) It introduced the term “Data Scientist”

B) It enabled distributed storage and processing of massive datasets on commodity hardware

C) It formalised Exploratory Data Analysis

D) It marked the beginning of Generative AI

Answer: B

Explanation: Hadoop (HDFS + MapReduce) democratised Big Data analytics by allowing organisations to process petabyte-scale data without expensive supercomputers.

2. In the Data Science lifecycle, a model shows high training accuracy but poor performance on the test set. The most appropriate next step is:

A) Immediately deploy the model and monitor

B) Increase the size of the training set only

C) Revisit feature engineering and check for overfitting using cross-validation

D) Switch to a more complex deep learning architecture without further analysis

Answer: C

3. Which statement correctly distinguishes Data Science from Artificial Intelligence?

A) All Data Science projects require AI techniques

B) Symbolic AI systems (e.g., Deep Blue) necessarily follow the full Data Science lifecycle

C) Data Science always depends on historical data, while AI may use rules without data

D) Data Science outputs are always autonomous systems

Answer: C

4. Among the 5 V’s of Big Data, which one is considered the most critical from a business perspective and why?

A) Volume – because storage cost is highest

B) Velocity – because real-time systems are hardest to build

C) Value – because data remains a cost and liability until actionable insights are extracted

D) Veracity – because noise is unavoidable

Answer: C

5. A bank wants to decide the optimal credit limit for each customer to maximise profit while controlling default risk. Which type of analytics is required?

A) Descriptive

B) Diagnostic

C) Predictive

D) Prescriptive

Answer: D

6. Which of the following is not a core responsibility typically owned by a Data Engineer?

A) Designing ETL/ELT pipelines

B) Building and tuning predictive models for business KPIs

C) Ensuring data reliability and SLA compliance

D) Managing data warehouse schemas and data lakes

Answer: B

7. The COMPAS recidivism algorithm controversy primarily highlights which challenge in Data Science?

A) Data velocity

B) Algorithmic bias and fairness

C) Lack of Big Data tools

D) Insufficient cloud computing resources

Answer: B

8. In Phase 6 (Model Evaluation), when is Recall preferred over Precision?

A) When false positives are very costly (e.g., spam filter)

B) When false negatives are very costly (e.g., cancer diagnosis)

C) When classes are perfectly balanced

D) When the goal is only to maximise overall accuracy

Answer: B

9. W.S. Cleveland’s 2001 paper is significant because it:

A) Coined the professional title “Data Scientist”

B) Proposed expanding statistics into a broader field called Data Science with six specific components

C) Introduced the CRISP-DM lifecycle

D) Released the first version of Apache Spark

Answer: B

10. Which combination correctly matches the technology with its primary strength?

A) Hadoop – in-memory iterative ML

B) Spark – high-throughput event streaming with replay capability

C) Kafka – distributed event streaming and message retention

D) MongoDB – columnar MPP analytics for petabyte SQL queries

Answer: C


Part B: Theory / Short Answer Questions

1. Explain why the Data Science lifecycle is described as iterative rather than linear. Give two concrete scenarios where an analyst would need to return to an earlier phase.

Model Answer:

The lifecycle is iterative because insights, model performance, or data quality issues discovered in later phases frequently invalidate assumptions made earlier.

Scenario 1: After modelling (Phase 5–6), high variance (overfitting) or poor metrics may require returning to Feature Engineering (Phase 5) or even Data Collection (Phase 2) for additional relevant features.

Scenario 2: During Monitoring (Phase 8), detection of data drift or concept drift forces the team to revisit Data Collection/Cleaning and retrain models.

2. Differentiate between Data Drift and Concept Drift with examples. How does each affect a deployed model?

Model Answer:

  • Data Drift: The statistical distribution of input features changes over time while the relationship to the target remains the same (e.g., user demographics shift after a marketing campaign). Model performance degrades because it sees unfamiliar feature values.
  • Concept Drift: The underlying relationship between features and target changes (e.g., COVID-19 suddenly alters demand patterns for products). Even correctly distributed inputs produce wrong predictions. Both require monitoring systems and retraining strategies.

3. Why is “Value” considered the most important of the 5 V’s? Illustrate with the Netflix recommendation example.

Model Answer:

Volume, Velocity, Variety and Veracity describe the nature of the data; Value determines whether the organisation gains competitive advantage or merely incurs storage, compute and compliance costs. Netflix’s recommendation engine (driven by collaborative filtering and deep learning) accounts for approximately 80 % of all content watched, directly translating massive viewing data into retention and revenue.

4. Compare the primary goals, outputs and human involvement of Data Science versus Artificial Intelligence in a tabular format (minimum four dimensions).

Model Answer: (Key rows)

  • Primary Goal: DS → extract insights for human decision-making; AI → build systems that autonomously exhibit intelligent behaviour.
  • Output: DS → dashboards, reports, predictive scores; AI → autonomous agents, generative models, robotic controllers.
  • Data Dependency: DS → always requires historical data; AI → may be rule-based (symbolic) or data-driven.
  • Human Involvement: DS → high (interpretation and final decisions); AI → can range from supervised to fully autonomous.

5. List the five core pillars of Data Science and briefly state why domain knowledge is indispensable even when technical skills are strong.

Model Answer:

Statistics & Mathematics, Programming & Computer Science, Machine Learning & AI, Data Engineering & Infrastructure, Domain Knowledge & Communication.

Domain knowledge is indispensable because it determines whether the right questions are asked, whether features are meaningful, whether results are actionable, and whether ethical/regulatory constraints are respected. Without it, technically perfect models often solve the wrong problem.


Part C: Long Answer / Essay-Type Questions (Tough)

1. Critically analyse the statement: “Not all Data Science is AI, and not all AI is Data Science.” Support your answer with historical and contemporary examples, and discuss the practical implications for a modern data team.

Model Answer Framework (expected depth):

  • Define the nested relationship (AI ⊃ ML ⊃ DL; DS intersects but is broader).
  • Historical counter-example for “AI without DS”: Deep Blue (1997) – hand-crafted evaluation functions and search trees, no training data or data pipeline.
  • Contemporary counter-example for “DS without AI”: A pure BI dashboard built in SQL + Power BI that tracks monthly revenue and churn rates using descriptive analytics only.
  • Overlap example: A churn-prediction neural network – the scientist follows the full DS lifecycle while simultaneously practising ML-based AI.
  • Practical implications: Teams must avoid forcing every problem into an ML solution; clear role boundaries (Analyst vs Scientist vs Engineer) prevent skill mismatch; governance and ethics frameworks differ for pure descriptive work versus high-risk autonomous systems (EU AI Act).
  • Conclusion: Understanding the distinction prevents both over-engineering and under-utilisation of available techniques.

2. A mid-sized e-commerce company wants to reduce customer churn by 15 % within nine months. Design an end-to-end Data Science project plan following the 8-phase lifecycle. For each phase, specify key activities, potential pitfalls, success metrics, and the primary role (Scientist / Analyst / Engineer) responsible. Finally, discuss how you would handle the ethical and regulatory challenges that are likely to arise.

Model Answer Framework:

  • Phase-by-phase breakdown with concrete activities (e.g., Phase 1: define churn precisely as “no purchase in 90 days”, set KPI = 15 % reduction, feasibility check on historical data).
  • Pitfalls: starting modelling before cleaning, ignoring class imbalance, deploying without monitoring, etc.
  • Role mapping: Engineer owns pipelines and feature store; Scientist owns modelling and evaluation; Analyst owns dashboards and business reporting.
  • Ethical/regulatory: demographic bias audit (equalised odds), data minimisation under DPDP/GDPR, right to explanation, model cards, human oversight for high-impact decisions (credit or discount offers).
  • Monitoring plan for drift and scheduled retraining.
  • Expected length: 800–1000 words equivalent in structured form.

3. Compare and contrast the three core data roles (Data Scientist, Data Analyst, Data Engineer) across at least eight dimensions. Then argue, with reference to real-world failure modes, why organisations that treat these roles as interchangeable frequently fail to realise value from data initiatives.

Model Answer Framework:

  • Detailed comparison table covering focus, analytics type, output, programming depth, maths depth, tools, questions answered, team interaction, entry requirements, and typical career progression.
  • Failure modes: – Asking a Data Analyst to build production-grade ML pipelines → technical debt and reliability issues. – Expecting a Data Engineer to interpret business KPIs and communicate to C-suite → poor stakeholder alignment. – Overloading Data Scientists with pure ETL work → high attrition and delayed insights.
  • Evidence from industry (Gartner 85–90 % project failure rate partly attributed to role confusion and unrealistic expectations).
  • Recommendation: clear RACI matrices, collaborative hand-offs, and specialised career ladders.

4. “Big Data technologies solved the Volume and Velocity problems, but Veracity and Value remain the hardest challenges.” Critically evaluate this claim with reference to Hadoop/Spark/Kafka ecosystems, real-world case studies (e.g., social-media bias, sensor noise, recommendation systems), and the organisational and ethical barriers that technology alone cannot address.

Model Answer Framework:

  • Acknowledge technological successes: HDFS/Spark for Volume, Kafka/Flink for Velocity.
  • Argue that Veracity requires domain expertise, quality frameworks, and human judgement (missingness mechanisms, label noise, adversarial data).
  • Value requires correct problem framing, organisational alignment, and ethical deployment – areas where technology is necessary but insufficient.
  • Case studies: Amazon hiring algorithm (bias amplification), social-media bot pollution, IoT sensor drift.
  • Organisational barriers: data silos, politics, unrealistic timelines.
  • Ethical/regulatory layer: GDPR fines, fairness metrics conflicts, accountability gaps.
  • Conclusion: Technology provides the substrate; people, process and governance determine whether Value is realised.

On this page