BTCE | 5th Sem
Data Science SubjectUnit 1

DS Unit 1: Complete Concept Guide

Unit 1: Foundations of Data Science -> Generated and Prepared By Thiruselvan (ThiruXD)

1.1 Evolution & Significance of Data Science

Definition of Data Science

Data Science is an interdisciplinary field that combines statistics, mathematics, computer science, and domain expertise to extract meaningful knowledge, patterns, and actionable insights from structured and unstructured data using scientific methods, algorithms, and systems.

What makes it interdisciplinary?

  • Statistics & Probability – foundation of inference and uncertainty
  • Computer Science – algorithms, programming, data structures
  • Mathematics – linear algebra, calculus, optimisation
  • Domain Knowledge – healthcare, finance, engineering context
  • Communication – turning insights into business decisions
  • Ethics & Law – responsible data handling (GDPR, DPDP Act)

What does a Data Scientist do?

  • Collects, cleans, and prepares raw data
  • Explores data visually and statistically (EDA)
  • Builds and trains predictive models (ML/DL)
  • Deploys models into production systems
  • Monitors model performance (MLOps)
  • Communicates findings to technical and non-technical teams

Why is Data Science important now?

  • 2.5 quintillion bytes of data created every day
  • Affordable cloud computing
  • Advances in ML (especially deep learning)
  • Every industry generates data
  • Evidence-based decision-making replaces intuition
  • WEF: Data roles among top 10 fastest-growing jobs globally

Evolution of Data Science – Part 1 (1663–1989)

  • 1663: John Graunt – Bills of Mortality (first quantitative population analysis; foundation of statistics and epidemiology)
  • 1805–1830: Legendre & Gauss – Method of Least Squares (foundation of regression)
  • 1936: Alan Turing – Turing Machine (theoretical foundation of computation)
  • 1960s: Database systems emerge (hierarchical and network models; precursor to SQL)
  • 1974: Peter Naur coins “Data Science” (alternative to computer science focused on data processing)
  • 1977: John Tukey – Exploratory Data Analysis (EDA becomes core practice)
  • 1989: Knowledge Discovery in Databases (KDD) workshop (formal start of data mining era)

Evolution of Data Science – Part 2 (1996–2023+)

  • 1996: First KDD Conference (data mining as a discipline)
  • 2001: W.S. Cleveland – “An Action Plan” (proposes expanding statistics into Data Science with six components)
  • 2006: Apache Hadoop released (distributed storage and processing of massive data)
  • 2008: “Data Scientist” title coined by DJ Patil and Jeff Hammerbacher
  • 2012: Harvard Business Review – “Data Scientist: The Sexiest Job of the 21st Century”
  • 2014–2016: Rise of Deep Learning & Cloud (GPU neural networks + AWS/Azure/GCP)
  • 2017–2022: AutoML, MLOps, Explainable AI (LIME, SHAP)
  • 2023+: Large Language Models & Generative AI (GPT-4, Gemini, Llama integrated into DS workflows)

Significance of Data Science

Economic: Organisations using DS outperform competitors; Amazon recommendation ~35% of revenue; Netflix ~80% of content watched via recommendations; Uber real-time surge pricing; McKinsey – data-driven companies 23× more likely to acquire customers.

Scientific & Research: Genomics, CERN (~15 PB/year), climate modelling, AlphaFold (200M+ protein structures), COVID-19 modelling, exoplanet detection.

Social & Governance: Smart cities, public health surveillance, fraud detection in tax systems, education policy, census, disaster response.

Career & Industry Demand: WEF top emerging roles; LinkedIn top growing titles; US BLS 35% growth 2022–2032; high salaries; demand exceeds supply by ~50%.

1.2 Data Science Lifecycle

CRISP-DM inspired 8-phase iterative process. Insights at any phase may require revisiting earlier steps.

Phase 1: Business / Problem Understanding

Most critical phase. Define problem & KPIs with domain experts. Translate goals into measurable questions. Conduct feasibility analysis. Deliverable: problem statement and project charter. Pitfall: starting with data instead of the problem.

Phase 2: Data Collection & Acquisition

Sources: databases, APIs, IoT, scraping, surveys, logs. Structured (SQL, CSV) and unstructured (text, images, social media). Document lineage and provenance. Check quantity, quality, and time coverage.

Phase 3: Data Cleaning & Preprocessing

60–80% of project time often spent here. Handle missing values (deletion, imputation), outliers (Z-score, IQR, DBSCAN), duplicates, encoding (one-hot/label), normalisation/standardisation. Deliverable: clean, consistent dataset.

Phase 4: Exploratory Data Analysis (EDA)

Understand distributions, anomalies, relationships. Descriptive statistics, univariate (histograms, box plots), bivariate (scatter, correlation matrices). Detect class imbalance. Tools: Pandas, Matplotlib, Seaborn, R, Tableau. Deliverable: EDA report with key insights.

Phase 5: Feature Engineering & Modelling

Create new features (interaction, polynomial, binning, TF-IDF, embeddings). Dimensionality reduction (PCA, LDA). Algorithm selection by problem type. Bias-variance trade-off. Cross-validation. Deliverable: trained model pipeline with hyperparameters.

Phase 6: Model Evaluation & Validation

Use held-out test set only. Metrics by type (see table below). Hyperparameter tuning (Grid/Random Search, Bayesian). Deliverable: performance report aligned to business metrics.

Phase 7: Deployment & Communication

Batch vs real-time (REST APIs with Flask/FastAPI). Containerisation (Docker), orchestration (Kubernetes), model registry (MLflow). Dashboards (Tableau/Streamlit). Explain model logic to stakeholders.

Phase 8: Monitoring & Maintenance

Data drift and concept drift detection. Tools: Evidently AI, Arize, Fiddler. Retraining strategies (trigger-based or scheduled). MLOps CI/CD. A/B testing. Deliverable: monitoring dashboard and audit trail.

Key Evaluation Metrics

Problem TypeMetricDefinition / FormulaWhen to Use
RegressionMAEMean ofyᵢ − ŷᵢ
RegressionRMSE√(Mean of (yᵢ−ŷᵢ)²)Penalise large errors
RegressionR²1 − SS_res / SS_totVariance explained
ClassificationAccuracyCorrect / TotalBalanced classes
ClassificationPrecisionTP / (TP + FP)False positives costly
ClassificationRecallTP / (TP + FN)False negatives costly
ClassificationF1-Score2×(P×R)/(P+R)Balance Precision & Recall
ClassificationAUC-ROCArea under ROC curveCompare classifiers (threshold-independent)
ClusteringSilhouette Score−1 to +1 (higher better)Cluster cohesion & separation

1.3 Data Science vs Artificial Intelligence

Conceptual Relationship

  • AI (broadest): systems that simulate human intelligence (reasoning, perception, language, action).
  • ML (subset of AI): systems learn from data without explicit programming.
  • DL (subset of ML): multi-layer neural networks for complex patterns.
  • Data Science: intersects AI/ML/DL but also includes data engineering, statistics, domain knowledge, and communication.

All ML-based AI uses Data Science pipelines. Not all Data Science uses AI (e.g., pure BI dashboards). Not all AI uses Data Science (e.g., symbolic chess engines like Deep Blue).

Detailed Comparison

DimensionData ScienceArtificial Intelligence
Primary GoalExtract insights & knowledge to support decisionsBuild systems that autonomously perform intelligent tasks
Core FocusData collection, cleaning, EDA, modelling, communicationReasoning, learning, perception, planning, autonomous action
OutputDashboards, reports, predictive models, recommendationsAutonomous systems, agents, generative models, robots
Data DependencyAlways depends on historical dataMay use data (ML) or rules (symbolic AI)
Human InvolvementHigh (interpretation & final decisions)Varies (supervised to fully autonomous)
ScopeIncludes non-ML work (SQL, BI, surveys)Broader autonomy goals; may skip data pipelines
Key ToolsPython, R, SQL, Pandas, Scikit-learn, Tableau, Power BITensorFlow, PyTorch, OpenAI API, Hugging Face, Prolog, ROS
Real ExamplesChurn prediction, fraud scores, demand forecastsGPT-4, AlphaGo, Tesla Autopilot, AlphaFold, Alexa

1.4 Big Data, BI & Analytics

Big Data & the 5 V’s

Datasets so large/complex that traditional tools fail.

  • Volume: TB to EB scale (Facebook 100+ TB/day interaction data; NASA ~2.4 TB/day). Solution: HDFS, object storage.
  • Velocity: Speed of generation and processing (Twitter 500k+ tweets/min; NYSE ~1 TB/day). Spectrum: batch → micro-batch → streaming. Solution: Kafka, Flink, Spark Streaming.
  • Variety: Structured, semi-structured (JSON/XML), unstructured (~80% of data – text, images, video).
  • Veracity: Trustworthiness and accuracy (noise, bias, sensor errors).
  • Value: Most important – data is a liability until insights are extracted (Netflix recommendations drive ~80% of viewing).

Big Data Technologies

  • Apache Hadoop: HDFS (distributed storage, 128 MB blocks, 3× replication), MapReduce, YARN. Batch-oriented; slower for iterative ML.
  • Apache Spark: In-memory processing (up to 100× faster than MapReduce for iterative algorithms). PySpark, MLlib, Streaming, GraphX.
  • Apache Kafka: High-throughput event streaming (producers → topics → consumers). Millions of events/sec, low latency. Use cases: real-time fraud, IoT, log aggregation.
  • NoSQL: MongoDB (document), Cassandra (wide-column), Neo4j (graph), Redis (key-value). Flexible schema, horizontal scale.
  • Cloud Data Warehouses: BigQuery, Redshift, Snowflake. Data Lakehouse (Delta Lake, Iceberg). Tools: dbt, Airflow, Fivetran.

Business Intelligence vs Data Science

DimensionBusiness Intelligence (BI)Data Science
Primary QuestionWhat happened? What is happening?Why? What will happen? What should we do?
Time OrientationRetrospectiveForward-looking
Analytics TypeDescriptive + basic DiagnosticPredictive + Prescriptive
Primary UsersManagers, C-suite, operationsData scientists, ML engineers, researchers
Data TypesMostly structured, clean SQL dataStructured + semi + unstructured
MethodsOLAP, dashboards, KPIs, data warehousingML/DL, EDA, statistical testing, A/B tests
Key ToolsPower BI, Tableau, Looker, SAP BOPython, R, TensorFlow, Scikit-learn, Spark
OutputScorecards, dashboards, reportsPredictive models, automated decisions, insights
Skill EmphasisSQL, ETL, report designStatistics, ML theory, programming, MLOps
Time to ValueDays to weeksWeeks to months

Four Types of Data Analytics

  1. Descriptive – What happened? Aggregation, pivot tables, dashboards.
  2. Diagnostic – Why did it happen? Root cause, correlation, hypothesis testing.
  3. Predictive – What is likely to happen? ML models (regression, Random Forest, XGBoost, LSTM). Probabilistic.
  4. Prescriptive – What action should we take? Optimisation, simulation, reinforcement learning (e.g., Google Maps routing).

1.5 Components & Data Scientist Role

Core Components (Five Pillars)

  1. Statistics & Mathematics: Probability distributions, hypothesis testing, linear algebra (PCA, neural nets), calculus (gradient descent), Bayesian reasoning, optimisation.
  2. Programming & Computer Science: Python (NumPy, Pandas, Scikit-learn, TensorFlow), R, SQL, data structures, Git, complexity analysis, APIs.
  3. Machine Learning & AI: Supervised (classification/regression), unsupervised (clustering, PCA), deep learning (CNN, RNN/LSTM/Transformers), NLP, Computer Vision, Reinforcement Learning.
  4. Data Engineering & Infrastructure: ETL/ELT, SQL/NoSQL warehouses, Hadoop/Spark/Kafka/Airflow, cloud (AWS/GCP/Azure), data lakehouse.
  5. Domain Knowledge & Communication: Industry context, regulations (GDPR, CCPA, DPDP), ethics/fairness/XAI, storytelling, visualisation (Matplotlib, Tableau, Streamlit).

Role of a Data Scientist – End-to-End Responsibilities

  • Problem framing & scoping (KPIs, feasibility, assumptions)
  • Data acquisition & wrangling (SQL, APIs, cleaning, lineage)
  • Exploratory analysis & insight generation
  • Model development & tuning (feature engineering, CV, hyperparameter search)
  • Deployment & communication (APIs, dashboards, model cards, non-jargon reports)
  • Ethics, monitoring & maintenance (bias audits, drift detection, compliance, post-deployment review)

1.6 Analyst, Engineer — Applications & Challenges

Data Analyst

Focus: Interpret existing data to answer business questions and produce reports/visualisations.

Responsibilities: Complex SQL, dashboards, trend/cohort analysis, KPI tracking, data quality validation.

Tools: SQL, Excel, Tableau, Power BI, Looker; basic Python/R.

Background: Business Analytics, Statistics, Economics, CS + tool certifications.

Data Engineer

Focus: Design, build, and maintain reliable data infrastructure.

Responsibilities: ETL/ELT pipelines, warehouses/lakes, reliability/availability SLAs, monitoring.

Tools: Spark, Kafka, Airflow, dbt, Snowflake/BigQuery/Redshift, Python/Scala/Java, cloud platforms.

Background: CS or Software Engineering + distributed systems knowledge.

Comparison of Three Core Roles

DimensionData ScientistData AnalystData Engineer
Core FocusPredictive/prescriptive models & insightsInterpret data; answer “what happened”Infrastructure & pipelines
Analytics TypePredictive & PrescriptiveDescriptive & DiagnosticEnabler of all analytics
Primary OutputML models, statistical reports, NLPDashboards, trend reports, KPIsETL pipelines, warehouses, streaming
Programming DepthAdvanced Python/R + ML frameworksBasic–intermediate SQL/Python + BIAdvanced Python/Scala/Java + systems
Maths/Stats DepthHighModerateLow (systems-oriented)
Key ToolsScikit-learn, TF/PyTorch, MLflow, DockerSQL, Excel, Tableau, Power BISpark, Kafka, Airflow, dbt, Kubernetes
Questions AnsweredWhat will happen? Why? Optimise what?What happened / is happening?How to collect, store, move data reliably
Team InteractionCross-functionalBusiness-facingEngineering-facing
Entry RequirementBSc/MSc CS/Stats/Maths/DSBSc Business/Stats/Econ/CS + toolsBSc CS/SE + distributed systems

Applications of Data Science

Healthcare: Disease prediction, medical imaging (CNNs), drug discovery (AlphaFold), personalised medicine, hospital operations, pandemic modelling, EHR analytics.

Finance: Real-time fraud detection, credit scoring, algorithmic trading, VaR/risk management, AML (graph ML), churn prediction, NLP chatbots.

Retail & E-Commerce: Recommendation engines (~35% Amazon revenue), demand forecasting, dynamic pricing, sentiment analysis, customer segmentation (RFM), visual search, supply-chain optimisation.

Transportation & Logistics: Route optimisation (UPS ORION), ride-hailing matching/surge, autonomous vehicles, predictive maintenance, fleet management, air-traffic prediction, last-mile delivery.

Social Media & Tech: Content recommendation (YouTube 70%+), fake-news detection, real-time ad bidding, behaviour analysis, hate-speech moderation, trend detection.

Agriculture: Precision farming (satellite/drone + ML), yield prediction, soil analysis, pest/disease detection, irrigation optimisation, hyper-local weather.

Education & EdTech: Adaptive learning, dropout prediction, intelligent tutoring, learning analytics, course recommendation, automated essay scoring.

Cybersecurity: Intrusion detection, malware classification, UEBA, phishing detection, threat intelligence (graph ML), vulnerability prioritisation.

Climate & Environment: Climate modelling, renewable energy forecasting, energy efficiency (DeepMind 40% cooling reduction), deforestation monitoring, ocean monitoring, disaster early-warning systems.

Challenges in Data Science

Technical

  • Data quality & completeness (missingness types MCAR/MAR/MNAR, label noise, sampling bias, inconsistency)
  • Data silos & integration (schema conflicts, organisational politics, legacy systems) → data mesh/fabric
  • Unrealistic expectations (85–90% of AI/ML projects fail to reach production – Gartner)

Organisational, Talent & Ethical

  • Algorithmic bias & fairness (COMPAS, Amazon hiring, facial recognition disparities). Metrics: demographic parity, equalised odds, etc. No single universal fairness definition.
  • Data privacy & regulation (GDPR, CCPA, India DPDP Act 2023). Privacy-preserving techniques (federated learning). Heavy fines possible.
  • Consent, transparency & accountability (opacity of automated decisions, EU AI Act risk classification, model cards, ethics boards).

Opportunities

  • Healthcare innovation (FDA-approved AI devices, AlphaFold, genomic medicine)
  • Climate action (energy efficiency, methane/deforestation monitoring, smart grids)
  • AutoML & democratisation (Google AutoML, SageMaker, ChatGPT Code Interpreter, pre-trained models)
  • Emerging technologies (Federated Learning, Edge AI, Quantum ML)
  • Career & economic growth (strong demand and salary progression)
  • Interdisciplinary research (computational biology, smart cities, brain-computer interfaces)

This completes the full coverage of Unit 1. All historical milestones, lifecycle phases, comparisons, 5 V’s, technologies, analytics types, pillars, roles, applications across domains, challenges, and opportunities are included exactly as presented in the source material. Use this as your complete concept reference.

On this page