BTCE | 5th Sem
No-SQL SubjectUnit 1

No-SQL Unit 1: Complete Concepts Guide

Unit 1: NoSQL Fundamentals -> Generated and Prepared By Thiruselvan (ThiruXD)

0. Basic Starting Point – What is a Database?

A database is an organized way to store, manage, and retrieve data.

  • Traditional (Relational / SQL) databases: Store data in tables (like Excel sheets) with fixed columns and rows. Examples: MySQL, PostgreSQL, Oracle.
  • NoSQL databases: A family of newer databases designed for modern apps that need flexibility, speed, and the ability to handle huge amounts of data across many computers.

NoSQL means “Not Only SQL”. It does not mean “no SQL at all”. Many NoSQL databases still support SQL-like queries.

Why did NoSQL appear?

Modern apps (Instagram, Amazon, Netflix, IoT sensors, social media) create data that is:

  • Very large in volume
  • Changing shape frequently
  • Distributed across many servers worldwide
  • Needs very fast response

Traditional single-server relational databases struggle with this → NoSQL was created.


1.1 Why NoSQL?

1.1.1 Relational Databases vs NoSQL – The Big Comparison

Relational Databases are still excellent and widely used. They are not outdated.

When to use Relational:

  • Data is highly structured and stable
  • You need perfect correctness (money transfers, accounting)
  • Complex joins and reports are required

When to use NoSQL:

  • Data is large, semi-structured, or changes often
  • You need high speed and horizontal scaling (add more machines)
  • Data is distributed across many servers

Complete Comparison Table (memorize this):

DimensionRelational DatabasesNoSQL Databases
Data ModelTables (rows + columns)Documents, Key-Value, Wide-Column, or Graphs
SchemaFixed and predefinedFlexible (fields can be different in each record)
Query LanguageStandard SQLVaries (JSON queries, CQL, Cypher, APIs, etc.)
TransactionsStrong ACIDOften BASE (or tunable)
ScalabilityMostly vertical (bigger server)Horizontal (add more servers)
JoinsExcellent supportUsually avoided (data is denormalized)
ConsistencyStrong by defaultStrong / Eventual / Tunable
Best ForBanking, ERP, payroll, structured reportsCatalogues, sessions, social feeds, IoT, recommendations
ExamplesMySQL, PostgreSQL, Oracle, SQL ServerMongoDB, Redis, Cassandra, HBase, Neo4j, DynamoDB

Key Takeaway: Relational and NoSQL are complementary, not enemies. Many companies use both (this is called Polyglot Persistence).

When to Choose Which – Quick Guide:

RequirementBetter ChoiceReason
Bank money transferRelationalMust be 100% correct and atomic
Product catalogue (different attributes)DocumentFlexible structure
Login session / shopping cartKey-ValueSuper fast lookup by key
Friends / recommendationsGraphRelationships are the main thing
Billions of sensor readingsWide-ColumnHandles massive writes + time-series

1.1.2 The Impedance Mismatch Problem (Very Important)

Simple Definition:

The difficulty of converting data between the way programmers think (objects with nested data) and the way relational databases store data (flat tables).

Why it happens:

  • In programming (Java, Python, etc.) → You work with rich objects that contain other objects and lists.
  • In Relational DB → You must break everything into separate tables connected by foreign keys.

Example – E-commerce Order:

In Application (Object view):

Order
 ├── Customer (name, email)
 ├── Shipping Address (city, pincode)
 ├── Items (list of products)
 └── Payment Status

In Relational Database: You need 5 tables: Orders, Customers, Addresses, OrderItems, Payments

→ You must use JOINs or ORM tools to put the order back together.

Document Databases solve this by storing the entire order as one nested document (JSON):

{
  "order_id": "ORD1025",
  "customer": {
    "name": "Ananya Rao",
    "email": "ananya@example.com"
  },
  "shipping_address": {
    "city": "Bengaluru",
    "pincode": "560001"
  },
  "items": [
    {"product": "Laptop", "qty": 1, "price": 62000},
    {"product": "Mouse", "qty": 2, "price": 800}
  ],
  "payment_status": "Paid"
}

Mismatch Areas Table:

Mismatch AreaObject ViewRelational ViewDifficulty
IdentityObject has identity in memoryPrimary keyMapping needed
StructureNested objects + listsFlat tablesSplit data across tables
RelationshipsDirect referencesForeign keys + joinsMultiple joins required
InheritanceClasses inheritNot naturalMapping strategies needed
Schema EvolutionEasy to add fieldsNeeds migrationSlows development
Query StyleNavigate objectsSQL set-based queriesTwo different thinking styles

Key Point: Document databases reduce this mismatch a lot, but they are not always the best choice if you need complex joins and strict multi-record transactions.


1.2 Theoretical Foundations of Distributed Data

Most NoSQL databases run on many computers (distributed systems) for better speed, scale, and reliability. This creates new challenges.

1.2.1 The CAP Theorem (Must Master)

CAP Theorem: In a distributed system, when a network partition (communication failure) happens, you can guarantee at most two of the following three:

  • C – Consistency: All users see the same, latest data.
  • A – Availability: Every request gets a response (system never goes down).
  • P – Partition Tolerance: System continues working even if network is broken between servers.

Important: Real distributed systems almost always need P. So the real choice is between CP or AP.

Simple Real-World Examples:

1. CP System (Consistency + Partition Tolerance)

Example: Bank ATM

  • You have ₹10,000.
  • Network between two ATMs breaks.
  • Second ATM refuses the withdrawal → “Please try again later”.
  • Better to show error than give wrong balance. Examples of CP: Banking systems, HBase, MongoDB (with strong settings), ZooKeeper.

2. AP System (Availability + Partition Tolerance)

Example: Instagram Likes

  • You like a photo → Delhi server shows 101 likes.
  • Chennai server still shows 100 for a few seconds.
  • App never becomes unavailable. Examples of AP: Cassandra, DynamoDB (eventual), Riak, social media feeds.

3. CA System (Consistency + Availability)

Only possible when there is no network partition (single server or very tightly controlled local setup).

Example: Single-computer college library system.

Key Takeaway:

  • Use CP when wrong data is unacceptable (money, inventory).
  • Use AP when staying online is more important than perfect immediate consistency (likes, feeds, recommendations).

1.2.2 BASE Properties (Opposite of ACID)

ACID (used in relational DBs): Atomicity, Consistency, Isolation, Durability → Perfect correctness.

BASE (used in many NoSQL systems):

  • B – Basically Available: System tries to respond even if some servers are down.
  • A – Soft State: Data can change by itself in the background (because of replication).
  • S – Eventually Consistent: All copies will become the same after some time (not immediately).

Easy Real-Life Understanding of BASE:

  1. Basically Available

    Amazon shopping: One server crashes → You can still search and order. Website stays usable.

  2. Soft State

    WhatsApp profile picture: You change it. Friend in Delhi sees new picture immediately. Friend in Chennai still sees old one for a few minutes. The data changed by itself without new user action.

  3. Eventually Consistent

    Instagram likes: After a few seconds, everyone sees the same number of likes.

ACID vs BASE Table:

AspectACIDBASE
Main GoalPerfect correctnessHigh availability + scale
ConsistencyImmediate and strongEventual (or tunable)
Failure HandlingReject or roll backKeep serving (temporary inconsistency OK)
Best ForPayments, accountingSocial feeds, logs, recommendations

Exam Tip: Use ACID when correctness is critical. Use BASE when high availability and scale matter more than temporary inconsistency.


1.3 The Four Types of NoSQL Databases

NoSQL is not one thing. There are four main types. Choosing the wrong type creates problems.

1. Document-Oriented Databases

  • Store data as documents (usually JSON or BSON).
  • Documents can have nested fields and different structures.
  • Grouped in collections.

Best for: Product catalogues, user profiles, content management, orders.

Example (Student document):

{
  "student_id": "CS501",
  "name": "Ravi Kumar",
  "skills": ["SQL", "Python", "MongoDB"],
  "address": {"city": "Mysuru", "state": "Karnataka"},
  "projects": [...]
}

Strengths: Flexible schema, natural for objects, good for nested data.

Limitations: Complex joins are hard; data duplication possible.

Examples: MongoDB, CouchDB, Couchbase, Amazon DocumentDB.

2. Key-Value Stores

  • Simplest form: Unique Key → Value.
  • Database does not look inside the value (opaque).
  • Like a dictionary/hash map.

Best for: Sessions, shopping carts, caching, OTP, counters.

Examples:

KeyValueUse
session:9a71{user_id: 105, role: student}Login session
cart:U102[PEN10, BOOK22]Shopping cart
otp:779921473819OTP cache

Strengths: Extremely fast.

Limitations: No complex queries.

Examples: Redis, DynamoDB, Riak, Memcached.

3. Column-Oriented / Wide-Column Databases

  • Data stored by Row Key + Column Families.
  • Each row can have different columns (good for sparse data).
  • Designed for massive write speed.

Best for: IoT sensors, time-series data, logs, user activity.

Example:

Row KeyProfile FamilyActivity Family
user_101name=Kiran, city=Hubballilogin_2026=true, clicks=85
sensor_220location=PlantA10:00=31.2, 10:01=31.4

Strengths: Huge write scalability, sparse data.

Limitations: Modelling is query-driven (hard for beginners).

Examples: Cassandra, HBase, Bigtable, ScyllaDB.

4. Graph Databases

  • Store Nodes (entities), Edges (relationships), Properties.
  • Relationships are first-class citizens.

Best for: Social networks, recommendations, fraud detection, path finding.

Example Elements:

  • Node: Alice, Bob, Product
  • Edge: FRIEND_OF, PURCHASED
  • Property: age, since:2023

Strengths: Excellent for relationship traversal.

Limitations: Not ideal for simple bulk lookups or pure logs.

Examples: Neo4j, JanusGraph, Amazon Neptune, TigerGraph.

Four Types Summary Table:

TypeData ModelBest ForCommon Examples
DocumentNested JSON documentsCatalogues, profiles, CMSMongoDB, CouchDB
Key-ValueKey → ValueCache, sessions, cartsRedis, DynamoDB
Wide-ColumnRow Key + flexible columnsIoT, time-series, logsCassandra, HBase
GraphNodes + Edges + PropertiesRelationships, paths, fraudNeo4j, Neptun

On this page