No-SQL Unit 1: Complete Concepts Guide
Unit 1: NoSQL Fundamentals -> Generated and Prepared By Thiruselvan (ThiruXD)
0. Basic Starting Point – What is a Database?
A database is an organized way to store, manage, and retrieve data.
- Traditional (Relational / SQL) databases: Store data in tables (like Excel sheets) with fixed columns and rows. Examples: MySQL, PostgreSQL, Oracle.
- NoSQL databases: A family of newer databases designed for modern apps that need flexibility, speed, and the ability to handle huge amounts of data across many computers.
NoSQL means “Not Only SQL”. It does not mean “no SQL at all”. Many NoSQL databases still support SQL-like queries.
Why did NoSQL appear?
Modern apps (Instagram, Amazon, Netflix, IoT sensors, social media) create data that is:
- Very large in volume
- Changing shape frequently
- Distributed across many servers worldwide
- Needs very fast response
Traditional single-server relational databases struggle with this → NoSQL was created.
1.1 Why NoSQL?
1.1.1 Relational Databases vs NoSQL – The Big Comparison
Relational Databases are still excellent and widely used. They are not outdated.
When to use Relational:
- Data is highly structured and stable
- You need perfect correctness (money transfers, accounting)
- Complex joins and reports are required
When to use NoSQL:
- Data is large, semi-structured, or changes often
- You need high speed and horizontal scaling (add more machines)
- Data is distributed across many servers
Complete Comparison Table (memorize this):
| Dimension | Relational Databases | NoSQL Databases |
|---|---|---|
| Data Model | Tables (rows + columns) | Documents, Key-Value, Wide-Column, or Graphs |
| Schema | Fixed and predefined | Flexible (fields can be different in each record) |
| Query Language | Standard SQL | Varies (JSON queries, CQL, Cypher, APIs, etc.) |
| Transactions | Strong ACID | Often BASE (or tunable) |
| Scalability | Mostly vertical (bigger server) | Horizontal (add more servers) |
| Joins | Excellent support | Usually avoided (data is denormalized) |
| Consistency | Strong by default | Strong / Eventual / Tunable |
| Best For | Banking, ERP, payroll, structured reports | Catalogues, sessions, social feeds, IoT, recommendations |
| Examples | MySQL, PostgreSQL, Oracle, SQL Server | MongoDB, Redis, Cassandra, HBase, Neo4j, DynamoDB |
Key Takeaway: Relational and NoSQL are complementary, not enemies. Many companies use both (this is called Polyglot Persistence).
When to Choose Which – Quick Guide:
| Requirement | Better Choice | Reason |
|---|---|---|
| Bank money transfer | Relational | Must be 100% correct and atomic |
| Product catalogue (different attributes) | Document | Flexible structure |
| Login session / shopping cart | Key-Value | Super fast lookup by key |
| Friends / recommendations | Graph | Relationships are the main thing |
| Billions of sensor readings | Wide-Column | Handles massive writes + time-series |
1.1.2 The Impedance Mismatch Problem (Very Important)
Simple Definition:
The difficulty of converting data between the way programmers think (objects with nested data) and the way relational databases store data (flat tables).
Why it happens:
- In programming (Java, Python, etc.) → You work with rich objects that contain other objects and lists.
- In Relational DB → You must break everything into separate tables connected by foreign keys.
Example – E-commerce Order:
In Application (Object view):
Order
├── Customer (name, email)
├── Shipping Address (city, pincode)
├── Items (list of products)
└── Payment StatusIn Relational Database: You need 5 tables: Orders, Customers, Addresses, OrderItems, Payments
→ You must use JOINs or ORM tools to put the order back together.
Document Databases solve this by storing the entire order as one nested document (JSON):
{
"order_id": "ORD1025",
"customer": {
"name": "Ananya Rao",
"email": "ananya@example.com"
},
"shipping_address": {
"city": "Bengaluru",
"pincode": "560001"
},
"items": [
{"product": "Laptop", "qty": 1, "price": 62000},
{"product": "Mouse", "qty": 2, "price": 800}
],
"payment_status": "Paid"
}Mismatch Areas Table:
| Mismatch Area | Object View | Relational View | Difficulty |
|---|---|---|---|
| Identity | Object has identity in memory | Primary key | Mapping needed |
| Structure | Nested objects + lists | Flat tables | Split data across tables |
| Relationships | Direct references | Foreign keys + joins | Multiple joins required |
| Inheritance | Classes inherit | Not natural | Mapping strategies needed |
| Schema Evolution | Easy to add fields | Needs migration | Slows development |
| Query Style | Navigate objects | SQL set-based queries | Two different thinking styles |
Key Point: Document databases reduce this mismatch a lot, but they are not always the best choice if you need complex joins and strict multi-record transactions.
1.2 Theoretical Foundations of Distributed Data
Most NoSQL databases run on many computers (distributed systems) for better speed, scale, and reliability. This creates new challenges.
1.2.1 The CAP Theorem (Must Master)
CAP Theorem: In a distributed system, when a network partition (communication failure) happens, you can guarantee at most two of the following three:
- C – Consistency: All users see the same, latest data.
- A – Availability: Every request gets a response (system never goes down).
- P – Partition Tolerance: System continues working even if network is broken between servers.
Important: Real distributed systems almost always need P. So the real choice is between CP or AP.
Simple Real-World Examples:
1. CP System (Consistency + Partition Tolerance)
Example: Bank ATM
- You have ₹10,000.
- Network between two ATMs breaks.
- Second ATM refuses the withdrawal → “Please try again later”.
- Better to show error than give wrong balance. Examples of CP: Banking systems, HBase, MongoDB (with strong settings), ZooKeeper.
2. AP System (Availability + Partition Tolerance)
Example: Instagram Likes
- You like a photo → Delhi server shows 101 likes.
- Chennai server still shows 100 for a few seconds.
- App never becomes unavailable. Examples of AP: Cassandra, DynamoDB (eventual), Riak, social media feeds.
3. CA System (Consistency + Availability)
Only possible when there is no network partition (single server or very tightly controlled local setup).
Example: Single-computer college library system.
Key Takeaway:
- Use CP when wrong data is unacceptable (money, inventory).
- Use AP when staying online is more important than perfect immediate consistency (likes, feeds, recommendations).
1.2.2 BASE Properties (Opposite of ACID)
ACID (used in relational DBs): Atomicity, Consistency, Isolation, Durability → Perfect correctness.
BASE (used in many NoSQL systems):
- B – Basically Available: System tries to respond even if some servers are down.
- A – Soft State: Data can change by itself in the background (because of replication).
- S – Eventually Consistent: All copies will become the same after some time (not immediately).
Easy Real-Life Understanding of BASE:
-
Basically Available
Amazon shopping: One server crashes → You can still search and order. Website stays usable.
-
Soft State
WhatsApp profile picture: You change it. Friend in Delhi sees new picture immediately. Friend in Chennai still sees old one for a few minutes. The data changed by itself without new user action.
-
Eventually Consistent
Instagram likes: After a few seconds, everyone sees the same number of likes.
ACID vs BASE Table:
| Aspect | ACID | BASE |
|---|---|---|
| Main Goal | Perfect correctness | High availability + scale |
| Consistency | Immediate and strong | Eventual (or tunable) |
| Failure Handling | Reject or roll back | Keep serving (temporary inconsistency OK) |
| Best For | Payments, accounting | Social feeds, logs, recommendations |
Exam Tip: Use ACID when correctness is critical. Use BASE when high availability and scale matter more than temporary inconsistency.
1.3 The Four Types of NoSQL Databases
NoSQL is not one thing. There are four main types. Choosing the wrong type creates problems.
1. Document-Oriented Databases
- Store data as documents (usually JSON or BSON).
- Documents can have nested fields and different structures.
- Grouped in collections.
Best for: Product catalogues, user profiles, content management, orders.
Example (Student document):
{
"student_id": "CS501",
"name": "Ravi Kumar",
"skills": ["SQL", "Python", "MongoDB"],
"address": {"city": "Mysuru", "state": "Karnataka"},
"projects": [...]
}Strengths: Flexible schema, natural for objects, good for nested data.
Limitations: Complex joins are hard; data duplication possible.
Examples: MongoDB, CouchDB, Couchbase, Amazon DocumentDB.
2. Key-Value Stores
- Simplest form: Unique Key → Value.
- Database does not look inside the value (opaque).
- Like a dictionary/hash map.
Best for: Sessions, shopping carts, caching, OTP, counters.
Examples:
| Key | Value | Use |
|---|---|---|
| session:9a71 | {user_id: 105, role: student} | Login session |
| cart:U102 | [PEN10, BOOK22] | Shopping cart |
| otp:779921 | 473819 | OTP cache |
Strengths: Extremely fast.
Limitations: No complex queries.
Examples: Redis, DynamoDB, Riak, Memcached.
3. Column-Oriented / Wide-Column Databases
- Data stored by Row Key + Column Families.
- Each row can have different columns (good for sparse data).
- Designed for massive write speed.
Best for: IoT sensors, time-series data, logs, user activity.
Example:
| Row Key | Profile Family | Activity Family |
|---|---|---|
| user_101 | name=Kiran, city=Hubballi | login_2026=true, clicks=85 |
| sensor_220 | location=PlantA | 10:00=31.2, 10:01=31.4 |
Strengths: Huge write scalability, sparse data.
Limitations: Modelling is query-driven (hard for beginners).
Examples: Cassandra, HBase, Bigtable, ScyllaDB.
4. Graph Databases
- Store Nodes (entities), Edges (relationships), Properties.
- Relationships are first-class citizens.
Best for: Social networks, recommendations, fraud detection, path finding.
Example Elements:
- Node: Alice, Bob, Product
- Edge: FRIEND_OF, PURCHASED
- Property: age, since:2023
Strengths: Excellent for relationship traversal.
Limitations: Not ideal for simple bulk lookups or pure logs.
Examples: Neo4j, JanusGraph, Amazon Neptune, TigerGraph.
Four Types Summary Table:
| Type | Data Model | Best For | Common Examples |
|---|---|---|---|
| Document | Nested JSON documents | Catalogues, profiles, CMS | MongoDB, CouchDB |
| Key-Value | Key → Value | Cache, sessions, carts | Redis, DynamoDB |
| Wide-Column | Row Key + flexible columns | IoT, time-series, logs | Cassandra, HBase |
| Graph | Nodes + Edges + Properties | Relationships, paths, fraud | Neo4j, Neptun |