No-SQL Unit 3: Complete Concepts Guide
Unit 3: Redis, Column-Family Stores & Cassandra -> Generated and Prepared By Thiruselvan (ThiruXD)
1. Introduction to Key-Value Stores
1.1 What is a Key-Value Store?
A key-value store is a type of NoSQL database that stores data as simple key-value pairs. Each key is unique and is used to retrieve the associated value. It is highly scalable, fast, and works well for simple lookups.
1.2 How It Works
| Key | Value |
|---|---|
user:101 | { "name": "Ananya", "age": 25 } |
product:200 | { "name": "Laptop", "price": 55000 } |
session:abc | "token_9382xyz" |
cart:500 | [ "item1", "item2", "item3" ] |
1.3 Real-World Examples
1. E-commerce — Shopping Cart Key-value stores are often used to store shopping cart data for each user. The user ID can be the key, and the cart items can be the value.
Key: cart:user123
Value: ["item1", "item2", "item3"]2. Caching — User Sessions Key-value stores are widely used for caching user session data in web applications. The session ID is the key and the session information is the value.
Key: session:abc123
Value: { "user_id": 101, "status": "active" }1.4 Advantages of Key-Value Stores
| # | Advantage | Explanation |
|---|---|---|
| 1 | Simple Model | Easy to understand and use |
| 2 | High Performance | Fast reads and writes |
| 3 | Scalable | Horizontally scalable |
| 4 | Flexible | Values can be any data type |
| 5 | Simple Lookups | O(1) access by key |
2. Redis — Introduction
2.1 What is Redis?
Redis (Remote Dictionary Server) is an open-source, in-memory key-value store. It is widely used for caching, session management, real-time analytics, and message queuing. Redis supports multiple data structures beyond simple strings.
2.2 Why Redis?
- In-memory: Extremely fast (microsecond latency)
- Data structures: Strings, Lists, Sets, Hashes, Sorted Sets, Streams
- Persistence: RDB snapshots and AOF logs
- Replication: Master-replica replication
- Pub/Sub: Messaging patterns
- Transactions: MULTI/EXEC blocks
- TTL: Automatic key expiration
2.3 Redis Data Structures Overview
| Data Structure | Description | Use Case |
|---|---|---|
| String | Basic key-value | Caching, counters |
| List | Ordered collection | Queues, logs |
| Set | Unordered unique | Tags, followers |
| Hash | Field-value pairs | User profiles |
| Sorted Set | Unique + score | Leaderboards |
| Stream | Append-only log | Event sourcing |
3. Redis Data Structures
3.1 String
A simple key-value pair where the value is a string (text, number, or binary data).
Redis Commands:
SET user:101 "Ananya"
GET user:101 # Returns "Ananya"
INCR counter # Increase by 1
DECR counter # Decrease by 1Example Use Case: Store user names, configuration settings, counters.
3.2 List
An ordered collection of string elements that allows insertion and removal from both ends (like a queue or stack).
Structure:
Left ← [A] [B] [C] [D] → RightRedis Commands:
LPUSH notifications "msg1" # Add to left
LPUSH notifications "msg2"
RPUSH notifications "msg3" # Add to right
LRANGE notifications 0 -1 # Get all
# Returns ["msg2", "msg1", "msg3"]
LPOP notifications # Remove from left
RPOP notifications # Remove from rightExample Use Case: Maintain recent activity logs, message queues, notification lists.
3.3 Set
An unordered collection of unique string elements (no duplicate values).
Structure:
A
B CRedis Commands:
SADD tags "redis"
SADD tags "nosql"
SADD tags "database"
SMEMBERS tags # Returns all
# ["redis", "nosql", "database"]
SISMEMBER tags "redis" # Check membership
SINTER tags otherSet # IntersectionExample Use Case: Store unique items such as tags, user IDs, or followers.
3.4 Hash
A collection of field-value pairs (like a dictionary or map) suitable for storing objects.
Structure:
Key: user:101
┌──────────┬──────────────────┐
│ Field │ Value │
├──────────┼──────────────────┤
│ name │ Ananya │
│ age │ 25 │
│ email │ ananya@gcu.edu │
│ role │ faculty │
└──────────┴──────────────────┘Redis Commands:
HSET user:101 name "Ananya"
HSET user:101 age 25
HGET user:101 name # Returns "Ananya"
HGETALL user:101 # Get all fields
HINCRBY user:101 age 1 # Increment numeric fieldExample Use Case: Store user profiles, product details, or application settings.
3.5 Sorted Set (ZSet)
A collection of unique elements, each with an associated score. Elements are automatically ordered by their score.
Structure:
┌─────────┬───────┐
│ Member │ Score │
├─────────┼───────┤
│ Charlie │ 60 │
│ Bob │ 80 │
│ Alice │ 100 │
└─────────┴───────┘Redis Commands:
ZADD leaderboard 100 "Alice"
ZADD leaderboard 80 "Bob"
ZADD leaderboard 60 "Charlie"
ZRANGE leaderboard 0 -1 WITHSCORES
# Returns ["Charlie", 60, "Bob", 80, "Alice", 100]
ZREVRANGE leaderboard 0 -1 # Descending order
ZSCORE leaderboard "Alice" # Get scoreExample Use Case: Leaderboards, rankings, priority queues, or scheduling.
3.6 Stream
An append-only log-like structure for events.
Redis Commands:
XADD mystream * event "login" user "Rahul"
XREAD COUNT 10 STREAMS mystream 0
XRANGE mystream - +Example Use Case: Event sourcing, activity feeds, real-time analytics.
4. Redis Key Management
4.1 Create a Key (SET)
SET username "Ananya"
# 127.0.0.1:6379> SET username "Ananya"
# OKIf the key already exists, SET will update its value.
4.2 Retrieve a Value (GET)
GET username
# 127.0.0.1:6379> GET username
# "Ananya"Returns the value stored at the key. If the key does not exist, returns (nil).
4.3 Set Expiry (EXPIRE)
EXPIRE username 60
# 127.0.0.1:6379> EXPIRE username 60
# (integer) 1The key will be automatically deleted after the specified time (in this case, 60 seconds).
4.4 Check TTL (TTL)
TTL username
# 127.0.0.1:6379> TTL username
# (integer) 42Returns the remaining time in seconds.
1: key exists but has no expiry2: key does not exist
4.5 Common Key Commands
| Command | Purpose |
|---|---|
SET key value | Create/update key |
GET key | Retrieve value |
DEL key | Delete key |
EXISTS key | Check existence |
EXPIRE key seconds | Set TTL |
TTL key | Check remaining TTL |
PERSIST key | Remove TTL |
KEYS pattern | Find keys by pattern |
RENAME old new | Rename key |
5. Redis Transactions
5.1 What is a Redis Transaction?
A Redis transaction groups multiple commands into a single atomic unit. Commands are queued and executed together.
5.2 MULTI — Start a Transaction
MULTI puts Redis into transaction mode. All subsequent commands are queued instead of being executed immediately. Redis replies with “QUEUED” for each command.
127.0.0.1:6379> MULTI
OK
127.0.0.1:6379(TX)> SET user:1 "Ananya"
QUEUED
127.0.0.1:6379(TX)> INCR login_count
QUEUED
127.0.0.1:6379(TX)> EXPIRE user:1 3600
QUEUED5.3 EXEC — Execute All Queued Commands
EXEC executes all the queued commands in the order they were added. All commands are executed as a single atomic unit. Returns an array with the result of each command.
127.0.0.1:6379(TX)> EXEC
1) OK
2) (integer) 1
3) (integer) 1Important: If EXEC is called, all queued commands are executed, even if one of them fails (runtime errors do not rollback).
5.4 WATCH — Optimistic Locking
WATCH monitors one or more keys for changes by other clients. If any watched key is modified before EXEC, the transaction is aborted (EXEC returns nil). Useful for implementing optimistic locking.
127.0.0.1:6379> WATCH balance:100
OK
127.0.0.1:6379> MULTI
OK
127.0.0.1:6379(TX)> DECRBY balance:100 50
QUEUED
127.0.0.1:6379(TX)> EXEC
(nil) # if balance:100 was modified by another client5.5 DISCARD — Cancel Transaction
127.0.0.1:6379(TX)> DISCARD
OK5.6 Transaction Use Case
Transfer money between two accounts (ensure both debit and credit happen together).
MULTI
DECRBY account:A 1000
INCRBY account:B 1000
EXEC5.7 Key Takeaways
| Command | Purpose |
|---|---|
| MULTI | Starts the transaction |
| EXEC | Executes all queued commands |
| DISCARD | Cancels the transaction |
| WATCH | Detects changes; prevents lost updates |
| UNWATCH | Stops watching keys |
Important: Redis transactions provide atomicity, but do not rollback on runtime errors. They are simple, fast, and useful for ensuring data consistency.
6. Caching Strategies — Cache-Aside
6.1 What is Cache-Aside?
The Cache-Aside strategy (Lazy Loading) means the application is responsible for loading data into the cache. Data is fetched from the cache first; if not found, it is loaded from the database and then stored in the cache for future requests.
6.2 How It Works
1. Application checks the cache for the data.
2. If found (cache hit), return directly from cache.
3. If not found (cache miss), fetch from database.
4. Store the data in cache (with TTL) and return to client.6.3 Read Flow (Cache Miss)
Client → Application → Redis Cache → (not found) → Database
↓
Store in cache (with TTL)
↓
Return data to clientAfter the first request, the data is cached, so subsequent requests will be faster.
6.4 Read Flow (Cache Hit)
Client → Application → Redis Cache → (found) → Return data (fast response)6.5 Write Flow (Cache-Aside)
Application → Update database → Invalidate/delete cache entryOn write, the application updates the database first, then invalidates (or deletes) the cache entry so the next read will load fresh data.
6.6 Example: E-commerce Product Details
When a user views a product (e.g., product ID 101), the application first checks Redis. If not found, it fetches the product details from the database, stores it in Redis (e.g., for 5 minutes), and returns it to the user.
Product ID: 101
Name: Running Shoes
Price: ₹2,999
Stock: 50
Cached for 5 minutes6.7 Key Benefits
| # | Benefit | Explanation |
|---|---|---|
| 1 | Reduces database load | Fewer DB queries |
| 2 | Improves application performance | Fast cache reads |
| 3 | Simple to implement | Application controls caching |
| 4 | Works well for read-heavy applications | Most reads hit cache |
| 5 | Gives control to the application | TTL and invalidation strategy |
6.8 Things to Consider
| # | Consideration | Explanation |
|---|---|---|
| 1 | First request is slower (cache miss) | Must fetch from DB |
| 2 | Need to handle cache invalidation | On data updates |
| 3 | Possibility of stale data | Until invalidated |
| 4 | Choose an appropriate TTL | Balance freshness vs performance |
7. Column-Family Store (Wide-Column Store)
7.1 What is a Column-Family Store?
A column-family store is a type of NoSQL database that stores data in column families (groups of columns) rather than fixed tables with rows and columns. It is optimized for large-scale, distributed storage and high write/read throughput.
7.2 Key Characteristics
- Stores data in column families, which are collections of columns
- Each row can have a different set of columns (sparse data)
- Designed to handle very large datasets across multiple machines
- Optimized for fast reads and writes of large amounts of data
- Examples: Apache Cassandra, HBase, Amazon DynamoDB (wide-column model)
7.3 Example: Column-Family Store
Table: users (Column Family)
| Row Key (user_id) | name | age | city | |
|---|---|---|---|---|
| u1 | Ananya | a@gcu.edu | 25 | Bangalore |
| u2 | Rahul | r@mail.com | — | Delhi |
| u3 | Meera | — | 30 | Chennai |
Each row can have a different set of columns. Missing values are not stored (sparse storage).
7.4 Comparison with Relational Database
Relational Table: users
| user_id | name | age | city | |
|---|---|---|---|---|
| u1 | Ananya | a@gcu.edu | 25 | Bangalore |
| u2 | Rahul | r@mail.com | NULL | Delhi |
| u3 | Meera | NULL | 30 | Chennai |
All rows follow the same fixed schema (same columns). Missing values are stored as NULL.
7.5 Key Differences
| Feature | Column-Family Store (Wide-Column) | Relational Database (RDBMS) |
|---|---|---|
| Data model | Column families with flexible columns (sparse) | Fixed tables with rows and columns |
| Schema | Dynamic (each row can have different columns) | Fixed schema (all rows have same columns) |
| Storage format | Stores data by column families, not rows | Stores data by rows in tables |
| Best for | Large-scale, distributed, high write/read workloads | Transactional applications, complex queries, OLTP |
| Scalability | Horizontally scalable (designed for clusters) | Typically vertically scalable (scale-up) |
| Query language | Uses its own query language (e.g., CQL for Cassandra) | Uses SQL |
| Joins | Does not support joins (denormalized data model) | Supports joins, normalization |
| Use cases | Time-series data, IoT, user profiles, logs, analytics | Banking, ERP, inventory, traditional business applications |
7.6 When to Use a Column-Family Store?
Use a Column-Family Store when:
- You need to store and manage very large amounts of data
- Your data is sparse or has changing attributes
- You need high write and read throughput and horizontal scalability
- Examples: Cassandra, HBase, DynamoDB
7.7 When to Use a Relational Database?
Use a Relational Database when:
- You need complex queries and joins
- Your data is highly structured and consistent
- You need ACID transactions
- Examples: MySQL, PostgreSQL, Oracle, SQL Server
Key Takeaway: Column-Family Stores offer flexibility and scale, while Relational Databases provide structure and strong transactional support.
8. Apache Cassandra — Core Concepts
8.1 Node
A node is a single instance of Cassandra running on a machine (physical or virtual). Each node stores a portion of the data (data partitions). Nodes are equal — there is no master or slave.
┌─────────────────────┐
│ Node │
│ (Cassandra instance)│
└─────────────────────┘8.2 Cluster
A cluster is a group of nodes that work together to store and manage data. All nodes in a cluster are peer-to-peer (equal). Provides high availability, fault tolerance, and horizontal scalability.
Cassandra Cluster
┌──────┐ ┌──────┐ ┌──────┐ ┌──────┐
│Node 1│ │Node 2│ │Node 3│ ... │Node N│
└──────┘ └──────┘ └──────┘ └──────┘8.3 Keyspace
A keyspace is a top-level container (similar to a database in relational systems). It defines replication settings and contains one or more tables (column families). You must create a keyspace before creating tables.
Keyspace: ecom
┌──────────┐ ┌──────────┐ ┌──────────┐
│ Table │ │ Table │ │ Table │
│ users │ │ orders │ │ products │
└──────────┘ └──────────┘ └──────────┘8.4 Partition
A partition is a subset of the total data, identified by a partition key. Rows with the same partition key are stored together in the same partition. Cassandra distributes (hashes) partitions across nodes in the cluster.
Partition Key (e.g., user_id = 101)
↓
┌─────────────────┐
│ Partition │
│ Row 1 │
│ Row 2 │
│ Row 3 │
│ ... │
└─────────────────┘8.5 Replication Factor (RF)
The replication factor defines how many copies of each partition are stored across different nodes in the cluster. Helps with fault tolerance and high availability.
Example: If RF = 3, each partition is stored on 3 different nodes (ideally in different racks/data centers). You can set the replication factor at the keyspace level.
Node 1 (Replica)
↑
┌──────┴──────┐
Node 4 │ RF = 3 │ Node 2 (Replica)
(Other │3 copies of│
data) │each part. │
└──────┬──────┘
↓
Node 3 (Replica)8.6 Cassandra Data Model Hierarchy
Cluster (Collection of nodes)
↓
Keyspace (Like a database)
↓
Table (Stores rows of data)9. CQL — Cassandra Query Language
9.1 What is CQL?
CQL (Cassandra Query Language) is the query language used to interact with Apache Cassandra. It is similar to SQL, but designed for distributed, scalable, NoSQL data storage. CQL allows us to define the data structure (keyspaces, tables) and perform operations like insert, select, update, and delete.
9.2 CQL vs SQL (Quick Comparison)
| CQL (Cassandra) | SQL (Relational DB) |
|---|---|
| Keyspace | Database |
| Table | Table |
| PRIMARY KEY | Primary Key |
| Designed for distributed systems | Designed for single server |
| No JOINs | Supports JOINs |
| Eventual consistency | Strong consistency |
9.3 Basic CQL Commands
a) Create a Keyspace
CREATE KEYSPACE university
WITH replication = {
'class': 'SimpleStrategy',
'replication_factor': 3
};- Keyspace name:
university - Replication strategy:
SimpleStrategy - Replication factor: 3 (number of copies stored across nodes)
b) Use a Keyspace
USE university;c) Create a Table
CREATE TABLE students (
student_id int,
name text,
department text,
age int,
PRIMARY KEY (student_id)
);- Table name:
students student_idis the primary key (which uniquely identifies each row)
d) Insert Data
INSERT INTO students (student_id, name, department, age)
VALUES (1, 'Ananya', 'CSE', 20);e) Query Data
SELECT * FROM students;
SELECT * FROM students WHERE student_id = 1;f) Update Data
UPDATE students SET age = 21 WHERE student_id = 1;g) Delete Data
DELETE FROM students WHERE student_id = 1;9.4 Verify the Creation
-- View keyspaces
DESCRIBE KEYSPACES;
-- View tables in the current keyspace
DESCRIBE TABLES;
-- View table structure
DESCRIBE TABLE students;Tip: Always create a keyspace first and then create tables inside it.
10. Unit Summary
NoSQL UNIT 3 — REDIS, COLUMN-FAMILY STORES & CASSANDRA
│
├── Key-Value Stores
│ ├── Definition and structure
│ ├── Real-world examples (cart, session)
│ └── Advantages
│
├── Redis — Introduction
│ ├── In-memory key-value store
│ ├── Data structures overview
│ └── Use cases
│
├── Redis Data Structures
│ ├── String (SET, GET, INCR, DECR)
│ ├── List (LPUSH, RPUSH, LPOP, RPOP, LRANGE)
│ ├── Set (SADD, SMEMBERS, SISMEMBER, SINTER)
│ ├── Hash (HSET, HGET, HGETALL, HINCRBY)
│ ├── Sorted Set (ZADD, ZRANGE, ZREVRANGE, ZSCORE)
│ └── Stream (XADD, XREAD, XRANGE)
│
├── Redis Key Management
│ ├── SET / GET
│ ├── EXPIRE / TTL
│ └── Common key commands
│
├── Redis Transactions
│ ├── MULTI (start)
│ ├── EXEC (execute)
│ ├── WATCH (optimistic locking)
│ ├── DISCARD (cancel)
│ └── Use case: money transfer
│
├── Caching — Cache-Aside
│ ├── Lazy loading strategy
│ ├── Read flow (hit/miss)
│ ├── Write flow (invalidate)
│ ├── E-commerce example
│ ├── Benefits
│ └── Considerations
│
├── Column-Family Store
│ ├── Definition (wide-column)
│ ├── Sparse storage
│ ├── Examples (Cassandra, HBase, DynamoDB)
│ ├── Comparison with RDBMS
│ └── When to use
│
├── Apache Cassandra — Core Concepts
│ ├── Node
│ ├── Cluster
│ ├── Keyspace
│ ├── Partition
│ ├── Partition Key
│ ├── Replication Factor
│ └── Data model hierarchy
│
└── CQL — Cassandra Query Language
├── CQL vs SQL
├── CREATE KEYSPACE
├── USE
├── CREATE TABLE
├── INSERT / SELECT / UPDATE / DELETE
└── DESCRIBE commands11. Key Commands — Quick Reference
Redis Commands
| Command | Purpose |
|---|---|
SET key value | Store a value |
GET key | Retrieve a value |
INCR key / DECR key | Increment/decrement |
LPUSH key value | Add to left of list |
RPUSH key value | Add to right of list |
LRANGE key start stop | Get list range |
LPOP key / RPOP key | Remove from list |
SADD key member | Add to set |
SMEMBERS key | Get all set members |
SISMEMBER key member | Check membership |
SINTER key1 key2 | Intersection |
HSET key field value | Set hash field |
HGET key field | Get hash field |
HGETALL key | Get all hash fields |
HINCRBY key field n | Increment hash field |
ZADD key score member | Add to sorted set |
ZRANGE key start stop | Get sorted range |
ZREVRANGE key start stop | Get reverse sorted range |
ZSCORE key member | Get score |
EXPIRE key seconds | Set TTL |
TTL key | Check TTL |
DEL key | Delete key |
EXISTS key | Check existence |
MULTI / EXEC / DISCARD | Transactions |
WATCH key | Optimistic locking |
CQL Commands
| Command | Purpose |
|---|---|
CREATE KEYSPACE name WITH replication = {...} | Create keyspace |
USE keyspace | Switch keyspace |
CREATE TABLE name (...) | Create table |
INSERT INTO table (...) VALUES (...) | Insert data |
SELECT * FROM table | Query data |
UPDATE table SET col = val WHERE ... | Update data |
DELETE FROM table WHERE ... | Delete data |
DESCRIBE KEYSPACES | List keyspaces |
DESCRIBE TABLES | List tables |
DESCRIBE TABLE name | Show table structure |
12. Exam-Focused Points
- Key-value store — Simple key → value mapping; fast lookups.
- Redis — In-memory key-value store with multiple data structures.
- Redis data structures — String, List, Set, Hash, Sorted Set, Stream.
- String commands — SET, GET, INCR, DECR.
- List commands — LPUSH, RPUSH, LPOP, RPOP, LRANGE.
- Set commands — SADD, SMEMBERS, SISMEMBER, SINTER.
- Hash commands — HSET, HGET, HGETALL, HINCRBY.
- Sorted Set commands — ZADD, ZRANGE, ZREVRANGE, ZSCORE.
- Key expiry — EXPIRE, TTL, PERSIST.
- Redis transactions — MULTI, EXEC, WATCH, DISCARD.
- WATCH — Optimistic locking; aborts transaction if key changes.
- Cache-Aside — Application manages cache; read: check cache → DB → cache; write: update DB → invalidate cache.
- Column-family store — Wide-column; sparse; flexible schema; horizontally scalable.
- Cassandra vs RDBMS — No joins vs joins; eventual vs strong consistency; distributed vs single server.
- Cassandra node — Single instance; peer-to-peer; no master.
- Cassandra cluster — Group of nodes; high availability.
- Keyspace — Top-level container (like database).
- Partition — Subset of data identified by partition key.
- Partition key — Determines data distribution across nodes.
- Replication factor — Number of copies of each partition.
- CQL — Cassandra Query Language; similar to SQL but for NoSQL.
- CQL vs SQL — Keyspace vs database; no joins; eventual consistency.
- CREATE KEYSPACE — Defines replication strategy and factor.
- USE — Switch to a keyspace.
- CREATE TABLE — Defines table with primary key.
- DESCRIBE — View keyspaces, tables, and structure.
- Cache hit vs cache miss — Found in cache vs not found.
- TTL — Time-to-live for cache entries.
- Cache invalidation — Delete/update cache on data changes.
- Sparse storage — Missing values not stored in column-family stores.
13. Glossary of Key Terms
| Term | Definition |
|---|---|
| Key-Value Store | NoSQL database storing data as key-value pairs |
| Redis | In-memory key-value store with multiple data structures |
| String | Basic Redis type; text, number, or binary |
| List | Ordered collection; allows duplicates |
| Set | Unordered collection; unique elements |
| Hash | Field-value pairs under one key |
| Sorted Set (ZSet) | Unique elements with scores; auto-sorted |
| Stream | Append-only log-like structure |
| TTL | Time-to-live; automatic key expiry |
| MULTI | Starts a Redis transaction |
| EXEC | Executes queued commands |
| WATCH | Monitors keys for changes (optimistic locking) |
| DISCARD | Cancels a transaction |
| Cache-Aside | Application-managed caching strategy |
| Cache Hit | Data found in cache |
| Cache Miss | Data not found in cache |
| Cache Invalidation | Removing/updating stale cache entries |
| Column-Family Store | Wide-column NoSQL database |
| Wide-Column | Storage model with flexible columns per row |
| Sparse Storage | Missing values not stored |
| Cassandra | Distributed wide-column NoSQL database |
| Node | Single Cassandra instance |
| Cluster | Group of Cassandra nodes |
| Keyspace | Top-level container (like database) |
| Table | Stores rows of data (column family) |
| Partition | Subset of data identified by partition key |
| Partition Key | Determines data distribution |
| Replication Factor | Number of copies of each partition |
| CQL | Cassandra Query Language |
| SimpleStrategy | Replication strategy for single data center |
| NetworkTopologyStrategy | Replication strategy for multiple data centers |
| Eventual Consistency | Data becomes consistent over time |
| Horizontal Scalability | Adding more nodes to scale |
| Vertical Scalability | Adding more resources to a single node |