No-SQL Unit 2: Complete Concepts Guide
Unit 2: Document Databases using MongoDB -> Generated and Prepared By Thiruselvan (ThiruXD)
1. INTRODUCTION TO DOCUMENT DATABASES
What are Document Databases?
Document databases are a major category of NoSQL databases that store data as flexible, self-describing documents instead of strict rows and columns. Each document contains field-and-value pairs, and the values may include:
- Strings
- Numbers
- Dates
- Arrays
- Embedded documents
- Other supported data types
Real-World Entities as Documents
Many real-world entities can be represented as structured documents. For example, a student record may include:
- Personal details
- Address
- Course details
- Marks
- Skills
Instead of splitting all of this data into separate tables, a document database can store much of it together in one document.
Common Applications Using Document Databases
- E-commerce platforms
- Learning management systems
- Mobile apps
- Content management systems
- Healthcare records
- IoT systems
- User profile management
2. MONGODB OVERVIEW
What is MongoDB?
MongoDB is one of the most widely used document database systems. It stores records as BSON documents, which are binary representations of JSON-like documents.
MongoDB Organization Structure
Database → Collections → Documents| Component | Description | Analogy to RDBMS |
|---|---|---|
| Database | Logical container that holds collections | Database |
| Collection | Group of MongoDB documents, similar to a table but without fixed row schema | Table |
| Document | Single record made of fields and values, similar to a JSON object | Row/Record |
Why Use MongoDB?
- Flexible document model
- Developer-friendly syntax (JSON-like documents)
- Powerful query language
- Secondary indexes
- Aggregation framework
- Rich tools ecosystem
- Distributed features (replication and sharding)
3. CORE DOCUMENT DATABASE FEATURES
| Feature | Description | Example |
|---|---|---|
| Document-based storage | Data stored as documents containing fields and values | Student document stores name, department, address, marks together |
| Flexible schema | Documents in same collection need not have identical fields | One student may have internship details while another may not |
| Nested documents | Documents can contain other documents as field values | address: {city: "Bengaluru", pin: 560001} |
| Arrays | Fields can store lists of values or lists of documents | skills: ["Python", "MongoDB", "Testing"] |
| Natural object mapping | Documents closely resemble objects in programming languages | JavaScript object stored with minimal transformation |
| Ad hoc querying | Users can query based on fields, ranges, patterns, conditions | Find students with score greater than 80 |
| Indexing | Indexes speed up frequently used queries | Index on department improves department-based searches |
| Aggregation | Data can be filtered, grouped, sorted, and summarized | Calculate average marks department-wise |
| Horizontal scaling | Distributed storage support | Large datasets spread across servers using sharding |
4. DOCUMENT DATABASE VS RELATIONAL DATABASE
| Dimension | Relational Database | Document Database / MongoDB |
|---|---|---|
| Data unit | Row in a table | Document in a collection |
| Structure | Fixed schema with columns and data types | Flexible document structure |
| Relationships | Managed using foreign keys and joins | Managed using embedding or references |
| Best suited for | Highly structured data with strong consistency and complex transactions | Flexible, hierarchical, and rapidly changing data |
| Query style | SQL queries | MongoDB Query Language and aggregation pipelines |
| Scaling pattern | Can scale vertically and horizontally depending on system | Designed with distributed scaling options such as replication and sharding |
5. MONGODB ADVANTAGES
| Advantage | Explanation | Classroom/Project Use |
|---|---|---|
| Flexible document model | Documents can store nested and variable data structures | Useful for student profiles, product catalogs, event records |
| Developer-friendly syntax | Documents resemble JSON objects used in web programming | Students can easily understand records and queries |
| Powerful query language | Supports filtering, comparison operators, logical operators, projection | Useful for searching by department, marks, city, or status |
| Secondary indexes | Indexes on commonly searched fields | Improves speed of repeated queries |
| Aggregation framework | Allows grouping, summarizing, and transforming data | Useful for reports such as average marks by department |
| Tools ecosystem | MongoDB Shell, Compass, Atlas, and drivers | Students can practice using GUI, shell, and programming languages |
| Distributed features | Replica sets for high availability, sharding for horizontal scaling | Introduces modern database architecture concepts |
6. ENVIRONMENT SETUP
Installation Options
| Installation Option | What It Means | When to Use |
|---|---|---|
| MongoDB Community Edition | Free self-managed MongoDB server installed locally | Best for offline labs, local practice, understanding server processes |
| MongoDB Atlas | Fully managed MongoDB cloud service | Best when students don’t want to install a local database server |
| MongoDB Compass | Graphical user interface for browsing databases and running queries | Best for beginners who prefer visual exploration |
| MongoDB Shell (mongosh) | Command-line shell for connecting and running MongoDB commands | Best for learning CRUD commands and database syntax |
General Local Installation Steps
- Visit the official MongoDB Community Server download/installation documentation
- Select the correct operating system (Windows, Ubuntu, macOS, or other supported platform)
- Download and install MongoDB Community Server using installer or package manager
- Install MongoDB Shell if not included
- Start the MongoDB server service (typically named
mongod) - Open terminal/command prompt and run
mongoshto connect - Verify connection with commands like
show dbsanddb.version()
Verification Commands:
show dbs
db.version()
db.stats()Connection Methods
| Connection Method | Example | Purpose |
|---|---|---|
| Local mongosh | mongosh | Connects to default local MongoDB server |
| Explicit local URI | mongosh "mongodb://localhost:27017" | Connects to local server using a URI |
| Atlas SRV URI | mongosh "mongodb+srv://cluster-url" | Connects to MongoDB Atlas cluster |
| Application driver | MongoClient("mongodb://localhost:27017") | Allows programs to connect from Python, Node.js, Java, etc. |
| Compass GUI | Paste connection string in MongoDB Compass | Connects through a graphical interface |
Important: Use mongodb+srv:// for SRV connection strings (Atlas clusters). Use mongodb:// for standard connection strings (local servers or explicitly listed hosts).
7. MONGODB SHELL (MONGOSH)
What is MongoDB Shell?
MongoDB Shell (mongosh) is an interactive command-line interface for MongoDB. It is useful for:
- Learning commands
- Testing queries
- Managing databases
- Running administrative operations
Essential mongosh Commands
| Command | Purpose |
|---|---|
show dbs | Displays available databases |
use databaseName | Switches to a database (creates it when data is inserted) |
show collections | Displays collections in the current database |
db | Shows the current database name |
db.collection.find() | Reads documents from a collection |
db.collection.countDocuments() | Counts documents in a collection |
db.collection.drop() | Drops a collection |
db.dropDatabase() | Deletes the current database |
cls | Clears the shell screen (in many terminals) |
exit | Exits the shell |
Common Beginner Shell Session
// 1. Start shell
mongosh
// 2. View databases
show dbs
// 3. Switch to a database
use collegeDB
// 4. Insert a sample document
db.students.insertOne({ name: "Sneha", department: "AIML", semester: 5 })
// 5. View collections
show collections
// 6. Read documents
db.students.find()
// 7. Count documents
db.students.countDocuments()8. CRUD OPERATIONS
Preparing a Sample Database
use collegeDB
db.students.insertMany([
{
name: "Asha",
department: "CSE",
semester: 5,
city: "Mysuru",
marks: { dbms: 86, python: 92 },
skills: ["Python", "MongoDB"]
},
{
name: "Ravi",
department: "ISE",
semester: 5,
city: "Bengaluru",
marks: { dbms: 78, python: 81 },
skills: ["Java", "Testing"]
},
{
name: "Meena",
department: "CSE",
semester: 3,
city: "Hubballi",
marks: { dbms: 91, python: 88 },
skills: ["Python", "Data Analysis"]
}
])8.1 CREATE Operations
Create operations add new documents to a collection.
| Method | Purpose | Example Use |
|---|---|---|
insertOne() | Inserts a single document | Add one new student |
insertMany() | Inserts multiple documents at once | Add a batch of students |
Syntax Examples:
// Insert one document
db.students.insertOne({
name: "Kiran",
department: "ECE",
semester: 5,
city: "Dharwad",
marks: { dbms: 74, python: 69 },
skills: ["C", "Electronics"]
})
// Insert multiple documents
db.students.insertMany([
{ name: "John", department: "CSE", semester: 3 },
{ name: "Jane", department: "ISE", semester: 5 }
])8.2 READ Operations
Read operations retrieve documents from a collection.
| Query Need | MongoDB Command | Meaning |
|---|---|---|
| All students | db.students.find() | Displays every document in the collection |
| One CSE student | db.students.findOne({ department: "CSE" }) | Returns one matching document |
| Marks > 80 | db.students.find({ "marks.dbms": { $gt: 80 } }) | Uses comparison operator $gt |
| Only selected fields | db.students.find({}, { name: 1, department: 1, _id: 0 }) | Uses projection |
| Sort by name | db.students.find().sort({ name: 1 }) | Sorts ascending by name |
Syntax Examples:
// Read all documents
db.students.find()
// Display in readable format
db.students.find().pretty()
// Find CSE students from semester 5
db.students.find({ department: "CSE", semester: 5 })
// Find students with DBMS marks > 80
db.students.find({ "marks.dbms": { $gt: 80 } })
// Projection: show only name and department
db.students.find({}, { name: 1, department: 1, _id: 0 })8.3 UPDATE Operations
Update operations modify existing documents.
| Method/Operator | Purpose | Example |
|---|---|---|
updateOne() | Updates one matching document | Change one student’s city |
updateMany() | Updates all matching documents | Add mentor field to all CSE students |
$set | Sets or changes a field value | { $set: { city: "Mysuru" } } |
$inc | Increments a numeric value | { $inc: { semester: 1 } } |
$push | Adds an item to an array | { $push: { skills: "AI" } } |
Syntax Examples:
// Update one document
db.students.updateOne(
{ name: "Ravi" },
{ $set: { city: "Mangaluru" } }
)
// Update many documents
db.students.updateMany(
{ department: "CSE" },
{ $set: { mentor: "Dr. Rao" } }
)
// Add a new skill to one student's skills array
db.students.updateOne(
{ name: "Asha" },
{ $push: { skills: "Cloud" } }
)
// Increment semester for all students
db.students.updateMany(
{},
{ $inc: { semester: 1 } }
)8.4 DELETE Operations
Delete operations remove documents from a collection.
| Method | Purpose |
|---|---|
deleteOne() | Removes the first matching document |
deleteMany() | Removes all matching documents |
Syntax Examples:
// Delete one matching document
db.students.deleteOne({ name: "Kiran" })
// Delete all students from a city
db.students.deleteMany({ city: "Hubballi" })
// Count remaining documents
db.students.countDocuments()Important: Use delete operations carefully - removed documents may not be recoverable unless backups are available.
9. DATA MODELLING PRINCIPLES
What is Data Modelling?
Data modelling means deciding how data will be organized inside the database. In MongoDB, good data modelling starts with access patterns - the designer should first ask:
- What data will the application read together?
- What data will be updated frequently?
- Which queries must be fast?
Key Data Modelling Principles
| Principle | Meaning | Example |
|---|---|---|
| Model for application queries | Design documents based on how the application reads and writes data | Store recent product reviews with product data if they are always shown together |
| Store related data together when appropriate | Use embedding when data is commonly accessed as one unit | Store address inside a student document |
| Use references for large or independent data | Store related data in separate collections and link using IDs | Store all exam attempts separately if the list grows too large |
| Avoid unlimited array growth | Very large arrays inside one document can become inefficient | Do not store thousands of log entries inside a single user document |
| Keep frequently updated data separate when needed | Heavy updates to one part of a document can affect performance | Separate live attendance logs from static student profile data |
| Use schema validation when required | MongoDB can enforce rules for fields and data types | Require name, department, and semester for student records |
Embedding vs Referencing
Embedding
Related data is stored inside the same document.
Use when:
- Related data is read together
- Data has a bounded size
- Data doesn’t change independently
Advantages:
- Fast single-document reads
- Natural document structure
- Atomic updates
Precautions:
- Avoid very large documents
- Avoid unbounded arrays
Example:
{
name: "Asha",
department: "CSE",
address: {
city: "Mysuru",
pin: 570001
}
}Referencing
Related data is stored in separate collections and linked using IDs.
Use when:
- Related data is large
- Data is shared between documents
- Data is frequently updated independently
Advantages:
- Reduces duplication
- Separates large datasets
- More flexible for independent updates
Precautions:
- May require multiple queries
- May need aggregation lookup
Example:
{
name: "Asha",
departmentId: ObjectId("...")
}Design Rule
Data that is accessed together should generally be stored together. However, if related data grows without limit or is updated independently, referencing may be a better design.
10. INDEXING STRATEGIES
What are Indexes?
Indexes are special data structures that make queries faster. Without an index, MongoDB may need to scan every document in a collection to find matches. With a suitable index, MongoDB can locate matching documents more efficiently.
Types of Indexes
| Index Type | Purpose | Example Command |
|---|---|---|
| Single-field index | Supports queries on one field | db.students.createIndex({ department: 1 }) |
| Compound index | Supports queries using multiple fields in defined order | db.students.createIndex({ department: 1, semester: 1 }) |
| Multikey index | Indexes values inside an array field | db.students.createIndex({ skills: 1 }) |
| Text index | Supports text search on string content | db.articles.createIndex({ title: "text", body: "text" }) |
| Unique index | Prevents duplicate values for a field | db.students.createIndex({ usn: 1 }, { unique: true }) |
Creating and Viewing Indexes
// Create a single-field index
db.students.createIndex({ department: 1 })
// Create a compound index
db.students.createIndex({ department: 1, semester: 1 })
// Create an index on an array field
db.students.createIndex({ skills: 1 })
// View existing indexes
db.students.getIndexes()
// Drop an index by name
db.students.dropIndex("department_1")Indexing Strategy Rules
- Create indexes for fields used frequently in filters, sorting, and joins/lookups
- Use compound indexes when queries commonly use multiple fields together
- Choose compound index field order carefully - the first field and prefix fields are important for query support
- Avoid creating too many indexes (each index uses storage and slows write operations)
- Use
explain()to check whether a query is using an index - Remove unused indexes after analyzing workload and performance needs
Analyzing Query Execution:
db.students.find({ department: "CSE" }).explain("executionStats")Best Practice
Indexes improve read performance but add overhead to insert, update, and delete operations because MongoDB must maintain the index structures. Create indexes based on real query requirements.
11. AGGREGATION FRAMEWORK
What is the Aggregation Framework?
The aggregation framework processes documents through a pipeline of stages. Each stage receives input documents, performs an operation, and passes the result to the next stage. Aggregation is used for:
- Reporting
- Analytics
- Grouping
- Filtering
- Sorting
- Reshaping documents
Common Aggregation Stages
| Stage | Purpose | Example Use |
|---|---|---|
$match | Filters documents based on conditions | Select only CSE students |
$project | Selects, removes, or reshapes fields | Show name and total marks only |
$group | Groups documents and computes totals, averages, counts, etc. | Average DBMS marks department-wise |
$sort | Sorts documents | Sort departments by average marks |
$limit | Limits number of output documents | Show top five results |
$unwind | Breaks array values into separate documents | Process each skill separately |
$lookup | Performs left outer join-like operation | Connect students with department details |
Example 1: Department-wise Average Marks
db.students.aggregate([
{
$group: {
_id: "$department",
averageDBMS: { $avg: "$marks.dbms" },
totalStudents: { $sum: 1 }
}
},
{
$sort: { averageDBMS: -1 }
}
])Example 2: Filter, Project, and Sort
db.students.aggregate([
{ $match: { semester: 5 } },
{
$project: {
_id: 0,
name: 1,
department: 1,
totalMarks: { $add: ["$marks.dbms", "$marks.python"] }
}
},
{ $sort: { totalMarks: -1 } }
])Explanation: This pipeline first selects fifth-semester students, then creates a new computed field called totalMarks, and finally sorts the output in descending order of total marks.
Example 3: Working with Array Data
db.students.aggregate([
{ $unwind: "$skills" },
{
$group: {
_id: "$skills",
numberOfStudents: { $sum: 1 }
}
},
{ $sort: { numberOfStudents: -1 } }
])Explanation: $unwind separates each skill in the skills array into an individual pipeline item. Then $group counts how many students have each skill.
Important Note
Aggregation pipelines usually do not modify the original documents. However, special stages such as
$mergeor$outcan write results to a collection, so use them carefully.
12. PERFORMANCE ANALYSIS
Performance Factors
| Performance Factor | Good Practice | Problem if Ignored |
|---|---|---|
| Document design | Store related bounded data together | Too many queries or very large documents |
| Indexes | Create indexes for frequent filters and sorts | Slow collection scans |
| Projection | Return only fields needed by the application | Unnecessary network and memory use |
| Aggregation order | Place selective $match stages early when possible | More documents processed in later stages |
| Array design | Keep arrays bounded or separate large events | Document growth and update overhead |
| Write workload | Avoid excessive indexes on write-heavy collections | Slower inserts and updates |
13. KEY TERMS GLOSSARY
| Term | Meaning |
|---|---|
| Document Database | A NoSQL database model that stores records as documents containing field-and-value pairs |
| MongoDB | A document-oriented NoSQL database that stores data records as BSON documents |
| BSON | Binary JSON; MongoDB uses BSON to represent documents internally and support additional data types |
| Database | A logical container that holds collections |
| Collection | A group of MongoDB documents, similar to a table in a relational database but without a fixed row schema |
| Document | A single record made of fields and values, often similar to a JSON object |
| CRUD | Create, Read, Update, and Delete operations performed on documents |
| mongosh | The MongoDB Shell used to connect to MongoDB and execute commands interactively |
| Index | A data structure that improves query performance by allowing faster lookup of documents |
| Aggregation Pipeline | A sequence of stages that processes documents and returns computed or transformed results |
| Embedding | Storing related data inside the same document |
| Referencing | Storing a link or identifier to related data stored in another collection |
14. CHAPTER SUMMARY
- Document databases store records as flexible documents instead of fixed rows and columns
- MongoDB organizes data into databases, collections, and BSON documents
- MongoDB supports CRUD operations through methods such as
insertOne(),find(),updateOne(), anddeleteOne() mongoshis an important tool for interactive MongoDB practice- Good MongoDB data modelling depends on application access patterns
- Embedding is useful when related data is read together and has a bounded size; referencing is useful when related data is large or independent
- Indexes improve query performance but add storage and write overhead
- Aggregation pipelines process documents through stages such as
$match,$group,$sort, and$project