BTCE | 5th Sem
No-SQL SubjectUnit 2

No-SQL Unit 2: Complete Concepts Guide

Unit 2: Document Databases using MongoDB -> Generated and Prepared By Thiruselvan (ThiruXD)

1. INTRODUCTION TO DOCUMENT DATABASES

What are Document Databases?

Document databases are a major category of NoSQL databases that store data as flexible, self-describing documents instead of strict rows and columns. Each document contains field-and-value pairs, and the values may include:

  • Strings
  • Numbers
  • Dates
  • Arrays
  • Embedded documents
  • Other supported data types

Real-World Entities as Documents

Many real-world entities can be represented as structured documents. For example, a student record may include:

  • Personal details
  • Address
  • Course details
  • Marks
  • Skills

Instead of splitting all of this data into separate tables, a document database can store much of it together in one document.

Common Applications Using Document Databases

  • E-commerce platforms
  • Learning management systems
  • Mobile apps
  • Content management systems
  • Healthcare records
  • IoT systems
  • User profile management

2. MONGODB OVERVIEW

What is MongoDB?

MongoDB is one of the most widely used document database systems. It stores records as BSON documents, which are binary representations of JSON-like documents.

MongoDB Organization Structure

Database → Collections → Documents
ComponentDescriptionAnalogy to RDBMS
DatabaseLogical container that holds collectionsDatabase
CollectionGroup of MongoDB documents, similar to a table but without fixed row schemaTable
DocumentSingle record made of fields and values, similar to a JSON objectRow/Record

Why Use MongoDB?

  • Flexible document model
  • Developer-friendly syntax (JSON-like documents)
  • Powerful query language
  • Secondary indexes
  • Aggregation framework
  • Rich tools ecosystem
  • Distributed features (replication and sharding)

3. CORE DOCUMENT DATABASE FEATURES

FeatureDescriptionExample
Document-based storageData stored as documents containing fields and valuesStudent document stores name, department, address, marks together
Flexible schemaDocuments in same collection need not have identical fieldsOne student may have internship details while another may not
Nested documentsDocuments can contain other documents as field valuesaddress: {city: "Bengaluru", pin: 560001}
ArraysFields can store lists of values or lists of documentsskills: ["Python", "MongoDB", "Testing"]
Natural object mappingDocuments closely resemble objects in programming languagesJavaScript object stored with minimal transformation
Ad hoc queryingUsers can query based on fields, ranges, patterns, conditionsFind students with score greater than 80
IndexingIndexes speed up frequently used queriesIndex on department improves department-based searches
AggregationData can be filtered, grouped, sorted, and summarizedCalculate average marks department-wise
Horizontal scalingDistributed storage supportLarge datasets spread across servers using sharding

4. DOCUMENT DATABASE VS RELATIONAL DATABASE

DimensionRelational DatabaseDocument Database / MongoDB
Data unitRow in a tableDocument in a collection
StructureFixed schema with columns and data typesFlexible document structure
RelationshipsManaged using foreign keys and joinsManaged using embedding or references
Best suited forHighly structured data with strong consistency and complex transactionsFlexible, hierarchical, and rapidly changing data
Query styleSQL queriesMongoDB Query Language and aggregation pipelines
Scaling patternCan scale vertically and horizontally depending on systemDesigned with distributed scaling options such as replication and sharding

5. MONGODB ADVANTAGES

AdvantageExplanationClassroom/Project Use
Flexible document modelDocuments can store nested and variable data structuresUseful for student profiles, product catalogs, event records
Developer-friendly syntaxDocuments resemble JSON objects used in web programmingStudents can easily understand records and queries
Powerful query languageSupports filtering, comparison operators, logical operators, projectionUseful for searching by department, marks, city, or status
Secondary indexesIndexes on commonly searched fieldsImproves speed of repeated queries
Aggregation frameworkAllows grouping, summarizing, and transforming dataUseful for reports such as average marks by department
Tools ecosystemMongoDB Shell, Compass, Atlas, and driversStudents can practice using GUI, shell, and programming languages
Distributed featuresReplica sets for high availability, sharding for horizontal scalingIntroduces modern database architecture concepts

6. ENVIRONMENT SETUP

Installation Options

Installation OptionWhat It MeansWhen to Use
MongoDB Community EditionFree self-managed MongoDB server installed locallyBest for offline labs, local practice, understanding server processes
MongoDB AtlasFully managed MongoDB cloud serviceBest when students don’t want to install a local database server
MongoDB CompassGraphical user interface for browsing databases and running queriesBest for beginners who prefer visual exploration
MongoDB Shell (mongosh)Command-line shell for connecting and running MongoDB commandsBest for learning CRUD commands and database syntax

General Local Installation Steps

  1. Visit the official MongoDB Community Server download/installation documentation
  2. Select the correct operating system (Windows, Ubuntu, macOS, or other supported platform)
  3. Download and install MongoDB Community Server using installer or package manager
  4. Install MongoDB Shell if not included
  5. Start the MongoDB server service (typically named mongod)
  6. Open terminal/command prompt and run mongosh to connect
  7. Verify connection with commands like show dbs and db.version()

Verification Commands:

show dbs
db.version()
db.stats()

Connection Methods

Connection MethodExamplePurpose
Local mongoshmongoshConnects to default local MongoDB server
Explicit local URImongosh "mongodb://localhost:27017"Connects to local server using a URI
Atlas SRV URImongosh "mongodb+srv://cluster-url"Connects to MongoDB Atlas cluster
Application driverMongoClient("mongodb://localhost:27017")Allows programs to connect from Python, Node.js, Java, etc.
Compass GUIPaste connection string in MongoDB CompassConnects through a graphical interface

Important: Use mongodb+srv:// for SRV connection strings (Atlas clusters). Use mongodb:// for standard connection strings (local servers or explicitly listed hosts).


7. MONGODB SHELL (MONGOSH)

What is MongoDB Shell?

MongoDB Shell (mongosh) is an interactive command-line interface for MongoDB. It is useful for:

  • Learning commands
  • Testing queries
  • Managing databases
  • Running administrative operations

Essential mongosh Commands

CommandPurpose
show dbsDisplays available databases
use databaseNameSwitches to a database (creates it when data is inserted)
show collectionsDisplays collections in the current database
dbShows the current database name
db.collection.find()Reads documents from a collection
db.collection.countDocuments()Counts documents in a collection
db.collection.drop()Drops a collection
db.dropDatabase()Deletes the current database
clsClears the shell screen (in many terminals)
exitExits the shell

Common Beginner Shell Session

// 1. Start shell
mongosh

// 2. View databases
show dbs

// 3. Switch to a database
use collegeDB

// 4. Insert a sample document
db.students.insertOne({ name: "Sneha", department: "AIML", semester: 5 })

// 5. View collections
show collections

// 6. Read documents
db.students.find()

// 7. Count documents
db.students.countDocuments()

8. CRUD OPERATIONS

Preparing a Sample Database

use collegeDB

db.students.insertMany([
  {
    name: "Asha",
    department: "CSE",
    semester: 5,
    city: "Mysuru",
    marks: { dbms: 86, python: 92 },
    skills: ["Python", "MongoDB"]
  },
  {
    name: "Ravi",
    department: "ISE",
    semester: 5,
    city: "Bengaluru",
    marks: { dbms: 78, python: 81 },
    skills: ["Java", "Testing"]
  },
  {
    name: "Meena",
    department: "CSE",
    semester: 3,
    city: "Hubballi",
    marks: { dbms: 91, python: 88 },
    skills: ["Python", "Data Analysis"]
  }
])

8.1 CREATE Operations

Create operations add new documents to a collection.

MethodPurposeExample Use
insertOne()Inserts a single documentAdd one new student
insertMany()Inserts multiple documents at onceAdd a batch of students

Syntax Examples:

// Insert one document
db.students.insertOne({
  name: "Kiran",
  department: "ECE",
  semester: 5,
  city: "Dharwad",
  marks: { dbms: 74, python: 69 },
  skills: ["C", "Electronics"]
})

// Insert multiple documents
db.students.insertMany([
  { name: "John", department: "CSE", semester: 3 },
  { name: "Jane", department: "ISE", semester: 5 }
])

8.2 READ Operations

Read operations retrieve documents from a collection.

Query NeedMongoDB CommandMeaning
All studentsdb.students.find()Displays every document in the collection
One CSE studentdb.students.findOne({ department: "CSE" })Returns one matching document
Marks > 80db.students.find({ "marks.dbms": { $gt: 80 } })Uses comparison operator $gt
Only selected fieldsdb.students.find({}, { name: 1, department: 1, _id: 0 })Uses projection
Sort by namedb.students.find().sort({ name: 1 })Sorts ascending by name

Syntax Examples:

// Read all documents
db.students.find()

// Display in readable format
db.students.find().pretty()

// Find CSE students from semester 5
db.students.find({ department: "CSE", semester: 5 })

// Find students with DBMS marks > 80
db.students.find({ "marks.dbms": { $gt: 80 } })

// Projection: show only name and department
db.students.find({}, { name: 1, department: 1, _id: 0 })

8.3 UPDATE Operations

Update operations modify existing documents.

Method/OperatorPurposeExample
updateOne()Updates one matching documentChange one student’s city
updateMany()Updates all matching documentsAdd mentor field to all CSE students
$setSets or changes a field value{ $set: { city: "Mysuru" } }
$incIncrements a numeric value{ $inc: { semester: 1 } }
$pushAdds an item to an array{ $push: { skills: "AI" } }

Syntax Examples:

// Update one document
db.students.updateOne(
  { name: "Ravi" },
  { $set: { city: "Mangaluru" } }
)

// Update many documents
db.students.updateMany(
  { department: "CSE" },
  { $set: { mentor: "Dr. Rao" } }
)

// Add a new skill to one student's skills array
db.students.updateOne(
  { name: "Asha" },
  { $push: { skills: "Cloud" } }
)

// Increment semester for all students
db.students.updateMany(
  {},
  { $inc: { semester: 1 } }
)

8.4 DELETE Operations

Delete operations remove documents from a collection.

MethodPurpose
deleteOne()Removes the first matching document
deleteMany()Removes all matching documents

Syntax Examples:

// Delete one matching document
db.students.deleteOne({ name: "Kiran" })

// Delete all students from a city
db.students.deleteMany({ city: "Hubballi" })

// Count remaining documents
db.students.countDocuments()

Important: Use delete operations carefully - removed documents may not be recoverable unless backups are available.


9. DATA MODELLING PRINCIPLES

What is Data Modelling?

Data modelling means deciding how data will be organized inside the database. In MongoDB, good data modelling starts with access patterns - the designer should first ask:

  • What data will the application read together?
  • What data will be updated frequently?
  • Which queries must be fast?

Key Data Modelling Principles

PrincipleMeaningExample
Model for application queriesDesign documents based on how the application reads and writes dataStore recent product reviews with product data if they are always shown together
Store related data together when appropriateUse embedding when data is commonly accessed as one unitStore address inside a student document
Use references for large or independent dataStore related data in separate collections and link using IDsStore all exam attempts separately if the list grows too large
Avoid unlimited array growthVery large arrays inside one document can become inefficientDo not store thousands of log entries inside a single user document
Keep frequently updated data separate when neededHeavy updates to one part of a document can affect performanceSeparate live attendance logs from static student profile data
Use schema validation when requiredMongoDB can enforce rules for fields and data typesRequire name, department, and semester for student records

Embedding vs Referencing

Embedding

Related data is stored inside the same document.

Use when:

  • Related data is read together
  • Data has a bounded size
  • Data doesn’t change independently

Advantages:

  • Fast single-document reads
  • Natural document structure
  • Atomic updates

Precautions:

  • Avoid very large documents
  • Avoid unbounded arrays

Example:

{
  name: "Asha",
  department: "CSE",
  address: {
    city: "Mysuru",
    pin: 570001
  }
}

Referencing

Related data is stored in separate collections and linked using IDs.

Use when:

  • Related data is large
  • Data is shared between documents
  • Data is frequently updated independently

Advantages:

  • Reduces duplication
  • Separates large datasets
  • More flexible for independent updates

Precautions:

  • May require multiple queries
  • May need aggregation lookup

Example:

{
  name: "Asha",
  departmentId: ObjectId("...")
}

Design Rule

Data that is accessed together should generally be stored together. However, if related data grows without limit or is updated independently, referencing may be a better design.


10. INDEXING STRATEGIES

What are Indexes?

Indexes are special data structures that make queries faster. Without an index, MongoDB may need to scan every document in a collection to find matches. With a suitable index, MongoDB can locate matching documents more efficiently.

Types of Indexes

Index TypePurposeExample Command
Single-field indexSupports queries on one fielddb.students.createIndex({ department: 1 })
Compound indexSupports queries using multiple fields in defined orderdb.students.createIndex({ department: 1, semester: 1 })
Multikey indexIndexes values inside an array fielddb.students.createIndex({ skills: 1 })
Text indexSupports text search on string contentdb.articles.createIndex({ title: "text", body: "text" })
Unique indexPrevents duplicate values for a fielddb.students.createIndex({ usn: 1 }, { unique: true })

Creating and Viewing Indexes

// Create a single-field index
db.students.createIndex({ department: 1 })

// Create a compound index
db.students.createIndex({ department: 1, semester: 1 })

// Create an index on an array field
db.students.createIndex({ skills: 1 })

// View existing indexes
db.students.getIndexes()

// Drop an index by name
db.students.dropIndex("department_1")

Indexing Strategy Rules

  1. Create indexes for fields used frequently in filters, sorting, and joins/lookups
  2. Use compound indexes when queries commonly use multiple fields together
  3. Choose compound index field order carefully - the first field and prefix fields are important for query support
  4. Avoid creating too many indexes (each index uses storage and slows write operations)
  5. Use explain() to check whether a query is using an index
  6. Remove unused indexes after analyzing workload and performance needs

Analyzing Query Execution:

db.students.find({ department: "CSE" }).explain("executionStats")

Best Practice

Indexes improve read performance but add overhead to insert, update, and delete operations because MongoDB must maintain the index structures. Create indexes based on real query requirements.


11. AGGREGATION FRAMEWORK

What is the Aggregation Framework?

The aggregation framework processes documents through a pipeline of stages. Each stage receives input documents, performs an operation, and passes the result to the next stage. Aggregation is used for:

  • Reporting
  • Analytics
  • Grouping
  • Filtering
  • Sorting
  • Reshaping documents

Common Aggregation Stages

StagePurposeExample Use
$matchFilters documents based on conditionsSelect only CSE students
$projectSelects, removes, or reshapes fieldsShow name and total marks only
$groupGroups documents and computes totals, averages, counts, etc.Average DBMS marks department-wise
$sortSorts documentsSort departments by average marks
$limitLimits number of output documentsShow top five results
$unwindBreaks array values into separate documentsProcess each skill separately
$lookupPerforms left outer join-like operationConnect students with department details

Example 1: Department-wise Average Marks

db.students.aggregate([
  {
    $group: {
      _id: "$department",
      averageDBMS: { $avg: "$marks.dbms" },
      totalStudents: { $sum: 1 }
    }
  },
  {
    $sort: { averageDBMS: -1 }
  }
])

Example 2: Filter, Project, and Sort

db.students.aggregate([
  { $match: { semester: 5 } },
  {
    $project: {
      _id: 0,
      name: 1,
      department: 1,
      totalMarks: { $add: ["$marks.dbms", "$marks.python"] }
    }
  },
  { $sort: { totalMarks: -1 } }
])

Explanation: This pipeline first selects fifth-semester students, then creates a new computed field called totalMarks, and finally sorts the output in descending order of total marks.

Example 3: Working with Array Data

db.students.aggregate([
  { $unwind: "$skills" },
  {
    $group: {
      _id: "$skills",
      numberOfStudents: { $sum: 1 }
    }
  },
  { $sort: { numberOfStudents: -1 } }
])

Explanation: $unwind separates each skill in the skills array into an individual pipeline item. Then $group counts how many students have each skill.

Important Note

Aggregation pipelines usually do not modify the original documents. However, special stages such as $merge or $out can write results to a collection, so use them carefully.


12. PERFORMANCE ANALYSIS

Performance Factors

Performance FactorGood PracticeProblem if Ignored
Document designStore related bounded data togetherToo many queries or very large documents
IndexesCreate indexes for frequent filters and sortsSlow collection scans
ProjectionReturn only fields needed by the applicationUnnecessary network and memory use
Aggregation orderPlace selective $match stages early when possibleMore documents processed in later stages
Array designKeep arrays bounded or separate large eventsDocument growth and update overhead
Write workloadAvoid excessive indexes on write-heavy collectionsSlower inserts and updates

13. KEY TERMS GLOSSARY

TermMeaning
Document DatabaseA NoSQL database model that stores records as documents containing field-and-value pairs
MongoDBA document-oriented NoSQL database that stores data records as BSON documents
BSONBinary JSON; MongoDB uses BSON to represent documents internally and support additional data types
DatabaseA logical container that holds collections
CollectionA group of MongoDB documents, similar to a table in a relational database but without a fixed row schema
DocumentA single record made of fields and values, often similar to a JSON object
CRUDCreate, Read, Update, and Delete operations performed on documents
mongoshThe MongoDB Shell used to connect to MongoDB and execute commands interactively
IndexA data structure that improves query performance by allowing faster lookup of documents
Aggregation PipelineA sequence of stages that processes documents and returns computed or transformed results
EmbeddingStoring related data inside the same document
ReferencingStoring a link or identifier to related data stored in another collection

14. CHAPTER SUMMARY

  • Document databases store records as flexible documents instead of fixed rows and columns
  • MongoDB organizes data into databases, collections, and BSON documents
  • MongoDB supports CRUD operations through methods such as insertOne(), find(), updateOne(), and deleteOne()
  • mongosh is an important tool for interactive MongoDB practice
  • Good MongoDB data modelling depends on application access patterns
  • Embedding is useful when related data is read together and has a bounded size; referencing is useful when related data is large or independent
  • Indexes improve query performance but add storage and write overhead
  • Aggregation pipelines process documents through stages such as $match, $group, $sort, and $project

On this page