B.Sc. (Computer Science) Information Retrieval Practical Syllabus - Mumbai University
This is the TY BSc Computer Science syllabus under NEP 2020, in force from the academic year 2026-27. The University still sets the earlier Choice Based papers alongside it — her Summer 2026 third-year timetables name that scheme — so check which scheme your exam form names before you revise.
Loading syllabus...
Syllabus for Information Retrieval Practical
This is the practical paper for Information Retrieval, which the University sets and examines separately.
Module I
- Text Preprocessing on Cranfield Collection
- Download the Cranfield document collection (1400 aeronautical abstracts).
- Perform: Tokenization, Stopword removal (NLTK stopword list), Porter stemming, Compute vocabulary size before and after preprocessing, Report top 20 frequent terms
- Inverted Index Construction for Cranfield Corpus Using the preprocessed Cranfield corpus build an inverted index, Store: term →
- {docID, term frequency}, Compute document frequency for each term, Print posting list for terms: “flow”, “boundary”, “shock”
- Positional Index & Phrase Query Processing Extend inverted index to store term positions. Process phrase queries: “boundary layer flow”, “supersonic shock wave”, Return matching document IDs.
- Boolean Retrieval using Cranfield Queries
- Using Cranfield query set to Implement Boolean retrieval, Process query: (boundary AND flow) AND NOT laminar, Return list of document IDs, Compare with relevance judgments file
- TF-IDF Weight Computation on Cranfield
- Compute TF, IDF, TF-IDF weights, Display top 10 highest weighted terms in Document ID 12 and Document ID 245. Dataset: Cranfield Collection
- Ranked Retrieval using Vector Space Model
- For Cranfield Query 1: Convert query and documents to TF-IDF vectors, Compute cosine similarity, Rank top 10 documents, Compare retrieved results with relevance judgments
- Evaluation of Boolean vs Vector Space Model
- For first 5 Cranfield queries: Retrieve top 10 results using Boolean model, Retrieve top 10 results using VSM, Compute Precision@10, Compare retrieval effectiveness
- Edit Distance Implementation for Query Correction
- Implement Levenshtein distance. Correct the following misspelled Cranfield query terms: “supersonc”, “aerodynamcs”, “turbulnce”. Suggest top 3 corrections from corpus vocabulary. Dataset: Cranfield vocabulary
- Integrated Spelling Correction in IR System
- Modify retrieval system so that when user enters: supersonc boundary layr. System automatically: Suggests corrected query, Performs retrieval using corrected query, Dataset: Cranfield Collection
- Precision, Recall and F1 on Cranfield Query 1
- Using Cranfield relevance judgments: Retrieve top 20 documents for Query 1
- Compute Precision, Recall, F1-score, Dataset: Cranfield Collection + qrels
Module II
- Average Precision and MAP
- For first 10 Cranfield queries: Compute Average Precision (AP) for each query, Compute Mean Average Precision (MAP), Compare with baseline Boolean model, Dataset: Cranfield Collection
- Precision-Recall Curve Visualization
- For Query 3 Compute precision at each recall level, Plot Precision–Recall curve, Compare Boolean vs VSM, Dataset: Cranfield Collection
- Naïve Bayes Classification using 20 Newsgroups
- Using 4 categories from 20 Newsgroups dataset: sci.space, comp.graphics, rec.sport.baseball, talk.politics.misc, Train Multinomial Naïve Bayes classifier and compute Accuracy, Confusion matrix, Macro F1-score, Dataset: 20 Newsgroups Dataset
- SVM Classification using 20 Newsgroups
- Using 20 Newsgroups dataset Train linear SVM, Compare performance with Naïve Bayes, Report Precision, Recall, F1, Dataset: 20 Newsgroups
- K-Means Clustering on BBC News Dataset
- Cluster BBC News dataset into 5 clusters. Analyze Top 10 terms per cluster, Compare clusters with actual categories, Dataset: BBC News Dataset
- Hierarchical Clustering on BBC Dataset
- Perform Agglomerative clustering. Generate dendrogram and Determine optimal number of clusters, Compare with K-Means, Dataset: BBC News Dataset
- Clustering Evaluation (Purity & Silhouette)
- For the above 2 practicals, Compute Cluster Purity, Silhouette Score For K = 3, 4, 5, Compare clustering quality., Dataset: BBC News Dataset
- Web Crawling & Link Analysis
- Web Crawling of a Website: Develop a crawler starting from a given URL, Crawl up to 50 pages, Extract page titles, Build inverted index, Respect robots.txt
- PageRank on Sample Web Graph
- Given web graph with 10 nodes Construct adjacency matrix, Implement iterative PageRank, Use damping factor = 0.85, Show convergence after 20 iterations Dataset:
- SNAP small web graph dataset (Stanford Network Analysis Project)
- Learning to Rank using LETOR Dataset
- Using LETOR 4.0 dataset Train RankSVM model, Evaluate using NDCG@10,
- Compare with TF-IDF baseline, Dataset: LETOR 4.0
Text Books
- 1 Ricardo Baeza-Yates and Berthier Ribeiro-Neto, ―Modern Information Retrieval: The Concepts and Technology behind Search, Second Edition, ACM Press Books
- 2 C. Manning, P. Raghavan, and H. Schütze, ―Introduction to Information Retrieval, Cambridge University Press
- 1 Ricci, F, Rokach, L. Shapira, B. Kantor, ―Recommender Systems Handbook‖, First Edition.
- 2 Bruce Croft, Donald Metzler, and Trevor Strohman, Search Engines: Information Retrieval in Practice, Pearson Education.
- 3 Stefan Buttcher, Charlie Clarke, Gordon Cormack, Information Retrieval: Implementing and Evaluating Search Engines, MIT Press.
Reproduced from the University of Mumbai syllabus for B.Sc. (Computer Science) under NEP 2020, in force from the academic year 2026-27. Wording is as printed in that syllabus. Module numbering is as printed there too.