munotes®

B.Sc. (Computer Science) Information Retrieval Practical Syllabus - Mumbai University

This is the TY BSc Computer Science syllabus under NEP 2020, in force from the academic year 2026-27. The University still sets the earlier Choice Based papers alongside it — her Summer 2026 third-year timetables name that scheme — so check which scheme your exam form names before you revise.

Information Retrieval Practical Syllabus.pdf
Major elective · Semester 6 · TY BSc Computer Science · 2 credits · 50 marks

Loading syllabus...

Syllabus for Information Retrieval Practical

Major elective · Semester 6 · TY BSc Computer Science · 2 credits · 50 marks

This is the practical paper for Information Retrieval, which the University sets and examines separately.

Module I

  • Text Preprocessing on Cranfield Collection
  • Download the Cranfield document collection (1400 aeronautical abstracts).
  • Perform: Tokenization, Stopword removal (NLTK stopword list), Porter stemming, Compute vocabulary size before and after preprocessing, Report top 20 frequent terms
  • Inverted Index Construction for Cranfield Corpus Using the preprocessed Cranfield corpus build an inverted index, Store: term →
  • {docID, term frequency}, Compute document frequency for each term, Print posting list for terms: “flow”, “boundary”, “shock”
  • Positional Index & Phrase Query Processing Extend inverted index to store term positions. Process phrase queries: “boundary layer flow”, “supersonic shock wave”, Return matching document IDs.
  • Boolean Retrieval using Cranfield Queries
  • Using Cranfield query set to Implement Boolean retrieval, Process query: (boundary AND flow) AND NOT laminar, Return list of document IDs, Compare with relevance judgments file
  • TF-IDF Weight Computation on Cranfield
  • Compute TF, IDF, TF-IDF weights, Display top 10 highest weighted terms in Document ID 12 and Document ID 245. Dataset: Cranfield Collection
  • Ranked Retrieval using Vector Space Model
  • For Cranfield Query 1: Convert query and documents to TF-IDF vectors, Compute cosine similarity, Rank top 10 documents, Compare retrieved results with relevance judgments
  • Evaluation of Boolean vs Vector Space Model
  • For first 5 Cranfield queries: Retrieve top 10 results using Boolean model, Retrieve top 10 results using VSM, Compute Precision@10, Compare retrieval effectiveness
  • Edit Distance Implementation for Query Correction
  • Implement Levenshtein distance. Correct the following misspelled Cranfield query terms: “supersonc”, “aerodynamcs”, “turbulnce”. Suggest top 3 corrections from corpus vocabulary. Dataset: Cranfield vocabulary
  • Integrated Spelling Correction in IR System
  • Modify retrieval system so that when user enters: supersonc boundary layr. System automatically: Suggests corrected query, Performs retrieval using corrected query, Dataset: Cranfield Collection
  • Precision, Recall and F1 on Cranfield Query 1
  • Using Cranfield relevance judgments: Retrieve top 20 documents for Query 1
  • Compute Precision, Recall, F1-score, Dataset: Cranfield Collection + qrels

Module II

  • Average Precision and MAP
  • For first 10 Cranfield queries: Compute Average Precision (AP) for each query, Compute Mean Average Precision (MAP), Compare with baseline Boolean model, Dataset: Cranfield Collection
  • Precision-Recall Curve Visualization
  • For Query 3 Compute precision at each recall level, Plot Precision–Recall curve, Compare Boolean vs VSM, Dataset: Cranfield Collection
  • Naïve Bayes Classification using 20 Newsgroups
  • Using 4 categories from 20 Newsgroups dataset: sci.space, comp.graphics, rec.sport.baseball, talk.politics.misc, Train Multinomial Naïve Bayes classifier and compute Accuracy, Confusion matrix, Macro F1-score, Dataset: 20 Newsgroups Dataset
  • SVM Classification using 20 Newsgroups
  • Using 20 Newsgroups dataset Train linear SVM, Compare performance with Naïve Bayes, Report Precision, Recall, F1, Dataset: 20 Newsgroups
  • K-Means Clustering on BBC News Dataset
  • Cluster BBC News dataset into 5 clusters. Analyze Top 10 terms per cluster, Compare clusters with actual categories, Dataset: BBC News Dataset
  • Hierarchical Clustering on BBC Dataset
  • Perform Agglomerative clustering. Generate dendrogram and Determine optimal number of clusters, Compare with K-Means, Dataset: BBC News Dataset
  • Clustering Evaluation (Purity & Silhouette)
  • For the above 2 practicals, Compute Cluster Purity, Silhouette Score For K = 3, 4, 5, Compare clustering quality., Dataset: BBC News Dataset
  • Web Crawling & Link Analysis
  • Web Crawling of a Website: Develop a crawler starting from a given URL, Crawl up to 50 pages, Extract page titles, Build inverted index, Respect robots.txt
  • PageRank on Sample Web Graph
  • Given web graph with 10 nodes Construct adjacency matrix, Implement iterative PageRank, Use damping factor = 0.85, Show convergence after 20 iterations Dataset:
  • SNAP small web graph dataset (Stanford Network Analysis Project)
  • Learning to Rank using LETOR Dataset
  • Using LETOR 4.0 dataset Train RankSVM model, Evaluate using NDCG@10,
  • Compare with TF-IDF baseline, Dataset: LETOR 4.0

Text Books

  • 1 Ricardo Baeza-Yates and Berthier Ribeiro-Neto, ―Modern Information Retrieval: The Concepts and Technology behind Search, Second Edition, ACM Press Books
  • 2 C. Manning, P. Raghavan, and H. Schütze, ―Introduction to Information Retrieval, Cambridge University Press
  • 1 Ricci, F, Rokach, L. Shapira, B. Kantor, ―Recommender Systems Handbook‖, First Edition.
  • 2 Bruce Croft, Donald Metzler, and Trevor Strohman, Search Engines: Information Retrieval in Practice, Pearson Education.
  • 3 Stefan Buttcher, Charlie Clarke, Gordon Cormack, Information Retrieval: Implementing and Evaluating Search Engines, MIT Press.

Reproduced from the University of Mumbai syllabus for B.Sc. (Computer Science) under NEP 2020, in force from the academic year 2026-27. Wording is as printed in that syllabus. Module numbering is as printed there too.

Report or request
Done!