B.Sc. (Computer Science) Information Retrieval Syllabus - Mumbai University
This is the TY BSc Computer Science syllabus under NEP 2020, in force from the academic year 2026-27. The University still sets the earlier Choice Based papers alongside it — her Summer 2026 third-year timetables name that scheme — so check which scheme your exam form names before you revise.
Loading syllabus...
Syllabus for Information Retrieval
The University sets the practical for this subject separately, in Information Retrieval Practical. It carries its own credits, so it is examined as a paper of its own.
Module I
- Introduction to Information Retrieval (IR) systems: Definition and goals of information retrieval, Components of an IR system, Challenges and applications of IR
- Document Indexing, Storage, and Compression: Inverted index construction and compression techniques, Document representation and term weighting, Storage and retrieval of indexed documents
- Retrieval Models: Boolean model: Boolean operators, query processing, Vector space model: TF-IDF, cosine similarity, query-document matching, Probabilistic model: Bayesian retrieval, relevance feedback
- Spelling Correction in IR Systems: Challenges of spelling errors in queries and documents, Edit distance and string similarity measures, Techniques for spelling correction in IR systems
- Performance Evaluation: Evaluation metrics: precision, recall, F-measure, average precision, Test collections and relevance judgments, Experimental design and significance testing
- Text Categorization and Filtering: Text classification algorithms: Naive Bayes, Support Vector Machines, Feature selection and dimensionality reduction, Applications of text categorization and filtering
- Text Clustering for Information Retrieval: Clustering techniques: K-means, hierarchical clustering, Evaluation of clustering results, Clustering for query expansion and result grouping
Module II
- Web Information Retrieval: Web search architecture and challenges, Crawling and indexing web pages, Link analysis and PageRank algorithm
- Link Analysis and its Role in IR Systems: Web graph representation and link analysis algorithms, HITS and PageRank algorithms, Applications of link analysis in IR systems
- Crawling and Near-Duplicate Page Detection: Web page crawling techniques: breadth-first, depth-first, focused crawling, Near-duplicate page detection algorithms, Handling dynamic web content during crawling
- Learning to Rank: Algorithms and Techniques, Supervised learning for ranking: RankSVM, RankBoost, Pairwise and listwise learning to rank approaches Evaluation metrics for learning to rank
- Advanced Topics in IR: Text Summarization: extractive and abstractive methods; Question Answering: approaches for finding precise answers; Recommender Systems: collaborative filtering, content-based filtering
- Cross-Lingual and Multilingual Retrieval: Challenges and techniques for cross- lingual retrieval, Machine translation for IR, Multilingual document representations and query translation, Evaluation Techniques for IR Systems
- User-based evaluation: user studies, surveys, Test collections and benchmarking, Online evaluation methods: A/B testing, interleaving experiments
Text Books
- 1 Ricardo Baeza-Yates and Berthier Ribeiro-Neto, ―Modern Information Retrieval: The Concepts and Technology behind Search, Second Edition, ACM Press Books
- 2 C. Manning, P. Raghavan, and H. Schütze, ―Introduction to Information Retrieval, Cambridge University Press
- 1 Ricci, F, Rokach, L. Shapira, B. Kantor, ―Recommender Systems Handbook‖, First Edition.
- 2 Bruce Croft, Donald Metzler, and Trevor Strohman, Search Engines: Information Retrieval in Practice, Pearson Education.
- 3 Stefan Buttcher, Charlie Clarke, Gordon Cormack, Information Retrieval: Implementing and Evaluating Search Engines, MIT Press.
Reproduced from the University of Mumbai syllabus for B.Sc. (Computer Science) under NEP 2020, in force from the academic year 2026-27. Wording is as printed in that syllabus. Module numbering is as printed there too.