munotes®

B.Sc. (Data Science) Large Language Models Syllabus - Mumbai University

This is the TY BSc Data Science syllabus under NEP 2020, in force from the academic year 2026-27. The University still sets the earlier Choice Based papers alongside it for ATKT candidates, so check which scheme your exam form names before you revise.

Large Language Models Syllabus.pdf
Major · Semester 6 · TY BSc Data Science · 2 credits · 50 marks

Loading syllabus...

Syllabus for Large Language Models

Major · Semester 6 · TY BSc Data Science · 2 credits · 50 marks

Module I: Foundations of Large Language Models

  • Introduction to Large Language Models
  • What is Language AI and Natural Language Processing evolution
  • Historical development of Language AI and the emergence of Generative AI
  • Representing language: Bag-of-Words and dense vector embeddings
  • Types of embeddings and contextual representations
  • Attention mechanism and transformer foundations
  • Encoder-only vs Decoder-only models
  • Training paradigm of Large Language Models
  • Applications of LLMs and responsible development
  • Interfacing with LLMs and generating initial outputs Tokens and Embeddings
  • LLM tokenization concepts
  • Preparing input text for language models
  • Tokenization methods: word, subword, character, and byte tokens
  • Comparison of trained LLM tokenizers
  • Token embeddings and contextual word embeddings
  • Sentence and document embeddings
  • Word embeddings beyond LLMs
  • Word2Vec and contrastive training approaches
  • Embeddings in recommendation systems Architecture of Large Language Models
  • Overview of transformer architecture
  • Inputs and outputs of transformer-based LLMs
  • Components of the forward pass
  • Sampling and decoding strategies
  • Parallel token processing and context windows
  • Key–value caching for faster generation
  • Internal structure of transformer blocks
  • Positional embeddings (RoPE) and architectural improvements Prompt Engineering
  • Using text generation models
  • Model selection and loading techniques
  • Controlling model output
  • Fundamentals of prompt engineering
  • Structure and components of effective prompts
  • Instruction-based prompting
  • Advanced prompting strategies
  • In-context learning and few-shot prompting
  • Chain-of-thought reasoning
  • Self-consistency and tree-of-thought prompting
  • Output verification and constrained generation

Module II: LLM Applications and Fine-Tuning

  • Advanced Text Generation and LLM Systems
  • Model I/O and loading quantized models using frameworks
  • Chains and prompt templates for LLM workflows
  • Multi-prompt chains and pipeline design
  • Conversation memory in LLM applications
  • Buffer memory and windowed conversation memory
  • Conversation summarization methods
  • LLM agents and multi-step reasoning systems
  • ReAct-based agent workflows Semantic Search and Retrieval-Augmented Generation
  • Overview of semantic search and retrieval systems
  • Semantic search using language models
  • Dense retrieval techniques
  • Reranking strategies for search results
  • Retrieval evaluation metrics
  • Retrieval-Augmented Generation (RAG) architecture
  • Transition from traditional search to RAG
  • Grounded generation with LLM APIs
  • RAG implementation with local models
  • Advanced RAG techniques and evaluation Fine-Tuning Representation Models
  • Supervised classification with transformer models
  • Fine-tuning pretrained BERT models
  • Freezing layers and transfer learning strategies
  • Few-shot classification methods
  • SetFit framework for efficient fine-tuning
  • Continued pretraining with masked language modeling
  • Named Entity Recognition tasks
  • Dataset preparation for NER
  • Fine-tuning models for NER applications Fine-Tuning Generation Models
  • LLM training pipeline: pretraining, supervised fine-tuning, preference tuning
  • Supervised fine-tuning (SFT) methods
  • Full fine-tuning vs parameter-efficient fine-tuning (PEFT)
  • Instruction tuning using QLoRA
  • Instruction dataset preparation and templating
  • Model quantization techniques
  • LoRA configuration and training setup
  • Training and merging model weights
  • Evaluation of generative models
  • Word-level metrics and benchmark evaluation
  • Human and automated evaluation approaches
  • Alignment techniques and RLHF
  • Reward models for preference learning
  • Direct Preference Optimization (DPO) 10 Text Books 1. Hands-On Large Language Models Language Understanding and Generation by Jay Alammar and Maarten Grootendorst, Oreilly Publication , Edition 01 , 2024 2. LLM Engineer’s Handbook by Paul Iusztin and Maxime Labonne, Pakt Publications, 2024 11 Reference Books 1. Natural Language Processing with Transformers Building Language Applications with Hugging Face, Oreilly, 2022 2. Building LLMs for Production by LOUIS-FRANÇOIS BOUCHARD et.al. 2024 12 Internal Continuous Assessment: 40% Semester End Examination: 60% 13 Continuous Evaluation through: 30 marks Semester End Examination Case study submission of 10 marks Quizzes/ Presentations/ Assignments: 10 marks Total: 20 marks 14 Format of Question Paper: (Semester End Examination: 30 Marks. Duration: 1 Hr ) Q1: Attempt any three (out of five/six) from Module 1 (15 Marks) Q2: Attempt any three (out of five/six) from Module 1 (15 Marks)

Text Books

  • 1 Hands-On Large Language Models Language Understanding and Generation by Jay Alammar and Maarten Grootendorst, Oreilly Publication , Edition 01 , 2024
  • 2 LLM Engineer’s Handbook by Paul Iusztin and Maxime Labonne, Pakt Publications, 2024
  • 1 Natural Language Processing with Transformers Building Language Applications with Hugging Face, Oreilly, 2022
  • 2 Building LLMs for Production by LOUIS-FRANÇOIS BOUCHARD et.al. 2024

Reproduced from the University of Mumbai syllabus for B.Sc. (Data Science) under NEP 2020, in force from the academic year 2026-27. Wording is as printed in that syllabus. Module numbering is as printed there too.

The complete syllabus

This subject is cut from the University circular for its year. Open a document here if you want the whole thing rather than a single subject.

PDF 2024 25 DS SEM I & II NEP NEP 2020 syllabus, in force from 2024-25 Read full PDF Read
PDF 2023 24 BSc Data Science Sem V & VI Earlier Choice Based syllabus, still set for ATKT candidates Read full PDF Read
PDF 2021 22 BSc Data Science Sem III & IV Earlier Choice Based syllabus, still set for ATKT candidates Read full PDF Read
Report or request
Done!