B.E. (Artificial Intelligence and Data Science) Natural Language Processing Syllabus - Mumbai University
This is the Fourth Year BE AI and DS syllabus under CBCS REV-2019 'C' Scheme, in force from the academic year 2023-24. The University has published no NEP 2020 syllabus for Semesters V to VIII of any engineering branch, so this is the scheme you are examined on — exam form 1T01817 and 1T01818. The first and second years of the degree are on NEP 2020.
Loading syllabus...
Syllabus for Natural Language Processing
Module 1 4 hours
- Introduction
- 1.1 Origin & History of NLP, The need of NLP, Generic NLP System, Levels of NLP, Knowledge in Language Processing, Ambiguity in Natural Language, Challenges of NLP, Applications of NLP.
Module 2 8 hours
- Word Level Analysis
- 2.1 Tokenization, Stemming, Segmentation, Lemmatization, Edit Distance, Collocations, Finite Automata, Finite State Transducers (FST), Porter
- Stemmer, Morphological Analysis, Derivational and Reflectional Morphology, Regular expression with types.
- 2.2 N –Grams, Unigrams/Bigrams Language Models, Corpora, Computing the Probability of Word Sequence, Training and Testing.
Module 3 8 hours
- Syntax analysis
- 3.1 Part-Of-Speech Tagging (POS) - Open and Closed Words. Tag Set for English (Penn Treebank), Rule Based POS Tagging, Transformation Based Tagging, Stochastic POS Tagging and Issues –Multiple Tags & Words, Unknown Words.
- 3.2 Introduction to CFG, Hidden Markov Model (HMM), Maximum Entropy, And Conditional Random Field (CRF).
Module 4 8 hours
- Semantic Analysis
- 4.1 Introduction, meaning representation; Lexical Semantics; Corpus study; Study of Various language dictionaries like WordNet, Babelnet; Relations among lexemes & their senses –Homonymy, Polysemy, Synonymy, Hyponymy; Semantic Ambiguity
- 4.2 Word Sense Disambiguation (WSD); Knowledge based approach (Lesk's Algorithm), Supervised (Naïve Bayes, Decision List), Introduction to Semi-supervised method (Yarowsky), Unsupervised (Hyperlex)
Module 5 6 hours
- Pragmatic & Discourse Processing
- 5.1 Discourse: Reference Resolution, Reference Phenomena, Syntactic & Semantic constraint on coherence; Anaphora Resolution using Hobbs and Cantering Algorithm
Module 6 5 hours
- Applications (preferably for Indian regional languages)
- 6.1 Machine Translation, Information Retrieval, Question Answers System, Categorization, Summarization, Sentiment Analysis, Named Entity Recognition.
- 6.2 Linguistic Modeling – Neurolinguistics Models- Psycholinguistic Models – Functional Models of Language – Research Linguistic Models- Common Features of Modern Models of Language.
Text Books
- 1 Daniel Jurafsky, James H. and Martin, Speech and Language Processing, Second Edition, Prentice Hall, 2008.
- 2 Christopher D.Manning and HinrichSchutze, Foundations of Statistical Natural Language Processing, MIT Press, 1999.
References
- 1 Siddiqui and Tiwary U.S., Natural Language Processing and Information Retrieval, Oxford University Press, 2008.
- 2 Daniel M Bikel and ImedZitouni " Multilingual natural language processing applications: from theory to practice, IBM Press, 2013.
- 3 Nitin Indurkhya and Fred J. Damerau, "Handbook of Natural Language Processing, Second Edition, Chapman and Hall/CRC Press, 2010.
Useful Links
- 1 https://onlinecourses.nptel.ac.in/noc21_cs102/preview
- 2 https://onlinecourses.nptel.ac.in/noc20_cs87/preview
- 3 https://nptel.ac.in/courses/106105158
Reproduced from the University of Mumbai syllabus for B.E. (Artificial Intelligence and Data Science), item 6.12 (N), under CBCS REV-2019 'C' Scheme, in force from the academic year 2023-24. Wording, module numbering and hours are as printed in that syllabus.
The complete syllabus
This subject is cut from the University circular for its year. Open a document here if you want the whole thing rather than a single subject.