munotes®

B.E. (Artificial Intelligence and Data Science) Big Data Analytics Syllabus - Mumbai University

This is the Fourth Year BE AI and DS syllabus under CBCS REV-2019 'C' Scheme, in force from the academic year 2023-24. The University has published no NEP 2020 syllabus for Semesters V to VIII of any engineering branch, so this is the scheme you are examined on — exam form 1T01817 and 1T01818. The first and second years of the degree are on NEP 2020.

Big Data Analytics.pdf
Semester 7 · Fourth Year BE AI and DS · 3 credits · CBCS REV-2019 'C' Scheme

Loading syllabus...

Syllabus for Big Data Analytics

Semester 7 · Fourth Year BE AI and DS · 3 credits · CBCS REV-2019 'C' Scheme

Module 01 04 hours

  • Introduction to Big Data & Hadoop 1.1 Introduction to Big Data, 1.2 Big Data characteristics, types of Big Data, 1.3 Traditional vs. Big Data business approach, 1.4 Case Study of Big Data Solutions. 1.5 Concept of Hadoop 1.6 Core Hadoop Components; Hadoop Ecosystem

Module 02 07 hours

  • Hadoop HDFS and Map Reduce 2.1 Distributed File Systems: Physical Organization of Compute Nodes, Large-Scale File-System Organization. 2.2 MapReduce: The Map Tasks, Grouping by Key, The Reduce Tasks, Combiners, Details of MapReduce Execution, Coping With Node Failures. 2.3 Algorithms Using MapReduce: Matrix-Vector Multiplication by MapReduce, Relational-Algebra Operations, Computing Selections by MapReduce, Computing Projections by MapReduce, Union, Intersection, and Difference by MapReduce 2.4 Hadoop Limitations s.

Module 03 05 hours

  • NoSQL 3.1 Introduction to NoSQL, NoSQL Business Drivers, 3.2 NoSQL Data Architecture Patterns: Key-value stores, Graph stores, Column family (Bigtable)stores, Document stores, Variations of NoSQL architectural patterns, NoSQL Case Study 3.3 NoSQL solution for big data, Understanding the types of big data problems; Analyzing big data with a shared-nothing architecture; Choosing distribution models: master-slave versus peer-to-peer; NoSQL systems to handle big data problems. peer-to-peer; Four ways that NoSQL systems handle big data problems

Module 04 09 hours

  • Mining Data Streams 4.1 The Stream Data Model: A Data-Stream-Management System, Examples of Stream Sources, Stream Queries, Issues in Stream Processing. 4.2 Sampling Data techniques in a Stream 4.3 Filtering Streams: Bloom Filter with Analysis. 4.4 Counting Distinct Elements in a Stream, Count-Distinct Problem, Flajolet-Martin Algorithm, Combining Estimates, Space Requirements 4.5 Counting Frequent Items in a Stream, Sampling Methods for Streams, Frequent Itemsets in Decaying Windows. 4.6 Counting Ones in a Window: The Cost of Exact Counts, The Datar-Gionis-Indyk-Motwani Algorithm, Query Answering in the DGIM Algorithm, Decaying Windows.

Module 05 06 hours

  • Finding Similar Items and Clustering 5.1 Distance Measures: Definition of a Distance Measure, Euclidean Distances, Jaccard Distance, Cosine Distance, Edit Distance, Hamming Distance. 5.2 CURE Algorithm, Stream-Computing , A Stream-Clustering Algorithm, Initializing & Merging Buckets, Answering Queries.

Module 06 08 hours

  • Real-Time Big Data Models 6.1 PageRank Overview, Efficient computation of PageRank: PageRank Iteration Using MapReduce, Use of Combiners to Consolidate the Result Vector. 6.2 A Model for Recommendation Systems, Content-Based Recommendations, Collaborative Filtering. 6.3 Social
  • Networks as Graphs, Clustering of Social-Network Graphs, Direct Discovery of Communities in a social graph.

Text Books

  • 1 Anand Rajaraman and Jeff Ullman "Mining of Massive Datasets", Cambridge University Press,
  • 2 Alex Holmes "Hadoop in Practice", Manning Press, Dreamtech Press.
  • 3 Dan Mcary and Ann Kelly "Making Sense of NoSQL" – A guide for managers and the rest of us, Manning Press.

References

  • 1 Bill Franks , "Taming The Big Data Tidal Wave: Finding Opportunities In Huge Data Streams With Advanced Analytics", Wiley
  • 2 Chuck Lam, "Hadoop in Action", Dreamtech Press
  • 3 Jared Dean, "Big Data, Data Mining, and Machine Learning: Value Creation for Business Leaders and Practitioners", Wiley India Private Limited, 2014.
  • 4 Jiawei Han and Micheline Kamber, "Data Mining: Concepts and Techniques", Morgan Kaufmann Publishers, 3rd ed, 2010.
  • 5 Lior Rokach and Oded Maimon, "Data Mining and Knowledge Discovery Handbook", Springer, 2nd edition, 2010.
  • 6 Ronen Feldman and James Sanger, "The Text Mining Handbook: Advanced Approaches in Analyzing Unstructured Data", Cambridge University Press, 2006.
  • 7 Vojislav Kecman, "Learning and Soft Computing", MIT Press, 2010

Reproduced from the University of Mumbai syllabus for B.E. (Artificial Intelligence and Data Science), item 6.12 (N), under CBCS REV-2019 'C' Scheme, in force from the academic year 2023-24. Wording, module numbering and hours are as printed in that syllabus.

The complete syllabus

This subject is cut from the University circular for its year. Open a document here if you want the whole thing rather than a single subject.

PDF B.E. Artificial Intelligence and Data Science - First Year, Semester I and II - NEP 2020 - Item 7.7 (R-A) NEP 2020, Semesters I and II, in force from 2024-25 Read full PDF Read
PDF B.E. Artificial Intelligence and Data Science - Second Year, Semester III and IV - NEP 2020 - Item 6.20 (N) NEP 2020, Semesters III and IV, in force from 2025-26 Read full PDF Read
PDF B.E. Artificial Intelligence and Data Science - Third Year, Semester V and VI - CBCS REV-2019 C Scheme - Item 6.42 (R) CBCS REV-2019 'C' Scheme, Semesters V and VI, in force from 2022-23 Read full PDF Read
PDF B.E. Artificial Intelligence and Data Science - Fourth Year, Semester VII and VIII - CBCS REV-2019 C Scheme - Item 6.12 (N) CBCS REV-2019 'C' Scheme, Semesters VII and VIII, in force from 2023-24 Read full PDF Read
Report or request
Done!