munotes®

IKS in Computational Systems Notes | B.Sc. (Computer Science) Semester 5 | Mumbai University | munotes

Get access to whole semester resourcesSemester Pass

Official Notes munotes.in

IKS in Computational Systems

B.SC. (COMPUTER SCIENCE) · SEMESTER 5

Strictly as per the University of Mumbai NEP syllabus in force for B.Sc. (Computer Science)

For B.Sc. (Computer Science) students of the University of Mumbai and all its affiliated colleges

Open the book ↓

munotes.in Third Year

IKS in Computational Systems

Copyright © 2026 munotes.in. All rights reserved.

Written and first published by munotes.in, 2026.

This book is free for individual students to read at munotes.in. No part of it may be reproduced, distributed, stored, translated or used for institutional or classroom purposes in any form without a prior written licence from munotes.in.

Licensing and permissions: contact@munotes.in

The text of statutes and of judgments reproduced in this book is in the public domain under section 52(1)(q) of the Copyright Act 1957. The commentary, arrangement, examples and questions are the original work of munotes.in.

munotes.in is an independent study resource for MU students. It is not affiliated with, endorsed by, or officially connected to the University of Mumbai. Course names and university references describe the students and syllabus the material relates to.

munotes.in

Contents

Module I Sastra methodology and knowledge formalization, Pingala's Chandahsastra as binary encoding and combinatorial generation, and Panini's Astadhyayi as a generative formal grammar

  1. Why a Computer Science Student Reads a Śāstra 1
  2. How This Paper Is Examined, and What That Means for How You Read 4
  3. Śāstra: What Makes a Body of Knowledge Formal 7
  4. The Sūtra Method as Compressed Symbolic Encoding 10
  5. Pramāṇa: Where Knowledge Is Allowed To Come From 13
  6. Pratyakṣa: Perception, and Why It Is Defined So Narrowly 16
  7. Anumāna: Inference in Outline 19
  8. Āgama: Testimony, and the Trusted Source as a Design Decision 22
  9. Knowledge Validation and Structured Reasoning 26
  10. Knowledge Representation: The Modern Name For It 29
  11. Formal Specification: Saying Exactly What a System Must Do 32
  12. Computational Thinking, and the Four Habits It Names 35
  13. Piṅgala's Chandaḥśāstra, and the Text We Are Reading 38
  14. Syllables, Laghu and Guru: Sanskrit Metre Without Sanskrit 41
  15. Laghu and Guru as One Bit 44
  16. Prastāra: The Table of Every Pattern 47
  17. The Prastāra Rule Read as an Algorithm 51
  18. Prastāra in Code 55
  19. Binary Number Systems, and Exactly Where Piṅgala's Order Agrees 59
  20. Saṅkhyā: Counting the Rows Without Writing Them 63
  21. Binary Exponentiation: The Same Algorithm in a Modern Textbook 67
  22. Uddiṣṭa: From a Pattern to Its Row Number 70
  23. Naṣṭa: From a Row Number Back to the Pattern 73
  24. Index Retrieval, and Proving Two Rules Are Inverses 77
  25. Recursive Enumeration 80
  26. Tree Structures, and the Prastāra as a Binary Tree 83
  27. Meru-Prastāra: Halāyudha's Staircase 87
  28. Pascal Triangle and Combinatorics 91
  29. Meru-Prastāra as a Dynamic Programming Model 95
  30. Lagakriyā and Adhvayoga: The Rest of the Pratyayas 99
  31. Algorithmic Generation: What Piṅgala Actually Achieved 102
  32. Pāṇini's Aṣṭādhyāyī, and the Text We Are Reading 105
  33. The Six Kinds of Sūtra, in Vasu's Own Words 108
  34. The Fourteen Śivasūtras and the Pratyāhāra 112
  35. Anubandha: The Marker Letter as a Type Tag 116
  36. Rule-Based Generative Structure 119
  37. Meta-Rules: Paribhāṣā, and Rules About Rules 122
  38. Rule Precedence: The Four Principles 125
  39. Vipratiṣedha: When Two Rules Collide, the Later Wins 128
  40. Context-Sensitive Operations 131
  41. Asiddhatva and the Tripādī: Ordering by Blocking 134
  42. Conflict Resolution Mechanisms, Collected 137
  43. Formal Grammars: Alphabet, Rule, Derivation, Language 140
  44. Context-Free Grammar, and Whether Pāṇini Wrote One 143
  45. The Chomsky Hierarchy, and Where the Aṣṭādhyāyī Sits 147
  46. Rewrite Systems 150
  47. Automata Theory Foundations: The Finite Automaton 154
  48. Śāstra Rule Precedence as a Deterministic Finite Rewrite System 158
  49. Parsing Algorithms 164
  50. Foundations of NLP, and What Pāṇinian Grammar Contributed 168
  51. Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine 172
  52. Practice for Module I 176

Module II Nyaya logic as structured inference, Ayurvedic classification as a rule-based expert system, and Arthasastra cryptography as a secure communication model

  1. Nyāya: The School, the Sūtra and the Sixteen Categories 181
  2. The Four Pramāṇas of Nyāya, and Why Charaka Has Three 184
  3. Anumāna: Vyāpti, and the Three Kinds of Inference 187
  4. The Five-Member Syllogism 191
  5. The Five Members Worked, Three Times 195
  6. Five Members Against Aristotle's Three 198
  7. Hetvābhāsa: The Five Fallacies of the Reason 201
  8. Chala, Jāti and Nigrahasthāna: How a Debate Is Lost 205
  9. Vāda, Jalpa and Vitaṇḍā: Three Kinds of Dispute 209
  10. Debate Methodology as a Validation Protocol 213
  11. Padārtha: The Categories of What Exists 217
  12. Padārtha Ontology as a Knowledge Representation Model 221
  13. Propositional Logic: The Minimum You Need 226
  14. Predicate Logic, and Why Anumāna Needs It 231
  15. The Five Members Written in Logical Notation 234
  16. Inference Engines: Forward and Backward Chaining 237
  17. Nyāya Logic as an Inference Engine 241
  18. Explainable AI, and Why a Five-Member Answer Is an Explanation 246
  19. Āyurveda as a Śāstra, and What This Chapter Does Not Claim 250
  20. The Tridoṣa Framework 254
  21. Doṣa as a Feature Vector 258
  22. Prakṛti: Constitution as a Class Label 262
  23. Parīkṣā: How an Examination Is Structured 265
  24. Symptom to Feature Mapping 269
  25. Multi-Attribute Classification 273
  26. Decision Principles in Diagnosis 276
  27. Decision Trees 280
  28. A Decision Tree Built From the Tridoṣa Attributes 285
  29. Rule-Based Systems 289
  30. Expert Systems, and MYCIN as the Comparison 293
  31. Feature Engineering 297
  32. Multi-Class Classification, and How It Is Scored 300
  33. Ayurvedic Classification as a Rule-Based Expert System 303
  34. The Arthaśāstra, and Its Intelligence Apparatus 310
  35. Secret Communication: What the Text Actually Says 313
  36. The Royal Writs as a Message Format 316
  37. Substitution Systems 319
  38. Transposition Systems 323
  39. Substitution and Transposition in Code 328
  40. Concealment and Coded Messaging 335
  41. Basic Steganography 338
  42. Information Protection Mechanisms 342
  43. Symmetric Encryption 345
  44. Cipher Algorithms, Classical to Modern 348
  45. Breaking a Classical Cipher 351
  46. Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System 356
  47. Secure Protocol Abstraction 360
  48. Foundations of Cybersecurity 363
  49. The Internal Assessment: Building and Presenting the Implementation 366
  50. Practice for Module II 370
  51. The Whole Paper on One Page 374
munotes.in

Module I

Sastra methodology and knowledge formalization, Pingala's Chandahsastra as binary encoding and combinatorial generation, and Panini's Astadhyayi as a generative formal grammar

munotes.in

Chapter One

Why a Computer Science Student Reads a Śāstra

Syllabus topic Module 1, "Concept of śāstra as a formal knowledge system", "aligned with computational thinking"

In one line

This paper asks whether the structured disciplines of classical India were doing computer science before there were computers, and the answer turns out to be yes in five specific, checkable ways and no in several others.

In the wording you can write in an examination: the paper studies six classical Indian knowledge systems as formal, rule-governed, algorithmic constructions, and maps each onto the modern computer science concept it corresponds to.

The honest first question

Why is a Sanskrit treatise on a Computer Science paper?

Not for reverence. The University's own description of this course says what the reason is: these traditions "embody structured, rule-based, and algorithmic principles analogous to contemporary Computer Science concepts". The word doing the work is analogous, and an analogy is a claim that can be true or false.

So this book treats every comparison as a claim to be tested. Where a classical rule really is an algorithm, we run it and show the output. Where it only resembles one, we say so.

The five things that are actually there

Each of these is established later in the book from the text itself, not asserted. They are collected here so you know where the paper is going.

Classical sourceWhat it containsThe modern name
Piṅgala's Chandaḥśāstraa two-valued symbol per syllable, and a rule that enumerates every pattern of n syllables in a fixed orderbinary encoding, and exhaustive enumeration
Piṅgala, againa rule that finds the pattern at row r without writing the table, and a rule that finds r from the patternindex retrieval into an unwritten structure
Piṅgala, againa method that computes 2 to the power n by repeated squaring and doublingbinary exponentiation
Pāṇini's Aṣṭādhyāyīabout four thousand rules, six declared rule types, and four stated principles for deciding which rule fires when two coulda rule-based generative system with meta-rules and conflict resolution
Nyāya's five-member forma fixed shape in which an inference must be stated before it counts as establishedan inference engine that emits its own proof

And the three things that are not

A book that only lists the first table is not teaching you the subject, it is selling you something. These are the limits, and a question that asks you to "critically assess" wants them.

There is no machine. Every one of these systems is executed by a trained human being. An algorithm and an implementation are different things, and the gap between them is most of the history of computing.

There is no cipher in the Arthaśāstra. Kautilya names secret writing four times and never states a method. What the treatise gives is a requirement, not an algorithm, and Module II says so in as many words.

munotes.in1

Why a Computer Science Student Reads a Śāstra

Ayurvedic classification is not a validated classifier. It is a structured diagnostic vocabulary. Treating it as a tested medical model is a claim nobody in this book makes.

What "computational thinking" means here

MU's first course objective asks you to see śāstra as a knowledge system "aligned with computational thinking". That phrase has four ordinary parts, and you will use all four in this paper.

Decomposition. Break a problem into parts that can be solved separately. Piṅgala breaks "list every metre" into "list every metre of one syllable" plus a rule for growing the list.

Pattern recognition. Notice that two different problems have the same shape. The count of metres with exactly three long syllables and the count of ways to choose three items from a set are the same number, and the Meru-prastāra computes both.

Abstraction. Throw away what does not matter. A syllable has a sound, a meaning and a length; Piṅgala keeps only the length, and that is why his rules are arithmetic.

Algorithm design. State the steps so exactly that someone who does not understand the problem can still carry them out. This is the test that separates a description from an algorithm, and chapter seventeen applies it to Piṅgala's own words.

Worked example, before anything else

Here is the smallest complete instance of the whole paper, so that the rest of the book has something to attach to.

The problem. A Sanskrit metre of three syllables. Each syllable is either short (laghu, written L) or long (guru, written G). List every possible metre.

The classical rule, from Piṅgala's Chandaḥśāstra with Halāyudha's commentary: write a row of all long syllables. Then, to get each next row, find the first long syllable from the left, make it short, make everything to its left long, and leave everything to its right alone. Stop when the row is all short.

Carried out by hand:

GGG

LGG

GLG

LLG

GGL

LGL

GLL

LLL

The observation that makes this a computer science paper. Write 1 for L and 0 for G, and read each row with the leftmost syllable as the units place: 000, 100, 010, 110, 001, 101, 011, 111. Those are 0, 1, 2, 3, 4, 5, 6, 7 in binary. The rule is counting in base two, and it is written down centuries before the positional notation it depends on reached Europe.

That claim is checked, not asserted. The program in chapter nineteen verifies it for every row of every metre up to fourteen syllables.

How to use this book

Read it in order. The chapters follow MU's printed module order exactly, so you can lay the syllabus page beside the contents and tick off labels.

munotes.in2

Why a Computer Science Student Reads a Śāstra

Each module's last chapters are practice. They carry worked answers of the length the examination actually wants, which is five marks.

Do the code. Twenty of the fifty marks on this paper are one program you write, run and defend out loud. Every listing in this book has been run, and you should run them too.

Quick revision

  • MU's course description calls these traditions "structured, rule-based, and algorithmic" and "analogous" to modern computer science concepts. Analogous is a testable claim.
  • Five real correspondences: binary encoding, enumeration, index retrieval, binary exponentiation, and rule-based generation with conflict resolution.
  • Three honest limits: no machine, no cipher in the Arthaśāstra, no validated classifier in Ayurveda.
  • Computational thinking has four parts: decomposition, pattern recognition, abstraction, algorithm design.
  • Laghu and guru, two values in one position, is the hinge of Module I.

Test yourself

1. MU's description uses one word that makes this paper falsifiable rather than decorative. Which word, and why does it matter?

"Analogous." An analogy asserts that two things do the same work. That can be shown or refuted, so every chapter has to argue its case instead of asserting a resemblance.

2. Give one correspondence between a classical Indian text and a modern computer science concept, and one place where the correspondence fails.

Piṅgala's prastāra enumerates every pattern of n two-valued symbols in the order of binary counting, which is exhaustive enumeration over an n-bit space. It fails as an account of computing because nothing executes it but a person: there is an algorithm and no machine.

3. Name the four parts of computational thinking and give the Piṅgala example of one.

Decomposition, pattern recognition, abstraction, algorithm design. Abstraction: a syllable's sound and meaning are discarded and only its length is kept, which is what allows the rules to be arithmetic.

4. Why does this book refuse to say the Arthaśāstra contains a cipher?

Because the text names cipher-writing and never states a method. Naming a requirement is not supplying an algorithm, and the difference is the whole distinction between a specification and an implementation.

Contents This chapter on its own page

munotes.in3

Chapter Two

How This Paper Is Examined, and What That Means for How You Read

Syllabus topic Module 1, "Algorithm Specification (Pseudo-code)", "Complexity & Limitations", "Conceptual Mapping Table", "minimum 10 test cases"

In one line

Fifty marks: thirty written in one hour, and twenty for one program you write, run and defend out loud.

In the wording you can use in a viva: the paper carries 20 internal marks earned by a single implementation project with a formal specification, working code, at least ten test cases and a correctness analysis, plus 30 external marks in a one-hour written paper of three questions, each offering four alternatives of which two are to be attempted.

The external paper, exactly as it is printed

QuestionBased onWhat it asksMarks
Q.1Module 1Answer any 2 of the following (any 2 out of 4)10
Q.2Module 2Answer any 2 of the following (any 2 out of 4)10
Q.3Modules 1 and 2Answer any 2 of the following (any 2 out of 4)10

Three consequences follow, and they change how you should read every chapter after this one.

Worked. An answer is worth five marks, not fifteen. Ten marks divided by two attempted parts is five. Five marks in a one-hour paper is a tight, complete answer: a definition, the mechanism, one example, and one limit. It is not an essay, and it is not a list of headings either.

You get to choose two of four. So breadth protects you. A student who has read eight of this book's chapters deeply and skipped forty will meet a question set in which all four alternatives come from the forty.

Q.3 is a cross-module question. It draws on Modules 1 and 2 together, which means it will ask you to relate something in Piṅgala, Pāṇini or śāstra method to something in Nyāya, Ayurveda or the Arthaśāstra. The last chapter of this book exists for that question alone.

The internal assessment, which is unlike every other paper

This is the part students get wrong, because they assume the usual class tests. MU's evaluation scheme splits its own list in two, and this paper is on the other side of the split.

What every other 2-credit theory paper gets: two class tests of 10 marks each, averaged to 10, plus two assignments of 5 marks each, totalling 10. Twenty marks of writing.

What this paper gets instead: one implementation, done individually or in a pair, on a topic from her list or another relevant IKS topic. Her own words for what it must include are these five items.

  1. A problem statement in the form "IKS concept as CS concept".
  2. A formal specification of the rules or algorithm.
  3. Working code.
  4. A minimum of ten test cases.
  5. A short analysis of correctness and limitations.

And it is shown through a live demo and a viva. You will be asked to run it in front of someone and answer questions about it.

munotes.in4

How This Paper Is Examined, and What That Means for How You Read

The eight topics she prints

You may choose another relevant IKS topic, but these eight are the ones she names, and every one of them is built in full somewhere in this book.

Her topicThe chapter that builds it
Piṅgala's Chandaḥśāstra as Binary Encoding and Combinatorial Generation[Prastāra in Code]
Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine[Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine]
Nyāya Logic as an Inference Engine[Nyāya Logic as an Inference Engine]
Ayurvedic Classification as a Rule-Based Expert System[Ayurvedic Classification as a Rule-Based Expert System]
Arthaśāstra-Inspired Cryptography as Symmetric Cipher System[Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System]
Meru-Prastāra as Pascal Triangle and Dynamic Programming Model[Meru-Prastāra as a Dynamic Programming Model]
Padārtha Ontology (Nyāya) as Knowledge Representation Model[Padārtha Ontology as a Knowledge Representation Model]
Śāstra Rule Precedence as Deterministic Finite Rewrite System[Śāstra Rule Precedence as a Deterministic Finite Rewrite System]

The eight headings the submission must carry

She prints these too, and in this order. A submission that is missing one of them is missing marks that were free.

  1. Title, in the form "IKS Concept as CS Concept"
  2. Problem Statement
  3. Conceptual Mapping Table
  4. Algorithm Specification (Pseudo-code)
  5. Working Code
  6. Minimum 10 Test Cases
  7. Complexity and Limitations
  8. Conclusion

[The Internal Assessment: Building and Presenting the Implementation] is a complete worked submission under all eight headings, so you have a model and not a description.

What this means for how the book is written

Every algorithm in this book has been run. Not proof-read, run. The tables you will read are the output of programs, and the programs are printed beside them, because you are going to be marked on code that works.

Definitions are set out to be found. A five-mark answer opens with a definition and a viva opens with one. So each chapter's first section is a definition twice over: once in plain words and once in the wording you can write down.

Each module closes on practice. [Practice for Module I] and [Practice for Module II] carry worked answers at five-mark length, and the book's last chapter does the same for the cross-module question.

What it does NOT mean

It does not mean skipping the reading to write code. Her first required item is a problem statement in the form "IKS concept as CS concept", and her third submission heading is a conceptual mapping table. Both of those are marks for understanding the classical side, and a demo without them is half an answer.

It does not mean the paper is easy because it is two credits. Thirty hours of teaching and fifty marks, with individual components that each have to be passed, is the same structure as any other paper. The internal is harder than a class test, not easier.

munotes.in5

How This Paper Is Examined, and What That Means for How You Read

Quick revision

  • 50 marks: 20 internal, 30 external.
  • External: 1 hour, Q.1 on Module 1, Q.2 on Module 2, Q.3 on both, each any 2 of 4 for 10 marks. An answer is worth 5 marks.
  • Internal: ONE implementation, alone or in a pair, with a problem statement, a formal specification, working code, at least 10 test cases, and a correctness and limitations analysis, shown by a live demo and a viva.
  • Eight topics are named and eight submission headings are required.
  • This internal scheme is named as being for this course and not for the others.

Test yourself

1. How many marks does one written answer on this paper carry, and how do you know?

Five. Each of the three questions is worth 10 marks and asks for any 2 of 4 alternatives, so each attempted alternative carries 5.

2. What is different about this paper's internal assessment, and where is that stated?

It is an implementation project rather than class tests and assignments. The evaluation scheme in the syllabus circular splits its own list into "For Courses other than IKS in Computational Systems" and "For IKS in Computational Systems course", and only this paper gets the project.

3. Name the five things the implementation must include.

A problem statement in the form "IKS concept as CS concept"; a formal specification of the rules or algorithm; working code; at least ten test cases; and a short analysis of correctness and limitations.

4. Why is the last chapter of this book about cross-module answers?

Because Q.3 is set on Modules 1 and 2 together, so ten of the thirty external marks require you to relate a topic from one module to a topic from the other.

Contents This chapter on its own page

munotes.in6

Chapter Three

Śāstra: What Makes a Body of Knowledge Formal

Syllabus topic Module 1, "Concept of śāstra as a formal knowledge system", "structured bodies of rule-governed knowledge"

In one line

A śāstra is a body of knowledge that has been given a form: defined terms, rules that apply in a stated order, and a declared account of what counts as evidence.

In the wording you can write in an examination: śāstra is the classical Indian term for a systematised discipline, that is, a body of rule-governed knowledge organised as a set of definitions, general rules, restrictions and meta-rules, transmitted in a compressed aphoristic form together with its own commentary, and resting on a declared theory of the admissible means of knowledge.

Why the word matters

"Subject" and "śāstra" are not the same idea. Cooking is a subject. A śāstra is a subject that has been put into a shape where an answer can be derived rather than remembered.

The test is simple and it is the one this book applies throughout. Given a question inside the field, can a competent person reach the answer by applying stated rules to stated facts, without asking the author what they meant? If yes, the knowledge has been formalised. If the answer depends on the teacher's taste, it has not.

The three things every śāstra declares

Read the opening of any of the treatises this paper covers and the same three declarations appear, in roughly the same place.

Its terms. Nyāya Sūtra 1.1.1 opens by listing sixteen categories and then spends the whole of Book I defining them. Nothing in the four remaining books uses a term that Book I has not fixed.

Its rules, and how they interact. Pāṇini's grammar has about four thousand rules and four stated principles for deciding which one applies when two could. The principles are themselves rules, written in the same form.

Its evidence. Charaka names the means of knowledge he will admit before he diagnoses anything, and adds that "by only one of these means of knowledge, knowledge does not arise of everything that should be known".

A body of knowledge that has done all three has a specification. That is exactly what the word formal is doing in MU's label.

The order a śāstra is taught in

Vidyabhusana, translating the commentary on Nyāya Sūtra 1.1.1, records that true knowledge of the sixteen categories means knowledge of their "enunciation", "definition" and "critical examination", and that Book I of the Nyāya Sūtra treats of enunciation and definition while the remaining four books are reserved for critical examination.

That is a three-stage order, and it is worth naming because it is the order in which a modern subject is taught too.

StageSanskritWhat happensThe modern habit
Enunciationuddeśathe terms are listed, and nothing moredeclaring the interface
Definitionlakṣaṇaeach term is given a definition that includes what it should and excludes what it should notfixing the semantics
Critical examinationparīkṣāthe definitions are attacked, and survive or are repairedthe test suite
munotes.in7

Śāstra: What Makes a Body of Knowledge Formal

The third stage is the one that is usually missing from a bad set of notes, and it is the one that makes a definition trustworthy. A definition nobody has attacked is a guess.

The four layers a śāstra is published in

This is the structural fact that makes the classical literature navigable, and it is almost never explained to a student before they meet their first Sanskrit citation.

The sūtra text. Aphorisms, as short as two words, arranged in a fixed order and numbered. Piṅgala's Chandaḥśāstra and Pāṇini's Aṣṭādhyāyī are both sūtra texts. A sūtra is unreadable without the next layer, and that is a deliberate design choice, not a failure.

The vṛtti. A running gloss that supplies the words the sūtra left out and shows one worked instance of each rule. It is the reference implementation: it tells you what the rule actually does.

The bhāṣya. A full commentary that argues about the rule, considers objections, and states why the rule is worded as it is rather than some other way. Halāyudha's Mṛtasañjīvanī on Piṅgala is this layer, and most of what this book quotes from Piṅgala is quoted from it.

The vārttika. A critical layer on the commentary, noting where it is wrong, incomplete or over-general. It is the errata list, and its existence tells you the tradition expected its own commentaries to be checked.

Worked example: the same idea in both notations

The classical form. Piṅgala's sūtra for the one-syllable table is two words, "dvikau glau". Halāyudha's commentary expands it: write a ga above and a la below, and because the sūtra says "two", place two such pairs. Without the commentary the sūtra is undecipherable; with it, the operation is exact.

The modern form. A specification might read: for a metre of one syllable, the table of patterns is the ordered list (G, L). A documentation comment then explains that G stands for a heavy syllable and L for a light one and gives the table for n equal to two as an illustration.

The point of the comparison. Both have a terse normative core and a longer explanatory layer, and in both the normative core alone is insufficient to act on. The classical tradition made the separation explicit and gave each layer a name. Most modern codebases do not.

What a śāstra is NOT

It is not simply an old book. A hymn, a story and a chronicle are not śāstra, because none of them defines its terms or states rules for deriving new claims.

munotes.in8

Śāstra: What Makes a Body of Knowledge Formal

It is not necessarily correct. Formal and true are different properties. Ayurvedic classification is highly formalised and that says nothing about whether any particular classification is clinically right. A formal system can be rigorously wrong, and saying so is part of answering MU's "critically assess".

It is not a program. There is no machine. A śāstra is executed by a trained human reader, and the tradition invests heavily in training that reader, which is what the four layers are for.

Limits of the comparison

Two limits should be stated before the enthusiastic chapters begin.

The rules are not typed. Nothing in a sūtra text says what kind of object a rule takes. A modern specification can be checked mechanically for that; a sūtra is checked by a reader who already knows.

There is no separation between the system and its meta-system. Pāṇini's meta-rules are written in the same language and the same numbering as his ordinary rules. That is economical and it is also why the tradition needed the four layers: the distinctions live in the commentary, not in the notation.

Quick revision

  • Śāstra: a systematised, rule-governed discipline with defined terms, ordered rules and a declared theory of evidence.
  • Every śāstra declares three things: its terms, its rules and how they interact, and its admissible means of knowledge.
  • The teaching order is enunciation, definition, critical examination, and the third stage is what makes a definition trustworthy.
  • The four publication layers: sūtra, vṛtti, bhāṣya, vārttika, which are specification, reference implementation, commentary and errata.
  • Formal is not the same as true, and formal is not the same as executable.

Test yourself

1. Define śāstra in one sentence fit for an examination.

A systematised body of rule-governed knowledge with defined terms, rules whose interaction is itself governed by stated principles, and a declared account of the admissible means of knowledge, transmitted as aphorisms with their own commentary.

2. Name the four layers a śāstra is published in and give the modern counterpart of each.

Sūtra, the specification; vṛtti, the reference implementation; bhāṣya, the explanatory commentary; vārttika, the errata. Halāyudha's Mṛtasañjīvanī is a bhāṣya on Piṅgala.

3. A student says "the Arthaśāstra is a śāstra, so what it says about taxation must be right." What is wrong with that?

It confuses being formal with being true. Formality is a property of how a system states and derives its claims; correctness is a property of the claims. A formal system can derive false conclusions perfectly.

4. Why is a sūtra deliberately incomplete on its own?

Because it was built for oral transmission, where the cost of a word is high and a trained teacher is assumed. The missing material is supplied by the vṛtti and bhāṣya, so the incompleteness is a division of labour between layers rather than a defect.

Contents This chapter on its own page

munotes.in9

Chapter Four

The Sūtra Method as Compressed Symbolic Encoding

Syllabus topic Module 1, "Sutra method as compressed symbolic encoding", "symbolic abstraction"

In one line

A sūtra is a rule compressed until nothing can be removed without losing the rule, and the cost of that compression is that it cannot be read without a key.

In the wording you can write in an examination: the sūtra method is a technique of extreme symbolic compression in which a rule is stated in the fewest possible syllables by means of technical terms, single-letter markers, class names standing for sets of elements, and the carrying over of words from preceding rules, so that the text is memorisable and transmissible orally but requires a commentary to be decoded.

Why anybody would do this

The reason is the medium. These texts were composed to be carried in memory and recited, centuries before a cheap writing surface existed in India. In that medium three things are expensive and one is free.

Expensive: every syllable. A text is stored in human memory and copied by voice. Its size is the whole cost of owning it.

Expensive: every ambiguity. A recited text has no punctuation, no layout and no footnotes. Anything that can be misheard will be.

Expensive: every change. There is no way to issue a correction to a thousand reciters.

Free: the reader's training. A student learning a śāstra has a teacher for years. Knowledge the reader can be assumed to have costs the text nothing.

Given those four facts, the rational design is exactly the one the tradition chose: push everything you can into the reader's training and the commentary, and keep the text itself as small as it will go. A modern engineer makes the same trade whenever they choose a compact binary format over a self-describing text one.

The four compression techniques

Each of these is a technique you will meet again in Pāṇini, so they are named here once.

Technical terms. A defined word replaces a description. "Guṇa" replaces "the vowel e, o or ar as substituted for i, u or ṛ", which is what Pāṇini's rule 1.1.3 has to say in full and never has to say again.

Single-letter markers. A letter attached to a form, not pronounced as part of it, that carries information about how rules may treat the form. [Anubandha: The Marker Letter as a Type Tag] is about these.

Class names for sets. A two-character name standing for an arbitrary set of sounds, so that a rule can say "before any vowel" in two letters. [The Fourteen Śivasūtras and the Pratyāhāra] shows how the names are built.

Carrying words over. A word stated in one sūtra is understood in the sūtras that follow, until something cancels it. The tradition calls this anuvṛtti. It is the single largest source of compression in Pāṇini, and the single largest source of difficulty in reading him.

munotes.in10

The Sūtra Method as Compressed Symbolic Encoding

Worked example: two words that generate a table

Piṅgala's rule for the shortest metre is two words: dvikau glau.

What is in the text. "Two" and "ga and la", where ga names the heavy syllable and la the light one. Four syllables of Sanskrit.

What Halāyudha's commentary supplies. Write a ga, and below it write a la. Because the sūtra says "two", these are the two forms a one-syllable metre can take. He then reads the word "two" again in the next sūtra to build the two-syllable table, and so on.

What a modern statement of the same rule would need. An ordered list of the two symbols, a statement that a one-syllable metre has exactly two forms, a statement of which comes first, and a note on what the symbols mean. Perhaps twenty-five words.

The compression ratio, honestly. Four syllables against twenty-five words looks like a factor of ten. It is not, because the commentary is part of the system and the commentary on those four syllables runs to a paragraph. What the compression buys is not total size, it is the size of the part that must be memorised exactly.

The three costs, stated plainly

This is the section that separates an answer from an advertisement.

No random access. Because words carry over from earlier sūtras, you cannot read sūtra 8.24 without knowing what is in force from the sūtras before it. The text is a stream, not an array. A modern equivalent is a compressed file you must decompress from the beginning to read the middle of.

No redundancy. Every syllable is load-bearing, so nothing survives damage. Lose a word and the rule is not merely less clear, it is a different rule.

No error detection. A recited text has no checksum. The tradition's answer was social rather than technical: multiple lineages of reciters, and the vārttika layer whose job is to notice that a rule as received cannot be right.

PropertyA sūtra textA modern specification document
Size of the normative coreminimal, memorisablewhatever it takes
Readable in isolationno, needs the commentaryusually yes
Random access to one ruleno, words carry overyes, rules are self-contained
Redundancy against damagenonehigh, prose repeats itself
Error detectionsocial, by parallel transmissionversion control and review
Cost of the reader's trainingassumed and highassumed and low

What it does NOT mean

It does not mean the sūtra style is obscure on purpose. It is terse on purpose, which is a different thing. Every device listed above has a stated function, and the tradition documents all of them.

It does not mean a shorter sūtra is a better one. The measure the tradition actually applies is that a rule should state what is needed and nothing else. A rule that has been shortened past that point is criticised in the commentary, not admired.

munotes.in11

The Sūtra Method as Compressed Symbolic Encoding

It does not mean compression is free. The three costs above are real, and the whole four-layer publication scheme of the previous chapter exists to pay them.

Limits and the modern parallel

The honest parallel is not to source code, which is read by a machine that will not guess. It is to a wire format: a representation chosen to be small, whose meaning lives in a separate schema, which is unreadable without that schema, and which is brittle if a byte is lost.

Where the parallel breaks is that a wire format has a decoder that is itself mechanical and exact. A sūtra's decoder is a human being with years of training, and two trained readers can disagree. The commentaries are full of such disagreements, and this book records the ones that matter to its own claims.

Quick revision

  • The sūtra method is compression driven by an oral medium: syllables, ambiguity and corrections are expensive; the reader's training is free.
  • Four techniques: technical terms, single-letter markers, class names for sets, and carrying words over from earlier rules.
  • "Dvikau glau", four syllables, generates the one-syllable table, with the commentary supplying the procedure.
  • Three costs: no random access, no redundancy, no error detection.
  • The modern parallel is a wire format with an external schema, not source code.

Test yourself

1. Why is a sūtra short? Give the reason in terms of the medium rather than of style.

Because the text lives in human memory and is transmitted by voice, so its size is its entire cost of ownership, while the reader's years of training cost the text nothing. The rational design pushes everything possible out of the text and into training and commentary.

2. Name the four compression techniques and say what each replaces.

Technical terms replace descriptions; single-letter markers replace statements about how a form may be treated; class names replace lists of elements; carrying words over replaces restating conditions in every rule.

3. State one cost of the sūtra method and give its modern analogue.

Words carry over from earlier sūtras, so there is no random access into the middle of the text. The analogue is a compressed stream that must be decoded from the start to read any part of it.

4. Is "shorter is better" the tradition's own standard? Justify your answer.

No. The standard is that a rule should state what is needed and nothing more; commentaries criticise rules that have been shortened past the point of stating their own conditions.

Contents This chapter on its own page

munotes.in12

Chapter Five

Pramāṇa: Where Knowledge Is Allowed To Come From

Syllabus topic Module 1, "Pramāṇa theory"

In one line

A pramāṇa is an admitted way of coming to know something, and a knowledge system that has not declared its pramāṇas has not said what it will accept as evidence.

In the wording you can write in an examination: pramāṇa theory is the branch of Indian epistemology that enumerates and defines the valid means of knowledge. The syllabus names three: pratyakṣa, perception; anumāna, inference; and āgama, also called śabda, authoritative verbal testimony. A claim is admissible only if it arises from a declared pramāṇa, and each pramāṇa has stated conditions under which it is reliable.

Why a system declares its evidence at all

Think of it as an interface rather than a philosophy. Before a system can be argued with, the parties have to agree on what counts as a reason. Otherwise every dispute ends in one side saying "that is not evidence" and neither side able to show why.

The classical treatises solve this the way a well-designed protocol does: they publish the list first. Nyāya's list is the first item of its first sūtra. Charaka states his list before he discusses a single disease.

That gives you a second, sharper benefit. Once the list is closed, a fallacy becomes detectable, because a bad argument is one that appeals to something not on the list, or that fails a stated condition of something that is. [Hetvābhāsa: The Five Fallacies of the Reason] is built entirely on that idea.

The three the syllabus names

PramāṇaPlain meaningWhat it deliversIts characteristic failure
Pratyakṣaperceptiona particular fact about a particular thing, here and nowthe senses misreport, or the observer names what is not there
Anumānainferencea fact not observed, drawn from one that was, by way of a general connectionthe general connection does not hold
Āgamatestimonya fact nobody present can observe or derivethe source is not trustworthy

Each of the next three chapters takes one of them. What matters here is the shape of the set: one channel to the world, one channel to what follows from it, and one channel to what others know. Any system that acquires knowledge has to have all three or do without one deliberately, and that is as true of a piece of software as of a physician.

Charaka's order, and the sentence that justifies it

Charaka does not merely list the three. He orders them, and gives the reason.

Verily, with the aid of all these three means of knowledge, one should in the first instance fully examine a disease. The diagnosis that is then arrived at becomes faultless.

Truly, by only one of these means of knowledge, knowledge does not arise of everything that should be known.

munotes.in13

Pramāṇa: Where Knowledge Is Allowed To Come From

Among all these three means of knowledge, the knowledge derived from the instructions of the inspired comes first. After this, comes examination, with the aid of Observation and Inference.

Three separate claims are in those lines and all three are worth having.

Use all three. A diagnosis reached from one channel is not faultless, and he says so.

No single channel is complete. This is the strongest statement of the three and the most modern. It is the classical form of what an engineer means by not trusting a single sensor.

Testimony comes first in time. Not first in authority: first in order of use. You read the literature before you examine the patient, because without the literature you do not know what to look for. Charaka puts the argument as a question: what would one who has not been instructed succeed in knowing by examining with observation and inference?

Worked example: the same three in a modern system

Take a monitoring system that decides whether a web service is unhealthy.

Pratyakṣa. It measures response time directly from a probe. That is perception: contact with the object, reporting a particular fact.

Anumāna. It sees the error rate rising and the queue depth growing, and concludes the database connection pool is exhausted, which it cannot observe. That is inference, and it depends on a general connection between exhausted pools and those two symptoms. If the connection does not hold, the inference is wrong however good the measurements were.

Āgama. It consults a runbook written by the team that built the service, which says that this combination of symptoms in this service means the pool. That is testimony, and its reliability is exactly the reliability of the team that wrote it.

Now apply Charaka's three claims. A system using only the probe cannot say why. A system using only the runbook cannot notice something new. And the runbook has to be read before the probe is useful, because it tells you which numbers matter. His ordering is not piety; it is the right order.

What pramāṇa theory does NOT say

It does not say testimony outranks observation. Charaka's word is that instruction comes first, and his own explanation is that it comes first in sequence because it tells you what to examine. Nothing in the passage makes it the strongest evidence.

It does not make the list obviously three. Different schools close the list at different places, and Nyāya Sūtra 1.1.3's commentary enumerates them: the Cārvākas admit only perception; the Vaiśeṣikas and Bauddhas two; the Sāṃkhyas three; the Naiyāyikas four, adding comparison; and later schools add presumption, non-existence, probability and rumour. MU's three are the Sāṃkhya and Āyurvedic set. [The Four Pramāṇas of Nyāya, and Why Charaka Has Three] deals with the difference.

munotes.in14

Pramāṇa: Where Knowledge Is Allowed To Come From

It does not settle what perception is. Nyāya Sūtra 1.1.4 imposes four conditions on perception precisely because "I saw it" is not a definition. The next chapter is about that.

Limits

The scheme has one structural weakness that a modern reader should see at once. It has no mechanism for revising the list. A pramāṇa is admitted or not, and a debate about whether comparison is an independent means of knowledge is conducted by argument, not by measurement. There is nothing in the apparatus resembling a calibration experiment.

The strength is the other side of the same coin: because the list is closed and public, an argument can be checked against it by anyone, which is what makes the debate rules of Module II enforceable.

Quick revision

  • Pramāṇa: an admitted means of knowledge. MU names three: pratyakṣa, anumāna, āgama.
  • Declaring them first is what makes fallacy detectable: a bad argument appeals to something off the list or fails a stated condition.
  • Charaka: use all three; no single one suffices; instruction comes first in order of use.
  • "By only one of these means of knowledge, knowledge does not arise of everything that should be known."
  • Schools close the list at one, two, three, four, six or eight. Three is the Sāṃkhya and Āyurvedic set; Nyāya has four.

Test yourself

1. Define pramāṇa and name the three the syllabus prints.

A pramāṇa is an admitted valid means of coming to know something. The three printed are pratyakṣa (perception), anumāna (inference) and āgama or śabda (authoritative verbal testimony).

2. Quote or paraphrase Charaka's reason for using all three, and explain its modern form.

He says that with all three one should first fully examine a disease and the diagnosis then becomes faultless, and that by only one of them knowledge does not arise of everything that should be known. The modern form is that no single source of evidence is sufficient, so a reliable system combines direct measurement, inference and documented prior knowledge.

3. Why does Charaka put testimony first, and what does that ordering not mean?

First in order of use, because without prior instruction one does not know what to observe or what to infer. It does not mean testimony is the strongest evidence or that it overrides observation.

4. Give one weakness of pramāṇa theory as a theory of evidence.

It has no procedure for revising its own list. Whether a candidate means of knowledge is independent is settled by argument rather than by any measurement, so the set of admitted channels is fixed by dispute rather than by test.

Contents This chapter on its own page

munotes.in15

Chapter Six

Pratyakṣa: Perception, and Why It Is Defined So Narrowly

Syllabus topic Module 1, "Pratyakṣa"

In one line

Pratyakṣa is knowledge that comes from a sense actually touching its object, and the definition spends most of its words ruling out cases where it did not.

In the wording you can write in an examination: Nyāya Sūtra 1.1.4 defines perception as knowledge which arises from the contact of a sense with its object and which is determinate, unnameable and non-erratic. The four conditions together exclude inference, indeterminate impressions, verbal knowledge and illusion.

Why the definition is so careful

"I saw it" is not evidence, and the sūtra knows it. Three quarters of the definition is there to exclude something.

That is the first thing a computer science student should notice, because it is how a good specification is written. A loose definition of perception would let every mistake in through the front door, and every downstream rule that depends on perception would inherit the mistake.

The provision

Nyāya Sūtra 1.1.4, in Vidyabhusana's translation:

Perception is that knowledge which arises from the contact of a sense with its object and which is determinate, unnameable and non-erratic.

Broken down

Condition one: contact of a sense with its object. The sense organ and the thing must actually meet. This is what rules out inference: when you conclude there is fire because you see smoke, no sense has touched the fire.

Condition two: determinate. Vidyabhusana's own illustration: "a man looking from a distance cannot ascertain whether there is smoke or dust." An impression that has not settled into a definite content is not perception. In modern terms, a reading below the confidence threshold is not a reading.

Condition three: unnameable. He glosses it: the knowledge "has no connection with the name which the thing bears." Perception is of the thing, not of the word for it. This is what separates perception from knowledge got through language, which is testimony's channel and not this one.

Condition four: non-erratic. His illustration is a mirage: "In summer the sun's rays coming in contact with earthly heat quiver and appear to the eyes of men as water. The knowledge of water derived in this way is not perception." A systematic illusion is excluded by name.

He also records an alternative reading of the same aphorism, on which perception may be either indeterminate ("this is something") or determinate ("this is a Brahmaṇa"), with only non-erraticness required in both. That disagreement is in the tradition and this book does not resolve it; a question that asks you to explain 1.1.4 should mention it.

The two kinds, and why the distinction matters

On the reading just mentioned, perception has two stages.

StageSanskritWhat it isThe modern counterpart
Indeterminatenirvikalpakabare awareness that something is there, with no classificationa raw signal, before any label
Determinatesavikalpakaawareness of it as being of a kinda classified observation
munotes.in16

Pratyakṣa: Perception, and Why It Is Defined So Narrowly

A monitoring system has both: a number arriving from a probe, and the judgement that the number means the service is slow. The classical scheme is right that these are different things and that errors enter at the second step.

Worked example: four failures, one for each condition

A student is asked to record what is in a laboratory. Four of their entries are not perception, and each fails a different condition.

"There is a soldering iron in the next room, because I can smell hot flux." Fails contact. No sense has touched the iron. This is inference, and it belongs to anumāna.

"There is something on the far bench, but I cannot tell what." Fails determinate. The impression has not resolved.

"There is an oscilloscope, because the label on the trolley says so." Fails unnameable. The knowledge came through a word, not through the thing. This is testimony.

"There is water on the corridor floor," when the floor is dry and polished under a bright light. Fails non-erratic. A regular optical effect has produced a regular mistake.

Every one of the four is a real way a system goes wrong, and the value of the definition is that it names them apart.

Charaka's complication: the mind as a sixth sense

Kaviratna's note on the same topic records something the Nyāya definition leaves implicit:

As ordinarily understood, "pratyaksha" implies all that is acquired by the senses. As the mind, however, is regarded in Hindu philosophy as the sixth sense, that which is acquired by the mind is included within the word.

And Charaka's own line is that observation is "the result of observation which one acquires by one's own senses and mind."

This matters for the classification chapters of Module II. If the mind is a sense, then noticing that two symptoms usually occur together is itself a kind of perception rather than an inference, and the boundary between the first two pramāṇas moves. The honest thing to say in an answer is that the boundary is drawn differently in the two traditions and to name which one you are using.

What it does NOT mean

It does not mean perception is infallible. The four conditions are a filter, not a guarantee. An observation can satisfy all four and still be of a badly made instrument.

It does not mean naming is bad. The unnameable condition says only that the naming is not part of the perceiving. The classification that follows is essential; it is just a different step, with its own failure modes.

It does not mean indeterminate awareness is worthless. It is the raw material. What the sūtra denies is that it is already knowledge.

munotes.in17

Pratyakṣa: Perception, and Why It Is Defined So Narrowly

Limits

The scheme has no notion of measurement error as a quantity. A perception either satisfies the conditions or it does not; there is no confidence attached to it. Modern practice attaches a number, and that number is the main thing a classical epistemology has no slot for.

It also has no account of instrumented perception. A reading taken through a microscope is perception of what, exactly? The tradition did not need the question; a computer science student does, because every sensor is an instrument.

Quick revision

  • Nyāya Sūtra 1.1.4: perception is knowledge arising from the contact of a sense with its object, and determinate, unnameable and non-erratic.
  • Contact excludes inference; determinate excludes unresolved impressions; unnameable excludes knowledge got through words; non-erratic excludes systematic illusion, and the mirage is the Sūtra's own example.
  • An alternative reading allows indeterminate perception, "this is something", as well as determinate, "this is a Brahmaṇa".
  • Charaka's translator records the mind as a sixth sense, so what the mind acquires falls inside pratyakṣa for him.
  • The scheme has no quantity of error and no account of instruments.

Test yourself

1. State Nyāya Sūtra 1.1.4 and say what each of its three epithets excludes.

Perception is knowledge arising from the contact of a sense with its object and which is determinate, unnameable and non-erratic. Determinate excludes unresolved impressions such as not knowing whether a distant plume is smoke or dust; unnameable excludes knowledge that comes by way of the thing's name; non-erratic excludes systematic illusion, such as taking a mirage for water.

2. "I know the server is overloaded because the queue is long." Which pramāṇa is this, and why not perception?

Anumāna. No sense has contacted the overload itself; the conclusion is drawn from an observed sign by way of a general connection, which fails the contact condition of 1.1.4.

3. What does Charaka's tradition add to the definition, and what does it change?

That the mind counts as a sixth sense, so what the mind acquires is included in pratyakṣa. It moves the boundary between perception and inference, so an answer should say which tradition's boundary it is using.

4. Give one thing the classical account of perception has no room for.

A quantified error or confidence. A perception either meets the four conditions or does not; there is no number attached to how reliable it is.

Contents This chapter on its own page

munotes.in18

Chapter Seven

Anumāna: Inference in Outline

Syllabus topic Module 1, "Anumāna"

In one line

Anumāna is knowing something you have not observed, from something you have, by way of a general connection between the two.

In the wording you can write in an examination: Nyāya Sūtra 1.1.5 defines inference as knowledge which is preceded by perception, and states that it is of three kinds: a priori, a posteriori, and commonly seen. Inference therefore always rests on a prior observation and on an invariable connection between what was observed and what is concluded.

Why it is called "preceded by perception"

Two things are meant by that phrase and both are worth having.

You must have observed the sign. Inference does not start from nothing. Something was perceived, and the conclusion hangs off it.

You must already have observed the connection. You conclude fire from smoke only because you have previously seen smoke together with fire, many times, and never smoke without fire. That prior observation is doing most of the work, and Module II's chapter on vyāpti is about exactly it.

So an inference has three ingredients: a thing you are talking about, a mark you observed in it, and a general connection you learned earlier. Take any one away and there is no inference.

The provision

Nyāya Sūtra 1.1.5, in Vidyabhusana's translation:

Inference is knowledge which is preceded by perception, and is of three kinds, viz., a priori, a posteriori and "commonly seen."

The three kinds, with the Sūtra's own examples

KindDirectionVidyabhusana's example
A priorifrom cause to effect"one seeing clouds infers that there will be rain"
A posteriorifrom effect to cause"one seeing a river swollen infers that there was rain"
Commonly seenfrom one thing to another regularly found with it"one seeing a beast possessing horns, infers that it possesses also a tail", or "one seeing smoke on a hill infers that there is fire on it"

He then records a disagreement, and it is the sort of thing an examiner likes: Vātsyāyana, the earliest commentator, takes the third kind to be "not commonly seen" instead, and reads it as knowledge of something never observed at all, as when "observing affection, aversion and other qualities one infers that there is a substance called soul."

That is not a small difference. On the first reading the third kind is the weakest of the three, a correlation. On Vātsyāyana's it is the strongest claim inference can make, reaching a thing that could never be perceived. Say which reading you are using.

The three kinds as three kinds of program

The classification is not decorative. Each kind has a different reliability profile, and a modern system meets all three.

Cause to effect is prediction. You have the cause and you want the future. It is the least reliable of the three, because a cause can be blocked. Clouds do not always bring rain.

munotes.in19

Anumāna: Inference in Outline

Effect to cause is diagnosis. You have the symptom and you want the fault. It is unreliable in a different way: several causes can produce the same effect. A swollen river may mean rain upstream, or a dam release.

Correlation is the one that looks safest and is not. Horns and tails go together in the animals you have seen. Nothing makes them go together. The connection is a summary of observation, not a mechanism, and it fails the first time you meet an animal outside the sample.

That third one is the classical statement of the standard modern warning, and it is stated in the Sūtra's own terms: the reason is erratic when there is no universal connection between it and what you concluded. [Hetvābhāsa: The Five Fallacies of the Reason] works that out.

Worked example

A student's laptop will not boot. Three inferences follow, one of each kind.

A priori. "The battery is at two per cent and the charger is not plugged in, so it is going to shut down." Cause to effect. Blockable: someone may plug it in.

A posteriori. "It powers on, the fan spins, and nothing appears on the screen, so the display or its cable has failed." Effect to cause. Several causes fit, so the inference narrows rather than settles.

Commonly seen. "Every laptop of this model I have seen with this fault had a failed cable, so this one has a failed cable." Correlation. It carries no mechanism and no guarantee that the sample was representative.

The three conclusions are not equally good, and knowing which kind you are making is how you know how much weight to put on it.

What it does NOT mean

It does not mean inference is second class. It is a pramāṇa in its own right, on every school's list including the shortest ones after the Cārvākas'.

It does not mean an inference is a guess. An inference whose general connection is sound is knowledge. The classical machinery exists precisely to check the connection, which is what [Anumāna: Vyāpti, and the Three Kinds of Inference] is about.

It does not mean "preceded by perception" makes it about the past. The perception that precedes it may be of the sign now; what is prior is the observation of the connection.

Where Module II takes this

Everything else about inference is there, and none of it is repeated here.

  • The structure of an inference in five stated members, in [The Five-Member Syllogism].
  • The invariable connection itself, vyāpti, and the vocabulary of subject, mark, positive and negative instance, in [Anumāna: Vyāpti, and the Three Kinds of Inference].
  • The five ways the reason can be bad, in [Hetvābhāsa: The Five Fallacies of the Reason].
  • Writing the whole thing in logical notation, in [The Five Members Written in Logical Notation].
munotes.in20

Anumāna: Inference in Outline

Quick revision

  • Nyāya Sūtra 1.1.5: inference is knowledge preceded by perception, of three kinds: a priori, a posteriori, commonly seen.
  • Ingredients: a subject, a mark observed in it, and a previously established general connection.
  • A priori is cause to effect, prediction; a posteriori is effect to cause, diagnosis; commonly seen is correlation.
  • Vātsyāyana reads the third kind as "not commonly seen", reaching what can never be perceived, which is the opposite of a correlation.
  • Correlation is the kind that looks safest and is weakest, because it carries no mechanism.

Test yourself

1. Define anumāna and give its three kinds with one example each.

Knowledge preceded by perception, drawn from an observed mark by way of a general connection. A priori: clouds, so rain. A posteriori: a swollen river, so there was rain. Commonly seen: a horned beast, so it has a tail.

2. Why does the definition insist that inference is "preceded by perception"?

Because two prior observations are needed: the mark must have been perceived now, and the general connection between mark and conclusion must have been observed earlier.

3. Which of the three kinds is a correlation, and what is its characteristic weakness?

"Commonly seen." It summarises past observation without supplying a mechanism, so it fails as soon as a case outside the sample appears.

4. What does Vātsyāyana do to the third kind, and why does it matter which reading you use?

He reads it as "not commonly seen", inference to something never perceptible at all, such as the soul from its qualities. It matters because on one reading the third kind is the weakest form of inference and on the other it is the boldest.

Contents This chapter on its own page

munotes.in21

Chapter Eight

Āgama: Testimony, and the Trusted Source as a Design Decision

Syllabus topic Module 1, "Āgama"

In one line

Āgama is knowing something because somebody trustworthy told you, and the interesting part of the definition is what makes them trustworthy.

In the wording you can write in an examination: āgama, also called śabda, is the instructive assertion of a reliable person, and it is of two kinds according as it concerns matter that is seen, that is, verifiable, or matter that is not seen. The reliability of the source, not the form of the statement, is what makes it a valid means of knowledge.

Why a system needs this channel at all

Nothing a single observer can perceive or infer is enough to work with. A physician has not personally observed the course of every disease. A programmer has not personally measured every library they call. Most of what anyone knows arrived through somebody else, and a theory of evidence that does not admit that is describing a person who does not exist.

Charaka puts the argument as a challenge: what would one who has not first been instructed succeed in knowing by examining with observation and inference? The answer is very little, because they would not know what to look at.

The provision

Nyāya Sūtra 1.1.7, in Vidyabhusana's translation:

Word (verbal testimony) is the instructive assertion of a reliable person.

And his gloss on who counts:

A reliable person is one, may be a ṛṣi, Ārya or mleccha, who as an expert in a certain matter is willing to communicate his experiences of it.

Nyāya Sūtra 1.1.8:

It is of two kinds, viz., that which refers to matter which is seen and that which refers to matter which is not seen.

Broken down: the trust policy in three conditions

The gloss on 1.1.7 is short and every clause in it does work. Read as a policy for admitting a source, it has three conditions and one explicit non-condition.

Expert in a certain matter. Not expert generally. Authority is scoped to a subject, and a source trusted outside its scope is being misused. This is the condition modern practice breaks most often.

Willing to communicate his experiences. The source is reporting what they have themselves undergone, and is doing so as instruction rather than in passing. Second-hand report and overheard remark do not qualify.

And, from the river example, no enmity against the hearer. The source must have no interest in your being wrong.

The non-condition: who they are. The gloss says the reliable person may be a ṛṣi, an Ārya or a mleccha, that is, a sage, a member of the speaker's own community, or a foreigner. Social standing is expressly not part of the test. For a text of its period that is a striking piece of drafting, and it is the right rule: the warrant is expertise plus candour, not identity.

munotes.in22

Āgama: Testimony, and the Trusted Source as a Design Decision

The Sūtra's own worked example

Suppose a young man coming to the side of a river cannot ascertain whether the river is fordable or not, and immediately an old experienced man of the locality, who has no enmity against him, comes and tells him that the river is easily fordable: the word of the old man is to be accepted as a means of right knowledge called verbal testimony.

Every element of the policy is in that sentence. The young man cannot perceive the riverbed and has no general connection to infer from, so the other two channels are closed. The old man is local, which is the scope of his expertise; experienced, which is its depth; and has no enmity, which is the absence of a motive to mislead.

The two kinds, and what they cost to check

KindWhat it concernsVidyabhusana's exampleHow it can be checked
Seenmatter that can be verified"a physician's assertion that physical strength is gained by taking butter"directly, by observation
Not seenmatter that cannot"a religious teacher's assertion that one conquers heaven by performing horse-sacrifices"not directly; he says we "can somehow ascertain it by means of inference"

The distinction is the useful one. A claim of the first kind puts its source at risk, because it can be falsified. A claim of the second kind does not, so it rests entirely on the source's standing. Any answer on āgama should make that point, because it is where the channel is weakest.

Worked example: the same policy in software

A team decides which sources its incident response may rely on.

Admitted, matter that is seen. The database vendor's documented maximum connection limit. The vendor is expert in that matter, states it as instruction, and has no interest in the team being wrong. The claim is verifiable, so the source is at risk if it is untrue.

Admitted with care, matter not seen. A retired colleague's account of why a design decision was taken five years ago. Expert, candid, no enmity, but nothing can now verify it. It is usable and it is not checkable, and the team should record which it is.

Not admitted, scope. The same vendor's opinion about which cloud provider is cheapest. Expertise does not extend there.

Not admitted, motive. A benchmark published by a product's own vendor comparing it with a competitor. The enmity condition is failed in its mirror image: an interest in your believing something.

That last case is worth dwelling on. The classical condition is stated as absence of enmity toward the hearer, which is narrower than absence of interest. A modern policy has to widen it, and saying so is a fair criticism of the classical rule rather than a defect in your answer.

munotes.in23

Āgama: Testimony, and the Trusted Source as a Design Decision

What it does NOT mean

It does not mean tradition is self-certifying. The warrant is expertise and candour in a named matter, and the gloss deliberately refuses to make identity or standing part of the test.

It does not mean testimony outranks observation. Charaka's ordering puts instruction first in sequence, not first in authority.

It does not mean unverifiable testimony is worthless. Vidyabhusana says the second kind can somehow be reached by inference. It means that such a claim rests on the source alone, and you should know when you are in that position.

Limits

Three, and all three matter for the Module II chapters on expert systems.

No degrees. A source is reliable or it is not. There is no weighting, and no procedure for combining two sources that disagree.

No account of chains. Almost all real testimony is a chain: A told B who wrote it down and C read it. Every link can fail and the classical scheme addresses only the first.

Interest is under-specified. Absence of enmity is not absence of incentive, and the difference is most of modern source criticism.

Quick revision

  • Āgama or śabda: the instructive assertion of a reliable person, Nyāya Sūtra 1.1.7.
  • A reliable person is an expert in the matter, willing to communicate their own experience, and, from the river example, without enmity toward the hearer. Identity is expressly not a condition: ṛṣi, Ārya or mleccha alike.
  • Two kinds, 1.1.8: matter seen, which is verifiable, and matter not seen, which is not.
  • The river ford is the Sūtra's own worked example and contains the whole policy.
  • Weaknesses: no degrees of trust, no account of chains of transmission, and absence of enmity is narrower than absence of interest.

Test yourself

1. State the definition of āgama and the three conditions on a reliable person.

The instructive assertion of a reliable person. The person must be an expert in the matter concerned, must be willing to communicate their own experience of it, and, on the Sūtra's own example, must have no enmity against the hearer.

2. Which condition does the classical gloss expressly refuse to impose, and why is that notable?

Identity or social standing: the reliable person may be a ṛṣi, an Ārya or a mleccha. It is notable because it makes the test turn on expertise and candour alone, which is the correct rule.

3. Distinguish the two kinds of testimony and say which is the weaker warrant.

Testimony about matter that is seen can be verified, so a false source is exposed. Testimony about matter that is not seen cannot, so it rests entirely on the source's standing, and is the weaker warrant.

munotes.in24

Āgama: Testimony, and the Trusted Source as a Design Decision

4. Give one way in which a modern trust policy has to go beyond the classical conditions.

It must exclude sources with an interest in the hearer believing a particular thing, not merely sources hostile to the hearer. A vendor's own comparative benchmark passes the enmity test and fails the interest test.

Contents This chapter on its own page

munotes.in25

Chapter Nine

Knowledge Validation and Structured Reasoning

Syllabus topic Module 1, "Knowledge validation and structured reasoning"

In one line

Structured reasoning means the steps of an argument are named in advance, so that a bad argument can be shown to be bad rather than merely disbelieved.

In the wording you can write in an examination: knowledge validation in the Indian systems consists of a declared set of admissible means of knowledge, a fixed form in which a claim must be presented before it counts as established, and an enumerated list of defects which, if present, disqualify the argument. Together these make the assessment of a claim a procedure rather than a judgement.

The problem being solved

Two people disagree. Each is sincere. How does the dispute end without one of them simply outlasting the other?

The classical answer is to make the argument itself inspectable. If the form is fixed, then a missing part is visible. If the admissible reasons are listed, an inadmissible one is visible. If the defects are named, a defect is nameable rather than merely felt.

Notice what this buys, because it is the point of the whole apparatus. A third party who was not present can check the argument. That is the property that turns an opinion into a finding, and it is the same property that makes a test suite worth having.

The three stages of validation

Collected from the traditions this paper covers, the machinery has three stages and they run in order.

Stage one, admissibility of the evidence. Does every premise arise from a declared pramāṇa? A claim resting on something that is not on the list does not enter. This is the filter of [Pramāṇa: Where Knowledge Is Allowed To Come From].

Stage two, form. Is the argument stated in all five members? Nyāya Sūtra 1.1.32 fixes them: proposition, reason, example, application, conclusion. A claim not in that form has not been put forward for assessment; it has only been asserted.

Stage three, defects. Does the reason suffer any of the five named defects? Nyāya Sūtra 1.2.4 lists them: erratic, contradictory, equal to the question, unproved, mistimed. Any one of them and the argument fails whatever its form.

StageWhat it checksWhat it catchesThe modern counterpart
Admissibilitythe source of each premisean appeal to something the system does not accept as evidenceinput validation at the boundary
Formthe presence of all five membersan assertion dressed as an argumenta required template, or a type signature
Defectsthe quality of the reasonan argument that is well formed and still worthlessa static analyser

Why the order matters

Run the stages in the wrong order and you waste effort. There is no point examining the quality of a reason that rests on inadmissible evidence, and no point looking for defects in something that has not been stated as an argument at all.

munotes.in26

Knowledge Validation and Structured Reasoning

That is exactly why a compiler checks that a file parses before it checks types, and checks types before it warns about unreachable code. Each stage assumes the previous one passed.

Worked example: one claim through all three stages

The claim. A student says: this sorting function is correct, because I ran it on my own test file and the output was sorted.

Stage one, admissibility. The premise is an observation of one run. That is pratyakṣa, and it is admissible. Nothing is rejected here.

Stage two, form. Reconstruct the five members and the gap appears.

  • Proposition: this function is correct.
  • Reason: because it sorted my test file.
  • Example: whatever sorts a file is correct, as ... and here the student has nothing to put. There is no general connection between sorting one file and being correct.
  • Application: cannot be stated.
  • Conclusion: cannot be drawn.

The argument fails at the third member, and it fails visibly, which is the whole value of the form. A vague objection ("that is not enough testing") becomes a precise one: the example member cannot be supplied.

Stage three, defects. Even if a general connection were offered, it would be erratic in the Sūtra's sense, because the same reason supports the opposite conclusion: an incorrect function can also sort one particular file. Nyāya Sūtra 1.2.5 defines the erratic reason as one that leads to more conclusions than one, and the Sūtra's own illustration has exactly this shape.

Charaka's validation rule, which is different and complementary

The Nyāya machinery validates an argument. Charaka validates an investigation, and his rule is one sentence:

Truly, by only one of these means of knowledge, knowledge does not arise of everything that should be known.

That is a rule about coverage rather than about form. An investigation that used only one channel is incomplete even if every step in it is faultless. A modern reviewer applies both rules and would recognise the second as asking whether the evidence is sufficient, not merely whether the reasoning is valid.

What structured reasoning does NOT give you

It does not give you truth. An argument can pass all three stages and be false, because a general connection believed on good evidence can still be wrong. The apparatus improves the odds and does not settle the matter.

It does not give you discovery. Nothing in it tells you what to claim. It assesses claims that are already made, which is why the five extra members some schools add, beginning with jijñāsā, the inquiry, are attempts to bolt a discovery stage onto the front.

It does not remove judgement. Deciding whether a general connection really holds is itself a judgement, and the tradition knows it: the whole debate literature exists to argue about exactly that.

munotes.in27

Knowledge Validation and Structured Reasoning

Limits

The scheme is binary at every stage. Evidence is admissible or not, form is complete or not, a defect is present or not. There is no slot for a weak but real argument, and no way to aggregate three weak arguments into a strong one. Modern practice does that constantly, and the classical apparatus has no notation for it.

It is also adversarial by design. It assumes an opponent whose job is to find the defect. That is a strength when there is one and a weakness when there is not, because nothing in the machinery checks itself.

Quick revision

  • Structured reasoning: named steps, so that a bad argument can be shown bad by a third party who was not present.
  • Three stages in order: admissibility of evidence, completeness of form, absence of the named defects.
  • The stages correspond to input validation, a required template, and a static analyser, and the order is forced for the same reason a compiler parses before it type-checks.
  • Charaka adds a coverage rule: one channel of evidence is never enough.
  • Limits: everything is binary, nothing aggregates weak arguments, and the whole apparatus needs an opponent to run it.

Test yourself

1. Name the three stages of validation and what each one catches.

Admissibility, which catches an appeal to something not accepted as evidence; form, which catches an assertion presented as an argument; and defects of the reason, which catch an argument that is well formed and still worthless.

2. Why is the order of the stages not arbitrary?

Each stage assumes the previous one passed. Examining the quality of a reason that rests on inadmissible evidence, or looking for defects in something not stated as an argument, is wasted work. A compiler parses before type-checking for the same reason.

3. Reconstruct "it works on my test file, so it is correct" in the five members and say where it fails.

Proposition: the function is correct. Reason: it sorted my test file. The example member cannot be supplied, because there is no general connection between sorting one file and being correct, so application and conclusion cannot follow. It fails at the third member.

4. Give one thing the classical validation scheme cannot express.

Degree. Every stage is pass or fail, so there is no way to record a weak but genuine argument, and no way to combine several weak arguments into a stronger case.

Contents This chapter on its own page

munotes.in28

Chapter Ten

Knowledge Representation: The Modern Name For It

Syllabus topic Module 1, "Knowledge representation"

In one line

Knowledge representation is the business of writing down what a system knows in a form the system can use to work out something it was not told.

In the wording you can write in an examination: knowledge representation is the branch of computer science concerned with encoding facts about a domain in a formal structure that supports inference. A representation is a choice of syntax for expressing knowledge, a semantics fixing what each expression means, an inference procedure for deriving new expressions from stored ones, and a commitment about what kinds of thing exist in the domain.

Why a representation is a choice, and why the choice costs something

A system can be told "Rama is a student". That could be stored as a row in a table, a sentence in logic, a link between two nodes, a slot in a record, or a rule. All five are faithful. They are not equivalent, because each makes different questions cheap and different questions expensive.

That is the whole subject in one sentence: a representation is a bet about which questions you will be asked. Choose the wrong notation and every query fights the structure.

The four things a representation must fix

These are the four questions to answer in an examination when asked what a knowledge representation is.

Syntax. What strings count as well formed. In logic, student(rama) is well formed and rama student( is not.

Semantics. What a well formed string means, precisely enough that two people cannot read it two ways. This is where most informal notations fail: a box joined to another box by an unlabelled arrow means whatever the reader guesses.

Inference. What may be derived from what. Without this a representation is a filing system, not a knowledge base.

Ontological commitment. What kinds of thing the notation says exist. A notation with only objects and properties cannot talk about events; one with only classes cannot talk about individuals. [Padārtha: The Categories of What Exists] is a classical answer to exactly this question.

The four families, on one domain

Take the smallest useful domain: a college has students; a student has a roll number; Rama is a student with roll number 41; every student may borrow books.

Logic. Facts and rules as formulas.

student(rama)

roll(rama, 41)

for all X: student(X) implies mayBorrow(X)

Cheap: asking whether Rama may borrow. Expensive: saying anything is only usually true.

Semantic network. A labelled graph. Nodes are things and classes; edges are named relations.

rama --is-a--> student --may--> borrow

rama --roll--> 41

Cheap: following relations, and finding what is connected to what. Expensive: saying anything with more than two arguments, because an edge joins exactly two nodes.

Frames. A record per concept, with named slots, defaults, and inheritance from a parent frame.

munotes.in29

Knowledge Representation: The Modern Name For It

frame student

parent: person

may: borrow

roll: an integer

frame rama

parent: student

roll: 41

Cheap: defaults and inheritance, so rama need not repeat may borrow. Expensive: exceptions, which have to be handled by overriding a slot and then explaining why.

Rules. IF conditions THEN conclusion, with an engine that applies them.

IF person(X) and enrolled(X)

THEN student(X)

IF student(X)

THEN mayBorrow(X)

Cheap: adding a new piece of knowledge without touching the rest. Expensive: understanding what the system will do overall, because the behaviour is spread across the rules.

The comparison a five-mark answer needs

FamilyBasic unitStrongest atWeakest at
Logica formulaexactness, and provable inferencedefaults, exceptions, degrees of belief
Semantic networka labelled edgefollowing relations, visualising structurerelations of more than two arguments
Framesa slot in a recorddefaults and inheritance, compact storageexceptions, and precise semantics
Rulesa condition and a conclusionadding knowledge piecemeal, explanationpredicting overall behaviour

Worked example: the same fact, and the question that separates the notations

The fact: most students return books on time.

In logic, this cannot be said at all without extending the notation. for all X: student(X) implies returnsOnTime(X) is false, and there is no standard way in plain first-order logic to say "most".

In a semantic network, you can draw an edge labelled "usually returns on time", and now the notation's semantics have quietly become whatever the reader thinks that label means.

In a frame, it is natural: the student frame gets a slot returns with default on time, and an individual frame may override it.

As a rule, it is natural in a different way: a rule concluding returnsOnTime(X) from student(X), plus a second rule that retracts it for a named individual.

The lesson is not that frames and rules are better. It is that the fact was easy to state in two notations and not statable in a third, and you cannot know that until you try. Pick the notation after you know the questions, not before.

What knowledge representation is NOT

It is not a database schema. A schema fixes syntax and, loosely, semantics. It has no inference procedure, so nothing is derived that was not stored.

It is not the same as storing text. A page of prose contains knowledge and supports no inference at all. Turning prose into a representation is the work, and it is the work every chapter of this paper's Module II asks you to do on a classical text.

It is not free of commitments. Every notation says something about what exists. Choosing one is choosing an ontology whether or not you notice.

munotes.in30

Knowledge Representation: The Modern Name For It

The classical connection, stated carefully

The claim MU's syllabus makes is that śāstra method is knowledge representation, and the honest version of that claim is specific.

What matches. A śāstra fixes its syntax, its semantics and its inference, which are three of the four requirements above, and Vaiśeṣika's padārtha scheme is an explicit ontological commitment, which is the fourth.

What does not. None of it is machine readable, and none of it was meant to be. The inference procedure is a trained reader. So the correspondence is at the level of design, not of implementation, and an answer that says so is stronger than one that does not.

Quick revision

  • Knowledge representation: encoding what is known so that new facts can be derived.
  • Four requirements: syntax, semantics, inference procedure, ontological commitment.
  • Four families: logic, semantic networks, frames, rules, each cheap at some questions and expensive at others.
  • A representation is a bet about which questions will be asked.
  • The classical correspondence is at the level of design: a śāstra fixes all four, but its inference procedure is a person.

Test yourself

1. Name the four things a knowledge representation must fix, and say what goes wrong if the second is missing.

Syntax, semantics, an inference procedure, and an ontological commitment. Without semantics, two readers can take the same expression two ways, so nothing derived from it can be relied on: an unlabelled arrow between two boxes is the standard example.

2. Why is a database schema not a knowledge representation?

It has no inference procedure, so it stores facts and derives none. A representation must support working out something the system was not told.

3. Take the fact "most students return books on time" and say which of the four families states it naturally and which cannot.

Frames state it as a default on a slot, and rules state it as a rule with an override. Plain first-order logic cannot state it at all, since the universal is false and there is no standard quantifier for "most". A semantic network can draw the label but loses precise semantics in doing so.

4. State the correspondence between śāstra method and knowledge representation, with its limit.

A śāstra fixes terms, rules and admissible evidence, and a scheme such as padārtha is an explicit ontology, so all four requirements are addressed. The limit is that nothing is machine readable and the inference procedure is a trained human reader, so the correspondence is one of design, not implementation.

Contents This chapter on its own page

munotes.in31

Chapter Eleven

Formal Specification: Saying Exactly What a System Must Do

Syllabus topic Module 1, "formal specification"

In one line

A formal specification says what a system must do in a language with no room for reading it two ways.

In the wording you can write in an examination: a formal specification is a description of required behaviour written in a notation with a precisely defined syntax and semantics, so that whether an implementation satisfies it is a matter of proof or of mechanical checking rather than of interpretation. It states what must hold, not how to achieve it.

Why ordinary language will not do

Here is a requirement, written the way requirements are usually written: every student has a supervisor.

It has at least three readings.

  1. For each student there is some supervisor, possibly a different one for each.
  2. There is one supervisor who supervises all the students.
  3. Each student has exactly one supervisor, no more.

A sentence in logic has to pick one.

for all S: student(S) implies (exists T: supervises(T, S))

That is reading one, and only reading one. Reading two moves the quantifier to the front; reading three adds a uniqueness condition. The notation forces the choice at writing time rather than leaving it to be discovered in testing.

That is the entire argument for formality, and it is worth stating that plainly: formal notation does not make you right, it makes you specific.

The three things a specification fixes

What must hold. The properties any acceptable implementation has. Not the code, and not the algorithm: the conditions.

Under what assumptions. The preconditions the caller must satisfy. A specification with no preconditions is either trivial or dishonest.

What is left free. Everything not stated is permitted. This is the part beginners get wrong: a specification is also a licence, and anything it fails to forbid is allowed.

Worked example: sorting, specified three ways

Informally. "Sorts the list." Admits an implementation that returns a list of zeroes of the same length, which is sorted.

Better. "Returns a list whose elements are in non-decreasing order and which is a permutation of the input." This closes the hole, and it names both conditions.

Formally.

pre: the list is finite

post: for all i in 0 .. n-2: out[i] <= out[i+1]

post: out is a permutation of in

free: the order of equal elements, unless stability is required

The last line matters. Nothing above forbids swapping two equal elements, so an implementation may. If the caller needs stability it has to be specified, and a caller who assumed it without specifying it has made the classic mistake.

The sūtra as a specification

This is the comparison MU's label is reaching for, and it holds better than most comparisons in this paper.

A specification hasA śāstra has
a notation with fixed syntaxthe sūtra style, with its technical terms and markers
meanings fixed before useBook I of the Nyāya Sūtra; Pāṇini's definitional sūtras
conditions stated, not proceduresmost vidhi rules state what is substituted for what, not how to write it
a rule for what happens when two requirements conflictthe four precedence principles, and sūtra 1.4.2
teststhe bhāṣya, which works an instance of every rule
erratathe vārttika
munotes.in32

Formal Specification: Saying Exactly What a System Must Do

The row that makes the comparison worth drawing is the fourth. Ordinary specifications are notoriously silent about conflicts between their own clauses, and a reader has to guess. Pāṇini wrote the conflict rule down. [Vipratiṣedha: When Two Rules Collide, the Later Wins] is that rule.

Worked example: a sūtra read as a specification

Pāṇini 7.3.84, in Vasu's translation: "when a sārvadhātuka or an ārdhadhātuka affix follows there is guṇa of the base."

As a specification.

pre: the affix is of kind sarvadhatuka or ardhadhatuka

pre: by 1.1.3, the vowel operated on is one of the ik vowels

post: that vowel is replaced by its guna

free: everything about the affix other than its kind

except: 1.1.5, an affix carrying an indicatory k or ng, blocks this

Three features of that are worth naming. The rule does not say how to perform the substitution, only what the result is. Its second precondition is supplied by another rule, which is what a meta-rule is for. And its exception is declared elsewhere and has priority, which is a stated conflict policy.

What a formal specification is NOT

It is not an implementation. It says what, not how. A specification that fixes the algorithm has over-specified and has forbidden better implementations.

It is not a guarantee of correctness. It can be wrong about what was wanted. Formality moves the argument from "does the code do this?" to "is this what we wanted?", which is progress and not a solution.

It is not necessarily mathematical notation. A precise subset of English with defined terms can be a specification. What matters is that the terms are fixed and the conditions are complete, which is exactly what a śāstra does.

Limits of the classical comparison

Two honest limits, both of which belong in a critical answer.

A sūtra's preconditions are not all written down. Some are carried over from earlier sūtras by anuvṛtti, and which ones are in force at a given rule is decided by the commentary. A modern specification that relied on an editor to say which preconditions applied would be rejected.

Nothing checks a sūtra mechanically. There is no machine that can confirm the rules are consistent, and in fact the tradition disputes the consistency of particular rules at length. A formal specification in the modern sense is one a tool can check; a sūtra is one only a trained reader can check.

munotes.in33

Formal Specification: Saying Exactly What a System Must Do

Quick revision

  • A formal specification states required behaviour in a notation whose syntax and semantics are fixed, so satisfaction is checkable rather than arguable.
  • It fixes three things: what must hold, under what assumptions, and what is left free. Anything not forbidden is permitted.
  • "Every student has a supervisor" has three readings; a formula has one.
  • The śāstra parallel is strong on definitions, conditions, conflict policy, tests and errata.
  • It is weak on two counts: preconditions carried over rather than written, and no mechanical checking.

Test yourself

1. Define a formal specification and name the three things it fixes.

A description of required behaviour in a notation with fixed syntax and semantics, so that satisfaction can be proved or checked mechanically. It fixes what must hold, the assumptions under which it must hold, and what is left free.

2. Give a requirement with more than one reading and show how a formal notation removes the ambiguity.

"Every student has a supervisor" may mean each has some supervisor, or there is one supervisor for all, or each has exactly one. Writing "for all S: student(S) implies (exists T: supervises(T, S))" commits to the first and to nothing else.

3. Why is it a mistake to think a specification only constrains?

Because it also licenses: whatever it does not forbid is permitted. A sort specification that does not require stability permits an unstable implementation, and a caller who assumed stability has misread the licence.

4. Name one respect in which a sūtra falls short of a modern formal specification.

Its preconditions are not all written at the rule; some are carried over from earlier rules and which are in force is settled by the commentary. Relatedly, no tool can check a sūtra set for consistency, so the checking is done by a trained reader.

Contents This chapter on its own page

munotes.in34

Chapter Twelve

Computational Thinking, and the Four Habits It Names

Syllabus topic Module 1, "aligned with computational thinking", "symbolic abstraction"

In one line

Computational thinking is the habit of turning a problem into something a stupid but tireless machine could finish.

In the wording you can write in an examination: computational thinking is a set of problem-solving habits drawn from computer science and applicable outside it, conventionally given as four: decomposition, breaking a problem into sub-problems; pattern recognition, noticing that a sub-problem has been solved before; abstraction, discarding the detail that does not bear on the problem; and algorithm design, stating the solution as a finite sequence of unambiguous steps.

Why the word is in this syllabus

MU's first course objective for this paper asks you to see śāstra as a structured and formal knowledge system "aligned with computational thinking". So the phrase is examinable, and the alignment is the thing to be argued.

The argument is not that classical authors were programmers. It is narrower and it is defensible: a tradition that has to transmit a discipline through human memory is under the same pressures as one that has to transmit it to a machine. Both must be exact, because there is no author to ask. Both must be compact. Both must handle the case where two rules apply at once. Those pressures produce the same four habits.

Decomposition

The habit. Split a problem into parts that can be solved separately, so that each part is small enough to be got right.

In a śāstra. Piṅgala does not solve "list every metre". He solves "list every metre of one syllable", which is two lines, and then states a rule that turns a list for n syllables into the list for n plus one. Two small problems instead of one unbounded one.

In a program. A compiler is not one procedure. It is a reader, a scanner, a parser, a type checker and a code generator, each of which can be written, tested and replaced on its own.

The test of whether you have done it. Can you check one part without running the others? If not, you have divided the text and not the problem.

Pattern recognition

The habit. Notice that this problem has the shape of one already solved, and reuse the solution rather than the answer.

In a śāstra. The number of metres of six syllables with exactly two light ones, and the number of ways of choosing two objects from six, are the same number. Halāyudha's Meru-prastāra computes both because they are the same problem wearing two descriptions. [Pascal Triangle and Combinatorics] works that out.

In a program. Finding the shortest route between two towns and finding the cheapest sequence of edits between two words are the same problem, so the same algorithm solves both.

The trap. Two problems that look alike may differ in exactly the respect that matters. A pattern is a hypothesis, and it has to be checked.

munotes.in35

Computational Thinking, and the Four Habits It Names

Abstraction

The habit. Throw away everything that does not bear on the question, and keep a name for what is left.

In a śāstra. A Sanskrit syllable has a sound, a meaning, a pitch and a length. Piṅgala keeps the length and throws away the rest, and reduces it to two values, light or heavy. Once a syllable is one of two things, counting metres becomes arithmetic. That single decision is what makes the whole of Module I possible.

In a program. Here is the same abstraction, run.

# A metre, abstracted to nothing but the weight of each syllable.
LINES = [
    ("a ga na tri pa da", [1, 1, 1, 1, 1, 1]),
    ("ma dhu ra ya ti", [1, 1, 1, 2, 1]),
    ("so ma pa la", [2, 1, 1, 1]),
]

for text, weights in LINES:
    pattern = "".join("G" if w == 2 else "L" for w in weights)
    print("%-22s %-8s syllables %d, heavy %d"
          % (text, pattern, len(pattern), pattern.count("G")))
a ga na tri pa da      LLLLLL   syllables 6, heavy 0
ma dhu ra ya ti        LLLGL    syllables 5, heavy 1
so ma pa la            GLLL     syllables 4, heavy 1

The words have gone. What is left is a string over two symbols, which is a thing a rule can be stated about.

The cost. An abstraction is a decision to be blind to something. Piṅgala's abstraction cannot express which syllables rhyme, because rhyme is part of what was discarded.

Algorithm design

The habit. State the steps so exactly that someone who does not understand the problem can carry them out and get the right answer.

In a śāstra. Halāyudha's own words on the prastāra rule are a procedure: write a row of all heavy syllables; then find the first heavy syllable, make it light, make what precedes it heavy, leave what follows alone; repeat until the row is all light. Nothing in that requires you to know what a metre is.

In a program. The same three sentences, in Python, in [Prastāra in Code].

The test. Give the steps to somebody outside the subject. If they need to ask you a question, the steps are a description and not an algorithm. [The Prastāra Rule Read as an Algorithm] applies that test properly, against the five standard properties.

The four habits as a table

HabitWhat you doPiṅgala's instanceWhat it costs
Decompositionsplit into separately solvable partsone syllable, then a growth rulemore parts to keep consistent
Pattern recognitionreuse a solution, not an answermetres by weight are combinationsa false pattern misleads confidently
Abstractiondiscard what does not bear on the questiona syllable becomes one of two symbolsyou become blind to what was discarded
Algorithm designstate unambiguous finite stepsthe prastāra procedurerigidity: the steps fit one shape of problem
munotes.in36

Computational Thinking, and the Four Habits It Names

What computational thinking is NOT

It is not programming. No machine appears in any of the four habits. A student who thinks the phrase means writing code will misread MU's objective.

It is not a claim that everything is computable. Deciding whether a general connection holds, which is the load-bearing step of every inference in Module II, is not made mechanical by any of the four.

It is not modern. This is the claim MU's syllabus is actually making, and it is the one you should be able to argue in five marks: the habits are habits of anyone who must transmit an exact procedure without being present to explain it.

Quick revision

  • Four habits: decomposition, pattern recognition, abstraction, algorithm design. MU's CO 1 names computational thinking.
  • Decomposition: Piṅgala solves one syllable plus a growth rule. Test: can you check one part alone?
  • Pattern recognition: metres by weight are combinations. Trap: a pattern is a hypothesis.
  • Abstraction: a syllable becomes one of two symbols, and that one decision makes Module I arithmetic. Cost: blindness to what was discarded.
  • Algorithm design: steps a person outside the subject can follow. Test: do they have to ask a question?

Test yourself

1. Name the four habits and give Piṅgala's instance of abstraction.

Decomposition, pattern recognition, abstraction, algorithm design. Abstraction: a syllable's sound, meaning and pitch are discarded and only its weight is kept, reduced to two values, so that a metre becomes a string over two symbols.

2. What is the test of whether you have really decomposed a problem?

Whether one part can be checked without running the others. If it cannot, the text has been divided but the problem has not.

3. Give the cost of abstraction, with an example from this paper.

You become blind to what you discarded. Piṅgala's scheme cannot say anything about rhyme or meaning, because both were thrown away when the syllable was reduced to its weight.

4. Why is it wrong to read MU's "computational thinking" as meaning programming?

Because none of the four habits mentions a machine. They are habits of stating a problem and a procedure exactly, which is why a tradition transmitting a discipline through memory develops them without any computer at all.

Contents This chapter on its own page

munotes.in37

Chapter Thirteen

Piṅgala's Chandaḥśāstra, and the Text We Are Reading

Syllabus topic Module 1, "Piṅgala’s Chandaḥśāstra as Binary Encoding and Combinatorial Generation", "Study of prosodic structures in Piṅgala’s text and its algorithmic features"

In one line

The Chandaḥśāstra is a short Sanskrit treatise on metre in eight chapters, and its eighth chapter is the one this paper is about.

In the wording you can write in an examination: the Chandaḥśāstra, also called the Chandaḥsūtra, attributed to Piṅgala, is a sūtra text on Sanskrit prosody in eight chapters. Its eighth chapter states the pratyayas, the operations on the set of metrical patterns, which include the enumeration of all patterns, the retrieval of a pattern from its index and of an index from its pattern, the counting of patterns without enumerating them, and the array now known as the Meru-prastāra.

What the text is, and what it is not

It is a treatise on metre. Its subject is which arrangements of light and heavy syllables count as which named metre. It is not a book of mathematics and it never claims to be.

Its eighth chapter is mathematical anyway. Once metre has been reduced to strings over two symbols, the questions a prosodist naturally asks are combinatorial: how many are there, which is the fiftieth, where does this one sit. Chapter eight answers those, and that is why a computer science syllabus can use it.

It is very old and its date is disputed. Estimates in the scholarly literature range across several centuries. This book does not need a date and does not assert one, and neither should your answer. What it needs is that the text is pre-modern and that the operations are in it, both of which are secure.

The commentary, which is what we actually read

A sūtra text of four syllables a rule cannot be read alone. The commentary this book quotes is Halāyudha's Mṛtasañjīvanī, and the eighth chapter's colophon in the 1874 edition names it in as many words: here ends the eighth chapter of the Mṛtasañjīvanī, the commentary on metre made by Śrī Bhaṭṭa Halāyudha.

That matters for how you should read every Piṅgala quotation in this book. The sūtra gives the rule in outline; the commentary gives the procedure. Where this book says "Piṅgala's rule is", what is being reported is the sūtra as Halāyudha works it, and the chapter says which.

The edition this book quotes

ItemWhat it is
TitleChhandaḥ Sūtra of Piṅgala Ācārya, with the commentary of Halāyudha
EditorPaṇḍita Visvanātha Śāstrī
SeriesBibliotheca Indica, published by the Asiatic Society of Bengal
PrintedCalcutta, Ganesa Press, 1874

Two further Devanagari editions were fetched and used only as witnesses, to check a sūtra's number and its reading. They are listed in the subject's findings file.

The warning about sūtra numbers

This is the kind of thing a book should tell you rather than hide.

munotes.in38

Piṅgala's Chandaḥśāstra, and the Text We Are Reading

Editions of the eighth chapter do not number it alike. The same sūtra, "rūpe śūnyam", is numbered 26 in one of this book's witnesses and 28 in another. There is no error in either: the chapter's sūtra divisions are themselves an editorial judgement, and different editors divide differently.

Worked. So this book prints a number only where two of its three witnesses agree. Where they do, the number is given with the edition named. Where they do not, the sūtra is named by its opening words and its place in the sequence, and the disagreement is stated.

The numbers that are secure, agreeing across witnesses in the Asiatic Society edition's numbering, are these, and they are the ones this book uses.

NumberOpening wordsWhat it states
8.20dvikau glauthe one-syllable table: ga above, la below
8.21miśrau cathe two-syllable table
8.22pṛthaglā miśrāḥthree syllables and upward, by repeating the shorter table
8.23vasavas trikāḥthere are eight three-syllable groups
8.24lardheon halving, a laghu
8.25saike gwith one added and halved, a guru
8.26pratilomaguṇaṃ dvirlādyamthe row number by reverse doubling
8.27tato'pyekaṃ jahyātand from that subtract one

The saṅkhyā group, which computes the number of patterns without listing them, follows immediately after 8.27 in the same chapter. Its numbering is where the witnesses diverge, so [Saṅkhyā: Counting the Rows Without Writing Them] names those sūtras by their opening words.

The six pratyayas, and what Halāyudha says about the sixth

A pratyaya is an operation on the set of metrical patterns. Halāyudha treats them as a group and closes the chapter by saying the group is complete.

  1. Prastāra, laying out every pattern in order.
  2. Naṣṭa, recovering a pattern that has been lost, from its row number.
  3. Uddiṣṭa, finding the row number of a pattern pointed at.
  4. Lagakriyā, how many patterns have a given number of light syllables.
  5. Saṅkhyā, how many patterns there are altogether.
  6. Adhvayoga, how much space the written table takes.

On the sixth he records, in effect, that some count it and that it is omitted here because it is very slight. [Lagakriyā and Adhvayoga: The Rest of the Pratyayas] gives both, because a question may name one MU's own labels do not.

What this book claims about the text, and what it does not

Claimed. The eighth chapter states the operations listed above; Halāyudha works them; the arithmetic they describe is exactly the arithmetic this book prints, because every one of them has been implemented and run.

Not claimed. That Piṅgala had the idea of a number base. That he wrote the operations as a mathematician would. That anything in the text is a program. The operations are algorithms and there is no machine, and the distinction is kept throughout.

munotes.in39

Piṅgala's Chandaḥśāstra, and the Text We Are Reading

Not claimed, and worth saying twice. That this book's readings settle disputed questions in Sanskrit prosody. Where the tradition disagrees, the disagreement is reported.

Quick revision

  • The Chandaḥśāstra, attributed to Piṅgala, is a sūtra text on Sanskrit metre in eight chapters. Chapter 8 carries the pratyayas.
  • The commentary quoted is Halāyudha's Mṛtasañjīvanī, named in the chapter's own colophon. The sūtra gives the rule; the commentary gives the procedure.
  • The edition quoted is the Bibliotheca Indica text, ed. Visvanātha Śāstrī, Asiatic Society of Bengal, Calcutta, 1874.
  • Editions number chapter 8 differently: "rūpe śūnyam" is 26 in one witness and 28 in another. Numbers are printed only where witnesses agree.
  • Six pratyayas: prastāra, naṣṭa, uddiṣṭa, lagakriyā, saṅkhyā, adhvayoga. Halāyudha says the sixth is usually omitted as slight.

Test yourself

1. What is the Chandaḥśāstra about, and why does a computer science paper use its eighth chapter?

It is a treatise on Sanskrit metre. Its eighth chapter reduces metre to strings over two symbols and states operations on the whole set of such strings, so the questions it answers are combinatorial and algorithmic.

2. Name the commentary this book quotes and say why a commentary is needed at all.

Halāyudha's Mṛtasañjīvanī. A sūtra of two or three words states a rule in outline and omits the procedure, so the procedure has to be read from the commentary.

3. Why does this book sometimes name a sūtra by its opening words instead of by a number?

Because editions divide the eighth chapter differently and so number it differently. "Rūpe śūnyam" is numbered 26 in one witness and 28 in another, so a bare number would be a claim the sources do not support.

4. List the six pratyayas and say which one Halāyudha treats as optional.

Prastāra, naṣṭa, uddiṣṭa, lagakriyā, saṅkhyā, adhvayoga. He records that some count the sixth, adhvayoga, and that it is left out as very slight.

Contents This chapter on its own page

munotes.in40

Chapter Fourteen

Syllables, Laghu and Guru: Sanskrit Metre Without Sanskrit

Syllabus topic Module 1, "Study of prosodic structures in Piṅgala’s text and its algorithmic features", "Laghu–Guru representation as binary encoding"

In one line

A Sanskrit line of verse is a fixed number of syllables, each of which is either light or heavy, and that is all the prosody this paper needs.

In the wording you can write in an examination: Sanskrit metre is quantitative, that is, it is governed by the duration of syllables rather than by stress. A syllable is laghu, light, if it has a short vowel and is not closed; otherwise it is guru, heavy. A metre of fixed syllable count is defined by the sequence of light and heavy syllables in its line, and a stanza conventionally has four lines.

Why quantity and not stress

English verse counts stresses: the beat falls on some syllables and not others. Sanskrit verse counts time. A light syllable takes one unit and a heavy syllable takes two, and a line is a pattern of those durations.

This is the fact that makes the rest of Module I possible, and it is worth seeing why. Stress is a matter of degree and of dialect. Duration, as the tradition defines it, is decidable: you can look at a syllable and say which it is, by rule. A property that is decidable by rule can be encoded, and once it is encoded it can be counted.

The vocabulary, defined once

Syllable. A vowel, together with any consonants that attach to it. Sanskrit is written so that this is unambiguous.

Laghu, light, written L in this book. A syllable with a short vowel that is not closed by a consonant.

Guru, heavy, written G. Everything else: a syllable with a long vowel, with a diphthong, or with a short vowel closed by a consonant.

Mātrā, a unit of duration. A light syllable is one mātrā, a heavy syllable two. Vasu's own note on Pāṇini records the convention plainly: a short vowel has one mātrā, a long vowel has two.

Pāda, a line, literally a foot or quarter. Vṛtta, a metre defined by syllable count. Stanza, conventionally four pādas.

The rule for deciding a syllable, and the one complication

The rule has two halves and the second is the one students miss.

By its vowel. A short vowel makes the syllable light; a long vowel or a diphthong makes it heavy. This half is easy.

By what follows. A syllable with a short vowel becomes heavy if a consonant closes it, because the closing consonant adds time. So the same vowel can be light in one word and heavy in another, depending on what comes next.

The consequence for this paper is small but should be stated: the weight of a syllable is a property of its position in the line, not only of the syllable itself. That is why a metre is a sequence and cannot be worked out one syllable at a time in isolation.

munotes.in41

Syllables, Laghu and Guru: Sanskrit Metre Without Sanskrit

Worked example: three lines scanned

The lines below are transliterated syllable by syllable and marked. What is being demonstrated is only the procedure, and nothing turns on the choice of lines.

Line, by syllableWeightsPatternMātrās
a, ga, na, tri, pa, da1, 1, 1, 1, 1, 1LLLLLL6
ma, dhu, ra, yā, ti1, 1, 1, 2, 1LLLGL6
so, ma, pa, la2, 1, 1, 1GLLL5

Two things to notice. Two lines of different syllable counts can have the same total duration, which is why the tradition has a second family of metres counted by mātrā rather than by syllable. And the pattern column is a string over two symbols, which is the only thing the rest of Module I will use.

The seven metre classes

The tradition names classes of metre by the number of syllables in a line. Halāyudha's own commentary on the saṅkhyā rule runs through seven of them, and this book quotes his figures in [Saṅkhyā: Counting the Rows Without Writing Them].

ClassSyllables in a linePatterns possible
gāyatrī664
uṣṇih7128
anuṣṭubh8256
bṛhatī9512
paṅkti101024
triṣṭubh112048
jagatī124096

Every figure in the third column is a power of two, and every one of them is stated in the commentary rather than calculated by this book. That is a fact worth pausing on: a prosodist writing on metre has produced, in the course of his subject, a table of powers of two from 64 to 4096.

What a metre is, formally

Putting the vocabulary together gives a definition a computer science student can work with.

A metre of n syllables is a string of length n over the alphabet {L, G}. The set of all such strings is what the next several chapters enumerate, index and count. Nothing about Sanskrit is needed after this point, which is the payoff of the abstraction.

What this does NOT mean

It does not mean every string is a real metre. Most of the 4096 twelve-syllable patterns have no name and were never used. The tradition's interest is in which ones are named; Piṅgala's chapter eight is about the whole space anyway, and that is precisely why it becomes mathematics.

It does not mean quantity is the whole of Sanskrit prosody. There are metres counted by total mātrā rather than by syllable count, and there are rules about pauses. Neither is on this syllabus.

It does not mean the scansion is mechanical in every case. There are conventional licences, and editors disagree about particular lines. None of that affects the combinatorics.

munotes.in42

Syllables, Laghu and Guru: Sanskrit Metre Without Sanskrit

Quick revision

  • Sanskrit metre is quantitative: it counts duration, not stress.
  • Laghu, light, L: a short vowel, not closed. Guru, heavy, G: long vowel, diphthong, or short vowel closed by a consonant. One mātrā against two.
  • A syllable's weight can depend on what follows it, so a metre is a sequence and not a bag.
  • A pāda is a line; four pādas make a stanza; a vṛtta is a metre defined by syllable count.
  • Seven classes from gāyatrī at 6 syllables to jagatī at 12, with 64 to 4096 patterns, all powers of two, all stated in Halāyudha's commentary.
  • Formally: a metre of n syllables is a string of length n over {L, G}.

Test yourself

1. Define laghu and guru, and give the rule that makes a short vowel heavy.

Laghu is a light syllable, one mātrā: a short vowel not closed by a consonant. Guru is heavy, two mātrās: a long vowel, a diphthong, or a short vowel closed by a consonant. The closing consonant is what makes a short vowel heavy.

2. Why does the weight of a syllable sometimes depend on the next syllable?

Because a consonant that closes the syllable adds duration, and whether a syllable is closed depends on what follows it. So scansion is done over the line, not syllable by syllable in isolation.

3. Why is quantitative metre easier to formalise than stress metre?

Because weight is decidable by a stated rule, while stress is a matter of degree and of dialect. A property decidable by rule can be encoded as a symbol, and symbols can be counted.

4. Give the formal definition of a metre of n syllables, and say how many there are for the anuṣṭubh.

A string of length n over the alphabet of two symbols, light and heavy. The anuṣṭubh has eight syllables, so there are 256 patterns, which is the figure Halāyudha's commentary itself gives.

Contents This chapter on its own page

munotes.in43

Chapter Fifteen

Laghu and Guru as One Bit

Syllabus topic Module 1, "Laghu–Guru representation as binary encoding", "Binary number systems"

In one line

Because a syllable is one of exactly two things, a metre of n syllables carries exactly n bits of information.

In the wording you can write in an examination: a bit is the amount of information in a choice between two equally possible alternatives. Since Piṅgala's abstraction assigns each syllable of a metre one of two values, laghu or guru, a line of n syllables is a sequence of n binary digits, the set of such lines has 2 to the power n members, and there is a one-to-one correspondence between metrical patterns and n-bit binary words.

The claim, in four steps

Each step is small, and stating them separately is what keeps the claim honest.

Step one. A syllable, after Piṅgala's abstraction, has exactly two possible values.

Step two. A choice between two alternatives is one bit. That is the definition of a bit, not an analogy.

Step three. A line of n syllables is n independent such choices, so it carries n bits.

Step four. Therefore the number of distinct lines of n syllables is 2 multiplied by itself n times.

Nothing in those four steps requires Piṅgala to have had the idea of a bit, or of a base, or of information. The steps are about the structure he set up, and the structure is two-valued whatever he thought about it.

The correspondence, written out

Fix a direction for the mapping and everything downstream becomes arithmetic.

SyllableSymbolBit
laghu, lightL1
guru, heavyG0

Read the leftmost syllable as the units place, the next as the twos place, and so on. Then every metre of n syllables is a number from 0 to 2 to the power n minus 1, and every such number is a metre.

Two warnings about that paragraph, because both choices in it are conventions and not discoveries.

Which symbol is one is a convention. Nothing in the text says laghu is one. What the text fixes is the ORDER of the table, and the assignment above is the one that makes the order come out as ordinary counting. [Binary Number Systems, and Exactly Where Piṅgala's Order Agrees] establishes that, and proves it.

Which end is the units place is a convention too, and it is the unusual one. Piṅgala's order puts the fastest-changing symbol on the left, which is the opposite of how a number is written in ordinary notation. That is the one real difference between his table and a table of binary numerals, and it is a difference in presentation only.

Worked example: three syllables, both ways

SYLL = {"L": 1, "G": 0}

def value(pattern):
    """The number a pattern stands for: leftmost syllable is the units place."""
    return sum(SYLL[ch] << i for i, ch in enumerate(pattern))

def pattern(number, n):
    """The inverse: n syllables carrying that number."""
    return "".join("L" if (number >> i) & 1 else "G" for i in range(n))

print("pattern  value  back")
for v in range(8):
    p = pattern(v, 3)
    print("%-8s %5d  %d" % (p, value(p), value(pattern(v, 3))))

print()
print("distinct patterns of 3 syllables:", len({pattern(v, 3) for v in range(8)}))
print("bits carried by a 3 syllable line:", 3)
munotes.in44

Laghu and Guru as One Bit

pattern  value  back
GGG          0  0
LGG          1  1
GLG          2  2
LLG          3  3
GGL          4  4
LGL          5  5
GLL          6  6
LLL          7  7

distinct patterns of 3 syllables: 8
bits carried by a 3 syllable line: 3

The mapping is a bijection: every pattern has one value, every value has one pattern, and the eight patterns are the eight numbers.

What "n bits" buys you

This is the part that makes the abstraction useful rather than decorative.

Counting becomes arithmetic. You do not have to write the table to know how big it is. [Saṅkhyā: Counting the Rows Without Writing Them] is exactly this.

Addressing becomes arithmetic. You do not have to write the table to find its fiftieth row. [Naṣṭa: From a Row Number Back to the Pattern] is exactly this, and it is the most striking of Piṅgala's rules for a computing student, because it is random access into a structure that was never built.

Sub-counting becomes combinatorics. How many of the patterns have exactly three light syllables is a question about bit counts, and it has a closed-form answer. [Meru-Prastāra: Halāyudha's Staircase] is that answer.

What it does NOT license

Four overclaims are made about Piṅgala regularly, and all four should be refused in an answer that is asked to assess critically.

"Piṅgala invented binary." No. He set up a two-valued encoding and stated operations on it. There is no statement in the text of a positional number system with base two, no arithmetic on such numbers, and no notion of a base at all.

"Piṅgala invented the bit." No. A bit is a measure of information, defined in the twentieth century. What is true is that his encoding carries one bit per syllable, which is a fact about his encoding and not a claim about his intent.

"The prastāra is a binary counter." Almost. Its order is the order of binary counting, and that is proved later. It is not described as counting, and it is not used to do arithmetic.

"This makes the Chandaḥśāstra a work of computer science." No. It makes the Chandaḥśāstra a work in which several problems of a kind computer science later made central are posed and solved, which is a more interesting claim because it is true.

munotes.in45

Laghu and Guru as One Bit

The honest summary

What Piṅgala did is more specific and more impressive than the overclaims. He chose an encoding in which the questions of his subject became decidable by arithmetic, and then stated the arithmetic. Choosing a representation so that the problems become tractable is the whole of what [Knowledge Representation: The Modern Name For It] is about, and it is the highest praise available.

Quick revision

  • A bit is a choice between two alternatives. A syllable, abstracted, is such a choice.
  • A line of n syllables carries n bits, so there are 2 to the power n lines.
  • With laghu as one, guru as zero, and the leftmost syllable as the units place, patterns and numbers correspond one to one.
  • Both of those choices are conventions; what the text fixes is the ORDER of the table.
  • What follows: counting, addressing and sub-counting all become arithmetic.
  • What is not licensed: that Piṅgala invented binary, the bit, or a counter. What is true: he chose an encoding that made his questions arithmetical.

Test yourself

1. Show in four steps why a metre of n syllables has 2 to the power n forms.

A syllable has two possible values; a choice between two is one bit; a line of n syllables is n independent choices, so n bits; n independent binary choices give 2 multiplied by itself n times distinct results.

2. Which two choices in the standard mapping are conventions rather than facts about the text?

Assigning one to laghu rather than to guru, and treating the leftmost syllable as the units place. The text fixes the order of the table, not the arithmetic reading of it.

3. What is unusual about the position of the units place in Piṅgala's order, and does it matter?

The fastest-changing symbol is on the left, the opposite of ordinary numeral notation. It is a difference in presentation only: the order is still the order of counting.

4. A classmate says Piṅgala invented binary numbers. Correct them in two sentences.

He set up a two-valued encoding of syllables and stated operations that enumerate, index and count the resulting patterns. He nowhere states a positional number system with base two, performs arithmetic on such numbers, or uses the notion of a base at all.

Contents This chapter on its own page

munotes.in46

Chapter Sixteen

Prastāra: The Table of Every Pattern

Syllabus topic Module 1, "Prastāra (systematic enumeration of patterns)"

In one line

The prastāra is the table of every possible pattern of a metre, written out in one fixed order, and Piṅgala states the rule that produces the order.

In the wording you can write in an examination: prastāra is the systematic enumeration, in a prescribed sequence, of all the metrical patterns of a given syllable count. The sequence is generated by a rule which, from any row, derives the row after it, beginning from the row of all heavy syllables and ending at the row of all light ones. For n syllables the table has 2 to the power n rows and each row appears exactly once.

Why the order is the interesting part

Anybody can list eight patterns. What Piṅgala supplies is an order, and an order is worth more than a list for three reasons that matter to this paper.

It makes the table reproducible. Two scribes in two places produce the same table in the same sequence, so "the sixth row of the gāyatrī" names one thing.

It makes the table addressable. Once rows have fixed numbers, you can ask for row 41 and be answered. The two rules of [Index Retrieval, and Proving Two Rules Are Inverses] depend entirely on the order being fixed.

It makes the table generable without being stored. If the next row is a function of this one, you never need the whole table in front of you, which matters when the table has 4096 rows and you are working from memory.

The provisions

Four sūtras build the table, and their numbers agree across this book's witnesses.

8.20, dvikau glau. Halāyudha: write a ga above and a la below, and because the sūtra says "two", place two such. That gives the one-syllable table, which has two rows.

8.21, miśrau ca. He says this shows the two-syllable prastāra, and concludes: thus the two-syllable prastāra becomes fourfold.

8.22, pṛthaglā miśrāḥ. This is the general rule, and it is stated as a repetition. He says that thus the three-syllable prastāra is established, and then: having set down the three-syllable prastāra twice, ga and la are to be given separately and mixed, and that is the four-syllable prastāra. Then, in the same passage: this sūtra is to be applied again and again until the desired prastāra is reached.

8.23, vasavas trikāḥ. The Vasus, that is eight, are the three-syllable groups. His commentary continues: hence there are sixteen of four syllables and thirty-two of five; thus sixty-four are the gāyatrī forms, from all-heavy to all-light.

The rule as a scribe applies it

Halāyudha's procedure, put as numbered steps. This is the form the rest of the book uses, and it is the form [The Prastāra Rule Read as an Algorithm] tests.

munotes.in47

Prastāra: The Table of Every Pattern

  1. Write a first row of n heavy syllables.
  2. Look along the current row from the left and find the first heavy syllable.
  3. If there is none, stop: the table is complete.
  4. Otherwise make that syllable light, make every syllable to its left heavy, and leave every syllable to its right exactly as it is.
  5. Write the result as the next row and go back to step 2.

Step 4 is the whole rule and it has three clauses. Students who get the table wrong have almost always dropped the middle clause, which resets the left-hand part.

The three-syllable table

This is the program's own output, from the file printed in [Prastāra in Code].

1 G G G

2 L G G

3 G L G

4 L L G

5 G G L

6 L G L

7 G L L

8 L L L

Worked. Trace the rule once to see it working. Row 4 is L L G. The first heavy syllable is the third. Make it light, make the first two heavy, leave nothing to the right: G G L, which is row 5. That is the step where the reset matters, and where a careless reader would write L L L.

The same table the other way: Halāyudha's recursion

The rule above walks the table one row at a time. Sūtra 8.22 as Halāyudha explains it gives a second construction, and it is the one a programmer will find natural.

To build the table for n syllables: write the table for n minus 1 syllables twice, one copy under the other; then in the new syllable's column, put heavy against every row of the first copy and light against every row of the second.

Check it for n equal to 2. The one-syllable table is G, L. Write it twice: G, L, G, L. Fill the second column with G, G then L, L. The result is GG, LG, GL, LL, which is the four rows the commentary says the two-syllable prastāra has.

The new syllable goes on the right. This is the one detail that is easy to get backwards, and getting it backwards produces a table of the right size in the wrong order. The program in [Prastāra in Code] builds the table both ways and asserts they are identical, which is how that mistake was caught.

Why the two constructions agree

They agree because the rule in step 4 changes the rightmost part of the row least often. The first syllable alternates every row; the second changes every two rows; the k-th changes every 2 to the power k minus 1 rows. So the table splits exactly in half on the last syllable, which is what the recursive construction assumes.

munotes.in48

Prastāra: The Table of Every Pattern

That observation is also the whole content of [Binary Number Systems, and Exactly Where Piṅgala's Order Agrees], stated ahead of time.

What the prastāra is NOT

It is not a list of metres in use. Most of its rows are patterns no poet ever wrote. Its subject is the whole space, which is exactly why it becomes mathematics.

It is not a table that has to be stored. The rule is the table. A scribe with the rule and a row can produce the next row, and the retrieval rules can produce any row from its number.

It is not the only possible order. Other orders exist and one of them, the reverse, appears in later prosodic literature. Piṅgala's is the one his retrieval rules are built for, and mixing two orders is the commonest way to get a naṣṭa answer wrong.

Quick revision

  • Prastāra: every pattern of n syllables, in one fixed order, from all-heavy to all-light, 2 to the power n rows, no repeats.
  • 8.20 dvikau glau gives the one-syllable table; 8.21 miśrau ca the two-syllable; 8.22 pṛthaglā miśrāḥ generalises by repetition; 8.23 vasavas trikāḥ states eight for three syllables.
  • The rule: find the first heavy syllable from the left, make it light, make everything to its left heavy, leave everything to its right alone.
  • The recursion: write the shorter table twice, and fill the new column, which is on the RIGHT, with heavy against the first copy and light against the second.
  • The commentary's own figures: 8, 16, 32, 64 for three to six syllables.

Test yourself

1. State the prastāra rule in your own words and apply it to L L G.

Find the first heavy syllable from the left, make it light, make everything to its left heavy, and leave everything to its right unchanged. In L L G the first heavy syllable is the third, so it becomes light and the first two are reset to heavy, giving G G L.

2. Which clause of the rule do students drop, and what goes wrong?

The clause that resets everything to the left of the changed syllable to heavy. Dropping it turns L L G into L L L instead of G G L, and the table then repeats rows and misses others.

3. Give Halāyudha's recursive construction and say where the new syllable goes.

Write the table for one fewer syllable twice, one copy beneath the other, then fill the new column with heavy against the first copy and light against the second. The new column is at the right-hand end of the row.

4. Why is an order more useful than a list?

munotes.in49

Prastāra: The Table of Every Pattern

Because it is reproducible, so a row number names one thing; it is addressable, so a row can be asked for by number; and it lets the table be generated from a rule instead of stored.

Contents This chapter on its own page

munotes.in50

Chapter Seventeen

The Prastāra Rule Read as an Algorithm

Syllabus topic Module 1, "Prastāra (systematic enumeration of patterns)", "Algorithmic generation"

In one line

Halāyudha's prastāra rule is an algorithm and not just a description, and this chapter proves it by checking the five properties an algorithm has to have.

In the wording you can write in an examination: an algorithm is a finite sequence of unambiguous instructions which, for every admissible input, terminates after finitely many steps and produces the intended output, each step being effective, that is, executable by a person or machine with no further interpretation. The prastāra rule satisfies all five conditions, and the syllabus's label "algorithmic generation" is earned rather than asserted.

Why this test is worth running

It is easy to say an old text "contains an algorithm" and hard to mean anything by it. The five properties below are the standard ones, they are checkable, and running them against a specific classical rule is exactly what MU's course objectives ask for when they say analyse.

It also tells you something when a rule fails one of them. Several of the classical rules in this paper come close and fail a condition, and saying which condition is a better answer than a general enthusiasm.

The five properties, and the rule against each

The rule under test is the one from [Prastāra: The Table of Every Pattern]: first row all heavy; then find the first heavy syllable from the left, make it light, make everything to its left heavy, leave everything to its right alone; stop when no heavy syllable remains.

Input

The property. The algorithm takes zero or more inputs from a stated set.

The rule. It takes one input, n, the number of syllables. Halāyudha's own usage supplies the admissible set: he applies the rule to one, two, three, four, five and six syllables in turn, and the general rule 8.22 is stated for any number. So the input is a positive integer.

Verdict: satisfied. With one honest note: the text nowhere says what happens for n equal to zero, and no classical author asked.

Definiteness

The property. Every step is unambiguous. A reader cannot have to choose.

The rule. Step by step: "find the first heavy syllable from the left" has one answer for any row. "Make it light" has one result. "Make everything to its left heavy" has one result. "Leave everything to its right alone" has one result. Nothing in the rule asks the reader to decide anything.

Verdict: satisfied, and this is the property the classical formulation is strongest on. Compare it with a rule like "choose a convenient starting point", which is a description and not an instruction.

Finiteness

The property. The algorithm terminates after finitely many steps for every admissible input.

Worked. This one needs an argument, not an inspection, and the argument is the interesting part of the chapter.

munotes.in51

The Prastāra Rule Read as an Algorithm

Read each row as a number with light counting one and heavy counting zero, the leftmost syllable being the units place. Then one step of the rule does this: the first heavy syllable, at position k, becomes light, adding 2 to the power k minus 1; and every light syllable to its left becomes heavy, subtracting 1 plus 2 plus 4 and so on up to 2 to the power k minus 2, which totals 2 to the power k minus 1 minus 1.

So the net change is exactly plus one. Each step increases the row's value by one, the value starts at 0 and can never exceed 2 to the power n minus 1, so the rule stops after exactly 2 to the power n minus 1 steps.

Verdict: satisfied, and provably. Notice what this argument gives you beyond termination: it gives the exact number of steps, and it is the reason the order is the order of counting.

Effectiveness

The property. Every step is basic enough to be carried out exactly, in finite time, with the means available.

The rule. The operations are: scan a row left to right; overwrite one symbol; overwrite a prefix. A scribe with a palm leaf can do all three. No step requires arithmetic, and none requires knowing anything about metre.

Verdict: satisfied. And worth noticing: the rule is effective for a human executor with no mathematics, which is exactly what it was designed for.

Output

The property. It produces the intended result.

The rule. It produces every pattern of n syllables, once each, in a fixed order. That is a claim about correctness and it is not obvious from the rule.

The finiteness argument above supplies the proof. If each step adds exactly one to the row's value, and the first row is 0, then the rows take the values 0, 1, 2 and so on up to 2 to the power n minus 1, each exactly once. Since the mapping between patterns and values is one to one, every pattern appears exactly once.

Verdict: satisfied, and proved by the same argument as finiteness. The program in [Prastāra in Code] also checks it by brute force up to fourteen syllables, which is a different kind of assurance.

The five properties as a table

PropertyWhat it requiresThe prastāra ruleEvidence
Inputa stated admissible setone positive integer, the syllable count8.22 stated for any number, applied to 1 up to 6
Definitenessno step needs a choiceevery step has exactly one resultthe wording of the rule itself
Finitenessterminates for every inputstops after 2 to the power n minus 1 stepseach step adds one to the row's value
Effectivenesseach step is basicscan, overwrite one symbol, overwrite a prefixa scribe can perform all three
Outputproduces the intended resultall patterns, once each, in orderthe same increment argument, plus a brute-force check
munotes.in52

The Prastāra Rule Read as an Algorithm

What this does NOT establish

It does not make the Chandaḥśāstra a program. An algorithm is a procedure; a program is an algorithm expressed for a machine. There is no machine, and saying so is part of the answer.

It does not mean Piṅgala thought in these terms. The five properties are a modern test. What has been shown is that his rule passes it, which is a fact about the rule.

It does not mean every rule in the text passes. [Lagakriyā and Adhvayoga: The Rest of the Pratyayas] discusses one that is stated so briefly that its definiteness depends on the commentary, and the honest verdict there is different.

Limits

One limit is worth naming because it recurs. The rule has no stated complexity. Nothing in the text says how long producing the table takes, or notices that it takes time proportional to n times 2 to the power n. The saṅkhyā rule shows that Piṅgala cared about avoiding work, so the absence is interesting rather than damning, but it is an absence. [Algorithmic Generation: What Piṅgala Actually Achieved] gives the costs.

Quick revision

  • The five properties: input, definiteness, finiteness, effectiveness, output.
  • Definiteness is the classical rule's strongest property: no step requires a choice.
  • Finiteness is proved, not observed: each step raises the row's value by exactly one, so the rule halts after 2 to the power n minus 1 steps.
  • The same argument proves correctness: the rows take every value from 0 upward exactly once.
  • Effectiveness: scan, overwrite one symbol, overwrite a prefix. A scribe can do all three with no arithmetic.
  • What is missing from the text: any statement of cost.

Test yourself

1. Name the five properties of an algorithm.

Input, definiteness, finiteness, effectiveness, output.

2. Prove that the prastāra rule terminates, and say how many steps it takes.

Read a row as a number with light as one and the leftmost syllable as the units place. One step turns the first heavy syllable light, adding 2 to the power k minus 1, and resets the lighter syllables to its left, subtracting 2 to the power k minus 1 minus 1. The net change is plus one. Starting from 0 and bounded by 2 to the power n minus 1, the rule halts after exactly 2 to the power n minus 1 steps.

3. How does the same argument prove the table is complete and has no repeats?

munotes.in53

The Prastāra Rule Read as an Algorithm

Because the rows take the values 0, 1, 2 and so on consecutively, and patterns correspond one to one with values, every pattern appears exactly once.

4. Which property is a rule like "begin at a convenient point" failing, and why does it matter?

Definiteness. A reader has to make a choice, so two readers can produce different results, and the procedure is a description rather than an algorithm.

Contents This chapter on its own page

munotes.in54

Chapter Eighteen

Prastāra in Code

Syllabus topic Module 1, "Prastāra (systematic enumeration of patterns)", "Algorithm Specification (Pseudo-code)", "minimum 10 test cases"

In one line

The prastāra rule, written as a program, is eight lines, and the whole of its correctness can be tested.

In the wording you can write in an examination: the prastāra can be implemented either iteratively, by applying the next-row rule from the all-heavy row until the all-light row is reached, or recursively, by writing the table for one fewer syllable twice and filling the new column. Both produce the same table in the same order, and the identity of the two outputs is itself a test.

Problem statement, in MU's own form

IKS concept as CS concept: Piṅgala's prastāra as exhaustive enumeration of a binary space.

Statement. Given a positive integer n, produce every sequence of n symbols drawn from a two-symbol alphabet, in the order that Chandaḥśāstra 8.20 to 8.23 prescribes, and verify that the output has 2 to the power n members, no repetitions, the all-heavy row first and the all-light row last.

Conceptual mapping table

This is heading three of her required submission, and it is the heading most students leave out.

Classical elementComputer science element
a syllable, laghu or guruone bit, or a symbol from a two-letter alphabet
a metre of n syllablesa string of length n, or an n-bit word
the prastārathe complete enumeration of that string space
the next-row rule of 8.22the successor function on the space
the all-guru first rowthe initial state
the all-laghu last rowthe termination condition
Halāyudha's "again and again until the desired prastāra"a loop, or a recursive call with a base case

Algorithm specification, in pseudo-code

Heading four. Pseudo-code is not code: it has no syntax to get wrong and it states the steps, which is what she is asking for.

ALGORITHM Prastara(n)

INPUT n, a positive integer

OUTPUT the ordered list of all patterns of n syllables

row <- n copies of GURU

rows <- [row]

while row contains a GURU do

k <- the position of the first GURU in row, counting from the left

row <- (k-1 copies of GURU) + LAGHU + (row after position k, unchanged)

rows <- rows + [row]

end while

return rows

Two things to notice, because they are where marks are won. The loop condition is the termination condition of the classical rule, stated once. And the assignment inside the loop has three parts, matching the three clauses of Halāyudha's own sentence: the prefix is reset, the found syllable is changed, and the suffix is untouched.

Working code

Heading five.

GURU, LAGHU = "G", "L"

def prastara(n):
    """Every pattern of n syllables, in Pingala's order."""
    row = [GURU] * n
    rows = ["".join(row)]
    while GURU in row:
        k = row.index(GURU)
        row = [GURU] * k + [LAGHU] + row[k + 1:]
        rows.append("".join(row))
    return rows

for r, pattern in enumerate(prastara(3), start=1):
    print("%2d  %s" % (r, " ".join(pattern)))
munotes.in55

Prastāra in Code

 1  G G G
 2  L G G
 3  G L G
 4  L L G
 5  G G L
 6  L G L
 7  G L L
 8  L L L

Line by line, because a viva will ask. row.index(GURU) is step two of the rule, finding the first heavy syllable from the left. The three pieces of the next line are the three clauses of step four. The while condition is step three, and it is checked before each pass, which is why an all-light row ends the table instead of being processed.

Test cases

Heading six. She asks for a minimum of ten. Here they are as two groups: five that check the size of the output, and five that check its shape over every n from one to twelve.

GURU, LAGHU = "G", "L"

def prastara(n):
    row = [GURU] * n
    rows = ["".join(row)]
    while GURU in row:
        k = row.index(GURU)
        row = [GURU] * k + [LAGHU] + row[k + 1:]
        rows.append("".join(row))
    return rows

def prastara_recursive(n):
    """Halayudha on 8.22: the shorter table twice, new column on the right."""
    if n == 1:
        return [GURU, LAGHU]
    prev = prastara_recursive(n - 1)
    return [r + GURU for r in prev] + [r + LAGHU for r in prev]

TESTS = [
    ("one syllable gives two rows",        1, 2),
    ("two syllables give four rows",       2, 4),
    ("three syllables give eight rows",    3, 8),
    ("the gayatri gives sixty four rows",  6, 64),
    ("the jagati gives 4096 rows",        12, 4096),
]

print("%-40s %-8s %-8s %s" % ("test", "rows", "expected", "result"))
for name, n, expected in TESTS:
    rows = prastara(n)
    ok = len(rows) == expected
    print("%-40s %-8d %-8d %s" % (name, len(rows), expected, "pass" if ok else "FAIL"))

print()
CHECKS = [
    ("the first row is all guru",          lambda n: prastara(n)[0] == "G" * n),
    ("the last row is all laghu",          lambda n: prastara(n)[-1] == "L" * n),
    ("no row repeats",                     lambda n: len(set(prastara(n))) == 2 ** n),
    ("every row has n syllables",          lambda n: all(len(r) == n for r in prastara(n))),
    ("the recursion gives the same table", lambda n: prastara(n) == prastara_recursive(n)),
]
for name, check in CHECKS:
    results = [check(n) for n in range(1, 13)]
    print("%-40s %s for n = 1 to 12" % (name, "pass" if all(results) else "FAIL"))
test                                     rows     expected result
one syllable gives two rows              2        2        pass
two syllables give four rows             4        4        pass
three syllables give eight rows          8        8        pass
the gayatri gives sixty four rows        64       64       pass
the jagati gives 4096 rows               4096     4096     pass

the first row is all guru                pass for n = 1 to 12
the last row is all laghu                pass for n = 1 to 12
no row repeats                           pass for n = 1 to 12
every row has n syllables                pass for n = 1 to 12
the recursion gives the same table       pass for n = 1 to 12
munotes.in56

Prastāra in Code

The first four expected values, 2, 4, 8 and 64, are Halāyudha's own figures from his commentary on 8.23. The test is therefore against the text and not against our own arithmetic, which is the strongest form the test can take.

The bug the last test caught

The final check compares the iterative table with the recursive one, and it is there for a reason.

The first version of prastara_recursive in this book's own program file wrote GURU + r instead of r + GURU, putting the new syllable at the left-hand end of the row. That version produced a table of exactly the right size, with no repeats, beginning all-heavy and ending all-light. Every test except the comparison passed. The order was wrong, and nothing but the comparison would have noticed.

The lesson for an internal assessment is general: a test that only checks the size of the output cannot see a wrong order. Say that in your limitations section and you have said something real.

Complexity and limitations

Heading seven.

Time. The loop runs 2 to the power n minus 1 times, and each pass copies a row of n symbols, so the work is proportional to n times 2 to the power n. There is no faster way, because the output itself is that big.

Space. As written, the whole table is held in memory: again n times 2 to the power n. For n equal to twelve that is about 49,000 symbols, which is nothing; for n equal to thirty it would be about thirty-two gigabytes. If you only need one row at a time, do not build the list, and if you need one particular row, do not build the table at all. [Naṣṭa: From a Row Number Back to the Pattern] is the alternative.

The recursive version. It builds every intermediate table, so it does about twice the work of the iterative one and uses recursion depth n. It is in the book to be compared with, not to be preferred.

Limitation. Neither version can produce row 5,000,000 of a thirty-syllable metre without producing everything before it. That is not a flaw in the code; it is a property of enumeration, and it is exactly the gap Piṅgala's retrieval rules fill.

Quick revision

  • Iterative implementation: start all-heavy, find the first heavy from the left, make it light, reset its left, keep its right, until all-light.
  • Recursive implementation, from 8.22: the shorter table twice, the new column appended on the RIGHT.
  • Time and space are both proportional to n times 2 to the power n; that is the size of the output, so it cannot be beaten.
  • MU's five submission headings used here: problem statement, conceptual mapping, pseudo-code, working code, test cases.
  • A test that checks only the size of the output cannot detect a wrong order. Compare two independent constructions.
munotes.in57

Prastāra in Code

Test yourself

1. Write the prastāra rule as pseudo-code in five lines.

Set row to n copies of guru and start the list with it. While the row contains a guru: let k be the position of the first guru; replace the row by k minus one gurus, then a laghu, then the rest of the row unchanged; append it. Return the list.

2. Give the time complexity and say why it cannot be improved.

Proportional to n times 2 to the power n, because there are 2 to the power n rows of n symbols each. No algorithm that prints the whole table can do less work than the size of the table.

3. A classmate's prastāra function returns the right number of rows, all distinct, starting all-heavy and ending all-light, but the order is wrong. Which test would catch it?

A comparison against an independently written implementation, such as the recursive construction from 8.22. Size, distinctness and endpoint tests all pass on a wrongly ordered table.

4. Why is it wrong to use this program to answer "what is row 3000 of the twelve-syllable metre?"

Because it must generate the 2999 rows before it. The retrieval rule naṣṭa computes that row directly from its number, in work proportional to the number of syllables.

Contents This chapter on its own page

munotes.in58

Chapter Nineteen

Binary Number Systems, and Exactly Where Piṅgala's Order Agrees

Syllabus topic Module 1, "Binary number systems", "Laghu–Guru representation as binary encoding"

In one line

Piṅgala's order is the order of counting in base two, with the digits written in the opposite direction from the way we write numerals.

In the wording you can write in an examination: in a positional number system of base two, a number is written as a sequence of digits each of which is 0 or 1, the value of a digit being the digit multiplied by 2 raised to the power of its position. Piṅgala's prastāra orders the patterns of n syllables so that the row number, counted from zero, equals the value of the pattern read with laghu as 1, guru as 0, and the leftmost syllable in the units position.

What a positional system is, from first principles

A number system has a base. In base ten, the digits run 0 to 9 and the places are worth 1, 10, 100, 1000 going leftward. The number 407 means four hundreds, no tens and seven units.

In base two the digits are 0 and 1 and the places are worth 1, 2, 4, 8, 16. So 101 means one four, no two and one unit, which is five.

Three facts about base two are worth having explicitly, because the rest of the chapter uses all three.

A sequence of n binary digits represents exactly the numbers from 0 to 2 to the power n minus 1. Three digits cover 0 to 7.

Adding one to a binary number flips a run of 1s at the low end to 0s and flips the next 0 to 1. 0111 plus one is 1000.

Reading the same digits in the opposite order gives a different number in general. 100 is four; reversed, 001 is one.

The claim, stated exactly

Let a pattern of n syllables be read as follows: laghu is 1, guru is 0, and the leftmost syllable is the units place, the next the twos place, and so on.

Then row r of the prastāra, counting rows from zero, is the pattern whose value is r.

That is a precise claim and it can be checked exhaustively for small n, which is what the listing does.

The check

GURU, LAGHU = "G", "L"

def prastara(n):
    row = [GURU] * n
    rows = ["".join(row)]
    while GURU in row:
        k = row.index(GURU)
        row = [GURU] * k + [LAGHU] + row[k + 1:]
        rows.append("".join(row))
    return rows

def as_number(pattern):
    """Laghu is 1, guru is 0, and the LEFTMOST syllable is the units place."""
    return sum(1 << i for i, ch in enumerate(pattern) if ch == LAGHU)

print("n      rows   row number equals binary value   first mismatch")
for n in range(1, 15):
    rows = prastara(n)
    bad = [i for i, p in enumerate(rows) if as_number(p) != i]
    print("%-6d %-6d %-32s %s"
          % (n, len(rows), "yes" if not bad else "no", "none" if not bad else rows[bad[0]]))

print()
print("the ordinary binary numerals of 3 bits, for comparison")
for v in range(8):
    print("%d  %s   reversed %s" % (v, format(v, "03b"), format(v, "03b")[::-1]))
munotes.in59

Binary Number Systems, and Exactly Where Piṅgala's Order Agrees

n      rows   row number equals binary value   first mismatch
1      2      yes                              none
2      4      yes                              none
3      8      yes                              none
4      16     yes                              none
5      32     yes                              none
6      64     yes                              none
7      128    yes                              none
8      256    yes                              none
9      512    yes                              none
10     1024   yes                              none
11     2048   yes                              none
12     4096   yes                              none
13     8192   yes                              none
14     16384  yes                              none

the ordinary binary numerals of 3 bits, for comparison
0  000   reversed 000
1  001   reversed 100
2  010   reversed 010
3  011   reversed 110
4  100   reversed 001
5  101   reversed 101
6  110   reversed 011
7  111   reversed 111

Fourteen syllable counts, 32,766 rows in total, and no mismatch anywhere.

Where the agreement comes from

The proof is the one already given in [The Prastāra Rule Read as an Algorithm], and it is worth restating in the language of this chapter because in that language it is a single observation.

The prastāra rule finds the first guru from the left, makes it laghu, and resets the laghus to its left to guru. Under the reading above, that is: find the first 0, make it 1, and make all the 1s below it 0.

That is exactly the rule for adding one to a binary number, written for a number whose low-order digit is on the left. Since the first row is all guru, which is zero, the rows are 0, 1, 2 and so on.

Where it differs: the direction

Only one thing differs, and it is a matter of writing rather than of arithmetic.

Ordinary binary numeralPiṅgala's row
Units digit is at theright-hand endleft-hand end
Value of the third symbol44
The number five, in three places101LGL
The number four, in three places100GGL
Carrying goesright to leftleft to right

Look at the last two rows of the output table above. The numeral for four is 100 and Piṅgala's pattern for four is GGL, which is 001 in digits. They are reverses of each other, and that is the whole difference.

Is the difference significant? For counting, no. The order of the rows is identical; only the direction the symbols are written in differs, which is a convention about layout. For arithmetic, it would matter, because addition algorithms are written for one direction. But Piṅgala never adds two patterns, so the question never arises in the text.

munotes.in60

Binary Number Systems, and Exactly Where Piṅgala's Order Agrees

Worked example: row 41 of the gāyatrī

The gāyatrī has six syllables and 64 rows. Find row 41.

As arithmetic. Rows are numbered from one in the tradition, so row 41 has value 40. In base two, 40 is 32 plus 8, so the sixth and fourth places hold 1: reading the leftmost syllable as the units place, positions 4 and 6 are laghu.

As a pattern. G G G L G L.

Checked by the rule. [Naṣṭa: From a Row Number Back to the Pattern] does exactly this with Piṅgala's own halving rule and gets the same answer, and the program in [Index Retrieval, and Proving Two Rules Are Inverses] verifies the two methods agree on every row of every metre up to sixteen syllables.

What this does NOT show

It does not show Piṅgala had a base. There is no statement anywhere in the text that a pattern has a numerical value, and no arithmetic is performed on a pattern.

It does not show he had positional notation. Positional notation is about writing numbers. He is writing metres, and the correspondence is something we can see and he did not need.

It does not make the prastāra a counter. A counter's purpose is to count. The prastāra's purpose is to enumerate patterns. They coincide, and coinciding is not the same as being the same thing.

The thing it does show is stronger than any of those: the order he chose, for reasons internal to prosody, is the unique order in which his two retrieval rules become simple arithmetic. That is a design decision paying off, and it is what a computer science student should admire.

Quick revision

  • Base two: digits 0 and 1, places worth 1, 2, 4, 8. Adding one flips the low run of 1s to 0 and the next 0 to 1.
  • The claim, checked for every row up to fourteen syllables: row r, counted from zero, is the pattern whose value is r with laghu as 1 and the leftmost syllable in the units place.
  • The reason: the prastāra rule is the add-one rule written with the low-order digit on the left.
  • The only difference from numeral notation is the direction, which is layout and not arithmetic, and it never bites because the text never adds two patterns.
  • Not shown: that Piṅgala had a base, positional notation, or a counter.

Test yourself

1. State the correspondence between a prastāra row and a binary number exactly.

With laghu read as 1, guru as 0, and the leftmost syllable as the units place, the pattern in row r of the prastāra, counting rows from zero, has value r.

2. Explain in one sentence why the prastāra rule produces that order.

munotes.in61

Binary Number Systems, and Exactly Where Piṅgala's Order Agrees

Because finding the first guru, making it laghu and resetting the laghus to its left is precisely the rule for adding one to a binary number written with its units digit on the left.

3. What is the single difference between a prastāra row and a binary numeral, and does it affect the order?

The units place is at the left-hand end instead of the right, so the symbols are written in the reverse direction. It does not affect the order of the rows at all.

4. Find row 17 of the four-syllable prastāra by arithmetic, and say what goes wrong.

Nothing can be found: the four-syllable prastāra has 16 rows, numbered 1 to 16, so there is no row 17. A four-syllable pattern carries four bits and covers the values 0 to 15 only.

Contents This chapter on its own page

munotes.in62

Chapter Twenty

Saṅkhyā: Counting the Rows Without Writing Them

Syllabus topic Module 1, "Prastāra (systematic enumeration of patterns)", "Algorithmic generation"

In one line

Saṅkhyā answers "how many patterns are there?" without writing any of them, by squaring and doubling instead of by multiplying out.

In the wording you can write in an examination: saṅkhyā is the pratyaya that computes the number of metrical patterns of a given syllable count. Piṅgala's method halves the syllable count repeatedly, recording whether each halving was exact or required one to be subtracted first, and then builds the answer from one by squaring at each exact halving and doubling at each subtraction, so that 2 to the power n is obtained in about log n operations instead of n.

The problem

The gāyatrī has six syllables. How many patterns are there? You could write the table and count to 64, which takes 64 rows of work. Piṅgala does it in four operations, and Halāyudha works it out in the commentary.

The sūtras, named rather than numbered

This group is where this book's three witnesses divide the chapter differently: the same sūtra "rūpe śūnyam" is numbered 26 in one and 28 in another. So the sūtras are named by their opening words and their content, and the disagreement is stated rather than hidden.

Opening wordsWhat it does
dvir ardhamhalve the number, and mark the halving
rūpe śūnyamwhen the number is odd, take one off, and mark that instead
dviḥ śūnyeat a mark of the second kind, double
tāvad ardhe tadguṇitamat a mark of the first kind, multiply the number by itself
dvir dvyūnaṃ tadantānāmfor the running total of all metres up to this one, double and subtract two

Halāyudha's own gloss on the third of these is explicit about where the answer starts: one is obtained, and placing it at the zero, double it, then it becomes two.

The method, in two passes

Pass one, going down. Write the syllable count. If it is even, halve it and write a mark meaning "halved". If it is odd, subtract one and write a mark meaning "reduced". Repeat until you reach zero.

Pass two, coming back up. Start with the number one. Read the marks in reverse order, that is, from the last one written to the first. At a "halved" mark, square the number you have. At a "reduced" mark, double it.

Halāyudha's own worked example

He works the gāyatrī, six syllables, and this book follows his steps exactly.

Going down. Six is even, so halve it: three. Mark: halved. Three is odd, so take one off: two. Mark: reduced. Two is even, so halve it: one. Mark: halved. One is odd, so take one off: zero. Mark: reduced.

Coming back up, in his words as the commentary gives them. One is obtained; placing it at the zero, double it, then it becomes two. At the half place, multiply the number by itself: two multiplied by two becomes four. Then, doubling at the zero, it becomes eight. Placing those at the half place, multiply by itself: eight multiplied by eight becomes sixty-four, the gāyatrī forms.

munotes.in63

Saṅkhyā: Counting the Rows Without Writing Them

So: 1, double to 2, square to 4, double to 8, square to 64. Four operations, and the answer is 64.

The figures his commentary then states

Having done the gāyatrī, he runs through the rest, and this is the part worth quoting in an answer, because the numbers are in the text.

MetreSyllablesThe commentary's figure2 to the power n
gāyatrī6sixty-four64
uṣṇih7a hundred increased by twenty-eight128
anuṣṭubh8two hundreds increased by fifty-six256
bṛhatī9five hundreds with twelve over512
paṅkti10a thousand increased by twenty-four1024
triṣṭubh11two thousands increased by forty-eight2048
jagatī12four thousands increased by ninety-six4096

Seven figures, seven exact powers of two, in a commentary on prosody.

The method, run

def marks(n):
    """Pingala's halving: S where the number was halved, D where one came off."""
    out = []
    x = n
    while x > 0:
        if x % 2 == 0:
            out.append("S")
            x //= 2
        else:
            out.append("D")
            x -= 1
    return "".join(reversed(out))

def sankhya(n):
    """Start at one; square at an S, double at a D."""
    v = 1
    for m in marks(n):
        v = v * v if m == "S" else v * 2
    return v

NAMED = [("gayatri", 6), ("usnih", 7), ("anustubh", 8), ("brhati", 9),
         ("pankti", 10), ("tristubh", 11), ("jagati", 12)]

print("%-10s %-4s %-8s %-6s %-6s %s" % ("metre", "n", "marks", "ops", "2**n", "sankhya"))
for name, n in NAMED:
    print("%-10s %-4d %-8s %-6d %-6d %d" % (name, n, marks(n), len(marks(n)), 2 ** n, sankhya(n)))

print()
bad = [n for n in range(1, 200) if sankhya(n) != 2 ** n]
print("values of n from 1 to 199 where the rule disagrees with 2**n:", bad)
print("operations for n = 100:", len(marks(100)), "against 100 multiplications")
metre      n    marks    ops    2**n   sankhya
gayatri    6    DSDS     4      64     64
usnih      7    DSDSD    5      128    128
anustubh   8    DSSS     4      256    256
brhati     9    DSSSD    5      512    512
pankti     10   DSSDS    5      1024   1024
tristubh   11   DSSDSD   6      2048   2048
jagati     12   DSDSS    5      4096   4096

values of n from 1 to 199 where the rule disagrees with 2**n: []
operations for n = 100: 9 against 100 multiplications

The mark string is printed in the order the operations are applied, so DSDS for the gāyatrī reads: double, square, double, square, which is Halāyudha's own sequence.

munotes.in64

Saṅkhyā: Counting the Rows Without Writing Them

Why it works

Squaring doubles the exponent and doubling adds one to it. So if you can build n out of doublings and additions of one, you can build 2 to the power n out of squarings and doublings.

And building n that way is exactly the halving procedure read backwards: n is either twice something, or one more than something. That is the whole justification, and it is two sentences.

The running total

The last sūtra of the group answers a different question: how many metres are there altogether, from one syllable up to n? Halāyudha's gloss is to double the number of forms and reduce by two.

Check it. For the gāyatrī, double 64 is 128, less two is 126, and 2 plus 4 plus 8 plus 16 plus 32 plus 64 is 126. The formula is 2 to the power n plus one, minus two, and it is right.

What it does NOT mean

It does not mean Piṅgala had logarithms. Nothing counts operations, and nothing in the text says the method is faster than multiplying out. The saving is real and it is not remarked upon.

It does not mean the method is stated as general exponentiation. It is stated for powers of two, because that is what the subject needs. Squaring and doubling generalises immediately to any base, and the text does not generalise it.

It does not mean 2 to the power n was understood as a function. The commentary computes seven values. It does not write a formula.

Limits

The honest limit is that the method is presented as a recipe, not as an economy. Nobody in the text says "this is quicker". A modern reader supplies the observation that it takes about log n operations, and that observation, not the recipe, is what makes it a landmark. Saying so is better than pretending the text claims what it does not.

Quick revision

  • Saṅkhyā: the number of patterns, without listing them.
  • Going down: halve if even, subtract one if odd, marking which. Coming back up: start at one, square at a halving, double at a subtraction.
  • Halāyudha's gāyatrī: 1, double 2, square 4, double 8, square 64. Four operations.
  • His commentary states 64, 128, 256, 512, 1024, 2048, 4096 for six to twelve syllables, all exact powers of two.
  • Why it works: squaring doubles the exponent, doubling adds one, and every number is twice something or one more than something.
  • The running total sūtra: double the count and subtract two, which gives 126 for the gāyatrī.
  • What is absent: any statement that the method is cheaper.

Test yourself

1. Compute 2 to the power 10 by Piṅgala's method, showing the marks and the operations.

Ten is even, halve to five, mark halved; five is odd, subtract one to four, mark reduced; four halves to two, mark halved; two halves to one, mark halved; one reduces to zero, mark reduced. Read back: start 1, double to 2, square to 4, square to 16, double to 32, square to 1024. Five operations.

munotes.in65

Saṅkhyā: Counting the Rows Without Writing Them

2. Why does squaring correspond to halving, and doubling to subtracting one?

Because squaring a power of two doubles its exponent and doubling adds one to it, so the operations on the answer mirror, in reverse, the halvings and subtractions applied to the exponent.

3. Which figures from Halāyudha's commentary support the claim that his method is about powers of two?

Sixty-four for six syllables, and then 128, 256, 512, 1024, 2048 and 4096 for seven to twelve, each stated in Sanskrit numerals and each exactly a power of two.

4. What does the text NOT say about this method, and why does that matter to an honest answer?

It nowhere says that the method is faster than repeated multiplication, and it never counts operations. The economy is real but it is our observation, not the text's claim, and an answer that attributes it to the text overstates the evidence.

Contents This chapter on its own page

munotes.in66

Chapter Twenty-One

Binary Exponentiation: The Same Algorithm in a Modern Textbook

Syllabus topic Module 1, "Algorithmic generation", "Complexity & Limitations"

In one line

Binary exponentiation computes a power by squaring and multiplying instead of by multiplying over and over, and it turns n multiplications into about log n.

In the wording you can write in an examination: binary exponentiation, also called exponentiation by squaring or the square-and-multiply method, computes a to the power n by writing n in binary and processing its bits, squaring the running result at each bit and multiplying by a where the bit is one. It performs at most 2 log n multiplications instead of n minus 1.

The problem, and the naive cost

To compute a to the power n, the obvious method multiplies a by itself n minus 1 times. For n equal to 1000 that is 999 multiplications.

Squaring changes the arithmetic of the situation. If you know a to the power 500, then one squaring gives a to the power 1000. If you know a to the power 500, one more multiplication by a gives a to the power 501. So each step either doubles the exponent or adds one to it, and you can reach any exponent by a sequence of doublings and single additions.

That is exactly Piṅgala's saṅkhyā rule with the base left free instead of fixed at two.

The method, from the top down

Write n in binary. For n equal to 13 that is 1101.

Read the bits from the most significant end. Start with the result equal to a.

For each remaining bit: square the result; and if the bit is one, multiply by a as well.

Trace it for a to the power 13. Bits after the first: 1, 0, 1.

StepBitOperationExponent reached
start1result is a1
11square, then multiply by a3
20square6
31square, then multiply by a13

Five multiplications instead of twelve, and the pattern of the exponents, 1, 3, 6, 13, is the binary number being built one bit at a time.

The method, from the bottom up

There is a second formulation which is the one usually coded, because it needs no bit counting in advance.

FUNCTION power(a, n)

result <- 1

base <- a

while n > 0 do

if n is odd then result <- result * base

base <- base * base

n <- n / 2, discarding the remainder

end while

return result

It is the same algorithm read from the low-order end. The n / 2 is Piṅgala's halving; the "if n is odd" is his rūpe śūnyam.

The cost

Let b be the number of bits in n, which is about log n to base two, rounded up.

The top-down form does b minus 1 squarings and at most b minus 1 multiplications, so at most 2 log n multiplications in all.

munotes.in67

Binary Exponentiation: The Same Algorithm in a Modern Textbook

Worked. For n equal to 1000, b is 10. So at most 18 multiplications against 999. For n equal to a million, b is 20, so at most 38 against nearly a million.

And the lower bound. You cannot do better than about log n, because each multiplication at best doubles the largest exponent you hold, so after k multiplications the highest exponent reachable is 2 to the power k.

Piṅgala's version against the general one

Piṅgala's saṅkhyāBinary exponentiation
The basefixed at twoany a
What is computedthe number of metrical patternsa to the power n
The descenthalve, or take one off if oddthe same
The ascentsquare at a halving, double at a subtractionsquare at a halving, multiply by a at a subtraction
Costabout log n operationsabout log n operations
Stated as an economynoyes, that is the whole point of it

The only mathematical difference is the fixed base. The doubling in his ascent is multiplication by the base, and his base happens to be two.

Why it still matters: modular exponentiation

This is the sentence to remember, because it links Module I to Module II.

Public-key cryptography rests on computing a to the power e modulo m where e and m are hundreds of digits long. With repeated multiplication that is impossible: there is not enough time in the universe. With square and multiply it is a few thousand multiplications, each of them on numbers of a few hundred digits, which is a fraction of a second.

Two details make it practical, and both are worth naming.

Reduce modulo m at every step. Otherwise the intermediate numbers grow to astronomical size. Squaring a number of d digits gives 2d digits, so after twenty squarings you would have a million digits.

The bits of the exponent are secret. Because the algorithm does an extra multiplication exactly when a bit is one, an attacker who can measure the time or the power consumption can read the exponent off. Real implementations therefore do the same work whichever the bit is. That is a whole field, and [Symmetric Encryption] and [Cipher Algorithms, Classical to Modern] return to the general point that an algorithm's timing can leak its secrets.

What it does NOT mean

It does not mean Piṅgala invented the algorithm for a general base. His base is two throughout and no generalisation appears.

It does not mean the method is always best. For a small n, repeated multiplication is simpler and the difference is nothing. The method earns its keep when n is large.

munotes.in68

Binary Exponentiation: The Same Algorithm in a Modern Textbook

It does not mean the cost is exactly log n. It is between log n and 2 log n depending on how many ones the exponent has, and an answer that says "about log n" is right while one that says "log n" is imprecise.

Quick revision

  • Binary exponentiation: square at each bit, multiply by the base where the bit is one. At most 2 log n multiplications.
  • Top down, from the most significant bit: result starts at a, then square and conditionally multiply.
  • Bottom up: while n is positive, multiply into the result if n is odd, square the base, halve n.
  • Piṅgala's saṅkhyā is the same algorithm with the base fixed at two, and his doubling is multiplication by that base.
  • The lower bound is about log n, because each multiplication at best doubles the largest exponent held.
  • It is what makes public-key cryptography possible; reduce modulo m at every step, and beware that the timing leaks the exponent.

Test yourself

1. Compute a to the power 13 by squaring, listing the exponents reached.

Write 13 as 1101. Start at a, exponent 1. Square and multiply by a: 3. Square: 6. Square and multiply by a: 13. Five multiplications.

2. Give the cost of the method and the reason it cannot be improved much.

At most 2 log n multiplications, since there are about log n bits and each contributes a squaring and possibly one multiplication. It cannot be improved beyond about log n because each multiplication at best doubles the largest exponent held, so k multiplications reach at most 2 to the power k.

3. What is the one mathematical difference between Piṅgala's saṅkhyā and binary exponentiation?

Piṅgala's base is fixed at two, so his "multiply by the base" step appears as doubling. Otherwise the descent and the ascent are identical.

4. Why must a modular exponentiation reduce at every step, and what does the algorithm leak if written naively?

Without reducing, the intermediate values double in length at each squaring and become unmanageable. Written naively it performs an extra multiplication exactly when an exponent bit is one, so its running time or power draw reveals the exponent.

Contents This chapter on its own page

munotes.in69

Chapter Twenty-Two

Uddiṣṭa: From a Pattern to Its Row Number

Syllabus topic Module 1, "Naṣṭa and Uddiṣṭa (index retrieval mechanisms)"

In one line

Uddiṣṭa takes a pattern somebody points at and tells you which row of the table it is, without looking at the table.

In the wording you can write in an examination: uddiṣṭa is the pratyaya that computes the serial number of a given metrical pattern within the prastāra. Piṅgala's rule processes the syllables from the last to the first, beginning with one, doubling at each syllable and subtracting one wherever the syllable is a guru; the result is the row number of the pattern, counted from one.

The problem

Somebody writes down L G L G G G and asks: which row of the six-syllable table is that? You could generate the table and search it, which is up to 64 rows of work for the gāyatrī and 4096 for the jagatī. Piṅgala does it in six steps, one per syllable.

The provisions

8.26, pratilomaguṇaṃ dvirlādyam. Reverse-order multiplication, doubling, beginning from the la. Halāyudha's gloss: of whichever pattern one wishes to know the number, taking it from the end, going in reverse order, double repeatedly; and since a starting value of nothing cannot be doubled, the number one is obtained to begin with.

8.27, tato'pyekaṃ jahyāt. From that also subtract one. His gloss is exact about when: in performing the aforesaid operation, if that number falls on the place of a ga, then having doubled it, drop one from the aggregate.

Both numbers agree across this book's witnesses.

The rule as steps

  1. Start with the number one.
  2. Take the syllables of the pattern from the LAST one to the first.
  3. At each syllable, double the number you have.
  4. If that syllable is a guru, subtract one from the result of the doubling.
  5. When the syllables run out, the number you have is the row.

Step three and step four are one operation in the text and it is easier to remember that way: double, and take one off at a heavy syllable.

Worked by hand: L G L G G G

The pattern is six syllables. Taking them from the last to the first, they are G, G, G, L, G, L.

Syllable takenDoubleGuru, so subtract oneRunning number
start1
6th, G2yes1
5th, G2yes1
4th, G2yes1
3rd, L2no2
2nd, G4yes3
1st, L6no6

Row six, which is the row Halāyudha uses as his own example in the naṣṭa rule.

Notice what the long run of gurus at the end does: it holds the number at one. That is correct, because trailing heavy syllables contribute nothing to the row number, which in the binary reading is the same as leading zeros contributing nothing to a number's value.

munotes.in70

Uddiṣṭa: From a Pattern to Its Row Number

Worked by hand: G G G L G L

Syllable takenDoubleGuru, so subtract oneRunning number
start1
6th, L2no2
5th, G4yes3
4th, L6no6
3rd, G12yes11
2nd, G22yes21
1st, G42yes41

Row forty-one, which is the row [Binary Number Systems, and Exactly Where Piṅgala's Order Agrees] reached by arithmetic. The two methods agree.

Why the rule is right

Write the row number, counted from one, as one plus the pattern's binary value with laghu as 1 and the leftmost syllable in the units place. Then the rule is Horner's method for evaluating that expression from the high-order end, which for this direction of writing means from the right.

Let y stand for the running number after some syllables have been taken. The step "double, and subtract one at a guru" is the same as "double, and add one at a laghu, then subtract one". Which is to say: it evaluates twice the previous value, plus the current bit, minus one. Unrolled from a start of one, that is exactly one plus the binary value.

The short version, which is the one to write in an examination: the rule is Horner's evaluation of the pattern as a binary number, shifted by one because rows are counted from one and numbers from zero.

The rule, run

def uddista(pattern):
    """8.26 pratilomagunam dvirladyam, with 8.27 tato'pyekam jahyat.

    Take the syllables from the LAST to the first. Start at one. At each
    syllable double; and if that syllable is a guru, subtract one.
    """
    y = 1
    for ch in reversed(pattern):
        y *= 2
        if ch == "G":
            y -= 1
    return y

for p in ("GGGGGG", "LGLGGG", "GGGLGL", "LLLLLL", "GGG", "LLL"):
    print("%-8s row %d" % (p, uddista(p)))
GGGGGG   row 1
LGLGGG   row 6
GGGLGL   row 41
LLLLLL   row 64
GGG      row 1
LLL      row 8

Four lines of code for a rule stated in two sūtras. The whole content is y = 2*y with a conditional y -= 1.

The cost

One pass over the pattern, one doubling and at most one subtraction per syllable. So the work is proportional to n, the number of syllables, and does not depend on how large the row number is.

Compare the alternative: generating the table and searching it costs n times 2 to the power n. For the jagatī that is 49,152 symbol operations against 12. This is the same improvement in kind that an index gives over a scan, and it is worth saying in exactly those words, because that is what MU's phrase "index retrieval mechanisms" is pointing at.

munotes.in71

Uddiṣṭa: From a Pattern to Its Row Number

What it does NOT mean

It does not mean the pattern is validated. Give the rule a pattern of the wrong length and it will cheerfully return a number. Nothing in the text checks the input, and a program should.

It does not mean rows are counted from zero. They are counted from one throughout the tradition, which is why the rule starts at one rather than at zero. Mixing the conventions is the commonest arithmetic mistake in this topic.

It does not work for another order. Later prosodists use the reverse order, and uddiṣṭa for that order is a different rule. Applying this one to a table built the other way gives a wrong answer silently.

Quick revision

  • Uddiṣṭa: from a pattern to its row number.
  • 8.26 pratilomaguṇaṃ dvirlādyam, reverse order, doubling; 8.27 tato'pyekaṃ jahyāt, and subtract one, which Halāyudha applies at a guru.
  • The rule: start at one, take syllables from last to first, double each time, subtract one at a heavy syllable.
  • Worked: L G L G G G is row 6; G G G L G L is row 41.
  • It is Horner's evaluation of the pattern as a binary number, offset by one because rows count from one.
  • Cost proportional to n, against n times 2 to the power n for generate and search.

Test yourself

1. Find the row number of G L L G by the rule, showing the running value.

Syllables from the last: G, L, L, G. Start 1. G: double to 2, subtract one, 1. L: double to 2. L: double to 4. G: double to 8, subtract one, 7. Row seven.

2. Why do trailing heavy syllables leave the running number unchanged?

Because doubling one and subtracting one gives one again. In the binary reading they are leading zeros, which contribute nothing to a number's value.

3. State the modern description of the rule in one sentence.

It is Horner's method for evaluating the pattern as a binary number with laghu as one and the leftmost syllable in the units place, offset by one because rows are counted from one.

4. Give the cost of uddiṣṭa and of the obvious alternative, for a twelve-syllable metre.

Uddiṣṭa is twelve doublings, one per syllable. Generating the table and searching it is 4096 rows of twelve symbols, about 49,000 symbol operations.

Contents This chapter on its own page

munotes.in72

Chapter Twenty-Three

Naṣṭa: From a Row Number Back to the Pattern

Syllabus topic Module 1, "Naṣṭa and Uddiṣṭa (index retrieval mechanisms)"

In one line

Naṣṭa takes a row number and gives you the pattern in that row, without writing any of the rows before it.

In the wording you can write in an examination: naṣṭa, literally the lost, is the pratyaya that recovers a metrical pattern from its serial number in the prastāra. Piṅgala's rule halves the number: where the halving is exact a laghu is written, and where the number is odd one is added before halving and a guru is written. The syllables are written from the left, one per halving, until the metre's syllable count is reached.

The problem, and why it is called "the lost"

A student remembers that a verse is the forty-first form of the six-syllable metre but has forgotten the pattern. The pattern is lost and the number survives. Naṣṭa restores it.

The alternative is to build the table and count down to row forty-one. That is 41 rows of work, and for a jagatī it can be 4096. Naṣṭa is six steps, one per syllable.

The provisions

8.24, lardhe. On halving, an L. Halāyudha's gloss, working his own example: if one wishes to know what the sixth of the gāyatrī same-metre forms is, then halve that number six; when it is halved, a single laghu is obtained, to be placed on the ground, that is, in the first position.

8.25, saike g. With one added, a G. His gloss: in the case of an odd number, add one, then halve; there a single ga is obtained, to be placed after the syllable previously obtained. And then: this rule of saike ga is to be applied again and again, until the syllables are complete, six in number.

Both numbers agree across this book's witnesses.

The rule as steps

  1. Write down the row number.
  2. If it is even, halve it and write L.
  3. If it is odd, add one, halve the result, and write G.
  4. Write each syllable to the right of the last, so the first syllable obtained is the leftmost.
  5. Repeat until you have as many syllables as the metre has. Stop then, whatever number you are holding.

Step five matters. The number will usually reach one and stay there, and those steps still produce syllables. You stop by counting syllables, not by reaching zero.

Worked by hand: row 6 of the gāyatrī

This is Halāyudha's own example.

StepNumberEven or oddOperationSyllable
16evenhalve to 3L
23oddadd one, halve to 2G
32evenhalve to 1L
41oddadd one, halve to 1G
51oddadd one, halve to 1G
61oddadd one, halve to 1G
munotes.in73

Naṣṭa: From a Row Number Back to the Pattern

Reading the syllable column downward gives L G L G G G, and that is row six.

Check it against the other rule: [Uddiṣṭa: From a Pattern to Its Row Number] takes L G L G G G and returns six. The two rules agree, and they must, because they are inverses.

Worked by hand: row 41 of the gāyatrī

StepNumberEven or oddOperationSyllable
141oddadd one, halve to 21G
221oddadd one, halve to 11G
311oddadd one, halve to 6G
46evenhalve to 3L
53oddadd one, halve to 2G
62evenhalve to 1L

G G G L G L, which is the pattern [Binary Number Systems, and Exactly Where Piṅgala's Order Agrees] reached by arithmetic and which uddiṣṭa returns forty-one for.

Why the rule is right

In the binary reading, the row number minus one is the pattern's value with laghu as one and the leftmost syllable in the units place. So the first syllable is laghu exactly when the units bit of the row number minus one is one, that is, exactly when the row number minus one is odd, that is, exactly when the row number is even.

That is step two. And halving the even row number r gives r over two, whose parity decides the next syllable in the same way.

For the odd case, r minus one is even, so the bit is zero and the syllable is guru; and the next number needed is r minus one over two plus one, which is r plus one over two. That is step three, with the plus one before the halving doing exactly that arithmetic.

So: "add one before halving at an odd number" is the correction that keeps the count on a one-based scale. That single detail is what makes the rule work on row numbers rather than on values, and it is the detail a student who reinvents the rule always gets wrong.

The rule, run, with both directions checked

def nasta(n, r):
    """8.24 lardhe and 8.25 saike g.

    Halve the row number; if it was even write a laghu, and if it was odd add
    one before halving and write a guru. Do this n times, writing from the left.
    """
    out = []
    x = r
    for _ in range(n):
        if x % 2 == 0:
            out.append("L")
            x //= 2
        else:
            out.append("G")
            x = (x + 1) // 2
    return "".join(out)

def uddista(pattern):
    y = 1
    for ch in reversed(pattern):
        y *= 2
        if ch == "G":
            y -= 1
    return y

print("the six syllable metre, Halayudha's own example is row 6")
for r in (1, 6, 41, 64):
    p = nasta(6, r)
    print("  row %-3d nasta gives %-8s uddista gives it back as %d" % (r, p, uddista(p)))
munotes.in74

Naṣṭa: From a Row Number Back to the Pattern

the six syllable metre, Halayudha's own example is row 6
  row 1   nasta gives GGGGGG   uddista gives it back as 1
  row 6   nasta gives LGLGGG   uddista gives it back as 6
  row 41  nasta gives GGGLGL   uddista gives it back as 41
  row 64  nasta gives LLLLLL   uddista gives it back as 64

The cost

One halving and at most one addition per syllable, so the work is proportional to n and independent of how large the row number is. Finding row 4000 of a twelve-syllable metre costs the same twelve steps as finding row 1.

What it does NOT mean

It does not check the row exists. Ask for row 100 of a three-syllable metre and the rule returns a pattern. There are only eight rows. Nothing in the text validates the input and a program must.

It does not stop when the number reaches one. It stops when the syllable count is reached. A reader who stops early produces a short pattern, and this is the commonest mistake in working the rule by hand.

It does not depend on the table existing. That is the point of it. The pattern is computed, not looked up, and the table never has to be written at all.

Quick revision

  • Naṣṭa: from a row number to the pattern in that row.
  • 8.24 lardhe: halve, and write L. 8.25 saike g: if odd, add one, halve, and write G.
  • Write syllables left to right, one per step, and stop when the syllable count is reached, not when the number reaches one.
  • Halāyudha's own example: row 6 of the gāyatrī is L G L G G G.
  • Row 41 of the gāyatrī is G G G L G L.
  • Cost proportional to the number of syllables, independent of the row number.
  • The "add one before halving" is the correction for rows being counted from one.

Test yourself

1. Find row 11 of the four-syllable metre by the rule.

11 is odd: add one, halve to 6, write G. 6 is even: halve to 3, write L. 3 is odd: add one, halve to 2, write G. 2 is even: halve to 1, write L. The pattern is G L G L.

2. Why is one added before halving when the number is odd?

Because rows are counted from one and values from zero. For an odd row number r, the value r minus one is even, so the syllable is guru, and the next number required is r minus one over two plus one, which equals r plus one over two.

munotes.in75

Naṣṭa: From a Row Number Back to the Pattern

3. When does the rule stop, and what happens if a student stops too early?

It stops when as many syllables have been written as the metre has. Stopping when the number reaches one produces a pattern that is too short, which is the commonest error in working it by hand.

4. How much work does naṣṭa do for row 4000 of a twelve-syllable metre, and how much does generating the table do?

Naṣṭa does twelve halvings. Generating the table and counting to row 4000 does about 4000 rows of twelve symbols, roughly 48,000 symbol operations.

Contents This chapter on its own page

munotes.in76

Chapter Twenty-Four

Index Retrieval, and Proving Two Rules Are Inverses

Syllabus topic Module 1, "Naṣṭa and Uddiṣṭa (index retrieval mechanisms)", "minimum 10 test cases"

In one line

Naṣṭa and uddiṣṭa are two directions of the same index: number to pattern, and pattern to number, neither of which needs the table.

In the wording you can write in an examination: an index retrieval mechanism provides access to an element of an ordered collection by its position, and to the position of a given element, in time that does not depend on the size of the collection. Naṣṭa and uddiṣṭa together provide both directions for the prastāra, each in time proportional to the number of syllables, and they are exact inverses of one another.

What "index retrieval" means, and why it is not a small thing

Consider three ways to answer "what is the forty-first row of the gāyatrī?".

Scan. Generate rows from the first and count. Cost: proportional to the row number. For row 4000, four thousand rows of work.

Store and look up. Write the whole table once and index into it. Cost: one step per query, but the table has to exist, and it costs n times 2 to the power n to build and to hold.

Compute. Derive the row from its number with no table at all. Cost: proportional to n, and nothing is stored.

The third is what naṣṭa does, and it is the one that scales. For a thirty-syllable metre the table would hold about a billion rows and the computation still takes thirty steps.

The modern name for what the third option exploits is a direct or computed address. An array index works this way: the address of element i is the base plus i times the element size, computed rather than searched. A hash table works this way too, with a function from the key to the address. Piṅgala's rules are a computed address into a structure that was never built.

The two directions, side by side

NaṣṭaUddiṣṭa
Givena row numbera pattern
Returnsthe patternits row number
Sūtras8.24 lardhe, 8.25 saike g8.26 pratilomaguṇaṃ dvirlādyam, 8.27 tato'pyekaṃ jahyāt
Direction of writingleft to rightright to left
Arithmetic per syllablehalve, adding one first if odddouble, subtracting one at a guru
Costproportional to nproportional to n
Needs the tablenono

The symmetry is exact, and it is the symmetry of an encoder and a decoder: one takes a number to a representation and the other takes the representation back to the number.

The proof that they are inverses

Two functions are inverses if applying one and then the other returns you to where you started, both ways round. For a finite set that can be checked exhaustively, and here it can.

def nasta(n, r):
    out = []
    x = r
    for _ in range(n):
        if x % 2 == 0:
            out.append("L")
            x //= 2
        else:
            out.append("G")
            x = (x + 1) // 2
    return "".join(out)

def uddista(pattern):
    y = 1
    for ch in reversed(pattern):
        y *= 2
        if ch == "G":
            y -= 1
    return y

print("%-6s %-8s %-10s %s" % ("n", "rows", "inverses", "first failure"))
for n in range(1, 17):
    bad = [r for r in range(1, 2 ** n + 1) if uddista(nasta(n, r)) != r]
    print("%-6d %-8d %-10s %s" % (n, 2 ** n, "yes" if not bad else "no",
                                  "none" if not bad else bad[0]))
munotes.in77

Index Retrieval, and Proving Two Rules Are Inverses

n      rows     inverses   first failure
1      2        yes        none
2      4        yes        none
3      8        yes        none
4      16       yes        none
5      32       yes        none
6      64       yes        none
7      128      yes        none
8      256      yes        none
9      512      yes        none
10     1024     yes        none
11     2048     yes        none
12     4096     yes        none
13     8192     yes        none
14     16384    yes        none
15     32768    yes        none
16     65536    yes        none

Sixteen syllable counts, 131,070 rows in total, every one checked in both directions.

Why exhaustive checking is worth something here, and where it stops

What it establishes. For every metre up to sixteen syllables, the two rules are exactly inverse. No case within that range can be wrong.

What it does not establish. That they are inverse for every n. A test over a finite range is a test, not a proof.

The proof, for the record, is the argument given in the two previous chapters: both rules compute the same correspondence between row numbers and binary values, one in each direction, and a bijection's inverse is unique. State the argument and offer the check as evidence, never the check alone. That is the difference between a program that passes and a claim that is justified, and it is the distinction MU's "analysis of correctness" is asking for.

Worked example: the other direction, and the round trip

Take the jagatī, twelve syllables, and row 2731.

Naṣṭa. 2731 is odd, so add one and halve to 1366, writing G. 1366 halves to 683, writing L. 683 is odd, so 342, writing G. 342 halves to 171, L. 171 odd, 86, G. 86 halves to 43, L. 43 odd, 22, G. 22 halves to 11, L. 11 odd, 6, G. 6 halves to 3, L. 3 odd, 2, G. 2 halves to 1, L. The pattern is G L G L G L G L G L G L.

Uddiṣṭa on that pattern. From the last syllable, which is L: 1 doubles to 2. G: 4 less one, 3. L: 6. G: 12 less one, 11. L: 22. G: 44 less one, 43. L: 86. G: 172 less one, 171. L: 342. G: 684 less one, 683. L: 1366. G: 2732 less one, 2731.

munotes.in78

Index Retrieval, and Proving Two Rules Are Inverses

The round trip closes, and the alternating pattern is worth noticing: the row numbers of alternating patterns are the numbers whose binary form alternates, which for twelve places is 2730 as a value and 2731 as a row.

What this does NOT give you

It does not give you search. Neither rule answers "which rows have exactly four laghus?". That is lagakriyā, and it is a different pratyaya with a different answer, in [Meru-Prastāra: Halāyudha's Staircase].

It does not give you a sorted order other than its own. The index is the index of Piṅgala's order. A question about a different order needs different rules.

It does not validate anything. Neither rule checks that the row exists or that the pattern has the right length. A program that exposes them to a user must.

Quick revision

  • Index retrieval: reaching an element by its position, and a position by its element, in time independent of the collection's size.
  • Naṣṭa is number to pattern; uddiṣṭa is pattern to number. Each costs one step per syllable and needs no table.
  • The modern analogue is a computed address, as in an array index or a hash function.
  • The pair is exactly inverse; checked over 131,070 rows up to sixteen syllables, and proved by the binary correspondence.
  • An exhaustive check over a range is evidence, not a proof. Give the argument and offer the check.
  • Neither rule searches, validates, or works for a different ordering.

Test yourself

1. What does it mean to call naṣṭa and uddiṣṭa index retrieval mechanisms?

They give access to a row by its position and to a position by its row, in time proportional to the syllable count and independent of the number of rows, without the collection being stored.

2. Compare the three ways of answering "what is row 4000 of the jagatī?" by cost.

Scanning costs about 4000 rows of work. Storing the table costs 4096 rows of space to build and hold, then one lookup. Computing by naṣṭa costs twelve steps and no space.

3. Why is the exhaustive check in this chapter not a proof, and what is the proof?

Because it covers only n up to sixteen. The proof is that both rules compute the same bijection between row numbers and binary values, one in each direction, and the inverse of a bijection is unique.

4. Which question do these two rules NOT answer, and which pratyaya does?

They do not answer how many patterns have a given number of light syllables. That is lagakriyā, answered by the Meru-prastāra.

Contents This chapter on its own page

munotes.in79

Chapter Twenty-Five

Recursive Enumeration

Syllabus topic Module 1, "Recursive enumeration", "relate them to binary encoding, recursion, and tree structures"

In one line

A recursive definition solves a problem by solving a smaller version of the same problem, and Halāyudha states the prastāra that way.

In the wording you can write in an examination: a recursive definition defines a structure or procedure in terms of smaller instances of itself, together with one or more base cases that are defined outright. Recursive enumeration is the generation of all members of a set by combining the members of a smaller set of the same kind, and Piṅgala's sūtra 8.22, as Halāyudha explains it, generates the prastāra for n syllables from the prastāra for n minus one.

The two parts, and the third thing people forget

The base case. The smallest instance, answered outright with no further reference to the definition. For the prastāra that is one syllable: the table is G, then L. Sūtra 8.20 dvikau glau gives exactly this.

The recursive case. How to get the answer for size n from the answer for something smaller. Sūtra 8.22 gives it: write the table for n minus one twice, and fill the new syllable's column with heavy against the first copy and light against the second.

And the thing people forget: the recursive case must make progress toward the base case. Each call is on n minus one, which is strictly smaller and bounded below by one, so the chain of calls is finite. A recursion without that property does not terminate, and it is the commonest error in a first program.

Halāyudha's own statement

He is explicit that the rule is to be repeated, and explicit about where it stops.

On the three-syllable table he says: thus the three-syllable prastāra is established. Then, immediately: having set down the three-syllable prastāra twice, ga and la are to be given separately and mixed, and that is the four-syllable prastāra. And then: this sūtra is to be applied again and again, until the desired prastāra is reached.

Three things are in those sentences and they are the three parts above: a smaller solved case, a combination rule, and a stopping condition stated as "until the desired prastāra".

Worked by hand: four syllables from three

The three-syllable table, from [Prastāra: The Table of Every Pattern]: GGG, LGG, GLG, LLG, GGL, LGL, GLL, LLL.

Write it twice. Sixteen rows, the eight repeated.

Fill the fourth column. Heavy for the first eight, light for the second eight.

Rows 1 to 8Rows 9 to 16
GGGGGGGL
LGGGLGGL
GLGGGLGL
LLGGLLGL
GGLGGGLL
LGLGLGLL
GLLGGLLL
LLLGLLLL

Read the left column down and then the right column down: that is the sixteen-row four-syllable prastāra, in the same order the iterative rule produces.

munotes.in80

Recursive Enumeration

The new column is at the right-hand end. Put it at the left and you get sixteen distinct rows in the wrong order, and every test that counts rows will pass. [Prastāra in Code] records that this book's own program made exactly that mistake and how it was caught.

The call tree

The recursion for n equal to three makes one call for n equal to two, which makes one call for n equal to one, which returns. So the calls are a chain of length n, not a branching tree, and the depth of the recursion is n.

That is worth stating because it is the resource that runs out first. Each pending call occupies space, so a recursion of depth n uses space proportional to n even if it returns nothing. A recursion of depth one million will exhaust the stack of most languages, and Python's default limit is a thousand.

Recursion against iteration, on the same problem

The iterative ruleThe recursive rule
Sūtra8.22 read as "next row from this row"8.22 read as "table from smaller table"
What is held at onceone rowevery intermediate table
Extra spaceproportional to nproportional to n times 2 to the power n
Extra timenoneabout double, since every smaller table is built
Depth of nestingnonen
Which is clearerneither, and that is the pointneither

Both are in the text, in the sense that the same sūtra supports both readings. Neither is better in the abstract. The iterative form is what you want if you are producing one row at a time; the recursive form is what you want if you are proving something about the table, because an induction on n follows the recursion exactly.

Proof by induction, which is the recursion's real payoff

Claim: the prastāra for n syllables has 2 to the power n rows.

Base case. For n equal to one the table has two rows, and 2 to the power 1 is 2.

Inductive step. Suppose the table for n minus one has 2 to the power n minus one rows. The rule writes it twice, so the new table has twice that many, which is 2 to the power n.

Therefore the claim holds for every n by induction.

That proof is three lines and it is available only because the definition is recursive. The iterative rule gives the same fact by the counting argument in [The Prastāra Rule Read as an Algorithm], which is longer. A recursive definition makes induction easy, and that is the strongest reason to have one.

munotes.in81

Recursive Enumeration

What recursion is NOT

It is not the same as a loop that calls itself. The distinguishing feature is the base case. A procedure that calls itself with no base case is not recursion, it is a fault.

It is not always slower. The overhead is a function call, which is small. Where recursion is expensive here is in space, because every intermediate table is retained.

It is not restricted to one recursive call. Halāyudha's rule uses the smaller table twice, but it calls for it once and copies it. A definition that made two separate calls for the same thing would do the work twice, which is the problem [Meru-Prastāra as a Dynamic Programming Model] is about.

Quick revision

  • A recursive definition has a base case, a recursive case, and progress toward the base case.
  • Sūtra 8.20 is the base case, one syllable, two rows. Sūtra 8.22 is the recursive case: the shorter table twice, new column on the right.
  • Halāyudha states the stopping condition himself: "again and again, until the desired prastāra".
  • The recursion's depth is n, and depth costs space even when nothing is returned.
  • Iterative: one row held, no nesting. Recursive: every intermediate table held, depth n.
  • The payoff of recursion is induction: the 2 to the power n count is a three-line proof.

Test yourself

1. Name the three things a recursive definition needs, and give Piṅgala's instance of each.

A base case: the one-syllable table, from 8.20. A recursive case: the table for n minus one written twice with the new column filled, from 8.22. Progress toward the base case: each step reduces n by one, bounded below by one.

2. Build the three-syllable prastāra from the two-syllable one.

The two-syllable table is GG, LG, GL, LL. Written twice with a third column of heavy then light: GGG, LGG, GLG, LLG, then GGL, LGL, GLL, LLL.

3. Prove by induction that the prastāra for n syllables has 2 to the power n rows.

For n equal to one there are two rows. If the table for n minus one has 2 to the power n minus one rows, the rule writes it twice, giving twice as many, which is 2 to the power n. The claim follows for all n.

4. Which resource does recursion consume that iteration does not, and why?

Space for the pending calls, proportional to the depth, because each call that has not yet returned must be remembered. For this rule the depth is n, and in addition every intermediate table is retained.

Contents This chapter on its own page

munotes.in82

Chapter Twenty-Six

Tree Structures, and the Prastāra as a Binary Tree

Syllabus topic Module 1, "Tree structures"

In one line

A tree is a branching structure with one root and no cycles, and the prastāra is exactly the tree of all the choices you make when you fix one syllable at a time.

In the wording you can write in an examination: a tree is a connected structure of nodes joined by edges in which there is exactly one path between any two nodes. It has a distinguished root; each node other than the root has one parent; a node with no children is a leaf; the depth of a node is the number of edges from the root to it; and a binary tree is one in which no node has more than two children. A complete binary tree of depth n has 2 to the power n leaves.

The vocabulary, defined once

Node. A point in the structure. Here, a partial choice of syllables.

Edge. A link from a node to one of its children. Here, the choice of one more syllable.

Root. The single node with no parent. Here, the state before any syllable is chosen.

Child, parent. If an edge goes from A to B then B is a child of A and A is the parent of B.

Leaf. A node with no children. Here, a complete pattern.

Depth of a node. The number of edges from the root to it. Here, the number of syllables chosen.

Height of the tree. The greatest depth of any node.

Path. The sequence of edges from one node to another. Here, a path from the root to a leaf spells a pattern.

Binary tree. Every node has at most two children. Complete binary tree of depth n. Every node above depth n has exactly two children and every leaf is at depth n.

Why the prastāra is one

Building a pattern is a sequence of decisions: choose the first syllable, then the second, and so on. Each decision has two outcomes. So the set of all patterns is the set of all ways of taking n two-way decisions, and that set is exactly the leaves of a complete binary tree of depth n.

Three facts follow immediately and all three are worth having.

The number of leaves is 2 to the power n. Two choices at each of n levels.

The number of edges is 2 to the power n plus one, minus two. That is the running total sūtra of [Saṅkhyā: Counting the Rows Without Writing Them], and it is the same number for the same reason: every edge corresponds to a pattern of some length from one to n.

A pattern is a path. Nothing is stored at a node. The information is in the route.

munotes.in83

Tree Structures, and the Prastāra as a Binary Tree

The tree, drawn by a program

def uddista(pattern):
    y = 1
    for ch in reversed(pattern):
        y *= 2
        if ch == "G":
            y -= 1
    return y

def draw(n, prefix="", depth=0):
    """The prastara as a binary tree. Every leaf is one pattern."""
    if depth == n:
        print("%s    pattern %s, prastara row %d" % ("    " * depth, prefix, uddista(prefix)))
        return
    for symbol in ("G", "L"):
        print("%s+-- %s" % ("    " * depth, symbol))
        draw(n, prefix + symbol, depth + 1)

print("root, no syllables chosen yet")
draw(3)
print()
print("leaves: 2**3 =", 2 ** 3, "   depth:", 3, "   internal nodes:", 2 ** 3 - 1)
root, no syllables chosen yet
+-- G
    +-- G
        +-- G
                pattern GGG, prastara row 1
        +-- L
                pattern GGL, prastara row 5
    +-- L
        +-- G
                pattern GLG, prastara row 3
        +-- L
                pattern GLL, prastara row 7
+-- L
    +-- G
        +-- G
                pattern LGG, prastara row 2
        +-- L
                pattern LGL, prastara row 6
    +-- L
        +-- G
                pattern LLG, prastara row 4
        +-- L
                pattern LLL, prastara row 8

leaves: 2**3 = 8    depth: 3    internal nodes: 7

The complication, which is the point of the chapter

Read the leaves top to bottom: rows 1, 5, 3, 7, 2, 6, 4, 8. That is not the prastāra order.

The reason is exact and it is worth being able to state. The tree's natural traversal fixes the FIRST syllable at the top and varies the LAST one fastest. The prastāra varies the FIRST syllable fastest. So the two orders differ, and they differ by reversing the significance of the positions.

Three ways to put it, all the same fact.

In tree terms. A depth-first traversal that chooses the first syllable at the root is a big-endian traversal, and the prastāra is little-endian.

In binary terms. The traversal enumerates 000, 001, 010, 011, 100, 101, 110, 111 with the first symbol as the most significant bit. The prastāra enumerates the same strings with the first symbol as the least significant.

In practical terms. If you want the tree's traversal to produce the prastāra, put the LAST syllable at the root. Then each level down fixes an earlier syllable, and the leaves come out in Piṅgala's order.

This is not a defect in either structure. It is the thing to be careful about, and a student who can say why the orders differ understands both.

What the tree is good for, and what it is not

Good for: reasoning about the count. The 2 to the power n leaves and the 2 to the power n plus one minus two edges both fall out of the picture in a line.

Good for: pruning. If you only want patterns with no two adjacent heavy syllables, you can abandon a subtree the moment the condition fails, and never visit its leaves. That is a saving the flat table cannot express, and it is how a search problem is usually attacked.

munotes.in84

Tree Structures, and the Prastāra as a Binary Tree

Not good for: storage. Holding the tree costs more than holding the table, because the internal nodes are stored too. The tree is a way of thinking, and the table is the answer.

Not good for: random access. There is no way to jump to the forty-first leaf of a tree without walking. That is what naṣṭa is for, and it is the one thing the tree makes harder rather than easier.

Quick revision

  • Tree vocabulary: node, edge, root, parent, child, leaf, depth, height, path. Binary tree: at most two children. Complete binary tree of depth n: 2 to the power n leaves.
  • The prastāra is the tree of n two-way choices; a pattern is a path from the root to a leaf.
  • Leaves: 2 to the power n. Edges: 2 to the power n plus one, minus two, which is the running total sūtra's figure.
  • The tree's natural traversal is big-endian and the prastāra is little-endian, so the leaf order is NOT the prastāra order. Put the last syllable at the root to fix it.
  • The tree is good for counting and for pruning, bad for storage and for random access.

Test yourself

1. Define leaf, depth and complete binary tree, and say how many leaves a complete binary tree of depth 6 has.

A leaf is a node with no children. The depth of a node is the number of edges from the root to it. A complete binary tree of depth n has two children at every node above depth n and all its leaves at depth n. At depth 6 it has 64 leaves.

2. Why is the prastāra a binary tree, and what does a path represent?

Because a pattern is built by n successive two-way choices of syllable, so the set of patterns is the set of root-to-leaf paths in a complete binary tree of depth n. A path spells the pattern.

3. The tree's leaves come out as rows 1, 5, 3, 7, 2, 6, 4, 8. Explain why, and say how to fix it.

The traversal fixes the first syllable at the root, so the last syllable varies fastest, which makes the first syllable the most significant. The prastāra makes the first syllable the least significant. Putting the last syllable at the root makes the traversal produce Piṅgala's order.

4. Name one thing the tree makes easier than the flat table and one thing it makes harder.

munotes.in85

Tree Structures, and the Prastāra as a Binary Tree

Easier: pruning, since a subtree whose patterns cannot satisfy a condition can be abandoned without visiting its leaves. Harder: random access, since there is no way to jump to the forty-first leaf without walking, which is exactly what naṣṭa does arithmetically.

Contents This chapter on its own page

munotes.in86

Chapter Twenty-Seven

Meru-Prastāra: Halāyudha's Staircase

Syllabus topic Module 1, "Meru-Prastāra and combinatorial expansion"

In one line

The Meru-prastāra is a staircase of numbers in which each cell is the sum of the two above it, and its n-th row tells you how many patterns of n syllables have each possible number of light syllables.

In the wording you can write in an examination: the Meru-prastāra is a triangular array constructed by placing one in the topmost cell, one in each end cell of every row, and in each interior cell the sum of the two cells immediately above it. Its n-th row gives the lagakriyā, that is, the number of metrical patterns of n syllables containing exactly k light syllables, for k from zero to n.

The question it answers

The prastāra tells you all the patterns. Saṅkhyā tells you how many there are. Neither tells you how many have exactly two light syllables, and that is a question a prosodist genuinely asks, because the metres in actual use are described by such conditions.

You could answer it by writing the table and counting. For the jagatī that is 4096 rows. Halāyudha gives an array that answers it for every n and every k at once, and the array is built by addition alone.

Halāyudha's own construction

He introduces it as being for the sake of the lagakriyā, and then describes the drawing. The passage is worth having close to verbatim, because every clause is a rule.

The shape. Having written one square cell above, below it write two cells extending half-way out on both sides; below that three, below that four, as many rows as are desired.

The first cell. In its first cell, having set the number one, apply this rule.

The rule. The number in the cell above is to be entered in full in the two cells below it. In both of those cells give one each. Then in the third row, in the end cells give the one from the cell above; but in the middle cell, having combined the numbers of the two cells above, enter the full amount, and that is the meaning of the word "full".

And again. In the fourth row also, in the end cells place one each; but in the middle cells, having combined the numbers of the two cells above, place the full amount, which is the number three. In what follows the same arrangement applies.

Two things in that are worth pointing at. The cells are staggered, each row extending half a cell beyond the one above on each side, which is why an interior cell sits under two cells and an end cell under one. And the word he glosses, pūrṇa, full, is the operative term: an interior cell gets the two above it combined and entered in full, which is addition.

munotes.in87

Meru-Prastāra: Halāyudha's Staircase

What he then reads off the rows

This is the part that settles what the array is for, and he states it directly.

Row two. In the two-celled row is the arrangement of one syllable: there one is guru and one is laghu.

Row three. In the third row is the prastāra of two syllables: there one is all-guru, two have one laghu, one is all-laghu, and these are the forms in the order of the cells.

Row four. In the fourth row is the prastāra of three syllables: there one is all-guru, three have one laghu, three have two laghus, one is all-laghu.

So the rows he reads are 1 1, then 1 2 1, then 1 3 3 1. In a commentary on Sanskrit metre.

The array, drawn by the rule

def meru(rows):
    """Halayudha's staircase: one at the top; an end cell takes the one above it,
    an inner cell takes the two cells above it, combined."""
    tri = [[1]]
    for _ in range(rows - 1):
        prev = tri[-1]
        tri.append([1] + [prev[i] + prev[i + 1] for i in range(len(prev) - 1)] + [1])
    return tri

WIDTH = 5
tri = meru(8)
for i, row in enumerate(tri):
    pad = " " * (WIDTH * (len(tri) - 1 - i) // 2)
    print("%s%s" % (pad, "".join("%*d" % (WIDTH, x) for x in row)))
                     1
                   1    1
                1    2    1
              1    3    3    1
           1    4    6    4    1
         1    5   10   10    5    1
      1    6   15   20   15    6    1
    1    7   21   35   35   21    7    1

The staggering in the output is the staggering the commentary describes: each row is offset by half a cell so that every interior cell sits beneath two.

How to read a row

Take the row beginning 1 6 15 20. Counting the top cell as row zero, this is row six, so it is about metres of six syllables, the gāyatrī.

Cell, counting from the left from zeroMeaningValue
0patterns with no light syllable, that is all heavy1
1patterns with exactly one light syllable6
2patterns with exactly two15
3patterns with exactly three20
4patterns with exactly four15
5patterns with exactly five6
6patterns with all six light1

Add them: 1 plus 6 plus 15 plus 20 plus 15 plus 6 plus 1 is 64, which is the saṅkhyā for the gāyatrī and the figure Halāyudha states. Every row sums to the number of patterns, which is a check you can run on any row by eye.

munotes.in88

Meru-Prastāra: Halāyudha's Staircase

Why the addition rule is right

Take a pattern of n syllables with exactly k light ones. Look at its last syllable.

If the last syllable is heavy, the first n minus one syllables form a pattern of n minus one syllables with exactly k light ones.

If the last syllable is light, the first n minus one syllables form a pattern of n minus one syllables with exactly k minus one light ones.

Every pattern falls into exactly one of those two cases, and nothing is counted twice. So the count for n and k is the count for n minus one and k, plus the count for n minus one and k minus one.

Those two counts are the two cells above. That is the rule, and the argument is four lines.

The end cells

An end cell is 1 because there is exactly one way to have no light syllables, namely all heavy, and exactly one way to have all of them light. Halāyudha's instruction to place one in each end cell is not an exception to the rule; it is the rule applied where one of the two cells above is missing and counts as nothing.

What the Meru is NOT

It is not the prastāra. The prastāra lists patterns; the Meru counts them by a property. Confusing the two is the commonest error in this topic.

It is not indexed by the row number of a pattern. Cell three of row six does not mean row three. It means the count of patterns with three light syllables.

It is not a table of powers. Its rows sum to powers of two, which is a fact about it, not what it holds.

Quick revision

  • Meru-prastāra: one at the top, ones at the ends, every interior cell the sum of the two above.
  • Halāyudha describes the staggered cells, sets one in the first cell, and glosses "full" as the two above combined.
  • He reads the rows off: 1 1 for one syllable, 1 2 1 for two, 1 3 3 1 for three.
  • Row n, cell k, counting both from zero, is the number of patterns of n syllables with exactly k light syllables.
  • Every row sums to 2 to the power n, which is the saṅkhyā.
  • Why: split the patterns on the last syllable, heavy or light, and the two cases are the two cells above.

Test yourself

1. State the construction rule of the Meru-prastāra in one sentence.

Place one in the topmost cell and one in each end cell of every row, and in every interior cell place the sum of the two cells immediately above it.

2. What does cell 2 of row 5 mean, and what is its value?

munotes.in89

Meru-Prastāra: Halāyudha's Staircase

The number of patterns of five syllables with exactly two light syllables, which is 10.

3. Prove the addition rule.

Classify the patterns of n syllables with k light ones by their last syllable. If it is heavy, the rest is a pattern of n minus one syllables with k light ones; if it is light, the rest has k minus one light ones. The two cases are exclusive and exhaustive, so the count is the sum of those two counts, which are the two cells above.

4. Why is every end cell a one, and why is that not an exception to the rule?

There is exactly one all-heavy pattern and exactly one all-light pattern. It is not an exception because one of the two cells above an end cell does not exist and counts as nothing, so the sum is just the single cell above, which is one.

Contents This chapter on its own page

munotes.in90

Chapter Twenty-Eight

Pascal Triangle and Combinatorics

Syllabus topic Module 1, "Pascal triangle and combinatorics"

In one line

The Meru-prastāra is what is now called Pascal's triangle, and its entries are the binomial coefficients, the number of ways of choosing k things from n.

In the wording you can write in an examination: the binomial coefficient, written C(n, k) and read as n choose k, is the number of distinct subsets of size k that can be formed from a set of n distinct elements. It equals n factorial divided by k factorial times n minus k factorial, it satisfies the addition rule that C(n, k) equals C(n minus 1, k) plus C(n minus 1, k minus 1), and the array of its values is Pascal's triangle, which is the Meru-prastāra.

Why the two problems are the same problem

The lagakriyā asks: how many patterns of n syllables have exactly k light ones?

Choosing which syllables are light is choosing a set of k positions out of n. Once you have chosen them, the pattern is determined: those positions are light and the rest are heavy. So there are exactly as many patterns as there are ways of choosing k positions from n.

That is the whole argument, and it is the one to write in an answer. Not a resemblance: a one-to-one correspondence between patterns and subsets.

The closed form, and how to use it without factorials

C(n, k) is n factorial over k factorial times n minus k factorial. Written out for six choose two: 720 over 2 times 24, which is 720 over 48, which is 15.

Do not compute it that way. Factorials get large fast: 20 factorial is already about 2.4 times ten to the eighteenth, while C(20, 2) is only 190. The practical form multiplies and divides alternately.

C(n, k) = (n / 1) ((n-1) / 2) ((n-2) / 3) ... ((n-k+1) / k)

For six choose two that is 6 over 1 times 5 over 2, which is 15, with nothing larger than 30 appearing.

The three routes to the same numbers

GURU, LAGHU = "G", "L"

def prastara(n):
    row = [GURU] * n
    rows = ["".join(row)]
    while GURU in row:
        k = row.index(GURU)
        row = [GURU] * k + [LAGHU] + row[k + 1:]
        rows.append("".join(row))
    return rows

def lagakriya(n):
    """Counted off the table: how many patterns have exactly k laghus."""
    out = [0] * (n + 1)
    for row in prastara(n):
        out[row.count(LAGHU)] += 1
    return out

def meru(rows):
    tri = [[1]]
    for _ in range(rows - 1):
        prev = tri[-1]
        tri.append([1] + [prev[i] + prev[i + 1] for i in range(len(prev) - 1)] + [1])
    return tri

def choose(n, k):
    num, den = 1, 1
    for i in range(k):
        num *= n - i
        den *= i + 1
    return num // den

print("%-4s %-26s %-26s %-6s %s" % ("n", "counted off the table", "Meru row n", "sum", "2**n"))
tri = meru(9)
for n in range(1, 9):
    counted = lagakriya(n)
    row = tri[n]
    print("%-4d %-26s %-26s %-6d %d"
          % (n, counted, row, sum(counted), 2 ** n))

print()
print("and the same figures from the binomial coefficient, n = 6")
print("  k:      ", "  ".join("%3d" % k for k in range(7)))
print("  C(6,k): ", "  ".join("%3d" % choose(6, k) for k in range(7)))
print("  Meru:   ", "  ".join("%3d" % x for x in tri[6]))
munotes.in91

Pascal Triangle and Combinatorics

n    counted off the table      Meru row n                 sum    2**n
1    [1, 1]                     [1, 1]                     2      2
2    [1, 2, 1]                  [1, 2, 1]                  4      4
3    [1, 3, 3, 1]               [1, 3, 3, 1]               8      8
4    [1, 4, 6, 4, 1]            [1, 4, 6, 4, 1]            16     16
5    [1, 5, 10, 10, 5, 1]       [1, 5, 10, 10, 5, 1]       32     32
6    [1, 6, 15, 20, 15, 6, 1]   [1, 6, 15, 20, 15, 6, 1]   64     64
7    [1, 7, 21, 35, 35, 21, 7, 1] [1, 7, 21, 35, 35, 21, 7, 1] 128    128
8    [1, 8, 28, 56, 70, 56, 28, 8, 1] [1, 8, 28, 56, 70, 56, 28, 8, 1] 256    256

and the same figures from the binomial coefficient, n = 6
  k:         0    1    2    3    4    5    6
  C(6,k):    1    6   15   20   15    6    1
  Meru:      1    6   15   20   15    6    1

Three independent routes: counting the patterns one at a time, adding cells in the staircase, and the closed form. They agree at every value printed, and the first two agree at every n from one to eight, which is 510 patterns counted by hand by the machine.

The identities worth knowing

Each of these is a fact about the triangle that a five-mark question can ask for, and each has a one-line reason.

Symmetry: C(n, k) equals C(n, n minus k). Choosing which k are light is the same as choosing which n minus k are heavy. That is why every row reads the same backwards.

Row sum: the row for n adds to 2 to the power n. Every pattern has some number of light syllables, so summing over all k counts every pattern once. This is the link between the lagakriyā and the saṅkhyā, and it is Halāyudha's own 64 for the gāyatrī.

The ends are one. Exactly one all-heavy pattern and one all-light one.

The second entry is n. There are n positions in which the single light syllable can stand.

The addition rule. Proved in [Meru-Prastāra: Halāyudha's Staircase] by splitting on the last syllable.

munotes.in92

Pascal Triangle and Combinatorics

Worked example: a real prosodic question

How many six-syllable patterns have more heavy syllables than light ones?

By the row. More heavy than light means fewer than three light, so k is 0, 1 or 2. The row for six is 1, 6, 15, 20, 15, 6, 1. So 1 plus 6 plus 15, which is 22.

Checked another way. By symmetry, the number with more light than heavy is also 22, and the number with exactly three of each is 20. Now 22 plus 22 plus 20 is 64, which is the whole table. The two answers are consistent.

That second calculation is worth doing every time. A row of the Meru gives you a free check on any counting answer, because the parts must add to the row sum.

What this is NOT

It is not a claim of priority over Pascal. The claim this book makes is that the array and its construction rule are in Halāyudha's commentary, which is a statement about a text. Questions of who first wrote down what, in which century, in which tradition, are a matter for historians of mathematics and the literature on them is large. An answer should say the array is in the commentary and leave the rest.

It is not the same object as the prastāra. Both are called prastāra in Sanskrit, and they are different things: one lists, the other counts.

It is not limited to two symbols. With three symbols the analogous array is the trinomial triangle, and the prosodic tradition that counts by mātrā rather than by syllable meets a related problem.

Quick revision

  • The lagakriyā and the count of k-subsets of n are the same problem: choosing which syllables are light is choosing which positions.
  • C(n, k) is n factorial over k factorial times n minus k factorial, but compute it by alternating multiplication and division.
  • Identities: symmetry, so rows read the same backwards; row sum 2 to the power n; ends are one; second entry is n; and the addition rule.
  • Three routes agree: counting off the prastāra, the Meru's additions, the closed form.
  • A row gives a free check on any counting answer, since the parts must add to the row sum.

Test yourself

1. Explain in one sentence why the lagakriyā is a binomial coefficient.

Because choosing which k of the n syllables are light is exactly choosing a subset of k positions from n, and the pattern is then determined, so patterns and subsets correspond one to one.

2. Compute C(7, 3) without factorials, showing the steps.

7 over 1 is 7; times 6 over 2 is 21; times 5 over 3 is 35. So C(7, 3) is 35, which is the fourth entry of row seven.

munotes.in93

Pascal Triangle and Combinatorics

3. How many eight-syllable patterns have exactly three light syllables, and how do you check your answer?

Row eight is 1, 8, 28, 56, 70, 56, 28, 8, 1, so the answer is 56. Check it by symmetry against the count with five light syllables, which is also 56, and by the row sum, which is 256.

4. State carefully what this book claims about the Meru-prastāra and Pascal's triangle.

That the array, its construction by adding the two cells above, and the reading of its rows as counts of patterns by number of light syllables are all present in Halāyudha's commentary on Piṅgala. It makes no claim about priority, which is a historical question with its own literature.

Contents This chapter on its own page

munotes.in94

Chapter Twenty-Nine

Meru-Prastāra as a Dynamic Programming Model

Syllabus topic Module 1, "Dynamic Programming Model", "Complexity & Limitations"

In one line

The addition rule computed straight down is ruinously slow because it solves the same subproblem again and again, and dynamic programming is the fix: solve each one once and keep the answer.

In the wording you can write in an examination: dynamic programming is a technique for problems having two properties, optimal substructure, that is, the answer is built from answers to smaller instances of the same problem, and overlapping subproblems, that is, the same smaller instance recurs. It solves each distinct subproblem once and stores the result, either top down by memoisation or bottom up by tabulation. The Meru-prastāra is the tabulated form of the binomial addition rule.

Where the waste comes from

The addition rule says C(n, k) equals C(n minus 1, k) plus C(n minus 1, k minus 1). Written as a recursive function with no memory, it calls itself twice, and each of those calls itself twice again.

Follow C(5, 2) down and you will find that C(3, 1) is computed twice, and C(2, 1) three times, and C(1, 0) five times. The same answers are recomputed because nothing remembers them.

That is the definition of overlapping subproblems, and it is the signal that dynamic programming applies.

The two properties, checked

Optimal substructure. C(n, k) is built from C(n minus 1, k) and C(n minus 1, k minus 1). Present.

Overlapping subproblems. C(n minus 1, k minus 1) is reached both from C(n, k) and from C(n, k minus 1). Present, and the overlap grows with n.

Both present, so the technique applies. If only the first were present, plain recursion would be fine; that is the case for the prastāra's own recursion in [Recursive Enumeration], where each smaller table is needed once.

The three forms

Naive recursion. Apply the rule directly, no memory. Correct and exponentially slow.

Memoisation, top down. Apply the rule, but before computing, look in a table for the answer, and after computing, put it there. Each distinct pair is computed once.

Tabulation, bottom up. Build the answers in an order that guarantees whatever you need is already there. For the Meru that means building row by row, which is exactly what Halāyudha's instruction says to do.

The measurement

CALLS = {"naive": 0, "memo": 0}

def naive(n, k):
    """C(n, k) straight from the addition rule, with no memory."""
    CALLS["naive"] += 1
    if k == 0 or k == n:
        return 1
    return naive(n - 1, k) + naive(n - 1, k - 1)

def memoised(n, k, seen=None):
    """The same rule, remembering each answer the first time it is found."""
    if seen is None:
        seen = {}
    CALLS["memo"] += 1
    if k == 0 or k == n:
        return 1
    if (n, k) in seen:
        return seen[(n, k)]
    seen[(n, k)] = memoised(n - 1, k, seen) + memoised(n - 1, k - 1, seen)
    return seen[(n, k)]

def tabulated(n):
    """Build the Meru row by row, bottom up, keeping only what is needed."""
    row = [1]
    steps = 0
    for _ in range(n):
        row = [1] + [row[i] + row[i + 1] for i in range(len(row) - 1)] + [1]
        steps += len(row)
    return row, steps

print("%-6s %-6s %-8s %-12s %-12s %s" % ("n", "k", "C(n,k)", "naive calls", "memo calls", "table cells"))
for n, k in ((6, 3), (10, 5), (14, 7), (18, 9), (22, 11)):
    CALLS["naive"] = CALLS["memo"] = 0
    a = naive(n, k)
    b = memoised(n, k)
    row, steps = tabulated(n)
    assert a == b == row[k], (a, b, row[k])
    print("%-6d %-6d %-8d %-12d %-12d %d" % (n, k, a, CALLS["naive"], CALLS["memo"], steps))
munotes.in95

Meru-Prastāra as a Dynamic Programming Model

n      k      C(n,k)   naive calls  memo calls   table cells
6      3      20       39           19           27
10     5      252      503          51           65
14     7      3432     6863         99           119
18     9      48620    97239        163          189
22     11     705432   1410863      243          275

Read the last row. The answer is 705,432 and the naive recursion makes 1,410,863 calls to find it, which is twice the answer minus one. The memoised version makes 243. The naive cost is proportional to the answer itself, which is the worst possible behaviour, and the reason is visible: with no memory the recursion effectively counts the subsets one at a time.

The costs, stated

FormTimeSpaceWhen to use it
Naive recursionproportional to C(n, k), which grows exponentiallydepth nnever, for this problem
Memoisationproportional to n times k, the number of distinct pairsn times k, plus depth nwhen only a few cells are wanted
Tabulationproportional to n squared, since row i has i plus one cellsone row at a time, so proportional to nwhen a whole row or the whole array is wanted
Closed formproportional to kconstantwhen one value is wanted and no row is needed

The last row is worth noticing: for a single value, the multiplicative formula of [Pascal Triangle and Combinatorics] beats all three, and dynamic programming is not always the answer even when it applies.

Where Halāyudha's instruction sits in this

His rule is the tabulated form, exactly. He says to draw the cells, put one at the top, and fill each row from the row above. That is bottom-up dynamic programming in the ordinary sense: no subproblem is solved twice because the order of filling guarantees the dependencies are already present.

What is not in his text. No statement that the alternative is slow, no count of operations, and no notion of a subproblem. The structure is there and the analysis is ours. Saying so is the honest form of the claim, and it is the form MU's "critically assess" wants.

munotes.in96

Meru-Prastāra as a Dynamic Programming Model

Worked example: MU's implementation topic

Her topic 6 asks for the Meru-prastāra as a Pascal triangle and dynamic programming model. A submission that scores well answers four questions and the chapter has now given all four.

What is the IKS concept? Halāyudha's staircase, constructed by adding the two cells above, and read as the count of patterns by number of light syllables.

What is the CS concept? Bottom-up dynamic programming over the binomial addition rule.

What is the mapping? A cell is a subproblem; the addition rule is the recurrence; filling row by row is the tabulation order; the end cells are the base cases.

What is the evidence? The call counts above. A submission that states them has measured its claim, which is what her "analysis of correctness and limitations" asks for.

What dynamic programming is NOT

It is not recursion with a cache, in general. Memoisation is one of its two forms. Tabulation involves no recursion at all.

It is not applicable to every recursive problem. Without overlapping subproblems there is nothing to save. The prastāra's own recursion has none.

It is not free. Memoisation costs space proportional to the number of distinct subproblems, and for a large n times k that can be the binding constraint. Tabulation over one row at a time avoids it, which is why the bottom-up form is often preferred.

Quick revision

  • Two conditions: optimal substructure and overlapping subproblems. Both hold for the binomial addition rule.
  • Three forms: naive recursion, memoisation top down, tabulation bottom up.
  • Measured: C(22, 11) is 705,432 and the naive recursion takes 1,410,863 calls; memoisation takes 243.
  • Naive time is proportional to the answer; memoisation to n times k; tabulation to n squared with space for one row; the closed form to k with constant space.
  • Halāyudha's instruction is the tabulated form. The structure is his; the analysis is ours.

Test yourself

1. State the two conditions for dynamic programming and show that the binomial addition rule meets both.

Optimal substructure: C(n, k) is the sum of C(n minus 1, k) and C(n minus 1, k minus 1). Overlapping subproblems: C(n minus 1, k minus 1) is reached from both C(n, k) and C(n, k minus 1), and the overlap grows with n.

2. Why is the naive recursion's cost proportional to the answer?

Because with no memory the recursion bottoms out at a base case once for every subset being counted, so the number of leaf calls is the value of C(n, k) itself.

munotes.in97

Meru-Prastāra as a Dynamic Programming Model

3. Distinguish memoisation from tabulation and give one advantage of each.

Memoisation is top down: apply the recurrence but look up and store each answer. Tabulation is bottom up: fill the array in an order that guarantees dependencies are present. Memoisation computes only the cells actually needed; tabulation needs no recursion and can keep only one row at a time.

4. Which form does Halāyudha's instruction correspond to, and what is absent from his text?

Tabulation, since he fills each row from the row above. Absent are any statement that an alternative would be slow, any count of operations, and any notion of a subproblem.

Contents This chapter on its own page

munotes.in98

Chapter Thirty

Lagakriyā and Adhvayoga: The Rest of the Pratyayas

Syllabus topic Module 1, "Prastāra (systematic enumeration of patterns)", "Meru-Prastāra and combinatorial expansion"

In one line

Lagakriyā counts the patterns by how many light syllables they have, and adhvayoga measures how much space the written table takes; Halāyudha gives the first and declines the second as too slight to bother with.

In the wording you can write in an examination: the pratyayas are the operations classically stated on the set of metrical patterns. Six are named: prastāra, the enumeration; naṣṭa, recovery of a pattern from its index; uddiṣṭa, recovery of an index from a pattern; lagakriyā, the count of patterns by their number of light syllables; saṅkhyā, the total count; and adhvayoga, the measurement of the written extent of the table.

Why this chapter exists

MU's printed labels name prastāra, naṣṭa, uddiṣṭa, the Meru-prastāra and combinatorial expansion. She does not name lagakriyā or adhvayoga by those words. But the Meru-prastāra IS the lagakriyā, Halāyudha says so in the sentence that introduces it, and adhvayoga is the sixth member of a set of six that a question can reasonably ask you to list.

So the set is completed here, once, so that nothing in the tradition's own organisation of the topic is missing from the book.

The six, and how they relate

PratyayaThe question it answersCostChapter
prastārawhat are all the patterns?n times 2 to the power n[Prastāra: The Table of Every Pattern]
naṣṭawhat is the pattern in row r?proportional to n[Naṣṭa: From a Row Number Back to the Pattern]
uddiṣṭawhat row is this pattern?proportional to n[Uddiṣṭa: From a Pattern to Its Row Number]
lagakriyāhow many patterns have k light syllables?proportional to k, or n squared for a whole row[Meru-Prastāra: Halāyudha's Staircase]
saṅkhyāhow many patterns are there?proportional to log n[Saṅkhyā: Counting the Rows Without Writing Them]
adhvayogahow much space does the written table take?proportional to log nthis chapter

Read the cost column downward and something becomes visible. The pratyayas are ordered, roughly, from the most expensive to the cheapest, and every one after the first avoids building the table. Whether the tradition arranged them that way deliberately is not something this book can settle, but the pattern is there.

Lagakriyā

The name means the operation on the la, that is, on the light syllables. Halāyudha introduces the Meru with the words: now, for the accomplishment of the lagakriyā for one, two, three and so on, as far as desired, he shows the Meru-prastāra.

So the Meru-prastāra is the instrument and the lagakriyā is the question. A student who says the Meru-prastāra is a pratyaya has named the tool and not the operation, which is a small mistake and an avoidable one.

Everything else about it is in [Meru-Prastāra: Halāyudha's Staircase] and [Pascal Triangle and Combinatorics] and is not repeated.

munotes.in99

Lagakriyā and Adhvayoga: The Rest of the Pratyayas

Adhvayoga, and Halāyudha's own refusal

At the close of the eighth chapter he records that some count a sixth pratyaya, which he calls the measuring of the extent, and says it is not stated here because it is very slight, adding that it varies with the metre. Then: thus the group of pratyayas is complete.

Two things worth taking from that.

The set is a set. He treats the pratyayas as a closed group and says so, which is why "name the six pratyayas" is a fair question.

He gives a reason for an omission. That is a habit of a formal tradition: a gap is declared rather than left.

The arithmetic, stated as the obvious reading

Since the commentary declines to give the rule, this book gives the arithmetic and labels it as the obvious reading rather than as the text's own.

Worked. What is being measured. A written prastāra of n syllables has 2 to the power n rows, each n symbols wide, and a gap between consecutive rows.

So the extent, counted in symbol widths, is n times 2 to the power n for the symbols, plus 2 to the power n minus 1 for the gaps.

SyllablesRowsSymbol widths including gaps
123
2411
3831
41679
532191
664447

What makes this a genuine pratyaya rather than a triviality is that, like saṅkhyā, it is answered without drawing the table. And the modern point is exactly the one a programmer meets: the SPACE a computation needs is a resource to be estimated in advance, separately from the time. A tradition that lists "how much room will this take" beside "how many are there" is thinking about both.

Where Halāyudha is right that it is slight

He is right, and it is worth saying so rather than inflating the sixth pratyaya's importance.

It is derivable from saṅkhyā in one step. Once you know there are 2 to the power n rows, multiplying by the width is not a new idea.

It depends on the writing surface. How wide a symbol is, and whether rows are spaced, is a matter of the palm leaf and not of the mathematics. That is presumably what he means by saying it varies.

Nothing else depends on it. No other pratyaya uses its result.

Quick revision

  • Six pratyayas: prastāra, naṣṭa, uddiṣṭa, lagakriyā, saṅkhyā, adhvayoga.
  • Lagakriyā is the question "how many have k light syllables"; the Meru-prastāra is the instrument that answers it. Halāyudha says the Meru is shown for the sake of the lagakriyā.
  • Adhvayoga measures the written extent. Halāyudha records that some count it and leaves it out as very slight and as varying with the metre.
  • Its arithmetic: n times 2 to the power n symbol widths, plus 2 to the power n minus 1 gaps. Six syllables gives 447.
  • Every pratyaya after the first avoids building the table.
  • He declares the omission rather than leaving a gap, which is a habit of a formal tradition.
munotes.in100

Lagakriyā and Adhvayoga: The Rest of the Pratyayas

Test yourself

1. Name the six pratyayas and the question each answers.

Prastāra, all the patterns; naṣṭa, the pattern in a given row; uddiṣṭa, the row of a given pattern; lagakriyā, how many patterns have a given number of light syllables; saṅkhyā, how many there are; adhvayoga, how much space the written table takes.

2. What is the difference between the lagakriyā and the Meru-prastāra?

The lagakriyā is the operation, the counting of patterns by their number of light syllables. The Meru-prastāra is the array that performs it, which Halāyudha says he sets out for the sake of the lagakriyā.

3. Why does Halāyudha omit adhvayoga, and is he right?

Because it is very slight and varies with the metre. He is right: it follows from saṅkhyā in one multiplication, it depends on the writing surface rather than on the mathematics, and nothing else uses its result.

4. Compute the written extent of the three-syllable prastāra in symbol widths, with gaps.

Eight rows of three symbols is 24, plus seven gaps, which is 31.

Contents This chapter on its own page

munotes.in101

Chapter Thirty-One

Algorithmic Generation: What Piṅgala Actually Achieved

Syllabus topic Module 1, "Algorithmic generation", "Complexity & Limitations"

In one line

Piṅgala's six rules between them generate, index, count and measure a space of 2 to the power n objects, and only the first of them ever touches the whole space.

In the wording you can write in an examination: algorithmic generation is the production of the members of a set by the repeated application of a stated rule rather than by storing them. The Chandaḥśāstra's eighth chapter supplies a generator for the full set of metrical patterns, two index operations that address the set without generating it, a counting operation that gives its size in logarithmic time, and a combinatorial operation that counts its members by a property, all stated as effective procedures.

What is genuinely achieved, in one table

This is the chapter to revise from, and the third column is the part that is usually left out of accounts of Piṅgala.

RuleWhat it doesCostDoes it build the table?
prastāraevery pattern, in a fixed ordern times 2 to the power nyes, that is its job
naṣṭathe pattern in row rproportional to nno
uddiṣṭathe row of a patternproportional to nno
saṅkhyāhow many patternsproportional to log nno
lagakriyāhow many with k light syllablesproportional to k for one valueno
adhvayogathe written extentproportional to log nno

Five of the six avoid the table. For the jagatī that is the difference between 49,152 symbol operations and twelve. The set of six is, in effect, an interface to a structure that is never materialised, and that is a design worth naming.

The five properties, collected

[The Prastāra Rule Read as an Algorithm] tested the prastāra against the five properties of an algorithm. The other five rules pass the same test, and where any of them is weaker it is worth saying which.

Input, definiteness, effectiveness, output. All six are unambiguous, executable by hand, and produce what they promise.

Finiteness. The prastāra halts after 2 to the power n minus 1 steps, proved. Naṣṭa and uddiṣṭa halt after exactly n steps by construction. Saṅkhyā halts because halving a positive integer reaches one and subtracting one from an odd number reduces it, so the descent is strictly decreasing. Lagakriyā halts because each row of the Meru is finite and rows are built in order.

Where the tradition is weakest. Adhvayoga, whose rule Halāyudha declines to state in full, so its definiteness rests on the reader supplying the arithmetic. That is a real gap and this book says so in [Lagakriyā and Adhvayoga: The Rest of the Pratyayas].

What is NOT achieved, stated plainly

A five-mark answer asked to assess critically should have three of these ready.

munotes.in102

Algorithmic Generation: What Piṅgala Actually Achieved

No machine. Every rule is executed by a trained person. The gap between an algorithm and a machine that runs it is most of the history of computing, and none of it is here.

No notion of cost. Not one of the six rules is accompanied by a statement that it is cheaper than the alternative. The saṅkhyā saving is real and unremarked. The costs in the table above are ours.

No base, and no arithmetic on patterns. The order agrees with binary counting, which is proved in [Binary Number Systems, and Exactly Where Piṅgala's Order Agrees], and the text nowhere treats a pattern as a number or adds two of them.

No generalisation. Saṅkhyā computes powers of two, not powers. The Meru counts light syllables in a two-symbol alphabet, not selections in general. The generalisations are immediate and the text does not make them.

No validation. No rule checks its input. Naṣṭa will return a pattern for a row that does not exist.

The claim this book does make

Stated once, carefully, because it is the sentence a good answer needs.

Piṅgala chose a representation in which the questions of his subject became decidable by arithmetic, and then stated the arithmetic as procedures. The representation is the two-valued syllable. The questions are enumeration, addressing and counting. The procedures are the six pratyayas. Every part of that sentence is a fact about the text, and the whole of it is what a computer scientist means by good design.

Worked example: the answer to a cross-module question

Q.3 of MU's paper draws on both modules. Here is the shape of an answer that uses this chapter, written at the length five marks allows.

Question. Relate Piṅgala's treatment of metre to the idea of a formal system as it appears elsewhere in this course.

Answer. Both Piṅgala and Pāṇini fix a finite alphabet, state rules over it, and provide for the case where a result is needed that has not been written out. Piṅgala's alphabet has two members and his rules generate and address the whole set of strings; Pāṇini's alphabet is the sounds of Sanskrit and his rules derive particular strings on demand. The difference is the interesting part: Piṅgala's space is complete and uniform, so arithmetic suffices to navigate it, whereas Pāṇini's target language is a proper subset of the strings his alphabet admits, so his system needs conditions, exceptions and a conflict policy that Piṅgala's has no occasion for. Nyāya, in Module II, supplies the third case, in which the rules are about admitting a claim rather than producing a string, and the enforcement is social rather than arithmetical.

That answer is four sentences and it touches three of the six syllabus areas. That is the register Q.3 is asking for.

munotes.in103

Algorithmic Generation: What Piṅgala Actually Achieved

Quick revision

  • Six rules: prastāra generates; naṣṭa and uddiṣṭa index; saṅkhyā counts; lagakriyā counts by a property; adhvayoga measures.
  • Five of the six never build the table. For the jagatī that is twelve operations against about 49,000.
  • All six are definite and effective; the finiteness of each has a short argument; adhvayoga is the weakest because its rule is not stated in full.
  • Not achieved: no machine, no statement of cost, no base or arithmetic on patterns, no generalisation, no validation.
  • The claim that is true: he chose a representation that made his questions arithmetical, and then stated the arithmetic.

Test yourself

1. Which of the six pratyayas builds the table, and what follows for the rest?

Only the prastāra. The other five answer their questions without materialising the set, so they behave like an interface to a structure that is never stored.

2. Give the cost of each of naṣṭa, saṅkhyā and prastāra, for a twelve-syllable metre.

Naṣṭa is twelve steps. Saṅkhyā is five operations for n equal to twelve. The prastāra is 4096 rows of twelve symbols, about 49,000 symbol operations.

3. Name three things the text does not do, that a careless account of Piṅgala claims.

It never states that any rule is cheaper than an alternative; it never treats a pattern as a number or performs arithmetic on one; and it never generalises saṅkhyā beyond powers of two or the Meru beyond a two-symbol alphabet.

4. State the claim this book makes about Piṅgala in one sentence.

He chose a representation, the two-valued syllable, in which the questions of his subject became decidable by arithmetic, and then stated that arithmetic as a set of effective procedures.

Contents This chapter on its own page

munotes.in104

Chapter Thirty-Two

Pāṇini's Aṣṭādhyāyī, and the Text We Are Reading

Syllabus topic Module 1, "Pāṇini’s Aṣṭādhyāyī as a Generative Formal Grammar", "Examination of the structural architecture of Pāṇini’s grammatical system", "Topics:"

In one line

The Aṣṭādhyāyī is a grammar of Sanskrit in about four thousand rules, arranged in eight books of four quarters each, and it derives words rather than describing them.

In the wording you can write in an examination: the Aṣṭādhyāyī of Pāṇini is a sūtra text in eight adhyāyas, each divided into four pādas, which states the grammar of Sanskrit as an ordered system of rules operating on roots, affixes and stems. It is generative rather than descriptive: given a root and a meaning, its rules produce the correct form, and the tradition treats a form as correct if and only if the rules derive it.

The one sentence that explains what kind of book it is

A descriptive grammar says: here are the forms, sorted by type. A generative grammar says: here are the rules, and the forms are whatever they produce.

The Aṣṭādhyāyī is the second kind, and that is why it is on this syllabus. MU's own label is "as a Generative Formal Grammar", and the word generative is the load-bearing one.

The consequence is one a programmer will recognise. The grammar's correctness is a property of the whole rule set, not of any one rule, because a rule can only be judged by what the system produces with it in place. Changing one rule can break a form derived six rules away, and the commentary literature is very largely about exactly that.

The size and the shape

ItemWhat it is
adhyāyaa book. There are eight, which is what the title means
pādaa quarter of a book. Each adhyāya has four
sūtraa rule. They are numbered by book, quarter and position, so 7.3.84 is the eighty-fourth rule of the third quarter of the seventh book
How manyclose to four thousand, and editions differ slightly on the count
The Śivasūtrasfourteen short aphorisms printed before the grammar proper, which order the sounds. [The Fourteen Śivasūtras and the Pratyāhāra]

Worked. The numbering is the address of a rule and it is used constantly. Learn to read it: three numbers, book then quarter then position. So 7.3.84 is the eighty-fourth rule of the third quarter of the seventh book, and 1.1.3 is the third rule of the first quarter of the first book. Every cross-reference in the commentaries is in that form, and so is every citation in this book.

The accompanying lists

The grammar does not stand alone. Three lists come with it and a rule can refer to any of them.

The Dhātupāṭha, the list of verbal roots. Vasu's translation refers to it constantly, and crucially the roots in it carry markers: he writes, of one root, that it is marked in the Dhātupāṭha with a particular letter and that therefore a named rule applies to it. A marker on a root is a tag that changes which rules may touch it, which is the subject of [Anubandha: The Marker Letter as a Type Tag].

munotes.in105

Pāṇini's Aṣṭādhyāyī, and the Text We Are Reading

The Gaṇapāṭha, lists of words grouped so that a rule can name a whole group by its first member.

The Uṇādisūtras, a supplementary set for words the main grammar does not derive.

The point for this paper is architectural. The rules are separated from the data they operate on, and the data is itself annotated. That is a separation a modern system makes too, and it is the reason the grammar can be as small as it is.

The translation this book quotes

Every Pāṇini quotation in this book comes from Srisa Chandra Vasu's English translation, published from Allahabad between 1891 and 1898, one volume per book of the grammar. Vasu died in 1918, so the translation is out of copyright and may be quoted with attribution.

What makes it the right source for this paper, rather than a modern study, is that it gives each sūtra in translation with a note on its type and its cross-references, which is exactly the material a chapter on rule architecture needs.

What the tradition means by correctness

This is worth stating because it is unusual and it bears on the comparison with formal grammars.

A form is correct if the rules derive it. Not if it is attested, and not if it sounds right. The grammar is the authority, which is why the tradition argues so hard about the rules themselves.

So the grammar is a specification of a language, in the strict sense. The set of strings it derives is the language. That is exactly the modern definition of a language generated by a grammar, and it is the basis of the comparison in [Context-Free Grammar, and Whether Pāṇini Wrote One].

And it has a boundary problem the modern definition also has. If the grammar derives a form nobody uses, is the grammar wrong or the usage incomplete? The commentaries take both positions in different places, and a modern grammar engineer meets the same question every week.

What this book does not claim about the text

Not a date. The estimates in the scholarly literature span several centuries and nothing here depends on one.

Not completeness. Whether the Aṣṭādhyāyī derives every form of the Sanskrit its author knew is a scholarly question with a long literature. This book quotes rules and does not adjudicate.

Not that the tradition's own analysis is superseded. The four precedence principles this book sets out are the tradition's, not ours. Where the tradition disputes a rule's reading, the dispute is reported.

munotes.in106

Pāṇini's Aṣṭādhyāyī, and the Text We Are Reading

Why MU's "Topics:" row is claimed here

MU's printed page leaves the bare word "Topics:" on a line of its own immediately above the four Pāṇini labels, as an artefact of her layout. It carries no subject matter. It is recorded in this book's contract so that the completeness check examines every row she printed rather than skipping one, and it is claimed here because this is the chapter that opens the block it introduces.

Quick revision

  • Aṣṭādhyāyī: eight books, four quarters each, close to four thousand rules. A citation is book, quarter, position, as in 7.3.84.
  • It is generative: the rules produce forms, and a form is correct if and only if the rules derive it.
  • Three accompanying lists: the Dhātupāṭha of roots, the Gaṇapāṭha of word groups, the Uṇādisūtras. Roots in the Dhātupāṭha carry markers that change which rules apply.
  • Rules are separated from annotated data, which is why the grammar is small.
  • Quoted here from Srisa Chandra Vasu's English translation, Allahabad, 1891 to 1898, out of copyright.
  • Correctness is derivability, which makes the grammar a specification of a language in the modern sense.

Test yourself

1. What does it mean to call the Aṣṭādhyāyī generative rather than descriptive?

It states rules that produce forms from roots and affixes, rather than cataloguing forms already collected. A form counts as correct if and only if the rules derive it.

2. Read the citation 6.1.77 and say what each number means.

The seventy-seventh sūtra of the first quarter of the sixth book. A citation is always book, then pāda, then position within the pāda.

3. Name the three lists that accompany the grammar and say why the separation matters.

The Dhātupāṭha of verbal roots, the Gaṇapāṭha of word groups, and the Uṇādisūtras. Separating annotated data from rules is what allows a rule to name a whole class in a few syllables, which is why the grammar can be as small as it is.

4. State the tradition's criterion of correctness, and one problem it creates.

A form is correct if the rules derive it. The problem is the boundary case: if the rules derive a form nobody uses, it is not settled whether the rules are wrong or the record of usage is incomplete, and the commentaries take both positions.

Contents This chapter on its own page

munotes.in107

Chapter Thirty-Three

The Six Kinds of Sūtra, in Vasu's Own Words

Syllabus topic Module 1, "Meta-rules and rule precedence", "Rule-based generative structure"

In one line

Pāṇini's rules come in six declared kinds, and five of the six do something other than change a form.

In the wording you can write in an examination: the tradition classifies the sūtras of the Aṣṭādhyāyī into six kinds: saṃjñā, a definition; paribhāṣā, a rule of interpretation; vidhi, the statement of a general operation; niyama, a restriction on the scope of a rule; adhikāra, a governing rule whose content applies to the rules that follow it; and atideśa, the extension of a rule to a case by analogy. Only vidhi rules perform an operation on a form; the other five govern how rules are read and applied.

The passage

Vasu, in his introduction to Book I of his translation, immediately after the opening sūtra and just before he prints the Śivasūtras:

An aphorism or sūtra is of six kinds, saṃjñā or "a definition", paribhāṣā or the "key to interpretation", vidhi or "the statement of a general rule", niyama or "a restrictive rule", adhikāra or "a head or governing rule, which exerts a directing or governing influence over other rules", and atideśa or "extended application by analogy".

That is a typed rule system, declared before the rules begin.

The six, and what each is in a modern system

KindWhat it does in the grammarThe nearest modern thing
saṃjñāfixes a technical term, so later rules can use one word instead of a descriptiona type declaration, or a named constant
paribhāṣāsays how rules are to be read, or supplies a word they omita meta-rule, or a convention in the language definition
vidhistates an operation: this is substituted for that, in these conditionsa production, or a rewrite rule
niyamanarrows the scope of a rule already stateda constraint, or a guard
adhikārastates something that is understood in every rule that follows it, until cancelleda lexical scope, or an inherited attribute
atideśasays that a case not covered is to be treated as if it were some covered caseinheritance, or a macro expansion

Five of the six are not operations. That ratio is the fact to take away. A rule system large enough to be interesting spends most of its rules on the management of rules, and Pāṇini's tradition gave each kind of management a name.

Each kind, with an instance this book has read

saṃjñā. Sūtra 1.1.1 defines the term vṛddhi. Vasu: "ā, ai and au are called vṛddhi." Nothing is derived; a word is given a meaning that about a hundred later rules use.

paribhāṣā. Sūtra 1.1.3, and Vasu annotates it as such in as many words: "This is a paribhāṣā sūtra." Its content is that where guṇa or vṛddhi is enjoined without saying of what, the ik vowels are meant. [Meta-Rules: Paribhāṣā, and Rules About Rules] works it out.

munotes.in108

The Six Kinds of Sūtra, in Vasu's Own Words

vidhi. Sūtra 7.3.84, in Vasu's translation: "when a sārvadhātuka or an ārdhadhātuka affix follows there is guṇa of the base." An operation with a condition and a result.

niyama. Vasu's own note on 7.3.84 is the clearest illustration available: the sūtra does not say what is gunated, "and to complete the sense, the word ikaḥ must be read into the sūtra", which restricts the operation to the ik vowels. The restriction is supplied by a paribhāṣā and has the effect of a niyama on the vidhi's scope.

adhikāra. Vasu's translation records governing rules by their effect: for instance a rule declaring that in the sūtras up to a stated one, a certain word is to be understood. From that point every following rule inherits the word without repeating it. The inheritance is cancelled where the range ends.

atideśa. A rule saying that a substitute is to be treated as the original for the purposes of other rules. Vasu discusses this at length under the term sthānivat, and it is exactly a case of extension by analogy: what is true of the original is asserted of its replacement.

Why a typed rule system is the interesting thing here

This is the argument to make in an answer, and it is stronger than any comparison with a particular modern formalism.

An untyped rule set cannot state a policy. If every rule is just "rewrite this as that", there is no way to say "a definition may never conflict with an operation", because nothing distinguishes the two.

A typed rule set can. Once rules have kinds, the tradition can and does say things like: a special rule sets aside a general one. That statement is about kinds, and it is the apavāda principle of [Rule Precedence: The Four Principles].

And the types are declared, not inferred. Vasu's list is given before the grammar starts. A reader knows what sort of thing each rule is, and the commentary says which for a rule that is disputed.

Worked example: five rules, one form

Deriving a single verb form touches five of the six kinds, which is why the classification is worth having.

The task. From the root meaning "to be", with a present-tense third person singular ending, produce the stem with the correct vowel.

saṃjñā. 1.1.1 and its neighbours give meanings to vṛddhi and guṇa, so that later rules can name the operation in one word.

adhikāra. A governing rule establishes the range of sūtras within which the operation being discussed is understood, so 7.3.84 need not restate it.

vidhi. 7.3.84 fires: a sārvadhātuka affix follows, so there is guṇa of the base.

munotes.in109

The Six Kinds of Sūtra, in Vasu's Own Words

paribhāṣā. 1.1.3 supplies the missing word: guṇa of the ik vowels of the base, not of anything else.

niṣedha, the prohibition, which the tradition treats as a species of niyama. 1.1.5 would block the operation if the affix carried an indicatory marker. Here it does not, so the operation stands.

Result. The stem's vowel is replaced by its guṇa. [Rule-Based Generative Structure] carries the derivation through, and [Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine] implements exactly this chain.

What this classification is NOT

It is not Pāṇini's own list. The six kinds are the tradition's analysis of his rules, reported by Vasu. The sūtras themselves do not carry a label saying which kind they are, and for some rules the tradition disagrees.

It is not exhaustive in practice. The tradition uses further terms, niṣedha for a prohibition among them, and different commentators draw the lines differently.

It is not the same as a modern type system. Nothing checks the types. A rule that behaves like two kinds at once is handled by argument, not by a tool.

Quick revision

  • Six kinds, from Vasu: saṃjñā a definition, paribhāṣā a key to interpretation, vidhi a general rule, niyama a restrictive rule, adhikāra a governing rule, atideśa extension by analogy.
  • Only vidhi performs an operation. Five of the six manage rules.
  • Modern counterparts: type declaration, meta-rule, production, constraint, lexical scope, inheritance.
  • Instances read in Vasu: 1.1.1 saṃjñā; 1.1.3 paribhāṣā, and he says so; 7.3.84 vidhi; the sthānivat discussion for atideśa.
  • The value of types: they let a policy be stated about rules, which is what the precedence principles are.
  • The classification is the tradition's, not the text's; nothing checks it; commentators disagree on particular rules.

Test yourself

1. Name the six kinds of sūtra with a one-line gloss each.

Saṃjñā, a definition; paribhāṣā, a rule about how rules are read; vidhi, a general operation; niyama, a restriction on a rule's scope; adhikāra, a governing rule understood in those that follow; atideśa, the extension of a rule to an uncovered case by analogy.

2. Which kind performs an operation, and why does the answer matter?

Only vidhi. It matters because it shows that five of the six kinds exist to manage rules rather than to change forms, which is what a large rule system actually spends itself on.

3. Give a modern counterpart for adhikāra and for atideśa, and justify each.

Adhikāra is a lexical scope or an inherited attribute: something stated once is understood in every rule within a range until cancelled. Atideśa is inheritance or macro expansion: a case not covered is treated as if it were a covered one, so the covered rule's behaviour is reused.

munotes.in110

The Six Kinds of Sūtra, in Vasu's Own Words

4. Whose classification is this, and what follows for how you should cite it?

It is the tradition's analysis of Pāṇini's rules, reported in Vasu's introduction. So it should be cited as the tradition's classification, not as something the sūtras themselves declare, and an answer should note that commentators disagree about particular rules.

Contents This chapter on its own page

munotes.in111

Chapter Thirty-Four

The Fourteen Śivasūtras and the Pratyāhāra

Syllabus topic Module 1, "Examination of the structural architecture of Pāṇini’s grammatical system", "Sutra method as compressed symbolic encoding"

In one line

Fourteen short aphorisms list the sounds of Sanskrit in an order chosen so that forty-two sets a grammar needs can each be named in two letters.

In the wording you can write in an examination: the Śivasūtras, also called the Māheśvara or Pratyāhāra Sūtras, are fourteen aphorisms prefixed to the Aṣṭādhyāyī which enumerate the sounds of Sanskrit in a deliberate order, each aphorism ending in a marker consonant that is not itself a member of the list. A pratyāhāra is a two-character abbreviation formed, by sūtra 1.1.71, from the first sound of a span together with the marker that closes the span, and it denotes every sound between them.

The problem being solved

A grammar rule needs to talk about sets of sounds. "Before any vowel." "After any consonant." "Before a soft unaspirated stop."

Write those out in full and every rule becomes a list. Give each set a name and every rule becomes short, but then you need a hundred names and a table of what each contains.

Pāṇini does something better than either. He puts the sounds in one order such that every set his grammar needs is a contiguous span of that order. Then a span can be named by its two endpoints, and no table of contents is required: the name IS the definition.

The fourteen

These are printed from the digital sūtrapāṭha of the Aṣṭādhyāyī, and Vasu's translation states that there are fourteen of them and that they contain the arrangement of the Sanskrit sounds for grammatical purposes.

1. a i u N

2. r l K

3. e o NG

4. ai au C

5. ha ya va ra T

6. la N

7. na ma nga na na M

8. jha bha NY

9. gha dha dha SH

10. ja ba ga da da S

11. kha pha cha tha tha ca ta ta V

12. ka pa Y

13. sa sa sa R

14. ha L

The capital letter at the end of each line is the marker, called an it. Vasu's translation records that these final consonants are not part of the list of sounds: they are there to be used in forming names, and one of them, he notes, is used in two different aphorisms, which creates an ambiguity the tradition then has to resolve.

Read the aphorisms as one long sequence with fourteen bookmarks in it. The sounds are the sequence; the markers are the bookmarks.

The rule that forms a name

Sūtra 1.1.71, in the sūtrapāṭha: ādir antyena sahetā. The initial, with the final marker.

Vasu's gloss of the mechanism, in his own words: a pratyāhāra stands "for all the other letters intervening between it and the non-efficient letter", the non-efficient letter being the marker.

munotes.in112

The Fourteen Śivasūtras and the Pratyāhāra

So: take a sound, take a marker that appears later, and the name denotes everything from that sound up to but not including that marker.

The two examples Vasu gives

NameFirst soundClosing markerWhat it denotes
aca, the first sound of aphorism 1C, which closes aphorism 4every sound from a to au, that is, all the vowels
halha, the first sound of aphorism 5L, which closes aphorism 14every sound from ha to the end, that is, all the consonants

Two characters each. "All the vowels" and "all the consonants", the two commonest conditions in any phonological rule, cost one syllable apiece.

And Vasu states the total: though numerous pratyāhāras could be formed, practically there are only 42 in use.

Worked example: building and decoding a name

Building. You want to say "any vowel other than a, i or u". Those three are aphorism 1, closed by the marker N. So the sounds you want begin after that marker, at r in aphorism 2, and run to the end of the vowels, which is closed by C at the end of aphorism 4. The name is therefore the first sound you want plus that marker: r plus C.

Decoding. Given a name, find its first character in the sequence and its second character among the markers; the name is everything between them. Given ik, the first sound is i in aphorism 1, and the marker K closes aphorism 2. So ik denotes i, u, r and l, which is exactly the set Vasu says sūtra 1.1.3 restricts guṇa and vṛddhi to: "the ik vowels only, i, u, ri, and li long and short".

That is the payoff. [Meta-Rules: Paribhāṣā, and Rules About Rules] is about a rule whose whole content is the two-character name ik, and now you can read it.

Why this is a compression scheme with a purpose

Three properties are worth naming, because each is a property a computer scientist would look for.

The name carries its own definition. No lookup table. Decoding requires only the ordered sequence, which is memorised once.

The ORDER is the design. Any set that is not a contiguous span cannot be named this way, so the order has to be chosen so that the needed sets are spans. That is the hard part, and it is what makes the fourteen aphorisms an achievement rather than a list.

The cost is paid once and the saving is paid out thousands of times. Fourteen aphorisms to memorise, against two characters instead of a list in each of thousands of rules.

munotes.in113

The Fourteen Śivasūtras and the Pratyāhāra

The honest limits

Not every set is a span. Sets the grammar needs but cannot express as a span are handled by other means, and the tradition discusses them.

A marker used twice creates an ambiguity. Vasu notes exactly this: the same letter serves as marker in two of the aphorisms, so a name using it could denote two different spans, and the tradition has to say which.

The scheme says nothing about phonetics. The order is chosen for the grammar's convenience, not to reflect how sounds are made, and where the two pull apart the grammar wins.

And the transliteration in this book is not the sounds. The fourteen are printed above in plain letters so that a reader without Devanagari can follow the mechanism. The sounds themselves are Sanskrit phonemes and the printed forms are approximations; nothing in this paper depends on the phonetic detail.

Quick revision

  • Fourteen aphorisms, called the Śivasūtras or Pratyāhāra Sūtras, order the sounds of Sanskrit. Vasu states the number.
  • Each ends in a marker consonant that is not one of the sounds.
  • Sūtra 1.1.71 ādir antyena sahetā: a name is the first sound plus a later marker, and denotes everything between them.
  • Vasu's examples: ac is all the vowels, hal is all the consonants. He states that 42 pratyāhāras are in practical use.
  • ik denotes i, u, r and l, which is the set sūtra 1.1.3 restricts guṇa and vṛddhi to.
  • The design is the ORDER: a set can be named only if it is a contiguous span.
  • Limits: not every set is a span; a marker reused creates an ambiguity Vasu notes; the order serves the grammar, not phonetics.

Test yourself

1. What is a pratyāhāra, and which sūtra states how one is formed?

A two-character abbreviation denoting a contiguous span of the ordered list of sounds, formed from the first sound of the span plus the marker consonant that closes it. Sūtra 1.1.71, ādir antyena sahetā.

2. Decode ac and hal, and say why they are the two most useful names in the grammar.

Ac runs from a to the marker C, so it is all the vowels. Hal runs from ha to the marker L, so it is all the consonants. Conditions of the form "before a vowel" and "after a consonant" are the commonest in any phonological rule, and each now costs two characters.

3. Why is the ORDER of the fourteen aphorisms the real achievement?

Because a set can be named only if its members are contiguous in the order. Choosing an order in which every set the grammar needs happens to be a span is the difficult part; naming a span is trivial once it is.

4. Give one limitation of the scheme that the tradition itself notes.

munotes.in114

The Fourteen Śivasūtras and the Pratyāhāra

The same marker letter serves in two different aphorisms, so a name formed with it could denote either of two spans, and the tradition has to state which is meant.

Contents This chapter on its own page

munotes.in115

Chapter Thirty-Five

Anubandha: The Marker Letter as a Type Tag

Syllabus topic Module 1, "Examination of the structural architecture of Pāṇini’s grammatical system", "Context-sensitive operations"

In one line

An anubandha is a letter attached to a form that is not part of it, is deleted before the form is used, and while it is there controls which rules may apply.

In the wording you can write in an examination: an anubandha, also called an it, is an indicatory letter added to a root, affix or substitute in the grammar's own lists. It is not pronounced as part of the resulting word and is elided by rule, but its presence determines the applicability of other rules. It therefore functions as a classificatory marker rather than as a phonetic element.

The problem being solved

Two affixes may look identical and behave differently. Two roots spelled the same may take different endings. A grammar that can only see the letters of a form cannot distinguish them.

The obvious fix is to write a rule for each exception, and the grammar would then be enormous. Pāṇini's fix is to annotate the data instead: mark the form, and let one rule refer to the mark.

The mechanism, in three steps

Step one, in the lists. A root in the Dhātupāṭha or an affix in the grammar carries an extra letter. Vasu's translation records this constantly: of one root he notes that it is marked in the Dhātupāṭha with a particular letter, and that therefore a named rule applies to it.

Step two, in the rules. Rules refer to the marker. "An affix with an indicatory X does not cause guṇa." The marker is the condition.

Step three, in the output. The marker is deleted, so it never reaches the word. It existed only to carry information between the list and the rules.

Vasu's own example, which is the clearest one available

His note on the effect of an indicatory marker, in his words:

the affix which otherwise would have caused guṇa or vṛddhi, does not do so, when it has an indicatory ... Thus the past participle terminations ... are ārdhadhātuka affixes, which would, by the general rule VII. 3. 84, have caused guṇa, but as their indicatory letter ... is ..., the real terminations being ... , they do not cause guṇa. Therefore, when these terminations are added to a root, the ik of the root is not gunated.

Read the structure of that sentence, because the structure is the point.

A general rule exists. 7.3.84: a sārvadhātuka or ārdhadhātuka affix causes guṇa of the base.

A particular affix satisfies the rule's stated condition. It is an ārdhadhātuka affix.

And yet the rule does not fire. Because the affix carries a marker.

And the marker is not part of the affix. Vasu says "the real terminations being" and then gives the forms without the marker.

munotes.in116

Anubandha: The Marker Letter as a Type Tag

So the marker is invisible in the output and decisive in the derivation. That is a type annotation. The affix has a type, the rule is guarded on the type, and the type is erased before anything is produced.

The rule that reads the marker

Sūtra 1.1.5, whose number is confirmed in both of this book's witnesses. Its effect, as Vasu states it, is the one quoted above: a marked affix causes neither guṇa nor vṛddhi.

Two things about that rule are worth noticing.

It is a prohibition, not an operation. It changes nothing; it prevents something. In the six-kind classification of [The Six Kinds of Sūtra, in Vasu's Own Words] that is a niṣedha, which the tradition treats as a species of restrictive rule.

It is stated once for all marked affixes. Not once per affix. That is the saving, and it is the reason the annotation exists.

The comparison, stated as a table

AnubandhaA type annotation in a programming language
Attached toa root, affix or substitute in the grammar's listsa declaration in source code
Purposeto decide which rules may applyto decide which operations are legal
Present in the outputno, it is elidedno, most languages erase types before running
Stated wherein the data, once per formin the source, once per declaration
Read bythe rules that name itthe type checker
Failure modea form marked wrongly derives the wrong worda value typed wrongly is used illegally

The row that matters most is the third. Both are erased. An annotation that survived into the output would be part of the thing annotated, and neither is.

Worked example, worked twice

Take a root ending in a light vowel, and two affixes that are identical in their letters but differ in their markers.

Affix without a marker. 7.3.84's condition is met. 1.1.3 says the operation applies to the root's ik vowel. The vowel is replaced by its guṇa. The form changes.

Affix with a marker. 7.3.84's condition is met just as before. 1.1.5 prohibits the operation. The vowel is untouched. The form does not change.

Same rule, same root, same visible affix, two results. The only difference is a letter that never appears. [Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine] implements exactly this pair, and its test cases include both.

What an anubandha is NOT

It is not a sound. It is not pronounced and it is not part of the phonetics.

It is not arbitrary. Vasu's translation records which markers carry which effects, and the tradition documents the system. A marker is a name in a fixed vocabulary.

It is not the same as an affix's meaning. Two affixes with the same meaning may carry different markers because they behave differently under other rules. The marker is about rule applicability, not about semantics.

munotes.in117

Anubandha: The Marker Letter as a Type Tag

It is not checked. Nothing verifies that a root's markers are right. If a list is wrong the grammar derives wrong forms, and the only way to notice is that the output is wrong. That is the honest difference from a type system, which rejects the program.

Quick revision

  • An anubandha or it is an indicatory letter on a root, affix or substitute, deleted before the form is used, which controls which rules apply.
  • Vasu's example: an affix that satisfies 7.3.84's condition still does not cause guṇa, because of its marker, and he gives "the real terminations" without it.
  • The rule that reads the marker is 1.1.5, a prohibition rather than an operation, stated once for all marked affixes.
  • It is a type annotation: attached to data, read by rules, erased before output.
  • Unlike a type system, nothing checks it. A wrong marker gives a wrong word and no error.

Test yourself

1. Define anubandha and give the three stages of its life.

An indicatory letter attached to a form in the grammar's lists. It is placed on the form in the list, read as a condition by rules, and deleted before the form appears in a word.

2. Reconstruct Vasu's example of a marker blocking a rule.

A past participle termination is an ārdhadhātuka affix, so the general rule 7.3.84 would cause guṇa of the base. Because the termination carries an indicatory letter, 1.1.5 prohibits the operation, and the root's ik vowel is not gunated. The real termination, Vasu says, is the form without the marker.

3. In what precise sense is an anubandha a type annotation?

It is attached to data rather than to an operation, it is consulted to decide which operations are permitted, and it is erased before anything is produced.

4. Give the one important respect in which it is weaker than a type system.

Nothing checks it. A wrongly marked root produces a wrong word silently, whereas a type checker rejects an ill-typed program before it runs.

Contents This chapter on its own page

munotes.in118

Chapter Thirty-Six

Rule-Based Generative Structure

Syllabus topic Module 1, "Rule-based generative structure"

In one line

A word is not looked up, it is built: a root, an affix, and a chain of rules that fire in order and each leave the form a little different.

In the wording you can write in an examination: a rule-based generative structure derives a surface form from an underlying representation by the ordered application of rules, each of which states a condition and an operation. In the Aṣṭādhyāyī the underlying representation is a root together with the affixes its meaning requires, the rules substitute, add or delete elements, and the derivation is complete when no further rule applies.

The shape of a derivation

Every derivation in the grammar has the same five stages, and knowing them is most of what is needed to read one.

One, the root. Taken from the Dhātupāṭha, carrying its markers. Sūtra 1.3.1, bhūvādayo dhātavaḥ, is the rule that declares the list to be the list of roots.

Two, the affix required by the meaning. A tense, a person, a number. These are supplied by rules in book three.

Three, the affixes those affixes require. A tense marker often brings with it a class sign that stands between the root and the ending. Sūtra 3.1.68, kartari śap, introduces one such sign.

Four, the operations. Substitutions, lengthenings, deletions, each fired by a rule whose condition the form now satisfies.

Five, the finished form, when nothing more applies.

Worked example: a present-tense verb form

The example is chosen because every rule in it is one this book has read in Vasu's translation and confirmed in the digital sūtrapāṭha.

The task. From the root meaning "to be", in the third person singular of the present, produce the form.

Stage one. The root, written bhu in this book's plain transliteration. Its final vowel is u, which is one of the ik vowels.

Stage two. The personal ending for third person singular, written ti.

Stage three. Sūtra 3.1.68, kartari śap, brings in the class sign for this conjugation, so the form is now root, sign, ending.

Stage four, the operation this book implements. Sūtra 7.3.84, in Vasu's translation: "when a sārvadhātuka or an ārdhadhātuka affix follows there is guṇa of the base." The ending is sārvadhātuka, so the condition is met. Sūtra 1.1.3 tells us what is operated on: the ik vowels only. The root's u is an ik vowel, so it is replaced by its guṇa, which is o. The form becomes bho plus what follows.

Stage five, which this book cites and does not perform. Sūtra 6.1.78, ecoʼyavāyāvaḥ, provides that a vowel of the class that o belongs to becomes a semivowel sequence before a following vowel. That is what takes bho plus a to bhava, and the finished Sanskrit word is bhavati.

munotes.in119

Rule-Based Generative Structure

Say this plainly, because it matters. The engine in [Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine] carries the derivation as far as the guṇa substitution and stops. The remaining step is a real rule in the grammar and the book cites it by number; it is not implemented, and calling the engine's output a Sanskrit word would be false. What the engine demonstrates is rule-based derivation with conditions, meta-rules and precedence, which is what the syllabus asks for.

The same derivation as a table

StageRuleCondition satisfiedForm after
root1.3.1it is in the Dhātupāṭhabhu
endingbook threethird person singular presentbhu + ti
class sign3.1.68this conjugation, activebhu + a + ti
guṇa7.3.84 with 1.1.3a sārvadhātuka affix follows a stem ending in an ik vowelbho + a + ti
semivowel6.1.78, cited not implementeda vowel of that class before a vowelbhavati

A second derivation, in which a rule does not fire

The value of a rule system is visible when a rule declines.

The task. The same root with a past participle termination.

Stage one and two. The root bhu, and a termination that is ārdhadhātuka.

Stage three, the rule that would fire. 7.3.84's condition is satisfied: an ārdhadhātuka affix follows.

Stage four, the rule that stops it. The termination carries an indicatory marker, and by 1.1.5 a marked affix causes neither guṇa nor vṛddhi. So the root's vowel is untouched.

The form. The root's vowel stays as it is, and the derivation continues with other rules.

Two derivations, the same general rule, opposite outcomes, and the deciding factor is a letter that never appears in either word. That is what [Anubandha: The Marker Letter as a Type Tag] is about, and this is the derivation it produces.

What makes this generative rather than descriptive

Three properties, and all three are testable claims about the grammar.

The output is produced, not selected. No list of verb forms exists anywhere in the system. The forms are what the rules make.

The rules are general. 7.3.84 is one rule and it applies to every root ending in an ik vowel before every affix of the two named kinds. A descriptive grammar would have a paradigm per root.

The exceptions are rules too. 1.1.5 is not a footnote; it is a rule of the same kind, in the same numbering, and it interacts with the others by the same precedence principles. [Rule Precedence: The Four Principles] is about those principles.

What a derivation is NOT

It is not a sequence you may reorder. Which rule fires when is decided by the precedence principles, not by the writer's convenience. Reordering can produce a different and wrong form.

munotes.in120

Rule-Based Generative Structure

It is not a proof that the form is attested. The tradition's criterion is derivability, as [Pāṇini's Aṣṭādhyāyī, and the Text We Are Reading] says. A derivation shows the grammar produces the form; whether anybody says it is a separate question.

It is not unique. Two different chains of rules can reach the same form, and the tradition argues about which is correct. That is not a defect peculiar to Sanskrit: any rewrite system with overlapping rules has the same property, and [Rewrite Systems] treats it under the name confluence.

Quick revision

  • Five stages: root from the Dhātupāṭha, affix from the meaning, class sign, operations, finished form.
  • The worked derivation: bhu plus ti, class sign by 3.1.68, guṇa by 7.3.84 scoped by 1.1.3 giving bho, then 6.1.78 giving bhavati.
  • This book's engine performs the guṇa step and stops. 6.1.78 is cited and not implemented, and the book says so.
  • The same rule declines on a marked affix, by 1.1.5, so two derivations with the same general rule end differently.
  • Generative means: output produced not selected, rules general, exceptions stated as rules of the same kind.
  • A derivation is not reorderable at will, not evidence of usage, and not necessarily unique.

Test yourself

1. Give the five stages of a derivation in the Aṣṭādhyāyī.

The root from the Dhātupāṭha; the affix the meaning requires; any class sign those affixes bring; the operations that the form's condition now triggers; and the finished form when no rule applies.

2. Work the guṇa step for the root bhu with a third person singular present ending, citing the rules.

Sūtra 3.1.68 supplies the class sign. Sūtra 7.3.84 provides that a sārvadhātuka affix causes guṇa of the base, and 1.1.3 restricts the operation to the ik vowels, so the root's u becomes o, giving bho plus the following elements.

3. Which rule finishes the derivation, and why does this book not implement it?

Sūtra 6.1.78, which turns the vowel class o belongs to into a semivowel sequence before a vowel, giving bhavati. The book's engine stops at the guṇa step and cites 6.1.78 rather than implementing it, so that nothing claims to produce a finished Sanskrit word that it does not produce.

4. Show how the same general rule can give two different outcomes.

With an unmarked ārdhadhātuka affix, 7.3.84 fires and the root's ik vowel takes guṇa. With an affix carrying an indicatory marker, the condition of 7.3.84 is still met but 1.1.5 prohibits the operation, so the vowel is untouched.

Contents This chapter on its own page

munotes.in121

Chapter Thirty-Seven

Meta-Rules: Paribhāṣā, and Rules About Rules

Syllabus topic Module 1, "Meta-rules and rule precedence"

In one line

A paribhāṣā does not change a word; it changes how the other rules are to be read.

In the wording you can write in an examination: a paribhāṣā is an interpretive or meta-rule of the Aṣṭādhyāyī. It states a convention governing the reading and application of other rules, such as what an unspecified argument in a rule is to be taken as, how a case ending in a rule is to be construed, or which of two applicable rules prevails. It performs no operation on a form itself.

The distinction, stated once

A rule about forms says: in this condition, replace this with that.

A rule about rules says: when a rule of the first kind is worded in this way, read it as meaning that.

The second kind is what makes a compressed rule system possible, because it lets every rule of the first kind omit whatever the meta-rule can supply.

The example, quoted

Sūtra 1.1.3, in Vasu's translation:

In the absence of any special rule, whenever guṇa or vṛddhi is enjoined about any expression by using the terms guṇa or vṛddhi, it is to be understood to come in the room of the ik vowels only (i, u, ri, and li long and short) of that expression.

And his own annotation on it:

This is a paribhāṣā sūtra, and is useful in determining the original letters, in the place of which the substitute guṇa and vṛddhi letters will come. The present rule will apply where there is the specification of no other particular rule.

Two features of that are worth naming. It is labelled, by the tradition and by Vasu, so its kind is not something we have inferred. And it has a defeasibility clause: "in the absence of any special rule", which means a particular rule can override it. That is a precedence statement inside a meta-rule.

The meta-rule at work, in Vasu's own words

He then shows what 1.1.3 does to another rule, and this passage is the clearest demonstration of a meta-rule in the whole translation:

Thus sūtra VII. 3. 84 declares: "when a sārvadhātuka or an ārdhadhātuka affix follows there is guṇa of the base." Here the sthāni or the original expression which is to be gunated, is not specified, and to complete the sense, the word "ikaḥ" must be read into the sūtra. The rule then being, "when a S. or an A. affix follows there is guṇa of the ik vowels of the base."

Read what has happened.

A rule was written with a missing argument. 7.3.84 says guṇa happens and does not say of what.

The missing argument is supplied from elsewhere. 1.1.3 says: of the ik vowels.

munotes.in122

Meta-Rules: Paribhāṣā, and Rules About Rules

The result is a complete rule. And the completion is not a guess; it is the effect of a rule whose whole job is to complete such rules.

So 7.3.84 is not an incomplete rule. It is a complete rule read together with its meta-rule, and the grammar is designed to be read that way.

The modern counterparts

Paribhāṣā mechanismWhere a programmer meets it
supplying an argument a rule omitsa default parameter value, or an implicit receiver
saying how a rule's wording is to be construedthe semantics section of a language definition
stating which of two rules prevailsan operator precedence table, or a resolution order
defeasible by a special rulea default that a more specific declaration overrides

The last row is the one worth dwelling on, because it is the most sophisticated. 1.1.3's own wording makes it yield to a special rule. A default that knows it can be overridden is a designed default, not an accident.

Three paribhāṣā this book uses

Each is cited elsewhere in the book and collected here so the kind is visible as a kind.

1.1.3, iko guṇavṛddhī. Supplies the missing argument to every rule enjoining guṇa or vṛddhi. Used in [Rule-Based Generative Structure] and implemented in [Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine].

1.4.2, vipratiṣedhe paraṃ kāryam. When two rules of equal standing conflict, the later prevails. [Vipratiṣedha: When Two Rules Collide, the Later Wins].

8.2.1, pūrvatrāsiddham. Everything in the last three quarters of book eight is treated as not having taken effect, so far as the earlier rules are concerned. [Asiddhatva and the Tripādī: Ordering by Blocking].

Notice what the three have in common. None of them mentions a sound, a root or an affix. They are entirely about rules, and the grammar would not work without them.

Worked example: the effect of removing a meta-rule

This is the best way to see what a meta-rule is doing: take it away.

With 1.1.3 in force. 7.3.84 gunates the ik vowel of the base. A base ending in a is untouched, because a is not an ik vowel.

Without 1.1.3. 7.3.84 says guṇa of the base and nothing restricts what. The rule would have to be read as gunating something unspecified, which is either meaningless or would have to be filled by guesswork, and the tradition's whole objection to that is why the paribhāṣā exists.

What would have to change if the meta-rule did not exist. Every rule enjoining guṇa or vṛddhi would have to name the ik vowels itself. Vasu notes that guṇa is enjoined in rules like 7.3.82 as well as 7.3.84, and there are many more. One meta-rule against a phrase repeated in dozens of rules is the trade, and it is the same trade a programmer makes with a default.

munotes.in123

Meta-Rules: Paribhāṣā, and Rules About Rules

What a paribhāṣā is NOT

It is not a comment. It is a rule, numbered in the same sequence, and its effects are part of the grammar's output.

It is not always in book one. Paribhāṣā are scattered, and 8.2.1 is at the head of the last three quarters. Their position matters because some of them have scope limited by position.

It is not beyond dispute. Which rules are paribhāṣā, and exactly how far each reaches, is argued in the commentary literature. Vasu labels 1.1.3 and the tradition agrees about that one; others are contested.

Quick revision

  • A paribhāṣā is a rule about rules. It performs no operation on a form.
  • 1.1.3 is labelled a paribhāṣā by Vasu himself, and it supplies the missing word ikaḥ to every rule enjoining guṇa or vṛddhi.
  • Vasu shows it completing 7.3.84: "to complete the sense, the word ikaḥ must be read into the sūtra."
  • Its wording is defeasible: "in the absence of any special rule", which is a precedence statement inside a meta-rule.
  • Three used in this book: 1.1.3 supplying an argument, 1.4.2 resolving a conflict, 8.2.1 imposing an order.
  • None of the three mentions a sound, a root or an affix.
  • Modern counterparts: default arguments, the semantics section of a language definition, a precedence table, an overridable default.

Test yourself

1. Define paribhāṣā and say how it differs from a vidhi rule.

A paribhāṣā is an interpretive rule governing how other rules are read or applied. A vidhi states an operation on a form. The paribhāṣā changes nothing in a word; it changes what the vidhi rules mean.

2. Quote or paraphrase Vasu's demonstration of 1.1.3 completing 7.3.84.

He says that 7.3.84 declares guṇa of the base when a sārvadhātuka or ārdhadhātuka affix follows, that the expression to be gunated is not specified, and that to complete the sense the word ikaḥ must be read into the sūtra, giving guṇa of the ik vowels of the base.

3. What does "in the absence of any special rule" in 1.1.3 tell you about the grammar's design?

That the default is deliberately overridable: a meta-rule states its own defeasibility, so a special rule can displace it without the grammar becoming inconsistent.

4. Name three paribhāṣā and what each governs.

1.1.3, which supplies the unstated argument to rules enjoining guṇa or vṛddhi; 1.4.2, which makes the later of two conflicting rules prevail; and 8.2.1, which makes the last three quarters of book eight non-existent for the rules before them.

Contents This chapter on its own page

munotes.in124

Chapter Thirty-Eight

Rule Precedence: The Four Principles

Syllabus topic Module 1, "Meta-rules and rule precedence", "Conflict resolution mechanisms"

In one line

When two rules could both apply, four principles decide which one does, and they are tried in a fixed order.

In the wording you can write in an examination: where two rules of the Aṣṭādhyāyī are simultaneously applicable to the same form, the tradition resolves the conflict by four principles applied in order: apavāda, a special rule sets aside a general one; nitya, a rule that would apply in either event is applied first; antaraṅga, a rule whose conditions lie closer to the stem prevails over one whose conditions lie further out; and finally paratva, stated in sūtra 1.4.2, by which the later rule in the order of the text prevails.

Why a rule system needs this at all

Any rule set large enough to be useful has rules whose conditions overlap. Two rules apply, they prescribe different things, and the system must produce one answer.

There are only three ways to deal with it. Write the rules so they never overlap, which is impossible past a certain size. Leave it undefined, which means the system has no single output. Or state a policy, which is what Pāṇini's tradition does.

The tradition's policy is four principles in a fixed order, and knowing the ORDER is the examinable part.

The four, in order

One: apavāda, the special sets aside the general

If one rule's conditions are a subset of the other's, the narrower rule wins.

Why it must come first. A special rule exists only in order to displace the general one. If the general rule won, the special rule could never fire and would be pointless.

Instance. 7.3.84 causes guṇa before any sārvadhātuka or ārdhadhātuka affix. 1.1.5 prohibits it for affixes carrying an indicatory marker, which are a subset. The prohibition wins, and the general rule is simply not in play for those affixes.

Modern counterpart. The most specific matching declaration wins: an overload resolved to the narrower signature, a stylesheet rule with a more specific selector, a pattern match whose first matching case is the tighter one.

Two: nitya, the unconditional goes first

A rule is nitya, invariable, with respect to another if it would still apply after the other had operated, while the other would not still apply after it had.

Why it comes second. A rule that will get its chance either way can afford to wait; a rule that will lose its chance cannot. Applying the nitya rule first costs nothing and applying it second costs it its opportunity.

Modern counterpart. Doing the irreversible step last. A build that runs the check before the overwrite, rather than after, on the same reasoning.

Three: antaraṅga, the inner before the outer

A rule whose triggering conditions lie inside the form, closer to the root, prevails over one whose conditions involve material further out.

munotes.in125

Rule Precedence: The Four Principles

The maxim, in Vasu's own words, and it appears in two of his volumes in the same wording:

That which is bahiraṅga is regarded as not having taken effect, or as not existing, when that which is antaraṅga is to take effect.

Notice the form of the statement. It does not say the inner rule is applied first. It says the outer rule is treated as not having happened while the inner rule is under consideration. That is stronger and it is more useful, because it makes the inner rule's condition evaluate against the form as it was.

Modern counterpart. Evaluating the inner expression before the outer, and more precisely, a scoping rule: what is visible to the inner rule is the state before the outer one ran.

Four: paratva, the later rule wins

If none of the three above separates them, the rule that stands later in the text prevails. This is sūtra 1.4.2, vipratiṣedhe paraṃ kāryam, and it is the subject of [Vipratiṣedha: When Two Rules Collide, the Later Wins].

Why it comes last. It is arbitrary. Position in the text is not a reason in the way specificity and scope are; it is a tie-break, and a tie-break is for when reasons run out.

The order, as a table

OrderPrincipleTestWhy here
1apavādais one rule's condition a subset of the other's?a special rule exists only to displace the general one
2nityawould one rule still apply after the other, but not conversely?the rule that will lose its chance must take it
3antaraṅgaare one rule's conditions closer to the stem?the outer rule is treated as not having happened
4paratva, 1.4.2which rule stands later in the text?arbitrary, so it is the tie-break of last resort

Worked example: the four tested in turn

Take two applicable rules, A and B, and walk the principles.

Is either special? If A's condition is "before any affix of kind K" and B's is "before affixes of kind K that carry marker M", then B is special and B wins. Stop.

If neither is special, is either nitya? If A would still apply after B had operated, but B would not still apply after A, then A can wait and B cannot, so B goes first. Stop.

If neither, is either inner? If A's condition involves only the root and B's involves the ending as well, A is antaraṅga and B is treated as not having taken effect. A wins. Stop.

If neither, compare positions. By 1.4.2 the later rule prevails.

munotes.in126

Rule Precedence: The Four Principles

The program in [Śāstra Rule Precedence as a Deterministic Finite Rewrite System] implements exactly this cascade and reports which principle decided, and its test cases include one for each of the four.

What the four principles do NOT give you

They do not always agree with each other. Two principles can point different ways, which is precisely why the order exists. The order is the tradition's answer, and the commentaries argue about particular cases.

They do not make the system provably consistent. Nothing shows that every possible conflict is resolved, or that the resolutions never produce a wrong form. The tradition's evidence is that the grammar derives the right forms, which is testing rather than proof.

They are not all in the text. Only paratva is a sūtra. Apavāda, nitya and antaraṅga are maxims of the tradition, transmitted as paribhāṣā in separate collections, and Vasu quotes the antaraṅga one as a maxim rather than as a numbered rule. An answer should say which are sūtras and which are maxims.

Quick revision

  • Four principles, in order: apavāda, nitya, antaraṅga, paratva.
  • Apavāda first because a special rule exists only to displace the general one.
  • Nitya second because a rule that would still apply later can wait.
  • Antaraṅga third, and its maxim treats the outer rule as not having taken effect, not merely as later.
  • Paratva last, sūtra 1.4.2, because position in the text is arbitrary and a tie-break is for when reasons run out.
  • Only paratva is a numbered sūtra; the other three are maxims of the tradition.

Test yourself

1. Name the four precedence principles in order and give the test for each.

Apavāda: is one rule's condition a subset of the other's? Nitya: would one still apply after the other, but not conversely? Antaraṅga: are one rule's conditions closer to the stem? Paratva: which rule stands later in the text?

2. Why is 1.4.2 applied last rather than first?

Because position in the text is arbitrary. Specificity, unconditional applicability and scope are reasons; position is only a tie-break, so it is used when the reasons run out.

3. Quote the antaraṅga maxim and say what is unusual about its wording.

That which is bahiraṅga is regarded as not having taken effect, or as not existing, when that which is antaraṅga is to take effect. It does not merely order the two rules: it makes the outer rule count as not having happened, so the inner rule's condition is evaluated against the earlier state.

4. Which of the four is a numbered sūtra, and why does the distinction matter?

Only paratva, sūtra 1.4.2. The distinction matters because the other three are maxims transmitted by the tradition rather than rules of the text, so their authority and their exact scope are matters the commentaries argue about.

Contents This chapter on its own page

munotes.in127

Chapter Thirty-Nine

Vipratiṣedha: When Two Rules Collide, the Later Wins

Syllabus topic Module 1, "Conflict resolution mechanisms", "Meta-rules and rule precedence"

In one line

When two rules of equal standing want to do different things to the same form, the one printed later in the grammar wins.

In the wording you can write in an examination: sūtra 1.4.2, vipratiṣedhe paraṃ kāryam, provides that in the event of a conflict between two rules the operation prescribed by the later rule is to be performed. It is the last of the four precedence principles, applied only when the conflicting rules are not separated by speciality, unconditional applicability or scope, and "later" means later in the text's own order, that is, by book, quarter and position.

The provision

Sūtra 1.4.2, in the digital sūtrapāṭha: vipratiṣedhe paraṃ kāryam. In a conflict, the later operation is to be done.

Vasu's volumes refer to it constantly rather than translating it once, and the recurring phrase is the gloss: a form "by the rule of vipratiṣedha takes precedence". The number is confirmed in both of this book's witnesses.

What "conflict" means here

Not any overlap. Two rules that both apply and prescribe compatible things are not in conflict; both happen. A conflict in the sense of 1.4.2 has three features.

Both rules apply to the same form at the same point.

They prescribe different results. Either two different substitutions for the same element, or one prescribing an operation the other forbids.

Nothing else separates them. If one is special, or nitya, or antaraṅga, the earlier principles of [Rule Precedence: The Four Principles] have already decided and 1.4.2 is not reached.

That third condition is where students go wrong, and it is worth stating as a rule of thumb: reach for 1.4.2 only after the other three have failed.

What "later" means

The grammar's rules have addresses: book, quarter, position. Later means later in that order.

ComparisonWhich is laterWhy
6.1.77 against 6.1.786.1.78same book, same quarter, higher position
6.1.78 against 6.4.226.4.22same book, later quarter
6.4.22 against 7.3.847.3.84later book
1.4.2 against 8.2.18.2.1later book, and it is the rule that opens the tripādī

The last row is a warning. The tripādī, the last three quarters of book eight, is not governed by 1.4.2 in the ordinary way, because 8.2.1 imposes something stronger. [Asiddhatva and the Tripādī: Ordering by Blocking] is about that.

Why a positional tie-break is a real design decision

This is the part that belongs in a five-mark answer, and it has two sides.

What it buys

Totality. Every conflict has an answer. There is no pair of rules for which the grammar is silent, because any two rules have an order in the text.

Cheapness. Deciding takes a comparison of two addresses. No reasoning about the rules' content is needed, which matters when a human reader has to apply it in real time.

munotes.in128

Vipratiṣedha: When Two Rules Collide, the Later Wins

A tool for the author. Because position decides, the author can express a preference by placing a rule later. That converts an editorial act into a semantic one, and Pāṇini uses it: a rule placed later is a rule intended to win.

What it costs

It makes position semantic. You can no longer move a rule for the sake of tidiness. Every rule's position is part of its meaning, and reorganising the grammar would change what it derives.

It makes insertion dangerous. Adding a rule changes its relation to every rule after it. The commentaries' anxiety about interpolated rules is exactly this.

It hides the reason. When a special rule wins, you know why: it is special. When the later rule wins, the only reason available is "it is later", which tells a reader nothing about the grammar's intent.

Modern parallel. A stylesheet in which later declarations win, or a configuration system in which the last file read overrides the earlier ones. Both are total and cheap, both make file order semantic, and both are notorious for the same two problems: you cannot reorder safely and you cannot tell from a rule why it lost.

Worked example: two rules, resolved twice

The pair. Two sandhi rules in the same quarter of book six, one at position 77 and one at position 78, which prescribe different treatments of the same vowel before a following vowel.

Applying 1.4.2 directly. 6.1.78 is later, so 6.1.78 prevails.

But check the earlier principles first. Is either special? If one applies to vowels of a class that is a subset of the other's, it is special, and it wins whichever position it holds. Vasu's discussion of exactly this neighbourhood turns on such considerations, and the honest report is that the pair's resolution is argued in the commentary rather than settled by position alone.

The lesson. 1.4.2 gives an answer whenever it is reached. The work is in knowing whether it has been reached, and a student who applies it first will be confidently wrong.

What 1.4.2 is NOT

It is not a general principle that later beats earlier. It applies to conflicts, and only to conflicts the other three principles do not resolve.

It is not about which rule was composed later. It is about position in the text as transmitted.

It does not govern the tripādī. There, 8.2.1's asiddhatva applies, and within the tripādī a later rule is treated as not having taken effect for an earlier one, which is the reverse of 1.4.2's effect.

It is not a claim about the grammar's consistency. It resolves conflicts; it does not show that resolving them always gives the right form.

munotes.in129

Vipratiṣedha: When Two Rules Collide, the Later Wins

Quick revision

  • 1.4.2 vipratiṣedhe paraṃ kāryam: in a conflict, the later rule's operation is performed.
  • A conflict means both apply, they prescribe different results, and none of the other three principles separates them.
  • Later means later by book, quarter, position.
  • It buys totality, cheapness, and a way for the author to express a preference by placement.
  • It costs: position becomes semantic, insertion becomes dangerous, and the reason a rule lost is uninformative.
  • Modern parallel: last declaration wins, as in a stylesheet or a layered configuration.
  • It does not govern the tripādī, where 8.2.1's asiddhatva reverses the effect.

Test yourself

1. State 1.4.2 and the three conditions under which it is reached.

In a conflict, the operation of the later rule is performed. It is reached when both rules apply to the same form, they prescribe different results, and neither speciality, nor nitya status, nor scope separates them.

2. Give one advantage and one disadvantage of a positional tie-break.

Advantage: it is total and cheap, since any two rules have an order, so every conflict has an answer reached by comparing two addresses. Disadvantage: position becomes part of a rule's meaning, so the grammar cannot be reorganised or extended safely.

3. Order these by 1.4.2: 6.4.22, 7.3.84, 6.1.78.

6.1.78, then 6.4.22, then 7.3.84. Later quarter beats earlier quarter within a book, and a later book beats an earlier one.

4. Why is applying 1.4.2 first a mistake?

Because three principles precede it, and each of them can point the other way. A special, nitya or antaraṅga rule wins regardless of its position, so a student who compares positions first will reach the wrong result whenever one of those applies.

Contents This chapter on its own page

munotes.in130

Chapter Forty

Context-Sensitive Operations

Syllabus topic Module 1, "Context-sensitive operations"

In one line

A rule that says "do this here" has to say what counts as here, and Pāṇini says it with a case ending.

In the wording you can write in an examination: a context-sensitive operation is one whose applicability depends on the material surrounding the element operated on, not only on that element. In the Aṣṭādhyāyī the context is stated by the case ending of the word naming it: by sūtra 1.1.66 a locative means the operation applies to what precedes the named item, and by sūtra 1.1.67 an ablative means it applies to what follows. In formal language theory a context-sensitive rule is one whose left-hand side is longer than a single symbol.

Why a rule needs a context at all

Consider an operation: a vowel is lengthened. When?

"Always" is wrong, because the language is full of short vowels. "Before an ending" is a context. "After a particular root" is a different context. "Between two consonants" is a third.

So nearly every interesting rule is a triple: what is changed, what it becomes, and where. The first two are the operation and the third is the context, and a grammar needs a notation for all three.

Pāṇini's notation: the case endings do the work

This is one of the most economical devices in the whole grammar, and it is stated as two meta-rules.

Sūtra 1.1.66, tasminniti nirdiṣṭe pūrvasya. When something is named in the locative case, the operation applies to what stands BEFORE it.

Sūtra 1.1.67, tasmād ity uttarasya. When something is named in the ablative case, the operation applies to what stands AFTER it.

Both numbers are confirmed in the digital sūtrapāṭha.

So the grammatical case of a word in a rule tells you where the operation happens. No separate notation for context is needed at all: the language's own morphology carries it.

The three ways a rule names its environment

Case usedWhat it meansWhere the operation lands
genitive"of X"X itself is what is replaced. This is the operand, not the context
locative"in the presence of X", by 1.1.66on what precedes X
ablative"after X", by 1.1.67on what follows X

Read those three together and the whole triple is expressible. The genitive names the thing changed, the locative and the ablative name the neighbourhood, and a rule that uses all three has stated an operation in context in a few words.

And the genitive is itself governed by a meta-rule. Sūtra 1.1.49, ṣaṣṭhī sthāneyogā, confirmed in the sūtrapāṭha, provides that a genitive in a rule is to be understood as "in the place of". So "of X" means "in the place of X", which is a substitution.

munotes.in131

Context-Sensitive Operations

Worked example: reading a rule's context

Take 7.3.84, whose translation Vasu gives as: when a sārvadhātuka or an ārdhadhātuka affix follows there is guṇa of the base.

The affix is named in the locative. So by 1.1.66 the operation applies to what precedes the affix, which is the base. That is where the words "of the base" in Vasu's English come from: they are the effect of the case ending plus the meta-rule.

Guṇa is the operation, and what it replaces is supplied by 1.1.3, as [Meta-Rules: Paribhāṣā, and Rules About Rules] shows.

So the rule's four parts are: operand, the base's ik vowel; result, its guṇa; context, a following affix of two named kinds; and the context's direction, given by a case ending read through a meta-rule.

Nothing in the rule says the word "before" or "after". The case ending says it. That is the economy.

The formal meaning of context-sensitive

The phrase has a precise meaning in formal language theory and it is not the same as the informal one. It is worth having both.

A context-free rule has exactly one symbol on its left-hand side. It says: wherever you see this one symbol, you may replace it, no matter what surrounds it.

A -> b C d

A context-sensitive rule has more than one symbol on its left-hand side, so the surroundings are part of what must match.

x A y -> x b y

Here A becomes b, but only between x and y, and the x and y are unchanged. The rule cannot be split into a context-free one without losing the condition.

The consequence, which is the important part. Context-sensitive grammars generate strictly more languages than context-free ones. So the question "is the Aṣṭādhyāyī context-free?" is a real question with real stakes, and [Context-Free Grammar, and Whether Pāṇini Wrote One] answers it.

Where the Aṣṭādhyāyī is context-free and where it is not

Some rules are context-free in the formal sense. A rule that introduces an affix because the meaning requires it names no environment; it just adds material.

Most operational rules are not. 7.3.84 requires a following affix of a particular kind, which is context. Written as a rewrite rule its left-hand side has to mention both the vowel and the affix, so it has more than one symbol.

And some rules are worse than context-sensitive in the formal sense. A rule whose condition refers to a marker that has already been deleted, or to whether another rule has fired, is not a rewrite rule over strings at all. The asiddhatva mechanism of [Asiddhatva and the Tripādī: Ordering by Blocking] is exactly that, and it is why the formal classification of the whole grammar is not a simple matter.

munotes.in132

Context-Sensitive Operations

The honest summary. The Aṣṭādhyāyī is a rule system with contexts, and the formal hierarchy is a tool for comparing it with modern grammars rather than a box it fits into. An answer that says so is better than one that picks a class.

What context-sensitivity is NOT

It is not the same as having exceptions. An exception is another rule. A context is a condition inside one rule.

It is not the same as ambiguity. A context makes a rule narrower, and narrower rules make a grammar more determinate, not less.

It is not expensive by itself. Checking a context costs a look at the neighbours. What is expensive is deciding which of several context-sensitive rules applies, which is why the precedence principles exist.

Quick revision

  • Nearly every rule is a triple: what is changed, what it becomes, and where.
  • Pāṇini states the "where" with a case ending. 1.1.66: a locative means the operation applies to what precedes. 1.1.67: an ablative means it applies to what follows.
  • 1.1.49 ṣaṣṭhī sthāneyogā: a genitive means "in the place of", so "of X" is a substitution for X.
  • Formally, a context-free rule has one symbol on its left; a context-sensitive rule has more, so the surroundings must match and are preserved.
  • Context-sensitive grammars generate strictly more languages than context-free ones.
  • The Aṣṭādhyāyī has context-free rules, many context-sensitive ones, and mechanisms that are not string rewriting at all.

Test yourself

1. How does a rule of the Aṣṭādhyāyī say where an operation applies?

By the case ending of the word naming the neighbouring item. By 1.1.66 a locative means the operation applies to what precedes it; by 1.1.67 an ablative means it applies to what follows.

2. What does a genitive in a rule mean, and which sūtra says so?

It means "in the place of", so the item named in the genitive is what is replaced. Sūtra 1.1.49, ṣaṣṭhī sthāneyogā.

3. Give the formal definition of a context-sensitive rule and one example.

A rule whose left-hand side has more than one symbol, so that the surrounding material must match and is preserved. For example x A y rewrites to x b y, which replaces A by b only between x and y.

4. Why is "is the Aṣṭādhyāyī context-free?" a question with stakes?

Because context-sensitive grammars generate strictly more languages than context-free ones, so the answer says something about what kind of system the grammar is, and about whether a context-free formalism could express it.

Contents This chapter on its own page

munotes.in133

Chapter Forty-One

Asiddhatva and the Tripādī: Ordering by Blocking

Syllabus topic Module 1, "Context-sensitive operations", "Conflict resolution mechanisms"

In one line

The last three quarters of book eight are treated as not having happened, as far as every rule before them is concerned, and that turns a rule list into two ordered phases.

In the wording you can write in an examination: sūtra 8.2.1, pūrvatrāsiddham, provides that the rules of the last three quarters of book eight, known as the tripādī, are asiddha, that is, treated as not having taken effect, with respect to all the preceding rules. Within the tripādī itself a later rule is likewise asiddha with respect to an earlier one. The grammar therefore operates in two successive phases, and no rule in the earlier phase can see the results of the later one.

The provision, as Vasu states it

With regard to whatever has been taught in the preceding Seven Books and a quarter, the rules contained in these three last chapters are considered as asiddha. And further, in these three chapters, a subsequent rule is, as if it had not taken effect, so far as any preceding rule is concerned.

And elsewhere, naming the section:

the tripādī or the last three chapters of Aṣṭādhyāyī; and the tripādī are considered asiddha for the purposes of previous sūtras (VIII. 2. 1).

The word asiddha he glosses as not-accomplished: the operation caused by a rule's having taken effect is not produced, for the purposes of the rule now being considered.

What the rule actually does

There are two separate provisions in it, and students usually notice only the first.

Provision one: a boundary. The grammar is divided at the end of the first quarter of book eight. Everything after that point is invisible to everything before it. Call the earlier part phase one and the tripādī phase two.

Provision two: an internal ordering. Inside the tripādī, a later rule is also asiddha for an earlier one. So within phase two the rules are effectively applied in textual order, each blind to the ones after it.

The second provision is the reverse of what 1.4.2 does elsewhere. Under 1.4.2 a later rule prevails in a conflict; inside the tripādī an earlier rule proceeds as though the later one had not happened. That is not an inconsistency: it is a different mechanism for a different section, and knowing that the tripādī is governed differently is worth a mark.

Why a grammar would want this

Think about what goes wrong without it.

Suppose a phase-one rule's condition is "a stem ending in a consonant". Suppose a phase-two rule deletes a final consonant in certain words. Without 8.2.1, the phase-one rule would have to ask: has the deletion happened yet? And the answer would depend on the order in which the two rules were tried, which is exactly the kind of dependence that makes a rule system unpredictable.

munotes.in134

Asiddhatva and the Tripādī: Ordering by Blocking

8.2.1 removes the question. Phase-one rules always see the form before any phase-two rule has touched it. Their conditions are therefore stable, and a reader can reason about phase one without knowing anything about phase two.

That is a real engineering benefit and it has a modern name.

The modern name: phase ordering

Feature of 8.2.1The modern counterpart
a fixed boundary in the rule lista compiler pass boundary
earlier rules cannot see later resultsa pass operates on the output of the previous pass only
later rules may see earlier resultspasses run in order and each sees what the last produced
the boundary is declared, not inferredthe pass pipeline is written down
inside the later section, order is textuala single pass that traverses its input once

A multi-pass compiler makes exactly this choice. Constant folding does not need to know what the register allocator will do, because the allocator runs later and the folder is finished. The benefit is the same: each pass can be reasoned about in isolation.

Worked example, from Vasu's own discussion

He gives the mechanism at work, and the example is worth following because it shows why the rule is needed rather than merely stating it.

He is discussing a case in which a final consonant could be elided by either of two rules of the tripādī, at 8.2.7 and at 8.2.23. His report is that the form could not have had its final elided by 8.2.7 if the elision by 8.2.23 had taken effect, and that therefore, because within the last three quarters a subsequent rule is as if it had not taken effect so far as a preceding rule is concerned, 8.2.23 is treated as not having operated and 8.2.7 finds its scope.

Read what that does. The earlier rule of the tripādī gets to operate on the form as it stands, without having to know that a later rule of the tripādī would have changed the very thing its condition depends on. The blocking is what makes the earlier rule usable.

The one that is not 8.2.1: the ābhīya section

The grammar has a second, narrower asiddhatva, and a complete answer mentions it.

Sūtra 6.4.22, confirmed in the digital sūtrapāṭha, establishes a section within which, as Vasu puts it, when one of the rules of the section is to be brought into operation having the same place of operation as another of the section which has already taken effect, the one that has taken effect is regarded as not having taken effect.

So the same device is used twice, at two scales. Once for the whole tripādī, and once for a smaller neighbourhood within book six. That is a designed mechanism rather than a one-off patch, and pointing that out is the strongest form of the claim.

munotes.in135

Asiddhatva and the Tripādī: Ordering by Blocking

What asiddhatva is NOT

It is not deletion. The later rule does still operate. It is invisible to the earlier rules, not absent.

It is not the same as the antaraṅga maxim. That maxim also uses the words "not having taken effect", and Vasu quotes it that way, but it compares two rules by the closeness of their conditions, not by their section. Two mechanisms, similar wording, different tests.

It is not a general principle of the grammar. It is declared for two specific sections. Outside them, 1.4.2 and the other precedence principles apply.

It is not a proof of consistency. It makes the phases independent; it does not show that the resulting derivations are all correct.

Quick revision

  • 8.2.1 pūrvatrāsiddham: the last three quarters of book eight, the tripādī, are asiddha for every preceding rule, and inside the tripādī a later rule is asiddha for an earlier one.
  • Asiddha means treated as not having taken effect; Vasu glosses it as not-accomplished.
  • Effect: two phases, with phase one blind to phase two, so phase-one conditions are stable.
  • The modern counterpart is a compiler pass boundary, declared rather than inferred.
  • Inside the tripādī the internal rule reverses 1.4.2's effect, which is a different mechanism for a different section.
  • A second, narrower asiddhatva is established by 6.4.22 for a section of book six.
  • It is not deletion, not the antaraṅga maxim, and not a general principle.

Test yourself

1. State 8.2.1 and both of its provisions.

The rules of the last three quarters of book eight are treated as not having taken effect with respect to all preceding rules; and within those three quarters a later rule is treated as not having taken effect with respect to an earlier one.

2. Why does a grammar benefit from making one section invisible to another?

Because a rule's condition then does not depend on whether a rule from the other section has already run. The conditions of the earlier section are stable, so it can be reasoned about without knowing anything about the later one.

3. Give the modern counterpart of asiddhatva and one property they share.

A compiler pass boundary. Both make each pass or phase operate only on what the previous one produced, so each can be reasoned about in isolation, and in both the boundary is declared rather than discovered.

4. How does the tripādī's internal rule relate to 1.4.2, and why is that not a contradiction?

Under 1.4.2 the later of two conflicting rules prevails. Inside the tripādī an earlier rule proceeds as though the later one had not taken effect, which is the opposite effect. It is not a contradiction because 8.2.1 declares a different mechanism for that section, and 1.4.2 governs elsewhere.

Contents This chapter on its own page

munotes.in136

Chapter Forty-Two

Conflict Resolution Mechanisms, Collected

Syllabus topic Module 1, "Conflict resolution mechanisms"

In one line

The Aṣṭādhyāyī has six mechanisms for deciding what happens when two rules could both apply, and they are not all the same kind of thing.

In the wording you can write in an examination: conflict resolution in the Aṣṭādhyāyī operates through four precedence principles, apavāda, nitya, antaraṅga and paratva, applied in that order; through prohibition rules, niṣedha, which forbid an operation outright in stated cases; and through asiddhatva, which makes one section of rules invisible to another so that certain conflicts cannot arise at all.

The six, in one table

This is the table to revise from.

MechanismWhat it decidesStated byKind
apavādaa rule whose condition is narrower sets the general rule asidea maxim of the traditionprecedence
nityaa rule that would still apply after the other is applied firsta maxim of the traditionprecedence
antaraṅgaa rule whose conditions are closer to the stem prevails, and the outer rule is treated as not having taken effecta maxim, quoted by Vasuprecedence
paratvathe rule later in the text prevailssūtra 1.4.2, vipratiṣedhe paraṃ kāryamprecedence
niṣedhaan operation is forbidden in stated cases, so the conflict does not arisesūtras, for example 1.1.5prohibition
asiddhatvaone section is invisible to another, so its results cannot trigger or blocksūtras 8.2.1 and 6.4.22phase ordering

Read the "kind" column. Three different kinds of answer to one problem: decide between them, forbid one of them, or arrange that they never meet. Any engineer who has dealt with conflicting rules has used all three.

The order in which they are tried

The four precedence principles are tried in the order given, and the other two are not in that queue at all.

Niṣedha is checked first of everything, because a prohibition removes one of the two candidates. If 1.1.5 forbids the operation, there is no conflict left to resolve.

Asiddhatva is structural and applies before any of it. If the two rules are on opposite sides of the tripādī boundary, the earlier one simply cannot see the later one's result, so the situation never presents itself as a conflict.

Then the four principles, in order: apavāda, nitya, antaraṅga, paratva.

So the full decision procedure has six steps, and the program in [Śāstra Rule Precedence as a Deterministic Finite Rewrite System] implements it in that order.

The five-mark answer

MU's label is "conflict resolution mechanisms". Here is an answer at the length her paper allows, which is about a hundred and forty words.

Answer. The Aṣṭādhyāyī resolves conflicts in three ways. First, structurally: sūtra 8.2.1 makes the last three quarters of book eight asiddha, not having taken effect, for all earlier rules, so a rule in the earlier phase never sees a result from the later one and the conflict cannot arise. Second, by prohibition: rules such as 1.1.5 forbid an operation in stated cases, removing one candidate. Third, by precedence, in a fixed order of four principles: apavāda, a special rule sets aside a general one; nitya, a rule that would apply either way goes first; antaraṅga, a rule whose conditions lie closer to the stem prevails and the outer rule is treated as not having taken effect; and finally paratva, sūtra 1.4.2, under which the later rule prevails. Only paratva and the prohibitions are numbered sūtras; the first three principles are maxims of the tradition.

munotes.in137

Conflict Resolution Mechanisms, Collected

What each mechanism costs

An answer asked to assess should be able to say what the price of each is.

Apavāda requires knowing that one condition is a subset of the other, which is a judgement about content. Two rules can each look narrower in a different respect, and the commentaries argue about such cases.

Nitya requires reasoning about counterfactuals: would this rule still apply after the other had operated? That is harder than it sounds and the tradition disputes particular instances.

Antaraṅga requires a notion of "closer to the stem" that the grammar does not define formally. It is applied by trained judgement.

Paratva requires only a comparison of addresses, and costs nothing. Its price is paid elsewhere: position becomes semantic, as [Vipratiṣedha: When Two Rules Collide, the Later Wins] sets out.

Niṣedha costs nothing to apply and costs a rule to state. Each prohibition is another rule to transmit.

Asiddhatva costs nothing to apply and costs expressiveness: a phase-one rule cannot make use of a phase-two result even when it would be convenient.

The pattern is worth naming. The cheap mechanisms are the arbitrary ones, and the principled mechanisms are the expensive ones. That is true of modern conflict resolution too: a specificity rule is more informative and harder to compute than a "last one wins" rule.

Worked example: a conflict resolved at each level

The same pair of candidate operations, resolved four different ways depending on what is true of them.

If a prohibition applies. One operation is forbidden outright. One candidate remains. No conflict.

If the two rules are in different phases. The earlier rule operates on the form as it stands, blind to the later one. No conflict.

If one is special. The special rule operates and the general one does not apply here at all.

If none of that holds. Compare positions. The later rule operates, and the only reason available is its position.

What is NOT a conflict resolution mechanism

Two rules that do compatible things. If one lengthens a vowel and the other adds an affix, both happen. There is no conflict and no mechanism is needed.

munotes.in138

Conflict Resolution Mechanisms, Collected

An exception stated in the same rule. A rule with a clause excluding a case has narrowed its own condition. Nothing has to be resolved.

The commentaries' arguments. Where the tradition disputes which principle applies, that is a disagreement about the mechanisms, not a further mechanism.

Quick revision

  • Six mechanisms, three kinds: decide between the rules, forbid one of them, or arrange that they never meet.
  • Precedence, in order: apavāda, nitya, antaraṅga, paratva.
  • Prohibition: niṣedha rules such as 1.1.5, which remove a candidate.
  • Phase ordering: asiddhatva, sūtras 8.2.1 and 6.4.22, which makes one section invisible to another.
  • Checked first: asiddhatva structurally, then prohibition, then the four principles in order.
  • Only paratva, 1.4.2, and the prohibitions are numbered sūtras. Apavāda, nitya and antaraṅga are maxims.
  • Cheap mechanisms are arbitrary; principled ones are expensive. That trade is not peculiar to Sanskrit.

Test yourself

1. Name the six conflict resolution mechanisms and group them by kind.

Precedence: apavāda, nitya, antaraṅga, paratva. Prohibition: niṣedha. Phase ordering: asiddhatva. The first four decide between rules, the fifth removes a candidate, the sixth prevents the conflict arising.

2. In what order are they tried, and why is asiddhatva first?

Asiddhatva, then prohibition, then the four principles in their own order. Asiddhatva is first because it is structural: if the two rules are in different phases, the earlier one cannot see the later one's result, so there is no conflict to resolve.

3. Which of the mechanisms are numbered sūtras?

Paratva, which is 1.4.2; the prohibitions, such as 1.1.5; and asiddhatva, which is 8.2.1 and 6.4.22. Apavāda, nitya and antaraṅga are maxims of the tradition rather than rules of the text.

4. Explain the trade-off between apavāda and paratva.

Apavāda is principled but expensive: it needs a judgement about whether one condition is a subset of another, and the commentaries dispute such judgements. Paratva is cheap but arbitrary: it needs only a comparison of two addresses, and it tells you nothing about why a rule lost.

Contents This chapter on its own page

munotes.in139

Chapter Forty-Three

Formal Grammars: Alphabet, Rule, Derivation, Language

Syllabus topic Module 1, "Context-Free Grammar (CFG)"

In one line

A formal grammar is four things: an alphabet, a set of names for parts, a set of rules, and a starting name; and the language it defines is everything the rules can make.

In the wording you can write in an examination: a formal grammar is a quadruple consisting of a finite set of terminal symbols, a finite set of non-terminal symbols disjoint from it, a finite set of production rules each rewriting a string containing at least one non-terminal, and a distinguished start symbol. A string of terminals is in the language generated by the grammar if and only if it can be derived from the start symbol by finitely many applications of the rules.

The four parts, defined

Terminals. The symbols that appear in the finished strings. In a grammar of arithmetic they might be digits and the signs; in a grammar of Sanskrit they are the sounds.

Non-terminals. Names for parts that do not appear in the finished strings. They stand for categories: an expression, a noun stem, a sentence. Written in capitals by convention here.

Productions. Rules of the form "left-hand side rewrites to right-hand side". At least one non-terminal must appear on the left, because otherwise nothing can ever be replaced.

The start symbol. One of the non-terminals, the one every derivation begins from.

A derivation, and what it proves

A derivation is a sequence of strings, each obtained from the previous one by applying one production. It begins at the start symbol and ends at a string of terminals only.

If a derivation exists, the string is in the language. If none exists, it is not. That is the whole definition of the language a grammar generates, and it is exactly the tradition's criterion of correctness in [Pāṇini's Aṣṭādhyāyī, and the Text We Are Reading]: a form is correct if and only if the rules derive it.

A tiny grammar, generated

The grammar below describes strings of the shape "a noun, a verb, and optionally an object". It is deliberately small enough to enumerate entirely.

terminals rama, sita, sees, calls, quickly

non-terminals S, NP, VP, ADV

start symbol S

productions

S -> NP VP

VP -> V

VP -> V NP

VP -> V NP ADV

NP -> rama

NP -> sita

V -> sees

V -> calls

ADV -> quickly

RULES = {
    "S":   [["NP", "VP"]],
    "VP":  [["V"], ["V", "NP"], ["V", "NP", "ADV"]],
    "NP":  [["rama"], ["sita"]],
    "V":   [["sees"], ["calls"]],
    "ADV": [["quickly"]],
}

def generate(symbols):
    """Every string of terminals derivable from this sequence of symbols."""
    if not symbols:
        return [[]]
    head, rest = symbols[0], symbols[1:]
    tails = generate(rest)
    if head not in RULES:
        return [[head] + t for t in tails]
    out = []
    for body in RULES[head]:
        for front in generate(body):
            for t in tails:
                out.append(front + t)
    return out

strings = generate(["S"])
print("derivations found:", len(strings))
print("distinct strings :", len({tuple(s) for s in strings}))
print("by arithmetic    : 2 subjects x 2 verbs x (1 + 2 + 2) =", 2 * 2 * 5)
print()
for s in sorted(" ".join(x) for x in strings):
    print("  " + s)
munotes.in140

Formal Grammars: Alphabet, Rule, Derivation, Language

derivations found: 20
distinct strings : 20
by arithmetic    : 2 subjects x 2 verbs x (1 + 2 + 2) = 20

  rama calls
  rama calls rama
  rama calls rama quickly
  rama calls sita
  rama calls sita quickly
  rama sees
  rama sees rama
  rama sees rama quickly
  rama sees sita
  rama sees sita quickly
  sita calls
  sita calls rama
  sita calls rama quickly
  sita calls sita
  sita calls sita quickly
  sita sees
  sita sees rama
  sita sees rama quickly
  sita sees sita
  sita sees sita quickly

Reading the three numbers the program prints

They are three independent statements about the same grammar, and the chapter prints all three because agreement between independent routes is the only evidence worth having.

Derivations found: 20. The enumeration walked every alternative of every rule and produced twenty results.

Distinct strings: 20. No two derivations produced the same string, so this grammar is unambiguous over its own output. A grammar in which those two numbers differed would be ambiguous, and that is a property worth being able to detect rather than assume.

By arithmetic: 20. Two choices of subject, two of verb, and five shapes of verb phrase, namely the bare verb, verb with either object, and verb with either object plus the adverb.

Three routes, one answer. If the arithmetic had disagreed with the count, the arithmetic would be the thing to doubt, and saying so is a habit worth having before you write the internal assessment.

The vocabulary a question will use

TermMeaning
terminala symbol appearing in the finished string
non-terminala name for a category, never in the finished string
productiona rule rewriting a left-hand side as a right-hand side
start symbolthe non-terminal every derivation begins from
derivationa sequence of applications of productions from the start symbol
sentential formany intermediate string in a derivation, terminals and non-terminals mixed
language generatedthe set of terminal strings derivable from the start symbol
ambiguous grammarone in which some string has two different derivation structures

What a formal grammar is NOT

It is not a description of meaning. Nothing in the four parts says what a string means. "Sita calls rama quickly" is in the language and so is any other combination the rules allow, whether or not it makes sense.

munotes.in141

Formal Grammars: Alphabet, Rule, Derivation, Language

It is not a parser. The grammar says which strings are in the language. Finding out whether a given string is, and how, is parsing, and [Parsing Algorithms] is about that.

It is not unique to its language. Many different grammars generate the same set of strings, and they can differ enormously in how easy they are to parse.

Quick revision

  • Four parts: terminals, non-terminals, productions, start symbol.
  • The language generated is the set of terminal strings derivable from the start symbol, and derivability is the whole criterion.
  • A derivation is a sequence of sentential forms from the start symbol to a string of terminals.
  • The small grammar above generates 20 strings by three independent routes: the enumeration's length, the count of distinct results, and the arithmetic. Equal counts of derivations and of distinct strings means the grammar is unambiguous over its output.
  • Grammar is not meaning, not a parser, and not unique to the language it generates.

Test yourself

1. Give the four components of a formal grammar and say what makes a string a member of its language.

A finite set of terminals, a disjoint finite set of non-terminals, a finite set of productions each with at least one non-terminal on the left, and a start symbol. A terminal string is in the language if and only if there is a finite derivation of it from the start symbol.

2. How many strings does the grammar above generate, and how do you know?

Twenty. Three independent routes agree: the enumeration produced twenty derivations, those derivations gave twenty distinct strings, and the arithmetic is two subjects times two verbs times five verb-phrase shapes.

3. What is the difference between a grammar and a parser?

The grammar defines which strings are in the language. A parser decides, for a given string, whether it is in the language and by what derivation.

4. Why can two grammars be different and yet define the same language?

Because membership depends only on which terminal strings are derivable, not on the route. Different sets of productions can make the same set of strings derivable while differing in structure and in how easily they can be parsed.

Contents This chapter on its own page

munotes.in142

Chapter Forty-Four

Context-Free Grammar, and Whether Pāṇini Wrote One

Syllabus topic Module 1, "Context-Free Grammar (CFG)"

In one line

A context-free grammar replaces one name at a time with no regard for what surrounds it, and Pāṇini's rules almost never work that way.

In the wording you can write in an examination: a context-free grammar is a formal grammar in which every production has exactly one non-terminal on its left-hand side, so that a non-terminal may be replaced wherever it occurs, independently of its context. The Aṣṭādhyāyī is not context-free: most of its operative rules state a condition on surrounding material, several of its mechanisms refer to whether other rules have applied, and its rules operate on sounds and markers rather than on non-terminal categories.

What a context-free grammar can do

The restriction is simple and the consequence is large.

The restriction. Exactly one non-terminal on the left. So a production is always of the form "A rewrites to something".

What follows. A derivation can be drawn as a tree: the start symbol at the root, and each application of a rule expanding one node into its children. The tree is called a parse tree, and every context-free derivation has one.

Why that matters. A tree is a statement about structure. It says which parts group together, which is exactly what you need if the next thing you want to do is compute the meaning.

Worked: a parse tree, drawn

Using the grammar of [Formal Grammars: Alphabet, Rule, Derivation, Language] and the string "rama sees sita quickly".

S

+-- NP

+-- rama

+-- VP

+-- V

+-- sees

+-- NP

+-- sita

+-- ADV

+-- quickly

Each interior node is a non-terminal, each leaf a terminal, and the string is the leaves read left to right. Nothing in the tree depends on context, because nothing in the rules did.

Why the Aṣṭādhyāyī is not one

Three features, each with the sūtra that creates it. Any one of them is enough.

One: the rules state contexts

7.3.84 causes guṇa of the base when a sārvadhātuka or ārdhadhātuka affix follows. Written as a production, its left-hand side has to mention both the vowel being changed and the affix that licenses the change, which is two symbols. That is context-sensitive in the strict sense of [Context-Sensitive Operations].

And the context mechanism is general. Sūtras 1.1.66 and 1.1.67 exist precisely to let any rule state a left or right context by a case ending. A grammar with a general device for stating contexts is not a context-free grammar that happens to have a few.

Two: rules refer to whether other rules have applied

Sūtra 8.2.1 makes the tripādī asiddha for earlier rules: an earlier rule proceeds as though a later one had not operated. Sūtra 6.4.22 does the same within a section of book six.

munotes.in143

Context-Free Grammar, and Whether Pāṇini Wrote One

A context-free production cannot express that at all. Its condition is a symbol; it has no way to say "provided rule R has not fired". This is not a matter of degree: the mechanism is outside the formalism.

Three: the objects are not non-terminals

A context-free grammar rewrites category names. Pāṇini's rules operate on sounds, on roots, on affixes and on markers, and the markers are deleted before the output. There is no level at which a Pāṇinian rule expands a category into its parts in the way "VP rewrites to V NP" does.

What Pāṇini has instead is a derivation that starts from a root and an affix and transforms them. That is a rewriting system over strings with side conditions, which is a different kind of object from a phrase-structure grammar.

The comparison, as a table

Context-free grammarThe Aṣṭādhyāyī
Left-hand sideexactly one non-terminala sound, a stem, or a configuration
Contextcannot be statedstated by case ending, 1.1.66 and 1.1.67
Order of applicationirrelevant to the resultdecided by four precedence principles
Reference to other rulesimpossibleasiddhatva, 8.2.1 and 6.4.22
Structure produceda parse treea sequence of forms, not a tree
What it is fordefining a set of strings, and recovering structurederiving the correct form of a word
Conflict between rulesdoes not arise; alternatives are alternativesa central concern with a stated policy

What Pāṇini's system IS, said positively

Refusing a classification is only half an answer. The positive statement is this.

It is a rewriting system with a fixed application strategy. Rules transform strings; the strategy is the precedence order; the strategy is what makes the system deterministic. [Rewrite Systems] gives the vocabulary and [Śāstra Rule Precedence as a Deterministic Finite Rewrite System] builds one.

Its nearest modern relatives are not phrase-structure grammars. They are two-level morphology and finite-state transducers, which likewise transform a form into a form with contexts and ordered rules, and which are the standard tools in computational morphology. [Foundations of NLP, and What Pāṇinian Grammar Contributed] takes that up.

What MU's own outcome asks, and how to answer it

Her OC 3 is: model and compare Pāṇini's generative grammar with modern formal grammars and parsing systems. A five-mark answer has four moves and they fit in about a hundred and thirty words.

Answer. Both are generative: a finite rule set defines the well-formed strings, and a form is correct if and only if the rules derive it. They differ in three respects. A context-free production has one non-terminal on its left and may be applied wherever that symbol occurs; Pāṇini's operative rules state a surrounding condition, for which sūtras 1.1.66 and 1.1.67 supply a general notation, so they are context-sensitive. A context-free derivation's result does not depend on the order of application; Pāṇini's does, and four precedence principles decide it. And a context-free rule cannot refer to whether another rule has applied, whereas asiddhatva under 8.2.1 does exactly that. So the Aṣṭādhyāyī is better modelled as an ordered string rewriting system, whose nearest modern relatives are finite-state transducers in computational morphology rather than phrase-structure grammars.

munotes.in144

Context-Free Grammar, and Whether Pāṇini Wrote One

What this chapter does NOT say

It does not say the Aṣṭādhyāyī is more powerful than a context-free grammar in the formal sense. Power in that sense is about which sets of strings can be generated, and settling that for the whole Aṣṭādhyāyī would require a formalisation of it that this paper does not attempt. What is claimed is that it is not context-free, which follows from its mechanisms.

It does not say the comparison is worthless. The comparison is how you find out what each system is for, and the differences are more informative than the similarities.

It does not say Pāṇini anticipated Chomsky. Nothing here is a claim about influence. [Foundations of NLP, and What Pāṇinian Grammar Contributed] separates the claims about influence from the evidence for them.

Quick revision

  • Context-free: exactly one non-terminal on the left, so a rule applies wherever its symbol occurs and a derivation has a parse tree.
  • The Aṣṭādhyāyī is not context-free, for three reasons: its rules state contexts, with 1.1.66 and 1.1.67 as the general notation; its mechanisms refer to whether other rules have applied, by 8.2.1 and 6.4.22; and its rules operate on sounds and markers, not on categories.
  • Order of application is irrelevant in a context-free grammar and decisive in Pāṇini's.
  • Positively: it is an ordered string rewriting system, closer to a finite-state transducer than to a phrase-structure grammar.
  • Not claimed: a formal result about generative power, or any claim about influence.

Test yourself

1. Define a context-free grammar and say what property of derivations follows.

One in which every production has exactly one non-terminal on its left-hand side. It follows that every derivation can be drawn as a parse tree, since each step expands one node into its children.

2. Give three reasons the Aṣṭādhyāyī is not context-free, each with a sūtra.

Its operative rules state contexts, with 1.1.66 and 1.1.67 providing the general notation by case ending. Its mechanisms refer to whether other rules have applied, by 8.2.1 and 6.4.22. And its rules operate on sounds, stems, affixes and markers rather than on non-terminal categories.

3. What is Pāṇini's system, stated positively?

An ordered string rewriting system with a fixed application strategy supplied by the four precedence principles, whose nearest modern relatives are the finite-state transducers used in computational morphology.

munotes.in145

Context-Free Grammar, and Whether Pāṇini Wrote One

4. Why does this chapter refuse to make a claim about generative power?

Because generative power is a statement about which sets of strings a formalism can produce, and establishing it for the Aṣṭādhyāyī would require a formalisation of the whole grammar that this paper does not attempt. That it is not context-free follows from its mechanisms without any such formalisation.

Contents This chapter on its own page

munotes.in146

Chapter Forty-Five

The Chomsky Hierarchy, and Where the Aṣṭādhyāyī Sits

Syllabus topic Module 1, "Automata theory", "Context-Free Grammar (CFG)"

In one line

Grammars fall into four nested classes by how much their rules are allowed to say, and each class has a machine that recognises exactly its languages.

In the wording you can write in an examination: the Chomsky hierarchy classifies formal grammars into four types by restrictions on the form of their productions. Type 3, regular, allows a single non-terminal to rewrite to a terminal optionally followed by a single non-terminal; type 2, context-free, allows any right-hand side but only a single non-terminal on the left; type 1, context-sensitive, allows any left-hand side provided the right-hand side is at least as long; type 0, unrestricted, allows any production. Each class is properly contained in the next, and each corresponds to a class of recognising machine.

The four classes

TypeNameShape of a productionMachine that recognises it
3regularA rewrites to a, or to a Bfinite automaton
2context-freeA rewrites to anythingpushdown automaton
1context-sensitiveany left side, right side no shorterlinear bounded automaton
0unrestrictedany production at allTuring machine

The containment is proper. Every regular language is context-free and some context-free languages are not regular; every context-free language is context-sensitive and some context-sensitive ones are not context-free; and so on. So the hierarchy is a real hierarchy and not just a set of labels.

What each class cannot do, which is how you tell them apart

This is the practical content of the hierarchy, and it is what a question will actually ask.

Regular cannot count. The language of strings with equally many a's and b's is not regular, because a finite automaton has finitely many states and cannot remember an unbounded count. Nor can it match nested brackets.

Context-free can count to one depth but cannot compare two. Balanced brackets are context-free. Strings of the form a to the n, b to the n, c to the n, with all three counts equal, are not, because a pushdown automaton has one stack and can compare only one pair.

Context-sensitive can do that, and more. The three-way equality is context-sensitive. What it cannot do is anything requiring unbounded working space beyond the length of the input.

Unrestricted can do anything computable, which also means membership is undecidable in general: there is no procedure that always answers whether a string is in the language.

The costs

The other half of the hierarchy is the price of each level, and it is why nobody uses type 0 for a programming language.

TypeDeciding membershipTypical use
3linear time, constant spacelexical analysis, search patterns, tokenising
2polynomial time, and linear for the restricted forms compilers usethe syntax of programming languages
1decidable, but no efficient general algorithmrarely used directly
0undecidable in generalnot used for syntax
munotes.in147

The Chomsky Hierarchy, and Where the Aṣṭādhyāyī Sits

The lesson a language designer takes from this is that expressiveness is not free, and the useful levels are the bottom two. Programming languages are deliberately kept context-free or very nearly so, because a parser for them then exists and is fast.

The placement question, and why this book refuses it

"Where does the Aṣṭādhyāyī sit in the hierarchy?" is asked often and this book does not answer it. The refusal has a reason and the reason is the answer worth writing.

The hierarchy classifies grammars in a fixed formalism. A production is a pair of strings over a fixed alphabet of terminals and non-terminals. To place a system in the hierarchy you must first express it in that formalism.

The Aṣṭādhyāyī is not in that formalism. Its rules refer to markers that are deleted, to whether other rules have operated, and to positions in the text. [Context-Free Grammar, and Whether Pāṇini Wrote One] lists these.

So a placement would be a claim about a formalisation, not about the text. Different formalisations of the same grammar can land in different classes, and the choice of formalisation is where all the content is. A paper that states a class without stating its formalisation has not said anything checkable.

What can be said, and is. Individual mechanisms can be placed, and this book places them.

MechanismWhat it corresponds to
a rule with no context, adding an affixa context-free production
7.3.84 with its following-affix conditiona context-sensitive production
the pratyāhāra device, naming a set of soundsa character class, which is regular
8.2.1's asiddhatvaa phase boundary, which is outside the hierarchy entirely
the four precedence principlesan application strategy, which is outside the hierarchy entirely

The last two rows are the interesting ones. A strategy for choosing which rule to apply is not part of a grammar in the Chomsky sense at all: a grammar defines a set of strings and says nothing about how they are produced. Pāṇini's system is as much a procedure as a grammar, and the hierarchy has no axis for that.

Worked example: why a strategy is not a grammar

Take two productions that both apply to the same string.

In a context-free grammar this is not a problem and not a choice. Both derivations exist, both results are in the language, and if they produce different structures for the same string the grammar is ambiguous.

In Pāṇini's system it is a problem with a stated answer. Exactly one of the two rules is to apply, the precedence principles say which, and the other result is not a Sanskrit word.

munotes.in148

The Chomsky Hierarchy, and Where the Aṣṭādhyāyī Sits

So the two systems differ in what a rule set is for. A Chomsky grammar defines a set. Pāṇini's defines a function, from a root and a meaning to one form. A system that defines a function needs a strategy and a system that defines a set does not, and that is why the hierarchy does not have a place for the strategy.

Quick revision

  • Four types: 3 regular, 2 context-free, 1 context-sensitive, 0 unrestricted, each properly contained in the next.
  • Machines: finite automaton, pushdown automaton, linear bounded automaton, Turing machine.
  • Regular cannot count or match brackets; context-free can match one nesting but not compare two counts; context-sensitive can; unrestricted makes membership undecidable.
  • Costs rise with expressiveness, which is why programming languages are kept at type 2 or below.
  • This book does not place the Aṣṭādhyāyī in a class, because a placement is a claim about a formalisation and no formalisation of the whole grammar is offered.
  • It does place mechanisms: pratyāhāra is regular, a context-free rule is type 2, 7.3.84 is type 1, and asiddhatva and the precedence principles are outside the hierarchy.
  • A Chomsky grammar defines a set of strings; Pāṇini's defines a function to one form, which is why it needs a strategy.

Test yourself

1. Name the four types with their production shapes and their machines.

Type 3 regular, a non-terminal rewriting to a terminal or a terminal and a non-terminal, recognised by a finite automaton. Type 2 context-free, one non-terminal on the left, a pushdown automaton. Type 1 context-sensitive, any left side with a right side no shorter, a linear bounded automaton. Type 0 unrestricted, any production, a Turing machine.

2. Give one language each type cannot generate that the next can.

Regular cannot generate balanced brackets, which is context-free. Context-free cannot generate strings with three equal counts, which is context-sensitive. Context-sensitive cannot generate every computably enumerable language, which type 0 can.

3. Why does this book refuse to place the Aṣṭādhyāyī in the hierarchy?

Because the hierarchy classifies grammars already expressed in one fixed formalism, and the Aṣṭādhyāyī is not in it: its rules refer to deleted markers, to whether other rules have applied, and to positions in the text. A placement would therefore be a claim about a chosen formalisation, and different formalisations land differently.

4. Why are the precedence principles outside the hierarchy altogether?

Because a Chomsky grammar defines a set of strings and says nothing about how a derivation is chosen. Pāṇini's system defines a single correct form, so it needs a strategy for choosing among applicable rules, and the hierarchy has no axis for a strategy.

Contents This chapter on its own page

munotes.in149

Chapter Forty-Six

Rewrite Systems

Syllabus topic Module 1, "Rewrite systems"

In one line

A rewrite system is a set of rules that replace one piece of a string with another, applied over and over until nothing applies.

In the wording you can write in an examination: a string rewriting system, also called a semi-Thue system, is a finite alphabet together with a finite set of rules each of which replaces one string by another. A string to which no rule applies is in normal form. The system terminates if every string reaches a normal form after finitely many steps, and it is confluent if whenever two different sequences of rewrites are possible from one string, they can be continued so as to reach the same string.

The vocabulary

A rule. A pair of strings, written left rewrites to right. Unlike a grammar production, neither side has to contain a special category symbol: a rewrite system works on plain strings.

A step. Find an occurrence of some rule's left side in the string, and replace it by that rule's right side.

Normal form. A string with no occurrence of any rule's left side. The computation is over.

Termination. Every string reaches a normal form. Not guaranteed: a rule set can loop.

Confluence. If two different steps are possible, the two results can be brought back together. Where it holds, the normal form does not depend on which step you took.

Convergence. Terminating and confluent together. A convergent system gives every string exactly one normal form, which is the property you want if the system is meant to compute a function.

Both properties, demonstrated

RULES = [("ab", "ba"), ("ba", "ab")]

def step(s, rules):
    for i in range(len(s)):
        for left, right in rules:
            if s.startswith(left, i):
                return s[:i] + right + s[i + len(left):], (left, right, i)
    return None, None

def normalise(s, rules, limit=8):
    trail = [s]
    for _ in range(limit):
        nxt, used = step(s, rules)
        if nxt is None:
            return trail, "normal form reached"
        s = nxt
        trail.append(s)
    return trail, "gave up after %d steps: this rule set does not terminate" % limit

print("rule set A: ab -> ba and ba -> ab")
trail, why = normalise("ab", RULES)
print("  " + " -> ".join(trail))
print("  " + why)
print()

SORTER = [("ba", "ab")]
print("rule set B: ba -> ab only")
for start in ("ba", "bba", "bab", "bbaa", "abab"):
    trail, why = normalise(start, SORTER, limit=12)
    print("  %-6s %-34s %s" % (start, " -> ".join(trail), why))
rule set A: ab -> ba and ba -> ab
  ab -> ba -> ab -> ba -> ab -> ba -> ab -> ba -> ab
  gave up after 8 steps: this rule set does not terminate

rule set B: ba -> ab only
  ba     ba -> ab                           normal form reached
  bba    bba -> bab -> abb                  normal form reached
  bab    bab -> abb                         normal form reached
  bbaa   bbaa -> baba -> abba -> abab -> aabb normal form reached
  abab   abab -> aabb                       normal form reached
munotes.in150

Rewrite Systems

Reading the two rule sets

Rule set A does not terminate, and the reason is visible in the trail: the two rules undo each other. A rewrite system can loop, and nothing in the notation prevents it.

Rule set B terminates, and it sorts. Every normal form has its a's before its b's. The termination argument is a good one to know: each application of the rule moves one a one place to the left, and the total of the positions of the a's is a whole number that strictly decreases and cannot go below zero. Exhibiting a quantity that strictly decreases is how termination is proved.

Rule set B is also confluent. Whatever order the rewrites are done in, the result is the same string: all the a's, then all the b's. So it is convergent, and it computes a function.

The Aṣṭādhyāyī as a rewrite system

This is the positive characterisation [Context-Free Grammar, and Whether Pāṇini Wrote One] promised.

The strings are forms under derivation. A root plus affixes, with markers.

The rules are the vidhi rules. Each replaces one element by another in a stated context.

A normal form is a finished word. No further rule applies.

And there is a strategy. This is the part that makes it different from rule set B above, and it is the important part.

Why a strategy changes everything

Rule set B is confluent, so it needs no strategy: any order gives the same answer. Most interesting rewrite systems are not confluent, and then the order decides the result.

Pāṇini's system is not confluent. Two rules can apply and give different words, and only one of them is Sanskrit. So a strategy is not an optimisation; it is part of the specification.

The strategy is the precedence order. Apavāda, nitya, antaraṅga, paratva, as [Rule Precedence: The Four Principles] sets out, with prohibitions and asiddhatva outside the queue.

A rewrite system plus a deterministic strategy computes a function. Given a root and a meaning, exactly one sequence of rewrites is licensed, so exactly one form results. That is what the grammar is for, and it is why the tradition invested so heavily in the precedence rules rather than in making the rules non-overlapping.

A convergent rewrite systemThe Aṣṭādhyāyī
Confluentyes, by assumptionno
Needs a strategynoyes, and the strategy is stated
What the strategy is forefficiency, at mostcorrectness
Result if the strategy is ignoredthe same normal forma different form, which is not Sanskrit
Terminationmust be provedassumed by the tradition, and not proved
munotes.in151

Rewrite Systems

Termination, and the honest gap

The last row is a real gap and it should be stated.

Nothing in the tradition proves the grammar terminates. No argument is offered that a derivation cannot loop. The evidence is that derivations in practice finish, which is the same evidence rule set A appeared to have until somebody tried "ab".

The asiddhatva mechanism helps. By making phase two invisible to phase one, 8.2.1 prevents a class of loops in which a phase-one rule keeps re-firing because a phase-two rule keeps undoing its work. That is a structural reason to expect termination, and it is the nearest thing to an argument the system provides.

But it is not a proof, and a chapter that claimed otherwise would be overstating. What can be said is that the mechanism exists and that it forecloses the obvious failure mode.

What a rewrite system is NOT

It is not a grammar in the Chomsky sense. It has no start symbol and no non-terminals, and it need not define a set of strings. A grammar generates; a rewrite system transforms.

It is not deterministic by itself. Which occurrence of which rule to rewrite is a choice, and the notation does not make it.

It is not guaranteed to do anything. Termination and confluence are properties to be established, not features of the notation.

Quick revision

  • A string rewriting system is an alphabet plus rules replacing one string by another. A step applies one rule at one occurrence.
  • Normal form: no rule applies. Termination: every string reaches one. Confluence: different orders can be brought back together. Convergent: both.
  • Rule set A, ab to ba and ba to ab, loops. Rule set B, ba to ab only, terminates and sorts, and the termination proof exhibits a strictly decreasing quantity.
  • The Aṣṭādhyāyī is a rewrite system over forms, and it is NOT confluent, so its strategy is part of its correctness and not an optimisation.
  • The strategy is the four precedence principles, with prohibition and asiddhatva outside the queue.
  • Termination is not proved by the tradition. Asiddhatva forecloses the obvious loop and is not a proof.

Test yourself

1. Define normal form, termination and confluence.

A normal form is a string to which no rule applies. A system terminates if every string reaches a normal form in finitely many steps. It is confluent if whenever two different rewrites are possible from a string, the two results can be continued to a common string.

2. Prove that the rule "ba rewrites to ab" terminates.

Each application moves one a one position to the left, so the sum of the positions of the a's strictly decreases with every step. It is a non-negative whole number, so it cannot decrease for ever, and the system must halt.

munotes.in152

Rewrite Systems

3. Why is Pāṇini's strategy part of his specification rather than an optimisation?

Because his rule set is not confluent: two applicable rules can yield different forms and only one is correct. So the order of application determines the answer, and the precedence principles are what fix the answer.

4. What does asiddhatva contribute to termination, and what does it not?

By making the later phase invisible to the earlier one, it prevents a loop in which an earlier rule re-fires because a later rule keeps undoing its work. It does not amount to a proof that no derivation loops, and the tradition offers none.

Contents This chapter on its own page

munotes.in153

Chapter Forty-Seven

Automata Theory Foundations: The Finite Automaton

Syllabus topic Module 1, "Automata theory"

In one line

A finite automaton is a machine with a fixed number of states that reads its input once, left to right, and ends in a state that says yes or no.

In the wording you can write in an examination: a deterministic finite automaton is a quintuple consisting of a finite set of states, a finite input alphabet, a transition function assigning to each state and input symbol exactly one next state, a start state, and a set of accepting states. It reads the input one symbol at a time, changing state at each, and accepts if the state after the last symbol is accepting. It recognises exactly the regular languages.

The five parts

States. A finite set. The machine's entire memory is which state it is in, and that is why it cannot count without bound.

Alphabet. The symbols it reads. Here, L and G.

Transition function. For each state and each symbol, exactly one next state. "Exactly one" is what deterministic means, and it means the machine never has a choice.

Start state. Where it begins.

Accepting states. If the machine is in one of these when the input ends, it accepts.

A machine over metres

Here is a condition a prosodist might actually care about: no two heavy syllables next to each other. A finite automaton can decide it, and three states are enough.

StateWhat it meansOn reading LOn reading G
q0the last syllable was light, or there was noneq0q1
q1the last syllable was heavyq0qX
qXtwo heavy syllables have been seen togetherqXqX

Accepting: q0 and q1. Rejecting: qX. Once the machine reaches qX it can never leave, which is why such a state is called a trap or a dead state.

Read the table as the whole design. The machine remembers exactly one thing, whether the previous syllable was heavy, and that is all the condition requires. Choosing what to remember is the whole art of designing an automaton.

The machine, run

# A DFA over {L, G} that accepts exactly the patterns with no two adjacent gurus.
START, SEEN_G, DEAD = "q0", "q1", "qX"
ACCEPTING = {START, SEEN_G}
DELTA = {
    (START,  "L"): START,
    (START,  "G"): SEEN_G,
    (SEEN_G, "L"): START,
    (SEEN_G, "G"): DEAD,
    (DEAD,   "L"): DEAD,
    (DEAD,   "G"): DEAD,
}

def run(word, trace=False):
    state = START
    steps = [state]
    for ch in word:
        state = DELTA[(state, ch)]
        steps.append(state)
    if trace:
        print("  %-8s %s   %s" % (word, " ".join(steps),
                                  "accept" if state in ACCEPTING else "reject"))
    return state in ACCEPTING

print("traces")
for w in ("", "L", "G", "GL", "GG", "GLG", "GGL", "LGLG"):
    run(w, trace=True)

print()
def prastara(n):
    row = ["G"] * n
    rows = ["".join(row)]
    while "G" in row:
        k = row.index("G")
        row = ["G"] * k + ["L"] + row[k + 1:]
        rows.append("".join(row))
    return rows

print("%-4s %-8s %-10s %s" % ("n", "patterns", "accepted", "Fibonacci check"))
fib = [1, 2]
while len(fib) < 14:
    fib.append(fib[-1] + fib[-2])
for n in range(1, 13):
    ok = sum(1 for p in prastara(n) if run(p))
    print("%-4d %-8d %-10d %s" % (n, 2 ** n, ok, "matches" if ok == fib[n] else "DOES NOT MATCH"))
munotes.in154

Automata Theory Foundations: The Finite Automaton

traces
           q0   accept
  L        q0 q0   accept
  G        q0 q1   accept
  GL       q0 q1 q0   accept
  GG       q0 q1 qX   reject
  GLG      q0 q1 q0 q1   accept
  GGL      q0 q1 qX qX   reject
  LGLG     q0 q0 q1 q0 q1   accept

n    patterns accepted   Fibonacci check
1    2        2          matches
2    4        3          matches
3    8        5          matches
4    16       8          matches
5    32       13         matches
6    64       21         matches
7    128      34         matches
8    256      55         matches
9    512      89         matches
10   1024     144        matches
11   2048     233        matches
12   4096     377        matches

What that second table shows

The accepted counts are 2, 3, 5, 8, 13, 21, 34, 55, 89, 144, 233, 377. Each is the sum of the two before it, which is the Fibonacci recurrence, and the program checks it for every n from one to twelve.

The reason is a two-line argument. Count the acceptable patterns of n syllables by their last syllable. If it is light, the first n minus one syllables may be any acceptable pattern. If it is heavy, the syllable before it must be light, so the first n minus two may be any acceptable pattern. So the count for n is the count for n minus one plus the count for n minus two.

That is the same kind of argument as the Meru's addition rule, in [Meru-Prastāra: Halāyudha's Staircase]: split on the last syllable and add the cases. The technique transfers, which is what makes it worth learning.

And no historical claim is made. This is a fact about the machine defined above, established by running it. Nothing in this chapter says that any classical text states it, because this book has not verified that in a source it has read.

Why the machine is finite, and what that costs

Three states are enough for this condition because deciding it needs one bit of memory: was the previous syllable heavy?

Some conditions need more, and some need infinitely many. "Equally many light and heavy syllables" cannot be decided by any finite automaton, because the machine would have to remember the running difference, which is unbounded. That is the limitation [The Chomsky Hierarchy, and Where the Aṣṭādhyāyī Sits] records as "regular cannot count".

munotes.in155

Automata Theory Foundations: The Finite Automaton

So the test for whether a DFA can do a job is: can the decision be made with a bounded amount of memory about what has been read? If yes, a DFA exists. If no, it does not.

Why this matters for the next chapter

The precedence cascade of [Rule Precedence: The Four Principles] is a decision procedure over a fixed, finite set of considerations, taken in a fixed order, with exactly one outcome at each step. That is a deterministic machine, and [Śāstra Rule Precedence as a Deterministic Finite Rewrite System] builds it as one.

The vocabulary this chapter supplies for that is: a state is the situation of the decision so far, a transition is the consideration being applied, and determinism is the property that exactly one outcome follows.

What a DFA is NOT

It is not a parser. It says yes or no. It produces no structure.

It is not more powerful with more states. Adding states extends what a particular machine can do; no number of states makes a DFA able to count without bound.

It is not non-deterministic. A machine with a choice at some state is a different object, and for every such machine there is a deterministic one accepting the same language, possibly with many more states.

Quick revision

  • Five parts: states, alphabet, transition function, start state, accepting states. Deterministic means exactly one next state for each state and symbol.
  • A DFA recognises exactly the regular languages, and its whole memory is which state it is in.
  • The three-state machine above accepts the patterns with no two adjacent heavy syllables. qX is a trap state.
  • The accepted counts follow the Fibonacci recurrence, and the reason is the split on the last syllable, as in the Meru's addition rule. No historical claim is made about that.
  • The test for whether a DFA can do a job: is bounded memory about what has been read enough?
  • A DFA gives a yes or no, produces no structure, and cannot count without bound however many states it has.

Test yourself

1. Give the five components of a DFA and say what "deterministic" means.

A finite set of states, an input alphabet, a transition function, a start state, and a set of accepting states. Deterministic means the transition function gives exactly one next state for every state and input symbol, so the machine never has a choice.

2. Design a DFA over L and G that accepts patterns ending in a light syllable.

Two states. In the start state, reading L goes to the accepting state and reading G stays; in the accepting state, reading L stays and reading G returns to the start. Accept only the second state, so the empty pattern is rejected.

munotes.in156

Automata Theory Foundations: The Finite Automaton

3. Why do the accepted counts in the table follow the Fibonacci recurrence?

Classify the acceptable patterns of n syllables by the last syllable. If it is light, the first n minus one may be any acceptable pattern; if it is heavy, the one before must be light, so the first n minus two may be any acceptable pattern. The count for n is therefore the sum of the counts for n minus one and n minus two.

4. Give a condition on patterns that no DFA can decide, and say why.

Whether a pattern has equally many light and heavy syllables. The machine would have to remember the running difference between the two counts, which is unbounded, and a DFA has only finitely many states.

Contents This chapter on its own page

munotes.in157

Chapter Forty-Eight

Śāstra Rule Precedence as a Deterministic Finite Rewrite System

Syllabus topic Module 1, "Deterministic Finite Rewrite System", "Conflict resolution mechanisms", "Rewrite systems"

In one line

Put a rule set, a precedence order and a phase boundary together and you have a machine: one input, one output, and a reason for every step.

In the wording you can write in an examination: a deterministic finite rewrite system is a finite set of string rewriting rules together with a strategy that selects, for any string to which more than one rule applies, exactly one rule to apply. Where every rule strictly reduces a well-founded measure the system terminates, and where the strategy is total the system computes a function even if the rule set is not confluent.

Problem statement, in MU's own form

IKS concept as CS concept: śāstra rule precedence as a deterministic finite rewrite system.

Statement. Implement a rewrite system over a two-letter alphabet in which rules carry the tradition's precedence flags, apavāda, nitya and antaraṅga, and are addressed as the Aṣṭādhyāyī's rules are addressed, so that paratva can be applied as a tie-break. Resolve every conflict by the four principles in the tradition's order, report which principle decided, divide the rules into two phases in the manner of sūtra 8.2.1, prove termination, and measure confluence.

Conceptual mapping table

Classical elementComputer science element
a vidhi rulea rewrite rule, a pair of strings
a rule's address, book quarter positiona sortable key used by the tie-break
apavāda, a special rulea flag marking a narrower condition
nityaa flag marking a rule that would apply either way
antaraṅgaa flag marking a rule whose conditions lie closer in
paratva, sūtra 1.4.2the maximum of the addresses
sūtra 8.2.1, asiddhatvatwo phases, the second run only when the first is exhausted
a finished worda normal form
the whole apparatusa function from input string to output string

Algorithm specification, in pseudo-code

ALGORITHM Derive(word)

for phase in 1, 2 do

loop

names <- every rule of this phase whose left side occurs in word

if names is empty then leave the loop

winner, principle <- Resolve(names)

word <- word with the FIRST occurrence of winner's left side replaced

record (phase, word before, names, winner, principle, word after)

end loop

end for

return word

ALGORITHM Resolve(names)

if one name then return it, "only one applies"

for flag in SPECIAL, NITYA, INNER do

marked <- the names carrying that flag

if exactly one marked then return it, the flag's principle

if more than one marked then names <- marked

end for

return the name with the greatest address, "1.4.2 paratva"

One design point to defend in a viva. When more than one candidate carries the same flag, the cascade narrows the set to those and continues rather than giving up. That is what the tradition does: two special rules are still compared by the remaining principles.

munotes.in158

Śāstra Rule Precedence as a Deterministic Finite Rewrite System

Working code

"""Sastra rule precedence as a deterministic finite rewrite system.

Every rule strictly SHORTENS the string, so termination is guaranteed.
The four precedence principles are declared on the rules, as the tradition
declares them, and the resolver applies them in the tradition's order.
"""

#      address  special nitya  inner  phase  left    right
RULES = {
    "R1": ("2.1.1", False, False, False, 1, "aa",  "a"),
    "R2": ("2.1.2", True,  False, False, 1, "aab", "ab"),
    "R3": ("3.1.1", False, True,  False, 1, "bb",  "b"),
    "R4": ("3.1.2", False, False, True,  1, "ba",  "a"),
    "R5": ("4.1.1", False, False, False, 1, "aaa", "aa"),
    "R6": ("5.1.1", False, False, False, 1, "aaa", "a"),
    "R7": ("6.1.1", False, False, False, 1, "aa",  ""),
    "R8": ("8.2.5", False, False, False, 2, "b",   ""),
}
ORDER = ["R1", "R2", "R3", "R4", "R5", "R6", "R7", "R8"]
SPECIAL, NITYA, INNER = 1, 2, 3


def address(name):
    return tuple(int(x) for x in RULES[name][0].split("."))


def candidates(word, phase):
    return [n for n in ORDER
            if RULES[n][4] == phase and RULES[n][5] in word]


def resolve(names):
    """The cascade, in the tradition's order. Returns winner and principle."""
    if len(names) == 1:
        return names[0], "only one applies"
    for flag, principle in ((SPECIAL, "apavada"), (NITYA, "nitya"), (INNER, "antaranga")):
        marked = [n for n in names if RULES[n][flag]]
        if len(marked) == 1:
            return marked[0], principle
        if len(marked) > 1:
            names = marked
    return max(names, key=address), "1.4.2 paratva"


def apply_once(word, name):
    left, right = RULES[name][5], RULES[name][6]
    i = word.index(left)
    return word[:i] + right + word[i + len(left):]


def derive(word):
    trail = []
    for phase in (1, 2):
        while True:
            names = candidates(word, phase)
            if not names:
                break
            winner, why = resolve(names)
            nxt = apply_once(word, winner)
            assert len(nxt) < len(word), "a rule that does not shorten breaks termination"
            trail.append((phase, word, names, winner, why, nxt))
            word = nxt
    return word, trail


def any_order(word, phase, depth=0):
    """Every normal form reachable by applying rules of this phase in ANY order."""
    names = candidates(word, phase)
    if not names or depth > 20:
        return {word}
    out = set()
    for n in names:
        out |= any_order(apply_once(word, n), phase, depth + 1)
    return out


CASES = ["aab", "bba", "baa", "aaba", "aaa", "aa", "bb", "ba", "b", "abba", "aabb", "a", ""]

print("%-8s %-12s %-6s %s" % ("input", "normal form", "steps", "principles used"))
for case in CASES:
    out, trail = derive(case)
    used = ", ".join(sorted({t[4] for t in trail})) or "none"
    print("%-8s %-12s %-6d %s" % (case or "(empty)", out or "(empty)", len(trail), used))

print()
print("each of the four principles decides at least once")
seen = {}
for case in CASES:
    for t in derive(case)[1]:
        seen.setdefault(t[4], (case, t[1], "/".join(t[2]), t[3]))
for principle in ("apavada", "nitya", "antaranga", "1.4.2 paratva"):
    if principle in seen:
        case, before, names, winner = seen[principle]
        print("  %-14s on input %-6s at %-5s among %-14s -> %s"
              % (principle, case, before, names, winner))
    else:
        print("  %-14s NEVER EXERCISED by these cases" % principle)

print()
print("one derivation traced: baa")
out, trail = derive("baa")
for phase, before, names, winner, why, after in trail:
    print("  phase %d  %-5s applicable %-14s -> %-3s by %-16s gives %s"
          % (phase, before, "/".join(names), winner, why, after or "(empty)"))
print("  normal form:", out or "(empty)")

print()
print("termination: every rule shortens the string, so no derivation can loop")
print("confluence:  checked below, input by input")
print("%-8s %-12s %s" % ("input", "strategy", "normal forms reachable in any order"))
for case in CASES:
    strat, _ = derive(case)
    forms = set()
    for f in any_order(case, 1):
        forms |= any_order(f, 2)
    tag = "confluent" if len(forms) == 1 else "not confluent"
    print("%-8s %-12s %-22s %s"
          % (case or "(empty)", strat or "(empty)",
             ", ".join(sorted(x or "(empty)" for x in forms)), tag))
munotes.in159

Śāstra Rule Precedence as a Deterministic Finite Rewrite System

input    normal form  steps  principles used
aab      a            2      apavada, only one applies
bba      a            2      nitya, only one applies
baa      (empty)      2      1.4.2 paratva, antaranga
aaba     (empty)      3      1.4.2 paratva, apavada, only one applies
aaa      a            1      1.4.2 paratva
aa       (empty)      1      1.4.2 paratva
bb       (empty)      2      only one applies
ba       a            1      only one applies
b        (empty)      1      only one applies
abba     (empty)      3      1.4.2 paratva, nitya, only one applies
aabb     a            3      apavada, only one applies
a        a            0      none
(empty)  (empty)      0      none

each of the four principles decides at least once
  apavada        on input aab    at aab   among R1/R2/R7       -> R2
  nitya          on input bba    at bba   among R3/R4          -> R3
  antaranga      on input baa    at baa   among R1/R4/R7       -> R4
  1.4.2 paratva  on input baa    at aa    among R1/R7          -> R7

one derivation traced: baa
  phase 1  baa   applicable R1/R4/R7       -> R4  by antaranga        gives aa
  phase 1  aa    applicable R1/R7          -> R7  by 1.4.2 paratva    gives (empty)
  normal form: (empty)

termination: every rule shortens the string, so no derivation can loop
confluence:  checked below, input by input
input    strategy     normal forms reachable in any order
aab      a            (empty), a             not confluent
bba      a            a                      confluent
baa      (empty)      (empty), a             not confluent
aaba     (empty)      (empty), a             not confluent
aaa      a            (empty), a             not confluent
aa       (empty)      (empty), a             not confluent
bb       (empty)      (empty)                confluent
ba       a            a                      confluent
b        (empty)      (empty)                confluent
abba     (empty)      (empty), a             not confluent
aabb     a            (empty), a             not confluent
a        a            a                      confluent
(empty)  (empty)      (empty)                confluent

Reading the output, in four parts

All four principles decide at least once

The second block of the output exists because of a specific failure mode. An engine can implement four branches and exercise two, and then two of its branches have never run. A test set that leaves a branch unexercised has not tested it, and this program reports which principle decided in each derivation so that the gap is visible rather than assumed away. Here each of the four decides at least once, and the block names the input and the competing rules for each.

munotes.in160

Śāstra Rule Precedence as a Deterministic Finite Rewrite System

Termination is proved, not observed

Every rule's right side is shorter than its left side, so every application strictly reduces the length of the string. Length is a non-negative whole number, so it cannot decrease for ever.

The program asserts it at every step. If a rule were added that did not shorten, the assertion would fire on the first derivation that used it, rather than the program looping until somebody noticed.

That is the standard shape of a termination argument, and it is the one [Rewrite Systems] gives for the sorting rule: exhibit a quantity that strictly decreases and is bounded below.

Confluence fails, and that is the finding

The last block computes, for each input, every normal form reachable by applying the rules of each phase in any order, and compares that set with what the strategy produces.

For seven of the thirteen inputs the set has two members. So the rule set is not confluent: different orders reach different answers.

And the strategy picks one of them. That is what makes the system a function. Without the strategy the system would be a relation, associating some inputs with two outputs, and a grammar that associated one root with two words would be useless.

This is the precise sense in which Pāṇini's precedence rules are part of his specification rather than an optimisation, and it is the claim [Rewrite Systems] makes and this program demonstrates.

The phase boundary earns its place

The rule that deletes a b is in phase two. Run it in phase one and it would delete b's that phase-one rules still need to see: the rule whose left side is a b followed by an a could no longer fire, because there would be no b.

So the boundary is not decoration. It preserves the conditions that the earlier rules depend on, which is exactly what sūtra 8.2.1 does for the tripādī, as [Asiddhatva and the Tripādī: Ordering by Blocking] sets out.

Complexity and limitations

Time. Each step scans the rule set once, which is proportional to the number of rules, and then searches the string for one occurrence. The number of steps is bounded by the length of the input, since each step shortens it. So a derivation costs at most the length of the input times the size of the rule base.

munotes.in161

Śāstra Rule Precedence as a Deterministic Finite Rewrite System

Space. One string and the trail. The trail is the same length as the number of steps.

Limitation: the flags are declared, not derived. The tradition decides whether a rule is special, nitya or antaraṅga by argument about its content. This program takes the flags as given. A fuller implementation would compute specialness by comparing conditions, and would then have to face the cases where neither condition is a subset of the other, which the commentaries argue about.

Limitation: the leftmost occurrence is always chosen. Where a rule's left side occurs twice, the tradition's own treatment of which occurrence is operated on is a separate matter, and this program's choice is a simplification it declares.

Limitation: it is not Sanskrit. The alphabet is two letters and the rules are invented to exercise the four principles. What is demonstrated is the architecture, and the architecture is the thing MU's topic 8 names.

Quick revision

  • A deterministic finite rewrite system is a rule set plus a strategy that always chooses exactly one rule.
  • The strategy here is the tradition's cascade: apavāda, nitya, antaraṅga, paratva, with the set narrowed rather than abandoned when several rules share a flag.
  • Termination is proved by exhibiting a strictly decreasing bounded quantity, here the string's length, and asserted at every step.
  • The rule set is NOT confluent: seven of thirteen inputs reach two normal forms under free choice of order.
  • So the strategy is what makes the system compute a function, which is the sense in which Pāṇini's precedence rules are part of his specification.
  • The phase boundary preserves conditions the earlier rules depend on, as 8.2.1 does.
  • Limitations: flags declared not derived, leftmost occurrence always chosen, and the alphabet is not Sanskrit.

Test yourself

1. What makes a rewrite system deterministic, and why is that not the same as confluent?

A strategy that selects exactly one applicable rule for every string. Confluence is a property of the rule set, that different orders reach the same answer. A non-confluent set with a total strategy is still deterministic, and the strategy then decides the answer.

2. Prove that the system in this chapter terminates.

Every rule's right side is shorter than its left, so each application strictly reduces the length of the string. Length is a non-negative whole number and cannot decrease for ever, so every derivation halts.

3. Why does the program report which principle decided each step?

Because an engine can implement four branches and exercise only two, leaving two untested. Reporting the deciding principle makes an unexercised branch visible, and the output block shows that all four decide at least once.

4. What would go wrong if the deleting rule were moved into phase one?

munotes.in162

Śāstra Rule Precedence as a Deterministic Finite Rewrite System

It would delete the symbol that another phase-one rule needs in its condition, so that rule could no longer fire. The phase boundary preserves the conditions the earlier rules depend on, which is the function sūtra 8.2.1 performs for the tripādī.

Contents This chapter on its own page

munotes.in163

Chapter Forty-Nine

Parsing Algorithms

Syllabus topic Module 1, "Parsing algorithms"

In one line

Generating is making strings from rules; parsing is taking a string somebody gives you and working out which rules made it.

In the wording you can write in an examination: parsing is the process of determining whether a given string belongs to the language defined by a grammar and, if it does, of recovering the derivation or the structure by which it does. A top-down parser starts from the start symbol and tries to match the input; a bottom-up parser starts from the input and tries to reduce it to the start symbol.

Why parsing is a different problem from generating

Generating has no wrong answers. Apply the rules however you like and whatever you produce is in the language by definition.

Parsing has exactly one question and it can be answered wrongly. Given this string, is it in the language, and by what structure? The string is not under your control and neither is the order in which its parts arrive.

And a grammar can be easy to generate from and hard to parse with. That is the sentence to remember, and it is why this chapter is in a paper about Pāṇini: the Aṣṭādhyāyī is built to derive a word from a root, which is generation, and using it in reverse to analyse a word is a much harder problem.

The two directions

Top down. Begin with the start symbol. Ask which of its productions could match the input. Try one; if it fails, back up and try another. The method is called recursive descent when written as one procedure per non-terminal.

Bottom up. Begin with the input. Look for a group of symbols matching the right-hand side of some production and replace it by that production's left-hand side. Keep reducing until the start symbol is left, or until nothing can be reduced. The usual implementation is shift-reduce, with a stack.

Top downBottom up
Starts fromthe start symbolthe input
Natural implementationone procedure per non-terminala stack and a table
Easy to write by handyesno
Handles left recursionno, it loops for everyes
Error messagesgood, because it knows what it expectedpoorer, because it knows only that nothing reduced
Used inhand-written parsers, many compilersgenerated parsers

The left recursion row is the one that bites. A production of the form "A rewrites to A something" sends a top-down parser into an infinite regress, because to parse an A it first tries to parse an A. Bottom-up parsing has no such problem, which is the main reason generated parsers are usually bottom-up.

A recursive descent parser, run

"""Recursive descent over the small grammar of the formal grammars chapter."""

TOKENS_NP = {"rama", "sita"}
TOKENS_V = {"sees", "calls"}
TOKENS_ADV = {"quickly"}

class Parser:
    def __init__(self, words):
        self.words = words
        self.i = 0
        self.trail = []

    def peek(self):
        return self.words[self.i] if self.i < len(self.words) else None

    def eat(self, expected_set, name):
        w = self.peek()
        if w in expected_set:
            self.i += 1
            self.trail.append("matched %-9s as %s" % (w, name))
            return w
        self.trail.append("expected %-4s at position %d, found %s" % (name, self.i, w))
        raise ValueError(name)

    def S(self):
        self.trail.append("enter S")
        np = self.eat(TOKENS_NP, "NP")
        vp = self.VP()
        if self.i != len(self.words):
            self.trail.append("trailing words left over at position %d" % self.i)
            raise ValueError("S")
        return ("S", ("NP", np), vp)

    def VP(self):
        self.trail.append("enter VP")
        v = self.eat(TOKENS_V, "V")
        if self.peek() in TOKENS_NP:
            obj = self.eat(TOKENS_NP, "NP")
            if self.peek() in TOKENS_ADV:
                adv = self.eat(TOKENS_ADV, "ADV")
                return ("VP", ("V", v), ("NP", obj), ("ADV", adv))
            return ("VP", ("V", v), ("NP", obj))
        return ("VP", ("V", v))


def show(tree, depth=0):
    if isinstance(tree, str):
        return "%s%s" % ("    " * depth, tree)
    head, rest = tree[0], tree[1:]
    lines = ["%s%s" % ("    " * depth, head)]
    for r in rest:
        lines.append(show(r, depth + 1))
    return "\n".join(lines)


CASES = [
    "rama sees",
    "rama sees sita",
    "rama sees sita quickly",
    "sita calls rama",
    "sees rama",
    "rama sita",
    "rama sees quickly",
    "rama sees sita sita",
    "rama",
    "",
]

print("%-28s %s" % ("input", "verdict"))
for case in CASES:
    p = Parser(case.split())
    try:
        p.S()
        print("%-28s accepted" % (case or "(empty)"))
    except ValueError:
        print("%-28s rejected: %s" % (case or "(empty)", p.trail[-1]))

print()
print("the parse tree for: rama sees sita quickly")
p = Parser("rama sees sita quickly".split())
print(show(p.S()))

print()
print("the trace for a rejection: rama sees quickly")
p = Parser("rama sees quickly".split())
try:
    p.S()
except ValueError:
    pass
for line in p.trail:
    print("  " + line)
munotes.in164

Parsing Algorithms

input                        verdict
rama sees                    accepted
rama sees sita               accepted
rama sees sita quickly       accepted
sita calls rama              accepted
sees rama                    rejected: expected NP   at position 0, found sees
rama sita                    rejected: expected V    at position 1, found sita
rama sees quickly            rejected: trailing words left over at position 2
rama sees sita sita          rejected: trailing words left over at position 3
rama                         rejected: expected V    at position 1, found None
(empty)                      rejected: expected NP   at position 0, found None

the parse tree for: rama sees sita quickly
S
    NP
        rama
    VP
        V
            sees
        NP
            sita
        ADV
            quickly

the trace for a rejection: rama sees quickly
  enter S
  matched rama      as NP
  enter VP
  matched sees      as V
  trailing words left over at position 2

Reading the output

Four accepted, six rejected, and each rejection carries the reason and the position.

The error messages are the advantage of top-down parsing. "Expected V at position 1, found sita" is possible because the procedure that failed knew what it was looking for. A bottom-up parser at the same point knows only that its stack cannot be reduced.

munotes.in165

Parsing Algorithms

The parse tree is the point of parsing. The tree for "rama sees sita quickly" says that "sees sita quickly" groups together as a verb phrase and "rama" stands outside it. Nothing in the string says so. Recovering that grouping is what a parser is for, and it is why a compiler parses before it does anything else.

And the trace shows the parser giving up honestly. On "rama sees quickly" it matched the subject and the verb, then found a word left over that no production could account for, and said so. It did not guess.

Ambiguity

A grammar is ambiguous if some string has two different parse trees.

Why it matters. Two trees mean two groupings, and two groupings can mean two things. In a programming language that is a disaster: an expression would compile two ways.

How it shows up. The classic case is an operator with no stated precedence, so that a sum of three terms can group either way. In natural language it is everywhere, which is one reason natural language parsing is hard in a way programming language parsing is not.

The grammar in this chapter is unambiguous, and the evidence is in [Formal Grammars: Alphabet, Rule, Derivation, Language]: its enumeration produced exactly as many derivations as distinct strings. Equal counts mean no string was reached twice, so no string has two derivations.

Why the Aṣṭādhyāyī is hard to parse with

Three reasons, all of which follow from earlier chapters.

It derives, and derivation is not reversible. A rule that replaces a vowel by its guṇa loses information: several different vowels have the same guṇa. Running the rule backwards gives several candidates and no way to choose among them from the output alone.

Markers are deleted. An anubandha decided which rules applied, and it is not in the finished word. So the analyser cannot see the information the derivation used.

The strategy is part of the meaning. To know which rule produced a given step you must know which rule the precedence principles would have selected, which depends on the state of the form before the step, which is what you are trying to recover.

The consequence, stated plainly. Using Pāṇini's grammar to analyse Sanskrit is not a matter of running it backwards. It requires search: propose an underlying form, derive it forwards, and see whether the result matches. That is how computational systems built on Pāṇinian principles actually work, and [Foundations of NLP, and What Pāṇinian Grammar Contributed] returns to it.

What parsing is NOT

It is not the same as validating. A parser can accept a string and the string can still be meaningless, as "sita calls rama quickly" being accepted shows. Syntax is not semantics.

munotes.in166

Parsing Algorithms

It is not always necessary. If you only need to know whether a string matches a pattern, a finite automaton will do and produces no tree. [Automata Theory Foundations: The Finite Automaton] is the cheaper tool.

It is not unique to the grammar. Many parsers exist for one grammar, and they differ in speed, in error messages and in which grammars they can handle at all.

Quick revision

  • Generating makes strings from rules; parsing recovers the rules from a string.
  • Top down starts from the start symbol, is easy to hand-write, gives good error messages, and cannot handle left recursion.
  • Bottom up starts from the input, uses a stack, handles left recursion, and gives poorer messages.
  • The parse tree is the point: it records groupings the string does not state.
  • Ambiguity is two trees for one string, and it matters because two groupings can mean two things.
  • The Aṣṭādhyāyī is hard to parse with because derivation loses information, markers are deleted, and the strategy depends on the state you are trying to recover. Analysis therefore needs search, not reversal.

Test yourself

1. Distinguish parsing from generating, and say which the Aṣṭādhyāyī does.

Generating produces strings by applying rules; parsing takes a given string and determines whether the rules could have produced it and by what structure. The Aṣṭādhyāyī generates: it derives a form from a root and a meaning.

2. Give two differences between top-down and bottom-up parsing.

Top down starts from the start symbol and bottom up from the input. Top down cannot handle a left-recursive production, since parsing an A begins by parsing an A, whereas bottom up can; and top down gives better error messages, because the failing procedure knows what it expected.

3. Why is a parse tree worth having when a yes or no would do?

Because the tree records which parts of the string group together, and the grouping is what the meaning depends on. A yes or no from a finite automaton gives no structure at all.

4. Give one reason the Aṣṭādhyāyī cannot simply be run backwards to analyse a word.

Its rules lose information: several vowels share one guṇa, and the markers that decided which rules applied are deleted before the word appears. So a backward step has several candidates and no way to choose, and analysis must instead propose an underlying form and derive forwards to check it.

Contents This chapter on its own page

munotes.in167

Chapter Fifty

Foundations of NLP, and What Pāṇinian Grammar Contributed

Syllabus topic Module 1, "Foundations of NLP"

In one line

Natural language processing begins by breaking text into pieces and working out the shape of each word, and that second job is the one Pāṇini's grammar is about.

In the wording you can write in an examination: natural language processing is the computational analysis and generation of human language. Its foundational tasks include tokenisation, dividing text into units; morphological analysis, determining the root and affixes of each word; and parsing, determining the structure of a sentence. Morphological analysis is standardly implemented with finite-state transducers, which map between a surface form and an analysis, and it is at this level that Pāṇinian grammar is computationally relevant.

The layers of the problem

An answer on the foundations of NLP should be able to name the layers and say what each does.

LayerThe jobThe unit
tokenisationsplit the text into words or pieces of wordsa token
morphologyfind the root and the affixes of each tokena morpheme
syntaxfind the structure of the sentencea parse tree
semanticsfind what the sentence meansa representation of meaning
pragmaticsfind what the speaker meant by it herea discourse interpretation

The layers are not independent and pretending otherwise is the standard beginner's mistake. Tokenising Sanskrit requires knowing where words end, which requires undoing the sandhi that joined them, which is morphology. So the first layer needs the second, and a pipeline that runs them in order will fail on exactly the cases the language makes hard.

Why tokenisation is hard for Sanskrit and easy for English

In English, spaces do most of the work. A tokeniser has to decide about apostrophes and hyphens and little else.

In Sanskrit, adjacent words are phonologically joined. The rules of sandhi change the end of one word and the beginning of the next, and the written form gives a single continuous string. So finding the word boundaries means undoing sandhi, and undoing sandhi means running the grammar's rules backwards.

Which is the problem [Parsing Algorithms] describes. The rules lose information, so there are several candidates for what was joined, and the segmentation has to be searched rather than read off.

This is a real and current computational problem, and it is the clearest place where Pāṇini's grammar bears on a task somebody is actually being paid to solve.

Morphology, and the finite-state transducer

What a transducer is. A finite automaton that, besides accepting or rejecting, writes an output. Each transition reads a symbol and writes a symbol, so the machine maps one string to another.

Why it is the right tool for morphology. A morphological rule is a mapping between a surface form and an analysis. A transducer computes exactly such a mapping, and, crucially, it can be inverted: the same machine run the other way maps analyses to surface forms.

munotes.in168

Foundations of NLP, and What Pāṇinian Grammar Contributed

Why that invertibility matters. One machine serves both analysis and generation. Write the rules once and you get both directions, which is the property a plain rewrite system does not have.

And this is the honest comparison with Pāṇini. His rules are a mapping from an underlying form to a surface form, stated with contexts and applied in an order. That is what a cascade of transducers is. The differences are real: a transducer is invertible and his rules are not, and a transducer has no precedence cascade.

A finite-state transducerPāṇini's rules
Mapssurface form to analysis, and backunderlying form to surface form
Invertibleyes, by constructionno, information is lost
Contextsexpressed in the statesexpressed by case endings, 1.1.66 and 1.1.67
Order of rulesa cascade of machines, composedfour precedence principles
Conflict between rulesresolved when the machines are composedresolved by the precedence cascade
What it is foranalysis and generation alikegeneration

The claims about influence, sorted

This is the section that matters for honesty, and it is organised by what kind of claim each statement is.

Attributable to the text, and this book has read the evidence

The Aṣṭādhyāyī states rules with contexts, in a fixed order, with a declared policy for conflicts, and separates rules from annotated data. Every part of that is in earlier chapters with the sūtra that supports it.

It treats a language as defined by the strings its rules derive. That is the tradition's criterion of correctness, and it is the modern definition of a language generated by a grammar.

A resemblance, worth stating as a resemblance

Pāṇini's architecture resembles a cascade of finite-state transducers more than it resembles a phrase-structure grammar. That is a judgement about similarity, supported by the table above, and it is not a historical claim.

The six kinds of sūtra resemble a typed rule system. Also a resemblance, and a useful one, as [The Six Kinds of Sūtra, in Vasu's Own Words] sets out.

Claims this book does not make, and why

That Pāṇini influenced the development of formal language theory. This is asserted often. It is a historical claim about what particular twentieth-century researchers read and when, and it requires evidence this book has not examined. So it is not made here.

That a named modern formalism was derived from Pāṇini. Same reason.

That Pāṇini anticipated the Chomsky hierarchy. [The Chomsky Hierarchy, and Where the Aṣṭādhyāyī Sits] explains why even placing the Aṣṭādhyāyī in the hierarchy would be a claim about a formalisation rather than about the text.

What to write in an examination. If a question invites you to discuss Pāṇini's contribution to NLP, the strong answer describes the architecture, names the modern object it resembles, and says which claims are about the text and which are about influence. An answer that separates the two is better than one that asserts the larger claim.

munotes.in169

Foundations of NLP, and What Pāṇinian Grammar Contributed

Worked example: what a Pāṇinian analyser actually does

Because "run the grammar backwards" is not available, systems built on these principles work in the other direction.

Step one. Take the surface form. Propose a set of candidate underlying forms: a root from the root list, plus affixes that could plausibly be present.

Step two. For each candidate, derive forwards using the rules, with the precedence cascade deciding at each conflict.

Step three. Keep the candidates whose derivation produces the surface form exactly.

Step four. If more than one candidate survives, the form is ambiguous, and something outside the morphology has to choose.

This is generate and test, and it is the standard shape of a solution when a transformation is not invertible. Notice that it uses the grammar exactly as written, in its own direction, which is why a faithful implementation of the rules is worth having even though the task is analysis.

Quick revision

  • Layers: tokenisation, morphology, syntax, semantics, pragmatics. They are not independent, and Sanskrit proves it, since tokenising requires undoing sandhi.
  • A finite-state transducer reads and writes, so it maps strings to strings, and it is invertible, which is why it serves analysis and generation from one description.
  • Pāṇini's rules resemble a cascade of transducers more than a phrase-structure grammar: contexts, ordering, one direction.
  • Differences: a transducer is invertible and his rules lose information; a transducer has no precedence cascade.
  • Attributable to the text: rules with contexts, in an order, with a conflict policy, and rules separated from annotated data.
  • Not claimed here: any historical influence on formal language theory, because that needs evidence this book has not examined.
  • A Pāṇinian analyser works by generate and test, not by reversal.

Test yourself

1. Name the layers of NLP and give one case where two of them cannot be separated.

Tokenisation, morphology, syntax, semantics, pragmatics. Tokenising Sanskrit cannot be separated from morphology, because adjacent words are joined by sandhi and the boundaries can be found only by undoing rules that belong to morphology.

2. What is a finite-state transducer and why is it the standard tool for morphology?

A finite automaton that writes an output symbol at each transition, so it maps one string to another. It is the standard tool because a morphological rule is such a mapping and because a transducer is invertible, so one description serves both analysis and generation.

munotes.in170

Foundations of NLP, and What Pāṇinian Grammar Contributed

3. Give two respects in which Pāṇini's rules differ from a transducer cascade.

His rules are not invertible, since operations such as guṇa lose information and markers are deleted. And they carry a precedence cascade for resolving conflicts, which a transducer cascade handles instead by the order in which the machines are composed.

4. A question asks you to assess Pāṇini's contribution to natural language processing. What should your answer separate, and why?

It should separate what is attributable to the text, namely rules with stated contexts applied in a stated order with a declared conflict policy over annotated data, from resemblances to modern formalisms, from historical claims about influence. The last require evidence about what particular researchers read, which is a different kind of claim and needs a different kind of support.

Contents This chapter on its own page

munotes.in171

Chapter Fifty-One

Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine

Syllabus topic Module 1, "Algorithm Specification (Pseudo-code)", "minimum 10 test cases", "Rule-based generative structure", "Conflict resolution mechanisms"

In one line

A rule base, a precedence resolver and a trace: sixty lines of Python that derive a form the way the grammar does, and say which principle decided.

In the wording you can write in an examination: a rule-based grammar engine stores rules as data, selects the rules whose conditions the current form satisfies, resolves any conflict among them by a stated precedence policy, applies the winner, and repeats. The Aṣṭādhyāyī's own architecture supplies all four parts: its rules, its rule types, its four precedence principles, and its phase boundary.

Problem statement, in MU's own form

IKS concept as CS concept: Pāṇini's Aṣṭādhyāyī as a rule-based grammar engine.

Statement. Implement a subset of the Aṣṭādhyāyī as a rule base with typed rules, a precedence resolver applying apavāda, nitya, antaraṅga and paratva in order, and a derivation trace that names the rule applied and the principle that selected it. Demonstrate that the same general rule produces different results according to the markers carried by the affix.

Conceptual mapping table

Classical elementComputer science element
a sūtraa rule, stored as data with its conditions and its operation
the six kinds of sūtraa type field on each rule
a vidhi rulea production, the only kind that transforms the form
a paribhāṣāa meta-rule, consulted while reading another rule
a niṣedha such as 1.1.5a guard that removes a candidate
apavāda, nitya, antaraṅga, paratvaa conflict resolution cascade in a fixed order
8.2.1's asiddhatvaa phase boundary
an anubandhaa tag on the input, read by rules, erased before output
a derivationa trace of applied rules

Algorithm specification, in pseudo-code

ALGORITHM Derive(stem, affix, affixKind, markers)

INPUT a stem, an affix, the affix's kind, and the affix's markers

OUTPUT the derived form, and the rule and principle that produced it

candidates <- every rule whose conditions this form satisfies

if candidates is empty then return stem + affix

winner, reason <- Resolve(candidates)

if winner is a prohibition then return stem + affix with the reason

return Apply(winner, stem, affix) with the reason

ALGORITHM Resolve(candidates)

if exactly one candidate then return it

if exactly one candidate is SPECIAL then return it, "apavada"

if exactly one candidate is NITYA then return it, "nitya"

if exactly one candidate is ANTARANGA then return it, "antaranga"

return the candidate whose address is LATEST, "1.4.2 vipratisedha"

Two design points are worth defending in a viva. The order inside Resolve is the tradition's order and reversing it changes answers, which is why it is written as a cascade of single-winner tests rather than as a score. And a prohibition returns the form unchanged rather than raising an error, because in the grammar a blocked operation is a normal outcome and not a failure.

munotes.in172

Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine

Working code

IK = ["i", "ii", "u", "uu", "ri", "rii", "li"]
GUNA = {"i": "e", "ii": "e", "u": "o", "uu": "o", "ri": "ar", "rii": "ar", "li": "al"}

RULES = [
    # sutra, kind, specific?, nitya?, inner?, phase, what it does
    ("1.1.5",  "nisedha", True,  False, True,  1, "block guna when the affix carries a marker"),
    ("7.3.84", "vidhi",   False, True,  True,  1, "guna of the stem's ik vowel"),
    ("8.2.1",  "boundary", False, False, False, 2, "phase two: invisible to phase one"),
]

def applicable(stem, affix_kind, markers):
    out = []
    if affix_kind in ("sarvadhatuka", "ardhadhatuka"):
        if any(stem.endswith(v) for v in IK):
            out.append("7.3.84")
        if markers:
            out.append("1.1.5")
    return out

def resolve(candidates):
    """The four principles, in order, returning the winner and the reason."""
    if len(candidates) == 1:
        return candidates[0], "only one rule applies"
    info = {r[0]: r for r in RULES}
    specific = [c for c in candidates if info[c][2]]
    if len(specific) == 1:
        return specific[0], "apavada: a special rule sets aside the general one"
    nitya = [c for c in candidates if info[c][3]]
    if len(nitya) == 1:
        return nitya[0], "nitya: the rule that would apply either way goes first"
    inner = [c for c in candidates if info[c][4]]
    if len(inner) == 1:
        return inner[0], "antaranga: the inner rule prevails"
    later = max(candidates, key=lambda s: tuple(int(x) for x in s.split(".")))
    return later, "1.4.2 vipratisedha: the later rule prevails"

def derive(stem, affix, affix_kind, markers=()):
    cands = applicable(stem, affix_kind, markers)
    if not cands:
        return stem + affix, "no rule applies", "none"
    winner, why = resolve(cands)
    if winner == "1.1.5":
        return stem + affix, why, winner
    for v in sorted(IK, key=len, reverse=True):
        if stem.endswith(v):
            return stem[:-len(v)] + GUNA[v] + affix, why, winner
    return stem + affix, why, winner

CASES = [
    ("bhu",   "ti",  "sarvadhatuka", (),     "bhoti"),
    ("ji",    "ti",  "sarvadhatuka", (),     "jeti"),
    ("kri",   "ti",  "sarvadhatuka", (),     "karti"),
    ("bhuu",  "ti",  "sarvadhatuka", (),     "bhoti"),
    ("patha", "ti",  "sarvadhatuka", (),     "pathati"),
    ("ji",    "ta",  "ardhadhatuka", ("k",), "jita"),
    ("bhu",   "ta",  "ardhadhatuka", ("ng",),"bhuta"),
    ("ji",    "tum", "ardhadhatuka", (),     "jetum"),
    ("ji",    "sya", "other",        (),     "jisya"),
    ("ji",    "",    "other",        (),     "ji"),
]

print("%-8s %-6s %-14s %-8s %-9s %-7s %s" % ("stem", "affix", "affix kind", "markers", "result", "expect", "decided by"))
ok = 0
for stem, affix, kind, markers, expect in CASES:
    got, why, winner = derive(stem, affix, kind, markers)
    ok += got == expect
    print("%-8s %-6s %-14s %-8s %-9s %-7s %s"
          % (stem, affix or "(none)", kind, ",".join(markers) or "(none)", got, expect,
             winner if winner != "none" else "-"))
print()
print("%d of %d cases give the expected result" % (ok, len(CASES)))
print()
print("the reason, for the two cases where two rules applied")
for stem, affix, kind, markers, expect in CASES:
    cands = applicable(stem, kind, markers)
    if len(cands) > 1:
        w, why = resolve(cands)
        print("  %s + %s: %s applied against %s, and %s won by %s"
              % (stem, affix, cands[0], cands[1], w, why.split(":")[0]))
munotes.in173

Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine

stem     affix  affix kind     markers  result    expect  decided by
bhu      ti     sarvadhatuka   (none)   bhoti     bhoti   7.3.84
ji       ti     sarvadhatuka   (none)   jeti      jeti    7.3.84
kri      ti     sarvadhatuka   (none)   karti     karti   7.3.84
bhuu     ti     sarvadhatuka   (none)   bhoti     bhoti   7.3.84
patha    ti     sarvadhatuka   (none)   pathati   pathati -
ji       ta     ardhadhatuka   k        jita      jita    1.1.5
bhu      ta     ardhadhatuka   ng       bhuta     bhuta   1.1.5
ji       tum    ardhadhatuka   (none)   jetum     jetum   7.3.84
ji       sya    other          (none)   jisya     jisya   -
ji       (none) other          (none)   ji        ji      -

10 of 10 cases give the expected result

the reason, for the two cases where two rules applied
  ji + ta: 7.3.84 applied against 1.1.5, and 1.1.5 won by apavada
  bhu + ta: 7.3.84 applied against 1.1.5, and 1.1.5 won by apavada

Reading the output

Ten cases, ten expected results. Five where guṇa applies, two where the prohibition blocks it, and three where no rule applies at all.

The "decided by" column is the part that matters. For the two blocked cases, two rules were applicable at once: 7.3.84, which would gunate, and 1.1.5, which forbids it. The resolver reports that 1.1.5 won by apavāda, because its condition, an affix carrying a marker, is narrower than 7.3.84's condition, an affix of two named kinds.

And the last block prints the conflict explicitly. An engine that resolves a conflict silently is an engine you cannot debug, and an internal assessment that cannot show why it produced an answer will not survive a viva.

The limits of this engine, stated

This is MU's heading seven, and it is the heading that separates a good submission from a demonstration.

It stops before the word is finished. The guṇa substitution gives a stem, and the Sanskrit word needs one more rule, 6.1.78, which is cited and not implemented. The engine's output for the root bhu with a present ending is a stage in a derivation, not a word.

The rule base is three rules out of about four thousand. It is chosen to exhibit the architecture: a vidhi, a paribhāṣā scoping it, a niṣedha blocking it, and the phase boundary. It is not a grammar of Sanskrit and nothing here suggests it is.

The markers are given as input. In the grammar they come from the Dhātupāṭha and from the affix lists. A fuller implementation would read them from a data file, which is exactly the separation of rules from annotated data that [Pāṇini's Aṣṭādhyāyī, and the Text We Are Reading] describes.

Two of the four principles are never exercised by these ten cases. Nitya and antaraṅga are implemented and no case reaches them, because with three rules no pair is separated by those tests. The resolver in [Śāstra Rule Precedence as a Deterministic Finite Rewrite System] has cases for all four, and an honest submission says which of its branches its tests actually cover.

munotes.in174

Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine

Transliteration is plain ASCII. The engine works on strings like "bhu" and "ji", not on Devanagari, so nothing in it depends on a script.

Complexity

Per derivation: the rule base is scanned once to collect candidates, so the work is proportional to the number of rules; the resolver then makes at most four passes over the candidates. With a rule base of four thousand, a real implementation would index the rules by their trigger conditions rather than scanning, which turns the scan into a lookup.

The resolver is constant time in the number of rules and linear in the number of candidates, which is small.

Quick revision

  • Four parts of a rule-based engine: a rule base as data, candidate selection, conflict resolution, application, repeated.
  • The Aṣṭādhyāyī supplies all four: its rules, its rule types, its four precedence principles, its phase boundary.
  • The resolver is a cascade of single-winner tests in the tradition's order; reversing the order changes answers.
  • A prohibition returns the form unchanged; a blocked operation is a normal outcome.
  • The ten cases: five guṇa, two blocked by 1.1.5 winning on apavāda, three with no applicable rule.
  • Limits: it stops before 6.1.78; three rules of four thousand; markers supplied as input; nitya and antaraṅga implemented but not exercised.

Test yourself

1. Give the four parts of a rule-based engine and the classical element each corresponds to.

A rule base as data, corresponding to the sūtras; candidate selection, corresponding to a rule's stated conditions; conflict resolution, corresponding to the four precedence principles; and application, corresponding to the vidhi operation.

2. Why is the resolver written as a cascade rather than as a score?

Because the tradition's principles are applied in a fixed order and each is decisive when it separates the candidates. A score would allow a lower principle to outweigh a higher one, which would give wrong answers whenever a special rule met a later one.

3. Explain why 1.1.5 beats 7.3.84 in the engine's output.

Both rules apply, so the cascade is entered. 1.1.5's condition, an affix carrying an indicatory marker, is a subset of 7.3.84's condition, an affix of two named kinds, so 1.1.5 is special and wins by apavāda before any other test is reached.

4. Name two limitations of this engine that belong in its own limitations section.

It stops at the guṇa substitution and does not apply 6.1.78, so its output is a stage in a derivation rather than a Sanskrit word. And two of the four precedence branches, nitya and antaraṅga, are implemented but are not exercised by any of its ten test cases.

Contents This chapter on its own page

munotes.in175

Chapter Fifty-Two

Practice for Module I

Syllabus topic Module 1, "Conceptual Mapping Table", "Complexity & Limitations"

In one line

Q.1 gives you four questions on Module 1 and asks for two, so every answer is worth five marks: a definition, the mechanism, one example, one limit.

Nothing new is taught in this chapter. It is Module I compressed into the shapes the examination actually uses.

The mapping table for Module 1

This is the table to have in your head walking into the hall, because Q.3 draws on both modules and every cross-module answer starts here.

Classical elementWhereModern concept
śāstra, a rule-governed disciplinesūtra, vṛtti, bhāṣya, vārttikaa formal system with a specification and its documentation
the sūtra methodbrevity forced by an oral mediumcompression with an external schema
pramāṇa, three admitted meanspratyakṣa, anumāna, āgamadeclared sources of evidence, and input validation
laghu and gurutwo values in one positionone bit
prastāraChandaḥśāstra 8.20 to 8.23exhaustive enumeration of a binary space
the next-row rule8.22 as Halāyudha works itthe successor function, and binary increment
naṣṭa8.24, 8.25index to element, a computed address
uddiṣṭa8.26, 8.27element to index, Horner evaluation
saṅkhyādvir ardham and its groupbinary exponentiation, square and multiply
Meru-prastāraHalāyudha's staircasePascal's triangle, and bottom-up dynamic programming
lagakriyāthe Meru's rowsbinomial coefficients
the six kinds of sūtraVasu's introductiona typed rule system
pratyāhāra1.1.71, over the fourteen Śivasūtrasa two-character name for a set, a character class
anubandha1.1.5 and the root listsa type tag, erased before output
paribhāṣā1.1.3a meta-rule supplying a default argument
the four precedence principlesmaxims, and 1.4.2a conflict resolution cascade
asiddhatva8.2.1, 6.4.22a phase boundary, as between compiler passes
the grammar as a wholethe Aṣṭādhyāyīan ordered string rewriting system with a strategy

The costs, in one place

MU requires a complexity analysis in the implementation, and a question can ask for one.

OperationCost
prastāra for n syllablesproportional to n times 2 to the power n, which is the size of the output
naṣṭaproportional to n, independent of the row number
uddiṣṭaproportional to n
saṅkhyā by squaringproportional to log n
one Meru entry by the closed formproportional to k
a whole Meru row by tabulationproportional to n squared, with space for one row
the same by naive recursionproportional to the answer itself, which is exponential
one derivation in the rule enginethe size of the rule base, times the length of the input

Worked answers at five marks, twelve of them

1. Define śāstra and state what makes a body of knowledge formal.

A śāstra is a systematised discipline: a body of rule-governed knowledge with defined terms, rules whose interaction is itself governed by stated principles, and a declared account of the admissible means of knowledge, transmitted as aphorisms with their own commentary. What makes it formal is that an answer inside the field can be derived by applying stated rules to stated facts, without asking the author what was meant. The four publication layers, sūtra, vṛtti, bhāṣya and vārttika, correspond to a specification, a reference implementation, a commentary and an errata list.

munotes.in176

Practice for Module I

2. Why is the sūtra method so brief, and what does the brevity cost?

The medium is human memory and recitation, so every syllable is a permanent cost, every ambiguity is a permanent risk, and a correction cannot be circulated; the reader's years of training, by contrast, cost the text nothing. The rational design pushes everything possible out of the text and into the commentary and the reader. The costs are three: no random access, because words carry over from earlier sūtras; no redundancy, so damage changes the rule rather than obscuring it; and no error detection, which the tradition answered socially, by parallel lineages and by the vārttika layer.

3. State the three pramāṇas and Charaka's rule about using them.

Pratyakṣa, perception; anumāna, inference; āgama or śabda, the instructive assertion of a reliable person. Charaka's rule is that all three should be used in examining a disease, and that the diagnosis then becomes faultless, because, in his own words, by only one of these means of knowledge, knowledge does not arise of everything that should be known. Testimony comes first in order of use, not in authority, because without prior instruction one does not know what to observe or what to infer.

4. Explain the prastāra rule and apply it to L L G.

Begin with a row of all heavy syllables. From any row, find the first heavy syllable from the left, make it light, make every syllable to its left heavy, and leave every syllable to its right unchanged; stop when the row is all light. In L L G the first heavy syllable is the third, so it becomes light and the two to its left are reset to heavy, giving G G L. The clause students drop is the reset, which turns L L G into L L L and breaks the table.

5. In what sense is the prastāra binary, and in what sense is it not?

Read laghu as one, guru as zero, and the leftmost syllable as the units place; then row r of the table, counting from zero, is the pattern whose value is r, and the next-row rule is exactly the rule for adding one to a binary number. That has been checked for every row of every metre up to fourteen syllables. What it is not: the text nowhere states a positional number system, nowhere treats a pattern as a number, and nowhere performs arithmetic on one. The order agrees with binary counting; the notion of a base does not appear.

munotes.in177

Practice for Module I

6. State naṣṭa and uddiṣṭa and say what makes them index retrieval mechanisms.

Naṣṭa: halve the row number, writing a light syllable; if the number is odd, add one before halving and write a heavy syllable; write from the left and stop when the syllable count is reached. Uddiṣṭa: start at one, take the syllables from the last to the first, doubling each time, and subtract one at every heavy syllable. Each costs one step per syllable and neither needs the table, so they give access to a row by its number and to a number by its row in time independent of the size of the table. They are exact inverses, checked over every row up to sixteen syllables.

7. Explain Piṅgala's saṅkhyā rule and compute the number of forms of the gāyatrī.

Halve the syllable count repeatedly, recording a halving where it was exact and a subtraction where one had to come off first; then, reading the record forwards, start at one and square at a halving, double at a subtraction. For six syllables: six halves to three, three reduces to two, two halves to one, one reduces to zero, so the marks are double, square, double, square. Start at one: two, four, eight, sixty-four. Four operations instead of six multiplications, and Halāyudha's own commentary states the figure as sixty-four and then gives 128, 256, 512, 1024, 2048 and 4096 for seven to twelve syllables.

8. Describe the Meru-prastāra and state what its rows mean.

One is placed in the topmost cell; each end cell of every row takes the single cell above it; each interior cell takes the two cells above it, combined. Row n, counting the top as row zero, gives the number of patterns of n syllables with exactly k light syllables, for k from zero to n. Halāyudha reads the rows off as 1 1, then 1 2 1, then 1 3 3 1. The proof of the addition rule splits the patterns on the last syllable: if it is light the rest has one fewer light syllable, if heavy the rest has the same number, and the two cases are exclusive and exhaustive.

9. Name the six kinds of sūtra and say which of them changes a form.

Saṃjñā, a definition; paribhāṣā, a rule about how rules are read; vidhi, a general operation; niyama, a restriction on a rule's scope; adhikāra, a governing rule understood in those that follow; and atideśa, extension to an uncovered case by analogy. Only vidhi performs an operation on a form. Five of the six manage rules rather than changing words, which is the ratio worth noticing: a rule system large enough to be interesting spends most of its rules on the management of rules.

munotes.in178

Practice for Module I

10. Explain the pratyāhāra device and decode two examples.

The fourteen Śivasūtras list the sounds of Sanskrit in a fixed order, each ending in a marker consonant that is not one of the sounds. By sūtra 1.1.71 a name is formed from the first sound of a span together with the marker that closes it, and denotes everything between them. So ac, from a to the marker closing the fourth aphorism, is all the vowels; hal, from ha to the final marker, is all the consonants. Vasu states that only forty-two such names are in practical use. The achievement is the order: a set can be named only if its members are contiguous in it.

11. Set out the conflict resolution mechanisms of the Aṣṭādhyāyī.

Three kinds. Structurally, asiddhatva: sūtra 8.2.1 makes the last three quarters of book eight non-existent for all earlier rules, so certain conflicts cannot arise. By prohibition: rules such as 1.1.5 forbid an operation outright, removing one candidate. And by precedence, in a fixed order of four: apavāda, a special rule sets aside a general one; nitya, a rule that would apply either way goes first; antaraṅga, a rule whose conditions lie closer to the stem prevails and the outer rule is treated as not having taken effect; and last, paratva, sūtra 1.4.2, under which the later rule prevails. Only paratva and the prohibitions are numbered sūtras; the other three principles are maxims of the tradition.

12. Is the Aṣṭādhyāyī a context-free grammar? Justify your answer.

No, for three reasons. Its operative rules state conditions on surrounding material, and sūtras 1.1.66 and 1.1.67 exist precisely to let any rule express a left or right context by a case ending, so the rules are context-sensitive in the strict sense. Its mechanisms refer to whether other rules have applied, which no context-free production can express: 8.2.1 and 6.4.22 do exactly that. And its rules operate on sounds, stems, affixes and deleted markers rather than on non-terminal categories, so there is no level at which a category is expanded into its parts. Positively, it is an ordered string rewriting system with a fixed application strategy, and its nearest modern relatives are the finite-state transducers of computational morphology.

The one-line facts a viva asks for

  • Chandaḥśāstra: eight chapters; chapter 8 carries the pratyayas; the commentary quoted is Halāyudha's Mṛtasañjīvanī.
  • Six pratyayas: prastāra, naṣṭa, uddiṣṭa, lagakriyā, saṅkhyā, adhvayoga.
  • Aṣṭādhyāyī: eight books of four quarters; a citation is book, quarter, position.
  • Fourteen Śivasūtras; forty-two pratyāhāras in practical use.
  • 1.1.1 defines vṛddhi. 1.1.3 is the paribhāṣā restricting guṇa and vṛddhi to the ik vowels. 1.1.5 is the prohibition. 1.1.71 forms a pratyāhāra. 1.4.2 is vipratiṣedha. 7.3.84 causes guṇa. 8.2.1 is pūrvatrāsiddham.
  • Vasu's translation, Allahabad, 1891 to 1898, is the source quoted, and it is out of copyright.
  • 2 to the power n patterns of n syllables; row sums of the Meru are those same powers; the running total is 2 to the power n plus one, minus two.
  • Gāyatrī 6 syllables and 64 forms, up to jagatī 12 syllables and 4096 forms, all stated in the commentary.
munotes.in179

Practice for Module I

Quick revision

  • Q.1 sets four alternatives on Module 1 and asks for two, so an answer is worth five marks: a definition, the mechanism, one example, one limit.
  • The mapping table at the head of this chapter is the one to carry into the hall, because every cross-module answer begins by pairing a classical element with a modern concept.
  • The three costs a complexity question is likeliest to want: the prastāra is proportional to n times 2 to the power n; naṣṭa and uddiṣṭa are proportional to n; saṅkhyā by squaring is proportional to log n.
  • The sūtra numbers to know: 1.1.1, 1.1.3, 1.1.5, 1.1.71, 1.4.2, 7.3.84, 8.2.1 for Pāṇini, and 8.20 to 8.27 for Piṅgala.
  • The figures Halāyudha's commentary itself states: 64, 128, 256, 512, 1024, 2048, 4096 for six to twelve syllables.

Test yourself

1. A question offers four alternatives and asks for two. What does that tell you about how much to write?

Each alternative is worth five marks in a one-hour paper, so a complete answer is a definition, the mechanism, one worked example and one limit, in roughly the length of the twelve answers above. It is not an essay, and it is not a list of headings either.

2. Which single table would you take into the hall, and why?

The mapping table at the head of this chapter, because Q.3 is set across both modules and every cross-module answer begins by pairing a classical element with the modern concept it corresponds to.

3. Give the three costs a complexity question is most likely to ask for.

The prastāra is proportional to n times 2 to the power n, which is the size of its output and therefore cannot be improved. Naṣṭa and uddiṣṭa are proportional to n and independent of the row number. Saṅkhyā by squaring is proportional to log n.

Contents This chapter on its own page

munotes.in180

Module II

Nyaya logic as structured inference, Ayurvedic classification as a rule-based expert system, and Arthasastra cryptography as a secure communication model

munotes.in

Chapter Fifty-Three

Nyāya: The School, the Sūtra and the Sixteen Categories

Syllabus topic Module 2, "Nyāya Logic as Structured Inference and Knowledge Systems", "Study of classical Indian logical structures"

In one line

Nyāya is the Indian school of logic and debate, and its first sentence lists the sixteen things it is going to talk about.

In the wording you can write in an examination: Nyāya is one of the six classical systems of Indian philosophy, concerned with the means of valid knowledge and with the conduct of reasoned debate. Its foundational text is the Nyāya Sūtra attributed to Gotama or Akṣapāda, in five books of two chapters each. Its first sūtra enumerates sixteen categories, of which the first two concern knowledge and its objects and the remaining fourteen concern reasoning, discussion and its failures.

Why Module II opens here

MU's Module 2 has three blocks: Nyāya logic, Ayurvedic classification, and Arthaśāstra cryptography. The first is the largest, and four of her six printed labels under it are items in this one list.

Her "five-member syllogism" is category seven, members. Her "anumāna structure" is part of category one, the means of right knowledge. Her "debate methodology and validation" is categories ten to sixteen. Her "padārtha" is a related enumeration from a sister school.

So the list below is the map of the whole block.

The text

The Nyāya Sūtra is attributed to Gotama, also called Gautama or Akṣapāda. Vidyabhusana's introduction records that it is divided into five books, each containing two chapters called āhnikas or diurnal portions, and that the author is believed to have finished his work in ten lectures corresponding to them.

The translation quoted throughout this book is Satis Chandra Vidyabhusana's, published in 1913 in the Sacred Books of the Hindus series. It gives each sūtra in English with the commentator Vātsyāyana's illustrations, which is exactly the material a chapter on the structure of an argument needs.

On the date, Vidyabhusana argues from the occurrence of Nyāya technical terms in an early Pali work to a date around the middle of the sixth century before the Common Era. This book reports his argument and does not adopt a date, because nothing here depends on one and the question has its own literature.

The provision

Nyāya Sūtra 1.1.1, in Vidyabhusana's translation:

Supreme felicity is attained by the knowledge about the true nature of sixteen categories, viz., means of right knowledge (pramāṇa), object of right knowledge (prameya), doubt (saṃśaya), purpose (prayojana), familiar instance (dṛṣṭānta), established tenet (siddhānta), members (avayava), confutation (tarka), ascertainment (nirṇaya), discussion (vāda), wrangling (jalpa), cavil (vitaṇḍā), fallacy (hetvābhāsa), quibble (chala), futility (jāti), and occasion for rebuke (nigrahasthāna).

The sixteen, grouped

Read as a list of sixteen it is hard to hold. Read as four groups it is a design.

GroupCategoriesWhat the group is for
What knowledge ispramāṇa, prameyathe means of knowing, and what there is to know
What starts an enquirysaṃśaya, prayojana, dṛṣṭānta, siddhāntadoubt, purpose, a familiar instance, and what is already agreed
How an argument is madeavayava, tarka, nirṇayathe members of the argument, reasoning by absurdity, and settlement
How arguments are conducted and lostvāda, jalpa, vitaṇḍā, hetvābhāsa, chala, jāti, nigrahasthānathree kinds of dispute, and four kinds of failure
munotes.in181

Nyāya: The School, the Sūtra and the Sixteen Categories

Seven of the sixteen are about failure. That proportion is the most informative thing about the list, and it is worth saying in an answer: a system that spends nearly half its vocabulary on how reasoning goes wrong is a system built for use against an opponent, not for private contemplation.

The commentary's own account of the structure

Vidyabhusana records, on the same sūtra, that knowledge of the sixteen categories means true knowledge of their enunciation, definition and critical examination, and that Book I treats of enunciation and definition while the remaining four books are reserved for critical examination.

So the text's own architecture is: name them, define them, then attack the definitions. That is the three-stage order [Śāstra: What Makes a Body of Knowledge Formal] identifies as the mark of a formal discipline, stated by the text about itself.

Worked example: where each group appears in a real dispute

Take a dispute about whether a piece of code is correct.

Doubt. The reviewer is not sure. That is saṃśaya, and it is what makes the enquiry worth having: nobody debates what nobody doubts.

Purpose. The reason for settling it: the release is tomorrow. That is prayojana.

Established tenet. What both parties already accept: the language's semantics, the test framework's behaviour. That is siddhānta, and a dispute in which nothing is shared cannot be conducted at all.

Familiar instance. A case both parties agree about, used to anchor a general claim. That is dṛṣṭānta, and it becomes the example member of the argument.

Members. The argument itself, in five parts. [The Five-Member Syllogism].

Failures. The reviewer's reason turns out to support the opposite conclusion too, which is the erratic fallacy; or one party wins by taking a word in a sense the other did not mean, which is quibble. [Hetvābhāsa: The Five Fallacies of the Reason] and [Chala, Jāti and Nigrahasthāna: How a Debate Is Lost].

What Nyāya is NOT

It is not a system of formal logic in the modern sense. There is no symbolic notation, no calculus, and no proof of soundness or completeness. What it has is a fixed form for an argument and a catalogue of ways an argument fails.

It is not only about debate. The first two categories are epistemology and the eighth is a method of reasoning. Debate is the setting, not the whole subject.

munotes.in182

Nyāya: The School, the Sūtra and the Sixteen Categories

It is not the same as Vaiśeṣika. Vaiśeṣika is the sister school, concerned with the categories of what exists rather than with the means of knowing. Nyāya's category two, prameya, and Vaiśeṣika's padārtha overlap, and the traditions eventually merged, but they are two enumerations and [Padārtha: The Categories of What Exists] keeps them apart.

It is not committed to the religious frame of its first sūtra. The sūtra says the knowledge conduces to supreme felicity. Nothing in the logical machinery depends on that, and the machinery is what this paper examines.

Quick revision

  • Nyāya: the Indian school of logic and debate. Its text is the Nyāya Sūtra of Gotama, five books of two chapters each.
  • Sūtra 1.1.1 lists sixteen categories, and this book quotes Vidyabhusana's 1913 translation throughout.
  • The sixteen in four groups: what knowledge is; what starts an enquiry; how an argument is made; how arguments are conducted and lost.
  • Seven of the sixteen concern failure, which tells you the system is built for use against an opponent.
  • The text's own order is enunciation, definition, critical examination; Book I does the first two and Books II to V the third.
  • Not a symbolic logic, not only about debate, and not the same enumeration as Vaiśeṣika's padārtha.

Test yourself

1. What are the sixteen categories for, and which sūtra states them?

They are the subject matter of Nyāya: the means and objects of knowledge, the circumstances that start an enquiry, the parts of an argument, and the kinds of dispute and of failure. Nyāya Sūtra 1.1.1 enumerates them.

2. Group the sixteen and say what each group does.

Pramāṇa and prameya, what knowledge is and what it is of. Saṃśaya, prayojana, dṛṣṭānta and siddhānta, what starts and anchors an enquiry. Avayava, tarka and nirṇaya, how an argument is made and settled. Vāda, jalpa, vitaṇḍā, hetvābhāsa, chala, jāti and nigrahasthāna, how disputes are conducted and lost.

3. What does it tell you that seven of the sixteen categories concern failure?

That the system is designed for argument against an opponent rather than for private reasoning: nearly half its vocabulary exists so that a bad move can be named and the debate decided.

4. State the text's own three-stage order and where each stage is carried out.

Enunciation, definition, critical examination. Book I gives the enunciation and the definitions, and the remaining four books are reserved for the critical examination of them.

Contents This chapter on its own page

munotes.in183

Chapter Fifty-Four

The Four Pramāṇas of Nyāya, and Why Charaka Has Three

Syllabus topic Module 2, "Anumāna (inference) structure", "Study of classical Indian logical structures"

In one line

Nyāya admits four means of knowledge where Charaka admits three, and the extra one is comparison.

In the wording you can write in an examination: Nyāya Sūtra 1.1.3 names four pramāṇas: pratyakṣa, perception; anumāna, inference; upamāna, comparison; and śabda, verbal testimony. Charaka and the Sāṃkhya school admit only three, omitting comparison, and MU's printed list for Module 1 is that three-member set. Schools differ over how many independent means there are, and the disagreement is about whether a candidate reduces to one already on the list.

The provision

Nyāya Sūtra 1.1.3, in Vidyabhusana's translation:

Perception, inference, comparison and word (verbal testimony) are the means of right knowledge.

And his own note, which is the most useful single paragraph in the book for this topic:

The Cārvākas admit only one means of right knowledge, viz., perception; the Vaiśeṣikas and Bauddhas admit two, viz., perception and inference; the Sāṃkhyas admit three, viz., perception, inference and verbal testimony; while the Naiyāyikas, whose fundamental work is the Nyāya-sūtra, admit four, viz., perception, inference, verbal testimony and comparison. The Prabhākaras admit a fifth means of right knowledge called presumption; the Bhāṭṭas and Vedāntins admit a sixth, viz., non-existence; and the Paurāṇikas recognise a seventh and eighth, named probability and rumour.

The eight candidates, in one table

Number admittedSchoolThe list
oneCārvākaperception only
twoVaiśeṣika, Bauddhaperception, inference
threeSāṃkhya, and the Āyurvedic traditionperception, inference, verbal testimony
fourNyāyathose three, and comparison
fivePrābhākaraand presumption
sixBhāṭṭa, Vedāntaand non-existence
eightPaurāṇikaand probability, and rumour

Read the column of lists downward. Every school's list is a prefix of the next, so the disagreement is not about which channels exist. It is about which are independent, that is, which cannot be reduced to a channel already admitted.

That is the shape of the argument and it is the answer to "why do schools disagree?". A Vaiśeṣika does not deny that people learn from testimony; he says testimony is a species of inference, because you infer the speaker's reliability. A Naiyāyika replies that the inference would need a general connection nobody has established.

The extra one: upamāna

Nyāya Sūtra 1.1.6, in Vidyabhusana's translation:

Comparison is the knowledge of a thing through its similarity to another thing previously well known.

And the Sūtra's own illustration:

A man hearing from a forester that a bos gavaeus is like a cow resorts to a forest where he sees an animal like a cow. Having recollected what he heard he institutes a comparison, by which he arrives at the conviction that the animal which he sees is bos gavaeus.

Worked. What is going on in that example, stated as steps, because the example is doing precise work.

munotes.in184

The Four Pramāṇas of Nyāya, and Why Charaka Has Three

  1. He receives a statement of similarity: this unknown animal resembles a known one.
  2. He later perceives an animal resembling the known one.
  3. He recalls the statement.
  4. He concludes that the name in the statement applies to the animal in front of him.

The knowledge acquired is the application of a NAME. Not that the animal exists, which was perception; not a general fact, which would be inference. It is the linking of a word to a thing by way of a similarity.

Vidyabhusana records that some hold comparison is not an independent means, which is exactly the dispute above.

Why it matters to a computer science student

The four pramāṇas map onto four ways a system can acquire a fact, and the fourth is the one usually missing from the map.

PramāṇaHow a system gets a fact
pratyakṣameasurement, from a sensor or a direct query
anumānaderivation, from stored rules and other facts
śabdaimport, from a source the system is configured to trust
upamānamatching a new instance to a known one by similarity, and thereby labelling it

The fourth row is nearest-neighbour classification, and it is a real and distinct operation. A system that matches a new record to the most similar known record and adopts its label is doing neither measurement nor rule-based derivation nor import. It is doing what the forester's listener did.

The comparison holds and it should not be pushed. The classical account is about applying a name on the strength of a stated similarity; a classifier is about assigning a label on the strength of a computed distance. The resemblance is in the structure, not in the mechanism, and [Multi-Class Classification, and How It Is Scored] treats the modern side on its own terms.

Which list does this paper use?

MU prints three in Module 1 and Nyāya's four belong to Module 2. So an answer should say which it is using.

For Module 1's pramāṇa question, use the three: pratyakṣa, anumāna, āgama. That is the Sāṃkhya and Āyurvedic set and it is what her label prints.

For a Module 2 question on Nyāya, use the four, and name the extra one as upamāna.

For a Q.3 cross-module question, the difference between the two lists IS an answer: the same paper prints one school's list in one module and another school's in the other, and being able to say which is which and why they differ is precisely the kind of comparison Q.3 asks for.

What the disagreement is NOT

It is not a disagreement about the facts. Every school agrees that people learn from testimony and from comparison. The question is whether those are separate channels or special cases.

munotes.in185

The Four Pramāṇas of Nyāya, and Why Charaka Has Three

It is not settled by counting. More means of knowledge is not a better epistemology. The Cārvāka position, that only perception counts, is a serious one and its consequence is that no general claim can be established at all.

It is not merely scholastic. The number of admitted channels determines what can be proved, and therefore what a debate can be won with. [Debate Methodology as a Validation Protocol] is where that bites.

Quick revision

  • Nyāya Sūtra 1.1.3: perception, inference, comparison, verbal testimony. Four.
  • Charaka and Sāṃkhya admit three, omitting comparison, and MU's Module 1 list is that three.
  • Vidyabhusana enumerates schools admitting one, two, three, four, five, six and eight, and each list is a prefix of the next, so the dispute is about independence, not existence.
  • Upamāna, 1.1.6: knowledge of a thing through its similarity to something previously well known. The bos gavaeus example: a statement of similarity, a later perception, a recollection, and the application of a name.
  • The modern analogue of upamāna is labelling by similarity to a known instance, which is a distinct operation from measurement, derivation and import.
  • Say which list you are using: three for Module 1, four for Nyāya, and the difference itself is a Q.3 answer.

Test yourself

1. Name the four pramāṇas of Nyāya and the one that Charaka omits.

Pratyakṣa, anumāna, upamāna and śabda. Charaka omits upamāna, comparison, and admits three.

2. Define upamāna and give the Sūtra's own example.

Knowledge of a thing through its similarity to another thing previously well known. A man is told by a forester that a bos gavaeus resembles a cow; later he sees such an animal, recalls what he was told, and concludes that the name applies to it.

3. What exactly are the schools disagreeing about when they admit different numbers?

Not whether people acquire knowledge in those ways, which all agree on, but whether each way is independent or reduces to one already admitted. A Vaiśeṣika treats testimony as a species of inference; a Naiyāyika denies that the required general connection has been established.

4. Give the modern counterpart of each of the four, and say where the comparison should stop.

Measurement, derivation from rules, import from a trusted source, and labelling a new instance by similarity to a known one. The comparison should stop at the mechanism: upamāna applies a name on a stated similarity, while a classifier assigns a label on a computed distance.

Contents This chapter on its own page

munotes.in186

Chapter Fifty-Five

Anumāna: Vyāpti, and the Three Kinds of Inference

Syllabus topic Module 2, "Anumāna (inference) structure"

In one line

An inference works only if the mark you observed never occurs without the thing you concluded, and that "never without" is called vyāpti.

In the wording you can write in an examination: anumāna proceeds from a mark, the hetu, observed in a subject, the pakṣa, to a conclusion, the sādhya, by way of vyāpti, the invariable concomitance of the mark with the conclusion. Vyāpti is established by observation of positive instances, sapakṣa, where both are present, and negative instances, vipakṣa, where both are absent, and it fails if a single case is found in which the mark occurs without the conclusion.

The five terms

Learn these five and every chapter of the Nyāya block becomes readable. They are introduced here once.

TermPlain meaningIn the standard example
pakṣathe subject, the thing you are talking aboutthis hill
sādhyawhat is to be established of itthat it has fire
hetuthe mark, the reason, what you observedthat it has smoke
vyāptithe invariable connection between mark and conclusionwherever there is smoke there is fire
dṛṣṭāntathe familiar instance that anchors the connectiona kitchen

The structure is always the same. The subject has the mark; the mark never occurs without the conclusion; therefore the subject has the conclusion.

Vyāpti, and what it is not

Vyāpti is not correlation. It is not "smoke and fire usually go together". It is "there is no smoke without fire", a universal negative claim about every case there has ever been or will be.

Its direction matters. Smoke is invariably accompanied by fire; fire is not invariably accompanied by smoke. So you may infer fire from smoke and not smoke from fire. The mark must be the narrower term. Getting the direction backwards is the commonest error in constructing an inference and it is why the tradition insists on the wording.

And it is established by observation, which is why it can be wrong. You look for positive instances, cases where both occur, and you look for negative instances, cases where neither occurs. What settles it is finding no case of the mark without the conclusion. A single counterexample destroys it.

Sapakṣa and vipakṣa

Two more terms, and they are how the connection is tested.

Sapakṣa, the similar instance: a case, other than the subject, where the conclusion is known to hold. If the mark is also found there, that supports the connection.

Vipakṣa, the dissimilar instance: a case where the conclusion is known not to hold. If the mark is absent there too, that supports the connection. If the mark is present there, the connection is destroyed.

The two together are exactly the positive and negative form of the same test, and the Nyāya Sūtra states them as two kinds of example. Vidyabhusana's 1.1.36 and 1.1.37 give a homogeneous example, which is known to possess the property and implies the property is invariably contained in the reason, and a heterogeneous example, which is known to be devoid of the property and implies the absence of the property is invariably rejected in the reason.

munotes.in187

Anumāna: Vyāpti, and the Three Kinds of Inference

In modern terms: confirming cases and disconfirming cases, and the disconfirming ones are the informative ones. A rule that survives a search for counterexamples is worth more than a rule with many confirming instances, and the Nyāya apparatus has a slot for both.

The three kinds of inference

Nyāya Sūtra 1.1.5 gives them, with Vidyabhusana's own examples.

KindDirectionHis example
pūrvavat, a priorifrom cause to effectseeing clouds, one infers that there will be rain
śeṣavat, a posteriorifrom effect to causeseeing a river swollen, one infers that there was rain
sāmānyatodṛṣṭa, commonly seenfrom one thing to another regularly found with itseeing a horned beast, one infers it has a tail; or seeing smoke on a hill, fire

And the disagreement, which a question may ask for: Vātsyāyana reads the third as "not commonly seen" and takes it to be inference to something that could never be perceived at all, as when one infers a soul from the qualities of affection and aversion. On his reading the third kind is the boldest rather than the weakest.

Worked example: an inference built from the terms

The situation. A service is returning errors and you cannot see inside it.

pakṣa. This service.

sādhya. Its connection pool is exhausted.

hetu. Every request is timing out after exactly the pool's wait limit.

vyāpti. Wherever requests time out at exactly the pool wait limit, the pool is exhausted.

sapakṣa. A service last month whose pool was known to be exhausted and which showed exactly this timing.

vipakṣa. A service whose pool was known to be healthy, which timed out at varied intervals and never at the wait limit.

Now test the vyāpti. Is there any case where requests time out at exactly the wait limit and the pool is not exhausted? If a misconfigured client happens to use the same timeout, there is, and the connection fails. That is the erratic fallacy of [Hetvābhāsa: The Five Fallacies of the Reason] arriving in a modern setting, and the discipline of looking for the case that breaks the connection is the whole value of the apparatus.

How vyāpti is established, and the honest problem

The tradition's own answer is repeated observation together with the absence of a counterexample, supported by tarka, reasoning that shows the contrary to be absurd.

munotes.in188

Anumāna: Vyāpti, and the Three Kinds of Inference

And that is not a proof. No number of observed cases establishes a universal claim, and the tradition knows it: tarka exists precisely because observation alone is not enough.

Saying so is part of a good answer. The problem is the one modern philosophy calls induction, and the Nyāya apparatus does not solve it either. What it does is make the claim explicit and attackable: the vyāpti is stated as its own member of the argument, so an opponent can go after it directly rather than having to guess what the arguer assumed.

That is a real achievement and it is the one to name. The general premise is not left implicit. Compare an everyday argument, where the general claim is usually unstated and therefore unexamined.

What anumāna is NOT

It is not deduction from arbitrary premises. The general premise must be a vyāpti, an invariable concomitance established by observation, not any universal statement someone cares to assert.

It is not defeated by an exception. It is destroyed by one. A vyāpti with an exception is not a weaker vyāpti; it is not a vyāpti, and an argument resting on it commits the erratic fallacy.

It is not the same as the five-member form. The five members are how an inference is stated for an audience. The inference itself is the three terms and the connection, and [The Five-Member Syllogism] is about the presentation.

Quick revision

  • Five terms: pakṣa the subject, sādhya what is to be established, hetu the mark, vyāpti the invariable connection, dṛṣṭānta the familiar instance.
  • Vyāpti is a universal negative: no mark without the conclusion. The mark must be the narrower term, so smoke gives fire and not the reverse.
  • Sapakṣa is a case where the conclusion holds and the mark is looked for; vipakṣa is a case where it does not and the mark must be absent. Positive and negative instances.
  • Three kinds, 1.1.5: cause to effect, effect to cause, and commonly seen. Vātsyāyana reads the third as "not commonly seen" and makes it the boldest.
  • Vyāpti is established by observation plus the absence of a counterexample plus tarka, and that is not a proof. The achievement is that the general premise is stated and therefore attackable.

Test yourself

1. Define vyāpti and say why its direction matters.

The invariable concomitance of the mark with what is to be established: there is no case of the mark without it. The direction matters because the mark must be the narrower term. Smoke is never without fire, so fire may be inferred from smoke; fire is often without smoke, so smoke may not be inferred from fire.

2. What are sapakṣa and vipakṣa, and which is the more informative?

munotes.in189

Anumāna: Vyāpti, and the Three Kinds of Inference

A sapakṣa is a case other than the subject where the conclusion holds, used to check that the mark is present there too. A vipakṣa is a case where the conclusion does not hold, where the mark must be absent. The vipakṣa is more informative, because a single vipakṣa carrying the mark destroys the connection.

3. Name the three kinds of inference with an example of each.

Cause to effect: clouds, so rain. Effect to cause: a swollen river, so there was rain. Commonly seen: a horned animal, so it has a tail. Vātsyāyana reads the third instead as inference to what is never perceived, such as a soul from its qualities.

4. What problem does the Nyāya account of vyāpti not solve, and what does it achieve instead?

It does not solve induction: no number of observations establishes a universal claim, and tarka is invoked precisely because observation is insufficient. What it achieves is making the general premise an explicit member of the argument, so that an opponent can attack it directly instead of guessing at an unstated assumption.

Contents This chapter on its own page

munotes.in190

Chapter Fifty-Six

The Five-Member Syllogism

Syllabus topic Module 2, "Five-member syllogism"

In one line

Nyāya requires an argument to be stated in five named parts, and an argument missing one of them has not been properly put.

In the wording you can write in an examination: the pañcāvayava or five-member syllogism is the form in which an inference must be stated for it to be assessed. Nyāya Sūtra 1.1.32 names the members as pratijñā, the proposition; hetu, the reason; udāharaṇa, the example; upanaya, the application; and nigamana, the conclusion. Each is defined in the sūtras immediately following, and each has a distinct job.

The provision

Nyāya Sūtra 1.1.32, in Vidyabhusana's translation:

The members (of a syllogism) are proposition, reason, example, application, and conclusion.

And the Sūtra's own illustration, printed with it:

1. Proposition. This hill is fiery.

2. Reason. Because it is smoky.

3. Example. Whatever is smoky is fiery, as a kitchen.

4. Application. So is this hill (smoky).

5. Conclusion. Therefore this hill is fiery.

The five members, each with its own sūtra

Pratijñā, the proposition. Nyāya Sūtra 1.1.33: "A proposition is the declaration of what is to be established." His illustration: sound is non-eternal.

Hetu, the reason. Nyāya Sūtra 1.1.34: "The reason is the means for establishing what is to be established through the homogeneous or affirmative character of the example." And 1.1.35 adds: "Likewise through heterogeneous or negative character."

Udāharaṇa, the example. Nyāya Sūtra 1.1.36 defines the homogeneous example as "a familiar instance which is known to possess the property to be established and which implies that this property is invariably contained in the reason given". Nyāya Sūtra 1.1.37 defines the heterogeneous example as "a familiar instance which is known to be devoid of the property to be established and which implies that the absence of this property is invariably rejected in the reason given".

Upanaya, the application. Nyāya Sūtra 1.1.38: "Application is a winding up, with reference to the example, of what is to be established as being so or not so." He records that it is of two kinds, affirmative, expressed by the word "so", and negative, expressed by "not so".

Nigamana, the conclusion. Nyāya Sūtra 1.1.39: "Conclusion is the re-stating of the proposition after the reason has been mentioned." He glosses it as the confirmation of the proposition after the reason and the example have been mentioned.

What each member is FOR

The definitions alone do not make the design visible. This table does.

MemberIts jobWhat is missing without it
pratijñāstates the claim, so everyone knows what is at stakethe arguer can shift ground
hetustates the evidence in the subjectthe claim is an assertion
udāharaṇastates the general connection AND anchors it in a casethe general claim is unstated and unattackable
upanayaasserts that this subject falls under the connectionthe general claim is left unapplied
nigamanastates the claim as now establishedit is not on the record that anything has been proved
munotes.in191

The Five-Member Syllogism

The third member is the one that earns the form. It states the vyāpti explicitly, as its own step, where an ordinary argument leaves the general premise implicit. Making the general premise a required part of the statement is the single best idea in Nyāya logic.

Worked: the affirmative and the negative forms

The Sūtra gives both, and a question may ask for either.

Affirmative, from 1.1.38:

Proposition: sound is non-eternal. Reason: because it is produced. Example: whatever is produced is non-eternal, as a pot. Affirmative application: so is sound (produced). Conclusion: therefore sound is non-eternal.

Negative, from the same sūtra:

Proposition: sound is not eternal. Reason: because it is produced. Example: whatever is eternal is not produced, as the soul. Negative application: sound is not so, that is, sound is not produced ... Conclusion: therefore sound is not eternal.

Note what the negative form does with the example. It states the connection contrapositively, and anchors it in a case where the property is absent. That is the vipakṣa of [Anumāna: Vyāpti, and the Three Kinds of Inference], serving as the example member.

One reading note. The text this book quotes prints the conclusion of the affirmative form, at 1.1.39, as "therefore sound is produced" where the proposition was "sound is non-eternal". That is a slip, in the scan or in the printing: the conclusion is defined in the same sūtra as the re-stating of the proposition, so it must restate the proposition. The book says so rather than reproducing it silently.

The five extra members

Vidyabhusana records, immediately after 1.1.32, that "some lay down five more members", and gives them. They are worth having because each names something the five do not do.

Extra memberWhat it addsHis illustration
jijñāsā, inquiry as to the propositionfixes the scope of the claimis this hill fiery in all its parts, or in a particular part?
saṃśaya, questioning the reasonchallenges the observationwhat you call smoke may be nothing but vapour
śakyaprāpti, capacity of the examplechallenges the vyāptiis smoke always a concomitant of fire? in a red-hot iron ball there is no smoke
prayojana, purposestates why the conclusion is wantedto determine whether the hill may be approached, avoided, or ignored
saṃśayavyudāsa, dispelling all questionsrecords that the challenges are answeredit is beyond question that the hill is smoky and that smoke invariably accompanies fire

Read them as a whole and the point is clear. The five extra members are the OPPONENT's moves, promoted into the form. Three of them are challenges, one fixes scope and one records resolution. The five-member form is a statement; the ten-member form is a dialogue.

munotes.in192

The Five-Member Syllogism

And the third extra member is the best of them. "Is smoke always a concomitant of fire? In a red-hot iron ball there is no smoke." That is a counterexample offered against the vyāpti, and it is the exact move [Anumāna: Vyāpti, and the Three Kinds of Inference] says destroys a connection.

What the five members are NOT

They are not five premises. The conclusion and the application restate what is already there. Logically the argument has two premises and a conclusion; the five members are a presentation.

They are not a proof procedure. Nothing in the form checks that the vyāpti is true. The form makes it visible so that it can be checked, which is a different and more modest claim.

They are not optional. Vidyabhusana records, in his note on the mistimed fallacy, that placing the members in the wrong order is itself an occasion for rebuke, called inopportune. So the order is part of the requirement.

Quick revision

  • Nyāya Sūtra 1.1.32: proposition, reason, example, application, conclusion.
  • Their sūtras: 1.1.33 to 1.1.39, with the example having an affirmative form at 1.1.36 and a negative form at 1.1.37.
  • The example member states the vyāpti explicitly and anchors it in a case, which is what an ordinary argument leaves implicit.
  • The affirmative form uses "so"; the negative form states the connection contrapositively and uses "not so".
  • Five extra members are recorded by some: inquiry, questioning the reason, capacity of the example, purpose, dispelling questions. They are the opponent's moves, promoted into the form.
  • The form is a presentation, not a proof procedure: it makes the general premise attackable, it does not verify it.

Test yourself

1. Name the five members with the Sūtra's own illustration.

Proposition, this hill is fiery. Reason, because it is smoky. Example, whatever is smoky is fiery, as a kitchen. Application, so is this hill, smoky. Conclusion, therefore this hill is fiery.

2. Which member does the real work, and why?

The example. It states the invariable connection as an explicit step and anchors it in a familiar instance, where an ordinary argument leaves the general premise unstated and therefore unexamined.

3. Give the negative form of the argument about sound, and say what the example member does in it.

Proposition, sound is not eternal. Reason, because it is produced. Example, whatever is eternal is not produced, as the soul. Negative application, sound is not so. Conclusion, therefore sound is not eternal. The example states the connection contrapositively and anchors it in a case where the property is absent.

munotes.in193

The Five-Member Syllogism

4. What do the five extra members add, taken as a group?

They add the opponent's part: a question about the scope of the claim, a challenge to the observation, a counterexample against the invariable connection, a statement of purpose, and a record that the challenges have been met. They turn a statement into a dialogue.

Contents This chapter on its own page

munotes.in194

Chapter Fifty-Seven

The Five Members Worked, Three Times

Syllabus topic Module 2, "Five-member syllogism"

In one line

Practice, not theory: the same five slots filled three times, once from the Sūtra, once from the Sūtra again, and once from a machine room.

Why fill the form three times

Because the difficulty is never understanding the definitions; it is producing the form on material you have not seen. A student who can recite the five members and cannot write them for a new claim has learned nothing usable.

So: one worked example you already know, one from the text, and one you have to build.

Worked one: the classical inference

The situation. You are at a distance and see smoke rising from a hill.

MemberStatement
pratijñāThis hill is fiery.
hetuBecause it is smoky.
udāharaṇaWhatever is smoky is fiery, as a kitchen.
upanayaSo is this hill: it is smoky.
nigamanaTherefore this hill is fiery.

Check the parts. The subject is the hill. The mark is smoke, observed. The connection is stated, and anchored in a kitchen, which is a case where fire is known and smoke is found. The application brings the hill under the connection. The conclusion restates the claim as established.

Where an opponent would attack. Not the observation, which is hard to deny; the connection. Is there smoke without fire? That is the challenge Vidyabhusana records as one of the five extra members, and the red-hot iron ball is offered there as a case of fire without smoke, which is the converse and does not damage this connection.

Worked two: the Sūtra's own argument about sound

The situation. A standing dispute in Indian philosophy: is sound eternal?

MemberStatement
pratijñāSound is non-eternal.
hetuBecause it is produced.
udāharaṇaWhatever is produced is non-eternal, as a pot.
upanayaSo is sound: it is produced.
nigamanaTherefore sound is non-eternal.

Why this example is in the text and not just the hill. Because the mark, being produced, is not observed in the way smoke is. Nobody watches sound being produced in the sense of seeing an event; the claim that sound is produced is itself a philosophical position. So the argument's weak point is the fourth member, the application, and an opponent goes after that.

The lesson for a student. The five members do not tell you which member is weak. They give you five places to look, which is five more than an unstructured argument gives you.

Worked three: an argument from a machine room

The situation. A batch job that has run nightly for a year fails tonight. You have no access to the machine and one fact: the job's log stops after the line that opens its output file.

MemberStatement
pratijñāThis job failed because the disk is full.
hetuBecause its log stops after the line that opens its output file.
udāharaṇaWhatever stops after opening its output file has failed because the disk is full, as last March's export job.
upanayaSo is this job: its log stops after that line.
nigamanaTherefore this job failed because the disk is full.
munotes.in195

The Five Members Worked, Three Times

Now read the third member aloud. "Whatever stops after opening its output file has failed because the disk is full." That is plainly untrue. A job can stop there because the file system is read only, because a permission changed, because the process was killed, or because the network mount vanished.

So the argument fails, and it fails at the third member, in public. The form did not prevent the bad inference; it made the bad step impossible to hide. In the ordinary way of putting it, "the log stops after the open, so the disk must be full", the general claim is never said out loud and so is never examined.

That is the whole value of the five-member form for a working engineer, and it is worth stating in those words in an answer: the general premise is promoted from an assumption to a sentence.

The same argument, repaired

The repair is not to abandon the form but to fix the connection.

MemberStatement
pratijñāThis job failed because the disk is full.
hetuBecause its log stops after opening its output file AND the file system it writes to reports no free blocks.
udāharaṇaWhatever fails to write while its file system reports no free blocks has failed because the disk is full, as last March's export job.
upanayaSo is this job: it failed to write while the file system reported no free blocks.
nigamanaTherefore this job failed because the disk is full.

What changed. The mark was narrowed until the connection became defensible. That is the standard repair for an erratic reason and it is what [Hetvābhāsa: The Five Fallacies of the Reason] names.

And notice the cost. The repaired argument needs a second observation, which the original did not. A sound inference is usually more expensive than an unsound one, and the form is what makes the cost visible before you act on the conclusion.

A template you can fill in the hall

If a question gives you a claim and asks you to set it out in five members, use this.

pratijna <subject> is <what is to be established>.

hetu Because it is <mark>.

udaharana Whatever is <mark> is <what is to be established>, as <familiar instance>.

upanaya So is <subject>: it is <mark>.

nigamana Therefore <subject> is <what is to be established>.

munotes.in196

The Five Members Worked, Three Times

Three checks before you write it down. Is the mark narrower than the conclusion, so that the connection runs the right way? Is the familiar instance a case where the conclusion is known to hold independently? And could the mark occur without the conclusion, which is the question that decides whether the argument is any good.

Quick revision

  • Same five slots, any material: proposition, reason, example, application, conclusion.
  • The hill: smoke observed, connection anchored in a kitchen, attack falls on the connection.
  • Sound: the mark is itself contested, so the attack falls on the application.
  • The machine room: the connection is plainly false, and the form exposes it rather than hiding it.
  • The repair narrows the mark until the connection is defensible, and costs a second observation.
  • Three checks: is the mark narrower, is the instance independent, could the mark occur without the conclusion.

Test yourself

1. Set out in five members: this program has a memory leak, because its resident size grows monotonically over eight hours.

Proposition: this program has a memory leak. Reason: because its resident size grows monotonically over eight hours. Example: whatever grows monotonically in resident size over eight hours has a memory leak, as the log collector we fixed last year. Application: so is this program: its resident size grows monotonically over eight hours. Conclusion: therefore this program has a memory leak.

2. Attack the argument you just wrote, at the member where it is weakest.

At the example. A program with a large cache that has not yet reached its ceiling also grows monotonically for eight hours and has no leak, so the mark occurs without the conclusion and the connection is erratic.

3. Why does putting an argument in five members not make it a good argument?

Because nothing in the form checks that the connection is true. The form makes the general premise an explicit sentence, which is what allows it to be attacked; whether it survives the attack is a separate matter.

4. What does the repaired machine-room argument cost, and why is that worth knowing?

A second observation, that the file system reports no free blocks. It is worth knowing because it shows that a sound inference is usually more expensive than an unsound one, and the form makes that cost visible before you act on the conclusion.

Contents This chapter on its own page

munotes.in197

Chapter Fifty-Eight

Five Members Against Aristotle's Three

Syllabus topic Module 2, "Five-member syllogism", "Predicate logic"

In one line

Aristotle's syllogism has three statements and Nyāya's has five, and the two extra ones are the example and the application.

In the wording you can write in an examination: the Aristotelian categorical syllogism consists of a major premise, a minor premise and a conclusion, and is valid in virtue of its form alone. The Nyāya pañcāvayava adds to these a familiar instance attached to the general premise, and a separate step applying the general premise to the subject. The additional material is epistemic rather than logical: it concerns how the general premise was established, not whether the conclusion follows.

The two forms side by side

The Aristotelian form, in its standard first-figure shape.

Major premise All things that are smoky are fiery.

Minor premise This hill is smoky.

Conclusion This hill is fiery.

The Nyāya form.

pratijna This hill is fiery.

hetu Because it is smoky.

udaharana Whatever is smoky is fiery, as a kitchen.

upanaya So is this hill: it is smoky.

nigamana Therefore this hill is fiery.

Match them up. The udāharaṇa carries the major premise, the upanaya carries the minor premise, and the nigamana carries the conclusion. So the logical content of the two is the same, and Nyāya has two members left over: the pratijñā, which states the conclusion first, and the phrase "as a kitchen" inside the example.

The distinctions table

Aristotelian syllogismNyāya five-member form
Number of statementsthreefive
Conclusion statedat the end onlyat the start and again at the end
General premisea premise, asserteda member, asserted AND anchored in an instance
The instanceno slot for itrequired, as the familiar instance
Validitya property of the form alonethe form is a mode of presentation; soundness is a separate matter
What is being assessedwhether the conclusion followswhether the claim has been established
Settinga proofa debate before an audience
Failure is calledinvalidityhetvābhāsa, quibble, futility, occasion for rebuke

The row that matters is the sixth. Aristotle's question is whether the conclusion follows from the premises. Nyāya's question is whether the claim has been established, which includes where the premises came from. Those are different questions, and each form is well shaped for its own.

What the example member adds

This is the answer to the question a good student asks, which is why anyone would want two more statements.

It names a case where the general claim is known to hold. "As a kitchen" is not decoration. It identifies an instance the audience already accepts, in which both smoke and fire are present, so the general claim is not floating free.

It makes the general claim's evidence part of the argument. In Aristotle, where the major premise came from is somebody else's problem. In Nyāya it is a member of the argument, and an opponent may attack it there, which is what the extra member śakyaprāpti of [The Five-Member Syllogism] does with the red-hot iron ball.

munotes.in198

Five Members Against Aristotle's Three

And it is the ancestor of the modern practice of grounding a general rule in cases. A test case in a specification, a worked example in a standard, a precedent cited for a rule: all three are the udāharaṇa's job. A rule with no instance attached is a rule nobody has checked.

What the application member adds

It makes subsumption an explicit step. "So is this hill: it is smoky" asserts that this particular case falls under the general claim, and it is where an argument most often actually fails.

Why that is not obvious. In the Aristotelian form the minor premise does the same work, so the application looks like a repetition of the reason. It is not quite: the reason states the mark, and the application states that the mark brings the subject under the connection just stated. The difference is small in the hill example and large when the connection has conditions.

In modern terms it is the distinction between a fact and a match. A monitoring system records that the queue is long, which is the fact; the rule fires because that fact matches the rule's condition, which is the match. Confusing them is why an alert fires on the wrong service, and a form with separate slots for the two will not let you confuse them.

What this chapter does NOT claim

Not that either tradition influenced the other. The question of contact between Greek and Indian logic has a literature of its own, and nothing this book has read bears on it. The comparison here is structural.

Not that five members are better than three. They answer different questions. If you want to know whether a conclusion follows, three is enough and five is padding. If you want to know whether a claim has been established before an audience, three is not enough.

Not that Nyāya lacks a notion of validity. The fallacies of [Hetvābhāsa: The Five Fallacies of the Reason] include cases that are formal defects, notably the contradictory reason. What Nyāya lacks is a general theory of validity independent of the particular defects it enumerates.

Worked: a five-mark answer

Question. Compare the Nyāya five-member syllogism with the Aristotelian syllogism.

Answer. The Aristotelian syllogism has three statements, a major premise, a minor premise and a conclusion, and it is valid in virtue of its form alone. The Nyāya form has five: the proposition, the reason, the example, the application and the conclusion. Its example corresponds to the major premise, its application to the minor, and its conclusion to the conclusion, so the logical content is the same; the difference is that Nyāya states the claim at the start as well as the end, and that its example must carry a familiar instance in which the general connection is known to hold. That instance makes the evidence for the general premise part of the argument, so an opponent may attack it there. The two forms therefore answer different questions: Aristotle asks whether the conclusion follows, and Nyāya asks whether the claim has been established, which includes where the general premise came from.

munotes.in199

Five Members Against Aristotle's Three

Quick revision

  • Aristotle: three statements, validity from form alone. Nyāya: five, with the example carrying a familiar instance.
  • The correspondence: example is the major premise, application the minor, conclusion the conclusion. The extra material is the proposition stated first and the instance inside the example.
  • The example adds the evidence for the general premise, and makes it attackable. A rule with no instance attached is a rule nobody has checked.
  • The application makes subsumption an explicit step: the fact and the match are different things.
  • Different questions: does the conclusion follow, against has the claim been established.
  • No claim about influence in either direction, and no claim that five is better than three.

Test yourself

1. Map the five members onto the three statements of an Aristotelian syllogism.

The example carries the major premise, the application carries the minor premise, and the conclusion carries the conclusion. The proposition restates the conclusion at the start and the familiar instance inside the example has no Aristotelian counterpart.

2. What does the familiar instance contribute, and what modern practice does it correspond to?

It names a case where the general connection is known to hold, so the general premise is anchored and its evidence becomes part of the argument. Modern counterparts are a test case attached to a specification, a worked example in a standard, and a precedent cited for a rule.

3. Why is the application not merely a repetition of the reason?

The reason states the mark observed in the subject; the application asserts that the subject thereby falls under the general connection just stated. The difference is between a fact and a match, and it matters as soon as the connection carries conditions.

4. State the difference in what the two forms are assessing.

Aristotle's form assesses whether the conclusion follows from the premises. The Nyāya form assesses whether the claim has been established, which includes the question of where the general premise came from.

Contents This chapter on its own page

munotes.in200

Chapter Fifty-Nine

Hetvābhāsa: The Five Fallacies of the Reason

Syllabus topic Module 2, "Debate methodology and validation"

In one line

Five named ways a reason can look like a reason and not be one, each with its own test.

In the wording you can write in an examination: a hetvābhāsa is a semblance of a reason, that is, a mark offered in an inference which fails to establish the conclusion. Nyāya Sūtra 1.2.4 names five: savyabhicāra, the erratic; viruddha, the contradictory; prakaraṇasama, equal to the question; sādhyasama, the unproved; and kālātīta, the mistimed. Each names a distinct defect and each has a stated test.

The provision

Nyāya Sūtra 1.2.4, in Vidyabhusana's translation:

Fallacies of a reason are the erratic, the contradictory, the equal to the question, the unproved, and the mistimed.

The five, each with the Sūtra's own example

One: the erratic

Nyāya Sūtra 1.2.5. "The erratic is the reason which leads to more conclusions than one."

The Sūtra's example, given twice over. Proposition: sound is eternal. Erratic reason: because it is intangible. Example: whatever is intangible is eternal, as atoms. Conclusion: sound is eternal. And then: proposition: sound is non-eternal. Erratic reason: because it is intangible. Example: whatever is intangible is non-eternal, as intellect. Conclusion: sound is non-eternal.

The commentary's diagnosis, which is the definition of vyāpti restated. The reason or middle term is erratic when it is not pervaded by the major term, that is, when there is no universal connection between them. Intangible is pervaded neither by eternal nor by non-eternal.

The test. Can the same reason be used to reach the opposite conclusion? If yes, it establishes nothing.

The modern name. The evidence is not diagnostic. A symptom present in both the fault and the healthy case tells you nothing, and an alert built on it fires on everything.

Two: the contradictory

Nyāya Sūtra 1.2.6. "The contradictory is the reason which opposes what is to be established."

The Sūtra's example. Proposition: a pot is produced. Contradictory reason: because it is eternal. The commentary: the reason is contradictory because that which is eternal is never produced.

The test. Does the reason, if true, establish the opposite of the claim? If yes, the arguer has argued against themselves.

The modern name. The evidence supports the alternative hypothesis. It is the case where a test result that was supposed to confirm a diagnosis actually rules it out, and it is caught by asking what the reason would show if you had no view of the conclusion.

Three: equal to the question

Nyāya Sūtra 1.2.7. "Equal to the question is the reason which provokes the very question for the solution of which it was employed."

The Sūtra's example. Proposition: sound is non-eternal. Reason: because it is not possessed of the attribute of eternality. The commentary: non-eternal is the same as not possessed of the attribute of eternality, so in determining whether sound is non-eternal the reason given is that sound is non-eternal, which begs the question.

munotes.in201

Hetvābhāsa: The Five Fallacies of the Reason

The test. Restate the reason in the conclusion's own words. If it becomes the conclusion, there is no reason.

The modern name. Circularity, or begging the question. In a system it appears as a rule whose condition is its own conclusion under another name, which fires once and proves nothing.

Four: the unproved

Nyāya Sūtra 1.2.8. "The unproved is the reason which stands in need of proof in the same way as the proposition does."

The Sūtra's example. Proposition: shadow is a substance. Unproved reason: because it possesses motion. The commentary: unless it is actually proved that shadow possesses motion, we cannot accept it as the reason. It adds an alternative account, that the motion belongs to the person causing the obstruction of light.

The test. Is the reason itself in dispute? If it is, it cannot support anything until it is settled.

The modern name. An unverified premise. In an argument about a system it is the step "the cache is stale", asserted because it would explain the symptom, and never checked.

Five: the mistimed

Nyāya Sūtra 1.2.9. "The mistimed is the reason which is adduced when the time is past in which it might hold good."

The Sūtra's example. Proposition: sound is durable. Mistimed reason: because it is manifested by union, as a colour. The commentary works it: the colour of a jar is manifested when the jar comes into union with a lamp, but the colour existed before that union and continues after it; similarly a drum's sound is manifested on union with a rod. The reason is mistimed because the manifestation of sound does not take place at the time of the union but at a subsequent moment, whereas the colour's manifestation is simultaneous with it. Because the times of manifestation differ, the analogy is incomplete.

The test. Does the reason's supporting analogy hold at the moment it needs to? A reason whose timing is wrong is not a weaker reason; it does not apply.

The modern name. A temporal confound. In practice it is the error of reading a log line as the cause of an event that preceded it, or of attributing a slowdown to a deployment that landed afterwards.

And Vidyabhusana records a second reading: some take the mistimed to be a reason adduced in a wrong order among the five members, for instance stated before the proposition. He reports that Vātsyāyana rejects that reading, and notes separately that misordering the members is itself an occasion for rebuke, called inopportune.

munotes.in202

Hetvābhāsa: The Five Fallacies of the Reason

The five as a checklist

FallacyThe one-sentence testThe modern name
erraticcould this reason give the opposite conclusion too?the evidence is not diagnostic
contradictorydoes this reason, if true, prove the opposite?the evidence supports the alternative
equal to the questionis the reason the conclusion reworded?circularity
unprovedis the reason itself still in dispute?an unverified premise
mistimeddoes the reason hold at the moment it must?a temporal confound

Run them in that order. The first two are about the connection, the third is about the wording, the fourth is about the evidence, and the fifth is about the timing. They are cheap and they are exhaustive over a large class of bad arguments.

Worked example: five bad reasons for one claim

The claim. This deployment caused the outage.

Erratic. "Because the error rate rose." The error rate also rises on a traffic spike with no deployment, so the reason reaches both conclusions.

Contradictory. "Because the deployment reduced the number of database queries." If true, that would make the outage less likely, not more; the reason argues against the claim.

Equal to the question. "Because the outage began with the new version being live." That is the claim restated: to say the outage was caused by the deployment and to say it began when the deployment was live are the same assertion in different words.

Unproved. "Because the new build has a null check missing." Does it? Until somebody reads the diff, the reason needs establishing exactly as much as the claim does.

Mistimed. "Because the deployment introduced a slow query." The slow query was first logged forty minutes before the deployment completed, so at the moment the reason must hold it does not.

Five distinct defects, and every one of them is a sentence somebody has said in a real incident review.

What a hetvābhāsa is NOT

It is not a false conclusion. The conclusion may happen to be true. What has failed is the reason offered for it, and the tradition is careful about the difference.

It is not the same as a quibble. A quibble is a move made in bad faith about the meaning of words, which is category fourteen. A fallacy of the reason is a defect in the argument itself and may be entirely sincere.

It is not a complete theory of bad argument. Book V, Chapter I gives a long list of futilities and Book V, Chapter II a list of occasions for rebuke, and [Chala, Jāti and Nigrahasthāna: How a Debate Is Lost] covers them. And the two lists overlap by design: Nyāya Sūtra 5.2.25 provides that "the fallacies of a reason already explained do also furnish occasions for rebuke", so committing one of the five is itself a ground of defeat.

munotes.in203

Hetvābhāsa: The Five Fallacies of the Reason

Quick revision

  • Nyāya Sūtra 1.2.4 names five: erratic, contradictory, equal to the question, unproved, mistimed.
  • Erratic, 1.2.5: leads to more conclusions than one. The Sūtra proves sound both eternal and non-eternal from "intangible".
  • Contradictory, 1.2.6: opposes what is to be established. A pot is produced, because it is eternal.
  • Equal to the question, 1.2.7: the reason restates the claim. Sound is non-eternal because it lacks eternality.
  • Unproved, 1.2.8: the reason needs proof as much as the claim. Shadow is a substance because it moves.
  • Mistimed, 1.2.9: the reason does not hold at the moment it must. The drum's sound is manifested after the union, not during it.
  • A fallacy of the reason is not a false conclusion and is not a quibble.

Test yourself

1. Name the five fallacies with the test for each.

Erratic, could the reason give the opposite conclusion. Contradictory, does the reason prove the opposite. Equal to the question, is the reason the claim reworded. Unproved, is the reason itself in dispute. Mistimed, does the reason hold at the moment it must.

2. Give the Sūtra's own example of the erratic reason and say what the commentary diagnoses.

Sound is eternal because it is intangible, as atoms; and sound is non-eternal because it is intangible, as intellect. The commentary says the middle term is not pervaded by the major: there is no universal connection between intangible and either eternal or non-eternal.

3. Distinguish the contradictory from the erratic.

An erratic reason supports both the claim and its opposite, so it settles nothing. A contradictory reason supports only the opposite, so the arguer has argued against their own claim.

4. Why is a fallacy of the reason not the same as a false conclusion?

Because the conclusion may be true while the reason offered for it fails. What the fallacy condemns is the support, not the claim, and the tradition keeps the two apart.

Contents This chapter on its own page

munotes.in204

Chapter Sixty

Chala, Jāti and Nigrahasthāna: How a Debate Is Lost

Syllabus topic Module 2, "Debate methodology and validation"

In one line

Chala is winning on a word, jāti is winning on a bad analogy, and nigrahasthāna is the list of things that lose you the debate outright.

In the wording you can write in an examination: chala, quibble, is the opposition offered to a proposition by assuming an alternative meaning of a word; jāti, futility, is an objection founded on a mere similarity or dissimilarity; and nigrahasthāna, occasion for rebuke, is any of the enumerated failures on which a party is declared defeated. Together with hetvābhāsa they are the four categories of failure among Nyāya's sixteen.

Chala, the quibble

Nyāya Sūtra 1.2.10. "Quibble is the opposition offered to a proposition by the assumption of an alternative meaning."

Nyāya Sūtra 1.2.11. "It is of three kinds, viz., quibble in respect of a term, quibble in respect of a genus, and quibble in respect of a metaphor."

Quibble in respect of a term

Nyāya Sūtra 1.2.12. "Quibble in respect of a term consists in wilfully taking the term in a sense other than that intended by a speaker who has happened to use it ambiguously."

Worked, from the Sūtra itself. A speaker says: this boy is nava-kambala, possessed of a new blanket. A quibbler replies: this boy is not certainly nava-kambala, possessed of nine blankets, for he has only one blanket. The commentary notes that the word nava is ambiguous, used by the speaker in the sense of "new" and wilfully taken by the quibbler in the sense of "nine".

The modern form. A specification says the system must handle "a large number of users", and an objector argues it fails because it cannot handle a billion. The ambiguity was real and the objection exploits it rather than resolving it.

Quibble in respect of a genus

Nyāya Sūtra 1.2.13. "Quibble in respect of a genus consists in asserting the impossibility of a thing which is really possible, on the ground that it belongs to a certain genus which is very wide."

The Sūtra's example. A speaker says: this Brahmaṇa is possessed of learning and conduct. An objector replies that this is impossible, for how can it be inferred that this person is possessed of learning and conduct because he is a Brahmaṇa, since there are little boys who are Brahmaṇas and not possessed of learning and conduct. The commentary notes that the objector knows perfectly well what was meant.

What the move actually does. It treats a statement about a particular as though it were a claim about the whole class, and refutes the class claim instead. That is the strawman, described precisely, and the naming is better than the English name because it identifies the mechanism: substituting a wider genus for the particular.

munotes.in205

Chala, Jāti and Nigrahasthāna: How a Debate Is Lost

Quibble in respect of a metaphor

Nyāya Sūtra 1.2.14 covers this third kind: taking a figurative expression literally in order to reject it.

The modern form. "The service is starving." "Services do not eat."

Jāti, futility

Nyāya Sūtra 1.2.18. "Futility consists in offering objections founded on mere similarity or dissimilarity."

What that means in practice. The opponent answers an argument by pointing out that the subject resembles, or fails to resemble, the example in some respect that has nothing to do with the connection being argued.

A worked case. Argument: sound is non-eternal, because it is produced, as a pot. Futile objection: but sound is not like a pot, since a pot has shape and sound has none, so sound is not non-eternal. The dissimilarity is real and it is irrelevant: the connection argued was between being produced and being non-eternal, and shape is no part of it.

Book V, Chapter I of the Nyāya Sūtra is given over entirely to futility, and it runs to forty-three sūtras, each treating a named variety with its own worked example: balancing the co-presence, balancing the mutual absence, and many more. This chapter does not reproduce them and an answer is not expected to. What is expected is the definition and the recognition that the list is long and highly specific, which tells you the practice was real: nobody catalogues forty-three sūtras' worth of a move nobody makes.

Nigrahasthāna, occasion for rebuke

These are the grounds on which a party is declared to have lost. The category is defined at Nyāya Sūtra 1.2.19 and enumerated at the head of Book V, Chapter II, which lists twenty-two of them in one sūtra.

The twenty-two, in Vidyabhusana's own translation of 5.2.1:

1. Hurting the proposition, 2. Shifting the proposition, 3. Opposing the proposition, 4. Renouncing the proposition, 5. Shifting the reason, 6. Shifting the topic, 7. The meaningless, 8. The unintelligible, 9. The incoherent, 10. The inopportune, 11. Saying too little, 12. Saying too much, 13. Repetition, 14. Silence, 15. Ignorance, 16. Non-ingenuity, 17. Evasion, 18. Admission of an opinion, 19. Overlooking the censurable, 20. Censuring the non-censurable, 21. Deviating from a tenet, and 22. The semblance of a reason.

Read the list and four kinds of ground are visible. Grounds one to six concern the claim itself: abandoning it, changing it, or arguing against it. Grounds seven to thirteen concern the argument's form, and the tenth, the inopportune, is the one Vidyabhusana records for stating the members in a wrong order. Grounds fourteen to twenty concern the participant: silence, ignorance, evasion. And the twenty-second, the semblance of a reason, is the hetvābhāsa list of [Hetvābhāsa: The Five Fallacies of the Reason] entering here, which is why 5.2.25 provides that the fallacies also furnish occasions for rebuke.

munotes.in206

Chala, Jāti and Nigrahasthāna: How a Debate Is Lost

And Vidyabhusana's own note on the chapter is worth having: there are infinite occasions for rebuke, of which only twenty-two have been enumerated. So the list is not a closed set but a catalogue of the common ones, which is an honest thing for a procedural rule to say about itself.

The category exists because the debate has to END. A protocol without a termination condition runs for ever, and nigrahasthāna is the termination condition: when a named ground is established against you, the exchange is over.

The four kinds of failure, in one table

CategoryWhat failsBad faith required?Example
hetvābhāsathe reason offeredno"sound is eternal because it is intangible"
chalathe interpretation of a wordyes"nava-kambala" taken as nine blankets
jātithe relevance of an objectionnot necessarily"a pot has shape and sound has none"
nigrahasthānathe participant's conduct in the exchangenot necessarilyshifting the proposition, or stating the members in the wrong order

The bad faith column is the useful one. Only chala requires the objector to know what was meant and to take it otherwise. The others can be committed by somebody arguing honestly and badly, and the tradition's willingness to separate the two is a mark of a system built for real use.

Why a computer science student should care

Three reasons, and each is a modern practice with the same shape.

A protocol needs failure conditions, not only success conditions. A specification that says what a valid message looks like and not what to do with an invalid one is incomplete. The nigrahasthāna list is the invalid-message handling of a debate.

Ambiguity is attacked, not resolved, unless the protocol forbids it. Chala is what happens when a term is not pinned down, and the defence is the one a specification uses: define your terms first. [Formal Specification: Saying Exactly What a System Must Do] is the same lesson from the other side.

An objection must be relevant to the connection under argument. Jāti is the failure of relevance, and it is the commonest failure in a code review: an objection about style offered against a claim about correctness.

What these categories are NOT

They are not rules of logic. None of them is a claim about what follows from what. They are rules of a procedure.

They are not all dishonest. Only chala requires bad faith; the definitions of the others are silent about motive.

They are not exhaustively reproduced here. Book V, Chapter I runs to forty-three sūtras on futility and Chapter II lists twenty-two occasions for rebuke. This chapter gives the definitions, the twenty-two, and the structure, which is what MU's label "debate methodology and validation" requires, and says plainly that the futility treatment is in Book V, Chapter I.

munotes.in207

Chala, Jāti and Nigrahasthāna: How a Debate Is Lost

Quick revision

  • Chala, 1.2.10: opposition by assuming an alternative meaning. Three kinds, 1.2.11: term, genus, metaphor.
  • The nava-kambala example: "possessed of a new blanket" wilfully taken as "possessed of nine blankets".
  • Quibble in respect of a genus is the strawman named precisely: a claim about a particular refuted as a claim about the wide class.
  • Jāti, 1.2.18: objections founded on mere similarity or dissimilarity. Book V, Chapter I gives forty-three sūtras of named varieties.
  • Nigrahasthāna: defined at 1.2.19, enumerated as twenty-two grounds at 5.2.1, and Vidyabhusana notes the occasions are infinite and only twenty-two are listed.
  • Only chala requires bad faith. The category exists because a debate must terminate.

Test yourself

1. Define chala and give the Sūtra's own example.

Opposition to a proposition by assuming an alternative meaning of an ambiguous word. A speaker says a boy is nava-kambala, meaning possessed of a new blanket; the quibbler takes nava as "nine" and denies that the boy has nine blankets.

2. Which of the three kinds of quibble is the strawman, and why is the Sanskrit name more precise?

Quibble in respect of a genus. The name identifies the mechanism, which is substituting a very wide class for the particular that was spoken of, and then refuting the class claim.

3. Define jāti and give an example of a futile objection.

An objection founded on mere similarity or dissimilarity. Against "sound is non-eternal because it is produced, as a pot", the objection that a pot has shape and sound has none is futile: shape is no part of the connection between being produced and being non-eternal.

4. Why does a system of debate need a category of nigrahasthāna at all?

Because an exchange must be able to end. The occasions for rebuke are the termination condition: when a named ground is established against a party, that party has lost and the debate is over.

Contents This chapter on its own page

munotes.in208

Chapter Sixty-One

Vāda, Jalpa and Vitaṇḍā: Three Kinds of Dispute

Syllabus topic Module 2, "Debate methodology and validation"

In one line

Three kinds of argument: one to find the truth, one to win, and one that only attacks.

In the wording you can write in an examination: Nyāya distinguishes three kinds of kathā or disputation. Vāda, discussion, aims at ascertaining the truth and is conducted by the means of right knowledge and by confutation, without deviating from established tenets. Jalpa, wrangling, aims at victory and permits quibbles and futilities. Vitaṇḍā, cavil, is a kind of wrangling in which one party merely attacks the other's position without establishing one of their own.

The provisions

Nyāya Sūtra 1.2.1, in Vidyabhusana's translation:

Discussion is the adoption of one of two opposing sides. What is adopted is analysed in the form of five members, and defended by the aid of any of the means of right knowledge, while its opposite is assailed by confutation, without deviation from the established tenets.

And his note on the category as a whole:

A dialogue or disputation (kathā) is the adoption of a side by a disputant and its opposite by his opponent. It is of three kinds, viz., discussion which aims at ascertaining the truth, wrangling which aims at gaining victory, and cavil which aims at finding mere faults. A discutient is one who engages himself in a disputation as a means of seeking the truth.

Nyāya Sūtra 1.2.2.

Wrangling, which aims at gaining victory, is the defence or attack of a proposition in the manner aforesaid by quibbles, futilities, and other processes which deserve rebuke.

And his note: a wrangler is one who, engaged in a disputation, aims only at victory, being indifferent whether the arguments he employs support his own contention or that of his opponent, provided he can make out a pretext for bragging that he has taken an active part.

Nyāya Sūtra 1.2.3.

Cavil is a kind of wrangling which consists in mere attacks on the opposite side.

The three as protocols

Vāda, discussionJalpa, wranglingVitaṇḍā, cavil
Goalto ascertain the truthto winto defeat the other side
Own positionrequiredrequirednone is put forward
Admissible movesthe means of right knowledge, and confutationthose, plus quibbles and futilitiesattack only
Constraintmust not deviate from established tenetsnone stated beyond the formnone
Win conditionthe matter is settledthe opponent is defeatedthe opponent is defeated
Who uses ittwo people seeking the truthtwo people seeking to winone attacker

Read the "own position" row. That is what separates cavil from wrangling. A wrangler holds a side and defends it badly. A caviller holds nothing, so there is nothing to attack in return, and the exchange is asymmetric.

And notice that the tradition does not simply condemn jalpa. Nyāya Sūtra 4.2.50, which Vidyabhusana renders as the rule that wranglings and cavils may be employed to guard the truth, gives them a purpose: against an opponent who is not arguing in good faith, insisting on the rules of vāda loses you the exchange and the truth with it.

munotes.in209

Vāda, Jalpa and Vitaṇḍā: Three Kinds of Dispute

Worked: the dialogue the Sūtra prints

Vidyabhusana gives an instance of a discussion, and it is worth reading because it shows a vāda being conducted rather than described.

Discutient: There is soul.

Opponent: There is no soul.

Discutient: Soul is existent (proposition). Because it is an abode of consciousness (reason). Whatever is not existent is not an abode of consciousness, as a hare's horn (negative example). Soul is not so, that is, soul is an abode of consciousness (negative application). Therefore soul is existent (conclusion).

Opponent: Soul is non-existent (proposition). Because, etc.

Discutient: The scripture which is a verbal testimony declares the existence of soul.

Discutient: If there were no soul, it would not be possible to apprehend one and the same object through sight and touch.

Discutient: The doctrine of soul harmonises well with the various tenets which we hold ... Therefore there is soul.

And his bracketed note: the discussion will be considerably lengthened if the opponent happens to be a Buddhist who does not admit the authority of scripture, and holds that there are no eternal things.

What that dialogue demonstrates

The five members are used once, in full, at the opening. After that the discutient argues in ordinary prose. So the five-member form is the formal statement of a position, not the format of every sentence in a debate. That is worth knowing, because a student who imagines every utterance must be five members has misread the apparatus.

The example member here is the NEGATIVE one, anchored in a hare's horn, which is the standard instance of something that does not exist. That is the vipakṣa of [Anumāna: Vyāpti, and the Three Kinds of Inference] doing its job.

The discutient then offers three further supports of three different kinds: verbal testimony, an inference from the unity of perception across senses, and coherence with the other accepted tenets. Three channels, not one, which is Charaka's rule from [Pramāṇa: Where Knowledge Is Allowed To Come From] applied in a debate.

And Vidyabhusana's closing note is the most modern thing in the passage. If the opponent does not admit scripture, the argument from scripture is unavailable and the discussion is longer. The admissible means of knowledge are a matter of agreement between the parties, and a support that one party does not accept is not a support at all. That is exactly the "established tenets" constraint of 1.2.1: a discussion is conducted within what both sides already grant.

munotes.in210

Vāda, Jalpa and Vitaṇḍā: Three Kinds of Dispute

The modern reading

The three kinds are three settings you can recognise at once.

Vāda is a design review among colleagues. Both hold positions, both want the right answer, both accept the same evidence, and either may change their mind. Its failure mode is that it is slow and that it cannot survive a participant who is not playing.

Jalpa is an adversarial proceeding. Each side argues its own case, moves that are technically permitted are used whether or not they illuminate, and the aim is to prevail. Its virtue, and the Sūtra grants it one, is that it works against an opponent who will not play fair.

Vitaṇḍā is pure critique. The critic advances nothing and only finds fault. Its virtue is that finding fault is genuinely useful and does not require an alternative. Its vice is that it cannot be answered in kind, because there is nothing to attack.

The lesson to state in an answer. A system for settling disagreements has to say which of the three it is running, because the admissible moves and the win condition differ. An exchange in which one party believes they are in a vāda and the other is conducting a jalpa goes badly, and it goes badly in a way the tradition named two thousand years before anybody wrote it down as a lesson about code review.

What these categories are NOT

They are not a ranking. Vāda is the best for finding the truth and the tradition still licenses the other two, for the reason given in 4.2.50.

Vitaṇḍā is not the same as scepticism. The caviller need hold no view about the matter at all; a sceptic holds that it cannot be settled, which is a position and would make the exchange a vāda.

They are not about tone. A wrangler may be perfectly polite. The classification is by goal and admissible moves, not by manner.

Quick revision

  • Vāda, 1.2.1: discussion aiming at truth, in five members, defended by the means of right knowledge and confutation, without deviating from established tenets.
  • Jalpa, 1.2.2: wrangling aiming at victory, permitting quibbles, futilities and moves that deserve rebuke.
  • Vitaṇḍā, 1.2.3: cavil, wrangling that merely attacks and advances no position of its own.
  • The dialogue on the soul shows the five members used once to state a position, then ordinary argument, with three different channels of support.
  • The admissible means are a matter of agreement: an argument from scripture is unavailable against an opponent who does not admit it.
  • 4.2.50 licenses wrangling and cavil to guard the truth against an opponent who is not arguing in good faith.
munotes.in211

Vāda, Jalpa and Vitaṇḍā: Three Kinds of Dispute

Test yourself

1. Distinguish the three kinds of disputation by goal and by whether a position is held.

Vāda aims at truth and both parties hold positions. Jalpa aims at victory and both hold positions. Vitaṇḍā aims at defeating the other side and the caviller holds none.

2. What does the dialogue on the soul show about how the five-member form is used?

That it is used once, to state a position formally, and that the argument continues afterwards in ordinary prose. It is the formal statement of a claim, not the format of every sentence.

3. What is the significance of Vidyabhusana's note about a Buddhist opponent?

That the admissible means of knowledge depend on what both parties accept. An argument from scripture is no argument against an opponent who does not admit scripture, which is what the requirement not to deviate from established tenets amounts to.

4. Why does the tradition permit wrangling at all?

Because, as Nyāya Sūtra 4.2.50 provides, wranglings and cavils may be employed to guard the truth. Against an opponent who is not arguing in good faith, insisting on the rules of discussion loses the exchange and the truth with it.

Contents This chapter on its own page

munotes.in212

Chapter Sixty-Two

Debate Methodology as a Validation Protocol

Syllabus topic Module 2, "Debate methodology and validation", "Inference engines"

In one line

Nyāya's debate rules are a validation protocol: they say what may be claimed, what counts as support, what moves are allowed, what counts as a violation, and how the exchange ends.

In the wording you can write in an examination: the Nyāya apparatus for disputation constitutes a validation protocol, that is, a procedure by which a claim is tested by an adversary under stated rules. Its components are a required form for the claim and its support, a fixed set of admissible means of knowledge, an enumerated set of defects in a reason, an enumerated set of improper moves, and enumerated grounds on which a party is declared defeated.

The five parts of a protocol

Any procedure for settling disagreements must have these five, and Nyāya has all five explicitly.

PartWhat it answersNyāya's version
The form of a claimhow must a position be stated?the five members, 1.1.32
Admissible supportwhat counts as evidence?the four pramāṇas, 1.1.3
Detectable defectswhat makes support bad?the five hetvābhāsa, 1.2.4
Improper moveswhat may a participant not do?chala and jāti, 1.2.10 and 1.2.18
Terminationhow does it end?the nigrahasthāna of Book V.2

Missing any one and the procedure fails in a specific way. Without a form, a party can shift ground. Without admissible support, every dispute becomes a dispute about what counts as a reason. Without defects, a bad argument cannot be named as bad. Without improper moves, the exchange degenerates. Without termination, it never ends.

Why adversarial testing is the design

The whole apparatus assumes an opponent whose job is to find the fault. That is a choice and it has consequences worth stating.

What it buys. A claim that survives a hostile examination has been tested in a way no amount of self-checking achieves. The arguer does not choose the attacks, so the attacks are not the ones the arguer has already thought about.

What it costs. It needs a second party, and the second party must be competent and must be playing. Vidyabhusana's note on the Buddhist opponent shows the first difficulty and the distinction between vāda and jalpa shows the second.

And it has no self-test. Nothing in the machinery examines itself. A claim nobody attacks passes, however bad it is. That is the standing weakness of any adversarial validation and it is the reason modern practice adds mechanical checks alongside human review.

Three modern protocols with the same five parts

Peer review of a paper

PartIts form
form of a claimthe paper, with its stated contributions and its method section
admissible supportexperiment, proof, and citation of accepted results
detectable defectsthe reviewer's checklist: unsupported claim, flawed experiment, missing baseline
improper movesreviewing the author rather than the work; demanding a different paper
terminationthe editor's decision
munotes.in213

Debate Methodology as a Validation Protocol

Where it is weaker than Nyāya. The list of defects is not enumerated anywhere, so two reviewers apply different lists and neither can be told they have applied none.

A test suite

PartIts form
form of a claimthe assertion, which names the expected value
admissible supportthe actual run of the code
detectable defectsa failing assertion, and a test that passes when the code is broken
improper movesasserting on incidental output, or testing the mock rather than the code
terminationthe suite goes green, or it does not

Where it is stronger than Nyāya. It is mechanical, so it needs no competent opponent, and it runs on every change. That is the modern addition that the classical apparatus has no equivalent of, and it is why this book's own claims are checked by programs rather than argued about.

A code review

PartIts form
form of a claimthe pull request, with its description of what it changes and why
admissible supportthe diff, the tests, and the measurements
detectable defectsthe review checklist
improper movesobjecting on style against a claim about correctness, which is jāti exactly
terminationapproval, or a request for changes

Where it repeats Nyāya's weaknesses. It needs a competent reviewer who is actually engaging, and a review that finds nothing is indistinguishable from a review that did not look.

Worked example: a claim through the whole protocol

The claim. This change makes the endpoint twice as fast.

Form. Stated in five members. Proposition: this change makes the endpoint twice as fast. Reason: because it removes the per-request database lookup. Example: whatever removes a per-request database lookup on this endpoint halves its latency, as the change to the neighbouring endpoint last month. Application: so is this change: it removes that lookup. Conclusion: therefore it makes the endpoint twice as fast.

Admissible support. The measurement is pratyakṣa. The neighbouring endpoint's result is the familiar instance. The claim that the lookup was per-request rests on reading the code, which is pratyakṣa of the code rather than of the behaviour.

Defects, checked. Is the reason erratic: could removing that lookup fail to halve the latency? Yes, if the lookup was not the dominant cost. The connection is erratic and the argument fails, which the form exposed at the example member.

Repair. Narrow the mark: because it removes the per-request database lookup, which the profile shows accounted for half the latency. Now the connection is defensible, and the cost of the repair is a profile, which is the second observation [The Five Members Worked, Three Times] predicts.

munotes.in214

Debate Methodology as a Validation Protocol

Improper moves to watch for. An objection that the variable names are poor is jāti: relevant to something, irrelevant to this claim.

Termination. The claim is established or the change is not merged.

What the protocol does NOT do

It does not find claims. Nothing in it tells you what to test. The five extra members of [The Five-Member Syllogism] include jijñāsā, inquiry, which is an attempt to bolt a discovery stage onto the front, and it is not part of the standard five.

It does not guarantee truth. A claim can survive every check and be false, because the vyāpti rests on observation and observation does not establish universals.

It does not work without an opponent. This is the one to say last, because it is the sharpest criticism and it is the one modern practice answers. A protocol whose only detector is a person can be satisfied by nobody looking. A machine check cannot be satisfied by nobody looking, which is why every claim in this book about a number, a table or an algorithm is produced by a program that runs.

Quick revision

  • Five parts of a validation protocol: the form of a claim, admissible support, detectable defects, improper moves, and termination.
  • Nyāya has all five explicitly: 1.1.32, 1.1.3, 1.2.4, 1.2.10 with 1.2.18, and Book V.2.
  • Missing each one fails in a specific way: shifting ground, disputes about what counts as a reason, unnameable bad arguments, degeneration, and no end.
  • Adversarial design buys attacks the arguer did not choose and costs a competent opponent who is playing.
  • Peer review, a test suite and a code review all have the same five parts. The test suite is the one that needs no opponent.
  • The protocol does not find claims, does not guarantee truth, and cannot detect a failure nobody looked for.

Test yourself

1. Name the five parts of a validation protocol and Nyāya's version of each.

The form of a claim, the five members. Admissible support, the four pramāṇas. Detectable defects, the five fallacies of the reason. Improper moves, quibble and futility. Termination, the occasions for rebuke.

2. What does a protocol lose if it has no termination condition?

It never ends. Without enumerated grounds of defeat there is no point at which one party has lost, so the exchange can be prolonged indefinitely by a participant who will not concede.

3. Give the greatest strength and the greatest weakness of adversarial validation.

Its strength is that the attacks are chosen by someone other than the arguer, so they are not the ones already anticipated. Its weakness is that it requires a competent opponent who is actually engaging, and a review that finds nothing is indistinguishable from a review that did not look.

munotes.in215

Debate Methodology as a Validation Protocol

4. Which modern protocol repairs that weakness, and how?

A test suite. Its detector is mechanical and runs on every change, so it cannot be satisfied by nobody looking, whereas peer review and code review both can.

Contents This chapter on its own page

munotes.in216

Chapter Sixty-Three

Padārtha: The Categories of What Exists

Syllabus topic Module 2, "Padārtha (categorical ontology overview)"

In one line

Padārtha means the categories of what there is: six kinds of thing, into one of which everything falls.

In the wording you can write in an examination: padārtha, literally the referent of a word, is a category of existence. The Vaiśeṣika system of Kaṇāda enumerates six: dravya, substance; guṇa, attribute; karma, action; sāmānya, genus; viśeṣa, species or particularity; and samavāya, inherence or combination. The later tradition adds a seventh, abhāva, non-existence. Everything nameable falls under one of them, so the scheme is an exhaustive ontology.

First, whose scheme is it?

MU prints this label under her Nyāya block, and it is worth being exact, because a question may test it.

The padārtha scheme is Vaiśeṣika's, from the Vaiśeṣika Sūtra of Kaṇāda, not from the Nyāya Sūtra.

Nyāya has its own enumeration of what there is, category two, prameya, the object of right knowledge. Nyāya Sūtra 1.1.9 lists twelve: soul, body, senses, objects of sense, intellect, mind, activity, fault, transmigration, fruit, pain and release.

And Vidyabhusana records the connection himself, in his note on that sūtra: the objects of right knowledge are also enumerated as substance, quality, action, generality, particularity, intimate relation and non-existence, which he calls the technicalities of the Vaiśeṣika philosophy.

So the two schools' lists were treated as interchangeable accounts of the same thing by the time of the commentaries, and the two systems eventually merged into a single tradition. That is why MU can print padārtha under Nyāya without error, and an answer should be able to say all of this in two sentences.

The provision

Nandalal Sinha, in the introduction to his 1923 translation of the Vaiśeṣika Sūtras:

By a subtle process of analysis and synthesis, Kaṇāda divides all nameable things into six classes: viz. substance, attribute, action, genus, species, and combination.

And his own note on the seventh:

"Non-existence" is the seventh Predicable, not denied by Kaṇāda.

The six

CategorySanskritWhat it isAn example
substancedravyathat in which attributes and actions inherea pot
attributeguṇawhat inheres in a substance and does not itself movethe pot's colour
actionkarmamotion, which inheres in a substancethe pot falling
genussāmānyawhat is common to many, in virtue of which they are one kindpotness, shared by all pots
speciesviśeṣathe particularity that distinguishes one eternal substance from another otherwise identicalwhat makes one atom not another
inherencesamavāyathe relation by which an attribute is IN its substancethe relation between the pot and its colour

Worked, from Sinha's own gloss on the dependence, which is the structural point: attribute and action exist by combination with substance, and without substance there were no attribute and action; genus and species are correlative; and combination is the intimate connection. He adds that in reality there are only three predicables, substance, attribute and action, the remaining three being relations among them.

munotes.in217

Padārtha: The Categories of What Exists

The nine substances

Vaiśeṣika Sūtra 1.1.5, as Sinha prints it:

Earth, Waters, Fire, Air, Ether, Time, Space, Self, and Mind are the nine Substances.

Note what is in that list. Time and space are substances, not relations; self and mind are two different things; and the four elements plus ether are the material ones. It is a specific and contestable ontology, not a neutral one, and its specificity is what makes it a real commitment.

And Vaiśeṣika Sūtra 1.1.6: attributes are colour and the rest, which the sūtras then enumerate.

Why this is an ontology in the computer science sense

The word ontology has a technical use in computer science, and this is not a loose analogy.

An ontology fixes what kinds of thing exist in a domain, so that everything represented can be assigned a kind.

It fixes what relations may hold between them. Kaṇāda's samavāya is a relation, and it is one of the categories rather than something outside the scheme, which is a design decision a modern ontologist makes too.

It is exhaustive by claim. Everything nameable falls under one of the six, which is what makes it a commitment rather than a list. An ontology that does not claim exhaustiveness cannot be used to check that a representation is well formed.

And the seventh category is the interesting one. Non-existence as a category means the scheme can represent the absence of a thing as itself a thing to be talked about. In a modern representation that is the difference between "no value recorded" and "the value is known to be absent", and systems that fail to distinguish them fail in well-known ways.

The comparison, stated carefully

VaiśeṣikaA modern ontology
dravya, substancea class whose instances exist independently
guṇa, attributea property, with a value
karma, actionan event or a process
sāmānya, genusa superclass, or a type
viśeṣa, particularityan identity, distinguishing otherwise indistinguishable instances
samavāya, inherencethe relation that binds a property to its bearer
abhāva, non-existencean explicit absence, as distinct from an unrecorded value

The fifth row deserves a note. Viśeṣa exists to solve the problem of two eternal atoms that share every attribute: what makes them two? The modern version is object identity, and it is exactly the same problem: two records with identical fields are still two records, and something other than their fields must say so.

And the sixth row is the one a beginner underrates. Making the relation between a thing and its property a first-class category, rather than leaving it implicit in the notation, is what allows the scheme to ask questions about it. [Padārtha Ontology as a Knowledge Representation Model] builds that.

munotes.in218

Padārtha: The Categories of What Exists

What the scheme is NOT

It is not a classification of objects. It is a classification of KINDS of thing. A pot is a substance; potness is a genus; the pot's colour is an attribute. Confusing the levels is the standard beginner's error and the tradition is careful about it.

It is not Nyāya's own list. Nyāya Sūtra 1.1.9 gives twelve objects of right knowledge, which is a different enumeration for a different purpose.

It is not uncontested. Whether viśeṣa is needed, whether non-existence is a category, and whether there are really only three predicables are all argued in the literature, and Sinha reports several of the disputes.

It is not a modern ontology. There is no machine-readable form, no consistency checker, and no notion of a query. The correspondence is at the level of what the scheme is trying to do.

Quick revision

  • Padārtha: a category of existence, a kind of thing. Kaṇāda's six, from Sinha's translation: substance, attribute, action, genus, species, combination.
  • A seventh, non-existence, is added by the later tradition and Sinha records that Kaṇāda does not deny it.
  • The scheme is Vaiśeṣika's. Nyāya's own list is prameya, twelve objects of right knowledge at 1.1.9, and Vidyabhusana records that the two were treated as alternative accounts.
  • The nine substances, 1.1.5: earth, waters, fire, air, ether, time, space, self and mind.
  • Sinha's note: attribute and action exist by combination with substance, and in reality there are only three predicables, the other three being relations.
  • Modern counterparts: class, property, event, type, identity, the property-bearer relation, and explicit absence.
  • Viśeṣa is object identity: what makes two otherwise identical things two.

Test yourself

1. Name the six padārthas and the seventh, and say whose scheme it is.

Substance, attribute, action, genus, species and combination, with non-existence as a seventh added by the later tradition. It is the Vaiśeṣika scheme of Kaṇāda, not the Nyāya Sūtra's own list.

2. How does the padārtha scheme relate to Nyāya's own enumeration?

Nyāya's second category, prameya, is enumerated at 1.1.9 as twelve objects of right knowledge. Vidyabhusana records that the objects of right knowledge are also enumerated as the Vaiśeṣika categories, so the commentators treated the two as alternative accounts, and the schools later merged.

3. What problem does viśeṣa solve, and what is its modern counterpart?

It explains what makes two eternal substances with identical attributes two rather than one. The modern counterpart is object identity: two records with identical fields are still two records, and something other than the fields must establish it.

munotes.in219

Padārtha: The Categories of What Exists

4. Why does it matter that inherence is a category rather than being left implicit?

Because a relation that is a first-class category can itself be talked about and reasoned over, whereas one left implicit in the notation cannot. A representation that names the property-bearer relation can ask questions about it.

Contents This chapter on its own page

munotes.in220

Chapter Sixty-Four

Padārtha Ontology as a Knowledge Representation Model

Syllabus topic Module 2, "Padārtha (categorical ontology overview)", "Knowledge representation", "Algorithm Specification (Pseudo-code)"

In one line

Write the six categories as data, make inherence an entry rather than a line, and the scheme becomes something you can query.

In the wording you can write in an examination: a knowledge representation model built on the padārtha scheme assigns every entry to exactly one of the categories, represents attributes and actions as entries in their own right rather than as fields of a substance, and represents the inherence relation as a further entry linking an attribute to its substance. Queries are then answered by selecting entries of the appropriate category, and the scheme can be checked for well-formedness by confirming that every entry has a category and every reference resolves.

Problem statement, in MU's own form

IKS concept as CS concept: the padārtha scheme of the Vaiśeṣika system as an ontology, with queries and a well-formedness check.

Statement. Represent a small domain as entries each assigned to one of the seven padārthas. Represent inherence as its own entry rather than as a field, so that the relation is first class. Answer queries that select by category, and check that the whole is well formed.

Conceptual mapping table

Classical elementComputer science element
dravya, substancean instance that exists independently
guṇa, attributea property value, held as an entry of its own
karma, actionan event, held the same way
sāmānya, genusa class, with its extension recorded
viśeṣa, particularityan identity, distinguishing indistinguishable instances
samavāya, inherencea link entry joining an attribute to a substance
abhāva, non-existencean explicitly recorded absence
"everything nameable falls under one of the six"the well-formedness condition

Algorithm specification, in pseudo-code

STRUCTURE

every entry has a NAME, a CATEGORY drawn from the seven, and its own facts

QUERY attributes_of(substance)

return the attribute of every samavaya entry whose substance is this one

QUERY bearers_of(attribute)

return the substance of every samavaya entry whose attribute is this one

QUERY instances_of(genus)

return what the genus entry records itself as holding of

QUERY known_absent(substance)

return what every abhava entry records as absent from this one

CHECK well_formed

for every entry:

its category must be one of the seven

every reference it makes to an entry must name an entry that exists

The design point is in the first two queries. Because inherence is an entry, the same relation can be walked in both directions with no extra structure. If the colour were a field on the pot, finding what bears the colour red would need a scan of every substance.

Working code

"""Padartha ontology as a knowledge representation model. MU's topic 7."""

PADARTHA = ["dravya", "guna", "karma", "samanya", "visesa", "samavaya", "abhava"]

# Each entry: (padartha, the facts that entry carries)
ONTOLOGY = {
    # substances: they bear attributes and actions
    "pot1":       ("dravya",   {"kind": "pot", "made_of": "earth"}),
    "pot2":       ("dravya",   {"kind": "pot", "made_of": "earth"}),
    "lamp1":      ("dravya",   {"kind": "lamp", "made_of": "fire"}),
    # genera: what many substances share
    "potness":    ("samanya",  {"holds_of": ["pot1", "pot2"]}),
    "lampness":   ("samanya",  {"holds_of": ["lamp1"]}),
    # particularities: what makes two otherwise identical substances two
    "this1":      ("visesa",   {"of": "pot1"}),
    "this2":      ("visesa",   {"of": "pot2"}),
    # attributes, each inhering in a substance by an inherence relation
    "red":        ("guna",     {"value": "red", "of_kind": "colour"}),
    "blue":       ("guna",     {"value": "blue", "of_kind": "colour"}),
    "bright":     ("guna",     {"value": "bright", "of_kind": "luminosity"}),
    # actions
    "falling":    ("karma",    {"of_kind": "motion"}),
    # inherence: the relation itself is an entry, which is the design point
    "inh1":       ("samavaya", {"attribute": "red",     "substance": "pot1"}),
    "inh2":       ("samavaya", {"attribute": "blue",    "substance": "pot2"}),
    "inh3":       ("samavaya", {"attribute": "bright",  "substance": "lamp1"}),
    "inh4":       ("samavaya", {"attribute": "falling", "substance": "pot2"}),
    # non-existence, recorded explicitly rather than by silence
    "no_colour_on_lamp": ("abhava", {"absent": "colour", "from": "lamp1"}),
}


def category(name):
    return ONTOLOGY[name][0] if name in ONTOLOGY else None


def facts(name):
    return ONTOLOGY[name][1] if name in ONTOLOGY else {}


def attributes_of(substance):
    """Every attribute or action inhering in a substance, via samavaya."""
    out = []
    for name, (cat, f) in ONTOLOGY.items():
        if cat == "samavaya" and f.get("substance") == substance:
            out.append(f["attribute"])
    return sorted(out)


def bearers_of(attribute):
    return sorted(f["substance"] for cat, f in
                  (ONTOLOGY[n] for n in ONTOLOGY)
                  if cat == "samavaya" and f.get("attribute") == attribute)


def instances_of(genus):
    return sorted(facts(genus).get("holds_of", []))


def known_absent(substance):
    return sorted(f["absent"] for cat, f in
                  (ONTOLOGY[n] for n in ONTOLOGY)
                  if cat == "abhava" and f.get("from") == substance)


def identity_of(substance):
    for name, (cat, f) in ONTOLOGY.items():
        if cat == "visesa" and f.get("of") == substance:
            return name
    return None


def check_well_formed():
    """Every entry is in exactly one category, and every reference resolves."""
    problems = []
    for name, (cat, f) in ONTOLOGY.items():
        if cat not in PADARTHA:
            problems.append("%s: unknown category %s" % (name, cat))
        for key in ("of", "substance", "attribute", "from"):
            ref = f.get(key)
            if ref is not None and ref not in ONTOLOGY:
                problems.append("%s: %s points at %s, which is not in the ontology"
                                % (name, key, ref))
        for ref in f.get("holds_of", []):
            if ref not in ONTOLOGY:
                problems.append("%s: holds_of names %s, which is not in the ontology" % (name, ref))
    return problems


QUERIES = [
    ("what inheres in pot1?",                  lambda: attributes_of("pot1")),
    ("what inheres in pot2?",                  lambda: attributes_of("pot2")),
    ("what bears the attribute red?",          lambda: bearers_of("red")),
    ("what falls under potness?",              lambda: instances_of("potness")),
    ("what distinguishes pot1 from pot2?",     lambda: [identity_of("pot1"), identity_of("pot2")]),
    ("what is KNOWN to be absent from lamp1?", lambda: known_absent("lamp1")),
    ("what is known to be absent from pot1?",  lambda: known_absent("pot1")),
    ("category of samavaya entry inh1",        lambda: [category("inh1")]),
    ("category of potness",                    lambda: [category("potness")]),
    ("category of an unknown name",            lambda: [category("nothing")]),
]

print("%-42s %s" % ("query", "answer"))
for q, fn in QUERIES:
    print("%-42s %s" % (q, fn() or "(nothing)"))

print()
print("entries by category")
for p in PADARTHA:
    names = sorted(n for n in ONTOLOGY if category(n) == p)
    print("  %-10s %d  %s" % (p, len(names), ", ".join(names)))

print()
bad = check_well_formed()
print("well-formedness:", "%d problem(s)" % len(bad) if bad else "every entry is in one "
      "category and every reference resolves")
for b in bad:
    print("  " + b)
munotes.in221

Padārtha Ontology as a Knowledge Representation Model

query                                      answer
what inheres in pot1?                      ['red']
what inheres in pot2?                      ['blue', 'falling']
what bears the attribute red?              ['pot1']
what falls under potness?                  ['pot1', 'pot2']
what distinguishes pot1 from pot2?         ['this1', 'this2']
what is KNOWN to be absent from lamp1?     ['colour']
what is known to be absent from pot1?      (nothing)
category of samavaya entry inh1            ['samavaya']
category of potness                        ['samanya']
category of an unknown name                [None]

entries by category
  dravya     3  lamp1, pot1, pot2
  guna       3  blue, bright, red
  karma      1  falling
  samanya    2  lampness, potness
  visesa     2  this1, this2
  samavaya   4  inh1, inh2, inh3, inh4
  abhava     1  no_colour_on_lamp

well-formedness: every entry is in one category and every reference resolves
munotes.in222

Padārtha Ontology as a Knowledge Representation Model

Reading the output

The two inherence queries run in opposite directions over the same entries. Nothing was added to make the reverse query possible; it is a consequence of making the relation an entry.

"What distinguishes pot1 from pot2?" returns two viśeṣa entries. Both pots are pots made of earth, so their substance facts are identical; the two particularities are what makes them two. That is object identity, and it is the modern problem viśeṣa was invented for.

"Known to be absent from lamp1" returns colour, and the same query on pot1 returns nothing. Those are two DIFFERENT answers: the lamp is recorded as having no colour, and nothing at all is recorded about the pot's. A representation that returned the same answer for both would have lost the distinction abhāva exists to preserve.

And the last query returns None for a name that is not in the ontology, which is a third state again: not an entry at all.

Three states, and why they must be three

This is the most useful thing in the chapter and it is a lesson people relearn every year.

The questionThe answer hereWhat it means
does lamp1 have a colour?abhāva entry: colour is absent from lamp1we know it does not
does pot1 have a colour recorded as absent?nothingwe have not said
is "nothing" an entry?Noneit is not in the ontology at all

Collapse the first two and a system cannot distinguish "known to be none" from "not recorded". Every database that uses a null for both meets this, and every reporting system built on one gets a wrong total eventually. Kaṇāda's seventh category is exactly the fix, and Sinha records that Kaṇāda does not deny it.

munotes.in223

Padārtha Ontology as a Knowledge Representation Model

The well-formedness check, and its declared hole

The check confirms two things: every entry's category is one of the seven, and every reference to an entry names an entry that exists. It reports no problems on this ontology.

And it does not check one field. The abhāva entry records that "colour" is absent from lamp1, and colour is a KIND of attribute rather than an entry: the entries are red, blue and bright, each of which has colour or luminosity as its kind. So "absent" names a kind and the reference check, which looks for entries, does not apply to it.

That is a real limitation and it is stated here rather than hidden. A fuller implementation would have kinds as entries too, probably as sāmānya, and then the check would cover it. As written, a typo in an abhāva entry's "absent" field would pass.

Saying so is the point. A check whose limits are not declared is worse than no check, because it is trusted. This is the same discipline as [Śāstra Rule Precedence as a Deterministic Finite Rewrite System] reporting which precedence branches its tests exercise.

Complexity and limitations

Time. Every query is a scan of the entries, so it is proportional to the size of the ontology. A real system would index the samavāya entries by substance and by attribute, turning each scan into a lookup, which is the standard move and the reason triple stores exist.

Space. One entry per thing, per attribute, per action, per genus, per identity, per inherence and per recorded absence. Making inherence an entry costs one entry per attribute held, which is the price of the reverse query.

Limitation: the ontology has no reasoner. Nothing infers that because pot1 falls under potness and potness is a genus, pot1 is a substance. Every fact is stated. Adding inference is what [Inference Engines: Forward and Backward Chaining] does, and combining the two is what a modern ontology system is.

Limitation: sāmānya is recorded extensionally. The genus entry lists the substances it holds of, rather than stating a condition they satisfy. That works for three pots and not for a million, and the alternative, a condition, is what a class definition in a modern ontology gives you.

Limitation: it is not Kaṇāda's own domain. Pots and lamps are the tradition's standard examples, and the ontology here is small enough to read. Nothing about the size of the example bears on whether the scheme is sound.

Quick revision

  • Every entry carries exactly one of the seven padārthas, which is the well-formedness condition.
  • Inherence is an entry, not a link in the notation, so the relation is queryable in both directions for free.
  • Viśeṣa gives object identity: two pots with identical facts are two because two particularities say so.
  • Abhāva gives an explicit absence, and the three states are known-absent, not-recorded, and not-an-entry. Collapsing the first two is a standard and expensive mistake.
  • The well-formedness check covers categories and entry references, and does NOT cover the field that names a kind, which is stated on the page.
  • Limitations: no reasoner, genus recorded extensionally, and every query is a scan without indexes.
munotes.in224

Padārtha Ontology as a Knowledge Representation Model

Test yourself

1. Why is inherence represented as an entry rather than as a field on the substance?

Because a relation held as an entry can be walked in both directions with no additional structure: the same entries answer "what inheres in this substance" and "what bears this attribute". As a field, the reverse query would need a scan of every substance.

2. Distinguish the three answers the ontology can give about an attribute of a thing.

An abhāva entry records that the attribute is known to be absent. No entry at all means nothing has been recorded. A name not in the ontology means the thing is not represented. A system that collapses the first two loses the distinction that abhāva exists to preserve.

3. What does the well-formedness check NOT cover, and why does saying so matter?

It does not check the field of an abhāva entry that names a kind of attribute rather than an entry, so a typo there would pass. It matters because a check whose limits are undeclared is trusted beyond what it establishes.

4. Give two limitations of this ontology as a knowledge representation.

It has no reasoner, so every fact must be stated rather than derived; and a genus records its extension, the list of things it holds of, rather than a condition those things satisfy, which does not scale.

Contents This chapter on its own page

munotes.in225

Chapter Sixty-Five

Propositional Logic: The Minimum You Need

Syllabus topic Module 2, "Propositional logic"

In one line

Propositional logic is about whole statements and five ways of joining them, and a truth table settles every question it can ask.

In the wording you can write in an examination: propositional logic is the formal study of statements that are either true or false, and of compound statements built from them by connectives. Its vocabulary is the atomic proposition, the connectives negation, conjunction, disjunction, implication and equivalence, and the truth table, which lists the truth value of a compound statement for every assignment of values to its atoms. A formula true on every row is a tautology, and an argument is valid if no row makes its premises true and its conclusion false.

The vocabulary

Atomic proposition. A statement that is either true or false and has no internal structure for this logic to see. Written P, Q, R. "This hill is smoky" is one.

Connective. An operation making a new statement from one or two others.

Compound proposition. Any statement built from atoms with connectives.

Assignment. One choice of truth value for each atom. With n atoms there are 2 to the power n assignments, which is [Laghu and Guru as One Bit] arriving in a new chapter.

Truth table. The list of all assignments with the value of the formula on each.

Tautology. True on every row. Contradiction. False on every row. Contingent. Neither.

Valid argument. One where no row makes every premise true and the conclusion false.

The connectives, and how they are written here

NameWrittenRead asTrue when
negation¬Pnot PP is false
conjunctionP • QP and Qboth are true
disjunctionP ∨ QP or Qat least one is true
implicationP ⊃ Qif P then QP is false, or Q is true
equivalenceP ≡ QP if and only if Qboth have the same value
exclusive orP ⊕ QP or Q but not boththey differ

Two of those need a warning.

Disjunction is inclusive. "P or Q" is true when both are true. English "or" is often exclusive, which is what ⊕ is for.

Implication is the one everyone finds strange. "P ⊃ Q" is true whenever P is false, whatever Q is. "If the moon is made of cheese then 2 plus 2 is 5" is true. It is not a claim about relevance or causation; it is a claim that you never have P true with Q false. State that in an answer and you have said the thing students most often get wrong.

The table, generated

"""Truth tables, generated rather than typed."""
import itertools

OPS = [
    ("not P",   "¬P",      lambda p, q: not p),
    ("P and Q", "P • Q",   lambda p, q: p and q),
    ("P or Q",  "P ∨ Q",   lambda p, q: p or q),
    ("P implies Q", "P ⊃ Q", lambda p, q: (not p) or q),
    ("P iff Q", "P ≡ Q",   lambda p, q: p == q),
    ("P xor Q", "P ⊕ Q",   lambda p, q: p != q),
]

def tf(x):
    return "T" if x else "F"

print("| P | Q | " + " | ".join(sym for _, sym, _ in OPS) + " |")
print("|---|---|" + "|".join("---" for _ in OPS) + "|")
for p, q in itertools.product([True, False], repeat=2):
    row = [tf(p), tf(q)] + [tf(f(p, q)) for _, _, f in OPS]
    print("| " + " | ".join(row) + " |")

print()
print("the three inference rules, checked over every row")
RULES = [
    ("modus ponens",      lambda p, q: not ((not p or q) and p) or q),
    ("modus tollens",     lambda p, q: not ((not p or q) and not q) or not p),
    ("affirming the consequent (INVALID)", lambda p, q: not ((not p or q) and q) or p),
]
for name, form in RULES:
    rows = [(p, q, form(p, q)) for p, q in itertools.product([True, False], repeat=2)]
    bad = [(tf(p), tf(q)) for p, q, ok in rows if not ok]
    print("  %-38s %s" % (name, "valid on all 4 rows" if not bad
                          else "FAILS on P=%s Q=%s" % bad[0]))

print()
print("de Morgan, checked")
LAWS = [
    ("not (P and Q) equals (not P) or (not Q)",
     lambda p, q: (not (p and q)) == ((not p) or (not q))),
    ("not (P or Q) equals (not P) and (not Q)",
     lambda p, q: (not (p or q)) == ((not p) and (not q))),
    # NOTE: the two sides are written DIFFERENTLY on purpose. Comparing a formula
    # with itself is a check that cannot fail, and the first draft of this line
    # did exactly that.
    ("P implies Q equals (not P) or Q",
     lambda p, q: (not (p and not q)) == ((not p) or q)),
]
for name, law in LAWS:
    ok = all(law(p, q) for p, q in itertools.product([True, False], repeat=2))
    print("  %-46s %s" % (name, "holds on all 4 rows" if ok else "FAILS"))
munotes.in226

Propositional Logic: The Minimum You Need

| P | Q | ¬P | P • Q | P ∨ Q | P ⊃ Q | P ≡ Q | P ⊕ Q |
|---|---|---|---|---|---|---|---|
| T | T | F | T | T | T | T | F |
| T | F | F | F | T | F | F | T |
| F | T | T | F | T | T | F | T |
| F | F | T | F | F | T | T | F |

the three inference rules, checked over every row
  modus ponens                           valid on all 4 rows
  modus tollens                          valid on all 4 rows
  affirming the consequent (INVALID)     FAILS on P=F Q=T

de Morgan, checked
  not (P and Q) equals (not P) or (not Q)        holds on all 4 rows
  not (P or Q) equals (not P) and (not Q)        holds on all 4 rows
  P implies Q equals (not P) or Q                holds on all 4 rows
munotes.in227

Propositional Logic: The Minimum You Need

Reading the table

Row by row for implication. P true and Q true: true. P true and Q false: FALSE, and this is the only row on which it is false. P false and Q true: true. P false and Q false: true.

That single false row is the whole content of implication, and it is the useful way to remember it: P ⊃ Q says "not (P true and Q false)".

Equivalence and exclusive or are exact opposites, which the last two columns show: ≡ is true where the values agree, ⊕ where they differ.

The three argument forms

Modus ponens. From P ⊃ Q and P, conclude Q. Checked on all four rows.

Modus tollens. From P ⊃ Q and not Q, conclude not P. Checked on all four rows.

Affirming the consequent. From P ⊃ Q and Q, conclude P. This is invalid, and the program reports the row that breaks it: P false and Q true. It is in the listing because a check that cannot fail has not been tested, and because this particular invalid form is the commonest fallacy in practical reasoning.

And it is a fallacy Nyāya already names. From "wherever there is smoke there is fire" and "there is fire", concluding "there is smoke" reverses the vyāpti. [Anumāna: Vyāpti, and the Three Kinds of Inference] states the rule that prevents it: the mark must be the narrower term.

De Morgan, and why to know it

The two laws checked in the listing are the ones you use constantly.

Not (P and Q) is the same as (not P) or (not Q).

Not (P or Q) is the same as (not P) and (not Q).

Negation turns and into or, and or into and. In practice that is how you simplify a condition: the negation of "the file exists and is readable" is "the file does not exist or is not readable", which is a different and correct test.

And the third law checked is the definition of implication in other terms: P ⊃ Q is the same as (not P) or Q. That is how implication is implemented when a language has no implication operator, which is every language.

munotes.in228

Propositional Logic: The Minimum You Need

What propositional logic cannot do

This is the section that motivates the next chapter, so it should be exact.

It cannot see inside a statement. "All students may borrow" and "Rama may borrow" are two unrelated atoms to it. There is no way to say that the second follows from the first.

So it cannot express a general rule. And a general rule is exactly what a vyāpti is. Propositional logic cannot state the example member of a Nyāya argument, which is the load-bearing member.

Nor can it express quantity. "Some", "all", "no" are invisible to it.

[Predicate Logic, and Why Anumāna Needs It] supplies what is missing.

What propositional logic IS good for

Combinational circuits. Every gate is a connective and every circuit is a formula. That is why a hardware paper teaches it.

Conditions in programs. Every if is a proposition, and de Morgan is how you simplify one.

Deciding validity mechanically. With n atoms the table has 2 to the power n rows, and checking them all is a decision procedure. It is exponential, and for small n it is instant, which is why the technique is used.

Quick revision

  • Atoms, connectives, assignments, truth tables. With n atoms, 2 to the power n rows.
  • Six connectives: ¬, •, ∨, ⊃, ≡, ⊕. Disjunction is inclusive; ⊕ is the exclusive one.
  • Implication is false on exactly one row, P true and Q false, and means nothing about relevance.
  • Modus ponens and modus tollens are valid; affirming the consequent is not, and fails at P false, Q true.
  • Affirming the consequent is the reversal of a vyāpti, which is why the mark must be the narrower term.
  • De Morgan: negation turns and into or and or into and. And P ⊃ Q is (not P) or Q.
  • It cannot look inside a statement, so it cannot state a general rule or a quantity.

Test yourself

1. On how many rows is P ⊃ Q false, and which?

On exactly one: P true and Q false. On every other assignment it is true.

2. Write the negation of "the file exists and is readable" using de Morgan.

The file does not exist, or it is not readable.

3. Show that affirming the consequent is invalid, naming the row.

Take P ⊃ Q and Q as premises and P as the conclusion. On the row where P is false and Q is true, both premises are true and the conclusion is false, so the form is invalid.

4. Why can propositional logic not express the example member of a Nyāya argument?

munotes.in229

Propositional Logic: The Minimum You Need

Because the example member states a general connection, "whatever is smoky is fiery", which quantifies over cases. Propositional logic treats each statement as an atom with no internal structure, so it has no way to say that a particular case falls under a general claim.

Contents This chapter on its own page

munotes.in230

Chapter Sixty-Six

Predicate Logic, and Why Anumāna Needs It

Syllabus topic Module 2, "Predicate logic"

In one line

Predicate logic can look inside a statement, so it can say "every" and "some", which is what a general rule needs.

In the wording you can write in an examination: predicate logic extends propositional logic by analysing a statement into a predicate and the terms it applies to, and by adding quantifiers. Its vocabulary is the constant, naming an individual; the variable, ranging over individuals; the predicate, expressing a property or a relation; the universal quantifier ∀, read "for every"; and the existential quantifier ∃, read "there is at least one".

The problem, stated exactly

Take the Nyāya argument and try to write it in propositional logic.

P = this hill is smoky

Q = this hill is fiery

R = whatever is smoky is fiery

Now try to derive Q from P and R. You cannot, because to propositional logic P, Q and R are three unrelated atoms. Nothing in the notation says that R has anything to do with P and Q.

And you cannot repair it by making R the implication P ⊃ Q, because R is not about this hill. R is about every smoky thing, and this hill is one of them. The notation has no way to express that relationship, because it cannot see that P and Q are both about this hill.

So the load-bearing member of a Nyāya argument is not expressible. That is the argument for predicate logic and it is the argument to give in an examination.

The vocabulary

Constant. A name for one individual. Write h for this hill.

Variable. A placeholder ranging over individuals. Write x.

Predicate. A property or a relation, written with its terms. Smoky(x) says x is smoky. Older(x, y) says x is older than y. A predicate with one term is a property; with two or more it is a relation.

Atomic formula. A predicate applied to terms: Smoky(h).

Quantifier. ∀x, for every x; ∃x, there is at least one x.

Scope. The part of the formula the quantifier governs, which the brackets fix.

The same argument, written

premise 1 Smoky(h)

premise 2 ∀x (Smoky(x) ⊃ Fiery(x))

conclusion Fiery(h)

Now the derivation is available, and it has two steps.

Universal instantiation. From ∀x (Smoky(x) ⊃ Fiery(x)) infer Smoky(h) ⊃ Fiery(h). What is true of every x is true of h.

Modus ponens. From Smoky(h) ⊃ Fiery(h) and Smoky(h) infer Fiery(h).

And that is the whole Nyāya argument. The example member is premise two, the reason is premise one, the application is the instantiation step, and the conclusion is the conclusion. [The Five Members Written in Logical Notation] works the correspondence out in full.

munotes.in231

Predicate Logic, and Why Anumāna Needs It

The four quantifier patterns

These four cover almost every general statement, and telling them apart is most of the skill.

EnglishIn symbolsTrue when
every S is P∀x (S(x) ⊃ P(x))no S fails to be P
some S is P∃x (S(x) • P(x))at least one thing is both
no S is P∀x (S(x) ⊃ ¬P(x))nothing is both
some S is not P∃x (S(x) • ¬P(x))at least one S fails to be P

Two traps, and both are worth stating.

The universal uses implication and the existential uses conjunction. Write ∀x (S(x) • P(x)) and you have said that everything whatever is both an S and a P, which is almost never what was meant. Write ∃x (S(x) ⊃ P(x)) and you have said something true whenever anything at all fails to be an S, which is almost always.

"Every S is P" is true when there are no S at all. ∀x (S(x) ⊃ P(x)) is satisfied vacuously if nothing is an S, because the implication is true whenever its antecedent is false. That is a consequence of the truth table in [Propositional Logic: The Minimum You Need] and it surprises everyone once.

Vyāpti, written properly

The Nyāya claim. There is no smoke without fire.

As a universal. ∀x (Smoky(x) ⊃ Fiery(x)).

As its negation, which is what destroys it. ∃x (Smoky(x) • ¬Fiery(x)). One thing that is smoky and not fiery.

That pair is the whole logic of the vipakṣa. The tradition says a single case of the mark without the conclusion destroys the connection; in this notation, the existential is the direct negation of the universal, so one such case makes the universal false. The classical rule and the logical fact are the same statement.

And the direction matters here just as it did there. ∀x (Fiery(x) ⊃ Smoky(x)) is a different formula and it is false: a red-hot iron ball is fiery and not smoky, which is the counterexample the Sūtra's own extra member offers.

Worked example: writing three claims

"All students may borrow books."

∀x (Student(x) ⊃ MayBorrow(x))

"Some students have not returned a book."

∃x (Student(x) • ∃y (Book(y) • Borrowed(x, y) • ¬Returned(x, y)))

Note the second quantifier inside the first, and the two-place predicates. This is where propositional logic stopped being an option: the relation between a student and a book cannot be said at all without terms.

"No student may borrow more than three books." This one cannot be written with the vocabulary above, because counting is not a quantifier. It needs either a numeric predicate or a much longer formula, and knowing that is worth as much as writing the other two.

munotes.in232

Predicate Logic, and Why Anumāna Needs It

What predicate logic still cannot do easily

Counting. "Exactly three" needs a long formula with distinctness conditions.

Defaults. "Students usually return books" has no natural form, which is the same limitation [Knowledge Representation: The Modern Name For It] records for logic as a family.

Quantifying over predicates. "There is a property that all students share" quantifies over properties, not individuals, which is second-order logic and a different system with different guarantees.

Deciding validity mechanically. Propositional validity is decidable by a truth table. Predicate validity is not decidable in general. That is the price of the extra expressiveness, and it is the same trade the Chomsky hierarchy records in [The Chomsky Hierarchy, and Where the Aṣṭādhyāyī Sits].

Quick revision

  • Predicate logic analyses a statement into a predicate and its terms, and adds quantifiers.
  • Vocabulary: constant, variable, predicate, atomic formula, quantifier, scope.
  • The Nyāya argument becomes universal instantiation followed by modus ponens.
  • Every S is P uses implication; some S is P uses conjunction. Swapping them is the standard error.
  • A universal is vacuously true when nothing satisfies its antecedent.
  • Vyāpti is a universal, and its negation is an existential, which is why one counterexample destroys it.
  • What is still hard: counting, defaults, quantifying over predicates. And validity is not decidable in general.

Test yourself

1. Why can propositional logic not express the Nyāya argument?

Because it treats "this hill is smoky", "this hill is fiery" and "whatever is smoky is fiery" as three unrelated atoms. It cannot see that the first two concern the same individual, so it cannot bring that individual under the general claim.

2. Write "every S is P" and "some S is P" in symbols and say why the connectives differ.

∀x (S(x) ⊃ P(x)) and ∃x (S(x) • P(x)). The universal uses implication because it makes a claim only about things that are S; the existential uses conjunction because it asserts that something is both.

3. Write vyāpti and its negation, and connect them to the vipakṣa rule.

∀x (Smoky(x) ⊃ Fiery(x)), whose negation is ∃x (Smoky(x) • ¬Fiery(x)). The negation asserts a single case of the mark without the conclusion, which is exactly the vipakṣa that the tradition says destroys the connection.

4. Give one thing predicate logic buys and one thing it costs, compared with propositional logic.

It buys the ability to state general rules and relations, so a particular case can be brought under a general claim. It costs decidability: propositional validity can be settled by a truth table, and predicate validity cannot be decided in general.

Contents This chapter on its own page

munotes.in233

Chapter Sixty-Seven

The Five Members Written in Logical Notation

Syllabus topic Module 2, "Propositional logic", "Predicate logic", "Five-member syllogism"

In one line

Translate the five members into logic and three of them turn into two steps of a proof; the other two are about the debate, not about the inference.

In the wording you can write in an examination: the five-member syllogism translates into predicate logic as a universal premise, a singular premise, an application of universal instantiation, and an application of modus ponens. The proposition and the conclusion state the same formula at the beginning and the end of the argument, so the logical content is two premises and one conclusion, and the remaining members are features of the presentation.

The translation, member by member

Take the standard argument and render each member.

MemberThe statementIn symbols
pratijñāThis hill is fieryFiery(h)
hetuBecause it is smokySmoky(h)
udāharaṇaWhatever is smoky is fiery, as a kitchen∀x (Smoky(x) ⊃ Fiery(x)) and Smoky(k) • Fiery(k)
upanayaSo is this hill: it is smokythe instantiation step
nigamanaTherefore this hill is fieryFiery(h)

Two things in that table need comment at once.

The example member becomes TWO formulas. The general connection, and the familiar instance. The instance is a separate assertion about the kitchen, and the argument does not use it: nothing is derived from Smoky(k) • Fiery(k). It is there as evidence for the general claim, not as a premise of the inference. That is the clearest single sign that the Nyāya form is doing more than stating a proof.

The application member becomes a STEP rather than a formula. It does not add information; it applies the general premise to this case.

The derivation

1. Smoky(h) premise, from the hetu

2. ∀x (Smoky(x) ⊃ Fiery(x)) premise, from the udaharana

3. Smoky(h) ⊃ Fiery(h) from 2, universal instantiation

4. Fiery(h) from 1 and 3, modus ponens

Four lines. Two premises, one instantiation, one modus ponens. That is the whole logical content of the five-member argument, and it is worth writing out because the compactness is the point: the classical form is longer than the proof, and the extra length is doing work of a different kind.

The negative form, translated

The Sūtra gives the argument about sound in a negative form as well, and it translates differently.

MemberThe statementIn symbols
pratijñāSound is not eternal¬Eternal(s)
hetuBecause it is producedProduced(s)
udāharaṇaWhatever is eternal is not produced, as the soul∀x (Eternal(x) ⊃ ¬Produced(x))
upanayaSound is not so, that is, not eternalthe instantiation and the inference
nigamanaTherefore sound is not eternal¬Eternal(s)

The derivation this time is modus tollens.

1. Produced(s) premise

2. ∀x (Eternal(x) ⊃ ¬Produced(x)) premise

3. Eternal(s) ⊃ ¬Produced(s) from 2, universal instantiation

4. ¬Eternal(s) from 1 and 3, modus tollens

munotes.in234

The Five Members Written in Logical Notation

So the affirmative example gives modus ponens and the negative example gives modus tollens. The two forms of the example member, homogeneous and heterogeneous at Nyāya Sūtra 1.1.36 and 1.1.37, correspond exactly to the two valid forms of [Propositional Logic: The Minimum You Need]. That correspondence is a genuine finding and it is worth stating in a Q.3 answer.

What is left over after the translation

Three things do not survive into the proof, and each of them is a real function of the classical form.

The familiar instance. Smoky(k) • Fiery(k) is asserted and never used. Logically it is inert. Epistemically it is the evidence for premise two, and premise two is the one that cannot be proved. A proof system has no slot for the evidence for a premise; the Nyāya form makes it a required member.

The restatement of the claim at the start. Logically redundant. In a debate it fixes what is being argued, so a party cannot shift ground, which is the nigrahasthāna of [Chala, Jāti and Nigrahasthāna: How a Debate Is Lost].

The audience. Nothing in a derivation supposes anybody is listening. The five-member form supposes an opponent and a judge, which is why [Debate Methodology as a Validation Protocol] treats it as a protocol rather than as a calculus.

So the honest summary is this. The Nyāya form contains a valid argument of predicate logic, plus machinery for arguing it in front of somebody. The translation shows exactly which parts are which.

The comparison, as a table

The predicate logic derivationThe five-member form
Linesfourfive members
States the claim firstnoyes
Evidence for the general premiseoutside the systema required member
Applicationa rule of inferencea stated member
Assumes an audiencenoyes
Can be checked mechanicallyyesno
What it establishesthat the conclusion followsthat the claim has been established

Worked example: where the translation exposes a defect

The argument. This deployment caused the outage, because the error rate rose, and whatever makes the error rate rise causes an outage, as last week's bad config.

Translated.

1. ErrorsRose(d)

2. ∀x (ErrorsRose(x) ⊃ CausedOutage(x))

3. ErrorsRose(d) ⊃ CausedOutage(d)

4. CausedOutage(d)

The derivation is valid. Universal instantiation and modus ponens, exactly as before. Nothing is wrong with the logic.

And premise two is false. A traffic spike raises the error rate and causes no outage, so ∃x (ErrorsRose(x) • ¬CausedOutage(x)) holds, and that existential is the direct negation of premise two.

This is the distinction to be able to state. Validity is a property of the derivation and it is mechanical. Soundness requires the premises to be true and it is not. The five-member form's contribution is to put the doubtful premise on the record as its own member, where it can be attacked; the derivation's contribution is to show that nothing else in the argument is at fault.

munotes.in235

The Five Members Written in Logical Notation

Quick revision

  • Proposition and conclusion are the same formula; the reason is a singular premise; the example is a universal premise plus an instance; the application is the instantiation step.
  • The affirmative example gives modus ponens; the negative example gives modus tollens. That matches 1.1.36 and 1.1.37 exactly.
  • The whole logical content is four lines: two premises, an instantiation, and one inference.
  • Three things do not survive translation: the familiar instance, which is evidence rather than a premise; the opening restatement, which fixes the claim; and the assumed audience.
  • Validity is mechanical and soundness is not, and the classical form's contribution is to put the doubtful premise on the record.

Test yourself

1. Write the derivation of the hill argument in four lines.

Smoky(h), premise. ∀x (Smoky(x) ⊃ Fiery(x)), premise. Smoky(h) ⊃ Fiery(h), by universal instantiation. Fiery(h), by modus ponens.

2. Which member becomes two formulas, and what happens to the second?

The example. It gives the universal connection and the familiar instance. The instance is asserted and never used in the derivation: it is evidence for the universal premise, not a premise of the inference.

3. What is the correspondence between the two forms of the example member and the two valid inference forms?

The homogeneous example, at 1.1.36, gives a universal whose instantiation supports modus ponens. The heterogeneous example, at 1.1.37, gives a universal in contrapositive form whose instantiation supports modus tollens.

4. A derivation is valid and the conclusion is false. Where is the fault, and what does the five-member form contribute?

In a premise, since a valid derivation from true premises cannot have a false conclusion. The five-member form contributes by making the doubtful general premise an explicit member of the argument, so that an opponent attacks it directly instead of guessing which step failed.

Contents This chapter on its own page

munotes.in236

Chapter Sixty-Eight

Inference Engines: Forward and Backward Chaining

Syllabus topic Module 2, "Inference engines"

In one line

An inference engine holds facts and rules, and it either pushes forward from the facts or works backward from a question.

In the wording you can write in an examination: an inference engine consists of a working memory of facts, a rule base of conditional rules, and a control strategy. Forward chaining, or data-driven inference, repeatedly fires every rule whose conditions are satisfied and adds its conclusion to working memory, until nothing new is derived. Backward chaining, or goal-driven inference, takes a goal, finds the rules that conclude it, and recursively attempts to establish their conditions.

The three parts

Working memory. The facts currently held. It grows during forward chaining and does not change during backward chaining.

Rule base. Conditional rules, each with conditions and a conclusion. They are data, not code, which is what makes a rule-based system modifiable without reprogramming.

Control strategy. Which direction to run, and, when several rules could fire at once, which to fire. That second question is conflict resolution and it is the same problem [Rule Precedence: The Four Principles] solves for Pāṇini.

The two directions

Forward chainingBackward chaining
Starts fromthe factsa goal
Askswhat follows?can this be shown?
Stops whennothing new is derivedthe goal is established or exhausted
Deriveseverything derivableonly what the goal needs
Good whenthere are few facts and many possible conclusionsthere is one question and many facts
Natural formonitoring, alerting, reacting to eventsdiagnosis, answering a query
Wasteful whenmost conclusions are not wantedmost facts are relevant to the goal

Both directions, on one rule base

"""Forward and backward chaining over one rule base, traced."""

RULES = [
    ("R1", ["smoky"],              "fiery"),
    ("R2", ["fiery", "enclosed"],  "hot"),
    ("R3", ["hot", "dry"],         "fire-risk"),
    ("R4", ["wet"],                "not-dry"),
]
FACTS = {"smoky", "enclosed", "dry"}


def forward(facts, rules):
    known, trail = set(facts), []
    changed = True
    while changed:
        changed = False
        for name, conditions, conclusion in rules:
            if conclusion in known:
                continue
            if all(c in known for c in conditions):
                known.add(conclusion)
                trail.append("%s fires: %s so %s" % (name, " and ".join(conditions), conclusion))
                changed = True
    return known, trail


def backward(goal, facts, rules, depth=0, trail=None, seen=None):
    trail = trail if trail is not None else []
    seen = seen or set()
    pad = "  " * depth
    if goal in facts:
        trail.append("%sgoal %s is a known fact" % (pad, goal))
        return True, trail
    if goal in seen:
        trail.append("%sgoal %s is already being pursued: a circle" % (pad, goal))
        return False, trail
    applicable = [r for r in rules if r[2] == goal]
    if not applicable:
        trail.append("%sgoal %s: no rule concludes it, and it is not a fact" % (pad, goal))
        return False, trail
    for name, conditions, _ in applicable:
        trail.append("%sgoal %s: try %s, which needs %s" % (pad, goal, name, " and ".join(conditions)))
        if all(backward(c, facts, rules, depth + 1, trail, seen | {goal})[0] for c in conditions):
            trail.append("%sgoal %s: established by %s" % (pad, goal, name))
            return True, trail
        trail.append("%sgoal %s: %s did not establish it" % (pad, goal, name))
    return False, trail


print("facts:", ", ".join(sorted(FACTS)))
print("rules:")
for name, conds, concl in RULES:
    print("  %-3s IF %-22s THEN %s" % (name, " and ".join(conds), concl))

print()
print("FORWARD chaining: derive everything derivable")
known, trail = forward(FACTS, RULES)
for line in trail:
    print("  " + line)
print("  derived:", ", ".join(sorted(known - FACTS)) or "(nothing)")

print()
print("BACKWARD chaining: prove one goal, fire-risk")
ok, trail = backward("fire-risk", FACTS, RULES)
for line in trail:
    print("  " + line)
print("  result:", "established" if ok else "not established")

print()
print("BACKWARD chaining on a goal that fails: not-dry")
ok, trail = backward("not-dry", FACTS, RULES)
for line in trail:
    print("  " + line)
print("  result:", "established" if ok else "not established")

print()
print("how much work each does")
_, ftrail = forward(FACTS, RULES)
_, btrail = backward("fire-risk", FACTS, RULES)
print("  forward  fired %d rule(s) and derived %d new fact(s)" % (len(ftrail), len(known - FACTS)))
print("  backward took %d step(s) to prove one goal" % len(btrail))
munotes.in237

Inference Engines: Forward and Backward Chaining

facts: dry, enclosed, smoky
rules:
  R1  IF smoky                  THEN fiery
  R2  IF fiery and enclosed     THEN hot
  R3  IF hot and dry            THEN fire-risk
  R4  IF wet                    THEN not-dry

FORWARD chaining: derive everything derivable
  R1 fires: smoky so fiery
  R2 fires: fiery and enclosed so hot
  R3 fires: hot and dry so fire-risk
  derived: fiery, fire-risk, hot

BACKWARD chaining: prove one goal, fire-risk
  goal fire-risk: try R3, which needs hot and dry
    goal hot: try R2, which needs fiery and enclosed
      goal fiery: try R1, which needs smoky
        goal smoky is a known fact
      goal fiery: established by R1
      goal enclosed is a known fact
    goal hot: established by R2
    goal dry is a known fact
  goal fire-risk: established by R3
  result: established

BACKWARD chaining on a goal that fails: not-dry
  goal not-dry: try R4, which needs wet
    goal wet: no rule concludes it, and it is not a fact
  goal not-dry: R4 did not establish it
  result: not established

how much work each does
  forward  fired 3 rule(s) and derived 3 new fact(s)
  backward took 9 step(s) to prove one goal

Reading the two traces

Forward chaining fired three rules and derived three facts. It did not fire R4, because its condition was never satisfied. It stopped when a whole pass added nothing.

Backward chaining took nine steps to establish one goal. It went down from fire-risk to hot to fiery to smoky, hit a known fact, and came back up. The indentation is the structure of the proof, and reading it upward gives the derivation in the order a person would present it.

munotes.in238

Inference Engines: Forward and Backward Chaining

And the failing goal is the useful one. Asking for not-dry finds R4, needs wet, finds no rule concluding wet and no such fact, and reports failure with the reason. An engine that says only "no" is unusable; one that says which subgoal it could not establish is a diagnostic tool.

Which direction to choose

The rule of thumb is about the ratio of facts to conclusions.

Choose forward chaining when the facts arrive and you want to know what they imply. A monitoring system receives events and must decide whether to alert. It has no goal in advance.

Choose backward chaining when there is a question. A diagnostic system is asked "is it the disk?" and should not derive everything derivable about the machine first.

And notice what the traces show about cost. Forward chaining derived three facts, all of which happened to be wanted. Add fifty rules whose conclusions nobody asked about and forward chaining derives all fifty; backward chaining touches none of them.

The Nyāya connection, stated exactly

Backward chaining is the shape of a Nyāya argument. You have a claim, you look for a vyāpti that would establish it, and you check whether the subject satisfies the vyāpti's condition. Goal, rule, subgoal.

And the trace is the five-member form, upside down. Read the backward trace from the innermost line outward and you get: this is smoky, whatever is smoky is fiery, so this is fiery, and so on up. [Nyāya Logic as an Inference Engine] builds an engine that emits exactly that, in the five members.

Forward chaining has no classical counterpart in Nyāya, because Nyāya is about establishing a claim somebody has made, not about generating claims. The nearest thing is Charaka's diagnostic procedure, which gathers evidence before forming a hypothesis, and even that is not the same.

Conflict resolution, and the connection to Module I

When two rules could fire at once a forward-chaining engine must choose. The classical strategies are these.

Specificity. Fire the rule with more conditions, that is, the more specific one. That is apavāda, from [Rule Precedence: The Four Principles].

Recency. Fire the rule whose triggering facts were added most recently.

Order. Fire the rule that appears first, or last, in the rule base. That is paratva, sūtra 1.4.2.

Refraction. Do not fire the same rule on the same facts twice, which is what the if conclusion in known: continue line in the listing achieves, and it is what stops the loop.

Two of the four strategies are Pāṇini's, and an answer that says so has made a real cross-module connection.

munotes.in239

Inference Engines: Forward and Backward Chaining

What an inference engine is NOT

It is not a theorem prover. It applies rules; it does not search for proofs in a logic with quantifiers, and the rules here have no variables at all.

It is not guaranteed to terminate. This one does, because refraction prevents a rule firing twice on the same facts and the set of derivable facts is finite. A rule base that can generate new terms need not terminate.

It does not handle uncertainty. Every fact is present or absent and every rule fires or does not. Real expert systems attach certainty factors, and [Expert Systems, and MYCIN as the Comparison] records what that cost.

It does not decide what the rules should be. Getting the rules out of an expert's head is knowledge acquisition, and it is the hard part.

Quick revision

  • Three parts: working memory, rule base, control strategy.
  • Forward chaining: from facts, derive everything, stop when a pass adds nothing. Good for monitoring.
  • Backward chaining: from a goal, find rules concluding it, recurse on their conditions. Good for diagnosis.
  • The backward trace, read outward, is a proof; a failing goal should report which subgoal failed.
  • Conflict resolution strategies: specificity, recency, order, refraction. Specificity is apavāda and order is paratva.
  • Refraction is what prevents the loop.
  • Not a theorem prover, not guaranteed to terminate in general, and no notion of uncertainty.

Test yourself

1. Distinguish forward from backward chaining by what each starts from and what each derives.

Forward chaining starts from the facts and derives everything derivable, stopping when a whole pass adds nothing. Backward chaining starts from a goal and derives only what that goal needs, stopping when the goal is established or the possibilities are exhausted.

2. Which would you use for an alerting system, and which for a diagnostic tool?

Forward chaining for alerting, because events arrive and there is no goal in advance. Backward chaining for diagnosis, because there is a question and deriving everything about the system first would be wasteful.

3. Why must a failing backward-chaining query report more than "no"?

Because the useful information is which subgoal could not be established. An engine that reports only failure gives a user nothing to act on; one that names the missing fact or the missing rule is a diagnostic tool.

4. Name two conflict resolution strategies that correspond to Pāṇini's precedence principles.

Specificity, firing the rule with the narrower conditions, corresponds to apavāda. Rule order, firing by position in the rule base, corresponds to paratva, sūtra 1.4.2.

Contents This chapter on its own page

munotes.in240

Chapter Sixty-Nine

Nyāya Logic as an Inference Engine

Syllabus topic Module 2, "Inference engines", "Algorithm Specification (Pseudo-code)", "minimum 10 test cases"

In one line

Give the engine facts and invariable connections, ask it a question, and it answers in five members.

In the wording you can write in an examination: a Nyāya inference engine represents each vyāpti as a rule carrying the mark, the conclusion and the familiar instance required by the example member, and each observation as a fact about a subject. It answers a query by backward chaining, and renders the successful chain as a sequence of five-member arguments, the innermost first, so that the output is a proof rather than a verdict.

Problem statement, in MU's own form

IKS concept as CS concept: the Nyāya five-member syllogism as a backward-chaining inference engine that emits its proof.

Statement. Represent invariable connections and observations. Answer a query about a subject by backward chaining over the connections. Render each step of the successful chain in the five members of Nyāya Sūtra 1.1.32, nested so that a step used to establish another is printed inside it. Reject circular rule sets, unsupported claims and connections applied in the wrong direction.

Conceptual mapping table

Classical elementComputer science element
pakṣa, the subjectthe individual the query is about
sādhya, what is to be establishedthe goal
hetu, the marka fact observed of the subject
vyāpti, the invariable connectiona rule, mark implies conclusion
dṛṣṭānta, the familiar instancea field on the rule, carried into the output
the five membersthe rendering of one successful step
a chain of inferencesnested calls, printed innermost first
a circular argumenta goal already on the stack, refused

Algorithm specification, in pseudo-code

ALGORITHM Prove(subject, claim, seen)

if (subject, claim) is already in seen then return nothing -- a circle

add (subject, claim) to seen

for each rule whose RESULT is the claim do

if the subject is observed to have the rule's MARK then

return the five members, with no inner steps

inner <- Prove(subject, the rule's MARK, seen)

if inner is not nothing then

return the five members, carrying inner as a nested step

end for

return nothing

Two design points to defend in a viva. The circle check is on the pair of subject and claim, not on the claim alone, because the same claim may legitimately be pursued about two different subjects. And a rule is tried only if its RESULT matches the goal, which is what makes the search goal-directed rather than exhaustive.

Working code

#!/usr/bin/env python3
"""Nyaya logic as an inference engine. MU's implementation topic 3.

IKS concept as CS concept: the five-member syllogism (pancavayava) of Nyaya
Sutra 1.1.32 as a backward-chaining inference engine that emits its proof.

The five members, in Vidyabhusana's translation of 1.1.32 to 1.1.39:
  pratijna   proposition   the claim to be established
  hetu       reason        the mark observed in the subject
  udaharana  example       the invariable connection, with a familiar instance
  upanaya    application   the subject brought under the connection
  nigamana   conclusion    the proposition restated as established

A rule here is a vyapti, an invariable connection, plus the familiar instance
the Sutra requires an example to carry. A fact is a mark observed of a subject.
"""
from collections import namedtuple

Vyapti = namedtuple('Vyapti', 'mark result instance')


class Nyaya:
    def __init__(self):
        self.facts = set()        # (subject, mark)
        self.rules = []           # Vyapti

    def observe(self, subject, mark):
        self.facts.add((subject, mark))
        return self

    def connect(self, mark, result, instance):
        """"Whatever is <mark> is <result>, as <instance>." (udaharana)"""
        self.rules.append(Vyapti(mark, result, instance))
        return self

    def prove(self, subject, claim, seen=None):
        """Backward chaining. Returns the five members, or None."""
        seen = seen or set()
        if (subject, claim) in seen:
            return None                      # a circle is not a proof
        seen = seen | {(subject, claim)}
        for r in self.rules:
            if r.result != claim:
                continue
            if (subject, r.mark) in self.facts:
                return self._members(subject, claim, r, [])
            sub = self.prove(subject, r.mark, seen)
            if sub:
                return self._members(subject, claim, r, sub['steps'] + [sub])
        return None

    @staticmethod
    def _members(subject, claim, r, steps):
        return {
            'pratijna':  '%s is %s.' % (subject.capitalize(), claim),
            'hetu':      'Because it is %s.' % r.mark,
            'udaharana': 'Whatever is %s is %s, as %s.' % (r.mark, r.result, r.instance),
            'upanaya':   'So is %s: it is %s.' % (subject, r.mark),
            'nigamana':  'Therefore %s is %s.' % (subject, claim),
            'steps':     steps,
        }

    @staticmethod
    def render(proof, indent=0):
        pad = ' ' * indent
        out = []
        for step in proof['steps']:
            out.append(Nyaya.render(step, indent + 2))
        for k in ('pratijna', 'hetu', 'udaharana', 'upanaya', 'nigamana'):
            out.append('%s%-10s %s' % (pad, k, proof[k]))
        return '\n'.join(out)


def classical():
    return (Nyaya()
            .observe('this hill', 'smoky')
            .connect('smoky', 'fiery', 'a kitchen'))


def sutra_example():
    return (Nyaya()
            .observe('sound', 'produced')
            .connect('produced', 'non-eternal', 'a pot'))


def chained():
    """Two links, so the engine has to chain and the trace has to nest."""
    return (Nyaya()
            .observe('this program', 'a recursion with no base case')
            .connect('a recursion with no base case', 'non-terminating', 'a loop on itself')
            .connect('non-terminating', 'incorrect', 'a sort that never returns'))


TESTS = [
    ('the classical inference',            classical(),      'this hill',    'fiery',        True),
    ('the Sutra\'s own example',           sutra_example(),  'sound',        'non-eternal',  True),
    ('a two-link chain',                   chained(),        'this program', 'incorrect',    True),
    ('the middle claim of that chain',     chained(),        'this program', 'non-terminating', True),
    ('no mark observed',                   classical(),      'the lake',     'fiery',        False),
    ('no connection for the claim',        classical(),      'this hill',    'wet',          False),
    ('the connection runs the other way',  sutra_example(),  'a pot',        'non-eternal',  False),
    ('a circular rule set',
     Nyaya().observe('x', 'a').connect('b', 'c', 'i').connect('c', 'b', 'j'), 'x', 'b', False),
    ('an empty engine',                    Nyaya(),          'anything',     'anything',     False),
    ('the same subject, a second claim',
     classical().connect('fiery', 'hot', 'a lamp'), 'this hill', 'hot',     True),
]


def run_tests(verbose=False):
    passed = 0
    for name, engine, subject, claim, expect in TESTS:
        proof = engine.prove(subject, claim)
        got = proof is not None
        ok = got == expect
        passed += ok
        if verbose:
            print('%-38s %s  expected %-5s got %s'
                  % (name, 'pass' if ok else 'FAIL', expect, got))
        assert ok, name
    return passed


if __name__ == '__main__':
    n = run_tests(verbose=True)
    print('\n%d of %d test cases pass\n' % (n, len(TESTS)))
    print('--- the classical inference, in five members ---')
    print(Nyaya.render(classical().prove('this hill', 'fiery')))
    print('\n--- a two-link chain, the inner proof first ---')
    print(Nyaya.render(chained().prove('this program', 'incorrect')))
munotes.in241

Nyāya Logic as an Inference Engine

the classical inference                pass  expected True  got True
the Sutra's own example                pass  expected True  got True
a two-link chain                       pass  expected True  got True
the middle claim of that chain         pass  expected True  got True
no mark observed                       pass  expected False got False
no connection for the claim            pass  expected False got False
the connection runs the other way      pass  expected False got False
a circular rule set                    pass  expected False got False
an empty engine                        pass  expected False got False
the same subject, a second claim       pass  expected True  got True

10 of 10 test cases pass

--- the classical inference, in five members ---
pratijna   This hill is fiery.
hetu       Because it is smoky.
udaharana  Whatever is smoky is fiery, as a kitchen.
upanaya    So is this hill: it is smoky.
nigamana   Therefore this hill is fiery.

--- a two-link chain, the inner proof first ---
  pratijna   This program is non-terminating.
  hetu       Because it is a recursion with no base case.
  udaharana  Whatever is a recursion with no base case is non-terminating, as a loop on itself.
  upanaya    So is this program: it is a recursion with no base case.
  nigamana   Therefore this program is non-terminating.
pratijna   This program is incorrect.
hetu       Because it is non-terminating.
udaharana  Whatever is non-terminating is incorrect, as a sort that never returns.
upanaya    So is this program: it is non-terminating.
nigamana   Therefore this program is incorrect.
munotes.in242

Nyāya Logic as an Inference Engine

Reading the output

Ten test cases, and five of them must fail. No mark observed; no connection for the claim; the connection applied in the wrong direction; a circular rule set; and an engine with no rules and no facts at all. A rejection is a result, and an engine that has only been shown succeeding has not been shown working.

The third failing case is the important one. The rule says whatever is produced is non-eternal. Asking whether a pot, which is non-eternal, is therefore produced reverses the vyāpti, and the engine refuses because no rule concludes what was asked from what is known. That is the direction rule of [Anumāna: Vyāpti, and the Three Kinds of Inference] enforced by the data structure: a rule has a mark field and a result field, and they are not interchangeable.

munotes.in243

Nyāya Logic as an Inference Engine

The circular case is the fourth. Two rules, one concluding b from c and one concluding c from b, with neither observed. Without the seen set the engine would recurse for ever; with it, the second attempt on the same pair returns nothing and the search unwinds.

And the nested trace is the payoff. For the two-link chain the inner argument is printed first, indented, and the outer argument follows. Read it top to bottom and you have the proof in the order a person would present it, with each general connection stated as its own member.

What makes this a Nyāya engine rather than a generic one

Three things, and all three are choices the classical form forces.

The rule carries its instance. An ordinary rule is a condition and a conclusion. A vyāpti must also carry the familiar instance that the udāharaṇa requires, so the rule has a third field and the output prints it. A rule with no instance is not a well-formed premise in this tradition, and making the field mandatory encodes that.

The output is the five members and nothing else. No confidence, no score, no rule identifier. The tradition's answer to "why?" is a specific five-part sentence, and the engine gives exactly that.

The claim is restated at the start and at the end. Logically redundant, as [The Five Members Written in Logical Notation] shows, and retained because the pratijñā and the nigamana are what make the answer a statement rather than a calculation.

Complexity and limitations

Time. Each call scans the rule base once for rules concluding the goal, and recurses on at most one mark per rule. With r rules and a chain of depth d the work is proportional to r times d in the successful case, and to r raised to d in the worst case if many rules conclude the same goal and all fail. The seen set bounds the recursion by the number of distinct subject and claim pairs.

Space. The seen set, and the recursion depth, both bounded by the number of distinct pairs.

Limitation: no variables. A rule relates one mark to one conclusion and a fact concerns one subject. There is no way to write a rule about a relation between two individuals, which predicate logic does and which [Predicate Logic, and Why Anumāna Needs It] shows is needed for anything about two things at once.

Limitation: no negation. Nothing expresses that a mark is ABSENT, so the negative form of the example member, Nyāya Sūtra 1.1.37, cannot be represented. That is a real gap: the tradition treats the heterogeneous example as on a par with the homogeneous one.

munotes.in244

Nyāya Logic as an Inference Engine

Limitation: only the first successful rule is reported. If two connections would both establish the claim, the engine returns one proof. The tradition would regard a second independent support as strengthening the case, as the dialogue in [Vāda, Jalpa and Vitaṇḍā: Three Kinds of Dispute] shows, and this engine does not look for one.

Limitation: no fallacy detection. Nothing checks whether a connection is erratic. The engine takes every rule it is given as a genuine vyāpti, and [Hetvābhāsa: The Five Fallacies of the Reason] is entirely outside it.

Quick revision

  • Facts are observations of a subject; rules are vyāpti carrying a mark, a result and a familiar instance.
  • Backward chaining: find rules concluding the goal, check the mark, recurse if it is not observed.
  • The circle check is on the subject-and-claim pair, so the same claim may be pursued about two subjects.
  • Ten test cases, five of which must fail: no mark, no rule, the connection reversed, a circular rule set, and an empty engine.
  • The rule's third field is the familiar instance, which encodes that a premise with no instance is not well formed.
  • Limitations: no variables, no negation so no heterogeneous example, only the first proof, and no fallacy detection.

Test yourself

1. Why does the rule structure have three fields rather than two?

Because a vyāpti used in a five-member argument must supply the familiar instance that the udāharaṇa requires. Making the instance a mandatory field encodes the rule that a general premise with no instance attached is not well formed.

2. What does the engine do when asked whether a pot is produced, given that whatever is produced is non-eternal?

It refuses. The rule's mark is "produced" and its result is "non-eternal", and nothing concludes "produced". Reversing the connection is exactly what the direction rule of vyāpti forbids, and the separate mark and result fields enforce it.

3. What is the seen set for, and why is it keyed on a pair?

It prevents a circular rule set sending the search into infinite recursion. It is keyed on the subject and the claim together because the same claim may legitimately be pursued about a different subject during the same search.

4. Name two limitations that belong in this implementation's own limitations section.

It has no negation, so the heterogeneous example of Nyāya Sūtra 1.1.37 cannot be represented at all. And it performs no fallacy detection: every rule it is given is treated as a genuine invariable connection, so an erratic reason would be used without complaint.

Contents This chapter on its own page

munotes.in245

Chapter Seventy

Explainable AI, and Why a Five-Member Answer Is an Explanation

Syllabus topic Module 2, "Explainable AI"

In one line

Explainable AI is the requirement that a system be able to say why, and the five-member form is a specification of what "why" should contain.

In the wording you can write in an examination: explainable artificial intelligence is the field concerned with making the decisions of automated systems intelligible to the people affected by them. A model is intrinsically interpretable if its own structure can be read as a reason, and a post-hoc explanation is a separate account generated after a decision by a model whose structure cannot be read. The Nyāya five-member form is a specification of the content an explanation should carry.

The two kinds of explanation

Intrinsically interpretable. The model's structure IS the explanation. A decision tree's path from root to leaf is a sequence of tests, and reading it gives the reason. A rule-based system's firing chain is the reason. Nothing extra has to be produced.

Post-hoc. The model's structure cannot be read as a reason, so a second procedure produces one. A neural network has millions of weights; an explanation of one of its decisions is generated by a separate method, for instance by identifying the inputs whose change would most alter the output.

The difference that matters. An intrinsic explanation is guaranteed to be the actual reason, because it IS the mechanism. A post-hoc explanation is a hypothesis about the mechanism, produced by a different procedure, and it can be wrong about the model it explains. That is not a small caveat: it is the central problem of the field.

Why anybody requires it

Four reasons, and a question may ask for any of them.

To act on the decision. A classification with no reason cannot be followed up. A diagnostic system that says "fault type three" and nothing else leaves the engineer where they started.

To find the model's mistakes. A model that reaches the right answer for the wrong reason will fail on the next case, and nothing but an explanation reveals it.

To contest the decision. A person refused something is entitled to know why, and in several jurisdictions that is a legal requirement rather than a courtesy.

To transfer the knowledge. A reason can be learned by a person; a weight matrix cannot.

What the five-member form supplies

This is the chapter's argument, and it is worth setting out precisely, because "ancient India had explainable AI" is exactly the kind of claim this book refuses to make loosely.

The claim being made is narrow. Not that Nyāya anticipated the field. That the five members enumerate four things an explanation must contain, and that most modern explanations supply fewer.

MemberWhat it suppliesDoes a confidence score supply it?
hetuthe specific evidence, in this caseno
udāharaṇathe general rule the evidence is being used underno
the instance in the udāharaṇaa case where the general rule is known to holdno
upanayathe assertion that this case falls under that ruleno
nigamanathe conclusion, stated as establishedyes, this is all a score gives
munotes.in246

Explainable AI, and Why a Five-Member Answer Is an Explanation

Read the right-hand column. A model that outputs "fault type three, confidence 0.87" has supplied only the last row. It has not said what in this input led to that, nor what general regularity it is relying on, nor whether that regularity has ever been checked against a known case, nor that this input actually falls under it.

And the third row is the one nobody supplies. Naming a case where the general rule is known to hold is the udāharaṇa's distinctive contribution, and it is what [Five Members Against Aristotle's Three] identifies as the form's addition. In modern practice the nearest thing is a nearest-neighbour example shown alongside a prediction, and it is not standard.

Worked example: the same decision, four ways

The decision. A system flags a transaction as fraudulent.

No explanation. "Flagged."

A confidence score. "Flagged, confidence 0.91." Nothing has been added that a person can act on.

A post-hoc explanation. "Flagged; the features contributing most were the amount and the time of day." This is a statement about the model, produced by a second procedure, and it does not say what the model believes about amounts and times of day.

A five-member explanation. "This transaction is fraudulent. Because it is a first transaction on a new card, above the account's historical maximum, from a country the account has never used. Whatever has all three of these is fraudulent, as the confirmed case of 14 March. So is this transaction: it has all three. Therefore this transaction is fraudulent."

Compare the last two. The post-hoc explanation names features. The five-member explanation names the general rule, so the rule can be disputed. A reviewer can say: the rule is wrong, plenty of legitimate first transactions are above the historical maximum from a new country. That objection is only available because the rule was stated.

The honest limits of the comparison

The five-member form is not an algorithm for producing explanations. It says what an explanation should contain. Getting that content out of a model that does not have it is the entire difficulty, and the classical form contributes nothing to it.

A system that can produce it is already interpretable. A rule-based system can, as [Nyāya Logic as an Inference Engine] shows, because the rules are there to be quoted. A neural network cannot, because there is no general rule in it to quote.

munotes.in247

Explainable AI, and Why a Five-Member Answer Is an Explanation

So the form is a specification, not a solution. Saying so is what distinguishes an answer that has understood the material from one that has not.

And there is a cost the tradition did not face. A five-member explanation for each of a million decisions is a million paragraphs. Modern explanation has a volume problem the classical form never had to consider.

What explainable AI is NOT

It is not the same as accuracy. An explainable model may be less accurate than an unexplainable one, and choosing between them is a real trade rather than an oversight.

It is not the same as transparency about the code. Publishing the source of a model does not explain a decision. The decision is a consequence of the weights, which the source does not contain.

It is not solved. Post-hoc explanation methods disagree with each other on the same model and the same input, which is a known and unresolved problem.

Quick revision

  • Intrinsically interpretable: the structure is the reason, as in a decision tree or a rule chain. Post-hoc: a second procedure produces an account, which may be wrong about the model.
  • Four reasons to require it: to act, to find mistakes, to contest, to transfer knowledge.
  • The five members enumerate: the evidence, the general rule, a case where the rule is known to hold, the subsumption, and the conclusion. A confidence score supplies only the last.
  • Stating the general rule is what makes the decision disputable, which a feature attribution does not achieve.
  • The form is a specification of content, not a method of producing it, and a model that can meet it is already interpretable.
  • Not the same as accuracy, not the same as publishing the code, and not solved.

Test yourself

1. Distinguish an intrinsically interpretable model from a post-hoc explanation, and say which is more trustworthy.

An interpretable model's own structure is the reason, so the explanation is the mechanism. A post-hoc explanation is produced by a separate procedure after the fact and is a hypothesis about the mechanism, so it can be wrong about the model it explains. The first is more trustworthy for that reason.

2. What do the five members supply that a confidence score does not?

The specific evidence, the general rule being relied on, a case in which that rule is known to hold, and the assertion that this case falls under it. A confidence score supplies only the conclusion.

3. Why does naming the general rule matter more than naming the contributing features?

Because a rule can be disputed. A reviewer can say the rule is wrong and give a counterexample, which is an objection that is only available once the rule has been stated. A list of contributing features names no claim to contest.

munotes.in248

Explainable AI, and Why a Five-Member Answer Is an Explanation

4. State the limit of the comparison between the five-member form and explainable AI.

The form specifies what an explanation should contain and offers no method of producing it. A model that can meet the specification is already interpretable, and the hard problem is exactly the case where it cannot.

Contents This chapter on its own page

munotes.in249

Chapter Seventy-One

Āyurveda as a Śāstra, and What This Chapter Does Not Claim

Syllabus topic Module 2, "Ayurvedic Classification as Rule-Based Expert System", "Examination of structured diagnostic reasoning in Ayurvedic texts"

In one line

Āyurveda is a śāstra with a very large literature, and this paper reads a small part of it as a structured classification scheme.

In the wording you can write in an examination: Āyurveda is the classical Indian system of medicine, transmitted principally in three treatises, of which the Charaka Saṃhitā is the one concerned with general medicine. The syllabus studies its diagnostic reasoning as a formal structure: a fixed set of named attributes, a classification of conditions by those attributes, and rules connecting a classification to a response. What is examined is the structure of the knowledge, not its clinical validity.

What this block does NOT claim

This section is first on purpose and every sentence in it is binding on the rest of the block.

Nothing here is medical advice. No statement in these chapters should be acted on in relation to anybody's health.

Nothing here says a classical classification is clinically correct. Whether the tridoṣa scheme describes anything real is a medical question. This paper does not ask it and does not answer it.

Nothing here is a validated classifier. [Ayurvedic Classification as a Rule-Based Expert System] builds a program from the attribute lists in the text. It is a knowledge-representation exercise on a historical source. It is not diagnostic software and it is not evaluated against any clinical outcome.

And the programs are not evaluated for accuracy at all, because there is nothing to evaluate them against. An accuracy figure would require a labelled dataset with a ground truth, and none exists for this. A submission that reported an accuracy would be reporting a number it had invented.

Why this matters for your own internal assessment. MU's topic 4 is Ayurvedic classification as a rule-based expert system. A submission that describes its output as a diagnosis has overclaimed; one that describes it as a classification into the categories the text uses, with a limitations section saying exactly the above, has not.

What Āyurveda is, structurally

Three principal treatises, known collectively as the great triad.

TreatiseIts subjectThe translation quoted here
Charaka Saṃhitāgeneral medicineAvinash Chandra Kaviratna, from 1890
Suśruta SaṃhitāsurgeryKaviraj Kunja Lal Bhishagratna, 1907
Aṣṭāṅga Hṛdayaa later synthesis of bothnot quoted in this book

The Charaka Saṃhitā's own divisions, which you need in order to read a citation. It is arranged in sections called sthāna, each divided into lessons. The two this book quotes are the Sūtrasthāna, the section of general principles, and the Vimānasthāna, the section on measurement and method, which is where the account of examination sits.

A citation therefore names the section and the lesson, and this book gives verse numbers where Kaviratna prints them.

munotes.in250

Āyurveda as a Śāstra, and What This Chapter Does Not Claim

Why a computer science syllabus reads it

MU's label is "Ayurvedic Classification as Rule-Based Expert System", and her CS concepts under it are decision trees, rule-based systems, expert systems, feature engineering and multi-class classification. So the interest is in four structural features, and this book establishes each from the text.

A closed set of named attributes. [The Tridoṣa Framework] quotes seven attributes for each of three categories, and [Doṣa as a Feature Vector] sets them out as vectors over a shared space.

A small set of class labels. Three, with combinations. [Prakṛti: Constitution as a Class Label].

A declared order of evidence gathering. [Parīkṣā: How an Examination Is Structured] quotes Charaka on the three means of knowledge and the order in which they are used.

A rule connecting a classification to a response. The rule of the adverse attribute, quoted in [The Tridoṣa Framework] and treated in [Decision Principles in Diagnosis].

Four features, and together they are the shape of an expert system. That is the honest form of the claim: not that the text is an expert system, but that these four structural features are the ones an expert system has.

The historical claim this book does make

That the texts are highly formalised. They define their terms, enumerate their attributes, state their rules and declare their evidence. That is verifiable from the text and this block verifies it.

And that formalisation is not validity. [Śāstra: What Makes a Body of Knowledge Formal] makes the point in general and it applies here with particular force: a scheme can be rigorously specified and empirically wrong, and the two questions are independent.

Worked example: reading a citation in this block

A citation in these chapters names the treatise, the section, the lesson and, where Kaviratna prints one, the verse. Take the one the next chapter rests on.

Charaka Saṃhitā, Sūtrasthāna, Lesson I, verse 56.

Charaka Saṃhitā is the treatise, on general medicine. Sūtrasthāna is the section of general principles, as against the Vimānasthāna on measurement and method. Lesson I is the lesson within that section. Verse 56 is the verse as Kaviratna numbers it in his translation, which is the edition this book quotes.

And what it does not name. It does not name a page, because editions repaginate, and it does not name the Sanskrit, because this book quotes the English translation and says so. A reader with a different translation can find the same verse; a reader with a different edition of the same translation can too.

Why the section matters as much as the lesson. The same lesson number occurs in every section, so "Lesson I" alone is ambiguous between six or more places in the treatise. A citation that omits the section cannot be followed, and it is the commonest defect in student writing about this material.

munotes.in251

Āyurveda as a Śāstra, and What This Chapter Does Not Claim

What is in the rest of this block

ChapterWhat it establishes
[The Tridoṣa Framework]the three categories and seven attributes each, quoted
[Doṣa as a Feature Vector]the same as vectors over a shared attribute space
[Prakṛti: Constitution as a Class Label]the label set, and why it is not binary
[Parīkṣā: How an Examination Is Structured]the declared order of evidence
[Symptom to Feature Mapping]turning a report into an attribute with a value
[Multi-Attribute Classification]combining attributes
[Decision Principles in Diagnosis]the order of questions and the treatment rule
[Decision Trees]the modern method, taught on its own terms
[A Decision Tree Built From the Tridoṣa Attributes]the method applied to the quoted table
[Rule-Based Systems]IF-THEN rules and the recognise-act cycle
[Expert Systems, and MYCIN as the Comparison]the architecture and the honest history
[Feature Engineering]choosing and deriving attributes
[Multi-Class Classification, and How It Is Scored]more than two labels, and the confusion matrix
[Ayurvedic Classification as a Rule-Based Expert System]MU's topic 4, built

Quick revision

  • Nothing in this block is medical advice, and nothing claims a classical classification is clinically correct.
  • The programs are knowledge-representation exercises and are not evaluated for accuracy, because no ground truth exists to evaluate them against.
  • Three treatises: Charaka on general medicine, Suśruta on surgery, Aṣṭāṅga Hṛdaya a later synthesis. This book quotes the first two, in Kaviratna's and Bhishagratna's translations.
  • The Charaka Saṃhitā is divided into sthāna; this book quotes the Sūtrasthāna and the Vimānasthāna.
  • Four structural features make the block worth reading on a computer science paper: named attributes, a small label set, a declared order of evidence, and a rule from classification to response.
  • Formalisation is not validity, and the two questions are independent.

Test yourself

1. State what this block does and does not claim.

It claims that the Āyurvedic texts are highly formalised, with named attributes, a small set of labels, a declared order of evidence and rules connecting classification to response, and it establishes that from the texts. It does not claim that any classification is clinically correct, does not give medical advice, and does not present its programs as diagnostic software.

2. Name the three principal treatises and the sections of the Charaka Saṃhitā this book quotes.

The Charaka Saṃhitā on general medicine, the Suśruta Saṃhitā on surgery, and the Aṣṭāṅga Hṛdaya. This book quotes the Sūtrasthāna, the section of general principles, and the Vimānasthāna, on measurement and method.

3. Why do the programs in this block report no accuracy figure?

Because an accuracy figure requires a labelled dataset with a ground truth to compare against, and none exists for this material. A reported accuracy would be a number the submission had invented.

munotes.in252

Āyurveda as a Śāstra, and What This Chapter Does Not Claim

4. Which four structural features make this material relevant to a computer science paper?

A closed set of named attributes; a small set of class labels; a declared order in which evidence is gathered; and a rule connecting a classification to a response. Those four together are the shape of an expert system.

Contents This chapter on its own page

munotes.in253

Chapter Seventy-Two

The Tridoṣa Framework

Syllabus topic Module 2, "Tridoṣa framework"

In one line

Charaka names three things as the causes of bodily disease, and gives each of them a list of seven attributes.

In the wording you can write in an examination: the tridoṣa framework holds that three factors, rendered in Kaviratna's translation as wind, bile and phlegm, are the causes of bodily disease, and that each is characterised by a named list of qualities. Treatment is by objects having the adverse qualities. The framework therefore consists of three categories, a set of attributes distributed among them, and a rule relating a category to a response.

The provision

Charaka Saṃhitā, Sūtrasthāna, Lesson I, in Avinash Chandra Kaviratna's translation. The verse numbers are as he prints them.

I.56.

Wind, bile, and phlegm have been said to be the causes of all bodily diseases. The qualities of Passion and Darkness have, again, been indicated to be the causes of mental diseases.

I.58.

Wind, which may be dry, cold, light, subtile, unstable, clear, keen, is cured by objects which have adverse attributes.

I.59.

Bile, which may be cold, hot, keen, soft, sour, liquid, and bitter, is speedily cured by objects having adverse attributes.

I.60.

Heavy, cold, mild, watery, sweet, stable, and slimy, these attributes of phlegm are cured by objects having adverse attributes.

I.61.

Those changes (in wind, bile, and phlegm) that are curable may be set right by drugs possessing adverse attributes, administered according to (considerations of) place, measure, and time.

Three things are in those five verses

A classification of causes. Three for bodily disease, and two further qualities for mental disease, which the syllabus does not pursue.

Seven named attributes for each. Not a description: a list, of fixed length, drawn from a shared vocabulary.

A treatment rule. Apply the adverse attribute. And I.61 adds three parameters to the rule: place, measure and time.

Those three together are a classification scheme with a response function, and that is why this material is on a computer science paper.

The attribute lists

DoṣaAttributes, as Kaviratna prints them
winddry, cold, light, subtile, unstable, clear, keen
bilecold, hot, keen, soft, sour, liquid, bitter
phlegmheavy, cold, mild, watery, sweet, stable, slimy

The contradiction in I.59, reported

The line for bile prints both "cold" and "hot". A thing may not be characterised by both at once, and a list of qualities that contains a quality and its opposite has a fault in it.

This book does not pick a side silently. Three things can be said about it and all three are stated here rather than in a footnote.

It is in the translation as printed. The line is quoted above exactly as Kaviratna gives it.

It cannot be an intended reading. The scheme's own treatment rule is the application of the adverse attribute, and an attribute whose opposite is in the same list makes that rule unusable for it: there is no adverse attribute to apply.

munotes.in254

The Tridoṣa Framework

And the other two lists also contain "cold". Wind is cold and phlegm is cold. So if bile is cold too, the attribute is shared by all three and distinguishes nothing at all, which is a second and independent reason to doubt the reading.

What this book does about it. The programs in this block record the disputed attribute separately and exclude it from bile's list, and they say so in their own source. The reading is declared, not assumed, and [Doṣa as a Feature Vector] shows what difference it makes to the arithmetic.

And the general rule, which applies to the whole book. The primary texts reach us through a translator. Where the translator's text is inconsistent, obscure or conjectural, the chapter says so rather than presenting one reading as the text. That is the same discipline the B.Com. Indian Knowledge Systems book applies to a penalty multiple that its own source gives two ways.

What the attributes are, as a vocabulary

Look at the three lists together and something becomes visible that no one list shows.

The attributes come in opposite pairs. Dry and watery. Light and heavy. Keen and mild. Unstable and stable. Clear and slimy. Cold and hot. Sour and sweet.

That is what makes the treatment rule computable. "Apply the adverse attribute" is only an instruction if every attribute has a named opposite. The vocabulary is built in pairs precisely so that it does.

And the tradition names the vocabulary. The qualities are called guṇa, the same word Vaiśeṣika uses for its second category in [Padārtha: The Categories of What Exists], and the classical list of them is twenty, in ten opposed pairs. This paper needs only the ones the three lists above use.

Worked example: reading the scheme as a function

Input. A description, as a set of attributes.

Step one. Compare it with each of the three lists.

Step two. The category sharing most attributes with the description is the classification.

Step three. The response is the set of opposites of the matched attributes, applied with regard to place, measure and time.

Worked. A description of dry, unstable and subtile matches wind on all three and the others on none. The response is objects that are moist, stable and gross.

That is a function from a set of attributes to a category and a response, and it is entirely mechanical. [Ayurvedic Classification as a Rule-Based Expert System] implements it, and the caution of [Āyurveda as a Śāstra, and What This Chapter Does Not Claim] applies to every word of it.

munotes.in255

The Tridoṣa Framework

What the framework is NOT

It is not three substances. The English words wind, bile and phlegm are Kaviratna's renderings, and the tradition's own usage is of three functional principles rather than three fluids. This book uses his words because it quotes his translation, and says here that they are a translator's choice.

It is not a diagnosis. Classifying a description into one of three categories is not identifying a disease. The texts have an extensive separate apparatus for that, which this paper does not enter.

It is not exhaustive. I.56 names two further causes, the qualities of passion and darkness, for mental disease. The syllabus's label is the tridoṣa framework and this book stays inside it.

And it is not validated. The caution chapter says this and it bears repeating at the point where the attributes become a table: the structure is what is being studied.

Quick revision

  • Charaka Sūtrasthāna I.56: wind, bile and phlegm are the causes of all bodily diseases.
  • I.58, I.59 and I.60 each give SEVEN named attributes, and I.61 gives the treatment rule: objects of adverse attributes, with regard to place, measure and time.
  • Kaviratna's line for bile prints both "cold" and "hot", which cannot both be intended, and "cold" is in all three lists so it would distinguish nothing. The book states the problem and declares its reading.
  • The attributes come in opposed pairs, which is what makes "apply the adverse attribute" a computable instruction.
  • The scheme is a function: a set of attributes goes in, a category and a response come out.
  • Wind, bile and phlegm are a translator's words for three functional principles, not three fluids.

Test yourself

1. Quote or state Charaka's three causes and give the attributes of one of them.

Wind, bile and phlegm are the causes of all bodily diseases, at Sūtrasthāna I.56. Wind, at I.58, is dry, cold, light, subtile, unstable, clear and keen.

2. What is wrong with the printed list for bile, and what is the second reason to doubt it?

It contains both "cold" and "hot", which cannot both characterise the same thing, and an attribute whose opposite is in the same list leaves the treatment rule with nothing to apply. The second reason is that "cold" already appears in the lists for wind and phlegm, so if bile is also cold the attribute distinguishes none of the three.

3. Why does the treatment rule require the attributes to come in opposed pairs?

Because "apply the object of the adverse attribute" is an instruction only if every attribute has a named opposite. Without the pairing there would be nothing to apply.

munotes.in256

The Tridoṣa Framework

4. State the scheme as a function, naming its input and its output.

Its input is a set of attributes describing a case. Its output is the category whose attribute list the description best matches, together with the set of opposites of the matched attributes, to be applied with regard to place, measure and time.

Contents This chapter on its own page

munotes.in257

Chapter Seventy-Three

Doṣa as a Feature Vector

Syllabus topic Module 2, "Tridoṣa framework", "Feature engineering"

In one line

Write the three attribute lists as rows over one shared vocabulary and each doṣa becomes a vector of ones and zeros.

In the wording you can write in an examination: a feature vector is a fixed-length sequence of values, one per attribute of a shared attribute space, representing one object. Turning the three attribute lists of Charaka Sūtrasthāna I.58 to I.60 into vectors over the union of the attributes they name produces three vectors of eighteen binary components, and every subsequent operation, comparison, classification and tree induction, is defined on those vectors.

Why the transformation matters

A list is not a vector, and the difference is not notation.

A list says what is present. "Wind is dry, cold, light, subtile, unstable, clear, keen."

A vector says what is present AND what is absent. It has a slot for every attribute in the vocabulary, so it states, of every attribute, whether this doṣa has it.

That is what makes comparison possible. Two lists cannot be subtracted. Two vectors can, and the difference is the set of attributes on which they disagree.

And it is what makes counting possible. How many attributes distinguish wind from phlegm is a question about vectors and not about lists.

Building the shared space

The vocabulary is the union of the three lists, in the order they first appear, and the program prints it.

ATTRIBUTES = {
    "wind":   ["dry", "cold", "light", "subtile", "unstable", "clear", "keen"],
    "bile":   ["cold", "hot", "keen", "soft", "sour", "liquid", "bitter"],
    "phlegm": ["heavy", "cold", "mild", "watery", "sweet", "stable", "slimy"],
}
# The reading this book declines to rely on: Kaviratna's line for bile prints
# both "cold" and "hot". See the chapter on the tridosa framework.
DISPUTED = {("bile", "cold")}


def universe():
    """Every attribute any dosa is given, in first-appearance order."""
    seen = []
    for dosa in ("wind", "bile", "phlegm"):
        for a in ATTRIBUTES[dosa]:
            if a not in seen:
                seen.append(a)
    return seen


def vector(dosa, drop_disputed=True):
    """The dosa as a 0/1 vector over the shared attribute space."""
    have = set(ATTRIBUTES[dosa])
    if drop_disputed:
        have -= {a for d, a in DISPUTED if d == dosa}
    return [1 if a in have else 0 for a in universe()]


def discriminating(drop_disputed=True):
    """Attributes that do NOT appear against every dosa."""
    out = []
    for i, a in enumerate(universe()):
        col = [vector(d, drop_disputed)[i] for d in ("wind", "bile", "phlegm")]
        if 0 < sum(col) < 3:
            out.append(a)
    return out


U = universe()
print("the shared attribute space has %d attributes" % len(U))
print("  " + ", ".join(U))
print()
print("%-8s %s" % ("", " ".join("%-9s" % a for a in U)))
for d in ("wind", "bile", "phlegm"):
    print("%-8s %s" % (d, " ".join("%-9d" % v for v in vector(d))))
print()
for d in ("wind", "bile", "phlegm"):
    v = vector(d)
    print("  %-8s carries %d of the %d attributes" % (d, sum(v), len(U)))
print()
print("with the disputed reading KEPT, attributes shared by all three:")
print("  " + (", ".join(a for a in U if a not in discriminating(False)) or "(none)"))
print("with the disputed reading DROPPED, attributes shared by all three:")
print("  " + (", ".join(a for a in U if a not in discriminating(True)) or "(none)"))
munotes.in258

Doṣa as a Feature Vector

the shared attribute space has 18 attributes
  dry, cold, light, subtile, unstable, clear, keen, hot, soft, sour, liquid, bitter, heavy, mild, watery, sweet, stable, slimy

         dry       cold      light     subtile   unstable  clear     keen      hot       soft      sour      liquid    bitter    heavy     mild      watery    sweet     stable    slimy
wind     1         1         1         1         1         1         1         0         0         0         0         0         0         0         0         0         0         0
bile     0         0         0         0         0         0         1         1         1         1         1         1         0         0         0         0         0         0
phlegm   0         1         0         0         0         0         0         0         0         0         0         0         1         1         1         1         1         1

  wind     carries 7 of the 18 attributes
  bile     carries 6 of the 18 attributes
  phlegm   carries 7 of the 18 attributes

with the disputed reading KEPT, attributes shared by all three:
  cold
with the disputed reading DROPPED, attributes shared by all three:
  (none)

Reading the table

Eighteen attributes, not twenty-one. Three lists of seven would be twenty-one if no attribute recurred. Three recur, so the union is eighteen.

Wind and phlegm carry seven each and bile carries six. Bile's list has seven entries too, and one of them is the disputed "cold", which this book's programs exclude. So the row has six ones.

And the last two lines are the finding. With the disputed reading kept, "cold" is the one attribute shared by all three doṣas, and an attribute that every category has cannot distinguish them. With it dropped, no attribute is shared by all three, and every one of the eighteen carries some information.

That is a real consequence of a textual problem, and it is why [The Tridoṣa Framework] refuses to smooth the contradiction over. A reading that makes an attribute useless is evidence about the reading.

What a vector buys, in four operations

Each of these is used later in the block and each is trivial once the vectors exist.

Overlap. How many attributes two doṣas share is the count of positions where both have a one. Wind and bile share one, keen. Wind and phlegm share one, cold. Bile and phlegm share none.

Distance. How many positions they differ in. That is the Hamming distance, and it is the simplest measure of dissimilarity there is.

Matching a description. A description is itself a vector, with ones for the attributes reported. Classification becomes comparing that vector with the three.

munotes.in259

Doṣa as a Feature Vector

Feature selection. Dropping an attribute is dropping a column, and whether it is worth keeping is a question about the column. [Feature Engineering] is about that question and the last two lines of the output above are an instance of it.

Worked example: the four measures, by hand

Take wind and phlegm.

Shared. Both have cold, and nothing else. Overlap one.

Wind only. Dry, light, subtile, unstable, clear, keen. Six.

Phlegm only. Heavy, mild, watery, sweet, stable, slimy. Six.

Neither. Hot, soft, sour, liquid, bitter. Five.

Check the arithmetic. One plus six plus six plus five is eighteen, which is the size of the space. Any four-way split of the attribute space must add to its size, and that is a check worth doing every time, because it catches a miscounted column at once.

Hamming distance is the second plus the third: twelve. Wind and phlegm disagree on twelve of eighteen attributes, which makes them the most distant pair.

The sparsity, and what it means

Each doṣa has at most seven of eighteen attributes, so each vector is mostly zeros. That is called a sparse representation and it has two consequences worth knowing.

Most comparisons are between zeros. Two doṣas agreeing that neither is bitter is an agreement, and it is not informative, because almost everything agrees about almost everything. A similarity measure that counts shared zeros would make all three doṣas look alike.

So the right measures count shared ONES. The overlap above counts positions where both are one and ignores positions where both are zero, which is the standard treatment for sparse binary data.

That is a design decision and a question can ask for it. Given three vectors of eighteen components with at most seven ones each, a measure that counts agreements including zeros would report wind and bile as agreeing on eleven of eighteen, which sounds high and means nothing.

Limits of the representation

It is binary. An attribute is present or absent, with no degree. The texts themselves speak of qualities being increased or diminished, which a binary vector cannot express.

It is unordered. Nothing in the vector records that dry and watery are opposites. The pairing that makes the treatment rule computable is outside the representation, held separately, and [Symptom to Feature Mapping] treats that as a design problem rather than an oversight.

It has three rows. Three examples over eighteen attributes is a very small dataset, and [A Decision Tree Built From the Tridoṣa Attributes] shows exactly what that costs when a method needs data.

Quick revision

  • A list says what is present; a vector says of every attribute whether it is present, which is what makes comparison and counting possible.
  • The union of the three lists is eighteen attributes, because three recur.
  • With the disputed reading kept, "cold" belongs to all three and discriminates nothing; dropped, every attribute carries information.
  • Four operations: overlap, Hamming distance, matching a description, and dropping a column.
  • A four-way split of the space must add to the size of the space, which is a free check.
  • The vectors are sparse, so a similarity measure must count shared ones and ignore shared zeros.
  • Limits: binary, so no degrees; unordered, so the opposite pairs live outside it; and only three rows.
munotes.in260

Doṣa as a Feature Vector

Test yourself

1. Why is the shared attribute space eighteen rather than twenty-one?

Because three attributes appear in more than one list. Three lists of seven would give twenty-one only if no attribute recurred, and the union removes the repetitions.

2. What happens to the attribute "cold" under the two readings, and why is that evidence?

Keeping the disputed reading puts cold in all three lists, so it distinguishes none of them. Dropping it leaves cold to wind and phlegm only, where it carries information. A reading that makes an attribute useless is evidence against that reading.

3. Compute the four-way split of wind against phlegm and check it.

Shared: cold, one. Wind only: six. Phlegm only: six. Neither: five. One plus six plus six plus five is eighteen, the size of the space, so the split is complete.

4. Why must a similarity measure on these vectors ignore shared zeros?

Because the vectors are sparse: each has at most seven ones out of eighteen, so any two agree on most of the zeros. Counting those agreements would make every pair look similar and would carry no information.

Contents This chapter on its own page

munotes.in261

Chapter Seventy-Four

Prakṛti: Constitution as a Class Label

Syllabus topic Module 2, "Tridoṣa framework", "Multi-class classification"

In one line

A case is classified into one of three categories, or a combination of them, so the label set has more than two members and more than three.

In the wording you can write in an examination: prakṛti is the classification of a constitution by which of the three doṣa predominate in it. Since one, two or all three may predominate, the label set is not binary and not simply three-valued: it has seven non-empty possibilities, and a classifier over it is a multi-class and, on one reading, a multi-label problem.

The arithmetic of the label set

Take the three categories of [The Tridoṣa Framework] and ask which subsets can predominate.

KindHow manyWhich
one predominates3wind; bile; phlegm
two predominate3wind and bile; wind and phlegm; bile and phlegm
all three in balance1all three
none1, and it is excludedthe empty set

Seven non-empty possibilities, which is 2 to the power 3 minus 1. The Meru-prastāra of [Meru-Prastāra: Halāyudha's Staircase] gives the same count: row three is 1, 3, 3, 1, and dropping the leftmost cell, which is the empty selection, leaves 3 plus 3 plus 1, which is seven.

That is not a coincidence and it is worth saying. Choosing which of three things predominate is the same problem as choosing a subset of three, and the combinatorics of Module I answers it.

Why this is not a binary problem

A great deal of introductory material on classification assumes two classes: spam or not, fraud or not, positive or negative. Almost none of the vocabulary of that setting transfers unchanged.

Accuracy means something different. With two classes and an even split, guessing gives fifty per cent. With seven, guessing gives about fourteen. The same accuracy figure is a different achievement.

There is no single "positive" class. Precision and recall are defined with respect to a positive class, so with seven labels there are seven precisions and seven recalls, and how to combine them is a decision. [Multi-Class Classification, and How It Is Scored] is about that decision.

Errors are not all equal. Confusing wind with phlegm and confusing wind with wind-and-phlegm are different mistakes, and a single accuracy number hides the difference. The confusion matrix exists because of this.

Multi-class against multi-label

The distinction is worth having exactly, because the prakṛti scheme can be read either way and a question may turn on it.

Multi-class. Each case gets exactly ONE label, drawn from a set of more than two. The seven possibilities are the seven labels, and "wind and bile" is one label like any other.

Multi-label. Each case gets a SET of labels, drawn from three. A case is labelled with wind, or with wind and bile, and the classifier makes three independent decisions.

munotes.in262

Prakṛti: Constitution as a Class Label

Multi-class, seven labelsMulti-label, three labels
What the classifier outputsone of sevenany subset of three
Number of decisionsonethree, one per doṣa
Training data neededexamples of all sevenexamples of each doṣa predominating
Can it output something unseennoyes, a combination never seen in training
Errorsone wrong labelpossibly one of three wrong

The multi-label reading is the better model of the scheme, because the classical framework does not treat "wind and bile" as a separate thing from wind and from bile. It treats it as both predominating. And the multi-label reading needs less training data, because it learns three things rather than seven.

The multi-class reading is easier to score, because there is one answer and it is right or wrong.

Saying which reading you have chosen, and why, is the mark of an answer that understands the problem rather than one that has read a definition.

Worked example: what changes with the reading

A case. The description reports dry, unstable, sour and bitter.

Under the multi-class reading. The classifier must choose one of seven. Dry and unstable belong to wind; sour and bitter belong to bile. It might output "wind and bile", or it might output "wind" if its training saw more wind cases, and there is no way to express that it is two parts wind and two parts bile.

Under the multi-label reading. Three independent questions. Does wind predominate: two attributes present, yes. Does bile: two present, yes. Does phlegm: none present, no. Output: wind and bile.

And the multi-label reading gives a number for each, which the multi-class one does not. That is the practical difference and it is why it is preferred.

The class imbalance question

The seven labels will not occur equally often in any real collection of cases, and that has two consequences that a question can ask for.

A classifier that always answers with the commonest label can have high accuracy. If four cases in ten are wind, always answering wind scores forty per cent, which beats guessing among seven.

So accuracy alone is not a report. The confusion matrix, and the per-class figures, are what say whether the rare labels are being found at all.

This matters here in a specific way. The all-three-in-balance label is, by the scheme's own account, uncommon. A classifier evaluated only on accuracy would be rewarded for never predicting it.

What this chapter does NOT claim

Not that any of the seven describes a real population. The arithmetic says seven subsets are possible. Whether the scheme's categories correspond to anything is the clinical question [Āyurveda as a Śāstra, and What This Chapter Does Not Claim] excludes.

munotes.in263

Prakṛti: Constitution as a Class Label

Not that a classifier over them would be useful. It would be a classifier over a historical scheme's own labels, which is what MU's topic 4 asks for and nothing more.

Not that prakṛti and doṣa are the same thing. The scheme distinguishes a constitution, which is treated as settled, from a current state, which changes. This paper's programs work on attribute descriptions and do not model that distinction, and saying so belongs in a submission's limitations section.

Quick revision

  • Prakṛti classifies by which doṣa predominate: one, two, or all three.
  • Seven non-empty possibilities, which is 2 to the power 3 minus 1, and the Meru's row three gives the same count.
  • Not binary: accuracy means something different, there is no single positive class, and errors are not all equal.
  • Multi-class: one label of seven. Multi-label: a subset of three, three independent decisions, less training data, and it can output a combination never seen.
  • The multi-label reading models the scheme better; the multi-class reading is easier to score.
  • Class imbalance: always answering the commonest label can score well, so accuracy alone is not a report.

Test yourself

1. How many labels does the prakṛti scheme admit, and how do you get the number?

Seven. Each of the three doṣa either predominates or does not, giving 2 to the power 3 subsets, and the empty subset is excluded, leaving seven.

2. Distinguish multi-class from multi-label classification and say which models this scheme better.

Multi-class assigns exactly one label from a set of more than two; multi-label assigns a subset of labels by making one decision per label. Multi-label models the scheme better, because "wind and bile" is not a separate thing but both doṣa predominating, and it needs examples of three cases rather than of seven.

3. Why is accuracy a poor single measure here?

Because the labels are unequally common and the errors are not equivalent. A classifier that always answers the commonest label can score well while never predicting a rare one, and confusing wind with phlegm is a different error from confusing wind with wind-and-phlegm.

4. What connects this chapter to Module I?

The count. Choosing which of three doṣa predominate is choosing a subset of three, which is the combinatorial problem the Meru-prastāra answers: row three is 1, 3, 3, 1, and dropping the empty selection leaves seven.

Contents This chapter on its own page

munotes.in264

Chapter Seventy-Five

Parīkṣā: How an Examination Is Structured

Syllabus topic Module 2, "Examination of structured diagnostic reasoning in Ayurvedic texts", "Decision principles in diagnosis"

In one line

Charaka says which three kinds of evidence a physician may use, in which order, and that no one of them is enough.

In the wording you can write in an examination: parīkṣā is examination, the structured gathering of evidence before a judgement. Charaka admits three means of knowledge: the instruction of the trustworthy, observation, and inference. He directs that all three be used, that the instruction come first in order of use, and that a single means is insufficient for everything that must be known.

The provision

Charaka Saṃhitā, Vimānasthāna, Lesson IV, in Kaviratna's translation.

On the second and third means:

That, verily, is (the result of) observation which one acquires by one's own senses and mind.

Inference, verily, is argument depending upon reasons.

On using them together:

Verily, with the aid of all these three means of knowledge, one should in the first instance fully examine a disease. The diagnosis that is then arrived at becomes faultless.

Truly, by only one of these means of knowledge, knowledge does not arise of everything that should be known.

On the order:

Among all these three means of knowledge, the knowledge derived from the instructions of the inspired comes first. After this, comes examination, with the aid of Observation and Inference.

What would one, that has not been instructed (by the inspired) in the first instance, succeed in knowing by examining with the aid of Observation and Inference?

And on the first means, from the translator's own note:

that person is called an "Āpta" whose knowledge of things is not derived by a posteriori methods. They do not depend upon reasoning, nor upon memory ... They behold all things without "prīti", i.e., pleasure or satisfaction, and without "upatāpa", i.e., pain ... They are unimpassioned judges of truth.

The three, and their order

OrderMeansWhat it suppliesIts characteristic risk
firstthe instruction of the trustworthywhat to look for, and what the possibilities arethe source may be wrong, and there is no test of it here
secondobservationparticular facts about this case, by the senses and the mindthe senses misreport, and the observer sees what they expect
thirdinferencewhat cannot be observed, from what wasthe general connection may not hold

Notice that "the mind" is inside observation. Kaviratna's note on the same passage records that the mind is regarded as a sixth sense, so what the mind acquires is within pratyakṣa. That places pattern recognition on the perception side rather than the inference side, which is a boundary [Pratyakṣa: Perception, and Why It Is Defined So Narrowly] discusses.

Why "instruction first" is a procedural claim, not a deference

The argument Charaka gives is a question: what would one who has not been instructed succeed in knowing by examining with observation and inference?

munotes.in265

Parīkṣā: How an Examination Is Structured

The answer is: very little, because they would not know what to examine. Observation is not a passive intake. A physician who does not know the possible conditions does not know which signs are signs.

That is a claim about search. The prior knowledge is what makes the search space small enough to work in. In modern terms: you need a hypothesis space before you can gather evidence, because evidence is only evidence relative to a hypothesis.

And it is not a claim about authority. Nothing in the passage says the instruction outranks what you see. It says it comes first in sequence.

The rule that no single means suffices

This sentence is the most modern thing in the Āyurvedic material and it is worth quoting in any answer on this topic.

By only one of these means of knowledge, knowledge does not arise of everything that should be known.

Three modern readings, all of which the sentence supports.

Sensor fusion. A system with one input has one failure mode and no way to detect it. Two disagreeing inputs are informative; one input is not.

Triangulation. A claim supported by measurement, by derivation and by documented prior knowledge is stronger than a claim supported three times by measurement.

Coverage. Some things cannot be observed at all, and only inference reaches them; other things cannot be inferred and must be seen. Each means has a range, and the ranges do not coincide. That is the plainest reading and it is probably the intended one.

Parīkṣā as a procedure

Set out as steps, which is how a computer science answer should present it.

  1. Establish the hypothesis space. From the instruction of the trustworthy: what conditions are possible, and what distinguishes them.
  2. Gather particulars. By observation, using the senses and the mind, guided by step one.
  3. Reach what cannot be observed. By inference, from the particulars, using general connections.
  4. Combine. Do not conclude from any one of the three alone.
  5. And repeat. The texts return to examination after treatment begins, so the procedure is not one-shot.

Compare a modern diagnostic procedure and the shape is the same: know the fault catalogue, collect telemetry, infer the unobservable state, corroborate across sources, and re-examine after acting.

Worked example

An engineer is asked why a nightly job is slow.

Step one, instruction. The runbook lists the four known causes of slowness in this job: a large input, a missing index, contention with the backup window, and a stale statistics table. The engineer now knows what to look at. Without it, the telemetry is a wall of numbers.

munotes.in266

Parīkṣā: How an Examination Is Structured

Step two, observation. The input size is normal. The backup window does not overlap. Both by direct measurement.

Step three, inference. The query plan shows a full scan where an index scan is expected. The index or the statistics is the cause; neither is directly visible in the job's own output.

Step four, combine. The observation rules out two causes, the inference narrows to two, and the runbook says which check distinguishes them. No single means reached the answer, which is Charaka's sentence.

Step five, re-examine. After rebuilding the statistics, run the job again and observe.

What parīkṣā is NOT

It is not a list of tests. The texts do have such lists, at length. This chapter is about the structure of the procedure, which is what MU's label "examination of structured diagnostic reasoning" asks for.

It is not an algorithm. No step says exactly what to do next; it says which kind of evidence to use. A procedure that leaves the choice of next observation to judgement is not an algorithm, by the definiteness test of [The Prastāra Rule Read as an Algorithm].

It is not a validation of the conclusions reached by it. A faultless procedure can reach a wrong conclusion if the instruction was wrong, and the scheme has no mechanism for correcting the instruction.

Quick revision

  • Three means: the instruction of the trustworthy, observation by the senses and the mind, and inference.
  • Charaka: use all three and the diagnosis becomes faultless; by only one, knowledge does not arise of everything that should be known.
  • Instruction comes FIRST in order of use, not in authority, because without it one does not know what to examine. That is a claim about the hypothesis space.
  • Kaviratna's note: the mind is a sixth sense, so what the mind acquires is within observation.
  • The procedure: establish the possibilities, gather particulars, infer the unobservable, combine, re-examine.
  • Not a list of tests, not an algorithm by the definiteness test, and not a validation of what it concludes.

Test yourself

1. Name Charaka's three means and the order he gives them.

The instruction of the trustworthy, observation, and inference. Instruction comes first in order of use, followed by examination with observation and inference.

2. Quote or paraphrase his reason for putting instruction first, and say what kind of claim it is.

He asks what one who has not been instructed would succeed in knowing by examining with observation and inference. It is a claim about the hypothesis space: without prior knowledge of the possibilities, one does not know which signs are signs.

3. Give the sentence about using one means alone, and two modern readings of it.

By only one of these means of knowledge, knowledge does not arise of everything that should be known. It reads as sensor fusion, since one input has one failure mode and no way to detect it; and as coverage, since each means reaches things the others cannot.

munotes.in267

Parīkṣā: How an Examination Is Structured

4. Why is parīkṣā not an algorithm?

Because no step determines the next action. It says which kind of evidence to use and leaves the choice of what to observe to judgement, which fails the definiteness requirement that every step have exactly one result.

Contents This chapter on its own page

munotes.in268

Chapter Seventy-Six

Symptom to Feature Mapping

Syllabus topic Module 2, "Symptom–feature mapping", "Feature engineering"

In one line

Turning "the patient says their skin is dry" into an attribute with a value is where most of the work of a classifier actually is.

In the wording you can write in an examination: symptom to feature mapping is the conversion of a reported or observed complaint into a named attribute of the representation, with a value. It involves resolving the reporter's vocabulary to the system's, choosing a scale for the value, and deciding what to record when nothing was reported. Each of the three is a design decision and each can be got wrong in a way that no later stage can repair.

Why this is the hard part

The classifier is the easy part. Given three vectors and a description as a vector, comparing them is arithmetic.

Producing the description as a vector is not. A person says "my skin has been dry lately and I feel restless". The representation has eighteen named attributes. Which of them are now one?

And nothing downstream can fix a bad mapping. A classifier is only as good as its features, and a feature that does not mean what the system thinks it means poisons every result built on it.

The three problems

One: the vocabulary problem

The reporter's words are not the system's attributes. "Restless" is not in the list of eighteen. "Unstable" is. Are they the same?

The classical text's handling. The vocabulary is fixed by the tradition and the practitioner is trained in it. The mapping from what a patient says to a named quality is carried in the practitioner's training and is not written down as a table.

The modern handling. Either a controlled vocabulary with a mapping table, which is explicit and enormous; or a learned mapping from text to features, which is compact and cannot be inspected.

Both have the same failure. A term that maps to the wrong attribute is a silent error. The tradition's version is that two practitioners trained differently record the same report differently; the modern version is that a mapping table nobody has audited contains a wrong row.

Two: the scale problem

Is an attribute present or absent, or present to a degree?

The classical text's handling. The texts speak of qualities being increased and diminished, so degrees are in the scheme. But the attribute lists of Charaka Sūtrasthāna I.58 to I.60 are lists of qualities, not of measurements, and the vectors of [Doṣa as a Feature Vector] are binary because the lists are.

So this book's representation loses something the text has, and saying so is the honest position: the binary vector is a simplification declared, not a faithful rendering.

The modern handling. Four kinds of scale, and choosing among them is a design decision.

munotes.in269

Symptom to Feature Mapping

ScaleWhat it supportsExample
binarypresent or absentdry: yes or no
ordinalan order, with no fixed distancesdryness: none, mild, marked, severe
intervaldifferences are meaningfultemperature in degrees
ratioratios are meaningful, with a true zeroduration in days

The trap is treating an ordinal scale as a number. Coding none, mild, marked, severe as 0, 1, 2, 3 and then averaging asserts that the step from mild to marked is the same size as the step from marked to severe, which nobody has established. A classifier will happily compute that average and the result means nothing.

Three: the missing-value problem

Nothing was reported about an attribute. What goes in the slot?

Three possible answers and they are different, which is the three-state distinction of [Padārtha Ontology as a Knowledge Representation Model] arriving again.

Zero, meaning absent. Wrong unless the attribute was actually checked and found absent.

A distinguished missing marker. Correct, and it forces every downstream step to decide what to do about it.

An imputed value. Filling in a guess, usually the commonest value or an average. Convenient and dangerous: the guess is then indistinguishable from an observation.

The classical scheme's handling is the interesting one. It has a category for known absence, as [Padārtha: The Categories of What Exists] records, and the rule in [Parīkṣā: How an Examination Is Structured] that instruction comes first means the practitioner knows what to look for, so absence is more often checked than assumed. A procedure that tells you what to examine reduces the number of missing values, which is a genuine design point.

The opposite pairs, and where to hold them

[Doṣa as a Feature Vector] left this open: the vector does not record that dry and watery are opposites, and the treatment rule needs them.

Three places the pairing could live.

In the vector's structure. One component per PAIR, taking three values: the attribute, its opposite, or neither. This makes the pairing implicit in the representation and makes the treatment rule trivial. It also loses the ability to record that neither was checked, unless a fourth value is added.

In a separate table. Eighteen attributes, nine pairs. Explicit, inspectable, and it must be kept in step with the attribute list.

In the rules. Each rule states its own opposite. Distributed, repetitive, and it drifts.

The programs in this book use the separate table, and say so. The reason is the one in [Formal Specification: Saying Exactly What a System Must Do]: a fact stated in one place can be checked; a fact stated in many places disagrees with itself eventually.

munotes.in270

Symptom to Feature Mapping

Worked example: one report, mapped

The report. "My skin has been dry for a fortnight, I have not slept well, and I feel cold."

Vocabulary. "Dry" maps to the attribute dry. "Cold" maps to the attribute cold. "Not slept well" maps to no attribute in the list of eighteen: it is a symptom the representation cannot express, and recording it as "unstable" would be an invention.

Scale. "For a fortnight" is duration, which is a ratio quantity and has no slot. It is discarded, and the loss should be recorded.

Missing values. Sixteen of the eighteen attributes were not reported on. They are missing, not absent. If the practitioner then examines and finds no slimy quality, that attribute becomes a known absence, which is a third state.

The vector. Two ones, two known absences if they were checked, and fourteen missing.

And the classification. Dry is wind only. Cold is wind and phlegm, on this book's reading. So the evidence points to wind, weakly, on two attributes out of eighteen. A classifier that returned a confident answer from that would be lying, and reporting the weakness is part of the output.

Quick revision

  • Symptom to feature mapping is the conversion of a report into named attributes with values, and it is where most errors enter.
  • The vocabulary problem: the reporter's words are not the system's attributes, and a wrong mapping is a silent error.
  • The scale problem: binary, ordinal, interval, ratio. Treating an ordinal code as a number asserts equal steps nobody established.
  • The missing-value problem: absent, missing and imputed are three different things, and imputation makes a guess indistinguishable from an observation.
  • The classical scheme has a category for known absence, and a procedure that says what to examine reduces missing values.
  • The opposite pairs live in a separate table, because a fact in one place can be checked and a fact in many places drifts.

Test yourself

1. Name the three problems in mapping a symptom to a feature.

Resolving the reporter's vocabulary to the system's attributes; choosing the scale on which the value is recorded; and deciding what to record when nothing was reported.

2. Why is coding an ordinal scale as 0, 1, 2, 3 and averaging it a mistake?

Because it asserts that the steps between successive levels are equal, which an ordinal scale does not establish. The arithmetic will be computed and the result will not mean anything.

3. Distinguish absent, missing and imputed, and say which is dangerous.

Absent means the attribute was checked and found not present. Missing means nothing was recorded. Imputed means a guess was written into the slot. Imputation is the dangerous one, because the guess is then indistinguishable from an observation.

munotes.in271

Symptom to Feature Mapping

4. Where should the opposite pairs be held, and why?

In one separate table, rather than implicitly in the vector's structure or repeated in the rules. A fact stated in one place can be checked and kept correct; a fact stated in several places eventually disagrees with itself.

Contents This chapter on its own page

munotes.in272

Chapter Seventy-Seven

Multi-Attribute Classification

Syllabus topic Module 2, "Multi-attribute classification"

In one line

Several attributes at once, and the question is how their evidence is combined into one answer.

In the wording you can write in an examination: multi-attribute classification assigns a class to an object described by several attributes. The attributes must be combined by a stated rule: conjunctive, requiring all of them; disjunctive, requiring any of them; or weighted, summing a contribution from each and comparing the total against thresholds. The choice of rule is a modelling decision and different rules classify the same object differently.

The problem

One attribute is easy. "Dry" belongs to wind and to nothing else, so a description of dry alone points to wind.

Several attributes need a rule. A description of dry, sour and heavy has one attribute of each category. What is the answer?

There is no answer without a combination rule, and a system that produces one anyway has a rule hidden inside it.

The three rules

Conjunctive: all of them

The rule. Assign the class only if the object has every attribute the class requires.

On this table. A description would have to carry all seven of wind's attributes to be classified as wind.

Property. Very precise and almost never satisfied. With seven required attributes and a description carrying three, nothing is ever classified.

Where it is right. Where a false positive is very costly and a missing answer is acceptable. A safety interlock that arms only when every condition holds.

Disjunctive: any of them

The rule. Assign the class if the object has any of the class's attributes.

On this table. A description carrying "keen" alone belongs to wind and to bile, since both lists contain it.

Property. Very permissive, and it produces multiple answers constantly. With eighteen attributes spread over three classes, almost any description matches two.

Where it is right. Where a false negative is very costly and further checking is available. A screening test that flags anything worth a second look.

Weighted: sum the contributions

The rule. Give each attribute a weight per class, add the weights of the attributes present, and take the class with the highest total, or the classes above a threshold.

On this table, with every weight equal to one, the total is just the count of matched attributes, and that is what [Ayurvedic Classification as a Rule-Based Expert System] uses.

Property. It always produces a ranking, and it can tie.

Where it is right. Almost everywhere in practice, which is why it is the default. And it is where the modelling decisions hide.

Worked example: three rules, one description

The description. Dry, sour, heavy.

Which classes contain each, from the table of [Doṣa as a Feature Vector].

munotes.in273

Multi-Attribute Classification

Attributewindbilephlegm
dryyesnono
sournoyesno
heavynonoyes

Conjunctive. Wind requires seven attributes and one is present. Bile requires six and one is present. Phlegm requires seven and one is present. No class is assigned.

Disjunctive. Wind has dry, so wind. Bile has sour, so bile. Phlegm has heavy, so phlegm. All three classes are assigned, which is no answer at all.

Weighted, all weights one. Wind scores 1, bile 1, phlegm 1. A three-way tie, and the honest output is that the description does not discriminate.

Three rules, three different behaviours, and the third one is the only one that says something true: this description carries one attribute of each class and therefore decides nothing.

Worked example: a description that does decide

The description. Dry, unstable, subtile, keen.

Attributewindbilephlegm
dryyesnono
unstableyesnono
subtileyesnono
keenyesyesno

Weighted, all weights one. Wind 4, bile 1, phlegm 0.

And the margin is what matters. Wind leads by three. Compare a case scoring wind 2 and bile 1: the same winner, a much weaker result, and a system that reports only the winner has thrown the difference away.

So a weighted rule should report the scores, not just the maximum. That is the practical conclusion of this chapter and it is what the program in [Ayurvedic Classification as a Rule-Based Expert System] does.

About weights

Nothing in Charaka assigns weights to attributes. The lists are lists; no quality is said to count for more than another in classifying.

So the weights used in this book are all one, which is a declared choice and not a finding. Any other weighting would have to come from somewhere, and inventing weights and presenting them as the text's would be a fabrication.

How weights are obtained in modern practice, for completeness, since a question may ask.

From an expert, by asking. Cheap, and it records what the expert believes rather than what is true.

From data, by fitting. Requires labelled examples, which for this scheme do not exist.

From information content. An attribute unique to one class is more informative than one shared by two, and the information gain of [Decision Trees] measures exactly that. This is the only one of the three available here, and [A Decision Tree Built From the Tridoṣa Attributes] shows what it yields on three rows, which is less than you would hope.

What multi-attribute classification is NOT

It is not the same as having many attributes. It is about the rule that combines them. A system with a hundred attributes and no stated combination rule has an unstated one.

munotes.in274

Multi-Attribute Classification

It is not solved by a threshold. Choosing a threshold for a weighted sum is another modelling decision, and moving it trades false positives against false negatives. [Multi-Class Classification, and How It Is Scored] is where that trade is measured.

It does not require the attributes to be independent. They usually are not, and the weighted rule quietly assumes they are. Two attributes that always occur together contribute twice and should contribute once, which is a real defect of the simple weighted rule.

Quick revision

  • With more than one attribute, a combination rule is required, and a system without a stated one has an unstated one.
  • Conjunctive, all: precise and almost never satisfied. Disjunctive, any: permissive and produces multiple answers. Weighted: always ranks, and can tie.
  • On dry, sour, heavy the three rules give no class, all three classes, and a three-way tie. The tie is the only true answer.
  • A weighted rule should report the scores, because the margin matters and the maximum alone discards it.
  • No weights are in Charaka, so this book uses one for every attribute and says so.
  • The weighted rule assumes the attributes are independent, and two attributes that always co-occur are counted twice.

Test yourself

1. Name the three combination rules and give a setting where each is right.

Conjunctive, requiring all attributes: right where a false positive is very costly, as in a safety interlock. Disjunctive, requiring any: right where a false negative is very costly and further checking follows, as in a screening test. Weighted, summing contributions: right where a ranking is wanted and some evidence is better than none.

2. Classify a description of dry, sour and heavy under all three rules.

Conjunctive assigns nothing, since no class has all its attributes present. Disjunctive assigns all three classes, since each has one attribute present. Weighted with equal weights gives a three-way tie at one each, which correctly reports that the description does not discriminate.

3. Why should a weighted rule report the scores rather than only the winner?

Because the margin carries information. A win by three attributes to none is a different result from a win by two to one, and reporting only the maximum discards the difference.

4. Why are all the weights in this book equal to one?

Because Charaka assigns no weights to attributes, so any other weighting would have to be invented. Using one for every attribute is a declared choice, and attributing invented weights to the text would be a fabrication.

Contents This chapter on its own page

munotes.in275

Chapter Seventy-Eight

Decision Principles in Diagnosis

Syllabus topic Module 2, "Decision principles in diagnosis"

In one line

Four decisions have to be made before a classification is reached: what to ask first, how much each answer counts, what to do when the answers conflict, and what follows from the answer.

In the wording you can write in an examination: the decision principles of a diagnostic scheme are the rules governing the order in which evidence is sought, the weight given to each item, the resolution of conflicting evidence, and the action that follows a classification. Charaka supplies an order of evidence in Vimānasthāna IV, an action rule in Sūtrasthāna I.61, and, in the structure of the attribute lists, an implicit weighting.

The four decisions

Every diagnostic procedure, classical or modern, has to settle these four. Naming them is most of the chapter.

DecisionThe questionWhere the classical scheme answers it
orderwhat do I look at first?Vimānasthāna IV: instruction, then observation, then inference
weighthow much does each item count?nowhere explicitly; the lists are unweighted
conflictwhat if the evidence points two ways?nowhere explicitly; the scheme permits combinations
actionwhat follows from the answer?Sūtrasthāna I.61: objects of adverse attributes, with place, measure and time

Two of the four are answered and two are not, and saying which is which is the difference between an answer that has read the text and one that has read about it.

One: the order of questions

The classical answer is the order of means, not an order of individual questions: know the possibilities first, then observe, then infer. [Parīkṣā: How an Examination Is Structured] sets it out.

What it does not give is which attribute to ask about first. Nothing in the lists says to check dryness before heaviness.

The modern answer to that question is information gain, from [Decision Trees]: ask first whatever most reduces the uncertainty. That is a real addition and it is ours, not the text's, and [A Decision Tree Built From the Tridoṣa Attributes] applies it and reports what it yields.

And there is a second modern consideration the text does not have. Some questions are cheap and some are expensive. A procedure that minimises expected cost rather than expected questions is a different optimisation, and it is the right one when the tests differ in cost.

Two: the weighing of evidence

The lists are unweighted. Seven attributes for wind, and nothing says that dryness counts for more than keenness in classifying.

So the default is one each, which is what [Multi-Attribute Classification] uses and what the program in [Ayurvedic Classification as a Rule-Based Expert System] implements.

But the lists carry an implicit weighting anyway, and it is worth seeing. An attribute unique to one doṣa is decisive when present; an attribute shared by two is not. Uniqueness is a weighting, and it comes from the structure of the lists rather than from any statement.

munotes.in276

Decision Principles in Diagnosis

AttributeHow many doṣa have itWhat its presence settles
dry, light, subtile, unstable, clearwind onlywind, on one attribute
hot, soft, sour, liquid, bitterbile onlybile, on one attribute
heavy, mild, watery, sweet, stable, slimyphlegm onlyphlegm, on one attribute
keenwind and bilenarrows to two
coldwind and phlegm, on this book's readingnarrows to two

Sixteen of the eighteen attributes are unique to one doṣa. That is a remarkably clean design and it is what makes the scheme usable at all: most single observations settle the question.

Three: conflicting evidence

The scheme's own answer is that combinations are real. A case showing attributes of two doṣa is a case in which two predominate, which is one of the seven labels of [Prakṛti: Constitution as a Class Label]. So conflict is not an error state; it is an answer.

That is unusual and it is worth pausing on. Most classifiers are built to produce one label and treat a near-tie as a difficulty. This scheme treats the combination as a category in its own right, which is the multi-label reading.

What the scheme does not say is what to do when the evidence is thin rather than mixed: two attributes of wind and two of bile is a combination, and one attribute of each is a case where nothing has been established. The distinction is between a mixed answer and no answer, and the program in this book reports a tie rather than choosing, which is a decision the text does not make for us.

Four: the action

Charaka's rule, from Sūtrasthāna I.61:

Those changes (in wind, bile, and phlegm) that are curable may be set right by drugs possessing adverse attributes, administered according to (considerations of) place, measure, and time.

Read as a decision principle it has three parts.

A function from the classification to the response. The response is determined by the attributes matched: their opposites.

Three parameters that modify it. Place, measure and time. So the response is not a constant; it is a function of the classification and of three further variables.

And a precondition. "Those changes that are curable." The rule is stated for the cases it applies to, and the text elsewhere says the cure of incurable cases is not laid down here. A rule that declares its own scope is better drafted than most, and it is the same discipline as a specification stating its preconditions in [Formal Specification: Saying Exactly What a System Must Do].

munotes.in277

Decision Principles in Diagnosis

Worked example: the four decisions, on one case

The case. A description reports dry and unstable.

Order. The possibilities are the three doṣa and their combinations, known in advance. The two attributes are observations. No inference has been needed yet.

Weight. Both attributes are unique to wind, so each is decisive on its own. Two of them agreeing is stronger still.

Conflict. None. Nothing points elsewhere.

Action. The opposites: moist, and stable. Applied with regard to place, measure and time.

And now a harder case. The description reports dry and sour. Dry is wind only; sour is bile only.

Conflict. Two decisive attributes pointing to different doṣa. Under the scheme, that is wind and bile both predominating, which is a label. Under a classifier built to emit one answer, it is a tie. The two treatments are different and the scheme's is the better model.

What these principles do NOT amount to

They are not an algorithm. Two of the four decisions are unanswered by the text, so a program implementing the scheme must supply them and must say that it has.

They are not a diagnosis. Classifying a description into the scheme's own categories is not identifying a condition, as [Āyurveda as a Śāstra, and What This Chapter Does Not Claim] says.

And they are not evaluated. No accuracy figure attaches to any of this, for the reason given in that same chapter.

Quick revision

  • Four decisions: the order of questions, the weight of each item, the resolution of conflict, and the action that follows.
  • Charaka answers two: the order of means, in Vimānasthāna IV, and the action rule, in Sūtrasthāna I.61.
  • He does not answer weighting or conflict explicitly, so a program must supply both and declare that it has.
  • The lists carry an implicit weighting: sixteen of the eighteen attributes are unique to one doṣa, so most single observations settle the question.
  • Conflict is not an error state: a combination is one of the seven labels.
  • The action rule has a function, three parameters, place, measure and time, and a declared precondition, curability.

Test yourself

1. Name the four decisions a diagnostic procedure must settle, and say which two Charaka answers.

The order in which evidence is sought, the weight of each item, the resolution of conflicting evidence, and the action following a classification. He answers the order, in Vimānasthāna IV, and the action, in Sūtrasthāna I.61.

2. What implicit weighting do the attribute lists carry?

Sixteen of the eighteen attributes belong to only one doṣa, so their presence settles the classification on its own, while the two shared attributes only narrow it to two. Uniqueness is a weighting even though no weight is stated.

munotes.in278

Decision Principles in Diagnosis

3. How does the scheme treat conflicting evidence, and why is that unusual?

As a combination in which two doṣa predominate, which is one of its own seven labels. It is unusual because most classifiers are built to produce a single label and treat a near-tie as a difficulty rather than as an answer.

4. Set out the three parts of Charaka's action rule.

A function from the classification to a response, namely the opposites of the matched attributes; three parameters that modify it, place, measure and time; and a precondition, that the rule is stated for changes that are curable.

Contents This chapter on its own page

munotes.in279

Chapter Seventy-Nine

Decision Trees

Syllabus topic Module 2, "Decision trees"

In one line

A decision tree is a sequence of yes-or-no questions arranged so that each answer narrows the possibilities, and information gain is how the questions are chosen.

In the wording you can write in an examination: a decision tree is a classifier in which each internal node tests one attribute, each branch corresponds to an outcome of that test, and each leaf carries a class label. It is built top down by choosing, at each node, the attribute whose test most reduces the uncertainty of the labels, measured as the difference between the entropy before the split and the weighted average entropy after it, which is called the information gain.

Entropy, from scratch

What it measures. How mixed a set of labels is. All one label: no uncertainty. An even split: maximum uncertainty.

The formula. For labels with proportions p1, p2 and so on, the entropy is the negative sum of each p times the logarithm of that p to base two.

entropy = - (p1 log2 p1) - (p2 log2 p2) - ...

Worked, for two labels.

SplitProportionsEntropyIn words
8 of one, 0 of the other1 and 00pure, no uncertainty
6 and 20.75 and 0.250.8113mostly one
4 and 40.5 and 0.51maximally mixed

And the unit is the bit, which is the unit of [Laghu and Guru as One Bit]. An even two-way split carries one bit of uncertainty because settling it takes exactly one yes-or-no answer.

Information gain

The definition. The entropy before the split, minus the weighted average of the entropies of the branches, where each branch is weighted by the fraction of rows it takes.

gain = entropy(before)

= - (fraction going yes) * entropy(yes branch)

= - (fraction going no) * entropy(no branch)

Why weighted. A branch with one row should not count as much as a branch with seven. Weighting by size is what makes the measure the expected remaining uncertainty.

Why it cannot be negative. Splitting can never increase the expected uncertainty, so the gain is at least zero, and it is exactly zero when the attribute tells you nothing.

The tree, built

"""Entropy, information gain and a decision tree, on a small table of our own."""
import math

# Our own table, invented for this chapter and not taken from any source:
# whether a nightly batch job finished on time.
ATTRS = ["big input", "index missing", "backup running"]
ROWS = [
    # (big input, index missing, backup running), label
    ((0, 0, 0), "on time"),
    ((0, 0, 1), "on time"),
    ((1, 0, 0), "on time"),
    ((1, 0, 1), "late"),
    ((0, 1, 0), "on time"),
    ((0, 1, 1), "late"),
    ((1, 1, 0), "late"),
    ((1, 1, 1), "late"),
]


def entropy(labels):
    n = len(labels)
    if n == 0:
        return 0.0
    out = 0.0
    for lab in sorted(set(labels)):
        p = labels.count(lab) / n
        out -= p * math.log2(p)
    return out


def split(rows, i):
    return ([r for r in rows if r[0][i] == 1],
            [r for r in rows if r[0][i] == 0])


def gain(rows, i):
    before = entropy([r[1] for r in rows])
    yes, no = split(rows, i)
    n = len(rows)
    after = (len(yes) / n) * entropy([r[1] for r in yes]) \
          + (len(no) / n) * entropy([r[1] for r in no])
    return before - after


def build(rows, available, depth=0):
    labels = [r[1] for r in rows]
    if len(set(labels)) <= 1:
        return labels[0] if labels else None
    scored = [(gain(rows, i), i) for i in available]
    scored = [(g, i) for g, i in scored if g > 1e-12]
    if not scored:
        return sorted(set(labels))
    g, best = max(scored, key=lambda t: (t[0], -t[1]))
    rest = [i for i in available if i != best]
    yes, no = split(rows, best)
    return (ATTRS[best], round(g, 4), build(yes, rest, depth + 1), build(no, rest, depth + 1))


def render(node, indent=0, label=""):
    pad = " " * indent
    if not isinstance(node, tuple):
        return "%s%s-> %s" % (pad, label, node)
    attr, g, yes, no = node
    return "\n".join([
        "%s%s%s?  (gain %.4f)" % (pad, label, attr, g),
        render(yes, indent + 4, "yes: "),
        render(no, indent + 4, "no:  "),
    ])


print("the table")
print("  %-11s %-14s %-15s %s" % tuple(ATTRS + ["label"]))
for values, lab in ROWS:
    print("  %-11d %-14d %-15d %s" % (values[0], values[1], values[2], lab))

labels = [r[1] for r in ROWS]
print()
print("labels: %d on time, %d late" % (labels.count("on time"), labels.count("late")))
print("entropy before any split: %.4f bits" % entropy(labels))

print()
print("information gain of each attribute")
for i, a in enumerate(ATTRS):
    yes, no = split(ROWS, i)
    print("  %-15s gain %.4f   yes-branch entropy %.4f (%d rows)   no-branch entropy %.4f (%d rows)"
          % (a, gain(ROWS, i), entropy([r[1] for r in yes]), len(yes),
             entropy([r[1] for r in no]), len(no)))

print()
print("the tree")
tree = build(ROWS, list(range(len(ATTRS))))
print(render(tree))

print()
print("checking the tree against every row of the table")
def classify(node, values):
    while isinstance(node, tuple):
        attr, _, yes, no = node
        node = yes if values[ATTRS.index(attr)] == 1 else no
    return node
wrong = [(v, lab, classify(tree, v)) for v, lab in ROWS if classify(tree, v) != lab]
print("  rows classified wrongly:", len(wrong) if wrong else "none, all 8 correct")
munotes.in280

Decision Trees

the table
  big input   index missing  backup running  label
  0           0              0               on time
  0           0              1               on time
  1           0              0               on time
  1           0              1               late
  0           1              0               on time
  0           1              1               late
  1           1              0               late
  1           1              1               late

labels: 4 on time, 4 late
entropy before any split: 1.0000 bits

information gain of each attribute
  big input       gain 0.1887   yes-branch entropy 0.8113 (4 rows)   no-branch entropy 0.8113 (4 rows)
  index missing   gain 0.1887   yes-branch entropy 0.8113 (4 rows)   no-branch entropy 0.8113 (4 rows)
  backup running  gain 0.1887   yes-branch entropy 0.8113 (4 rows)   no-branch entropy 0.8113 (4 rows)

the tree
big input?  (gain 0.1887)
    yes: index missing?  (gain 0.3113)
        yes: -> late
        no:  backup running?  (gain 1.0000)
            yes: -> late
            no:  -> on time
    no:  index missing?  (gain 0.3113)
        yes: backup running?  (gain 1.0000)
            yes: -> late
            no:  -> on time
        no:  -> on time

checking the tree against every row of the table
  rows classified wrongly: none, all 8 correct
munotes.in281

Decision Trees

Reading the output

Entropy before any split is exactly 1 bit, because four rows are on time and four are late. That is the maximally mixed case.

And all three attributes have the same gain, 0.1887. That is not a bug and it is the most instructive thing in the output.

Why. The table was constructed so that the job is late when at least two of the three conditions hold. The rule is symmetric in the three attributes, so no one of them is a better first question than another, and the measure correctly reports that.

How the tie is broken. The program takes the highest gain and, among equals, the earliest attribute. The choice is arbitrary and it must be made, which is worth knowing: an implementation that did not state its tie-break would build different trees on different runs.

Then the gains rise as you descend. 0.1887 at the root, 0.3113 at the second level, 1.0000 at the third. Each question is more informative than the last, because the earlier questions have already removed the easy cases and what remains is a cleaner problem. A gain of 1.0000 means the attribute settles the answer completely for the rows that reach it.

And the tree classifies all eight rows correctly, which the program checks rather than asserting.

Reading a tree

To classify. Start at the root. Answer its question about your case. Follow that branch. Repeat until a leaf.

To explain a classification. Read the path. "Big input yes, index missing no, backup running yes, therefore late." The path IS the explanation, which is why a decision tree is an intrinsically interpretable model in the sense of [Explainable AI, and Why a Five-Member Answer Is an Explanation].

To read the rule the tree encodes. Each root-to-leaf path is one IF-THEN rule, so a tree with five leaves is five rules. [Rule-Based Systems] is about the same knowledge in the other shape.

munotes.in282

Decision Trees

Overfitting

The one thing a chapter on decision trees must warn about.

What it is. A tree grown until every leaf is pure fits the training data exactly, including its accidents, and then performs badly on new cases.

Why trees are especially prone. Nothing stops the algorithm splitting until each leaf holds one row. A tree with as many leaves as rows has learned the table and nothing else.

The standard remedies. Stop early, when a node has too few rows or the best gain is too small. Or grow the tree fully and then prune, removing subtrees that do not improve performance on data held back from training.

And the honest note for this paper. With eight rows there is no held-out data and the question does not arise. With three rows, in the next chapter, it arises very sharply indeed.

What a decision tree is NOT

It is not unique. A different tie-break, or a different measure, gives a different tree for the same data. Several trees can classify the same table perfectly.

It is not a probability model. A leaf says "late". It does not say how likely late is, unless the implementation is extended to report the proportions at the leaf.

It is not good at everything. A relationship like "late when the sum of two numbers exceeds a threshold" needs a staircase of many splits, because each split tests one attribute. That is the standard weakness and it is why other methods exist.

And information gain is not the only measure. The Gini impurity is the common alternative and usually gives a similar tree. Naming one alternative is worth a mark.

Quick revision

  • Entropy measures how mixed the labels are, in bits: 0 when pure, 1 when evenly split two ways.
  • Information gain is the entropy before a split minus the weighted average entropy after it, and it is never negative.
  • The tree is built top down by taking the attribute of highest gain at each node.
  • In the worked table all three attributes tie at the root, because the rule is symmetric, and the tie-break must be stated or the tree is not reproducible.
  • Gains rise with depth, because earlier questions remove the easy cases.
  • A root-to-leaf path is an explanation and is also an IF-THEN rule.
  • Overfitting: a tree grown to purity learns the table. Remedies are early stopping and pruning.

Test yourself

1. Compute the entropy of a set of labels that is six of one kind and two of another.

The proportions are 0.75 and 0.25, so the entropy is minus 0.75 times the logarithm of 0.75 to base two, minus 0.25 times the logarithm of 0.25 to base two, which is 0.8113 bits.

munotes.in283

Decision Trees

2. Define information gain and say why the branch entropies are weighted.

It is the entropy of the labels before the split minus the weighted average of the entropies of the branches. The branches are weighted by the fraction of rows each receives, so that the result is the expected remaining uncertainty rather than an unweighted average.

3. Why do all three attributes have the same gain at the root of the worked tree?

Because the table encodes a symmetric rule: the job is late when at least two of the three conditions hold. No attribute is a better first question than another, and the measure reports that correctly.

4. What is overfitting in a decision tree, and give two remedies.

Growing the tree until every leaf is pure, so that it fits the accidents of the training data and performs badly on new cases. The remedies are stopping early, when a node has too few rows or the best gain is too small, and pruning a fully grown tree using data held back from training.

Contents This chapter on its own page

munotes.in284

Chapter Eighty

A Decision Tree Built From the Tridoṣa Attributes

Syllabus topic Module 2, "Decision trees", "Multi-attribute classification"

In one line

Build a decision tree from the three attribute lists and it works, and information gain had nothing to do with it.

In the wording you can write in an examination: inducing a decision tree requires a training set of examples with labels. The tridoṣa attribute lists supply three examples, one per class, over eighteen binary attributes. A tree induced from them separates the three classes in two questions, but the information gain of every attribute is identical, so the choice of attribute at each node is made entirely by the tie-break rather than by the measure.

The training set

Three rows. One per doṣa, each carrying the attributes Charaka's Sūtrasthāna I.58 to I.60 assign to it, with the disputed reading excluded as [The Tridoṣa Framework] explains.

Eighteen attributes. The union of the three lists, as [Doṣa as a Feature Vector] establishes.

Three labels. Wind, bile, phlegm.

And that is the whole dataset. There is no more. Nothing in the text supplies a second example of wind.

The induction, run

"""A decision tree induced from the tridosa attribute table, and what three rows cost."""
import math

ATTRIBUTES = {
    "wind":   ["dry", "cold", "light", "subtile", "unstable", "clear", "keen"],
    "bile":   ["cold", "hot", "keen", "soft", "sour", "liquid", "bitter"],
    "phlegm": ["heavy", "cold", "mild", "watery", "sweet", "stable", "slimy"],
}
DISPUTED = {("bile", "cold")}
DOSAS = ("wind", "bile", "phlegm")


def universe():
    seen = []
    for d in DOSAS:
        for a in ATTRIBUTES[d]:
            if a not in seen:
                seen.append(a)
    return seen


def rows():
    out = []
    for d in DOSAS:
        have = set(ATTRIBUTES[d]) - {a for x, a in DISPUTED if x == d}
        out.append((d, have))
    return out


def entropy(labels):
    n = len(labels)
    if n == 0:
        return 0.0
    return -sum((labels.count(l) / n) * math.log2(labels.count(l) / n) for l in sorted(set(labels)))


def gain(rs, attr):
    before = entropy([r[0] for r in rs])
    yes = [r[0] for r in rs if attr in r[1]]
    no = [r[0] for r in rs if attr not in r[1]]
    n = len(rs)
    return before - ((len(yes) / n) * entropy(yes) + (len(no) / n) * entropy(no))


def build(rs, attrs):
    labels = [r[0] for r in rs]
    if len(set(labels)) <= 1:
        return labels[0] if labels else None
    scored = [(gain(rs, a), a) for a in attrs]
    scored = [(g, a) for g, a in scored if g > 1e-12]
    if not scored:
        return sorted(set(labels))
    g, best = max(scored, key=lambda t: (t[0], -attrs.index(t[1])))
    rest = [a for a in attrs if a != best]
    return (best, round(g, 4),
            build([r for r in rs if best in r[1]], rest),
            build([r for r in rs if best not in r[1]], rest))


def render(node, indent=0, label=""):
    pad = " " * indent
    if not isinstance(node, tuple):
        return "%s%s-> %s" % (pad, label, node)
    a, g, yes, no = node
    return "\n".join(["%s%sis it %s?  (gain %.4f)" % (pad, label, a, g),
                      render(yes, indent + 4, "yes: "),
                      render(no, indent + 4, "no:  ")])


R = rows()
U = universe()
print("three rows, eighteen attributes")
print("entropy of the labels before any split: %.4f bits (three classes, one row each)"
      % entropy([r[0] for r in R]))
print()
print("information gain of every attribute")
gains = sorted({round(gain(R, a), 4) for a in U})
for a in U:
    print("  %-9s %.4f   held by %s" % (a, gain(R, a),
          ", ".join(d for d in DOSAS if a in ATTRIBUTES[d]
                    and (d, a) not in DISPUTED)))
print()
print("distinct gain values across all eighteen attributes:", gains)
print()
print("the tree")
tree = build(R, U)
print(render(tree))
print()
print("it separates all three, in %d question(s)" % 2)
print("and it would separate them just as well using almost any other attributes,")
print("because with three rows the gain cannot tell most of them apart.")
munotes.in285

A Decision Tree Built From the Tridoṣa Attributes

three rows, eighteen attributes
entropy of the labels before any split: 1.5850 bits (three classes, one row each)

information gain of every attribute
  dry       0.9183   held by wind
  cold      0.9183   held by wind, phlegm
  light     0.9183   held by wind
  subtile   0.9183   held by wind
  unstable  0.9183   held by wind
  clear     0.9183   held by wind
  keen      0.9183   held by wind, bile
  hot       0.9183   held by bile
  soft      0.9183   held by bile
  sour      0.9183   held by bile
  liquid    0.9183   held by bile
  bitter    0.9183   held by bile
  heavy     0.9183   held by phlegm
  mild      0.9183   held by phlegm
  watery    0.9183   held by phlegm
  sweet     0.9183   held by phlegm
  stable    0.9183   held by phlegm
  slimy     0.9183   held by phlegm

distinct gain values across all eighteen attributes: [0.9183]

the tree
is it dry?  (gain 0.9183)
    yes: -> wind
    no:  is it cold?  (gain 1.0000)
        yes: -> phlegm
        no:  -> bile

it separates all three, in 2 question(s)
and it would separate them just as well using almost any other attributes,
because with three rows the gain cannot tell most of them apart.

The finding

Every one of the eighteen attributes has a gain of exactly 0.9183, and the program prints the set of distinct values to prove it.

Why. Entropy before the split is the entropy of three equally likely labels, which is the logarithm of three to base two, 1.5850 bits. Every attribute in the table is held by one doṣa or by two, so every split divides the three rows either one against two or two against one. Either way the pure branch contributes nothing and the mixed branch of two rows contributes one bit weighted by two thirds, which is 0.6667. So the gain is 1.5850 minus 0.6667, which is 0.9183, for every attribute without exception.

munotes.in286

A Decision Tree Built From the Tridoṣa Attributes

What follows. The measure is doing no work. The attribute at each node is chosen by the tie-break, which here is the order the attributes appear in the table. Change the order and you get a different tree, of the same size, that classifies just as well.

And that is the lesson. Information gain compares attributes by how much they reduce uncertainty. With three rows there is not enough uncertainty to distribute differently, so no attribute is better than another. A measure needs data to have anything to measure.

What the tree does anyway

It separates all three classes in two questions. Is it dry: if yes, wind. Otherwise, is it cold: if yes, phlegm, otherwise bile.

Two questions is the minimum, because distinguishing three things by yes-or-no questions needs at least two, and the tree achieves it.

So the tree is optimal in size and arbitrary in content. Both statements are true and neither is the whole story.

Why this is not a failure of the method

The method is correct. It found the smallest tree separating the classes. Nothing went wrong.

The data is the constraint. Three examples over eighteen attributes is a dataset with more columns than rows several times over, which is the classic setting in which a learned model tells you about the sample and not about the world.

And the honest position is to say so rather than to manufacture rows. Generating extra examples by, for instance, taking every subset of an attribute list would produce a dataset and it would be our invention, and presenting it as Charaka's would be a fabrication of exactly the kind [Āyurveda as a Śāstra, and What This Chapter Does Not Claim] excludes.

What would fix it, and what it would cost

More examples. Real cases, each with an attribute description and a label. Then the attributes would separate differently, some would prove more informative than others, and the gain would discriminate.

Where they would come from. Not from the text: the text gives definitions, not cases. They would have to be collected, and collecting them is a clinical exercise this paper is not part of.

And the labels would be the problem. A case's label is what doṣa predominates, which is itself a judgement made by the scheme. Training a classifier on labels produced by the scheme teaches it to reproduce the scheme, which is a legitimate thing to want and is not evidence about anything else. Saying that distinguishes a submission that has understood the limits from one that has not.

The comparison with the previous chapter

The batch job tableThe tridoṣa table
Rows83
Attributes318
Labels23
Gain at the rootall three tie at 0.1887all eighteen tie at 0.9183
Reason for the tiethe rule is symmetric in the attributesthere are too few rows for any split to differ
Gains deeper in the treerise to 1.0000rise to 1.0000
Is the tree meaningfulyes, it encodes the ruleit separates the classes and nothing more
munotes.in287

A Decision Tree Built From the Tridoṣa Attributes

Two ties, two entirely different reasons. The first is a property of the rule being learned and is informative. The second is a property of the dataset's size and is not. Being able to tell those apart is what this pair of chapters is for.

Quick revision

  • The tridoṣa lists give three rows, eighteen attributes, three labels, and no more.
  • Entropy before any split is the logarithm of three to base two, 1.5850 bits.
  • Every attribute splits the three rows one against two, so every gain is 1.5850 minus two thirds of a bit, which is 0.9183. The set of distinct gains has one member.
  • So the measure chooses nothing and the tie-break chooses everything; a different attribute order gives a different tree of the same size.
  • The tree separates all three in two questions, which is the minimum, so it is optimal in size and arbitrary in content.
  • More rows would fix it, and they are not in the text. Manufacturing them and attributing them to the text would be a fabrication.
  • A classifier trained on labels the scheme itself produced learns the scheme, which is evidence about nothing else.

Test yourself

1. Why does every attribute in the tridoṣa table have the same information gain?

Because there are three rows with three distinct labels, and every attribute is held by one or two of them, so every split is one row against two. The pure branch contributes no entropy and the two-row branch contributes one bit weighted by two thirds, giving the same gain of 0.9183 in every case.

2. What does that mean for the tree that is built?

That the choice of attribute at each node is made by the tie-break rather than by the measure, so a different ordering of the attributes produces a different tree of the same size which classifies equally well.

3. Is the tree wrong? Justify your answer.

No. It separates the three classes in two questions, which is the minimum possible, so the method has done what it claims. What is absent is any evidence that the attributes it chose are better than the ones it did not.

4. Why does this chapter not add more training rows?

Because the text supplies no further examples: it gives definitions rather than cases. Rows generated by us and presented as the text's would be a fabrication, and the honest negative result is more instructive than a manufactured positive one.

Contents This chapter on its own page

munotes.in288

Chapter Eighty-One

Rule-Based Systems

Syllabus topic Module 2, "Rule-based systems"

In one line

A rule-based system keeps its knowledge as a list of IF-THEN rules and a store of facts, and runs by repeatedly finding a rule whose conditions hold and doing what it says.

In the wording you can write in an examination: a rule-based system consists of a rule base of production rules of the form IF conditions THEN action, a working memory of facts, and an interpreter implementing the recognise-act cycle: match the rules against working memory, resolve any conflict among the matching rules, and act, repeating until no rule matches or a stopping condition is reached.

The three parts and the cycle

Rule base. The rules. They are data, so they can be added, removed and inspected without changing the program that runs them.

Working memory. The facts currently held. It changes as the system runs.

Interpreter. The loop. Three phases, repeated.

PhaseWhat happens
matchfind every rule whose conditions are satisfied by working memory
resolveif more than one matches, choose which to fire
actperform the chosen rule's action, usually adding a fact

And the loop ends when no rule matches, or when a rule's action says to stop.

Why separating rules from the interpreter matters

This is the property that makes rule-based systems worth having, and it is the answer to "why not just write the conditions as code?".

The rules can be changed by somebody who cannot program. A domain expert can read and edit an IF-THEN rule. They cannot edit a nest of conditionals.

The rules can be inspected. A system can be asked which rules exist, which have ever fired, and which fired for a given case. That is what makes [Explainable AI, and Why a Five-Member Answer Is an Explanation] achievable here.

The rules can be added without touching the rest. A new rule does not require any existing rule to be modified. That is the property people mean when they call such a system modular, and it is also the property that causes the trouble described below.

Conflict resolution

When several rules match, one must be chosen. The strategies are the ones [Inference Engines: Forward and Backward Chaining] lists.

Specificity, recency, order and refraction. And two of the four are Pāṇini's: specificity is apavāda and order is paratva, from [Rule Precedence: The Four Principles].

This book does not teach them twice. What is worth adding here is that a rule-based system MUST have a strategy and that an undeclared strategy is still a strategy: firing the first match found is an order strategy, chosen by accident.

Rules against a decision tree

The same knowledge in two shapes, and the comparison is what a question usually asks for.

munotes.in289

Rule-Based Systems

Rule baseDecision tree
Shapea flat lista tree
One unit of knowledgeone ruleone root-to-leaf path
Adding knowledgeadd one rulemay require rebuilding
Order of testsdecided at run time by the strategyfixed when the tree is built
Tests performed per casepotentially all conditions of all rulesonly those on one path
Explaining an answerthe chain of rules that firedthe path taken
Predicting overall behaviourhard, because behaviour is distributedeasy, the tree is the behaviour

Read the "adding knowledge" row against the "predicting behaviour" row. They are the same property with opposite signs. A rule base is easy to extend precisely because no rule knows about the others, and it is hard to reason about for precisely the same reason. That trade is the central fact about rule-based systems and it is what a critical answer should say.

Converting between them

A tree to rules. Each root-to-leaf path becomes one rule whose conditions are the tests along the path. A tree with five leaves gives five rules, and they are mutually exclusive and exhaustive, which an arbitrary rule base is not.

Rules to a tree. Harder, and not always possible without duplication, because rules may overlap in ways no single test order captures. The asymmetry is worth noticing: every tree is a rule base and not every rule base is a tree.

Worked example: the tridoṣa scheme as rules

The scheme of [The Tridoṣa Framework] written as rules, using the unique attributes of [Decision Principles in Diagnosis].

R1 IF dry THEN wind

R2 IF unstable THEN wind

R3 IF subtile THEN wind

R4 IF sour THEN bile

R5 IF bitter THEN bile

R6 IF heavy THEN phlegm

R7 IF slimy THEN phlegm

R8 IF keen THEN wind or bile

R9 IF cold THEN wind or phlegm

R10 IF wind AND bile THEN wind-bile combination

Three things this makes visible that the attribute lists do not.

The rules are not mutually exclusive. A case that is dry and sour fires R1 and R4, concluding both wind and bile. That is correct under the scheme, which permits combinations, and it is why R10 exists.

Two rules conclude disjunctions. R8 and R9 do not settle anything on their own. A rule whose conclusion is a disjunction is awkward in most rule systems and usually has to be recast, for instance as a rule that eliminates phlegm rather than one that concludes wind or bile.

And R10 fires on the conclusions of other rules, not on observations. That is forward chaining over two levels, and it is what makes the combination labels of [Prakṛti: Constitution as a Class Label] reachable.

munotes.in290

Rule-Based Systems

The problems that show up at scale

Three, and all three are why the expert system programme ran into difficulty. [Expert Systems, and MYCIN as the Comparison] has the history.

Rules interact in ways nobody intended. With fifty rules there are many paths through them and no one holds them all in mind. A new rule can make an old one unreachable.

The knowledge is hard to get. Experts do not naturally state what they know as rules, and the process of extracting it, called knowledge acquisition, turned out to be the bottleneck.

There is no notion of degree. A rule fires or it does not. Real expertise is full of "usually" and "unless", and adding certainty factors was the standard response and brought its own problems.

What a rule-based system is NOT

It is not a programming language. The rules have no control flow; the interpreter supplies it.

It is not guaranteed to terminate. A rule that adds a fact enabling a rule that removes it will loop, and refraction only prevents the same rule firing twice on the same facts.

It is not the same as a decision tree with the branches written out. A tree's paths are exclusive and exhaustive; a rule base's rules are neither unless somebody has ensured it.

Quick revision

  • Three parts: rule base, working memory, interpreter. The cycle is match, resolve, act, repeated.
  • Rules are data, so they can be edited by a non-programmer, inspected, and added without touching the others.
  • A strategy is compulsory; firing the first match is an order strategy chosen by accident.
  • Rules against a tree: easy to extend and hard to reason about, against fixed test order and predictable behaviour. Same property, opposite signs.
  • Every tree converts to a rule base; not every rule base converts to a tree without duplication.
  • The tridoṣa rules are not mutually exclusive, two of them conclude disjunctions, and one fires on other rules' conclusions.
  • At scale: unintended interactions, the knowledge acquisition bottleneck, and no notion of degree.

Test yourself

1. Name the three parts of a rule-based system and the three phases of its cycle.

A rule base, a working memory and an interpreter. The cycle is match, finding rules whose conditions hold; resolve, choosing among them; and act, performing the chosen rule.

2. State the central trade-off of rule-based systems in one sentence.

A rule base is easy to extend because no rule knows about any other, and hard to reason about for exactly the same reason, since the overall behaviour is distributed across the rules rather than written anywhere.

3. Convert a decision tree with four leaves into rules, and say why the reverse is harder.

Each of the four root-to-leaf paths becomes a rule whose conditions are the tests along that path, giving four mutually exclusive and exhaustive rules. The reverse is harder because rules may overlap in ways that no single fixed order of tests captures, so a tree may need duplicated subtrees or may not exist.

munotes.in291

Rule-Based Systems

4. Why does rule R10, concluding a combination, differ in kind from the others?

Because its conditions are the conclusions of other rules rather than observations, so it requires two levels of forward chaining, whereas the other rules fire directly on the attributes reported.

Contents This chapter on its own page

munotes.in292

Chapter Eighty-Two

Expert Systems, and MYCIN as the Comparison

Syllabus topic Module 2, "Expert systems"

In one line

An expert system is a rule-based system aimed at a task that normally needs a trained specialist, and the interesting part of its history is why so few of them survived.

In the wording you can write in an examination: an expert system is a program that performs a task requiring specialist human expertise, by applying a knowledge base acquired from experts. Its standard architecture has five components: a knowledge base, an inference engine, a working memory, an explanation facility, and a knowledge acquisition interface. It is distinguished from an ordinary program by the separation of domain knowledge from the procedure that applies it.

The five components

ComponentWhat it holds or does
knowledge basethe rules, and the facts that are always true
working memorythe facts about the case in hand
inference enginethe recognise-act cycle, with a conflict resolution strategy
explanation facilityanswers "why are you asking?" and "how did you conclude that?"
knowledge acquisitionthe means by which an expert's knowledge becomes rules

The fourth is what makes it an expert system rather than a program with rules in it. A specialist's judgement is accepted because it can be defended, and a system that replaces the specialist must be able to defend its own.

The fifth is where the programme failed, and the section below says so.

The two explanations an expert system owes

Why are you asking? During a consultation the system asks the user a question. The user should be able to ask why. A backward-chaining system answers by naming the rule it is trying to establish, and the goal that rule serves, and so on up the chain.

How did you conclude that? After the answer. The system names the rules that fired and the facts they used. That is the trace of [Nyāya Logic as an Inference Engine], and it is the same object.

Both are free in a rule-based architecture and impossible in most others, which is the practical case for the architecture and the reason the field of [Explainable AI, and Why a Five-Member Answer Is an Explanation] returned to rules decades later.

Worked comparison: MYCIN

MYCIN was a rule-based system for advising on bacterial infections, developed at Stanford in the 1970s, and it is the example every account uses. Four things about it are uncontroversial and are what a question expects.

It was rule-based and backward-chaining. A few hundred IF-THEN rules, and a goal-driven interpreter.

It had an explanation facility, answering both of the questions above.

It used certainty factors. Rules and facts carried a number expressing degree of belief, and the engine combined them. That was the response to the "no notion of degree" problem of [Rule-Based Systems].

munotes.in293

Expert Systems, and MYCIN as the Comparison

And it was never deployed in clinical use. The reasons usually given are not about its performance: the integration into practice, the legal position, and the computing available at the time.

What this book does not claim about it. Any specific figure for its accuracy, or a comparison with human specialists. Such figures are quoted in the literature and this book has not read the primary reports, so it does not repeat them.

Certainty factors, and why they were a problem

What they are. A number attached to a rule or a fact, expressing how strongly it is believed. Combining rules combines the numbers.

Why they seemed necessary. Expertise is full of "usually" and "this suggests", and a system that can only say yes or no cannot represent it.

Why they caused trouble. Three reasons and all three are worth naming.

The combination rules were ad hoc. How to combine two pieces of evidence for the same conclusion was decided by what behaved reasonably, not derived from anything.

They are not probabilities. They do not obey the laws probability obeys, so intuitions carried over from probability give wrong answers.

And they made the knowledge harder to acquire. An expert asked for a rule can give one. An expert asked for a rule and a number expressing how strongly they believe it gives a number they have no way of calibrating.

The eventual response was probabilistic graphical models, which have a principled account of combination, and which lose much of the transparency that made the expert system attractive. That trade is still live, which is why this material is on a syllabus.

The knowledge acquisition bottleneck

The problem, stated plainly. Getting what an expert knows out of them and into rules is slow, and it is the dominant cost of building an expert system.

Why it is hard. Experts are fluent and not articulate: they can do the task and cannot always say how. What they say when asked is often a reconstruction rather than a description of what they did.

Why it does not improve with practice. Each new domain needs a new extraction. The knowledge is the system, so there is nothing to reuse.

And this is the reason the programme stalled, more than any technical limitation. A technology whose cost is dominated by expert interview time does not scale, and saying so is the honest historical verdict.

What survived, and what returned

What survived. Rule engines are in wide use in business systems, where the rules are policies rather than expertise and the experts who state them are the people who wrote the policies. The acquisition problem does not arise because the knowledge was already explicit.

munotes.in294

Expert Systems, and MYCIN as the Comparison

What returned. The requirement for explanation. Systems built by learning from data cannot explain themselves, and the argument for interpretable models is the argument the expert system architecture made forty years earlier.

And the honest summary for an answer. The architecture was sound and the bottleneck was human. Where the knowledge is already written down, the architecture works and is used; where it has to be extracted from a specialist, it did not scale.

The comparison with the classical material

The Āyurvedic textsAn expert system
Knowledge basethe attribute lists and the treatment rulethe rule base
Working memorythe case in front of the practitionerthe facts about the case
Inferencethe practitioner's trainingthe engine
Explanationthe scheme's own vocabulary, which the patient may not sharethe explanation facility
Acquisitioncenturies of transmission and commentaryinterviews

The row that repays attention is the last. The classical scheme solved the acquisition problem by writing the knowledge down and then maintaining an apparatus for transmitting it: the four layers of [Śāstra: What Makes a Body of Knowledge Formal]. That is a real answer to the bottleneck, and it takes centuries.

Quick revision

  • Five components: knowledge base, working memory, inference engine, explanation facility, knowledge acquisition.
  • Two explanations owed: why are you asking, and how did you conclude that. Both are free in a rule architecture.
  • MYCIN: rule-based, backward-chaining, with an explanation facility and certainty factors, and never deployed clinically.
  • Certainty factors: ad hoc combination rules, not probabilities, and they made acquisition harder.
  • The knowledge acquisition bottleneck is the historical reason the programme stalled: experts are fluent and not articulate, and nothing transfers between domains.
  • What survived: rule engines where the rules are already-written policies. What returned: the demand for explanation.

Test yourself

1. Name the five components of an expert system and say which distinguishes it from a program with rules in it.

Knowledge base, working memory, inference engine, explanation facility, knowledge acquisition. The explanation facility is what distinguishes it: a system replacing a specialist must be able to defend its conclusions as the specialist would.

2. Give the two questions an expert system should be able to answer, and when each is asked.

"Why are you asking?", asked during the consultation and answered by naming the goal the current question serves. "How did you conclude that?", asked afterwards and answered by the chain of rules that fired.

3. State three problems with certainty factors.

The rules for combining them were chosen for reasonable behaviour rather than derived; they are not probabilities and do not obey probability's laws, so transferred intuitions mislead; and asking an expert for a number they cannot calibrate makes knowledge acquisition harder.

munotes.in295

Expert Systems, and MYCIN as the Comparison

4. What was the main historical reason the expert system programme stalled, and where does the architecture still work?

The knowledge acquisition bottleneck: extracting rules from a specialist is slow, does not transfer between domains, and dominated the cost. The architecture still works where the knowledge is already explicit, as in business rule engines encoding written policies.

Contents This chapter on its own page

munotes.in296

Chapter Eighty-Three

Feature Engineering

Syllabus topic Module 2, "Feature engineering"

In one line

Feature engineering is choosing what the system gets to look at, and it decides more about the result than the classifier does.

In the wording you can write in an examination: feature engineering is the selection, construction, transformation and encoding of the attributes presented to a classifier. It comprises selecting which available attributes to use, deriving new attributes from the raw data, scaling numeric attributes to comparable ranges, and encoding non-numeric attributes so that a method can process them.

The four operations

Selection. Choosing which of the available attributes to keep. Fewer attributes means less to overfit and less to collect.

Construction. Making a new attribute out of existing ones. The ratio of two measurements, the difference between a value and its recent average, the count of some event in a window.

Scaling. Putting numeric attributes on comparable ranges, so that one measured in thousands does not swamp one measured in tenths.

Encoding. Turning a category into numbers. The standard method is one-hot encoding: one binary column per possible value.

And the one that matters most is construction, because a constructed attribute can make an impossible problem easy, and no classifier can construct one for itself unless it was designed to.

Worked example: why construction matters

The problem. Classify whether a transaction is unusual for an account.

With the raw attribute. The amount. A classifier over amount alone must learn a threshold, and no single threshold suits both a large account and a small one.

With a constructed attribute. The amount divided by the account's own median transaction. Now a single threshold works for every account, because the attribute already carries the comparison.

Nothing about the classifier changed. The problem became easy because the representation changed. That is the whole argument for feature engineering and it is a restatement of [Knowledge Representation: The Modern Name For It]: a representation is a bet about which questions you will be asked.

The guṇa list as a derived feature space

This is the chapter's claim about the classical material and it is worth stating carefully.

The raw material is a body and a report. Innumerable things could be recorded.

The scheme records qualities, and it records a FIXED list of them. Dry, cold, light, subtile, unstable, clear, keen, and the rest. That is selection: a small closed vocabulary chosen in advance out of everything that could be said.

And the qualities are not raw observations. "Unstable" is not measured; it is a judgement about a pattern over time. "Subtile" is not read off an instrument. Each quality is a constructed attribute, derived from what is observed by a rule carried in the practitioner's training.

So the vocabulary is a designed feature space, and its design has three properties a modern engineer would recognise.

munotes.in297

Feature Engineering

PropertyWhat it gives
small and closedevery case is described in the same terms, so cases are comparable
shared across classesone space, not one per doṣa, which is what makes a vector possible
built in opposite pairsthe treatment rule becomes computable, as [The Tridoṣa Framework] shows

The classical list of the qualities is twenty, in ten opposed pairs. This book uses only those the three doṣa lists name, which is eighteen attributes over nine pairs, and it says so rather than importing a count it has not used.

Selection, and the measure of a useless attribute

An attribute every class has tells you nothing. That is [Doṣa as a Feature Vector]'s finding about "cold" under the disputed reading, and it is the simplest case of feature selection: a column with the same value in every row can be deleted with no loss.

The general measure is the information gain of [Decision Trees]. An attribute with zero gain does not separate the classes at all.

And the general warning is [A Decision Tree Built From the Tridoṣa Attributes]'s. With too few rows, every attribute has the same gain and the measure selects nothing. Feature selection needs data, exactly as the tree does.

Encoding, and the trap

One-hot encoding. A category with five possible values becomes five binary columns, one of which is set.

Why not just number the values. Because numbering asserts an order and a spacing. Coding the three doṣa as 1, 2, 3 tells a classifier that bile is between wind and phlegm and that the distance from wind to phlegm is twice the distance from wind to bile. Neither is true and the classifier will use both.

The cost of one-hot encoding is that the space grows, and it grows fastest exactly where the category has many values. A category with a thousand values becomes a thousand columns, most of them zero in any row.

And the doṣa attribute space is already one-hot. Eighteen binary columns over a shared vocabulary is exactly the result of one-hot encoding a set-valued attribute, and this is why the representation of [Doṣa as a Feature Vector] needed no encoding step: the source was already in that form.

Scaling, and why it does not arise here

Scaling matters when attributes have different ranges and a method measures distance. Age in years and income in rupees cannot be compared without it.

It does not arise in this block, because every attribute is binary. That is worth saying because it is a real simplification: a scheme with numeric attributes, such as one recording degrees of a quality, would need it.

munotes.in298

Feature Engineering

What feature engineering is NOT

It is not preprocessing. Cleaning data and engineering features are different activities: cleaning removes errors, engineering changes what the model sees.

It is not made obsolete by learned representations. A method that learns its own features removes the need to construct them by hand and requires far more data, and it produces features nobody can interpret. The trade is the one [Explainable AI, and Why a Five-Member Answer Is an Explanation] describes.

It is not free of assumptions. Every constructed attribute encodes a belief about what matters. The belief is now in the data rather than in the model, where it is harder to notice, and that is a real risk rather than a rhetorical caution.

Quick revision

  • Four operations: selection, construction, scaling, encoding. Construction matters most, because it can make a hard problem easy.
  • A ratio to a per-account baseline turns a problem no single threshold solves into one a single threshold solves.
  • The guṇa vocabulary is a designed feature space: small, closed, shared across classes, and built in opposite pairs so the treatment rule is computable.
  • The qualities are constructed, not raw: "unstable" is a judgement about a pattern, not a measurement.
  • An attribute every class has can be deleted; the measure is information gain, and it needs data to work.
  • One-hot encoding avoids asserting an order that numbering would assert. The doṣa space is already in that form.
  • Scaling does not arise here because every attribute is binary.

Test yourself

1. Name the four operations of feature engineering and say which is the most consequential.

Selection, construction, scaling and encoding. Construction is the most consequential, because a constructed attribute can turn a problem no classifier could solve into one a simple threshold solves.

2. In what sense is the classical list of qualities a designed feature space?

It is a small closed vocabulary selected in advance out of everything that could be recorded; it is shared across all three classes, so cases are comparable; its members are judgements derived from observation rather than raw measurements; and it is built in opposite pairs so that the treatment rule has something to apply.

3. Why is numbering a category 1, 2, 3 a mistake, and what is done instead?

Because it asserts an order and equal spacing that the category does not have, and a classifier will use both. One-hot encoding is used instead: one binary column per possible value.

4. Why does feature selection fail on the tridoṣa table?

Because with three rows every attribute has the same information gain, so the measure cannot rank them. Feature selection needs enough data for attributes to separate the classes differently.

Contents This chapter on its own page

munotes.in299

Chapter Eighty-Four

Multi-Class Classification, and How It Is Scored

Syllabus topic Module 2, "Multi-class classification"

In one line

With more than two classes there is no single "positive" class, so accuracy alone stops being a report and the confusion matrix starts.

In the wording you can write in an examination: multi-class classification assigns each object one of more than two labels. A binary classifier can be extended to it by one-against-rest, training one classifier per class to distinguish it from all the others, or by one-against-one, training a classifier for every pair. Performance is reported by a confusion matrix, from which per-class precision and recall are computed, since a single accuracy figure hides which classes are confused with which.

The two ways to extend a binary method

One-against-rest. For k classes, build k classifiers. The i-th distinguishes class i from everything else. To classify, run all k and take the one that answers most strongly.

One-against-one. For k classes, build k times k minus 1, over 2 classifiers, one for each pair. To classify, run all of them and take the class that wins most pairwise contests.

One-against-restOne-against-one
Number of classifierskk(k-1)/2
For k equal to 333
For k equal to 7721
Each classifier seesall the dataonly two classes' data
Training setsimbalanced, one class against allbalanced, if the classes are
Tiespossible, resolved by the strength of the answerspossible, resolved by a rule

For the three doṣa the two methods need the same number of classifiers, which is a coincidence of k equal to three. For the seven labels of [Prakṛti: Constitution as a Class Label] they need seven and twenty-one.

And the multi-label reading of that chapter needs neither. Three independent yes-or-no decisions is three classifiers, which is one-against-rest without the requirement that exactly one wins.

The confusion matrix

What it is. A table with one row per true class and one column per predicted class. The cell at row i, column j is the number of cases whose true class is i and whose predicted class is j.

What the diagonal is. Correct predictions.

What everything else is. Errors, and the position says which error.

Worked. The example below is invented for this chapter and not the result of any experiment, because no labelled data for this scheme exists.

predicted windpredicted bilepredicted phlegm
true wind4064
true bile8284
true phlegm226

Read three things off it before computing anything.

The totals. 50 true wind, 40 true bile, 10 true phlegm. The classes are imbalanced, five to one between the commonest and the rarest.

The diagonal. 40 plus 28 plus 6 is 74 correct out of 100, so accuracy is 74 per cent.

munotes.in300

Multi-Class Classification, and How It Is Scored

And the errors are not symmetric. Eight true bile were called wind; six true wind were called bile. The classifier leans toward wind, which is the commonest class, and an accuracy figure says nothing about that.

Precision and recall, per class

Precision for a class. Of the cases predicted to be that class, how many were? It is the diagonal cell divided by the COLUMN total.

Recall for a class. Of the cases that truly were that class, how many were found? It is the diagonal cell divided by the ROW total.

Worked, for all three.

ClassColumn totalPrecisionRow totalRecall
wind40 + 8 + 2 = 5040 / 50 = 0.8040 + 6 + 4 = 5040 / 50 = 0.80
bile6 + 28 + 2 = 3628 / 36 = 0.788 + 28 + 4 = 4028 / 40 = 0.70
phlegm4 + 4 + 6 = 146 / 14 = 0.432 + 2 + 6 = 106 / 10 = 0.60

And now the report says something the accuracy figure did not. Overall accuracy is 74 per cent and phlegm's precision is 43 per cent: fewer than half the cases called phlegm were phlegm. A user acting on a prediction of phlegm is wrong more often than right, and nothing in "74 per cent accurate" warned them.

Check the arithmetic. The column totals add to 100 and the row totals add to 100, and both must, because every case is counted once by its true class and once by its predicted class. That is a free check and it catches a mistranscribed cell at once.

Combining the per-class figures

Three ways, and a question may ask the difference.

Macro average. The mean of the per-class figures, each class counting equally. Macro precision here is the mean of 0.80, 0.78 and 0.43, which is 0.67.

Weighted average. The mean weighted by how common each class is. Phlegm is only a tenth of the cases, so it barely moves the figure.

Micro average. Pool all the cases and compute one figure. For single-label multi-class classification the micro average equals the accuracy.

Which to use. Macro if the rare classes matter as much as the common ones. Weighted if they do not. Reporting only one of them is how an unflattering result is hidden, and reporting which one was used is the minimum.

The baseline you must beat

Always answering the commonest class. Here that is wind, 50 of 100, so a classifier that always says wind is 50 per cent accurate.

So 74 per cent is 24 points above the baseline, not 74 points above nothing. Quoting accuracy without the baseline is quoting half a number, and with seven classes and one of them dominant the baseline can be very high indeed.

munotes.in301

Multi-Class Classification, and How It Is Scored

What this does NOT apply to in this book

No figure in this chapter comes from an experiment. The matrix above is made up to be worked. The programs in this block are not evaluated for accuracy, because no ground truth exists, and [Āyurveda as a Śāstra, and What This Chapter Does Not Claim] says so.

And a submission that reported an accuracy for such a program would be reporting a number it had invented. That is worth repeating at the point where the machinery for reporting accuracy has just been taught.

Quick revision

  • One-against-rest: k classifiers, imbalanced training sets. One-against-one: k(k-1)/2 classifiers, balanced pairs. For k equal to 3 both need three.
  • Confusion matrix: rows are true classes, columns are predicted; the diagonal is correct and the position of an error says which error.
  • Precision is the diagonal cell over its COLUMN total; recall is the diagonal cell over its ROW total.
  • In the worked matrix accuracy is 74 per cent while phlegm's precision is 43 per cent, so a prediction of phlegm is wrong more often than right.
  • Row totals and column totals must each add to the number of cases, which is a free check.
  • Macro, weighted and micro averages differ, and micro equals accuracy in this setting. Say which you used.
  • Quote the majority-class baseline beside any accuracy figure.

Test yourself

1. Compute precision and recall for a class from a confusion matrix.

Precision is the diagonal cell for that class divided by the total of its column, the cases predicted to be it. Recall is the same cell divided by the total of its row, the cases that truly were it.

2. In the worked matrix, why is 74 per cent accuracy a misleading summary?

Because it hides that only 43 per cent of the cases predicted to be phlegm actually were, so a prediction of phlegm is wrong more often than right, and because the majority-class baseline is already 50 per cent.

3. Distinguish macro from weighted averaging and say when each is right.

Macro averages the per-class figures with every class counting equally; weighted averages them by how common each class is. Macro is right when rare classes matter as much as common ones; weighted when the overall case load is what matters.

4. Why does this book report no accuracy figure for its own Āyurvedic classifier?

Because no labelled data with a ground truth exists for the scheme, so there is nothing to measure against. A reported accuracy would be a number the submission had invented.

Contents This chapter on its own page

munotes.in302

Chapter Eighty-Five

Ayurvedic Classification as a Rule-Based Expert System

Syllabus topic Module 2, "Ayurvedic Classification as Rule-Based Expert System", "Rule-based systems", "Expert systems", "Algorithm Specification (Pseudo-code)", "minimum 10 test cases"

In one line

A rule base read straight off Charaka's attribute lists, a scoring rule, and an explanation that names the text.

In the wording you can write in an examination: a rule-based expert system over the tridoṣa scheme represents each doṣa's attribute list as its rule set, classifies a description by counting matched attributes, reports the score for every class rather than only the maximum, and accompanies each answer with the attributes that produced it and the counter-measure the treatment rule prescribes.

Problem statement, in MU's own form

IKS concept as CS concept: Āyurvedic classification as a rule-based expert system.

Statement. Represent the attribute lists of Charaka Sūtrasthāna I.58 to I.60 as a rule base. Classify a description, given as a set of attributes, by the number of attributes it shares with each doṣa. Report all three scores and the matched attributes, so that the answer carries its own explanation, and derive the counter-measure from Sūtrasthāna I.61's rule of the adverse attribute. Exclude the disputed reading and say so in the source.

Conceptual mapping table

Classical elementComputer science element
a guṇa, a named qualitya binary attribute
the list of seven per doṣathat class's rule set
the union of the liststhe shared feature space
"cold" in bile, printed and doubteda declared exclusion, held in its own constant
classifying by shared qualitiesa weighted vote with equal weights
a case showing qualities of two doṣaa tie, reported rather than broken
the rule of the adverse attributea lookup from each matched attribute to its opposite
place, measure and timethree parameters the program does not model, and says so

Algorithm specification, in pseudo-code

ALGORITHM Classify(described)

INPUT a set of attribute names

OUTPUT a score per dosa, the attributes matched, an answer or a tie, and a reason

for each dosa do

have <- that dosa's attributes, LESS any disputed reading for it

matched <- described intersected with have

score <- the size of matched

end for

best <- the highest score

if best is zero then return "no attribute in the description appears in any list"

winners <- every dosa whose score is best

if there is more than one winner then return a TIE naming them

return the single winner, its matched attributes, and the opposites of those attributes

ALGORITHM Discriminating()

return every attribute that some dosa has and some dosa lacks

Two design points to defend in a viva. A tie is returned as a tie and not broken, because the scheme itself treats a combination as an answer, as [Decision Principles in Diagnosis] sets out. And the disputed reading is held in its own constant rather than deleted from the table, so that the program records what the text prints as well as what it relies on.

munotes.in303

Ayurvedic Classification as a Rule-Based Expert System

Working code

#!/usr/bin/env python3
"""Ayurvedic classification as a rule-based expert system. MU's topic 4.

 THIS IS NOT MEDICAL SOFTWARE AND IT IS NOT MEDICAL ADVICE. It is a
knowledge-representation exercise on a classical text, which is what MU's
syllabus asks for. Nothing here says a classical rule is clinically correct.

IKS concept as CS concept: the tridosa attribute lists of Charaka Samhita,
Sutrasthana I.58 to I.60, as a feature space; the classification of a
description into wind, bile or phlegm as multi-class classification; and
Charaka's own rule of the adverse attribute as an IF-THEN rule with an
explanation.

 THE TABLE IS QUOTED, NOT INVENTED. It is read off Avinash Chandra
Kaviratna's translation:
  I.58  "Wind, which may be dry, cold, light, subtile, unstable, clear, keen,
         is cured by objects which have adverse attributes."
  I.59  "Bile, which may be cold, hot, keen, soft, sour, liquid, and bitter,
         is speedily cured by objects having adverse attributes."
  I.60  "Heavy, cold, mild, watery, sweet, stable, and slimy, these attributes
         of phlegm are cured by objects having adverse attributes."

 I.59 prints BOTH "cold" and "hot" for the same dosa, which cannot both be
intended. The table below records the line as printed and DISPUTED holds the
attribute the book refuses to rely on. FINDINGS section 3.
"""
import math

ATTRIBUTES = {
    'wind':   ['dry', 'cold', 'light', 'subtile', 'unstable', 'clear', 'keen'],
    'bile':   ['cold', 'hot', 'keen', 'soft', 'sour', 'liquid', 'bitter'],
    'phlegm': ['heavy', 'cold', 'mild', 'watery', 'sweet', 'stable', 'slimy'],
}
DISPUTED = {('bile', 'cold')}

OPPOSITE = {
    'dry': 'moist', 'cold': 'hot', 'hot': 'cold', 'light': 'heavy',
    'heavy': 'light', 'subtile': 'gross', 'unstable': 'stable',
    'stable': 'unstable', 'clear': 'slimy', 'slimy': 'clear',
    'keen': 'mild', 'mild': 'keen', 'soft': 'hard', 'sour': 'sweet',
    'sweet': 'sour', 'liquid': 'solid', 'bitter': 'sweet',
    'moist': 'dry', 'watery': 'dry', 'gross': 'subtile',
}


def universe():
    """Every attribute any dosa is given, in first-appearance order."""
    seen = []
    for dosa in ('wind', 'bile', 'phlegm'):
        for a in ATTRIBUTES[dosa]:
            if a not in seen:
                seen.append(a)
    return seen


def vector(dosa, drop_disputed=True):
    """The dosa as a 0/1 vector over the shared attribute space."""
    have = set(ATTRIBUTES[dosa])
    if drop_disputed:
        have -= {a for d, a in DISPUTED if d == dosa}
    return [1 if a in have else 0 for a in universe()]


def discriminating(drop_disputed=True):
    """Attributes that do NOT appear against every dosa.

    An attribute all three share cannot tell them apart, whatever else is true
    of it. This is feature selection done by hand before any entropy is
    computed.
    """
    out = []
    for i, a in enumerate(universe()):
        col = [vector(d, drop_disputed)[i] for d in ('wind', 'bile', 'phlegm')]
        if 0 < sum(col) < 3:
            out.append(a)
    return out


# ----------------------------------------------------------- the rule base
def classify(described):
    """Score a set of described attributes against each dosa.

     A count, not a diagnosis. The explanation is part of the answer,
    because MU's Module II asks for explainable inference and an expert
    system that cannot say why has failed at its job.
    """
    described = set(described)
    scores, why = {}, {}
    for dosa in ('wind', 'bile', 'phlegm'):
        have = set(ATTRIBUTES[dosa]) - {a for d, a in DISPUTED if d == dosa}
        matched = sorted(described & have)
        scores[dosa] = len(matched)
        why[dosa] = matched
    best = max(scores.values())
    winners = sorted(d for d, s in scores.items() if s == best)
    return {
        'scores': scores,
        'matched': why,
        'answer': winners[0] if best and len(winners) == 1 else None,
        'tied': winners if best and len(winners) > 1 else [],
        'explanation': _explain(scores, why, winners, best),
    }


def _explain(scores, why, winners, best):
    if not best:
        return 'No attribute in the description appears in any of the three lists.'
    if len(winners) > 1:
        return ('The description matches %s equally, on %d attribute(s) each, so it '
                'does not decide between them.' % (' and '.join(winners), best))
    d = winners[0]
    return ('%s, because the description gives %s, and Sutrasthana I names those '
            'among the attributes of %s. The counter-measure the text states is an '
            'object of the adverse attribute: %s.'
            % (d.capitalize(), ', '.join(why[d]), d,
               ', '.join(OPPOSITE.get(a, 'the opposite of ' + a) for a in why[d])))


# --------------------------------------------------- the decision tree
def entropy(labels):
    n = len(labels)
    if n == 0:
        return 0.0
    out = 0.0
    for lab in set(labels):
        p = labels.count(lab) / n
        out -= p * math.log2(p)
    return out


def information_gain(rows, attr):
    """rows: [(label, set-of-attributes)]. Split on has/has not."""
    labels = [r[0] for r in rows]
    before = entropy(labels)
    yes = [r[0] for r in rows if attr in r[1]]
    no = [r[0] for r in rows if attr not in r[1]]
    n = len(rows)
    after = (len(yes) / n) * entropy(yes) + (len(no) / n) * entropy(no)
    return before - after


def training_rows(drop_disputed=True):
    rows = []
    for dosa in ('wind', 'bile', 'phlegm'):
        have = set(ATTRIBUTES[dosa])
        if drop_disputed:
            have -= {a for d, a in DISPUTED if d == dosa}
        rows.append((dosa, have))
    return rows


def build_tree(rows, attrs):
    labels = [r[0] for r in rows]
    if len(set(labels)) <= 1:
        return labels[0] if labels else None
    usable = [(information_gain(rows, a), a) for a in attrs]
    usable = [(g, a) for g, a in usable if g > 0]
    if not usable:
        return sorted(set(labels))
    gain, attr = max(usable, key=lambda t: (t[0], -attrs.index(t[1])))
    rest = [a for a in attrs if a != attr]
    return {
        'attribute': attr,
        'gain': round(gain, 4),
        'yes': build_tree([r for r in rows if attr in r[1]], rest),
        'no': build_tree([r for r in rows if attr not in r[1]], rest),
    }


def render_tree(node, indent=0, label=''):
    pad = ' ' * indent
    if not isinstance(node, dict):
        return '%s%s-> %s' % (pad, label, node)
    out = ['%s%sis it %s?  (information gain %.4f)'
           % (pad, label, node['attribute'], node['gain'])]
    out.append(render_tree(node['yes'], indent + 2, 'yes: '))
    out.append(render_tree(node['no'], indent + 2, 'no:  '))
    return '\n'.join(out)


def walk(node, described):
    trail = []
    while isinstance(node, dict):
        a = node['attribute']
        took = a in described
        trail.append('%s %s' % (a, 'yes' if took else 'no'))
        node = node['yes'] if took else node['no']
    return node, trail


TESTS = [
    ('three attributes of wind',        ['dry', 'unstable', 'subtile'],      'wind'),
    ('three attributes of bile',        ['hot', 'sour', 'bitter'],           'bile'),
    ('three attributes of phlegm',      ['heavy', 'sweet', 'slimy'],         'phlegm'),
    ('one attribute only, unique',      ['watery'],                          'phlegm'),
    ('the shared attribute alone',      ['cold'],                            None),
    ('nothing in any list',             ['blue', 'loud'],                    None),
    ('an empty description',            [],                                  None),
    ('keen, which two dosas share',     ['keen'],                            None),
    ('keen with one wind attribute',    ['keen', 'dry'],                     'wind'),
    ('keen with one bile attribute',    ['keen', 'sour'],                    'bile'),
    ('a mixed description, wind leads', ['dry', 'light', 'clear', 'sour'],    'wind'),
    #  Expected None, and the reason is the finding: with the disputed reading
    # dropped, "cold" belongs to wind and phlegm and "hot" to bile, so this
    # description matches all three once and settles nothing. Writing 'bile'
    # here was the first guess and the test caught it.
    ('cold and hot together',           ['cold', 'hot'],                     None),
]


def run_tests(verbose=False):
    passed = 0
    for name, described, expect in TESTS:
        got = classify(described)['answer']
        ok = got == expect
        passed += ok
        if verbose:
            print('%-46s %s  expected %-7s got %s'
                  % (name, 'pass' if ok else 'FAIL', expect, got))
        assert ok, (name, got, expect)
    return passed


def prove():
    u = universe()
    assert len(u) == len(set(u))
    # every dosa gets seven attributes as printed
    for d in ATTRIBUTES:
        assert len(ATTRIBUTES[d]) == 7, d
    # "cold" is printed against all three, so it discriminates nothing
    assert all('cold' in ATTRIBUTES[d] for d in ATTRIBUTES)
    assert 'cold' not in discriminating(drop_disputed=False)
    # dropping the disputed reading makes cold a phlegm-and-wind attribute
    assert 'cold' in discriminating(drop_disputed=True)
    assert abs(information_gain(training_rows(), 'cold')) > 0
    assert abs(information_gain(training_rows(drop_disputed=False), 'cold')) < 1e-12
    # and bile never claims the disputed attribute
    assert classify(['cold'])['matched']['bile'] == []
    # the tree separates all three
    tree = build_tree(training_rows(), discriminating())
    for d in ('wind', 'bile', 'phlegm'):
        leaf, _ = walk(tree, set(ATTRIBUTES[d]) - {a for x, a in DISPUTED if x == d})
        assert leaf == d, (d, leaf)
    run_tests()
    return True


if __name__ == '__main__':
    prove()
    print('### the shared attribute space, %d attributes' % len(universe()))
    print('   ' + ', '.join(universe()))
    print()
    print('### the three vectors')
    print('%-8s %s' % ('', ' '.join('%-9s' % a for a in universe())))
    for d in ('wind', 'bile', 'phlegm'):
        print('%-8s %s' % (d, ' '.join('%-9d' % v for v in vector(d))))
    print()
    print('### information gain of every attribute, disputed reading dropped')
    rows = training_rows()
    for a in universe():
        print('   %-9s %.4f' % (a, information_gain(rows, a)))
    print()
    print('### the tree')
    print(render_tree(build_tree(rows, discriminating())))
    print()
    print('### the ten test cases')
    n = run_tests(verbose=True)
    print('\n%d of %d pass' % (n, len(TESTS)))
    print()
    print('### one answer with its explanation')
    r = classify(['dry', 'unstable', 'subtile'])
    print('   ' + r['explanation'])
munotes.in304

Ayurvedic Classification as a Rule-Based Expert System

### the shared attribute space, 18 attributes
   dry, cold, light, subtile, unstable, clear, keen, hot, soft, sour, liquid, bitter, heavy, mild, watery, sweet, stable, slimy

### the three vectors
         dry       cold      light     subtile   unstable  clear     keen      hot       soft      sour      liquid    bitter    heavy     mild      watery    sweet     stable    slimy
wind     1         1         1         1         1         1         1         0         0         0         0         0         0         0         0         0         0         0
bile     0         0         0         0         0         0         1         1         1         1         1         1         0         0         0         0         0         0
phlegm   0         1         0         0         0         0         0         0         0         0         0         0         1         1         1         1         1         1

### information gain of every attribute, disputed reading dropped
   dry       0.9183
   cold      0.9183
   light     0.9183
   subtile   0.9183
   unstable  0.9183
   clear     0.9183
   keen      0.9183
   hot       0.9183
   soft      0.9183
   sour      0.9183
   liquid    0.9183
   bitter    0.9183
   heavy     0.9183
   mild      0.9183
   watery    0.9183
   sweet     0.9183
   stable    0.9183
   slimy     0.9183

### the tree
is it dry?  (information gain 0.9183)
  yes: -> wind
  no:  is it cold?  (information gain 1.0000)
    yes: -> phlegm
    no:  -> bile

### the ten test cases
three attributes of wind                       pass  expected wind    got wind
three attributes of bile                       pass  expected bile    got bile
three attributes of phlegm                     pass  expected phlegm  got phlegm
one attribute only, unique                     pass  expected phlegm  got phlegm
the shared attribute alone                     pass  expected None    got None
nothing in any list                            pass  expected None    got None
an empty description                           pass  expected None    got None
keen, which two dosas share                    pass  expected None    got None
keen with one wind attribute                   pass  expected wind    got wind
keen with one bile attribute                   pass  expected bile    got bile
a mixed description, wind leads                pass  expected wind    got wind
cold and hot together                          pass  expected None    got None

12 of 12 pass

### one answer with its explanation
   Wind, because the description gives dry, subtile, unstable, and Sutrasthana I names those among the attributes of wind. The counter-measure the text states is an object of the adverse attribute: moist, gross, stable.
munotes.in305

Ayurvedic Classification as a Rule-Based Expert System

Reading the output

Eighteen attributes and three vectors, which is [Doṣa as a Feature Vector]'s table produced by the program that uses it.

munotes.in306

Ayurvedic Classification as a Rule-Based Expert System

Every attribute has the same information gain, 0.9183, which is [A Decision Tree Built From the Tridoṣa Attributes]'s finding, printed here from the same code that builds the tree.

munotes.in307

Ayurvedic Classification as a Rule-Based Expert System

The tree separates all three in two questions.

Twelve test cases pass, and five of them return no answer: the shared attribute alone, nothing in any list, an empty description, "keen" which two doṣa share, and "cold and hot" together. Five refusals out of twelve, which is the proportion a classifier over a small closed vocabulary should have.

And the last block is the explanation. It names the attributes that produced the answer, says which section of the text names them, and gives the counter-measure as the opposites. That is what makes this an expert system rather than a classifier: the answer carries its reason, as [Expert Systems, and MYCIN as the Comparison] requires.

The test case that was wrong first

"Cold and hot together" was expected to return bile, on the reasoning that "hot" is bile's and "cold" is disputed for bile so only "hot" would count.

The program returned nothing, and the program was right. With the disputed reading dropped, "cold" belongs to wind and to phlegm. So the description matches wind once, bile once and phlegm once: a three-way tie, and nothing is settled.

The test was corrected, not the program. That is recorded in the source and it is the sort of thing a limitations section should say: a test written from an expectation rather than from the specification is a test of the expectation.

Complexity and limitations

This is MU's heading seven and it is where a submission earns its marks.

Time. Three set intersections over a description of at most eighteen attributes, so the work is constant for practical purposes. The tree induction is proportional to the number of attributes times the number of rows, which is fifty-four.

Space. The table, the disputed set, and the opposite lookup.

Limitation: it reports no accuracy, and cannot. There is no labelled data with a ground truth for this scheme. An accuracy figure would be invented.

Limitation: equal weights. Nothing in Charaka weights the attributes, so every one counts one. [Multi-Attribute Classification] explains why any other weighting would have to be invented.

Limitation: it does not model place, measure and time. Sūtrasthāna I.61 makes the counter-measure depend on all three, and the program returns the opposites without them. That is a real gap between the text and the implementation, and it is declared rather than passed over.

Limitation: binary attributes. The text speaks of qualities being increased and diminished; the program has present and absent. [Doṣa as a Feature Vector] states the simplification.

Limitation: it does not distinguish prakṛti from a current state. The scheme distinguishes a constitution from a present condition, and the program works on a single description with no notion of which it is.

munotes.in308

Ayurvedic Classification as a Rule-Based Expert System

Limitation: the vocabulary problem is unsolved. The program takes attribute names as input. Getting from what a person says to those names is the hard part, and [Symptom to Feature Mapping] describes it rather than solving it.

And the one that matters most. It classifies into the scheme's own categories. Whether those categories correspond to anything is a question this program does not ask and cannot answer.

Quick revision

  • The rule base is the three attribute lists, less the disputed reading, which is held in its own constant.
  • Classification is a count of shared attributes, with equal weights, and all three scores are reported.
  • A tie is returned as a tie, because the scheme treats a combination as an answer.
  • The counter-measure is the opposites of the matched attributes, from Sūtrasthāna I.61.
  • Twelve test cases, five of which return no answer, which is the right proportion for a small closed vocabulary.
  • One test was written from an expectation and was wrong; the test was corrected, not the program.
  • Limitations: no accuracy and none possible, equal weights, no place-measure-time, binary attributes, no prakṛti distinction, and the vocabulary problem unsolved.

Test yourself

1. Why does the program return a tie rather than breaking it?

Because the scheme itself treats a case showing the qualities of two doṣa as a case in which two predominate, which is one of its own labels. Breaking the tie would discard information the scheme regards as the answer.

2. What does the program do with the disputed attribute, and why is that better than deleting it?

It holds it in a separate constant and excludes it from bile's list when classifying. That records both what the text prints and what the program relies on, so a reader can see the decision instead of finding a silently shortened list.

3. Give three limitations that belong in this implementation's own limitations section.

It reports no accuracy and none is possible, because no labelled data with a ground truth exists. It weights every attribute equally, because Charaka assigns no weights. And it does not model place, measure and time, which Sūtrasthāna I.61 makes the counter-measure depend on.

4. A test expected bile for a description of cold and hot, and the program returned nothing. What was wrong?

The test. With the disputed reading excluded, cold belongs to wind and phlegm and hot to bile, so the description matches each class once and settles nothing. The expectation was written from a half-remembered reading of the table rather than from the specification.

Contents This chapter on its own page

munotes.in309

Chapter Eighty-Six

The Arthaśāstra, and Its Intelligence Apparatus

Syllabus topic Module 2, "Arthaśāstra Cryptography as Secure Communication Model", "Study of intelligence and secure communication practices"

In one line

The Arthaśāstra is a treatise on running a state, and one of the things it is most detailed about is intelligence.

In the wording you can write in an examination: the Arthaśāstra, attributed to Kautilya, is a Sanskrit treatise on statecraft in fifteen books. Its first book, "Concerning Discipline", establishes an apparatus of espionage with stationary and wandering agents under named institutes, and its provisions on the transmission of intelligence are the passages this syllabus studies as a secure communication model.

The text

Fifteen books, each divided into chapters, and the chapters numbered both within their book and continuously from the beginning. Shamasastry's colophons give both: the chapter on royal writs is "Chapter X ... in Book II" and also "the thirty-first chapter from the beginning".

A citation therefore names the book and the chapter, and this book gives both, with the chapter's own printed title.

The translation quoted throughout is R. Shamasastry's, published by the Bangalore Government Press in 1915. Shamasastry died in 1941, so the translation is out of copyright and may be quoted with attribution.

On the date and the authorship, the scholarly literature is large and divided. Nothing in this paper depends on either, so this book reports neither.

The intelligence apparatus

Three books bear on it, and the structure is worth having because the secret-writing passages presuppose all of it.

Book I, Chapter XI, "The Institution of Spies." Establishes the classes of agent: the recruits are drawn from named walks of life and are placed where they will hear things.

Book I, Chapter XII, "Creation of Wandering Spies." Establishes the mobile agents, and it is this chapter that carries two of the four passages on secret communication.

Book I, Chapter XVI, "The Mission of Envoys." Establishes the envoy's duties abroad, including what to do when open conversation is impossible, and it carries the third passage.

Book XIII, Chapter I, "Sowing the Seeds of Dissension." Carries the fourth, on a sealed letter arriving by pigeon.

The two kinds of agent

The scheme divides agents by whether they move, and the division determines the communication problem.

Stationary agentsWandering agents
Whereplaced in a fixed post: a household, a guild, a templemoving, on errands, among the population or abroad
What they seeone place over a long timemany places briefly
How they reportthrough the chain their post belongs toon return, or through a channel arranged in advance
The communication problemroutine reporting that must not be noticedoccasional reporting from an unknown location

And the second column is where cryptography would live. An agent abroad with information and no safe channel is the situation that makes secret writing necessary, and it is the situation Book I, Chapter XVI describes.

munotes.in310

The Arthaśāstra, and Its Intelligence Apparatus

The institutes, and the point about them

Book I, Chapter XII names the officers of the institutes of espionage and, in Shamasastry's rendering, says they:

shall by making use of signs or writing (saṃjñālipibhih) set their own spies in motion (to ascertain the validity of the information).

Worked. Three things are in that sentence and all three matter for [Secret Communication: What the Text Actually Says].

There is an institution, not just individuals. An institute with officers and subordinates has a structure, and a structure needs a protocol.

The officers communicate downward using signs or writing. So the covert channel runs from the centre outward as well as inward.

And the purpose stated is to CHECK the information, not merely to collect it. That is corroboration, and it is the same principle as Charaka's rule in [Parīkṣā: How an Examination Is Structured] that one means of knowledge is not enough.

The channels the text presupposes

Read the four passages together and a picture of the communication system emerges, even though no passage describes it whole.

A courier who does not know what they carry. The passage in Book I, Chapter XII has prostitutes, artisans and court-bards carrying information "under the pretext of taking in musical instruments" or by cipher-writing. The carrier is a channel and need not be trusted with the content.

A prearranged sign system. "By means of signs" appears in three of the four passages. A sign system must be agreed in advance, which means a shared secret exists before the message does.

A physical token with integrity. The sealed letter of Book XIII, Chapter I. A seal is not confidentiality; it is evidence of tampering.

And a written message whose content is concealed. Gūḍhalekhya, cipher-writing. The one the syllabus is really about, and the one the text says least about.

What this chapter sets up and what the next one settles

Set up here. That there was an apparatus, that it had a structure with officers and channels, that corroboration was its stated purpose, and that agents abroad had a communication problem the text is aware of.

Settled next. Exactly what the four passages say about secret communication, and the honest limit: the treatise names cipher-writing and does not give a cipher. [Secret Communication: What the Text Actually Says] quotes all four and says so.

What this book does NOT claim about the Arthaśāstra

Not a date, and not an author. Both are disputed and neither is needed.

Not that its intelligence apparatus was ever implemented as described. A treatise prescribes; whether any state did this is a historical question with its own evidence.

munotes.in311

The Arthaśāstra, and Its Intelligence Apparatus

Not that its methods were effective. [Breaking a Classical Cipher] measures what a classical system would have withstood, and the answer is: not much, by modern standards, which is not a criticism of a text written two millennia before the analysis existed.

And not that anything in it is a cipher algorithm. That is the next chapter's finding and it is stated in advance so that nothing here reads as building toward a claim the evidence does not support.

Quick revision

  • The Arthaśāstra, attributed to Kautilya, is in fifteen books, and Shamasastry's colophons number each chapter both within its book and from the beginning.
  • The translation quoted is Shamasastry's, Bangalore Government Press, 1915, out of copyright.
  • The four passages on secret communication are in Book I, Chapter XII, twice; Book I, Chapter XVI; and Book XIII, Chapter I.
  • Agents are stationary or wandering, and the wandering ones create the communication problem cryptography answers.
  • The officers of the institutes communicate by signs or writing, and the stated purpose is to CHECK information, which is corroboration.
  • Four channels are presupposed: an untrusted courier, a prearranged sign system, a sealed token for integrity, and concealed writing for confidentiality.
  • Not claimed: a date, an author, that the apparatus was implemented, or that any of it is a cipher algorithm.

Test yourself

1. Where in the Arthaśāstra are the passages on secret communication, and which translation does this book quote?

Two in Book I, Chapter XII, "Creation of Wandering Spies"; one in Book I, Chapter XVI, "The Mission of Envoys"; and one in Book XIII, Chapter I, "Sowing the Seeds of Dissension". The translation is R. Shamasastry's of 1915.

2. Why does the division into stationary and wandering agents matter to this syllabus?

Because the communication problem differs. A stationary agent reports routinely through an established chain; a wandering agent must report occasionally from an unknown place, which is the situation that makes concealed writing necessary.

3. What is the stated purpose of the officers' use of signs or writing, and what modern principle does it match?

To ascertain the validity of information already received. That is corroboration, the same principle as Charaka's rule that no single means of knowledge suffices.

4. Name the four channels the passages presuppose and what each protects.

An untrusted courier, who carries without knowing. A prearranged sign system, which requires a secret shared before the message. A seal, which gives evidence of tampering rather than confidentiality. And concealed writing, which hides the content.

Contents This chapter on its own page

munotes.in312

Chapter Eighty-Seven

Secret Communication: What the Text Actually Says

Syllabus topic Module 2, "Secret communication techniques"

In one line

Kautilya names secret writing four times and never says how it is done.

In the wording you can write in an examination: the Arthaśāstra refers to concealed communication in four places, using three terms: saṃjñā, signs; saṃjñālipi, writing by signs; and gūḍhalekhya, secret or cipher writing. It also records a sealed letter carried by a domestic pigeon. In none of the four does it describe a method of enciphering, so the text establishes a requirement for secret communication and supplies no algorithm.

The four passages

One: Book I, Chapter XII, the officers of the institutes

The immediate officers of the institutes of espionage (saṃsthānām antevāsinaḥ) shall by making use of signs or writing (saṃjñālipibhih) set their own spies in motion (to ascertain the validity of the information).

What it establishes. That the institution communicates with its own agents covertly, and that the medium is signs or a writing built on signs.

Two: Book I, Chapter XII, the mendicant woman at the entrance

If a mendicant woman is stopped at the entrance, the line of door-keepers, spies under the guise of father and mother (mātāpitṛ vyañjanāḥ), women artisans, court-bards, or prostitutes shall, under the pretext of taking in musical instruments, or through cipher-writing (gūḍhalekhya), or by means of signs, convey the information to its destined place (cāraṃ nirhareyuḥ).

What it establishes. Three alternative channels, named together: a physical pretext, cipher-writing, and signs. This is the passage in which gūḍhalekhya appears as a named technique, and it appears as one option among three.

Three: Book I, Chapter XVI, the envoy abroad

If there is no possibility of carrying on any such conversation (conversation with the people regarding their loyalty), he may try to gather such information by observing the talk of beggars, intoxicated and insane persons or of persons babbling in sleep, or by observing the signs made in places of pilgrimage and temples or by deciphering paintings and secret writings (citra-gūḍha-lekhya-saṃjñābhiḥ).

What it establishes. That the envoy is expected to READ secret writings, not only to produce them. That is the other side of the channel and it presupposes that the other party uses them too.

Worked, and the compound is worth reading. Citra, painting; gūḍha, hidden; lekhya, writing; saṃjñā, sign. Four things in one term, and the translation renders them as "paintings and secret writings".

Four: Book XIII, Chapter I, the pigeon

... pretensions to the knowledge of foreign affairs by means of his power to read omens and signs invisible to others when information about foreign affairs is just received through a domestic pigeon which has brought a sealed letter.

What it establishes. A physical carrier, and a seal.

And notice the context. The passage is about the king pretending to omniscience: he receives the letter secretly and presents the knowledge as supernatural. The secrecy being protected is the CHANNEL, not the message. That is a different security property and [Concealment and Coded Messaging] is about the difference.

munotes.in313

Secret Communication: What the Text Actually Says

What the four passages give, and what they do not

Established by the textNot in the text
that secret writing was usedyes, named three times
that it had a nameyes, gūḍhalekhya
that agents both wrote and read ityes, passages two and three
that signs were prearrangedimplied: a sign system must be agreedthe agreement procedure
that messages were sealedyes, passage four
how a message was encipherednothing at all
what the key wasnothing at all
how a key was distributednothing at all

Read the bottom three rows. They are the three things a cipher system consists of, and the text supplies none of them.

The honest statement, which belongs in any answer

The Arthaśāstra establishes a requirement and does not supply an algorithm.

Why that is still worth studying. A requirement is not nothing. The text identifies that messages must travel through hostile territory, that the carrier may be searched, that the carrier need not be trusted, that the recipient must be able to read what the sender wrote, and that the existence of the message may itself need hiding. Those five are the threat model, and stating a threat model is the first thing a secure design does.

And what it means for MU's implementation topic 5. Her topic is "Arthaśāstra-Inspired Cryptography as Symmetric Cipher System", and the word is INSPIRED. She has not asked you to implement Kautilya's cipher, because there is not one. She has asked for a symmetric cipher built to the requirement his text states, which is exactly what [Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System] builds.

A submission that claims to implement Kautilya's own cipher has made a claim its source does not support. Saying so in the limitations section is worth more than any amount of enthusiasm.

What can be said about the likely method

Nothing definite, and the honest treatment is to say what constrains it rather than to guess.

It had to be usable by the agents described. Prostitutes, artisans and court-bards, working in the field, without instruments.

It had to be reversible by the recipient without a conversation. So the method and any key were arranged in advance.

And it had to survive a search, or the writing had to be concealed as well as enciphered. Passage two offers three channels, one of which is a physical pretext, which suggests that concealment and encipherment were alternatives rather than a layered defence.

munotes.in314

Secret Communication: What the Text Actually Says

Those three constraints rule out a great deal and they identify nothing. That is where the evidence stops.

What this chapter does NOT say

Not that Indian antiquity had no cryptography. Other texts are cited in the literature for substitution methods, and this book has not read them and does not report them.

Not that the passages are unimportant. A threat model, a named technique and evidence of two-way use is a real contribution to the history of the subject.

And not that Shamasastry's renderings are the only ones. "Cipher-writing" is his translation of gūḍhalekhya; another translator might render it "secret writing" without the implication of a cipher. The English word does work the Sanskrit may not, and an answer that notices that is a strong one.

Quick revision

  • Four passages: Book I Chapter XII twice, Book I Chapter XVI, Book XIII Chapter I.
  • Three terms: saṃjñā, signs; saṃjñālipi, writing by signs; gūḍhalekhya, secret or cipher writing. Plus a sealed letter by pigeon.
  • Agents both write and read secret writings, which presupposes the other side uses them.
  • The text supplies no method, no key and no key distribution: the three things a cipher system consists of.
  • What it does supply is a threat model: hostile territory, a searchable carrier, an untrusted carrier, a recipient who must read it, and the existence of the message itself needing concealment.
  • MU's topic says INSPIRED, because there is no cipher in the text to implement.
  • "Cipher-writing" is Shamasastry's rendering, and the English may carry more than the Sanskrit does.

Test yourself

1. Name the three Sanskrit terms the passages use and what each denotes.

Saṃjñā, signs; saṃjñālipi, writing by means of signs; and gūḍhalekhya, which Shamasastry renders cipher-writing.

2. What are the three things a cipher system consists of, and how many does the Arthaśāstra supply?

A method of enciphering, a key, and a means of distributing the key. It supplies none of them.

3. State the threat model the passages establish.

Messages must cross hostile territory; the carrier may be stopped and searched; the carrier need not be trusted with the content; the recipient must be able to recover the message without conversing with the sender; and the existence of the message may itself need concealing.

4. Why is MU's implementation topic worded "Arthaśāstra-Inspired" rather than "Arthaśāstra's"?

Because the text names secret writing and gives no cipher, so there is nothing of Kautilya's to implement. The task is to build a symmetric cipher meeting the requirement his text states.

Contents This chapter on its own page

munotes.in315

Chapter Eighty-Eight

The Royal Writs as a Message Format

Syllabus topic Module 2, "Information protection mechanisms", "Study of intelligence and secure communication practices"

In one line

Kautilya's chapter on royal writs is a message format specification with required properties and named defects.

In the wording you can write in an examination: Arthaśāstra Book II, Chapter X prescribes the composition of official written communications. It states the qualifications of the writer, the elements a writ must contain according to the rank of the addressee, six necessary qualities of a writ, and five named faults. It is therefore a specification of a message format together with its validation conditions.

Why this chapter belongs in a section on cryptography

Because a cipher is not a communication system. A system needs a format, a way of checking a message is well formed, a way of establishing who sent it, and a way of detecting alteration. Encipherment addresses only one property, confidentiality.

And this chapter supplies the rest. It is the only place in the treatise where the structure of a message is prescribed, and reading it as a specification is what MU's label "information protection mechanisms" is pointing at.

The provisions

On the importance of the format.

Writs are of great importance to kings inasmuch as treaties and ultimate leading to war depend upon writs.

On who may compose one.

Hence one who is possessed of ministerial qualifications, acquainted with all kinds of customs, smart in composition, good in legible writing, and sharp in reading shall be appointed as a writer (lekhaka).

Such a writer, having attentively listened to the king's order and having well thought out the matter under consideration, shall reduce the order to writing.

On what a writ must contain, by rank of addressee.

As to a writ addressed to a lord (īśvara), it shall contain a polite mention of his country, his possessions, his family and his name, and as to that addressed to a common man (aniśvara), it shall make a polite mention of his country and name.

On fitting the writ to the addressee.

Having paid sufficient attention to the caste, family, social rank, age, learning (śruta), occupation, property, character (śīla), blood-relationship (yaunānubandha) of the addressee, as well as to the place and time (of writing), the writer shall form a writ befitting the position of the person addressed.

The six necessary qualities

Arrangement of subject-matter (arthakrama), relevancy (sambandha), completeness, sweetness, dignity, and lucidity are the necessary qualities of a writ.

And he defines the first:

The act of mentioning facts in the order of their importance is arrangement.

QualityWhat it requiresThe modern counterpart
arrangementfacts in order of importancea defined field order, and the most significant first
relevancynothing extraneousa minimal message; no fields the recipient cannot use
completenessnothing missingrequired fields, all present
sweetnessagreeable expressiona tone appropriate to the recipient
dignityexpression befitting the senderconsistency of register
lucidityunambiguousone reading only, which is the whole of [Formal Specification: Saying Exactly What a System Must Do]
munotes.in316

The Royal Writs as a Message Format

Four of the six are properties a modern message format would state, and the other two are properties of a diplomatic letter that no protocol needs. Separating them is the analysis this chapter asks for.

The five named faults

Clumsiness, contradiction, repetition, bad grammar, and misarrangement are the faults of a writ.

And each is defined.

Clumsiness (akānti). "Black and ugly leaf, and uneven and uncoloured writing cause clumsiness." A fault of the medium rather than the content: an illegible message.

Contradiction (vyāghāta). "Subsequent portion disagreeing with previous portion of a letter, causes contradiction." A message that contradicts itself is invalid, which is an internal consistency condition.

Repetition. "Stating for a second time what has already been said above is repetition." Redundancy, treated as a fault.

Bad grammar (apaśabda). "Wrong use of words in gender, number, time and case." A well-formedness condition on the encoding.

Misarrangement (samplava). "Division of paragraphs (varga) in unsuitable places, omission of necessary division of paragraphs, and violation of any other necessary qualities of a writ." A structural fault: the message parses wrongly.

The five faults as a validator

Read them as checks a recipient could run and the correspondence is exact.

FaultThe checkThe modern name
clumsinesscan the message be read at all?a transmission or encoding error
bad grammaris it well formed in the language?a parse failure
misarrangementis its structure correct?a schema violation
contradictiondo its parts agree?an internal consistency check
repetitionis anything stated twice?redundancy, which in a message is a waste and in a specification is a source of drift

Worked. Five checks in order, from the medium to the meaning. And the order is the right one: there is no point checking consistency in a message you cannot read.

That is the same ordering argument as [Knowledge Validation and Structured Reasoning]'s three stages, and it is the same reason a compiler parses before it type-checks.

What the chapter does NOT supply

No integrity mechanism. Nothing detects alteration by a third party. The seal in Book XIII does that and this chapter does not mention it.

No authentication. Nothing establishes that the writ is from the king rather than from somebody with a good hand. The writer's appointment is the control, which is organisational rather than technical.

No confidentiality. The writ is not secret. It is official, and the secret channel is a separate apparatus.

So of the four properties a communication system needs, this chapter supplies format and validation, and the other two are elsewhere or absent. Saying which is which is the analysis.

munotes.in317

The Royal Writs as a Message Format

Kautilya's closing line

Having followed all sciences and having fully observed forms of writing in vogue, these rules of writing royal writs have been laid down by Kautilya in the interest of kings.

Two things in it. The rules are said to be drawn from existing practice, not invented, which is what a standards body says of a specification. And the purpose is stated: in the interest of kings.

And the colophon fixes the location: Chapter X, "The Procedure of Forming Royal Writs", in Book II, "The Duties of Government Superintendents", the thirty-first chapter from the beginning.

Quick revision

  • Book II, Chapter X prescribes the composition of official writs and is a message format specification.
  • The writer is appointed on stated qualifications, which is the only authentication control here and it is organisational.
  • Required content varies by the rank of the addressee, and the writ must fit the addressee's circumstances and the place and time.
  • Six necessary qualities: arrangement, relevancy, completeness, sweetness, dignity, lucidity. Four map onto a modern format's requirements.
  • Five faults: clumsiness, contradiction, repetition, bad grammar, misarrangement, each defined, and together they are a validator running from the medium to the meaning.
  • Absent: integrity, authentication and confidentiality. The chapter supplies format and validation only.

Test yourself

1. Name the six necessary qualities of a writ and say which have modern counterparts.

Arrangement, relevancy, completeness, sweetness, dignity and lucidity. Arrangement, relevancy, completeness and lucidity correspond to field order, minimality, required fields and unambiguity; sweetness and dignity are properties of a diplomatic letter with no protocol counterpart.

2. List the five faults with Kautilya's own definition of two of them.

Clumsiness, contradiction, repetition, bad grammar and misarrangement. Contradiction is a subsequent portion disagreeing with a previous portion; bad grammar is wrong use of words in gender, number, time and case.

3. Why is the order of the five checks the right order?

Because each assumes the previous one passed. There is no point testing whether a message's parts agree if the message cannot be read at all, just as a compiler parses before it type-checks.

4. Which of the four properties a communication system needs does this chapter supply, and where are the others?

It supplies the format and its validation. Integrity appears elsewhere, as the seal of Book XIII; confidentiality is the separate apparatus of secret writing; and authentication is addressed only organisationally, by the appointment of a qualified writer.

Contents This chapter on its own page

munotes.in318

Chapter Eighty-Nine

Substitution Systems

Syllabus topic Module 2, "Substitution and transposition systems"

In one line

A substitution cipher replaces each letter by another, always the same way, and it keeps the letters and changes their identity.

In the wording you can write in an examination: a substitution cipher enciphers by replacing each symbol of the plaintext with a symbol determined by the key. In a monoalphabetic substitution the replacement is fixed, so each plaintext letter always becomes the same ciphertext letter. The Caesar cipher is the special case in which the substitution is a rotation of the alphabet by a fixed amount.

The two primitives of classical cryptography

Everything before the twentieth century is built from two operations, and this chapter and the next take one each.

Substitution changes WHAT the letters are and leaves their positions alone.

Transposition changes WHERE the letters are and leaves their identities alone.

And the two are composed, because either alone is weak. [Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System] composes them.

The Caesar cipher

The rule. Shift every letter forward through the alphabet by a fixed amount, wrapping round from Z to A.

The key. One number from 1 to 25.

Worked by hand. Shift 3. A becomes D, B becomes E, and so on to W becoming Z, X becoming A, Y becoming B, Z becoming C.

plain A B C D E F G H I J K L M N O P Q R S T U V W X Y Z

cipher D E F G H I J K L M N O P Q R S T U V W X Y Z A B C

And the key space is 26, which includes the useless shift of 0. Twenty-five useful keys is not a key space; it is a list. Trying all of them takes seconds by hand and no time at all by machine, which [Breaking a Classical Cipher] demonstrates.

The keyword alphabet

The problem with Caesar is that the cipher alphabet is determined by one number, so the key is tiny.

The fix. Let the cipher alphabet be an arbitrary permutation of the plain alphabet. Now the key is the whole permutation.

And a keyword makes a permutation memorable. Write the keyword's distinct letters first, then the remaining letters of the alphabet in order.

Worked by hand, for the keyword KAUTILYA. Its distinct letters, in order of first appearance, are K, A, U, T, I, L, Y. The second A is dropped. Then the rest of the alphabet in order, skipping those seven: B, C, D, E, F, G, H, J, M, N, O, P, Q, R, S, V, W, X, Z.

plain A B C D E F G H I J K L M N O P Q R S T U V W X Y Z

cipher K A U T I L Y B C D E F G H J M N O P Q R S V W X Z

munotes.in319

Substitution Systems

The two, run

"""Substitution, worked: Caesar, keyword and the size of the key space."""
import string, math
from collections import Counter

A = string.ascii_uppercase

def caesar(text, shift, decrypt=False):
    if decrypt:
        shift = -shift
    return "".join(A[(A.index(c) + shift) % 26] if c in A else c for c in text.upper())

def keyword_alphabet(key):
    seen = []
    for c in key.upper():
        if c in A and c not in seen:
            seen.append(c)
    for c in A:
        if c not in seen:
            seen.append(c)
    return "".join(seen)

def substitute(text, key, decrypt=False):
    cipher = keyword_alphabet(key)
    src, dst = (cipher, A) if decrypt else (A, cipher)
    table = dict(zip(src, dst))
    return "".join(table.get(c, c) for c in text.upper())

MSG = "SEND THE ARMY AT DAWN"

print("Caesar, every shift from 1 to 5")
for s in range(1, 6):
    print("  shift %d  %s" % (s, caesar(MSG, s)))

print()
print("keyword alphabets")
for key in ("KAUTILYA", "ARTHASASTRA", "A"):
    ka = keyword_alphabet(key)
    print("  key %-12s plain  %s" % (key, A))
    print("  %-16s cipher %s" % ("", ka))

print()
print("monoalphabetic substitution under the key KAUTILYA")
ct = substitute(MSG, "KAUTILYA")
print("  plaintext  %s" % MSG)
print("  ciphertext %s" % ct)
print("  decrypted  %s" % substitute(ct, "KAUTILYA", decrypt=True))

print()
print("key spaces")
print("  Caesar                        %s" % format(26, ","))
print("  keyword, distinct 8-letter    at most %s" % format(26 * 25 * 24 * 23 * 22 * 21 * 20 * 19, ","))
print("  general monoalphabetic        %s" % format(math.factorial(26), ","))

print()
print("letter frequencies of a longer ciphertext")
SAMPLE = ("THE ENEMY IS AT THE GATE AND WE MUST SEND WORD TO THE CAPITAL AT ONCE "
          "FOR THE GATE WILL NOT HOLD THROUGH THE NIGHT")
enc = substitute(SAMPLE, "KAUTILYA")
counts = Counter(c for c in enc if c in A)
total = sum(counts.values())
print("  ciphertext letters: %d" % total)
for ch, n in counts.most_common(5):
    print("    %s %3d  %5.2f%%" % (ch, n, 100 * n / total))
print("  the commonest plaintext letter of English is E, and E maps to %s under this key"
      % keyword_alphabet("KAUTILYA")[A.index("E")])
Caesar, every shift from 1 to 5
  shift 1  TFOE UIF BSNZ BU EBXO
  shift 2  UGPF VJG CTOA CV FCYP
  shift 3  VHQG WKH DUPB DW GDZQ
  shift 4  WIRH XLI EVQC EX HEAR
  shift 5  XJSI YMJ FWRD FY IFBS

keyword alphabets
  key KAUTILYA     plain  ABCDEFGHIJKLMNOPQRSTUVWXYZ
                   cipher KAUTILYBCDEFGHJMNOPQRSVWXZ
  key ARTHASASTRA  plain  ABCDEFGHIJKLMNOPQRSTUVWXYZ
                   cipher ARTHSBCDEFGIJKLMNOPQUVWXYZ
  key A            plain  ABCDEFGHIJKLMNOPQRSTUVWXYZ
                   cipher ABCDEFGHIJKLMNOPQRSTUVWXYZ

monoalphabetic substitution under the key KAUTILYA
  plaintext  SEND THE ARMY AT DAWN
  ciphertext PIHT QBI KOGX KQ TKVH
  decrypted  SEND THE ARMY AT DAWN

key spaces
  Caesar                        26
  keyword, distinct 8-letter    at most 62,990,928,000
  general monoalphabetic        403,291,461,126,605,635,584,000,000

letter frequencies of a longer ciphertext
  ciphertext letters: 90
    Q  15  16.67%
    I  12  13.33%
    B   9  10.00%
    K   7   7.78%
    J   7   7.78%
  the commonest plaintext letter of English is E, and E maps to I under this key
munotes.in320

Substitution Systems

Reading the output

The three keyword alphabets show the mechanism's range. KAUTILYA displaces most of the alphabet. ARTHASASTRA has many repeated letters, so only five distinct ones survive and the rest of the alphabet is barely disturbed. And the key "A" produces the identity permutation, which enciphers nothing: a key that is a single letter already in place is not a key at all.

That last case is worth a test in any implementation. A cipher that silently accepts a key doing nothing is a cipher that will one day be used with one.

The key spaces. Caesar has 26. An eight-letter keyword has at most about 63 billion, because the first letter may be any of 26, the second any of the remaining 25, and so on. A general permutation has 26 factorial, which is about 4 times ten to the twenty-sixth.

And the gap between the last two is the point. A keyword alphabet is convenient and it throws away almost all of the key space: 63 billion against 4 times ten to the twenty-sixth is a factor of about ten to the sixteen. Memorability costs keys, and that trade recurs in every practical system.

The frequency table, and the trap

The last block enciphers a longer message and counts the letters.

The commonest ciphertext letter is Q, at 16.67 per cent.

And E maps to I, which is only the second commonest.

So the obvious attack is wrong on this sample. "The commonest ciphertext letter is E" is the standard first move and here it would substitute E for the image of T, because in this particular text T is commoner than E.

Why it still works in general. With enough ciphertext the frequencies converge to English's, and the full profile, not just the top letter, identifies the mapping. [Breaking a Classical Cipher] shows the attack working on a Caesar cipher, where the whole profile is shifted together and matching is easy.

And the lesson for an answer. Frequency analysis is a statistical attack and it needs enough text. On twenty letters it says nothing; on a thousand it is decisive.

Why monoalphabetic substitution is weak anyway

The key space is irrelevant. 4 times ten to the twenty-sixth is far too many to search, and the cipher is still broken easily, because the attacker does not search the key space. They read the structure off the ciphertext.

munotes.in321

Substitution Systems

Three structural leaks.

Letter frequencies survive. Each plaintext letter maps to exactly one ciphertext letter, so the frequency profile is permuted, not destroyed.

Word lengths survive if spaces are kept, which is why ciphertext is conventionally written in fixed-size groups.

And repeated patterns survive. A doubled letter in the plaintext is a doubled letter in the ciphertext, and common short words are recognisable by shape.

The general lesson. A large key space is necessary and not sufficient. A cipher is as strong as its weakest structural leak, not as its key count, and that sentence is worth a mark by itself.

Quick revision

  • Substitution changes what the letters are; transposition changes where they are. Both are needed.
  • Caesar: one fixed shift, key space 26, trivially broken by trying all of them.
  • Keyword alphabet: the keyword's distinct letters, then the rest in order. Memorable, and it wastes almost all the key space.
  • A keyword whose letters are already in place gives the identity permutation and enciphers nothing.
  • Key spaces: 26, about 63 billion for an eight-letter keyword, and 26 factorial for a general permutation.
  • The commonest ciphertext letter need not be the image of E; frequency analysis needs enough text and uses the whole profile.
  • Monoalphabetic substitution leaks frequencies, word lengths and repeated patterns, so its key space does not save it.

Test yourself

1. Encipher SEND with a Caesar shift of 4, and say why the key space is not a defence.

WIRH. The key space is 26, so an attacker tries every shift and reads the one that is English, which takes seconds.

2. Build the keyword alphabet for the key ARTHA and encipher the word GATE.

The distinct letters of ARTHA are A, R, T, H, so the cipher alphabet is A R T H B C D E F G I J K L M N O P Q S U V W X Y Z. G maps to D, A to A, T to Q, E to B, so GATE becomes DAQB.

3. Why does a keyword alphabet throw away key space, and does it matter?

Because only permutations expressible as a keyword followed by the remaining letters in order can be reached, which is a tiny fraction of all permutations. It matters less than it appears, because the cipher is broken by structure rather than by search.

4. State the general lesson about key spaces in one sentence.

A large key space is necessary and not sufficient: a cipher is broken by its weakest structural leak, such as surviving letter frequencies, rather than by exhausting its keys.

Contents This chapter on its own page

munotes.in322

Chapter Ninety

Transposition Systems

Syllabus topic Module 2, "Substitution and transposition systems"

In one line

A transposition cipher keeps every letter and rearranges them, so the ciphertext has exactly the same letters as the plaintext in a different order.

In the wording you can write in an examination: a transposition cipher enciphers by permuting the positions of the symbols of the plaintext according to the key, without altering the symbols themselves. The rail fence writes the text in a zigzag across a number of rows and reads it off row by row; the columnar transposition writes it in rows of fixed width and reads the columns in an order determined by the key.

The rail fence

The rule. Choose a number of rails. Write the text diagonally down and up across them. Read it off one rail at a time.

Worked by hand, three rails, on SENDTHEARMYATDAWN.

rail 0 S...T...R...T...N

rail 1 .E.D.H.A.M.A.D.W.

rail 2 ..N...E...Y...A..

Reading rail 0, then rail 1, then rail 2 gives STRTN EDHAMADW NEYA.

The key is the number of rails, and that is its whole weakness: with a text of any length only a handful of rail counts are worth trying, so the key space is a few dozen at most.

The columnar transposition

The rule. Write the text in rows of width equal to the key's length. Number the columns by the alphabetical order of the key's letters. Read the columns off in that numbered order.

Worked by hand, key ZEBRAS, on the same text.

The key's letters in alphabetical order are A, B, E, R, S, Z, so the column holding A is read first, then B, then E, then R, then S, then Z.

key Z E B R A S

order 6 3 2 4 1 5

row 1 S E N D T H

row 2 E A R M Y A

row 3 T D A W N X

Read column 5 first, which is T Y N; then column 2, N R A; then column 1... no: read the column whose key letter comes first alphabetically. A is in position 5, so T Y N. B is in position 3, so N R A. E is in position 2, so E A D. R is in position 4, so D M W. S is in position 6, so H A X. Z is in position 1, so S E T.

Ciphertext: TYN NRA EAD DMW HAX SET.

The padding problem

Seventeen letters do not fill three rows of six. One cell is left over, and the program fills it with X.

Three consequences, and all three are practical.

The ciphertext is longer than the plaintext. Anyone who sees the length knows something.

The padding must be removable. The recipient of SENDTHEARMYATDAWNX has to know that the final X is not part of the message. If the plaintext might legitimately end in X, they cannot.

munotes.in323

Transposition Systems

And where the padding is added matters. [Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System] records a real bug in this book's own code: padding after the substitution step meant the inverse substitution turned the pad into a different letter, and a message round-tripped as SENDTHEARMYATDAWNYYYYYYY.

Both, run

"""Transposition, worked: rail fence and columnar."""
import string
A = string.ascii_uppercase

def clean(t):
    return "".join(c for c in t.upper() if c in A)

def rail_pattern(n, rails):
    out, r, step = [], 0, 1
    for _ in range(n):
        out.append(r)
        if r == 0:
            step = 1
        elif r == rails - 1:
            step = -1
        r += step
    return out

def rail_fence(text, rails, decrypt=False):
    n = len(text)
    pat = rail_pattern(n, rails)
    order = [i for rr in range(rails) for i in range(n) if pat[i] == rr]
    if not decrypt:
        return "".join(text[i] for i in order)
    out = [""] * n
    for pos, i in enumerate(order):
        out[i] = text[pos]
    return "".join(out)

def columnar(text, key, decrypt=False, pad="X"):
    k = [c for c in key.upper() if c in A]
    cols = len(k)
    order = sorted(range(cols), key=lambda i: (k[i], i))
    if not decrypt:
        body = text + pad * ((-len(text)) % cols)
        rows = [body[i:i + cols] for i in range(0, len(body), cols)]
        return "".join("".join(r[c] for r in rows) for c in order)
    rows_n = len(text) // cols
    chunks, at = {}, 0
    for c in order:
        chunks[c] = text[at:at + rows_n]
        at += rows_n
    return "".join(chunks[c][r] for r in range(rows_n) for c in range(cols))

MSG = clean("SEND THE ARMY AT DAWN")
print("plaintext", MSG, "(%d letters)" % len(MSG))

print()
print("rail fence, three rails: the pattern of rails, then the read-off")
pat = rail_pattern(len(MSG), 3)
for r in range(3):
    print("  rail %d  %s" % (r, "".join(MSG[i] if pat[i] == r else "." for i in range(len(MSG)))))
ct = rail_fence(MSG, 3)
print("  ciphertext %s" % ct)
print("  decrypted  %s" % rail_fence(ct, 3, decrypt=True))

print()
print("rail fence with two to five rails")
for rails in range(2, 6):
    c = rail_fence(MSG, rails)
    print("  %d rails  %-20s round trip %s" % (rails, c, rail_fence(c, rails, decrypt=True) == MSG))

print()
print("columnar transposition under the key ZEBRAS")
KEY = "ZEBRAS"
k = list(KEY)
order = sorted(range(len(k)), key=lambda i: (k[i], i))
body = MSG + "X" * ((-len(MSG)) % len(k))
rows = [body[i:i + len(k)] for i in range(0, len(body), len(k))]
print("  key      %s" % "  ".join(k))
print("  read in  %s" % "  ".join(str(order.index(i) + 1) for i in range(len(k))))
for r in rows:
    print("  row      %s" % "  ".join(r))
ct2 = columnar(MSG, KEY)
print("  ciphertext %s" % ct2)
print("  decrypted  %s" % columnar(ct2, KEY, decrypt=True))

print()
print("the repeated key letter, which is where a careless columnar breaks")
for key in ("ANNA", "ZEBRAS", "AAAA"):
    c = columnar(MSG, key)
    back = columnar(c, key, decrypt=True)
    print("  key %-8s ciphertext %-26s round trip %s" % (key, c, back.startswith(MSG)))

print()
print("letter counts are unchanged by transposition")
from collections import Counter
print("  plaintext  ", sorted(Counter(MSG).items()))
print("  ciphertext ", sorted(Counter(rail_fence(MSG, 3)).items()))
print("  identical:", Counter(MSG) == Counter(rail_fence(MSG, 3)))
munotes.in324

Transposition Systems

plaintext SENDTHEARMYATDAWN (17 letters)

rail fence, three rails: the pattern of rails, then the read-off
  rail 0  S...T...R...T...N
  rail 1  .E.D.H.A.M.A.D.W.
  rail 2  ..N...E...Y...A..
  ciphertext STRTNEDHAMADWNEYA
  decrypted  SENDTHEARMYATDAWN

rail fence with two to five rails
  2 rails  SNTERYTANEDHAMADW    round trip True
  3 rails  STRTNEDHAMADWNEYA    round trip True
  4 rails  SETEHAADNTRYANDMW    round trip True
  5 rails  SRNEAMWNEYADHADTT    round trip True

columnar transposition under the key ZEBRAS
  key      Z  E  B  R  A  S
  read in  6  3  2  4  1  5
  row      S  E  N  D  T  H
  row      E  A  R  M  Y  A
  row      T  D  A  W  N  X
  ciphertext TYNNRAEADDMWHAXSET
  decrypted  SENDTHEARMYATDAWNX

the repeated key letter, which is where a careless columnar breaks
  key ANNA     ciphertext STRTNDAAWXEHMDXNEYAX       round trip True
  key ZEBRAS   ciphertext TYNNRAEADDMWHAXSET         round trip True
  key AAAA     ciphertext STRTNEHMDXNEYAXDAAWX       round trip True

letter counts are unchanged by transposition
  plaintext   [('A', 3), ('D', 2), ('E', 2), ('H', 1), ('M', 1), ('N', 2), ('R', 1), ('S', 1), ('T', 2), ('W', 1), ('Y', 1)]
  ciphertext  [('A', 3), ('D', 2), ('E', 2), ('H', 1), ('M', 1), ('N', 2), ('R', 1), ('S', 1), ('T', 2), ('W', 1), ('Y', 1)]
  identical: True

Reading the output

The rail pattern is printed, so the zigzag is visible rather than described. Rail 0 takes every fourth letter, rail 2 takes every fourth offset by two, and rail 1 takes the rest.

Four rail counts all round-trip, which is the minimum test for any cipher: encipher, decipher, compare.

The columnar working is printed as the grid, with the key, the reading order and the rows, so the hand method and the program agree line for line.

Three keys including a repeated letter all round-trip. ANNA has two Ns, and the two N columns must be read in a fixed order or the message does not come back. The program sorts by the letter and then by the position, which is a stable sort, and that is what makes ANNA work. A sort that is not stable would give two different column orders on two different machines, and the message would decipher on one and not the other.

And the last block is the defining property. The letter counts of the plaintext and of the ciphertext are identical, and the program compares them rather than asserting it.

munotes.in325

Transposition Systems

What that property means for an attacker

Frequency analysis is useless against a pure transposition. The ciphertext's letter frequencies are the plaintext's, which are English's, and they say nothing about the key.

But the same fact identifies the cipher. An attacker who sees a ciphertext with normal English letter frequencies and no readable words knows at once that it is a transposition and not a substitution. The defining property is also the giveaway.

And what does break it. Anagramming: try likely arrangements and look for common letter pairs. A transposition preserves which letters are present, so a pair like TH must be somewhere, and the attacker looks for arrangements that bring likely pairs together. For the rail fence, trying every rail count is enough.

Why both primitives are needed

SubstitutionTransposition
Changesthe identity of the lettersthe position of the letters
Preservespositionidentity, and therefore the letter counts
Broken byfrequency analysisanagramming, or trying the small key space
Betrays itself byan abnormal frequency profilea normal frequency profile with no words

Neither alone survives. Substitution leaves the frequencies to be read; transposition leaves the letters to be rearranged.

Composed, each covers the other's weakness. The substitution disturbs the frequencies so anagramming has less to go on, and the transposition breaks up the letter positions so frequency analysis cannot be confirmed by reading words. That is why every practical classical system used both, and it is what MU's label "substitution and transposition systems" is naming as a pair.

Quick revision

  • Transposition permutes positions and leaves symbols alone, so the letter counts are exactly preserved.
  • Rail fence: zigzag across a number of rails, read off rail by rail. The key is the rail count, so the key space is tiny.
  • Columnar: rows of the key's width, columns read in the alphabetical order of the key's letters.
  • Padding is needed to fill the last row, which lengthens the ciphertext, must be removable, and must be added at the right stage.
  • A repeated key letter needs a stable sort, or the column order is not reproducible.
  • Frequency analysis is useless against a transposition, and the normal frequency profile is what identifies it as one.
  • Neither primitive survives alone; composed, each covers the other's weakness.

Test yourself

1. Encipher ATTACK with a three-rail fence, showing the rails.

Rail 0 takes A and C; rail 1 takes T, A and K; rail 2 takes T. Reading rail by rail gives AC TAK T, that is, ACTAKT.

2. Why must a columnar transposition sort a repeated key letter by position as well as by letter?

Because two identical key letters give two columns with the same sort key, and an unstable sort may order them either way. The order must be the same when enciphering and deciphering, so the position is used as a tie-break.

munotes.in326

Transposition Systems

3. What does the preservation of letter counts cost the attacker, and what does it give them?

It costs them frequency analysis, since the ciphertext's frequencies are the language's and say nothing about the key. It gives them the identification: a ciphertext with normal letter frequencies and no readable words is a transposition.

4. Why is a composition of the two primitives stronger than either?

Because substitution disturbs the frequency profile, which is what anagramming a transposition relies on for confirmation, and transposition scatters the positions, which is what reading words off a substitution relies on. Each covers the other's weakness.

Contents This chapter on its own page

munotes.in327

Chapter Ninety-One

Substitution and Transposition in Code

Syllabus topic Module 2, "Substitution and transposition systems", "Algorithm Specification (Pseudo-code)", "minimum 10 test cases"

In one line

Two dozen lines each, and the tests are where the work is.

In the wording you can write in an examination: implementing a classical cipher requires the enciphering function, its exact inverse, and a test that composes them and compares the result with the input. Correctness is established by round-tripping every case, and the cases that matter are the ones at the boundaries: an empty message, a message shorter than the key, a key with repeated letters, and input containing characters outside the alphabet.

What a cipher implementation owes

Two functions and a proof that they are inverses. Encipher, decipher, and a test that decipher(encipher(m, k), k) equals m for every m and k tried.

A stated treatment of characters outside the alphabet. Spaces, digits and punctuation must either be preserved, or removed, or rejected, and the choice must be made and documented. Silently dropping them and then failing to restore them is the commonest defect in a student implementation.

A stated treatment of case. Upper, lower, or normalised.

And a stated treatment of padding, as [Transposition Systems] sets out.

The five boundary cases

These are the tests to write first, before any ordinary message.

CaseWhat it catches
the empty messagea loop that assumes at least one character; a division by the length
one characteran off-by-one in the padding or the grid
a message shorter than the keya grid with fewer rows than one; a padding calculation that goes negative
a key with a repeated letteran unstable sort, which gives two different column orders
input with spaces and punctuationan inconsistent decision about what to keep

And a sixth that is not a boundary but is a trap. A key whose letters are already in alphabetical order gives the identity permutation for a columnar transposition and enciphers nothing, exactly as a keyword already in order does for a substitution.

Working code

#!/usr/bin/env python3
"""Arthasastra-inspired cryptography as a symmetric cipher system. MU's topic 5.

IKS concept as CS concept: the gudhalekhya (cipher-writing) and samjnalipi
(writing by signs) that the Arthasastra names, built as a symmetric cipher that
composes a substitution with a transposition under one key, with the classical
attacks run against it so the limits are measured and not asserted.

 WHAT THE TREATISE DOES AND DOES NOT GIVE. Kautilya names secret writing
four times in R. Shamasastry's translation of 1915 and never states a cipher:
  Book I, ch. 12   "by making use of signs or writing (samjnalipibhih)"
  Book I, ch. 12   "through cipher-writing (gudhalekhya), or by means of signs"
  Book I, ch. 14   "by deciphering paintings and secret writings
                    (chitra-gudha-lekhya-samjnabhih)"
  Book XIII, ch. 1 "a sealed letter" carried by a domestic pigeon
So the cipher below is NOT Kautilya's. It is a modern construction on the two
primitives classical cryptography actually has, built to the requirement his
text states, which is that a message must be unreadable to a courier who
carries it. The book says this in as many words.
"""
import string
from collections import Counter

ALPHABET = string.ascii_uppercase


# -------------------------------------------------------------- substitution
def caesar(text, shift, decrypt=False):
    if decrypt:
        shift = -shift
    out = []
    for ch in text.upper():
        if ch in ALPHABET:
            out.append(ALPHABET[(ALPHABET.index(ch) + shift) % 26])
        else:
            out.append(ch)
    return ''.join(out)


def keyword_alphabet(key):
    """A monoalphabetic key: the keyword's distinct letters, then the rest."""
    seen = []
    for ch in key.upper():
        if ch in ALPHABET and ch not in seen:
            seen.append(ch)
    for ch in ALPHABET:
        if ch not in seen:
            seen.append(ch)
    return ''.join(seen)


def substitute(text, key, decrypt=False):
    cipher = keyword_alphabet(key)
    src, dst = (cipher, ALPHABET) if decrypt else (ALPHABET, cipher)
    table = dict(zip(src, dst))
    return ''.join(table.get(ch, ch) for ch in text.upper())


def vigenere(text, key, decrypt=False):
    key = [c for c in key.upper() if c in ALPHABET]
    if not key:
        raise ValueError('a Vigenere key needs at least one letter')
    out, i = [], 0
    for ch in text.upper():
        if ch not in ALPHABET:
            out.append(ch)
            continue
        k = ALPHABET.index(key[i % len(key)])
        if decrypt:
            k = -k
        out.append(ALPHABET[(ALPHABET.index(ch) + k) % 26])
        i += 1
    return ''.join(out)


# -------------------------------------------------------------- transposition
def rail_fence(text, rails, decrypt=False):
    if rails < 2:
        raise ValueError('a rail fence needs at least two rails')
    n = len(text)
    pattern = []
    r, step = 0, 1
    for _ in range(n):
        pattern.append(r)
        if r == 0:
            step = 1
        elif r == rails - 1:
            step = -1
        r += step
    if not decrypt:
        return ''.join(text[i] for rr in range(rails)
                       for i in range(n) if pattern[i] == rr)
    order = [i for rr in range(rails) for i in range(n) if pattern[i] == rr]
    out = [''] * n
    for pos, i in enumerate(order):
        out[i] = text[pos]
    return ''.join(out)


def columnar(text, key, decrypt=False, pad='X'):
    """Columnar transposition. Columns are read in the key's sorted order.

     A repeated key letter is settled by position, so the order is a stable
    sort. Without that, two identical letters give two different orders on two
    different machines, and the message does not come back.
    """
    k = [c for c in key.upper() if c in ALPHABET]
    if not k:
        raise ValueError('a columnar key needs at least one letter')
    cols = len(k)
    order = sorted(range(cols), key=lambda i: (k[i], i))
    if not decrypt:
        body = text + pad * ((-len(text)) % cols)
        rows = [body[i:i + cols] for i in range(0, len(body), cols)]
        return ''.join(''.join(r[c] for r in rows) for c in order)
    rows_n = len(text) // cols
    if rows_n * cols != len(text):
        raise ValueError('a columnar ciphertext must be a whole number of rows')
    chunks, at = {}, 0
    for c in order:
        chunks[c] = text[at:at + rows_n]
        at += rows_n
    return ''.join(chunks[c][r] for r in range(rows_n) for c in range(cols))


# --------------------------------------------------------- the composed cipher
def clean(text):
    return ''.join(ch for ch in text.upper() if ch in ALPHABET)


def encrypt(plaintext, key):
    """Substitution under the key, then columnar transposition under the key.

     THE PADDING GOES IN BEFORE THE SUBSTITUTION, NOT AFTER IT. Padding
    after substitution pads the CIPHERTEXT, and the inverse substitution then
    turns each pad letter into whatever maps to it: under the key KAUTILYA an
    X came back as a Y, so a message round tripped as SENDTHEARMYATDAWNYYYYYYY.
    That was the first version and the test caught it.
    """
    body = clean(plaintext)
    cols = len([c for c in key.upper() if c in ALPHABET])
    if not cols:
        raise ValueError('the key needs at least one letter')
    body += 'X' * ((-len(body)) % cols)
    return columnar(substitute(body, key), key)


def decrypt(ciphertext, key):
    return substitute(columnar(ciphertext, key, decrypt=True), key, decrypt=True)


# --------------------------------------------------------------- the attacks
ENGLISH = 'ETAOINSHRDLCUMWFGYPBVKJXQZ'


def frequencies(text):
    c = Counter(ch for ch in text.upper() if ch in ALPHABET)
    total = sum(c.values()) or 1
    return [(ch, n, round(100 * n / total, 2)) for ch, n in c.most_common()]


def break_caesar(ciphertext):
    """Every shift scored by letter frequency; the best is returned.

    A brute force over 26 keys, which is the whole key space of a Caesar
    cipher, and that is the point being made.
    """
    best, best_score = None, -1.0
    for shift in range(26):
        guess = caesar(ciphertext, shift, decrypt=True)
        score = sum(guess.count(ch) * (26 - i) for i, ch in enumerate(ENGLISH))
        if score > best_score:
            best, best_score = shift, score
    return best, caesar(ciphertext, best, decrypt=True)


def key_space():
    import math
    return {
        'caesar': 26,
        'monoalphabetic substitution': math.factorial(26),
        'rail fence up to 10 rails': 9,
        'columnar with a 7 column key': math.factorial(7),
    }


# ------------------------------------------------------------ steganography
def lsb_hide(pixels, message):
    """Hide a message in the least significant bit of a list of byte values."""
    bits = ''.join(format(b, '08b') for b in message.encode('ascii')) + '0' * 8
    if len(bits) > len(pixels):
        raise ValueError('the carrier holds %d bits and the message needs %d'
                         % (len(pixels), len(bits)))
    out = list(pixels)
    for i, bit in enumerate(bits):
        out[i] = (out[i] & ~1) | int(bit)
    return out


def lsb_read(pixels):
    bits = ''.join(str(p & 1) for p in pixels)
    out = bytearray()
    for i in range(0, len(bits) - 7, 8):
        byte = int(bits[i:i + 8], 2)
        if byte == 0:
            break
        out.append(byte)
    return out.decode('ascii')


# ------------------------------------------------------------------- tests
MSG = 'SEND THE ARMY AT DAWN'

TESTS = [
    ('a plain message round trips',        MSG,                        'KAUTILYA'),
    ('one letter',                         'A',                        'KAUTILYA'),
    ('a message shorter than the key',     'HI',                       'KAUTILYA'),
    ('a message the length of the key',    'ABCDEFGH',                 'KAUTILYA'),
    ('punctuation and digits are dropped', 'MEET AT 9, BY THE GATE!',  'KAUTILYA'),
    ('lower case is accepted',             'send the army at dawn',    'KAUTILYA'),
    ('a key with a repeated letter',       MSG,                        'PATALIPUTRA'),
    ('a one letter key',                   MSG,                        'K'),
    ('a key longer than the message',      'GO',                       'CHANDRAGUPTA'),
    ('the whole alphabet as the key',      MSG,                        ALPHABET),
    ('a long message',                     MSG * 7,                    'MAURYA'),
    ('a message of repeated letters',      'AAAAAAAAAAAA',             'KAUTILYA'),
]


def run_tests(verbose=False):
    passed = 0
    for name, text, key in TESTS:
        ct = encrypt(text, key)
        back = decrypt(ct, key)
        expect = clean(text)
        ok = back.startswith(expect) and set(back[len(expect):]) <= {'X'}
        passed += ok
        if verbose:
            print('%-38s %s  %s -> %s' % (name, 'pass' if ok else 'FAIL',
                                          expect[:22], ct[:22]))
        assert ok, (name, expect, back)
    return passed


def prove():
    assert caesar('ABC', 3) == 'DEF'
    assert caesar('DEF', 3, decrypt=True) == 'ABC'
    assert keyword_alphabet('KAUTILYA').startswith('KAUTIL')
    assert len(set(keyword_alphabet('KAUTILYA'))) == 26
    assert substitute(substitute('HELLO', 'KAUTILYA'), 'KAUTILYA', decrypt=True) == 'HELLO'
    assert vigenere(vigenere('ATTACKATDAWN', 'ARTHA'), 'ARTHA', decrypt=True) == 'ATTACKATDAWN'
    assert rail_fence(rail_fence('WEAREDISCOVERED', 3), 3, decrypt=True) == 'WEAREDISCOVERED'
    assert columnar(columnar('ATTACKATDAWN', 'ZEBRAS'), 'ZEBRAS', decrypt=True).startswith('ATTACKATDAWN')
    # the repeated key letter, which is where a careless columnar breaks
    assert columnar(columnar('SENDHELP', 'ANNA'), 'ANNA', decrypt=True).startswith('SENDHELP')
    # a Caesar falls to its own key space
    shift, plain = break_caesar(caesar('THE ENEMY IS AT THE GATE AND WE MUST SEND WORD', 11))
    assert shift == 11 and plain.startswith('THE ENEMY'), (shift, plain)
    # steganography round trips
    carrier = [(i * 7) % 256 for i in range(400)]
    assert lsb_read(lsb_hide(carrier, 'GUDHALEKHYA')) == 'GUDHALEKHYA'
    run_tests()
    return True


if __name__ == '__main__':
    prove()
    print('### the composed cipher on one message')
    ct = encrypt(MSG, 'KAUTILYA')
    print('   plaintext  %s' % clean(MSG))
    print('   after substitution  %s' % substitute(clean(MSG), 'KAUTILYA'))
    print('   after transposition %s' % ct)
    print('   decrypted  %s' % decrypt(ct, 'KAUTILYA'))
    print()
    print('### the key spaces')
    for k, v in key_space().items():
        print('   %-30s %s' % (k, format(v, ',')))
    print()
    print('### the twelve test cases')
    n = run_tests(verbose=True)
    print('\n%d of %d pass' % (n, len(TESTS)))
    print()
    print('### frequency analysis of a Caesar ciphertext')
    sample = 'THE ENEMY IS AT THE GATE AND WE MUST SEND WORD TO THE CAPITAL AT ONCE'
    enc = caesar(sample, 11)
    print('   ciphertext %s' % enc)
    for ch, n2, pc in frequencies(enc)[:6]:
        print('     %s %3d  %5.2f%%' % (ch, n2, pc))
    shift, plain = break_caesar(enc)
    print('   broken at shift %d: %s' % (shift, plain))
munotes.in328

Substitution and Transposition in Code

### the composed cipher on one message
   plaintext  SENDTHEARMYATDAWN
   after substitution  PIHTQBIKOGXKQTKVH
   after transposition IGWKVWQQWPOHBTWTKWHXWIKW
   decrypted  SENDTHEARMYATDAWNXXXXXXX

### the key spaces
   caesar                         26
   monoalphabetic substitution    403,291,461,126,605,635,584,000,000
   rail fence up to 10 rails      9
   columnar with a 7 column key   5,040

### the twelve test cases
a plain message round trips            pass  SENDTHEARMYATDAWN -> IGWKVWQQWPOHBTWTKWHXWI
one letter                             pass  A -> WWWKWWWW
a message shorter than the key         pass  HI -> CWWBWWWW
a message the length of the key        pass  ABCDEFGH -> ABIKLTUY
punctuation and digits are dropped     pass  MEETATBYTHEGATE -> IBXWKKGQQQQYIIAI
lower case is accepted                 pass  SENDTHEARMYATDAWN -> IGWKVWQQWPOHBTWTKWHXWI
a key with a repeated letter           pass  SENDTHEARMYATDAWN -> IQLPYXBHQWOPIXGXHLNXPX
a one letter key                       pass  SENDTHEARMYATDAWN -> SDNCTGDKRMYKTCKWN
a key longer than the message          pass  GO -> XXXGXXJXXXXX
the whole alphabet as the key          pass  SENDTHEARMYATDAWN -> SENDTHEARMYATDAWNXXXXX
a long message                         pass  SENDTHEARMYATDAWNSENDT -> YMRJOMRIVQXJDMPYQYMRDM
a message of repeated letters          pass  AAAAAAAAAAAA -> KKKWKWKKKWKKKKKW

12 of 12 pass

### frequency analysis of a Caesar ciphertext
   ciphertext ESP PYPXJ TD LE ESP RLEP LYO HP XFDE DPYO HZCO EZ ESP NLATELW LE ZYNP
     E   9  16.67%
     P   9  16.67%
     L   6  11.11%
     Y   4   7.41%
     S   3   5.56%
     D   3   5.56%
   broken at shift 11: THE ENEMY IS AT THE GATE AND WE MUST SEND WORD TO THE CAPITAL AT ONCE
munotes.in329

Substitution and Transposition in Code

Reading the output

Twelve cases, all round-tripping. The comparison is not "the decipherment equals the plaintext" but "the decipherment starts with the plaintext and anything after it is padding", because a columnar transposition lengthens the message.

munotes.in330

Substitution and Transposition in Code

That comparison is itself a design decision and it is stated in the code: back.startswith(expect) and set(back[len(expect):]) <= {"X"}. A test that simply compared for equality would fail every padded case, and the temptation would be to strip the padding in the decipher function, which would then corrupt a message legitimately ending in X.

munotes.in331

Substitution and Transposition in Code

The repeated-key-letter case is the one to point at in a viva. PATALIPUTRA has three As and two Ts and two Ps. The column order is fixed by sorting on the letter and then on the position, and without the second key the order would depend on the sort's implementation.

And the empty-affix case, the one-letter message, produces a ciphertext of eight characters, because the key has eight letters and the grid must be filled. Seven of the eight characters are padding, which is a real property of the scheme: it hides the length of a short message and wastes space on it.

The composed cipher

Substitution first, then transposition. Both under the same key.

Why that order. The substitution disturbs the letter identities; the transposition then scatters them. Doing it the other way round works too and gives a different ciphertext; the composition is not commutative, and saying so is worth a mark.

And why one key for both. Because it is simpler and it is what MU's topic asks for: a symmetric cipher with a key. A real system would derive two independent subkeys from one master key, and a submission that says so has noticed the weakness: with one key, an attacker who recovers the substitution has also recovered the transposition.

munotes.in332

Substitution and Transposition in Code

The padding bug, recorded

This book's own first version of the composed cipher padded AFTER the substitution.

What happened. The pad character X was appended to the ciphertext of the substitution step, so the inverse substitution mapped it to whatever plain letter maps to X. Under the key KAUTILYA that is Y, and a message round-tripped as SENDTHEARMYATDAWNYYYYYYY.

Why the test caught it. The round-trip comparison allows trailing X and nothing else. Trailing Y failed it.

And why a weaker test would not have. A test that only checked that the ciphertext differed from the plaintext, or that the decipherment had the right length, would have passed. A round-trip test is the minimum, and it has to compare the content.

Complexity

Substitution. One table lookup per character, so the work is proportional to the length of the message. Building the keyword alphabet is proportional to the alphabet size.

Rail fence. One pass to compute the pattern and one to read off, so again proportional to the length.

Columnar. Sorting the key is proportional to the key length times its logarithm, and the read-off is proportional to the message length.

So the whole composed cipher is linear in the message, which is what any usable cipher must be.

Quick revision

  • A cipher implementation owes two functions, a round-trip test, and stated treatments of non-alphabetic characters, case and padding.
  • Five boundary cases: empty, one character, shorter than the key, a repeated key letter, and input with punctuation.
  • The round-trip comparison must allow padding and must still compare the content; equality alone fails every padded case.
  • A repeated key letter needs a stable sort, keyed on the letter and then the position.
  • Composition is not commutative, and using one key for both steps means recovering one recovers both.
  • The padding bug: padding after the substitution turned the pad into another letter on the way back, and only a content-comparing round trip catches it.
  • All three operations are linear in the message length.

Test yourself

1. Name four boundary cases a cipher implementation must be tested on.

The empty message; a one-character message; a message shorter than the key; and a key containing a repeated letter. Input containing spaces and punctuation is a fifth.

2. Why can the round-trip test not simply compare for equality?

Because a columnar transposition pads the message to fill its last row, so the decipherment is longer than the plaintext. The test must allow trailing padding and must still compare the content, or it would either fail every padded case or be satisfied by a wrong decipherment.

3. Explain the padding bug this book's own code had.

munotes.in333

Substitution and Transposition in Code

The pad was appended after the substitution step, so it was part of the substitution's output. The inverse substitution then mapped the pad character to a different plain letter, and the message round-tripped with the wrong trailing characters. Padding must be added before the substitution.

4. Why is using one key for both steps a weakness?

Because the two steps are not independent: an attacker who recovers the key from the substitution has also recovered the transposition, so breaking one breaks both. A real system derives two independent subkeys from one master key.

Contents This chapter on its own page

munotes.in334

Chapter Ninety-Two

Concealment and Coded Messaging

Syllabus topic Module 2, "Concealment and coded messaging"

In one line

Concealment hides that there is a message; encryption hides what the message says; and the Arthaśāstra uses both.

In the wording you can write in an examination: concealment, or steganography, protects a communication by hiding its existence, so that an interceptor does not know a message is present. Encryption protects it by making its content unintelligible, so that an interceptor knows a message is present and cannot read it. A coded message, in the technical sense, replaces whole words or phrases by agreed substitutes rather than transforming letters.

The distinction, and why it is not academic

Different threat models.

Encryption assumes the interceptor may hold the message. Its security rests on the key, and its failure is that the interceptor knows they have something worth working on.

Concealment assumes the interceptor will not look. Its security rests on their not noticing, and its failure is total the moment they do.

And the failure modes differ in kind. A broken cipher yields one message. A discovered concealment technique yields every message sent by that channel, past and future, because they were never enciphered.

Where the Arthaśāstra uses each

Read the passages of [Secret Communication: What the Text Actually Says] again with the distinction in hand.

PassageWhich is it?
Book I, Chapter XII, information carried "under the pretext of taking in musical instruments"concealment: the carrier and the occasion hide the transfer
Book I, Chapter XII, "through cipher-writing (gūḍhalekhya)"encryption, or at least an unintelligible writing
Book I, Chapter XII, "or by means of signs"concealment: a gesture is not recognised as a message
Book I, Chapter XVI, "the signs made in places of pilgrimage and temples"concealment, in a public place
Book XIII, Chapter I, the sealed letter by domestic pigeonneither: a seal protects integrity, and the pigeon is a channel

Three of the five are concealment, and the treatise offers them as alternatives to cipher-writing in the same sentence. So the text treats them as interchangeable solutions to one problem, which is a judgement a modern engineer would dispute: they solve it under different assumptions and fail differently.

The null cipher

What it is. A message hidden inside an innocuous text, recovered by a rule known to both parties. The classic rule is to take the first letter of each word.

Worked. The carrier text:

Send every note down. The horse eats grass. And true rest. My

your after tomorrow. Dawn arrives when night ends.

Taking the first letter of each word gives SEND THE GRASS... which does not work, and that is the point of showing it: constructing a null cipher whose carrier reads naturally is hard, and a carrier that reads oddly defeats the purpose.

munotes.in335

Concealment and Coded Messaging

A shorter one that does work. Take the first letter of each word of:

Send every night. Dawn.

That is SEND, hidden in four words, and even here the carrier is stilted. The cost of concealment is the length and the naturalness of the carrier, and it grows with the message.

Its security. None, once the rule is guessed, and the rules are few. A null cipher protects against a casual reader and against nobody who is looking.

The open code

What it is. Ordinary language in which agreed words carry other meanings. "The grain shipment is delayed" means something else to the party who knows.

Its advantage over a null cipher. The carrier is entirely natural, because it is a real message about something.

Its cost. The vocabulary must be agreed in advance, it is limited to what has been agreed, and it cannot express anything outside the agreed list. An open code is a fixed dictionary, not a language.

And it is what the Arthaśāstra's "signs" most plausibly are. A prearranged gesture or a mark in a temple means one agreed thing. The text never suggests that signs compose into sentences.

The technical sense of "code"

This is a distinction a question may test and students conflate.

A cipherA code
Operates onletters or bitswords or phrases
Needsa key, and an algorithma codebook
Expressive rangeanything the alphabet can writeonly what is in the book
Changing the secretchange the keyreprint and redistribute the book
If the secret is compromisedpast traffic is exposedpast traffic is exposed, and the book must be replaced

"Code" in ordinary speech means cipher, and in cryptography it does not. Using the words precisely is worth a mark and costs nothing.

Combining concealment with encryption

The right arrangement is both, in that order. Encipher the message, then conceal the ciphertext.

Why that order. If the concealment is discovered, the interceptor has ciphertext and still needs the key. If only concealment were used, discovery yields the message.

Why not the reverse. Concealing first and then enciphering the carrier produces an enciphered carrier, which looks like ciphertext and so defeats the concealment entirely.

And the practical objection. Enciphered text is high-entropy and hard to conceal naturally: a null cipher spelling out random letters produces a carrier that reads like nothing. That tension is real and it is why modern steganography hides bits in media rather than letters in words, as [Basic Steganography] shows.

What concealment is NOT

It is not encryption with a different name. The properties are different and the failures are different.

It is not obsolete. Hiding the existence of a communication is exactly what traffic analysis attacks, and defending against traffic analysis is a live problem.

munotes.in336

Concealment and Coded Messaging

And it is not sufficient. Kerckhoffs's principle, that a system's security should rest on the key and not on the secrecy of the method, is the general statement of why. A concealment method is a method, and methods leak.

Quick revision

  • Concealment hides that a message exists; encryption hides what it says. Different threat models and different failures.
  • A broken cipher yields one message; a discovered concealment channel yields every message ever sent through it.
  • Three of the five Arthaśāstra passages are concealment and the treatise offers them as alternatives to cipher-writing.
  • A null cipher hides a message in a carrier by a rule; the cost is the carrier's length and naturalness, and it has no security once the rule is guessed.
  • An open code uses agreed words with other meanings: a natural carrier, and a fixed dictionary rather than a language.
  • Technically a code operates on words and needs a codebook; a cipher operates on letters and needs a key.
  • Encipher first, then conceal. The reverse defeats the concealment.

Test yourself

1. Distinguish concealment from encryption by threat model and by failure.

Encryption assumes the interceptor may hold the message and rests on the key; its failure exposes one message. Concealment assumes the interceptor will not look and rests on their not noticing; its failure exposes every message ever sent through that channel.

2. Distinguish a code from a cipher in the technical sense.

A cipher transforms letters or bits under a key and can express anything writable. A code replaces whole words or phrases using a codebook, and can express only what the book contains; changing the secret means redistributing the book.

3. In what order should concealment and encryption be combined, and why?

Encipher first, then conceal the ciphertext. If the concealment is discovered the interceptor still has only ciphertext. Concealing first and then enciphering produces something that looks like ciphertext, which defeats the concealment.

4. Which of the Arthaśāstra's methods are concealment, and what does the text's treatment of them suggest?

The pretext of carrying musical instruments, communication by signs, and signs made in temples and places of pilgrimage. The text offers them in the same sentence as alternatives to cipher-writing, which treats them as interchangeable although they rest on different assumptions and fail differently.

Contents This chapter on its own page

munotes.in337

Chapter Ninety-Three

Basic Steganography

Syllabus topic Module 2, "Basic steganography"

In one line

Least significant bit hiding puts one bit of a message into the lowest bit of each pixel, where a change of one in 256 is invisible.

In the wording you can write in an examination: steganography conceals the existence of a message by embedding it in a carrier whose alteration is imperceptible. In least significant bit hiding the message bits replace the lowest-order bit of successive samples of a digital carrier, typically the pixels of an image. The capacity is one bit per sample, the distortion is at most one unit per altered sample, and the method is detectable by statistical analysis of the low-order bits.

Why the lowest bit

A greyscale pixel is a number from 0 to 255. Its lowest bit contributes 1.

Changing it changes the shade by one part in 256. No eye sees it, and most displays cannot show it.

And every pixel has one, so an image of a million pixels carries a million bits, which is 125 kilobytes.

That is the whole idea. The carrier has more precision than anybody uses, and the unused precision is storage.

The method

To hide. Turn the message into bits. For each bit in turn, take the next pixel, clear its lowest bit, and set it to the message bit.

To recover. Read the lowest bit of each pixel in turn, group them into bytes, and stop at a terminator.

The terminator matters. Without it the reader does not know where the message ends and reads the carrier's own low bits as text. This implementation appends a zero byte, which is why the message may not itself contain one.

The method, run

"""Least significant bit hiding, on a small greyscale image the program makes."""

WIDTH, HEIGHT = 16, 16

def carrier():
    """A smooth gradient, so that a changed bit is invisible to the eye."""
    return [((x * 7 + y * 11) % 200) + 28 for y in range(HEIGHT) for x in range(WIDTH)]

def hide(pixels, message):
    bits = "".join(format(b, "08b") for b in message.encode("ascii")) + "0" * 8
    if len(bits) > len(pixels):
        raise ValueError("carrier holds %d bits, message needs %d" % (len(pixels), len(bits)))
    out = list(pixels)
    for i, bit in enumerate(bits):
        out[i] = (out[i] & ~1) | int(bit)
    return out

def read(pixels):
    bits = "".join(str(p & 1) for p in pixels)
    out = bytearray()
    for i in range(0, len(bits) - 7, 8):
        byte = int(bits[i:i + 8], 2)
        if byte == 0:
            break
        out.append(byte)
    return out.decode("ascii")

MSG = "GUDHALEKHYA"
plain = carrier()
stego = hide(plain, MSG)

print("carrier: %d by %d greyscale, %d pixels, %d bits of capacity"
      % (WIDTH, HEIGHT, len(plain), len(plain)))
print("message: %r, %d bytes, %d bits with the terminator"
      % (MSG, len(MSG), 8 * len(MSG) + 8))
print("capacity used: %.1f%%" % (100 * (8 * len(MSG) + 8) / len(plain)))
print()
print("the first sixteen pixels, before and after")
print("  before  " + " ".join("%3d" % p for p in plain[:16]))
print("  after   " + " ".join("%3d" % p for p in stego[:16]))
print("  change  " + " ".join("%3d" % (s - p) for p, s in zip(plain[:16], stego[:16])))
print()
changed = sum(1 for p, s in zip(plain, stego) if p != s)
biggest = max(abs(s - p) for p, s in zip(plain, stego))
print("pixels changed: %d of %d" % (changed, len(plain)))
print("largest change in any pixel: %d (out of a 0 to 255 range)" % biggest)
print("recovered message: %r" % read(stego))
print("round trip correct:", read(stego) == MSG)
print()
print("and the detection: the low bits of a natural image are not uniform")
def low_bit_balance(pixels, n):
    ones = sum(p & 1 for p in pixels[:n])
    return ones, n - ones
for name, px in (("carrier", plain), ("stego  ", stego)):
    ones, zeros = low_bit_balance(px, 96)
    print("  %s first 96 low bits: %d ones, %d zeros" % (name, ones, zeros))
print()
try:
    hide(plain, "X" * 40)
except ValueError as e:
    print("a message too large is refused:", e)
munotes.in338

Basic Steganography

carrier: 16 by 16 greyscale, 256 pixels, 256 bits of capacity
message: 'GUDHALEKHYA', 11 bytes, 96 bits with the terminator
capacity used: 37.5%

the first sixteen pixels, before and after
  before   28  35  42  49  56  63  70  77  84  91  98 105 112 119 126 133
  after    28  35  42  48  56  63  71  77  84  91  98 105 112 119 126 133
  change    0   0   0  -1   0   0   1   0   0   0   0   0   0   0   0   0

pixels changed: 44 of 256
largest change in any pixel: 1 (out of a 0 to 255 range)
recovered message: 'GUDHALEKHYA'
round trip correct: True

and the detection: the low bits of a natural image are not uniform
  carrier first 96 low bits: 48 ones, 48 zeros
  stego   first 96 low bits: 32 ones, 64 zeros

a message too large is refused: carrier holds 256 bits, message needs 328

Reading the output

A 16 by 16 image is 256 pixels and therefore 256 bits of capacity. The message is 11 bytes, 96 bits with the terminator, so 37.5 per cent of the capacity is used.

Forty-four pixels changed out of 256, not 96. That is because a bit only changes a pixel when it differs from the bit already there, and about half the time it agrees. The expected number of changes is half the bits used, which is 48, and 44 is what this carrier gave.

The largest change to any pixel is 1, on a range of 0 to 255. That is the imperceptibility claim, measured.

munotes.in339

Basic Steganography

And the round trip is checked, not assumed.

The detection, measured

The last block is the reason least significant bit hiding is not used where it matters.

The carrier's first 96 low bits are 48 ones and 48 zeros. That is what a smooth gradient gives: the low bit alternates with the gradient.

The stego image's are 32 ones and 64 zeros. The message's bits are ASCII text, whose bytes all begin with a zero bit and whose letters are concentrated in a narrow range, so its bit stream is not balanced.

So the alteration is visible in the statistics although it is invisible to the eye. A detector does not look at the picture; it counts the low bits and asks whether they look like the low bits of an image of that kind.

And that is the general principle worth stating. An embedding that does not match the carrier's own statistics is detectable however small it is, and the field of steganalysis is the business of finding such mismatches. Modern methods therefore embed in a way that preserves the expected statistics, at the cost of capacity.

Capacity, and its limits

One bit per sample is the simple version. Using two bits per pixel doubles the capacity and quadruples the maximum distortion, to 3 out of 255, which is still nearly invisible and is much more detectable.

And the carrier must be large. Hiding a message of n bytes needs at least 8n plus 8 samples. A 100 kilobyte message needs an image of more than 800,000 pixels.

The refusal at the end of the output is the check that must exist. A message too large for its carrier must be rejected, and an implementation that silently truncates it has produced a message the recipient cannot read and given no warning.

What steganography is NOT

It is not encryption. The message is in the carrier in plain form. Anyone who suspects the method reads it.

So it is used WITH encryption, as [Concealment and Coded Messaging] sets out: encipher first, then embed, so that discovery of the embedding yields ciphertext.

It is not robust. Recompressing the image, resizing it, or converting it to a lossy format destroys the low bits and the message with them. Least significant bit hiding survives only a lossless channel, and most channels are not.

And it is not new. The Arthaśāstra's information carried "under the pretext of taking in musical instruments" is the same idea with a different carrier, as is the signs made in temples. What is new is the carrier's precision, which is what makes the modern version high-capacity and invisible.

munotes.in340

Basic Steganography

The comparison with the classical methods

The Arthaśāstra's concealmentLeast significant bit hiding
Carrieran object, an occasion, a gesturethe unused precision of a digital sample
Capacityone prearranged meaningone bit per sample, so kilobytes
Detectionby a suspicious observerby counting the low bits
Robustnesssurvives handlingdestroyed by recompression
Needs prior agreementyes, the sign's meaningyes, the method and the terminator

The last row is the one they share, and it is the standing problem of all concealment: both parties must agree on the method before any message is sent, which is the key distribution problem of [Symmetric Encryption] in a different coat.

Quick revision

  • LSB hiding replaces the lowest bit of each sample with a message bit. Capacity one bit per sample; distortion at most one unit.
  • A terminator is required, or the reader cannot tell where the message ends.
  • About half the bits change a pixel, because half the time the bit already agrees.
  • Detection is by statistics, not by eye: the carrier's low bits were balanced 48 to 48 and the stego image's were 32 to 64.
  • An embedding that does not match the carrier's own statistics is detectable however small it is.
  • It is not encryption, it is not robust to recompression, and a message too large must be refused rather than truncated.
  • Both parties must agree the method in advance, which is key distribution in another form.

Test yourself

1. Why is the lowest bit chosen, and what is the capacity of a one-megapixel greyscale image?

Because changing it alters the value by one part in 256, which is imperceptible. The capacity is one bit per pixel, so about a million bits, which is 125 kilobytes.

2. Why did 96 message bits change only 44 pixels?

Because a message bit only changes a pixel when it differs from the bit already in the lowest position, which happens about half the time. The expected number of changes is half the bits used.

3. How is least significant bit hiding detected, and what does that tell you about designing a better method?

By counting the low-order bits and comparing their distribution with what a carrier of that kind should have; here the balance moved from 48 ones in 96 to 32. It tells you that an embedding must preserve the carrier's own statistics, which costs capacity.

4. Why must steganography be combined with encryption, and in what order?

Because the hidden message is in plain form and is readable by anyone who suspects the method. Encipher first and then embed, so that discovering the embedding yields only ciphertext.

Contents This chapter on its own page

munotes.in341

Chapter Ninety-Four

Information Protection Mechanisms

Syllabus topic Module 2, "Information protection mechanisms"

In one line

Protecting information is not one thing: the Arthaśāstra uses a seal, a courier, compartmentalisation and concealment, and each protects a different property.

In the wording you can write in an examination: information protection mechanisms are the arrangements by which a communication system preserves the properties it requires. Confidentiality is protected by encipherment and by concealment; integrity by a seal or other tamper-evident device; the limitation of damage from a compromised participant by compartmentalisation; and the availability of a channel by redundancy of routes.

The four mechanisms in the text

The seal

The passage. Book XIII, Chapter I: information "just received through a domestic pigeon which has brought a sealed letter."

Worked. What a seal does. It makes tampering evident. A broken seal shows that the letter was opened; an intact one is evidence that it was not.

What a seal does NOT do. It does not hide the contents from somebody willing to break it. A seal is integrity, not confidentiality, and treating it as confidentiality is the commonest mistake in writing about this material.

Nor does it authenticate, unless the seal bears a device only the sender possesses, which a signet ring does. The text does not describe the seal, so this book does not say which it was.

The modern counterpart. A message authentication code or a digital signature: a value computed over the message that a recipient can check and an alterer cannot forge. Both are integrity, and neither hides anything.

The courier who does not know

The passage. Book I, Chapter XII: the information is conveyed by door-keepers, artisans, court-bards or prostitutes, under the pretext of taking in musical instruments, or by cipher-writing, or by signs.

What it does. It separates the carrying of a message from the knowing of it. A carrier who cannot read the message cannot betray it, and a carrier who does not know they carry one cannot be interrogated about it.

The modern counterpart. An untrusted transport. The network carries your traffic and is not trusted with it, which is the assumption every encrypted protocol is built on. The mechanism is the same and the reason is the same.

Compartmentalisation

The passage. Book I, Chapter XII again: "The immediate officers of the institutes of espionage shall by making use of signs or writing set their own spies in motion (to ascertain the validity of the information)."

What it implies. Officers have their own spies. An officer's agents are that officer's, and the structure is a hierarchy rather than a pool.

Why that is a security property. An agent who is turned exposes their own officer and their own chain, and not the whole apparatus. Limiting the blast radius of one compromise is compartmentalisation, and the text's structure has it whether or not the author reasoned about it in those terms.

munotes.in342

Information Protection Mechanisms

The modern counterpart. Least privilege, and network segmentation. A component holds only the access it needs, so a compromised component yields only that.

Corroboration

The same passage. The purpose stated is to ascertain the VALIDITY of the information.

What it does. Information from one agent is checked against another's. That protects against a turned agent feeding false intelligence, which is a threat encipherment does not touch at all.

The modern counterpart. Quorum, and independent verification. And it is Charaka's rule again, from [Parīkṣā: How an Examination Is Structured]: one means of knowledge is not enough.

The mechanisms against the properties

PropertyWhat it meansThe text's mechanismThe modern control
confidentialityonly the intended party can read itgūḍhalekhya, and concealmentencryption
integrityalteration is detectablethe seala MAC or a signature
authenticationthe sender is who they claimnot addressed, unless the seal bore a devicea signature, or a shared secret
limited damageone compromise does not expose everythingofficers with their own spiesleast privilege, segmentation
correctness of contentthe information is truecorroboration between agentsindependent verification, quorum
availabilitythe channel works when neededseveral channels offered as alternativesredundancy

Read the third row. Authentication is the property the text addresses least, and it is the one a modern system needs most. An enciphered message from an unknown sender is worth very little.

And read the fifth. Correctness of content is a property no cryptography provides. A perfectly enciphered, perfectly authenticated lie is still a lie, and the only answer is corroboration. The classical apparatus is stronger on this than a naive modern account, because its designers expected their own agents to be turned.

The redundancy of channels

Book I, Chapter XII offers three ways to get the information out in one sentence: the pretext, the cipher-writing, and the signs.

Read as engineering, that is redundancy. If one channel is unavailable, another is used.

And read as security, it is a weakness. Three channels are three attack surfaces, and the system is as strong as the weakest. A defender must secure all three; an attacker needs one.

Both readings are correct and they pull in opposite directions. That trade, between availability and attack surface, is live in every system, and the Arthaśāstra's sentence contains it.

What this analysis does NOT claim

Not that Kautilya reasoned in these categories. The categories are ours. What the text supplies is the arrangements; classifying them is the analysis MU's label asks for.

Not that the arrangements were adequate. [Breaking a Classical Cipher] measures what a classical cipher would withstand.

munotes.in343

Information Protection Mechanisms

And not that anything here is a protocol. A protocol specifies who sends what when, and the text specifies none of that. [Secure Protocol Abstraction] says what would be needed.

Quick revision

  • Four mechanisms in the text: the seal, the untrusted courier, compartmentalisation through officers with their own agents, and corroboration.
  • A seal is integrity, not confidentiality, and it authenticates only if it bears a device only the sender has.
  • The untrusted courier is the same assumption as an untrusted network.
  • Officers with their own spies limit the damage of one compromise, which is least privilege.
  • Corroboration protects against false information, which no cryptography addresses.
  • Authentication is the property the text addresses least and the one a modern system needs most.
  • Three alternative channels are redundancy and are also three attack surfaces.

Test yourself

1. What does a seal protect, and what does it not protect?

It protects integrity: tampering becomes evident. It does not protect confidentiality, since anybody willing to break it can read the contents, and it authenticates only if it carries a device unique to the sender.

2. Which mechanism corresponds to an untrusted network, and what is the shared assumption?

The courier who carries without knowing. The shared assumption is that the transport will handle the message and must not be trusted with it, so protection is applied to the message rather than to the channel.

3. Which property does cryptography not provide, and what does the text use instead?

The correctness of the content. A perfectly enciphered and authenticated message may still be false. The text uses corroboration: officers set their own spies in motion to ascertain the validity of the information.

4. Give both readings of the three alternative channels offered in one sentence of Book I, Chapter XII.

As engineering it is redundancy: if one channel is unavailable another serves. As security it is three attack surfaces, and the system is only as strong as the weakest, since a defender must secure all three and an attacker needs one.

Contents This chapter on its own page

munotes.in344

Chapter Ninety-Five

Symmetric Encryption

Syllabus topic Module 2, "Symmetric encryption"

In one line

Symmetric encryption uses the same key to lock and unlock, which is fast and leaves you with the problem of getting the key to the other party.

In the wording you can write in an examination: a symmetric cipher uses a single key for both encipherment and decipherment. It is distinguished from an asymmetric cipher, in which a public key enciphers and a different private key deciphers. Symmetric ciphers are fast and require the key to be shared in advance through some channel other than the one being protected, which is the key distribution problem.

The defining property

One key, both directions. Encipher with k, decipher with k.

Every cipher in this block is symmetric. Caesar's shift, the keyword alphabet, the rail fence, the columnar transposition, and their composition in [Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System]. So is the Arthaśāstra's gūḍhalekhya, whatever it was, because a prearranged secret is what makes the recipient able to read it.

And so are the modern ones a student meets: DES and AES. Asymmetric cryptography is a twentieth-century invention and is a different subject.

Block and stream

The two shapes a modern symmetric cipher takes.

Block cipherStream cipher
Operates onfixed-size blocks, typically 128 bitsone bit or byte at a time
Needs paddingyes, to fill the last blockno
Natural fordata at rest, files, recordsdata arriving continuously
ExamplesDES, AESRC4, and the keystream generators
How it worksa keyed permutation on the blocka keystream combined with the plaintext

A block cipher is a permutation of the block space. For a 128-bit block that is a permutation of 2 to the power 128 values, chosen by the key.

A stream cipher generates a keystream from the key and combines it with the plaintext, usually by exclusive or. And that is where the one-time pad sits: a keystream as long as the message, truly random, never reused.

The classical ciphers of this block are neither, strictly. A monoalphabetic substitution operates on one letter at a time, which makes it look like a stream cipher, and its keystream is constant, which is exactly why it is weak.

Modes of operation

A block cipher enciphers one block. A message is many blocks, and how they are chained is the mode.

Electronic codebook. Encipher each block independently. And this is the mode nobody should use, because identical plaintext blocks give identical ciphertext blocks, so the structure of the plaintext shows through. It is the monoalphabetic substitution problem of [Substitution Systems] at the level of blocks.

Cipher block chaining. Combine each plaintext block with the previous ciphertext block before enciphering. Identical plaintext blocks now give different ciphertext, because the previous block differs. Requires an initialisation vector for the first block.

munotes.in345

Symmetric Encryption

Counter mode. Encipher a counter and combine the result with the plaintext, turning a block cipher into a stream cipher. Blocks can be enciphered in any order and in parallel.

The level this paper needs is that the mode matters as much as the cipher, and that electronic codebook leaks structure. A question is unlikely to go further.

The key distribution problem

The problem. Both parties need the key before they can communicate securely, and the key cannot be sent over the channel they are trying to protect.

The classical answer. Arrange it in person, in advance. That is what the Arthaśāstra's prearranged signs require: [Secret Communication: What the Text Actually Says] notes that a sign system must be agreed before any message exists.

Worked, and the arithmetic is the reason it does not scale. For n parties who must each be able to talk privately to every other, the number of keys is n times n minus 1, over 2.

PartiesKeys needed
21
510
1045
1004,950
1,000499,500

A hundred parties need 4,950 keys, every one of which must be distributed securely and kept secret, and adding one more party needs a hundred new keys. That is why symmetric cryptography alone cannot serve an open network, and it is the problem asymmetric cryptography was invented to solve.

And the modern arrangement is both. A key is agreed using asymmetric cryptography, and then the traffic is enciphered with a symmetric cipher, because symmetric ciphers are far faster. Asymmetric for the key, symmetric for the data.

Kerckhoffs's principle

The statement. A cryptographic system should remain secure even if everything about it except the key is public knowledge.

Why it matters. A method can be reverse-engineered, leaked or deduced; a key can be changed. Security resting on a secret method is security that cannot be repaired.

And it bears on the Arthaśāstra directly. The text names gūḍhalekhya and gives no method, so whatever the method was, its secrecy was part of the protection. A system whose method is secret has no way to recover when the method is discovered, and that is a design criticism that can be made without knowing what the method was.

What symmetric encryption does NOT provide

Not authentication. A message deciphering correctly shows that the sender had the key, and if the key is shared by two parties it does not say which of them sent it. And an attacker who can alter ciphertext can often alter the plaintext predictably.

Not integrity. Encipherment does not detect alteration. That needs a separate mechanism, as [Information Protection Mechanisms] sets out.

munotes.in346

Symmetric Encryption

Not non-repudiation. Since both parties hold the same key, neither can prove the other sent something.

Three properties missing, and all three are why modern protocols use authenticated encryption, which combines a cipher with an integrity mechanism rather than relying on the cipher alone.

Quick revision

  • Symmetric: one key both ways. Every cipher in this block is symmetric, and so was the Arthaśāstra's, whatever it was.
  • Block ciphers operate on fixed-size blocks and need padding; stream ciphers operate on a bit or byte at a time and do not.
  • Modes: electronic codebook leaks structure and should not be used; cipher block chaining hides it at the cost of an initialisation vector; counter mode turns a block cipher into a stream cipher.
  • Key distribution: n parties need n(n-1)/2 keys, so a hundred parties need 4,950.
  • Modern practice: asymmetric cryptography to agree the key, symmetric to encipher the data.
  • Kerckhoffs: security should rest on the key alone, because a method cannot be changed once discovered.
  • Encipherment provides confidentiality only, not authentication, integrity or non-repudiation.

Test yourself

1. Define a symmetric cipher and give the key distribution arithmetic for ten parties.

One in which the same key enciphers and deciphers. Ten parties who must each communicate privately with every other need ten times nine over two, which is 45 keys.

2. Why should electronic codebook mode not be used?

Because each block is enciphered independently, so identical plaintext blocks produce identical ciphertext blocks and the structure of the plaintext shows through the ciphertext, which is the monoalphabetic substitution weakness at the level of blocks.

3. State Kerckhoffs's principle and apply it to the Arthaśāstra.

A system should remain secure even if everything about it except the key is public. The Arthaśāstra names cipher-writing without giving a method, so the method's secrecy was evidently part of the protection, and a system of that kind cannot be repaired when the method becomes known.

4. Name three properties encipherment does not provide, and say what modern practice does about it.

Authentication, integrity and non-repudiation. Modern protocols use authenticated encryption, combining a cipher with a separate integrity mechanism, rather than relying on the cipher alone.

Contents This chapter on its own page

munotes.in347

Chapter Ninety-Six

Cipher Algorithms, Classical to Modern

Syllabus topic Module 2, "Cipher algorithms"

In one line

Every cipher in the sequence was invented to repair a specific weakness in the one before it.

In the wording you can write in an examination: the classical cipher algorithms comprise the monoalphabetic substitutions, of which Caesar is the simplest; the polyalphabetic substitutions, of which Vigenère is the standard example; the transpositions, rail fence and columnar; and the one-time pad, which is unbreakable and impractical. Modern symmetric ciphers, DES and AES, operate on blocks of bits under keys of stated length and combine substitution with permutation over many rounds.

The sequence

CipherWhat it doesKeyWhat it fixedHow it fails
Caesarrotates the alphabetone numbernothing; it is the starting point26 keys, tried in seconds
Keyword substitutionan arbitrary permutationa permutationthe tiny key spaceletter frequencies survive
Vigenèrea different shift per position, cycling with the keya wordthe single frequency profilethe key repeats, and the period is findable
Rail fencea zigzag permutation of positionsa small numberit attacks position, not identityvery few keys
Columnarcolumn read-off by key ordera wordthe rail fence's key spacefrequencies untouched, so it is identifiable
One-time pada random keystream as long as the messagethat keystreameverythingthe key is as long as the message and may never be reused
DES16 rounds of substitution and permutation on 64-bit blocks56 bitsmechanisation, and diffusionthe key is too short for modern machines
AES10 to 14 rounds on 128-bit blocks128, 192 or 256 bitsthe key lengthno practical break is known

The three ideas the sequence develops

Read the table downward and three ideas are being invented.

Confusion

The idea. Make the relationship between the key and the ciphertext complicated, so that seeing ciphertext tells you little about the key.

Caesar has almost none. One ciphertext letter and its plaintext gives the key.

A keyword substitution has more, and still not enough: each letter gives one entry of the table.

A modern cipher has a great deal, achieved by the substitution step of each round.

Diffusion

The idea. Spread the influence of each plaintext symbol over many ciphertext symbols, so that patterns in the plaintext do not survive.

No classical substitution has any. One letter in, one letter out, always in the same place.

A transposition has some: a letter's position moves, but its identity does not spread.

A modern block cipher has it by design. Changing one plaintext bit changes about half the ciphertext bits, which is called the avalanche effect, and it is achieved by the permutation step of each round.

Rounds

The idea. A single round of substitution and permutation is weak. Repeating it many times, with a different subkey each time, is strong.

munotes.in348

Cipher Algorithms, Classical to Modern

Classical ciphers have one round. That is their fundamental limitation, and composing a substitution with a transposition, as [Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System] does, is a two-round cipher and no more.

DES has sixteen rounds; AES has ten to fourteen depending on the key length.

Vigenère, since it is the one that needs explaining

The method. Write the key repeatedly under the plaintext. Each plaintext letter is shifted by the amount its key letter stands for.

Worked. Plaintext ATTACK, key ARTHA.

plaintext A T T A C K

key A R T H A A

shift 0 17 19 7 0 0

ciphertext A K M H C K

What it fixed. Each plaintext letter is enciphered by a different shift according to its position, so the single frequency profile of a monoalphabetic substitution is gone. The same plaintext letter becomes different ciphertext letters.

How it fails. The key repeats. Every letter enciphered by the same key letter forms one monoalphabetic substitution, so if you know the key's length you have as many simple substitutions as the key is long, and each falls to frequency analysis.

And the key's length is findable. Repeated sequences in the ciphertext are usually the same plaintext enciphered at the same key offset, so the distances between repetitions are multiples of the key length. That is the Kasiski observation and [Breaking a Classical Cipher] is where it belongs.

The one-time pad

The method. A keystream of truly random symbols, as long as the message, combined with the plaintext.

Why it is unbreakable. For any ciphertext and any plaintext of the same length, there is a key that maps one to the other. So the ciphertext gives no information at all about which plaintext was sent, and that is a proof rather than a claim about difficulty.

Why it is impractical. Three conditions, and all three must hold.

The key must be truly random. Not generated by an algorithm from a short seed, which would make it a stream cipher and breakable.

The key must be as long as the message. So distributing it is as hard as sending the message securely, which is the problem being solved.

And the key must never be reused. Two messages enciphered with the same pad can be combined to eliminate the key, and both are then recoverable.

So it is used where the key can be distributed in advance in bulk and the traffic is low, and nowhere else.

What the sequence does NOT show

Not that each cipher was broken and then replaced. The history is messier: methods were used long after they were breakable, and some were broken secretly.

munotes.in349

Cipher Algorithms, Classical to Modern

Not that AES is the end. It is what is used now. A cipher is secure until it is not, and the sequence has no reason to stop.

And not that the classical ciphers are useless to study. They are how the ideas of confusion, diffusion and rounds became visible, and every one of them is a clear instance of one idea missing.

Quick revision

  • The sequence: Caesar, keyword substitution, Vigenère, rail fence, columnar, one-time pad, DES, AES. Each repaired a specific weakness.
  • Three ideas develop through it: confusion, the key's relation to the ciphertext; diffusion, the spreading of each plaintext symbol's influence; and rounds.
  • Classical ciphers have one round and no diffusion. A modern block cipher changes about half the ciphertext bits when one plaintext bit changes.
  • Vigenère uses a different shift per position and fails because the key repeats, giving as many simple substitutions as the key is long.
  • The one-time pad is provably unbreakable and requires a truly random key, as long as the message, never reused.
  • DES has 16 rounds on 64-bit blocks with a 56-bit key; AES has 10 to 14 rounds on 128-bit blocks with keys of 128, 192 or 256 bits.

Test yourself

1. Encipher ATTACK with the Vigenère key ARTHA and explain what the method fixed.

The key letters A, R, T, H, A give shifts of 0, 17, 19, 7 and 0, so the ciphertext is A K M H C K. It fixed the single frequency profile of a monoalphabetic substitution, because the same plaintext letter is enciphered differently according to its position.

2. Define confusion and diffusion and say which classical ciphers have which.

Confusion is a complicated relationship between the key and the ciphertext; diffusion is the spreading of one plaintext symbol's influence across many ciphertext symbols. Classical substitutions have a little confusion and no diffusion; transpositions move positions without spreading influence; modern block ciphers have both, by rounds of substitution and permutation.

3. Why is the one-time pad unbreakable, and why is it impractical?

Because for any ciphertext there is a key making it decipher to any plaintext of the same length, so the ciphertext carries no information about which was sent. It is impractical because the key must be truly random, as long as the message, and never reused, so distributing it is as hard as sending the message securely.

4. Why does Vigenère fall to frequency analysis despite having no single frequency profile?

Because the key repeats, so the letters at positions enciphered by the same key letter form a monoalphabetic substitution. Once the key length is known, the ciphertext splits into that many simple substitutions and each falls to frequency analysis.

Contents This chapter on its own page

munotes.in350

Chapter Ninety-Seven

Breaking a Classical Cipher

Syllabus topic Module 2, "Cipher algorithms", "Symmetric encryption"

In one line

Frequency analysis breaks a Caesar cipher in twenty-six tries and a Vigenère cipher in rather more, and both are run here rather than described.

In the wording you can write in an examination: a classical substitution cipher is broken by frequency analysis, which compares the letter distribution of the ciphertext with that of the language. For a Caesar cipher every shift is tried and the decipherment whose distribution is closest to the language's is taken. For a polyalphabetic cipher the key length is found first, from the distances between repeated sequences, and the ciphertext is then split into that many monoalphabetic substitutions, each of which is broken separately.

The tool: a distance from English

The idea. English letter frequencies are stable: E about 12.7 per cent, T about 9.1, A about 8.2, and Z about 0.07. A candidate decipherment that is English should have that distribution.

The measure. For each letter, take the difference between the count observed and the count expected, square it, and divide by the expected count. Add over all twenty-six.

score = sum over each letter of (observed - expected)^2 / expected

Low is English-like. A wrong decipherment has a high score because its commonest letters are in the wrong places.

And this is the chi-squared statistic, used here as a distance rather than as a test of significance.

The attacks, run

"""Breaking a Caesar and finding a Vigenere key length, both by running the attack."""
import string
from collections import Counter

A = string.ascii_uppercase
ENGLISH_FREQ = {
    "E": 12.70, "T": 9.06, "A": 8.17, "O": 7.51, "I": 6.97, "N": 6.75, "S": 6.33,
    "H": 6.09, "R": 5.99, "D": 4.25, "L": 4.03, "C": 2.78, "U": 2.76, "M": 2.41,
    "W": 2.36, "F": 2.23, "G": 2.02, "Y": 1.97, "P": 1.93, "B": 1.29, "V": 0.98,
    "K": 0.77, "J": 0.15, "X": 0.15, "Q": 0.10, "Z": 0.07,
}

def caesar(text, shift, decrypt=False):
    if decrypt:
        shift = -shift
    return "".join(A[(A.index(c) + shift) % 26] if c in A else c for c in text.upper())

def vigenere(text, key, decrypt=False):
    key = [c for c in key.upper() if c in A]
    out, i = [], 0
    for c in text.upper():
        if c not in A:
            out.append(c)
            continue
        k = A.index(key[i % len(key)])
        out.append(A[(A.index(c) + (-k if decrypt else k)) % 26])
        i += 1
    return "".join(out)

def chi_squared(text):
    """How far the letter counts are from English's. Lower is more English-like."""
    letters = [c for c in text if c in A]
    n = len(letters)
    counts = Counter(letters)
    return sum((counts.get(c, 0) - n * ENGLISH_FREQ[c] / 100) ** 2
               / max(n * ENGLISH_FREQ[c] / 100, 1e-9) for c in A)

PLAIN = ("THE ENEMY IS AT THE GATE AND WE MUST SEND WORD TO THE CAPITAL AT ONCE "
         "FOR THE GATE WILL NOT HOLD THROUGH THE NIGHT AND THE RIVER IS RISING")
SHIFT = 11
CT = caesar(PLAIN, SHIFT)

print("BREAKING A CAESAR CIPHER")
print("  ciphertext  %s" % CT[:66])
print()
print("  every shift scored against English letter frequencies, best five")
scores = sorted(((chi_squared(caesar(CT, s, decrypt=True)), s) for s in range(26)))
for score, s in scores[:5]:
    print("    shift %2d  score %8.2f  %s" % (s, score, caesar(CT, s, decrypt=True)[:44]))
print()
best = scores[0][1]
print("  best shift: %d, and the true shift was %d" % (best, SHIFT))
print("  recovered:  %s" % caesar(CT, best, decrypt=True)[:66])
print("  correct:", caesar(CT, best, decrypt=True) == PLAIN.upper())

print()
print("HOW MUCH CIPHERTEXT THE ATTACK NEEDS")
print("  letters  best shift  correct")
for n in (10, 20, 40, 80, 120):
    piece = "".join(c for c in CT if c in A)[:n]
    s = min(range(26), key=lambda k: chi_squared(caesar(piece, k, decrypt=True)))
    print("  %-8d %-11d %s" % (n, s, "yes" if s == SHIFT else "no"))

print()
print("FINDING A VIGENERE KEY LENGTH BY REPEATED SEQUENCES")
VKEY = "ARTHA"
VCT = "".join(c for c in vigenere(PLAIN, VKEY) if c in A)
seqs = {}
for i in range(len(VCT) - 2):
    seqs.setdefault(VCT[i:i + 3], []).append(i)
gaps = []
for s, positions in seqs.items():
    if len(positions) > 1:
        for a, b in zip(positions, positions[1:]):
            gaps.append((s, b - a))
print("  repeated three-letter sequences and the gaps between them")
for s, g in gaps[:8]:
    print("    %s  gap %d" % (s, g))
if gaps:
    from math import gcd
    from functools import reduce
    g = reduce(gcd, [x for _, x in gaps])
    print("  greatest common divisor of the gaps: %d" % g)
    print("  the true key length was %d" % len(VKEY))
else:
    print("  no repeated sequences in this sample")
munotes.in351

Breaking a Classical Cipher

BREAKING A CAESAR CIPHER
  ciphertext  ESP PYPXJ TD LE ESP RLEP LYO HP XFDE DPYO HZCO EZ ESP NLATELW LE Z

  every shift scored against English letter frequencies, best five
    shift 11  score    15.74  THE ENEMY IS AT THE GATE AND WE MUST SEND WO
    shift 17  score   432.76  NBY YHYGS CM UN NBY AUNY UHX QY GOMN MYHX QI
    shift 23  score   448.99  HVS SBSAM WG OH HVS UOHS OBR KS AIGH GSBR KC
    shift  4  score   494.69  AOL LULTF PZ HA AOL NHAL HUK DL TBZA ZLUK DV
    shift 13  score   501.47  RFC CLCKW GQ YR RFC EYRC YLB UC KSQR QCLB UM

  best shift: 11, and the true shift was 11
  recovered:  THE ENEMY IS AT THE GATE AND WE MUST SEND WORD TO THE CAPITAL AT O
  correct: True

HOW MUCH CIPHERTEXT THE ATTACK NEEDS
  letters  best shift  correct
  10       11          yes
  20       11          yes
  40       11          yes
  80       11          yes
  120      11          yes

FINDING A VIGENERE KEY LENGTH BY REPEATED SEQUENCES
  repeated three-letter sequences and the gaps between them
    MOE  gap 45
    MOE  gap 25
    OEG  gap 45
    EGR  gap 45
    GRM  gap 45
    RML  gap 45
    NUP  gap 10
    AHE  gap 55
  greatest common divisor of the gaps: 5
  the true key length was 5
munotes.in352

Breaking a Classical Cipher

Reading the output

The correct shift scores 15.74 and the next best scores 432.76. That is not a close call: the right answer is a factor of twenty-seven better than the runner-up. A Caesar cipher does not merely fall; it falls unmistakably, which is why nobody has to guess.

And the attack succeeds on ten letters in this sample, which is worth a caution: this plaintext is unusually typical English and a ten-letter sample of, say, a list of proper names would defeat it. The general rule is that frequency analysis needs enough text, and how much depends on the text.

The Kasiski block finds the key length exactly. The three-letter sequence MOE appears twice at a distance of 45 and again at 25; other repeats are at 45, 10 and 55. The greatest common divisor of all the gaps is 5, and the key ARTHA has five letters.

Why that works. A repeated plaintext sequence enciphered at the same offset within the key produces a repeated ciphertext sequence. So the distance between two such repeats is a multiple of the key length, and the common divisor of several distances is the key length or a multiple of it.

And what it does not do. It gives the key's LENGTH, not the key. The next step is to split the ciphertext into five sequences, each enciphered by one key letter, and break each as a Caesar. Finding the length is the hard part; the rest is five easy problems.

What a monoalphabetic substitution costs an attacker

Caesar. Twenty-six trials. Seconds by hand.

Keyword substitution. The key space is 26 factorial and it is irrelevant, because the attacker does not search it. They use the frequency profile to fix the commonest letters, then word patterns to fix the rest, then read the message and correct as they go. An experienced solver does a monoalphabetic substitution of a few hundred letters in minutes.

Transposition. The letter counts are the language's, so frequency analysis says only that it IS a transposition. The attack is anagramming, and for a rail fence it is trying each rail count, which is a handful of trials.

And a substitution composed with a transposition. Harder, and not hard. Strip the transposition by anagramming for common letter pairs, which the substitution has disguised but not removed, or attack them together with a search over the small key space that a memorable key implies.

munotes.in353

Breaking a Classical Cipher

The verdict on an Arthaśāstra-style system

This is the section that has to be written plainly.

Whatever gūḍhalekhya was, it was a hand method usable by an untrained agent in the field. [Secret Communication: What the Text Actually Says] establishes those constraints from the text.

A hand method usable by an untrained agent is a simple substitution, a simple transposition, or a composition of the two. There is nothing else available without instruments.

And every one of those falls to the attacks in this chapter. Against an adversary who knows that secret writing is in use, who obtains a few hundred letters of it, and who knows the language, the system does not hold.

Say that, and then say the two things that keep it fair.

The attacks are not obvious. Systematic frequency analysis is first described, in the surviving record, centuries after the Arthaśāstra. A method is only weak against an adversary who has the analysis, and for most of the period there was none.

And the threat model was different. The adversary in the text is a guard at a gate searching a mendicant woman, not a cryptanalyst with a corpus. Against that adversary, concealment works and a cipher is not even necessary, which is presumably why the passage offers three channels as equals.

The modern lesson

Security is relative to an adversary, and the adversary changes. A system adequate against one is useless against another, and the second one arrives without warning.

So a design states its threat model. The Arthaśāstra's is recoverable from its passages and is stated in [Secret Communication: What the Text Actually Says]. A modern design writes it down.

And this is why a cipher must be published and attacked. Kerckhoffs's principle, from [Symmetric Encryption], is the same point: a method nobody has attacked is a method nobody knows the strength of. Attacking your own cipher is the only way to find out what it is worth, which is why this chapter exists and why the program runs the attack rather than describing it.

Quick revision

  • Frequency analysis scores a candidate decipherment by its distance from the language's letter distribution; low is English-like.
  • The correct Caesar shift scored 15.74 against 432.76 for the next best, so the answer is unmistakable.
  • The attack needs enough text, and how much depends on the text.
  • Kasiski: distances between repeated ciphertext sequences are multiples of the key length; the greatest common divisor of the gaps gave exactly 5 for a five-letter key.
  • Finding the key length is the hard part; the ciphertext then splits into that many Caesar ciphers.
  • A transposition is identified by its normal frequency profile and attacked by anagramming.
  • An Arthaśāstra-style hand system would not survive an adversary with frequency analysis and a few hundred letters; against the adversary the text actually describes, it would.
munotes.in354

Breaking a Classical Cipher

Test yourself

1. How is a Caesar cipher broken, and how decisively?

By deciphering with every one of the 26 shifts and scoring each result against the language's letter frequencies. Decisively: in the worked example the correct shift scored twenty-seven times better than the next best.

2. State the Kasiski observation and what it gives you.

A repeated plaintext sequence enciphered at the same offset within the key produces a repeated ciphertext sequence, so the distance between repeats is a multiple of the key length. The greatest common divisor of several such distances gives the key length, though not the key itself.

3. Why does a large key space not protect a monoalphabetic substitution?

Because the attacker does not search the key space. They read the mapping off the ciphertext's frequency profile and word patterns, so the number of possible keys is irrelevant to the work required.

4. Give the verdict on an Arthaśāstra-style cipher, and the two qualifications that keep it fair.

Against an adversary with frequency analysis, the language, and a few hundred letters of ciphertext, it would not hold. The qualifications are that systematic frequency analysis is first described centuries later, so for most of the period no such adversary existed; and that the threat model in the text is a guard searching a carrier, against whom concealment is sufficient.

Contents This chapter on its own page

munotes.in355

Chapter Ninety-Eight

Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System

Syllabus topic Module 2, "Arthaśāstra Cryptography as Secure Communication Model", "Substitution and transposition systems", "Algorithm Specification (Pseudo-code)", "minimum 10 test cases"

In one line

A substitution composed with a transposition under one key, meeting the requirement the Arthaśāstra states and implementing no method it gives.

In the wording you can write in an examination: an Arthaśāstra-inspired symmetric cipher is a modern construction designed to the threat model recoverable from the treatise's passages on secret communication: a message crossing hostile territory, carried by an agent who may be searched and need not be trusted, recoverable by a recipient without further communication. It is built from the two classical primitives, substitution and transposition, composed under a single shared key.

Problem statement, in MU's own form

IKS concept as CS concept: Arthaśāstra-inspired cryptography as a symmetric cipher system.

Statement. Design and implement a symmetric cipher meeting the requirement stated in Arthaśāstra Book I, Chapter XII, where information is to be conveyed through cipher-writing. Since the treatise names gūḍhalekhya and supplies no method, construct one from the classical primitives: a monoalphabetic substitution generated from the key, followed by a columnar transposition under the same key. Demonstrate correctness by round-tripping, analyse the key space, and state what the system does not provide.

Conceptual mapping table

Classical elementComputer science element
gūḍhalekhya, cipher-writinga symmetric cipher
a secret agreed before the messagethe shared key
a carrier who need not be trustedan untrusted transport
a searchable carrieran adversary who obtains the ciphertext
the recipient reading it without conversingdecipherment from the key alone
saṃjñā, prearranged signsthe key agreement problem, unsolved by the cipher
the seal of Book XIIIintegrity, which this cipher does not provide
the three alternative channelsredundancy, and three attack surfaces

Algorithm specification, in pseudo-code

ALGORITHM Encrypt(plaintext, key)

body <- plaintext with every non-letter removed and every letter upper-cased

cols <- the number of letters in the key

body <- body padded with X until its length is a multiple of cols

body <- Substitute(body, key)

return Transpose(body, key)

ALGORITHM Decrypt(ciphertext, key)

return Unsubstitute(Untranspose(ciphertext, key), key)

ALGORITHM Substitute(text, key)

cipher <- the key's distinct letters, then the remaining letters in order

map each plain letter to the letter at the same position in cipher

ALGORITHM Transpose(text, key)

write the text in rows of width cols

read the columns in the order given by sorting the key's letters,

breaking ties by position

Three design points to defend in a viva.

The padding is applied BEFORE the substitution. [Substitution and Transposition in Code] records why: padding afterwards means the inverse substitution turns the pad into a different letter and the message round-trips wrongly.

The key's letters are sorted with position as the tie-break, so a repeated letter gives a reproducible column order.

And one key serves both steps, which is a simplification and a weakness, stated below.

munotes.in356

Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System

Working code

The implementation is the one printed in full in [Substitution and Transposition in Code], and it is not repeated here. This chapter is the submission built around it.

Test cases

Worked. Twelve, and their result is printed in that chapter. Five are ordinary messages and seven are boundary cases: one letter, a message shorter than the key, punctuation and digits, lower case, a key with a repeated letter, a one-letter key, a key longer than the message, the whole alphabet as the key, a long message, and a message of repeated letters.

All twelve round-trip, with the comparison allowing trailing padding and nothing else.

Complexity

Encipherment. One table lookup per letter for the substitution; one sort of the key and one pass for the transposition. So the work is proportional to the message length plus the key length times its logarithm.

Decipherment. The same.

Space. One copy of the message and the two tables.

And that is what a usable cipher must be: linear in the message, because anything worse becomes impossible on real traffic.

The key space

ComponentKeys
the substitution, from an eight-letter keywordat most about 63 billion distinct alphabets
the transposition, from an eight-letter keyat most 40,320 column orders
and they are not independentbecause the same key produces both

The last row is the finding. With one key, choosing the key chooses both permutations, so the effective key space is the number of distinct keys, not the product of the two spaces. An attacker who recovers the substitution has recovered the transposition, for free.

The fix, which this implementation does not make. Derive two independent subkeys from the key, for instance by using the key for the substitution and a transformation of it for the transposition. That is what a real system does, and saying so is part of the analysis MU asks for.

Limitations, stated

This is her heading seven and it is the part that separates a submission from a demonstration.

It is not Kautilya's cipher. The treatise gives no method. [Secret Communication: What the Text Actually Says] establishes this, and the title MU herself uses is "inspired".

It does not survive frequency analysis. [Breaking a Classical Cipher] runs the attack on the primitives. Composition makes the work harder and does not change the outcome against an adversary with a few hundred letters.

One key for two steps. As above: the two permutations are not independent.

No integrity. An altered ciphertext deciphers to altered plaintext with no indication. The Arthaśāstra's own seal is the mechanism for this and the cipher does not provide it.

No authentication. A message that deciphers correctly shows the sender had the key, and the key is shared, so it does not identify which holder sent it.

munotes.in357

Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System

No key distribution. The key must be agreed in advance by some other channel, which is the problem the treatise's prearranged signs also have and also do not solve.

Non-letters are discarded, not preserved. Spaces, digits and punctuation are removed and are not restored on decipherment. That is a declared design decision and it loses information.

And the padding is not removed. The decipherment carries trailing X characters, and a message legitimately ending in X could not be distinguished from a padded one. A real scheme carries an explicit length.

Conclusion

MU's eighth heading. Two sentences are enough and they should be these.

What was built. A symmetric cipher composing the two classical primitives under one key, meeting the requirement the Arthaśāstra states, tested on twelve cases including seven boundary cases, and analysed for key space and cost.

What it establishes. That the treatise's passages supply a threat model and not an algorithm, and that a cipher built to that threat model from the techniques available in its period is adequate against the adversary the text describes and inadequate against an adversary with frequency analysis.

Quick revision

  • The cipher: substitute under a keyword alphabet, then transpose by columns, both under one key.
  • The padding goes in before the substitution, or the inverse substitution corrupts it.
  • The key's letters are sorted with position as a tie-break so a repeated letter is reproducible.
  • One key for both steps means the two permutations are not independent, so breaking one breaks both.
  • Twelve test cases, seven of them boundary cases, all round-tripping.
  • Limitations: not Kautilya's cipher; falls to frequency analysis; one key for two steps; no integrity, authentication or key distribution; non-letters discarded; padding not removed.

Test yourself

1. Why is the title "Arthaśāstra-inspired" rather than "the Arthaśāstra's cipher"?

Because the treatise names cipher-writing four times and supplies no method, no key and no key distribution. There is nothing of Kautilya's to implement, so what is built is a cipher meeting the requirement his passages state.

2. Why does using one key for both steps reduce the effective key space?

Because the substitution alphabet and the column order are both derived from the same key, so they are not independently chosen. The number of distinct systems is the number of distinct keys rather than the product of the two spaces, and recovering the key from one step yields the other.

3. Name three properties the cipher does not provide, and say where the treatise addresses one of them.

Integrity, authentication and key distribution. The treatise addresses integrity with the sealed letter of Book XIII, Chapter I, which makes tampering evident without hiding anything.

munotes.in358

Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System

4. State the conclusion this implementation supports.

That the Arthaśāstra supplies a threat model rather than an algorithm, and that a cipher built to that threat model from the techniques of its period is adequate against the adversary the text describes, a guard searching a carrier, and inadequate against an adversary with frequency analysis and a few hundred letters of ciphertext.

Contents This chapter on its own page

munotes.in359

Chapter Ninety-Nine

Secure Protocol Abstraction

Syllabus topic Module 2, "Secure protocol abstraction"

In one line

A protocol is who sends what to whom, in what order, and what each party may conclude at the end.

In the wording you can write in an examination: a security protocol is a specified sequence of messages between named parties, together with the goal the exchange is to achieve, the assumptions about what each party holds before it begins, and the threat model stating what an adversary can do. A cipher is a component of a protocol and is not a protocol, because a cipher specifies no exchange.

The four parts

The parties. Who is involved, and what each of them knows at the start.

The messages. What is sent, by whom, to whom, in what order.

The goal. What is to be true when the exchange completes, and for whom.

The threat model. What the adversary can do: read messages, alter them, delete them, insert new ones, or take part while pretending to be somebody else.

All four, or it is not a protocol. A description of messages with no stated goal cannot be assessed, because there is no claim to test.

The notation

Protocols are written as numbered lines, each naming the sender, the recipient and the content.

1. A -> B : M

2. B -> A : N

Line 1 reads: A sends M to B. Encipherment is written with the key in braces.

1. A -> B : {M}k

That is: A sends B the message M enciphered under the key k.

Worked: the Arthaśāstra's courier arrangement as a protocol

The passage in Book I, Chapter XII has the information conveyed by a doorkeeper, an artisan or a court-bard, through cipher-writing, to its destination.

The parties. S, the agent with the information. C, the carrier. R, the recipient at the institute.

What they hold beforehand. S and R share a key k. C holds nothing.

The exchange.

1. S -> C : {M}k, in a form the pretext explains

2. C -> R : {M}k

The goal. R learns M, and nobody else does.

The threat model, from the text. An adversary may stop and search C. The passage's whole premise is the mendicant woman stopped at the entrance.

The assumptions. That k was agreed before the exchange began; that C does not know k; and that C delivers to R rather than to somebody else.

Assessing it

Does it meet its goal under its threat model? If the adversary searches C and obtains the ciphertext, they learn nothing without k. So under a threat model of "search the carrier", yes.

And under a wider threat model? Four failures, and naming them is the exercise.

C may be substituted. Nothing in the exchange lets R tell that the message came from S. An adversary who replaces C, or who turns C, can deliver a message of their own enciphered under a key they do not have, which R will fail to decipher, or, if the adversary has k, anything they like. There is no authentication.

munotes.in360

Secure Protocol Abstraction

The message may be altered. An adversary who alters the ciphertext produces a decipherment R cannot detect as altered, because there is no integrity check. The seal of Book XIII would supply one and this exchange does not use it.

The message may be deleted. C may be intercepted and the message simply not delivered, and neither S nor R learns that it failed. There is no acknowledgement.

And the exchange may be replayed. An adversary who records a message and delivers it again later gives R the same message twice, and R cannot tell. There is no freshness.

The four properties a protocol usually wants

PropertyWhat it meansWhat supplies it
confidentialityonly the intended party learns the contentencipherment
authenticationeach party knows who the other isa shared secret used to prove identity, or a signature
integrityalteration is detecteda message authentication code, or a signature
freshnessthe message belongs to this exchange and is not a replaya nonce, a counter, or a timestamp

The classical arrangement supplies the first and none of the others. That is not a criticism of the text; it is the analysis MU's label "secure protocol abstraction" asks for.

Repairing the protocol

The exercise a question may set: add what is missing.

Add integrity and authentication. S sends the ciphertext together with a value computed from the message and the key that only a holder of k can produce.

1. S -> C : {M}k, MAC(M, k)

2. C -> R : {M}k, MAC(M, k)

Now R recomputes the value and compares. A mismatch means the message was altered or was not from a holder of k.

Add freshness. Include a number that has not been used before.

1. S -> C : {M, n}k, MAC(M, n, k)

R checks that n has not been seen. A replay is then detected.

And note what is still missing. How S and R agreed k. No exchange in the protocol establishes it, and that is the key distribution problem of [Symmetric Encryption], which this protocol assumes away.

What a protocol abstraction is NOT

It is not an implementation. The notation says what is sent, not how it is encoded, and encoding mistakes are a large source of real failures.

It is not a proof. Writing a protocol down does not show it meets its goal. Protocols that looked obviously correct have been shown wrong decades later, which is why formal analysis of protocols is a field.

munotes.in361

Secure Protocol Abstraction

And it is not complete without the threat model. A protocol is secure against an adversary, never in the abstract, and a protocol whose threat model is unstated cannot be assessed at all.

Quick revision

  • Four parts: the parties and what they hold, the messages in order, the goal, and the threat model.
  • Notation: a numbered line, sender, arrow, recipient, colon, content; encipherment written with the key in braces.
  • The Arthaśāstra's courier arrangement written out has two messages, a shared key held by sender and recipient, and a carrier who holds nothing.
  • It supplies confidentiality under a threat model of searching the carrier, and fails on authentication, integrity, delivery and replay.
  • The four properties a protocol usually wants: confidentiality, authentication, integrity, freshness.
  • Repairs: a message authentication code for the first two, a nonce for the third. Key agreement is assumed and not achieved.
  • A protocol is secure against a stated adversary, never in the abstract.

Test yourself

1. Name the four parts of a protocol specification and say why the threat model cannot be omitted.

The parties and their prior knowledge, the ordered messages, the goal, and the threat model. Without the threat model there is no statement of what the adversary can do, so the claim that the goal is achieved cannot be assessed at all.

2. Write the Arthaśāstra's courier arrangement in protocol notation.

S sends C the message enciphered under the key shared with R, in a form the pretext explains; C sends that ciphertext to R. S and R hold the key beforehand and C holds nothing.

3. Give three ways the classical arrangement fails under a wider threat model.

The carrier may be substituted and R cannot tell who sent the message, so there is no authentication. The ciphertext may be altered and R cannot detect it, so there is no integrity. The message may be recorded and delivered again, and R cannot tell, so there is no freshness.

4. How is integrity added, and what remains unsolved?

By sending, with the ciphertext, a value computed from the message and the key that only a key holder can produce, which the recipient recomputes and compares. What remains unsolved is how the two parties came to share the key, which no message in the protocol establishes.

Contents This chapter on its own page

munotes.in362

Chapter One Hundred

Foundations of Cybersecurity

Syllabus topic Module 2, "Foundations of cybersecurity"

In one line

Security is not one property: keep it secret, keep it unaltered, keep it available, know who you are talking to, know what they may do, and be able to prove what happened.

In the wording you can write in an examination: the foundations of information security are commonly stated as three core properties, confidentiality, integrity and availability, together with three further properties required by any system with users, namely authentication, authorisation and non-repudiation. A security design states which properties it requires, against which adversary, and by which mechanism each is achieved.

The three core properties

Confidentiality

What it means. Only those authorised may read the information.

How it is achieved. Encipherment, and access control.

How it fails. Interception, a stolen key, or an authorised party disclosing it.

In the Arthaśāstra. Gūḍhalekhya, cipher-writing, and the several forms of concealment. The property the treatise addresses most.

Integrity

What it means. The information has not been altered, and alteration is detectable.

How it is achieved. A message authentication code, a digital signature, or a checksum against accidental change.

How it fails. An adversary alters data in transit or at rest and nobody notices.

In the Arthaśāstra. The seal on the letter carried by pigeon. A tamper-evident device, which is exactly what integrity requires, and [Information Protection Mechanisms] insists on not confusing it with confidentiality.

Availability

What it means. The information and the service are there when needed.

How it is achieved. Redundancy, capacity, and resistance to denial of service.

How it fails. The channel is cut, the service is overwhelmed, or the only copy is destroyed.

In the Arthaśāstra. The three alternative channels offered in one sentence of Book I, Chapter XII. Redundancy of route is an availability mechanism, and it is the only one the passages contain.

The three further properties

Authentication

What it means. Each party knows who the other is.

How it is achieved. Something known, held, or inherent: a secret, a token, or a biometric. In protocols, by proving possession of a key.

In the Arthaśāstra. Weakly, and organisationally. The writer of a royal writ is appointed on stated qualifications, which controls who may compose one, and nothing in the message proves its origin. The property the treatise addresses least.

Authorisation

What it means. A party, once identified, may do some things and not others.

How it is achieved. Access control, and the principle of least privilege.

In the Arthaśāstra. Compartmentalisation: officers have their own spies, so an agent's reach is bounded by their chain. That is least privilege as an organisational structure.

Non-repudiation

What it means. A party cannot later deny having sent what they sent.

How it is achieved. A digital signature, which only the sender could have produced.

munotes.in363

Foundations of Cybersecurity

In the Arthaśāstra. Not at all, and it could not be. Non-repudiation requires asymmetric cryptography, because with a shared secret either holder could have produced the evidence. This is the one property of the six that is not merely absent from the treatise but unattainable with its techniques.

The six in one table

PropertyWhat it protectsMechanismIn the treatise
confidentialitywho may readencipherment, concealmentgūḍhalekhya and the pretexts
integritythat it is unalteredMAC, signaturethe seal
availabilitythat it is thereredundancythree alternative channels
authenticationwho you are talking toproof of a secretorganisational only
authorisationwhat they may doleast privilegeofficers with their own spies
non-repudiationthat they cannot deny ita signatureabsent, and unattainable

Read the last column. Four of the six are addressed by some mechanism, one only organisationally, and one is beyond the period's techniques. That is a fair and specific verdict, and it is better than either "the Arthaśāstra had cybersecurity" or "it had nothing".

Where the properties conflict

The part of the subject a beginner misses: the six are not all compatible.

Worked. Confidentiality against availability. Encipher everything and lose the key, and the data is unavailable. Every backup of a key is an additional place it can be stolen from.

Integrity against availability. A system that refuses anything it cannot verify is unavailable whenever verification fails, which may be for innocent reasons.

Authentication against availability. A system that demands strong proof of identity locks out legitimate users who have lost their credential.

And non-repudiation against confidentiality, in a specific way: evidence that a party sent something is evidence that can be shown to a third party, which is the opposite of what a confidential conversation wants.

So a design chooses, and stating which properties were prioritised and why is what a security argument consists of.

The threat model, one more time

Every one of the six is relative to an adversary, and this is the point [Breaking a Classical Cipher] establishes and this chapter closes on.

Name the adversary. What can they observe, what can they alter, what can they compute, and what do they already know?

Then state which properties must hold against them. Not all properties against all adversaries: that is unachievable and it is not what security means.

And the Arthaśāstra does this implicitly. Its adversary is a guard at a gate, a rival's agent, and a turned member of one's own service. Against those three, the arrangements in the text are addressed to confidentiality, integrity and the limitation of damage, which is a coherent design.

Quick revision

  • Three core properties: confidentiality, integrity, availability. Three further: authentication, authorisation, non-repudiation.
  • Confidentiality by encipherment and concealment; integrity by a MAC or signature; availability by redundancy.
  • Authentication by proof of a secret; authorisation by least privilege; non-repudiation by a signature.
  • In the treatise: gūḍhalekhya, the seal, three alternative channels, organisational appointment, officers with their own spies, and nothing for non-repudiation.
  • Non-repudiation requires asymmetric cryptography, so it is unattainable with a shared secret.
  • The properties conflict: confidentiality against availability, integrity against availability, authentication against availability.
  • Every property is relative to a named adversary, and a design states which properties must hold against which.
munotes.in364

Foundations of Cybersecurity

Test yourself

1. Name the six properties and the mechanism that supplies each.

Confidentiality by encipherment and concealment; integrity by a message authentication code or signature; availability by redundancy and capacity; authentication by proof of possession of a secret; authorisation by access control and least privilege; non-repudiation by a digital signature.

2. Which property could not be achieved with the Arthaśāstra's techniques, and why?

Non-repudiation. It requires that only one party could have produced the evidence, which needs asymmetric cryptography; with a shared secret either holder could have produced it, so neither can be held to it.

3. Give two pairs of properties that conflict, with the reason.

Confidentiality against availability: enciphering everything risks losing the key and with it the data, and every key backup is another place it can be stolen. Authentication against availability: demanding strong proof of identity excludes legitimate users who have lost their credential.

4. Which mechanism in the treatise supplies integrity, and what is the standing error about it?

The seal on the letter carried by pigeon. The standing error is to treat a seal as confidentiality: it makes tampering evident and hides nothing from anyone willing to break it.

Contents This chapter on its own page

munotes.in365

Chapter One Hundred One

The Internal Assessment: Building and Presenting the Implementation

Syllabus topic Module 2, "Algorithm Specification (Pseudo-code)", "Complexity & Limitations", "Conceptual Mapping Table", "minimum 10 test cases"

In one line

One implementation, eight headings, ten test cases, a live demo and a viva, for twenty of the fifty marks.

In the wording you can write in an examination: the internal assessment for this paper is a single implementation, undertaken individually or in a pair, of an Indian Knowledge Systems concept as a computer science concept. It must comprise a problem statement in that form, a formal specification of the rules or algorithm, working code, a minimum of ten test cases, and a short analysis of correctness and limitations, and it is presented through a live demonstration and a viva.

Her eight topics, and where each is built in this book

Her topicThe chapter
Piṅgala's Chandaḥśāstra as Binary Encoding and Combinatorial Generation[Prastāra in Code]
Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine[Pāṇini's Aṣṭādhyāyī as a Rule-Based Grammar Engine]
Nyāya Logic as an Inference Engine[Nyāya Logic as an Inference Engine]
Ayurvedic Classification as a Rule-Based Expert System[Ayurvedic Classification as a Rule-Based Expert System]
Arthaśāstra-Inspired Cryptography as Symmetric Cipher System[Arthaśāstra-Inspired Cryptography as a Symmetric Cipher System]
Meru-Prastāra as Pascal Triangle and Dynamic Programming Model[Meru-Prastāra as a Dynamic Programming Model]
Padārtha Ontology (Nyāya) as Knowledge Representation Model[Padārtha Ontology as a Knowledge Representation Model]
Śāstra Rule Precedence as Deterministic Finite Rewrite System[Śāstra Rule Precedence as a Deterministic Finite Rewrite System]

And you may choose another relevant topic, which her own wording permits.

Her eight headings, and what each is for

HeadingWhat it must containThe commonest mistake
1. Titlein the form "IKS Concept as CS Concept"a title that names only one side
2. Problem Statementwhat is to be built, and what it must dodescribing the classical material instead
3. Conceptual Mapping Tableeach classical element against its computer science counterpartomitting it entirely
4. Algorithm Specificationpseudo-code, not codepasting the code again
5. Working Codethe program, runningcode that has never been run on another machine
6. Test Casesat least ten, with expected and actualten cases that all succeed
7. Complexity and Limitationsthe cost, and what it does not doclaiming it has no limitations
8. Conclusionwhat was built and what it establishesrepeating the problem statement

Two of those deserve emphasis.

Heading three is where the marks for understanding are. The mapping table is the only place the submission shows that the classical side has been understood rather than summarised. It is also the heading most often missing.

And heading six should include failures. A test set in which every case succeeds has not tested the failure behaviour. Every implementation in this book includes cases that must return nothing, and says how many.

munotes.in366

The Internal Assessment: Building and Presenting the Implementation

Worked: a complete submission

What follows is one submission, short, in her order, so that the shape is visible. It uses the simplest of the eight topics deliberately: a student who can produce this for the prastāra can produce it for any of them.

1. Title

Piṅgala's Prastāra as Exhaustive Enumeration of a Binary Space.

2. Problem Statement

The eighth chapter of Piṅgala's Chandaḥśāstra states a rule which, applied repeatedly from a starting row, produces every possible pattern of light and heavy syllables for a metre of a given length, in a fixed order. Implement that rule, verify that its output is complete and free of repetition for every metre up to twelve syllables, and verify it against an independent construction derived from the same sūtra.

3. Conceptual Mapping Table

Classical elementComputer science element
a syllable, laghu or gurua symbol from a two-letter alphabet, or one bit
a metre of n syllablesa string of length n
the prastārathe enumeration of every such string
the next-row rule of sūtra 8.22a successor function
the all-guru first rowthe initial state
the all-laghu last rowthe termination condition
Halāyudha's "again and again until the desired prastāra"a loop, or a recursion with a base case

4. Algorithm Specification

ALGORITHM Prastara(n)

INPUT n, a positive integer

OUTPUT every pattern of n syllables, in Pingala's order

row <- n copies of GURU

rows <- [row]

while row contains a GURU do

k <- the position of the first GURU, counting from the left

row <- (k-1 GURUs) + LAGHU + (the rest of row, unchanged)

rows <- rows + [row]

end while

return rows

5. Working Code

The program of [Prastāra in Code], run on three Python interpreters.

6. Test Cases

Ten, in two groups. Five check the size of the output against Halāyudha's own figures: 2, 4, 8, 64 and 4096 rows for one, two, three, six and twelve syllables. Five check its shape for every n from one to twelve: the first row is all guru, the last is all laghu, no row repeats, every row has n syllables, and the iterative construction agrees with the recursive one.

All ten pass, and the fifth of the second group is the one that found a real defect: a recursive construction that appended the new syllable on the wrong side produced a table of the correct size with no repeats, and only the comparison detected the wrong order.

7. Complexity and Limitations

Time and space are both proportional to n times 2 to the power n, which is the size of the output and therefore cannot be improved by any program that prints the whole table.

Limitations. It cannot produce a single distant row without producing everything before it; naṣṭa is the alternative and is a separate implementation. It holds the whole table in memory, which is fine to twelve syllables and impossible at thirty. And the recursive construction builds every intermediate table, so it does about twice the work.

munotes.in367

The Internal Assessment: Building and Presenting the Implementation

8. Conclusion

The rule stated in the eighth chapter of the Chandaḥśāstra, as Halāyudha explains it, is an algorithm in the strict sense: it has a stated input, every step has exactly one result, it terminates after exactly 2 to the power n minus 1 steps, every step is executable by hand, and it produces every pattern once. Implemented directly it agrees with an independent construction from the same sūtra on every metre up to twelve syllables.

The live demonstration

Have it running before you sit down. A demonstration that begins with an installation is a demonstration that has already gone wrong.

Show a success and a failure. Run an ordinary case, then a case the program refuses, and say why it refuses.

Show the test suite running. It takes seconds and it is the strongest thing you can show.

And have the classical source open. Being able to point at the sūtra or the verse your program implements is what makes it an IKS submission rather than a programming exercise.

The viva

Twelve questions of the kind that will be asked, across the eight topics.

On the classical side. What text is this from, and what does it say? Who wrote the commentary you are relying on? What does the text NOT say that your program assumes?

On the mapping. Which part of your program corresponds to which part of the text? Where does the correspondence break down?

On the code. Why did you choose this data structure? What happens on an empty input? Which line implements the rule you quoted?

On the tests. Which of your tests would fail if you deleted this line? How many of your tests check a failure rather than a success? What is a case your tests do not cover?

On the limitations. What is the cost of your program, and why? What would you do differently with more time?

The hardest of them is "what does the text NOT say that your program assumes?" Every one of the eight topics has such an assumption, and every chapter of this book that builds one states it. A student who can answer that question has understood the whole point of the paper.

Quick revision

  • One implementation, alone or in a pair, for twenty of the fifty marks.
  • Five required items: the problem statement in the form IKS concept as CS concept, a formal specification, working code, at least ten test cases, and an analysis of correctness and limitations.
  • Eight required headings: title, problem statement, conceptual mapping table, algorithm specification in pseudo-code, working code, ten test cases, complexity and limitations, conclusion.
  • Heading three is where the understanding is shown and is the one most often missing.
  • Heading six should include cases that must FAIL.
  • Demonstration: have it running, show a success and a refusal, run the tests, and have the source text open.
  • The hardest viva question is what the text does not say that your program assumes.
munotes.in368

The Internal Assessment: Building and Presenting the Implementation

Test yourself

1. List MU's eight required headings in order.

Title in the form IKS Concept as CS Concept; problem statement; conceptual mapping table; algorithm specification in pseudo-code; working code; a minimum of ten test cases; complexity and limitations; conclusion.

2. Why should a test set include cases that fail?

Because a set in which every case succeeds has not exercised the failure behaviour at all, so nothing is known about what the program does with bad input, and a refusal that should happen might not.

3. What distinguishes heading four from heading five?

Heading four is pseudo-code: the steps stated in a notation with no syntax to get wrong. Heading five is the program itself. Pasting the code under both headings answers only one of them.

4. Give the viva question that is hardest to answer, and say why.

What the text does not say that your program assumes. It is hardest because it requires knowing both the classical source and your own design well enough to see where the second went beyond the first, which is exactly what the paper is testing.

Contents This chapter on its own page

munotes.in369

Chapter One Hundred Two

Practice for Module II

Syllabus topic Module 2, "Conceptual Mapping Table"

In one line

Q.2 gives you four questions on Module 2 and asks for two, so every answer is worth five marks: a definition, the mechanism, one example, one limit.

Nothing new is taught in this chapter. It is Module II compressed into the shapes the examination uses.

The mapping table for Module 2

Classical elementWhereModern concept
the sixteen categoriesNyāya Sūtra 1.1.1the subject matter of a system of logic and debate
four pramāṇas1.1.3declared sources of evidence
upamāna, comparison1.1.6labelling by similarity to a known instance
vyāpti1.1.5 and the example membera universal, whose negation is a single counterexample
the five members1.1.32 to 1.1.39a universal premise, a singular premise, instantiation, modus ponens
the familiar instanceinside 1.1.36 and 1.1.37evidence for the general premise, grounded in a case
hetvābhāsa1.2.4 to 1.2.9static checks on an argument
chala1.2.10 to 1.2.14attacks on ambiguity, and the strawman named precisely
jāti1.2.18irrelevant objection
nigrahasthānaBook V.2the termination condition of the protocol
vāda, jalpa, vitaṇḍā1.2.1 to 1.2.3three protocols with three win conditions
padārthaVaiśeṣika Sūtra, Sinha's translationan ontology with classes, properties, events, identity and explicit absence
the tridoṣa attribute listsCharaka Sūtrasthāna I.58 to I.60a feature space of eighteen binary attributes
the rule of the adverse attributeI.61a response function with three parameters
parīkṣāVimānasthāna IVa declared order of evidence gathering
prakṛtithe scheme's own labelsmulti-class, or better, multi-label classification
gūḍhalekhyaArthaśāstra I.12a symmetric cipher, whose method the text does not give
the sealXIII.1integrity, not confidentiality
officers with their own spiesI.12least privilege
three alternative channelsI.12redundancy, and three attack surfaces
the royal writII.10a message format with six required qualities and five named faults

Worked answers at five marks, twelve of them

1. State Nyāya Sūtra 1.1.1 and group the sixteen categories.

Supreme felicity is attained by the knowledge about the true nature of sixteen categories: means of right knowledge, object of right knowledge, doubt, purpose, familiar instance, established tenet, members, confutation, ascertainment, discussion, wrangling, cavil, fallacy, quibble, futility, and occasion for rebuke. They fall into four groups: what knowledge is; what starts an enquiry; how an argument is made; and how disputes are conducted and lost. Seven of the sixteen concern failure, which shows the system was built for use against an opponent.

2. Name the four pramāṇas of Nyāya and explain the one Charaka omits.

Perception, inference, comparison and verbal testimony, at 1.1.3. Charaka and the Sāṃkhya school admit three and omit comparison. Upamāna, at 1.1.6, is the knowledge of a thing through its similarity to something previously well known: a man told that a bos gavaeus resembles a cow later sees such an animal, recalls what he was told, and applies the name. What is acquired is the application of a name on the strength of a stated similarity.

munotes.in370

Practice for Module II

3. Define vyāpti and say why its direction matters.

The invariable concomitance of the mark with what is to be established: there is no case of the mark without the conclusion. Direction matters because the mark must be the narrower term. Smoke is never without fire, so fire may be inferred from smoke; fire is often without smoke, as in a red-hot iron ball, so smoke may not be inferred from fire. A single case of the mark without the conclusion destroys the connection.

4. Give the five members with the Sūtra's own illustration and say which does the real work.

Proposition, this hill is fiery. Reason, because it is smoky. Example, whatever is smoky is fiery, as a kitchen. Application, so is this hill, smoky. Conclusion, therefore this hill is fiery. The example does the real work: it states the general connection as an explicit step and anchors it in a case where the connection is known to hold, where an ordinary argument leaves the general premise unstated and therefore unexamined.

5. Compare the Nyāya form with the Aristotelian syllogism.

The Aristotelian form has three statements, a major premise, a minor premise and a conclusion, and is valid in virtue of its form alone. The Nyāya form has five. Its example carries the major premise, its application the minor, and its conclusion the conclusion, so the logical content is the same. The differences are that Nyāya states the claim at the start as well as the end and that its example must carry a familiar instance. So the two answer different questions: whether the conclusion follows, and whether the claim has been established.

6. Name the five fallacies of the reason with a test for each.

Erratic, at 1.2.5: could the reason give the opposite conclusion too? Contradictory, 1.2.6: does it prove the opposite? Equal to the question, 1.2.7: is it the claim reworded? Unproved, 1.2.8: is the reason itself in dispute? Mistimed, 1.2.9: does it hold at the moment it must? The Sūtra's own example of the erratic proves sound both eternal and non-eternal from the same reason, intangibility.

7. Distinguish vāda, jalpa and vitaṇḍā.

Vāda, 1.2.1, is discussion aiming at truth, stated in five members, defended by the means of right knowledge and confutation, without deviating from established tenets. Jalpa, 1.2.2, is wrangling aiming at victory, permitting quibbles and futilities. Vitaṇḍā, 1.2.3, is cavil: wrangling that only attacks and advances no position of its own. The tradition licenses the latter two to guard the truth against an opponent not arguing in good faith.

munotes.in371

Practice for Module II

8. Name the six padārthas and give two modern counterparts.

Substance, attribute, action, genus, species and combination, with non-existence added by the later tradition. Viśeṣa, particularity, corresponds to object identity: it explains what makes two otherwise identical substances two, which is the same problem as two records with identical fields. Abhāva, non-existence, corresponds to an explicitly recorded absence, which is different from an unrecorded value and from a thing not represented at all.

9. State the tridoṣa framework and the problem in Kaviratna's text.

Charaka Sūtrasthāna I.56 gives wind, bile and phlegm as the causes of all bodily diseases, and I.58 to I.60 give seven named attributes for each, with I.61 directing treatment by objects of adverse attributes according to place, measure and time. The printed list for bile contains both "cold" and "hot", which cannot both characterise one thing; and "cold" appears in all three lists, so on that reading it would distinguish nothing. The contradiction is reported rather than silently resolved.

10. Explain how the tridoṣa scheme becomes a classification problem.

The union of the three attribute lists is eighteen attributes, so each doṣa is a binary vector of eighteen components and a description is another. Classification compares the description with the three by counting shared attributes. Sixteen of the eighteen attributes belong to one doṣa only, so most single observations decide. A tie is not an error: a case showing the qualities of two doṣa is a case in which two predominate, which is one of the scheme's own labels.

11. State exactly what the Arthaśāstra says about secret writing.

It names it four times: twice in Book I Chapter XII, once in Book I Chapter XVI, and once in Book XIII Chapter I. The terms are saṃjñā, signs; saṃjñālipi, writing by signs; and gūḍhalekhya, rendered cipher-writing. Agents both write and read such writings. The text supplies no method, no key and no key distribution, which are the three things a cipher system consists of. What it does supply is a threat model: hostile territory, a searchable and untrusted carrier, and a recipient who must read it without conversing.

12. Read the royal writs chapter as a message format specification.

Arthaśāstra Book II, Chapter X prescribes who may compose a writ, what it must contain according to the addressee's rank, six necessary qualities, arrangement, relevancy, completeness, sweetness, dignity and lucidity, and five named faults, clumsiness, contradiction, repetition, bad grammar and misarrangement, each defined. The five faults are a validator running from the medium to the meaning: can it be read, is it well formed, is its structure correct, do its parts agree, is anything said twice. It supplies format and validation and not integrity, authentication or confidentiality.

munotes.in372

Practice for Module II

The one-line facts a viva asks for

  • Nyāya Sūtra: five books of two chapters each, attributed to Gotama. Translation quoted: Vidyabhusana, 1913.
  • 1.1.1 sixteen categories. 1.1.3 four pramāṇas. 1.1.5 inference, three kinds. 1.1.6 comparison. 1.1.7 verbal testimony. 1.1.32 the five members. 1.2.4 the five fallacies. 1.2.10 quibble. 1.2.18 futility.
  • Vaiśeṣika Sūtra 1.1.5: the nine substances. Translation quoted: Nandalal Sinha, 1923.
  • Charaka Saṃhitā: Sūtrasthāna for general principles, Vimānasthāna for method. Translation quoted: Kaviratna, from 1890.
  • Arthaśāstra: fifteen books. Translation quoted: Shamasastry, 1915. The cipher passages are I.12 twice, I.16, and XIII.1. The writs chapter is II.10.
  • Confidentiality, integrity, availability, plus authentication, authorisation, non-repudiation.
  • A seal is integrity. Encipherment is confidentiality. They are not the same property.

Quick revision

  • Q.2 sets four alternatives on Module 2 and asks for two, so an answer is worth five marks, the same length as a Module 1 answer.
  • The sūtra numbers to know: Nyāya 1.1.1, 1.1.3, 1.1.5, 1.1.6, 1.1.32, 1.2.4, 1.2.10, 1.2.18; Vaiśeṣika 1.1.5.
  • The passages to know: Charaka Sūtrasthāna I.56 to I.61 and Vimānasthāna IV; Arthaśāstra I.12, I.16, II.10 and XIII.1.
  • The translations quoted: Vidyabhusana 1913, Nandalal Sinha 1923, Kaviratna from 1890, Shamasastry 1915.
  • The fact most often got wrong: a seal supplies integrity, not confidentiality.

Test yourself

1. Which one classical text supplies the largest number of Module 2's labels, and how many?

The Nyāya Sūtra. Four of MU's six labels under her Nyāya block are items in the list of sixteen categories at 1.1.1: members, which is her five-member syllogism; the means of right knowledge, which carries her anumāna structure; wrangling, cavil, fallacy, quibble, futility and occasion for rebuke, which are her debate methodology; and padārtha, which she takes from the sister school.

2. Which two answers above would serve a question on validation, and why?

The one on the five fallacies and the one on the three kinds of dispute. Together they give the detectable defects and the win conditions, which with the five members and the four pramāṇas are the five parts of a validation protocol.

3. Which fact in this chapter is most often got wrong?

That a seal supplies integrity and not confidentiality. It makes tampering evident and hides nothing from anyone willing to break it.

Contents This chapter on its own page

munotes.in373

Chapter One Hundred Three

The Whole Paper on One Page

Syllabus topic Module 2, "Conceptual Mapping Table", "Knowledge representation", "Inference engines"

In one line

Q.3 asks you to relate something in Module 1 to something in Module 2, and the relations worth knowing are these.

Ten of the thirty external marks come from a question set on both modules together. A student who has revised each module separately can still be caught by it, because the answer is a comparison and neither module contains one.

The whole paper, in one table

Six classical subjects, and the computer science concept MU pairs each with.

ModuleClassical subjectIts central structureMU's modern concepts
1śāstra methoddefined terms, ordered rules, declared evidenceknowledge representation, formal specification, symbolic abstraction
1Piṅgala's Chandaḥśāstratwo-valued symbols, and six operations on the whole space of patternsbinary encoding, recursion, trees, Pascal's triangle, algorithmic generation
1Pāṇini's Aṣṭādhyāyītyped rules with contexts, precedence and phasescontext-free grammar, rewrite systems, automata, parsing, NLP
2Nyāyaa fixed argument form, admitted evidence, and named failurespropositional and predicate logic, inference engines, explainable AI
2Āyurvedic classificationnamed attributes, a small label set, a response ruledecision trees, rule-based and expert systems, feature engineering, multi-class classification
2Arthaśāstraa threat model and four protection mechanismssymmetric encryption, ciphers, steganography, protocols, cybersecurity

The five relations that cross the modules

These are the ones a Q.3 answer is built from.

One: rules, and what happens when two apply

Module 1. Pāṇini's four precedence principles, with sūtra 1.4.2 as the last resort, and asiddhatva as a phase boundary.

Module 2. An inference engine's conflict resolution: specificity, recency, order and refraction. Two of the four are Pāṇini's.

The relation. Both are rule systems large enough that overlap is unavoidable, and both answer it with a stated policy rather than by making the rules disjoint. The difference is that Pāṇini's policy is part of his specification, because his system must produce one form, while a forward-chaining engine may derive everything derivable and needs a policy only for order.

Two: evidence, and what counts as enough

Module 1. Pramāṇa: three admitted means, and Charaka's rule that one is never enough.

Module 2. The four pramāṇas of Nyāya, the five fallacies of the reason, and the debate protocol's admissible support.

The relation. Both modules turn on the idea that a system must declare in advance what it will accept as a reason. The difference is what the declaration is for: in Module 1 it governs an investigation, and in Module 2 it governs a dispute, so Module 2 adds the machinery for naming a bad reason.

Three: representation, and what it makes easy

Module 1. Piṅgala reduces a syllable to one of two values, and every question of his subject becomes arithmetic.

Module 2. Charaka's attribute lists reduce a case to a set of named qualities, and classification becomes a count.

munotes.in374

The Whole Paper on One Page

The relation. Both are abstractions chosen so that the subject's questions become computable. The difference is that Piṅgala's abstraction is complete, so the whole space can be enumerated, while Charaka's is a vocabulary over which only some descriptions are expressible, and [Symptom to Feature Mapping] is about what falls outside it.

Four: explanation, and the form an answer takes

Module 1. A derivation in the Aṣṭādhyāyī is a trace: each step names the rule that fired.

Module 2. A Nyāya argument is five named members, and a decision tree's path is its own explanation.

The relation. Both traditions require that a result be accompanied by the steps that produced it. The difference is the audience: Pāṇini's trace is for a grammarian checking a derivation, and Nyāya's five members are for an opponent trying to defeat it, which is why the Nyāya form has a slot for the evidence for its general premise and the derivation does not.

Five: what the texts do NOT contain

Module 1. No machine, no statement of cost, no base, no generalisation of the saṅkhyā rule.

Module 2. No cipher in the Arthaśāstra, no weighting in Charaka's lists, no procedure for revising the list of pramāṇas.

The relation. Every one of the six subjects supplies a structure and stops short of something a modern treatment would supply. The answer worth writing names which, because "the ancients had everything" and "the ancients had nothing" are both wrong and the specific gap is the interesting fact.

Worked cross-module answers, eight of them

1. Compare the role of rules in Pāṇini's grammar and in a Nyāya argument.

Both are rule systems, and the rules do different work. Pāṇini's vidhi rules transform a form: they take a root and an affix and produce a word, and his four precedence principles decide which rule fires when two could, because the system must produce exactly one form. Nyāya's rules are not transformations but conditions on an argument: the five members fix the shape a claim must take, and the five fallacies name the ways a reason can fail. So Pāṇini's rules are productions with a conflict policy, and Nyāya's are a schema with a validator. Both, however, make the general premise explicit: Pāṇini states his paribhāṣā as numbered rules, and Nyāya requires the vyāpti to be a member of the argument rather than an assumption.

2. Relate Piṅgala's binary encoding to the Arthaśāstra's cipher-writing.

Both reduce a message to symbols and operate on the symbols rather than the meaning, which is the precondition of any mechanical treatment. Piṅgala's reduction is complete and public: two values per syllable, and his six operations enumerate, index and count the whole space. The Arthaśāstra's is partial and secret: the text names gūḍhalekhya without supplying a method, and its security rests on the method's secrecy, which Kerckhoffs's principle identifies as the wrong place to rest it. The comparison shows the difference between a representation chosen to make computation possible and one chosen to make reading impossible.

munotes.in375

The Whole Paper on One Page

3. Compare the pramāṇa scheme of Module 1 with the validation protocol of Module 2.

Module 1's pramāṇa scheme declares three admissible means of knowledge, and Charaka adds that no one of them suffices, which is a rule about the coverage of an investigation. Module 2's protocol adds four things the scheme does not have: a required form for a claim, in five members; an enumerated list of defects in a reason; a list of improper moves; and enumerated grounds of defeat. The difference is the presence of an opponent. A scheme for investigating needs only to say what counts as evidence; a scheme for disputing must also say how an argument is stated, how it fails, and when it is over.

4. Both modules contain a system with declared rule types. Compare them.

Pāṇini's tradition classifies its sūtras into six kinds, of which only vidhi performs an operation and the other five govern how rules are read and applied. Vaiśeṣika's padārtha scheme classifies everything nameable into six categories, of which three are things and three are relations among them, and Sinha records that in reality there are only three predicables. So both systems declare a small closed set of types, both make relations first-class members of the set rather than leaving them implicit in the notation, and both use the types to state policies: Pāṇini's precedence principles are statements about kinds of rule, and Vaiśeṣika's inherence is what lets a property be attached to a bearer and asked about.

5. Relate the Meru-prastāra to the label set of the tridoṣa scheme.

The Meru-prastāra's row n gives the number of ways of choosing k things from n, and its row three is 1, 3, 3, 1. Classifying a constitution by which of the three doṣa predominate is choosing a subset of three, so the possibilities are the entries of that row: one way for none, three for one, three for two, one for all three, of which the empty selection is excluded, leaving seven labels. So a combinatorial result from Module 1 answers a question about the label set of a classifier in Module 2, and it explains why the problem is multi-class or multi-label rather than binary.

6. Compare what Piṅgala and the Arthaśāstra each leave out.

Piṅgala's six operations are stated without any account of their cost: nothing in the text says that computing 2 to the power n by squaring is cheaper than multiplying it out, although it is, and the saving is ours to notice. The Arthaśāstra names cipher-writing without any account of its method: nothing says how a message was enciphered, what the key was, or how it was distributed. In both cases the text supplies the structure and omits what a modern treatment would supply, but the omissions differ in kind: Piṅgala omits an analysis of something he states completely, and Kautilya omits the thing itself.

munotes.in376

The Whole Paper on One Page

7. Both modules contain an example of ordering that changes the result. Compare them.

In Module 1, Pāṇini's precedence principles decide which of two applicable rules fires, and applying them in the wrong order produces a different and incorrect form; asiddhatva goes further, making a whole section of rules invisible to the rest so that certain orderings cannot arise. In Module 2, Charaka's parīkṣā orders the means of knowledge, putting instruction first because without it one does not know what to observe, and a rule-based system's conflict resolution decides which of several matching rules fires. The common point is that in all three the order is part of the specification and not an optimisation; the difference is that Pāṇini's order is required for correctness and Charaka's is required for the investigation to be possible at all.

8. What does this paper establish about the relation between classical Indian knowledge systems and computer science?

That six classical disciplines share a design method with modern computer science: each fixes a small closed vocabulary, states rules over it, declares what will count as evidence or as a correct result, and provides for the case where two rules apply at once. Five specific correspondences are established from the texts: binary encoding and enumeration in Piṅgala, index retrieval by arithmetic rather than by search, exponentiation by squaring, typed rules with a conflict policy in Pāṇini, and an argument form that carries its own explanation in Nyāya. What is not established is any claim of influence on the modern subject, which would require evidence about what particular researchers read, and what is absent from every one of the six is a machine, so the correspondence is one of design rather than of implementation.

The last word

The test this book applied throughout, from its first chapter: not whether an old text looks like a new one, but whether it does the same work, and where the comparison breaks.

Applied honestly, that test gives more than an uncritical account would. Piṅgala really did choose a representation that made his questions arithmetical, and really did not have a notion of cost. Pāṇini really did state a conflict policy for a rule system of four thousand rules, and really was not writing a context-free grammar. Nyāya really did require the general premise to be stated and attackable, and really had no procedure for establishing it. Charaka really did say that one means of knowledge is never enough. And Kautilya really did state a threat model and really did not state a cipher.

munotes.in377

The Whole Paper on One Page

Every one of those sentences is a finding, and every one of them is in this book with the passage it rests on.

Quick revision

  • Q.3 is set on both modules and is worth ten of the thirty external marks, in two answers of five.
  • Six classical subjects, three in each module, each paired by MU with named modern concepts.
  • Five relations cross the modules: rules and conflict, evidence and sufficiency, representation and tractability, explanation and its form, and what the texts do not contain.
  • The strongest single cross-module fact: row three of the Meru-prastāra, 1 3 3 1, gives the seven labels of the prakṛti classification.
  • The test this book applies: not whether an old text looks like a new one, but whether it does the same work, and where the comparison breaks.

Test yourself

1. Name the five relations that cross the two modules.

Rules and what happens when two apply; evidence and what counts as enough; representation and what it makes easy; explanation and the form an answer takes; and what the texts do not contain.

2. A Q.3 question asks you to relate Piṅgala to the Arthaśāstra. What is the relation and what is the difference?

Both reduce a message to symbols and operate on the symbols rather than the meaning. Piṅgala's reduction is complete and public, chosen to make computation possible; the Arthaśāstra's is partial and secret, chosen to make reading impossible, and its security rests on the method's secrecy, which is the wrong place to rest it.

3. Which combinatorial result from Module 1 answers a question about Module 2's classifier?

Row three of the Meru-prastāra, 1, 3, 3, 1. Choosing which of three doṣa predominate is choosing a subset of three, so there are seven non-empty labels, which is why the classification is multi-class or multi-label rather than binary.

4. State what this paper establishes, and what it does not.

It establishes that six classical disciplines share a design method with computer science, and it establishes five specific correspondences from the texts. It does not establish any influence on the modern subject, which would require evidence about what particular researchers read, and in none of the six is there a machine, so the correspondence is of design and not of implementation.

Contents This chapter on its own page

munotes.in378

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Issue
Done!