Ayurvedic Classification as a Rule-Based Expert System
Chapter Eighty-Five
Syllabus topic Module 2, "Ayurvedic Classification as Rule-Based Expert System", "Rule-based systems", "Expert systems", "Algorithm Specification (Pseudo-code)", "minimum 10 test cases"
Pages 303 to 309 of 378
In one line
A rule base read straight off Charaka's attribute lists, a scoring rule, and an explanation that names the text.
In the wording you can write in an examination: a rule-based expert system over the tridoṣa scheme represents each doṣa's attribute list as its rule set, classifies a description by counting matched attributes, reports the score for every class rather than only the maximum, and accompanies each answer with the attributes that produced it and the counter-measure the treatment rule prescribes.
Problem statement, in MU's own form
IKS concept as CS concept: Āyurvedic classification as a rule-based expert system.
Statement. Represent the attribute lists of Charaka Sūtrasthāna I.58 to I.60 as a rule base. Classify a description, given as a set of attributes, by the number of attributes it shares with each doṣa. Report all three scores and the matched attributes, so that the answer carries its own explanation, and derive the counter-measure from Sūtrasthāna I.61's rule of the adverse attribute. Exclude the disputed reading and say so in the source.
Conceptual mapping table
| Classical element | Computer science element |
|---|---|
| a guṇa, a named quality | a binary attribute |
| the list of seven per doṣa | that class's rule set |
| the union of the lists | the shared feature space |
| "cold" in bile, printed and doubted | a declared exclusion, held in its own constant |
| classifying by shared qualities | a weighted vote with equal weights |
| a case showing qualities of two doṣa | a tie, reported rather than broken |
| the rule of the adverse attribute | a lookup from each matched attribute to its opposite |
| place, measure and time | three parameters the program does not model, and says so |
Algorithm specification, in pseudo-code
ALGORITHM Classify(described)
INPUT a set of attribute names
OUTPUT a score per dosa, the attributes matched, an answer or a tie, and a reason
for each dosa do
have <- that dosa's attributes, LESS any disputed reading for it
matched <- described intersected with have
score <- the size of matched
end for
best <- the highest score
if best is zero then return "no attribute in the description appears in any list"
winners <- every dosa whose score is best
if there is more than one winner then return a TIE naming them
return the single winner, its matched attributes, and the opposites of those attributes
ALGORITHM Discriminating()
return every attribute that some dosa has and some dosa lacks
Two design points to defend in a viva. A tie is returned as a tie and not broken, because the scheme itself treats a combination as an answer, as [Decision Principles in Diagnosis] sets out. And the disputed reading is held in its own constant rather than deleted from the table, so that the program records what the text prints as well as what it relies on.
Ayurvedic Classification as a Rule-Based Expert System
Working code
#!/usr/bin/env python3
"""Ayurvedic classification as a rule-based expert system. MU's topic 4.
THIS IS NOT MEDICAL SOFTWARE AND IT IS NOT MEDICAL ADVICE. It is a
knowledge-representation exercise on a classical text, which is what MU's
syllabus asks for. Nothing here says a classical rule is clinically correct.
IKS concept as CS concept: the tridosa attribute lists of Charaka Samhita,
Sutrasthana I.58 to I.60, as a feature space; the classification of a
description into wind, bile or phlegm as multi-class classification; and
Charaka's own rule of the adverse attribute as an IF-THEN rule with an
explanation.
THE TABLE IS QUOTED, NOT INVENTED. It is read off Avinash Chandra
Kaviratna's translation:
I.58 "Wind, which may be dry, cold, light, subtile, unstable, clear, keen,
is cured by objects which have adverse attributes."
I.59 "Bile, which may be cold, hot, keen, soft, sour, liquid, and bitter,
is speedily cured by objects having adverse attributes."
I.60 "Heavy, cold, mild, watery, sweet, stable, and slimy, these attributes
of phlegm are cured by objects having adverse attributes."
I.59 prints BOTH "cold" and "hot" for the same dosa, which cannot both be
intended. The table below records the line as printed and DISPUTED holds the
attribute the book refuses to rely on. FINDINGS section 3.
"""
import math
ATTRIBUTES = {
'wind': ['dry', 'cold', 'light', 'subtile', 'unstable', 'clear', 'keen'],
'bile': ['cold', 'hot', 'keen', 'soft', 'sour', 'liquid', 'bitter'],
'phlegm': ['heavy', 'cold', 'mild', 'watery', 'sweet', 'stable', 'slimy'],
}
DISPUTED = {('bile', 'cold')}
OPPOSITE = {
'dry': 'moist', 'cold': 'hot', 'hot': 'cold', 'light': 'heavy',
'heavy': 'light', 'subtile': 'gross', 'unstable': 'stable',
'stable': 'unstable', 'clear': 'slimy', 'slimy': 'clear',
'keen': 'mild', 'mild': 'keen', 'soft': 'hard', 'sour': 'sweet',
'sweet': 'sour', 'liquid': 'solid', 'bitter': 'sweet',
'moist': 'dry', 'watery': 'dry', 'gross': 'subtile',
}
def universe():
"""Every attribute any dosa is given, in first-appearance order."""
seen = []
for dosa in ('wind', 'bile', 'phlegm'):
for a in ATTRIBUTES[dosa]:
if a not in seen:
seen.append(a)
return seen
def vector(dosa, drop_disputed=True):
"""The dosa as a 0/1 vector over the shared attribute space."""
have = set(ATTRIBUTES[dosa])
if drop_disputed:
have -= {a for d, a in DISPUTED if d == dosa}
return [1 if a in have else 0 for a in universe()]
def discriminating(drop_disputed=True):
"""Attributes that do NOT appear against every dosa.
An attribute all three share cannot tell them apart, whatever else is true
of it. This is feature selection done by hand before any entropy is
computed.
"""
out = []
for i, a in enumerate(universe()):
col = [vector(d, drop_disputed)[i] for d in ('wind', 'bile', 'phlegm')]
if 0 < sum(col) < 3:
out.append(a)
return out
# ----------------------------------------------------------- the rule base
def classify(described):
"""Score a set of described attributes against each dosa.
A count, not a diagnosis. The explanation is part of the answer,
because MU's Module II asks for explainable inference and an expert
system that cannot say why has failed at its job.
"""
described = set(described)
scores, why = {}, {}
for dosa in ('wind', 'bile', 'phlegm'):
have = set(ATTRIBUTES[dosa]) - {a for d, a in DISPUTED if d == dosa}
matched = sorted(described & have)
scores[dosa] = len(matched)
why[dosa] = matched
best = max(scores.values())
winners = sorted(d for d, s in scores.items() if s == best)
return {
'scores': scores,
'matched': why,
'answer': winners[0] if best and len(winners) == 1 else None,
'tied': winners if best and len(winners) > 1 else [],
'explanation': _explain(scores, why, winners, best),
}
def _explain(scores, why, winners, best):
if not best:
return 'No attribute in the description appears in any of the three lists.'
if len(winners) > 1:
return ('The description matches %s equally, on %d attribute(s) each, so it '
'does not decide between them.' % (' and '.join(winners), best))
d = winners[0]
return ('%s, because the description gives %s, and Sutrasthana I names those '
'among the attributes of %s. The counter-measure the text states is an '
'object of the adverse attribute: %s.'
% (d.capitalize(), ', '.join(why[d]), d,
', '.join(OPPOSITE.get(a, 'the opposite of ' + a) for a in why[d])))
# --------------------------------------------------- the decision tree
def entropy(labels):
n = len(labels)
if n == 0:
return 0.0
out = 0.0
for lab in set(labels):
p = labels.count(lab) / n
out -= p * math.log2(p)
return out
def information_gain(rows, attr):
"""rows: [(label, set-of-attributes)]. Split on has/has not."""
labels = [r[0] for r in rows]
before = entropy(labels)
yes = [r[0] for r in rows if attr in r[1]]
no = [r[0] for r in rows if attr not in r[1]]
n = len(rows)
after = (len(yes) / n) * entropy(yes) + (len(no) / n) * entropy(no)
return before - after
def training_rows(drop_disputed=True):
rows = []
for dosa in ('wind', 'bile', 'phlegm'):
have = set(ATTRIBUTES[dosa])
if drop_disputed:
have -= {a for d, a in DISPUTED if d == dosa}
rows.append((dosa, have))
return rows
def build_tree(rows, attrs):
labels = [r[0] for r in rows]
if len(set(labels)) <= 1:
return labels[0] if labels else None
usable = [(information_gain(rows, a), a) for a in attrs]
usable = [(g, a) for g, a in usable if g > 0]
if not usable:
return sorted(set(labels))
gain, attr = max(usable, key=lambda t: (t[0], -attrs.index(t[1])))
rest = [a for a in attrs if a != attr]
return {
'attribute': attr,
'gain': round(gain, 4),
'yes': build_tree([r for r in rows if attr in r[1]], rest),
'no': build_tree([r for r in rows if attr not in r[1]], rest),
}
def render_tree(node, indent=0, label=''):
pad = ' ' * indent
if not isinstance(node, dict):
return '%s%s-> %s' % (pad, label, node)
out = ['%s%sis it %s? (information gain %.4f)'
% (pad, label, node['attribute'], node['gain'])]
out.append(render_tree(node['yes'], indent + 2, 'yes: '))
out.append(render_tree(node['no'], indent + 2, 'no: '))
return '\n'.join(out)
def walk(node, described):
trail = []
while isinstance(node, dict):
a = node['attribute']
took = a in described
trail.append('%s %s' % (a, 'yes' if took else 'no'))
node = node['yes'] if took else node['no']
return node, trail
TESTS = [
('three attributes of wind', ['dry', 'unstable', 'subtile'], 'wind'),
('three attributes of bile', ['hot', 'sour', 'bitter'], 'bile'),
('three attributes of phlegm', ['heavy', 'sweet', 'slimy'], 'phlegm'),
('one attribute only, unique', ['watery'], 'phlegm'),
('the shared attribute alone', ['cold'], None),
('nothing in any list', ['blue', 'loud'], None),
('an empty description', [], None),
('keen, which two dosas share', ['keen'], None),
('keen with one wind attribute', ['keen', 'dry'], 'wind'),
('keen with one bile attribute', ['keen', 'sour'], 'bile'),
('a mixed description, wind leads', ['dry', 'light', 'clear', 'sour'], 'wind'),
# Expected None, and the reason is the finding: with the disputed reading
# dropped, "cold" belongs to wind and phlegm and "hot" to bile, so this
# description matches all three once and settles nothing. Writing 'bile'
# here was the first guess and the test caught it.
('cold and hot together', ['cold', 'hot'], None),
]
def run_tests(verbose=False):
passed = 0
for name, described, expect in TESTS:
got = classify(described)['answer']
ok = got == expect
passed += ok
if verbose:
print('%-46s %s expected %-7s got %s'
% (name, 'pass' if ok else 'FAIL', expect, got))
assert ok, (name, got, expect)
return passed
def prove():
u = universe()
assert len(u) == len(set(u))
# every dosa gets seven attributes as printed
for d in ATTRIBUTES:
assert len(ATTRIBUTES[d]) == 7, d
# "cold" is printed against all three, so it discriminates nothing
assert all('cold' in ATTRIBUTES[d] for d in ATTRIBUTES)
assert 'cold' not in discriminating(drop_disputed=False)
# dropping the disputed reading makes cold a phlegm-and-wind attribute
assert 'cold' in discriminating(drop_disputed=True)
assert abs(information_gain(training_rows(), 'cold')) > 0
assert abs(information_gain(training_rows(drop_disputed=False), 'cold')) < 1e-12
# and bile never claims the disputed attribute
assert classify(['cold'])['matched']['bile'] == []
# the tree separates all three
tree = build_tree(training_rows(), discriminating())
for d in ('wind', 'bile', 'phlegm'):
leaf, _ = walk(tree, set(ATTRIBUTES[d]) - {a for x, a in DISPUTED if x == d})
assert leaf == d, (d, leaf)
run_tests()
return True
if __name__ == '__main__':
prove()
print('### the shared attribute space, %d attributes' % len(universe()))
print(' ' + ', '.join(universe()))
print()
print('### the three vectors')
print('%-8s %s' % ('', ' '.join('%-9s' % a for a in universe())))
for d in ('wind', 'bile', 'phlegm'):
print('%-8s %s' % (d, ' '.join('%-9d' % v for v in vector(d))))
print()
print('### information gain of every attribute, disputed reading dropped')
rows = training_rows()
for a in universe():
print(' %-9s %.4f' % (a, information_gain(rows, a)))
print()
print('### the tree')
print(render_tree(build_tree(rows, discriminating())))
print()
print('### the ten test cases')
n = run_tests(verbose=True)
print('\n%d of %d pass' % (n, len(TESTS)))
print()
print('### one answer with its explanation')
r = classify(['dry', 'unstable', 'subtile'])
print(' ' + r['explanation'])Ayurvedic Classification as a Rule-Based Expert System
### the shared attribute space, 18 attributes
dry, cold, light, subtile, unstable, clear, keen, hot, soft, sour, liquid, bitter, heavy, mild, watery, sweet, stable, slimy
### the three vectors
dry cold light subtile unstable clear keen hot soft sour liquid bitter heavy mild watery sweet stable slimy
wind 1 1 1 1 1 1 1 0 0 0 0 0 0 0 0 0 0 0
bile 0 0 0 0 0 0 1 1 1 1 1 1 0 0 0 0 0 0
phlegm 0 1 0 0 0 0 0 0 0 0 0 0 1 1 1 1 1 1
### information gain of every attribute, disputed reading dropped
dry 0.9183
cold 0.9183
light 0.9183
subtile 0.9183
unstable 0.9183
clear 0.9183
keen 0.9183
hot 0.9183
soft 0.9183
sour 0.9183
liquid 0.9183
bitter 0.9183
heavy 0.9183
mild 0.9183
watery 0.9183
sweet 0.9183
stable 0.9183
slimy 0.9183
### the tree
is it dry? (information gain 0.9183)
yes: -> wind
no: is it cold? (information gain 1.0000)
yes: -> phlegm
no: -> bile
### the ten test cases
three attributes of wind pass expected wind got wind
three attributes of bile pass expected bile got bile
three attributes of phlegm pass expected phlegm got phlegm
one attribute only, unique pass expected phlegm got phlegm
the shared attribute alone pass expected None got None
nothing in any list pass expected None got None
an empty description pass expected None got None
keen, which two dosas share pass expected None got None
keen with one wind attribute pass expected wind got wind
keen with one bile attribute pass expected bile got bile
a mixed description, wind leads pass expected wind got wind
cold and hot together pass expected None got None
12 of 12 pass
### one answer with its explanation
Wind, because the description gives dry, subtile, unstable, and Sutrasthana I names those among the attributes of wind. The counter-measure the text states is an object of the adverse attribute: moist, gross, stable.Ayurvedic Classification as a Rule-Based Expert System
Reading the output
Eighteen attributes and three vectors, which is [Doṣa as a Feature Vector]'s table produced by the program that uses it.
Ayurvedic Classification as a Rule-Based Expert System
Every attribute has the same information gain, 0.9183, which is [A Decision Tree Built From the Tridoṣa Attributes]'s finding, printed here from the same code that builds the tree.
Ayurvedic Classification as a Rule-Based Expert System
The tree separates all three in two questions.
Twelve test cases pass, and five of them return no answer: the shared attribute alone, nothing in any list, an empty description, "keen" which two doṣa share, and "cold and hot" together. Five refusals out of twelve, which is the proportion a classifier over a small closed vocabulary should have.
And the last block is the explanation. It names the attributes that produced the answer, says which section of the text names them, and gives the counter-measure as the opposites. That is what makes this an expert system rather than a classifier: the answer carries its reason, as [Expert Systems, and MYCIN as the Comparison] requires.
The test case that was wrong first
"Cold and hot together" was expected to return bile, on the reasoning that "hot" is bile's and "cold" is disputed for bile so only "hot" would count.
The program returned nothing, and the program was right. With the disputed reading dropped, "cold" belongs to wind and to phlegm. So the description matches wind once, bile once and phlegm once: a three-way tie, and nothing is settled.
The test was corrected, not the program. That is recorded in the source and it is the sort of thing a limitations section should say: a test written from an expectation rather than from the specification is a test of the expectation.
Complexity and limitations
This is MU's heading seven and it is where a submission earns its marks.
Time. Three set intersections over a description of at most eighteen attributes, so the work is constant for practical purposes. The tree induction is proportional to the number of attributes times the number of rows, which is fifty-four.
Space. The table, the disputed set, and the opposite lookup.
Limitation: it reports no accuracy, and cannot. There is no labelled data with a ground truth for this scheme. An accuracy figure would be invented.
Limitation: equal weights. Nothing in Charaka weights the attributes, so every one counts one. [Multi-Attribute Classification] explains why any other weighting would have to be invented.
Limitation: it does not model place, measure and time. Sūtrasthāna I.61 makes the counter-measure depend on all three, and the program returns the opposites without them. That is a real gap between the text and the implementation, and it is declared rather than passed over.
Limitation: binary attributes. The text speaks of qualities being increased and diminished; the program has present and absent. [Doṣa as a Feature Vector] states the simplification.
Limitation: it does not distinguish prakṛti from a current state. The scheme distinguishes a constitution from a present condition, and the program works on a single description with no notion of which it is.
Ayurvedic Classification as a Rule-Based Expert System
Limitation: the vocabulary problem is unsolved. The program takes attribute names as input. Getting from what a person says to those names is the hard part, and [Symptom to Feature Mapping] describes it rather than solving it.
And the one that matters most. It classifies into the scheme's own categories. Whether those categories correspond to anything is a question this program does not ask and cannot answer.
Quick revision
- The rule base is the three attribute lists, less the disputed reading, which is held in its own constant.
- Classification is a count of shared attributes, with equal weights, and all three scores are reported.
- A tie is returned as a tie, because the scheme treats a combination as an answer.
- The counter-measure is the opposites of the matched attributes, from Sūtrasthāna I.61.
- Twelve test cases, five of which return no answer, which is the right proportion for a small closed vocabulary.
- One test was written from an expectation and was wrong; the test was corrected, not the program.
- Limitations: no accuracy and none possible, equal weights, no place-measure-time, binary attributes, no prakṛti distinction, and the vocabulary problem unsolved.
Test yourself
1. Why does the program return a tie rather than breaking it?
Because the scheme itself treats a case showing the qualities of two doṣa as a case in which two predominate, which is one of its own labels. Breaking the tie would discard information the scheme regards as the answer.
2. What does the program do with the disputed attribute, and why is that better than deleting it?
It holds it in a separate constant and excludes it from bile's list when classifying. That records both what the text prints and what the program relies on, so a reader can see the decision instead of finding a silently shortened list.
3. Give three limitations that belong in this implementation's own limitations section.
It reports no accuracy and none is possible, because no labelled data with a ground truth exists. It weights every attribute equally, because Charaka assigns no weights. And it does not model place, measure and time, which Sūtrasthāna I.61 makes the counter-measure depend on.
4. A test expected bile for a description of cold and hot, and the program returned nothing. What was wrong?
The test. With the disputed reading excluded, cold belongs to wind and phlegm and hot to bile, so the description matches each class once and settles nothing. The expectation was written from a half-remembered reading of the table rather than from the specification.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.