Knowledge Representation: The Modern Name For It
Chapter Ten
Syllabus topic Module 1, "Knowledge representation"
Pages 29 to 31 of 378
In one line
Knowledge representation is the business of writing down what a system knows in a form the system can use to work out something it was not told.
In the wording you can write in an examination: knowledge representation is the branch of computer science concerned with encoding facts about a domain in a formal structure that supports inference. A representation is a choice of syntax for expressing knowledge, a semantics fixing what each expression means, an inference procedure for deriving new expressions from stored ones, and a commitment about what kinds of thing exist in the domain.
Why a representation is a choice, and why the choice costs something
A system can be told "Rama is a student". That could be stored as a row in a table, a sentence in logic, a link between two nodes, a slot in a record, or a rule. All five are faithful. They are not equivalent, because each makes different questions cheap and different questions expensive.
That is the whole subject in one sentence: a representation is a bet about which questions you will be asked. Choose the wrong notation and every query fights the structure.
The four things a representation must fix
These are the four questions to answer in an examination when asked what a knowledge representation is.
Syntax. What strings count as well formed. In logic, student(rama) is well formed and rama student( is not.
Semantics. What a well formed string means, precisely enough that two people cannot read it two ways. This is where most informal notations fail: a box joined to another box by an unlabelled arrow means whatever the reader guesses.
Inference. What may be derived from what. Without this a representation is a filing system, not a knowledge base.
Ontological commitment. What kinds of thing the notation says exist. A notation with only objects and properties cannot talk about events; one with only classes cannot talk about individuals. [Padārtha: The Categories of What Exists] is a classical answer to exactly this question.
The four families, on one domain
Take the smallest useful domain: a college has students; a student has a roll number; Rama is a student with roll number 41; every student may borrow books.
Logic. Facts and rules as formulas.
student(rama)
roll(rama, 41)
for all X: student(X) implies mayBorrow(X)
Cheap: asking whether Rama may borrow. Expensive: saying anything is only usually true.
Semantic network. A labelled graph. Nodes are things and classes; edges are named relations.
rama --is-a--> student --may--> borrow
rama --roll--> 41
Cheap: following relations, and finding what is connected to what. Expensive: saying anything with more than two arguments, because an edge joins exactly two nodes.
Frames. A record per concept, with named slots, defaults, and inheritance from a parent frame.
Knowledge Representation: The Modern Name For It
frame student
parent: person
may: borrow
roll: an integer
frame rama
parent: student
roll: 41
Cheap: defaults and inheritance, so rama need not repeat may borrow. Expensive: exceptions, which have to be handled by overriding a slot and then explaining why.
Rules. IF conditions THEN conclusion, with an engine that applies them.
IF person(X) and enrolled(X)
THEN student(X)
IF student(X)
THEN mayBorrow(X)
Cheap: adding a new piece of knowledge without touching the rest. Expensive: understanding what the system will do overall, because the behaviour is spread across the rules.
The comparison a five-mark answer needs
| Family | Basic unit | Strongest at | Weakest at |
|---|---|---|---|
| Logic | a formula | exactness, and provable inference | defaults, exceptions, degrees of belief |
| Semantic network | a labelled edge | following relations, visualising structure | relations of more than two arguments |
| Frames | a slot in a record | defaults and inheritance, compact storage | exceptions, and precise semantics |
| Rules | a condition and a conclusion | adding knowledge piecemeal, explanation | predicting overall behaviour |
Worked example: the same fact, and the question that separates the notations
The fact: most students return books on time.
In logic, this cannot be said at all without extending the notation. for all X: student(X) implies returnsOnTime(X) is false, and there is no standard way in plain first-order logic to say "most".
In a semantic network, you can draw an edge labelled "usually returns on time", and now the notation's semantics have quietly become whatever the reader thinks that label means.
In a frame, it is natural: the student frame gets a slot returns with default on time, and an individual frame may override it.
As a rule, it is natural in a different way: a rule concluding returnsOnTime(X) from student(X), plus a second rule that retracts it for a named individual.
The lesson is not that frames and rules are better. It is that the fact was easy to state in two notations and not statable in a third, and you cannot know that until you try. Pick the notation after you know the questions, not before.
What knowledge representation is NOT
It is not a database schema. A schema fixes syntax and, loosely, semantics. It has no inference procedure, so nothing is derived that was not stored.
It is not the same as storing text. A page of prose contains knowledge and supports no inference at all. Turning prose into a representation is the work, and it is the work every chapter of this paper's Module II asks you to do on a classical text.
It is not free of commitments. Every notation says something about what exists. Choosing one is choosing an ontology whether or not you notice.
Knowledge Representation: The Modern Name For It
The classical connection, stated carefully
The claim MU's syllabus makes is that śāstra method is knowledge representation, and the honest version of that claim is specific.
What matches. A śāstra fixes its syntax, its semantics and its inference, which are three of the four requirements above, and Vaiśeṣika's padārtha scheme is an explicit ontological commitment, which is the fourth.
What does not. None of it is machine readable, and none of it was meant to be. The inference procedure is a trained reader. So the correspondence is at the level of design, not of implementation, and an answer that says so is stronger than one that does not.
Quick revision
- Knowledge representation: encoding what is known so that new facts can be derived.
- Four requirements: syntax, semantics, inference procedure, ontological commitment.
- Four families: logic, semantic networks, frames, rules, each cheap at some questions and expensive at others.
- A representation is a bet about which questions will be asked.
- The classical correspondence is at the level of design: a śāstra fixes all four, but its inference procedure is a person.
Test yourself
1. Name the four things a knowledge representation must fix, and say what goes wrong if the second is missing.
Syntax, semantics, an inference procedure, and an ontological commitment. Without semantics, two readers can take the same expression two ways, so nothing derived from it can be relied on: an unlabelled arrow between two boxes is the standard example.
2. Why is a database schema not a knowledge representation?
It has no inference procedure, so it stores facts and derives none. A representation must support working out something the system was not told.
3. Take the fact "most students return books on time" and say which of the four families states it naturally and which cannot.
Frames state it as a default on a slot, and rules state it as a rule with an override. Plain first-order logic cannot state it at all, since the universal is false and there is no standard quantifier for "most". A semantic network can draw the label but loses precise semantics in doing so.
4. State the correspondence between śāstra method and knowledge representation, with its limit.
A śāstra fixes terms, rules and admissible evidence, and a scheme such as padārtha is an explicit ontology, so all four requirements are addressed. The limit is that nothing is machine readable and the inference procedure is a trained human reader, so the correspondence is one of design, not implementation.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.