Feature Engineering
Chapter Eighty-Three
Syllabus topic Module 2, "Feature engineering"
Pages 297 to 299 of 378
In one line
Feature engineering is choosing what the system gets to look at, and it decides more about the result than the classifier does.
In the wording you can write in an examination: feature engineering is the selection, construction, transformation and encoding of the attributes presented to a classifier. It comprises selecting which available attributes to use, deriving new attributes from the raw data, scaling numeric attributes to comparable ranges, and encoding non-numeric attributes so that a method can process them.
The four operations
Selection. Choosing which of the available attributes to keep. Fewer attributes means less to overfit and less to collect.
Construction. Making a new attribute out of existing ones. The ratio of two measurements, the difference between a value and its recent average, the count of some event in a window.
Scaling. Putting numeric attributes on comparable ranges, so that one measured in thousands does not swamp one measured in tenths.
Encoding. Turning a category into numbers. The standard method is one-hot encoding: one binary column per possible value.
And the one that matters most is construction, because a constructed attribute can make an impossible problem easy, and no classifier can construct one for itself unless it was designed to.
Worked example: why construction matters
The problem. Classify whether a transaction is unusual for an account.
With the raw attribute. The amount. A classifier over amount alone must learn a threshold, and no single threshold suits both a large account and a small one.
With a constructed attribute. The amount divided by the account's own median transaction. Now a single threshold works for every account, because the attribute already carries the comparison.
Nothing about the classifier changed. The problem became easy because the representation changed. That is the whole argument for feature engineering and it is a restatement of [Knowledge Representation: The Modern Name For It]: a representation is a bet about which questions you will be asked.
The guṇa list as a derived feature space
This is the chapter's claim about the classical material and it is worth stating carefully.
The raw material is a body and a report. Innumerable things could be recorded.
The scheme records qualities, and it records a FIXED list of them. Dry, cold, light, subtile, unstable, clear, keen, and the rest. That is selection: a small closed vocabulary chosen in advance out of everything that could be said.
And the qualities are not raw observations. "Unstable" is not measured; it is a judgement about a pattern over time. "Subtile" is not read off an instrument. Each quality is a constructed attribute, derived from what is observed by a rule carried in the practitioner's training.
So the vocabulary is a designed feature space, and its design has three properties a modern engineer would recognise.
Feature Engineering
| Property | What it gives |
|---|---|
| small and closed | every case is described in the same terms, so cases are comparable |
| shared across classes | one space, not one per doṣa, which is what makes a vector possible |
| built in opposite pairs | the treatment rule becomes computable, as [The Tridoṣa Framework] shows |
The classical list of the qualities is twenty, in ten opposed pairs. This book uses only those the three doṣa lists name, which is eighteen attributes over nine pairs, and it says so rather than importing a count it has not used.
Selection, and the measure of a useless attribute
An attribute every class has tells you nothing. That is [Doṣa as a Feature Vector]'s finding about "cold" under the disputed reading, and it is the simplest case of feature selection: a column with the same value in every row can be deleted with no loss.
The general measure is the information gain of [Decision Trees]. An attribute with zero gain does not separate the classes at all.
And the general warning is [A Decision Tree Built From the Tridoṣa Attributes]'s. With too few rows, every attribute has the same gain and the measure selects nothing. Feature selection needs data, exactly as the tree does.
Encoding, and the trap
One-hot encoding. A category with five possible values becomes five binary columns, one of which is set.
Why not just number the values. Because numbering asserts an order and a spacing. Coding the three doṣa as 1, 2, 3 tells a classifier that bile is between wind and phlegm and that the distance from wind to phlegm is twice the distance from wind to bile. Neither is true and the classifier will use both.
The cost of one-hot encoding is that the space grows, and it grows fastest exactly where the category has many values. A category with a thousand values becomes a thousand columns, most of them zero in any row.
And the doṣa attribute space is already one-hot. Eighteen binary columns over a shared vocabulary is exactly the result of one-hot encoding a set-valued attribute, and this is why the representation of [Doṣa as a Feature Vector] needed no encoding step: the source was already in that form.
Scaling, and why it does not arise here
Scaling matters when attributes have different ranges and a method measures distance. Age in years and income in rupees cannot be compared without it.
It does not arise in this block, because every attribute is binary. That is worth saying because it is a real simplification: a scheme with numeric attributes, such as one recording degrees of a quality, would need it.
Feature Engineering
What feature engineering is NOT
It is not preprocessing. Cleaning data and engineering features are different activities: cleaning removes errors, engineering changes what the model sees.
It is not made obsolete by learned representations. A method that learns its own features removes the need to construct them by hand and requires far more data, and it produces features nobody can interpret. The trade is the one [Explainable AI, and Why a Five-Member Answer Is an Explanation] describes.
It is not free of assumptions. Every constructed attribute encodes a belief about what matters. The belief is now in the data rather than in the model, where it is harder to notice, and that is a real risk rather than a rhetorical caution.
Quick revision
- Four operations: selection, construction, scaling, encoding. Construction matters most, because it can make a hard problem easy.
- A ratio to a per-account baseline turns a problem no single threshold solves into one a single threshold solves.
- The guṇa vocabulary is a designed feature space: small, closed, shared across classes, and built in opposite pairs so the treatment rule is computable.
- The qualities are constructed, not raw: "unstable" is a judgement about a pattern, not a measurement.
- An attribute every class has can be deleted; the measure is information gain, and it needs data to work.
- One-hot encoding avoids asserting an order that numbering would assert. The doṣa space is already in that form.
- Scaling does not arise here because every attribute is binary.
Test yourself
1. Name the four operations of feature engineering and say which is the most consequential.
Selection, construction, scaling and encoding. Construction is the most consequential, because a constructed attribute can turn a problem no classifier could solve into one a simple threshold solves.
2. In what sense is the classical list of qualities a designed feature space?
It is a small closed vocabulary selected in advance out of everything that could be recorded; it is shared across all three classes, so cases are comparable; its members are judgements derived from observation rather than raw measurements; and it is built in opposite pairs so that the treatment rule has something to apply.
3. Why is numbering a category 1, 2, 3 a mistake, and what is done instead?
Because it asserts an order and equal spacing that the category does not have, and a classifier will use both. One-hot encoding is used instead: one binary column per possible value.
4. Why does feature selection fail on the tridoṣa table?
Because with three rows every attribute has the same information gain, so the measure cannot rank them. Feature selection needs enough data for attributes to separate the classes differently.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.