Intrusion Detection: Statistical Anomaly and Rule-Based
Chapter Eighty-Five
Syllabus topic Module 2, "Intrusion Detection"
Pages 571 to 576 of 678
In one line
Intrusion detection watches what happens on a system and looks for signs of attack in two ways: by comparing behaviour with what is normal for that user or system (statistical anomaly detection), or by matching it against rules and patterns that describe misuse and known attacks (rule-based detection).
In the words an answer should use: intrusion detection is based on the assumption that the behaviour of an intruder differs from that of a legitimate user in ways that can be measured, although the two overlap. Statistical anomaly detection collects data on the behaviour of legitimate users over a period and applies statistical tests to decide whether new behaviour is legitimate; it comes as threshold detection, with thresholds on the frequency of events independent of the user, and profile-based detection, with a profile of activity for each user. Rule-based detection defines a set of rules that decide whether behaviour is that of an intruder; it comes as rule-based anomaly detection, rules drawn from past usage patterns, and rule-based penetration identification, rules that describe known attacks and suspicious behaviour.
What intrusion detection is for
NIST's guide, SP 800-94 (2007), defines it as "the process of monitoring the events occurring in a computer system or network and analyzing them for signs of possible incidents". Prevention fails sometimes, so detection matters for three reasons. An intrusion found quickly can be stopped before much damage is done. A system known to be watched deters some intruders. And what detection records shows how the intruder got in, which is how prevention improves.
Intrusion prevention goes one step further: it is "the process of performing intrusion detection and attempting to stop detected possible incidents", for example by dropping the attacker's traffic.
Statistical anomaly detection
Threshold detection counts events of a type over a period, such as failed logins in a minute, and raises an alarm above a fixed limit, whoever the user is. Dorothy Denning's 1987 paper "An Intrusion-Detection Model", where this approach was set out, calls it the operational model and gives the example of password failures "where more than 10, say, suggests an attempted break-in". It is simple and crude: a limit right for one user is wrong for another.
Profile-based detection builds a profile of each user's normal behaviour from their own past records, and flags departures from it. Denning's profiles are built from three kinds of metric:
| Metric | Measures | Example from Denning |
|---|---|---|
| Event counter | how many audit records of a kind occur in a period | number of logins in an hour; password failures in a minute |
| Interval timer | the time between two related events | time between successive logins to an account |
| Resource measure | the quantity of a resource used by an action | pages printed a day; CPU time used by a program |
Intrusion Detection: Statistical Anomaly and Rule-Based
Textbooks add a fourth, the gauge, a value that can go up as well as down, such as the number of connections a program has open.
To decide whether a new value is abnormal, Denning describes five statistical models:
- Operational: compare with fixed limits, as above.
- Mean and standard deviation: a value is abnormal if it lies more than d standard deviations from the mean. Whatever the data's shape, Chebyshev's inequality says at most 1/d^2 of normal values can fall outside; Denning's example is d = 4, "at most .0625".
- Multivariate: the same, on correlations between two or more metrics, such as CPU time against input and output.
- Markov process: treats each kind of event as a state and learns how often one follows another; a sequence of commands whose transitions are improbable is abnormal.
- Time series: uses the order and timing of values as well, and so can see a gradual shift in behaviour.
Its strength and its weaknesses. Anomaly detection can catch what nobody has seen before: SP 800-94 says it "can be very effective at detecting previously unknown threats". But the profile may be learned while an attacker is already present, and an attacker can "perform small amounts of malicious activity occasionally, then slowly increase the frequency" until a profile that keeps adjusting accepts it as normal.
Rule-based detection
Rule-based anomaly detection writes the past as rules instead of statistics: rules generated from historical audit records describe each user's usual patterns, and current behaviour that matches no rule is flagged.
Rule-based penetration identification writes rules about attacks instead: an expert system with rules describing known penetrations, known weaknesses and suspicious behaviour, such as a user reading files in other users' home directories, or a program that a user has never run appearing in their session. SP 800-94 calls such patterns signatures: "A signature is a pattern that corresponds to a known threat." Signature-based detection "is very effective at detecting known threats but largely ineffective at detecting previously unknown threats", and at variants of known ones.
The run: a profile, a limit, signatures, and a threshold
The listing builds one clerk's profile from 60 past sessions (synthetic data, made up for the illustration), applies Denning's mean and standard deviation model and operational model, runs three signatures taken from SP 800-94's own examples, and then moves the threshold of the anomaly detector to show what every choice costs.
# Intrusion detection two ways: statistical anomaly detection on a user's profile (Denning, 1987),
# and signatures (NIST SP 800-94's own examples). Then the trade-off every detector must choose.
import math, random
rng = random.Random(1987)
# ---- 1. a profile: records a clerk reads per session, over 60 past sessions -----------------
history = [max(0, round(rng.gauss(40, 8))) for _ in range(60)]
n = len(history)
total, squares = sum(history), sum(x * x for x in history)
mean = total / n
stdev = math.sqrt((squares - n * mean * mean) / (n - 1))
D = 4 # Denning's example: 4 standard deviations
low, high = mean - D * stdev, mean + D * stdev
print('1. MEAN AND STANDARD DEVIATION MODEL (60 past sessions)')
print(' mean %.1f records, standard deviation %.1f' % (mean, stdev))
print(' normal range at d = %d: %.1f to %.1f records' % (D, low, high))
print(' Chebyshev: at most %.2f%% of honest sessions can fall outside it' % (100 / D ** 2))
for x in (44, 71, 400):
verdict = 'normal' if low <= x <= high else 'ALARM'
print(' a session reading %3d records: %s' % (x, verdict))
# ---- 2. the operational model: a fixed limit, no history needed -----------------------------
print()
print('2. OPERATIONAL MODEL: more than 10 password failures in a minute')
for failures in (2, 11, 37):
print(' %2d failures: %s' % (failures, 'ALARM' if failures > 10 else 'normal'))
# ---- 3. signatures: patterns of known attacks (SP 800-94, section 2.3.1) ---------------------
SIGNATURES = {
'telnet login as root': lambda e: e['service'] == 'telnet' and e['user'] == 'root',
'known malware attachment': lambda e: e.get('attachment') == 'freepics.exe',
'auditing disabled (status 645)': lambda e: e.get('status') == 645,
}
events = [
{'service': 'ssh', 'user': 'chetan'},
{'service': 'telnet', 'user': 'root'},
{'service': 'mail', 'user': 'asha', 'attachment': 'freepics.exe'},
{'service': 'mail', 'user': 'asha', 'attachment': 'freepics2.exe'},
{'service': 'log', 'user': 'system', 'status': 645},
]
print()
print('3. SIGNATURE-BASED DETECTION')
for e in events:
hits = [name for name, rule in SIGNATURES.items() if rule(e)]
what = (e.get('attachment') or ('status %d' % e['status'] if 'status' in e else '')
or e['service'] + ' as ' + e['user'])
print(' %-16s %s' % (what, ', '.join(hits) if hits else 'no signature matches'))
# ---- 4. the trade-off: move the threshold, and both rates move --------------------------------
honest = [max(0, round(rng.gauss(40, 8))) for _ in range(10_000)]
attacks = [max(0, round(rng.gauss(75, 25))) for _ in range(500)]
print()
print('4. ONE DETECTOR, FIVE THRESHOLDS (10,000 honest sessions, 500 intrusions)')
print(' d alarm above intrusions caught honest sessions flagged')
for d in (1, 1.5, 2, 3, 4):
limit = mean + d * stdev
caught = sum(x > limit for x in attacks) / len(attacks)
flagged = sum(x > limit for x in honest) / len(honest)
print(' %-5s %11.1f %18.1f%% %24.2f%%' % (d, limit, 100 * caught, 100 * flagged))Intrusion Detection: Statistical Anomaly and Rule-Based
1. MEAN AND STANDARD DEVIATION MODEL (60 past sessions)
mean 40.3 records, standard deviation 8.3
normal range at d = 4: 7.0 to 73.5 records
Chebyshev: at most 6.25% of honest sessions can fall outside it
a session reading 44 records: normal
a session reading 71 records: normal
a session reading 400 records: ALARM
2. OPERATIONAL MODEL: more than 10 password failures in a minute
2 failures: normal
11 failures: ALARM
37 failures: ALARM
3. SIGNATURE-BASED DETECTION
ssh as chetan no signature matches
telnet as root telnet login as root
freepics.exe known malware attachment
freepics2.exe no signature matches
status 645 auditing disabled (status 645)
4. ONE DETECTOR, FIVE THRESHOLDS (10,000 honest sessions, 500 intrusions)
d alarm above intrusions caught honest sessions flagged
1 48.6 87.8% 14.67%
1.5 52.7 84.6% 5.98%
2 56.9 81.2% 2.13%
3 65.2 69.2% 0.10%
4 73.5 55.6% 0.00%Intrusion Detection: Statistical Anomaly and Rule-Based
What the run establishes, in order.
The profile catches the large departure and passes the small one. At d = 4 the clerk's normal range is 7 to 73.5 records a session. A session reading 400 records, a misfeasor copying the database, raises an alarm; one reading 71 does not, although it is well above average. Chebyshev guarantees that at most 6.25 per cent of honest sessions can fall outside the range, whatever the shape of the data.
A fixed limit needs no history at all. Eleven password failures in a minute raise an alarm for anyone. That is the operational model's appeal and its weakness.
A signature catches exactly what it describes, and nothing else. Telnet as root, the known attachment and the log code for auditing switched off each match. The attachment renamed "freepics2.exe" matches nothing, which is SP 800-94's own illustration of how easily a signature is evaded.
Every threshold is a trade. At d = 1 the detector catches 87.8 per cent of intrusions but flags 14.67 per cent of honest sessions; at d = 4 it flags none of 10,000 honest sessions but catches only 55.6 per cent of intrusions. Plotting the share of intrusions caught against the share of honest activity flagged, for every possible threshold, gives what engineers call the receiver operating characteristic, or ROC curve. There is no setting that is right for everyone: an operator chooses a point on the curve, and the base-rate arithmetic of the intruders chapter shows why that point must lie very near zero false alarms.
Four kinds of IDS, by what they watch
SP 800-94 sorts intrusion detection and prevention systems by "the types of events that they monitor and the ways in which they are deployed":
| Kind | Watches | Sees |
|---|---|---|
| Network-based | the traffic on a network segment | attacks carried in network and application protocols |
| Wireless | the radio traffic of wireless networks | attacks on the wireless protocols themselves, rogue access points |
| Network behaviour analysis | the pattern of flows across a network | unusual traffic flows, such as denial of service, some malware, policy violations |
| Host-based | one computer and the events on it | what happens inside that host: logs, files, processes |
Intrusion Detection: Statistical Anomaly and Rule-Based
"For most environments, a combination of network-based and host-based IDPS technologies is needed", SP 800-94 advises. The next chapter follows the host-based side into its audit records and the network side into distributed IDS and honeypots.
Distinctions that carry marks
| Statistical anomaly detection | Rule-based penetration identification | |
|---|---|---|
| Compares behaviour with | what is normal for this user or system | patterns of known attacks and misuse |
| Needs | a period of learning | experts to write rules |
| Catches unknown attacks | yes | no |
| Main weakness | false alarms; profiles poisoned or slowly drifted | misses new attacks and simple variants |
| Threshold detection | Profile-based detection | |
|---|---|---|
| Limits set | once, for everyone | from each user's own history |
| Denning's model | operational | mean and standard deviation, multivariate, Markov, time series |
What beginners get wrong here
Confusing rule-based anomaly detection with penetration identification. Both use rules, but the first describes normal use and flags departures; the second describes attacks and flags matches.
Saying anomaly detection finds only known attacks. It is the signature approach that needs to know the attack; anomaly detection's advantage is that it does not.
Believing a stricter threshold is simply better. Lowering the threshold catches more intrusions and raises more false alarms; the two always move together.
Treating an IDS as protection. An intrusion detection system reports; only an intrusion prevention system tries to stop an attack, and even it cannot stop what it does not detect.
Quick revision
- Intrusion detection: monitoring events and analysing them "for signs of possible incidents" (SP 800-94). Why: stop intrusions early, deter, learn.
- Statistical anomaly: threshold detection (fixed limits, all users) and profile-based (per user).
- Denning, 1987: metrics event counter, interval timer, resource measure (textbooks add gauge); models operational, mean and standard deviation, multivariate, Markov process, time series. Chebyshev: at most 1/d^2 outside d standard deviations.
- Rule-based: anomaly detection (rules of past behaviour) and penetration identification (rules of known attacks, the signatures).
- Signatures miss variants: "freepics2.exe".
- ROC: intrusions caught against honest activity flagged, for every threshold.
- SP 800-94's four kinds: network-based, wireless, network behaviour analysis, host-based.
Test yourself
1. Distinguish statistical anomaly detection from rule-based detection. Statistical anomaly detection collects data on the behaviour of legitimate users over time and applies statistical tests to decide whether new behaviour is normal; it can detect attacks never seen before but raises false alarms and can be fooled by gradual change. Rule-based detection applies a set of rules: either rules derived from past usage, flagging departures, or rules describing known attacks and suspicious behaviour, flagging matches. The second form catches known attacks reliably but misses new ones and simple variants.
Intrusion Detection: Statistical Anomaly and Rule-Based
2. Name and explain the metrics used in profile-based intrusion detection. Denning defines three. An event counter is the number of audit records of some kind in a period, such as logins in an hour. An interval timer is the time between two related events, such as successive logins. A resource measure is the quantity of a resource used by an action, such as pages printed in a day or CPU time used by a program. Textbooks add a gauge, a value that can rise and fall, such as the number of open connections.
3. Explain the mean and standard deviation model with its guarantee. From past observations of a metric the system computes the mean and standard deviation, and defines a new observation as abnormal if it lies more than d standard deviations from the mean. By Chebyshev's inequality, whatever the distribution of the data, at most 1/d^2 of normal observations can lie outside that interval; for d = 4 that is at most 6.25 per cent. The model needs no prior knowledge of normal activity, learns from each user's own data, and so can treat as normal for one user what is abnormal for another.
4. What is rule-based penetration identification, and what is its weakness? It is an expert-system approach in which rules describe known attacks, known weaknesses and suspicious behaviour, and activity matching a rule is reported as a likely intrusion; these rules are what NIST calls signatures. Its weakness is that it can only find what its rules describe: a new attack, or a small variant of a known one such as a renamed file, matches no rule and passes unnoticed.
5. What does moving the threshold of an anomaly detector do? A lower threshold flags more activity, catching more intrusions but also flagging more honest activity as false alarms; a higher threshold does the reverse. Each threshold gives one pair of rates, and the curve through all of them is the receiver operating characteristic. Because intrusions are rare compared with honest activity, the threshold must usually be set to keep false alarms very low, at the cost of missing some intrusions.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.