Directory Traversal
Chapter Eighty
Syllabus topic Module 2, "Web Server Vulnerabilities and Hardening Techniques: ... directory traversal attacks"
Pages 378 to 381 of 578
In one line
Directory traversal escapes the folder a web server is meant to serve from, by putting "go up a level" sequences into input that the application turns into a file path, reaching files elsewhere on the server. The robust fix is not to build file paths from user input at all.
In examination wording: directory traversal, or path traversal, is an attack that accesses files outside the intended directory by supplying input containing path-navigation sequences that the application incorporates into a file path without adequate validation, disclosing or manipulating files the application never intended to expose.
The mechanism, at the level of the path
A web server is meant to serve files from one folder, the web root, and its subfolders. A request for a document maps, somehow, to a file within that area.
The vulnerability arises when the application builds a file path from user input without checking it. If a parameter names a file to serve, and the application takes that parameter and turns it into a path, then an attacker can include the "go up a level" sequence (written ../ on most systems) to climb out of the intended folder.
At the level of the path: if the application intends to serve files from a documents folder, and it constructs the path by joining that folder with the user's parameter, then a parameter packed with repeated ../ sequences walks up out of the documents folder, up out of the web root, and back down to a file elsewhere on the server, a configuration file, a password file, application source, anything the server's account can read.
The essence, and the reason the fix is what it is: the server followed a path it should never have been asked to construct. It did what it was told, and it was told to build a path from input that should not have been trusted to name a path.
What it exposes
The consequence is reading, and sometimes writing, files outside the intended area:
- Configuration files, which may contain database credentials and keys (the stored-credentials finding from earlier chapters, reached through the web).
- Source code, revealing the application's logic and its other flaws.
- System files, such as the file listing user accounts.
- Logs and other data.
In the worst cases, where the flaw allows writing rather than only reading, an attacker can place a file that leads to further compromise. Even read-only, the disclosure of configuration and source is serious, because it hands the attacker credentials and a map of the application.
Why filtering is the wrong fix
The tempting fix is to filter the input: reject any parameter containing ../. This is fragile, and understanding why is the chapter's transferable lesson.
Directory Traversal
An attacker who controls the input can encode the dangerous sequence in ways that a naive filter does not recognise but the system still interprets as "go up": different encodings, doubled or mixed separators, and platform-specific variations. A filter that blocks the plain ../ is bypassed by a form of it the filter did not anticipate but the file system still resolves. This is a recurring theme: filtering a dangerous pattern is weak when the attacker controls the encoding, which the input-validation chapter states as a general principle and which appeared already in the Log4Shell case.
So filtering is a supplementary measure at best. The robust fixes work at a different level.
The robust fixes
The fixes remove the ability to build an arbitrary path from input, rather than trying to catch the bad inputs:
Do not build file paths from user input at all. The strongest fix. Instead of taking a filename from the user and turning it into a path, map the user's choice to a fixed set of allowed files: the user supplies an identifier, and the application looks up which file that identifier corresponds to from a list it controls. The raw input never becomes part of a path, so there is no path to traverse. This removes the class entirely and is the preferred approach.
Canonicalise and validate. Where a path must be derived from input, resolve it to its canonical, absolute form (which collapses all the ../ sequences and encodings into a definite path) and then confirm that the result still lies inside the intended folder before serving it. Anything that resolves outside the folder is rejected. The key is that the check happens after resolution, on the final path, so no encoding trick survives, because they have all been resolved away.
Run the server with least privilege. So that even a successful traversal reaches only what the server's account can read, which should be very little. This does not prevent the traversal but limits its reach, and it is defence in depth: the web-server account should not be able to read the whole file system.
The combination, and the order: map inputs to allowed files where possible; where a path must be built, canonicalise and validate against the intended folder; and run with least privilege as a backstop. Filtering ../ is at most a supplementary measure and never the primary defence.
A worked example, framed defensively
An assessor reviews a document-download feature, using a test account and the application's own files.
- The feature takes a
fileparameter and joins it with the documents folder to build a path. Supplying a value with repeated../sequences resolves to a file outside the documents folder, and the assessor confirms (on a test file they placed for the purpose, not on real sensitive files) that the traversal reaches outside the intended area. Finding: directory traversal. - Testing a filter: the application blocks a plain
../, but an encoded form is not blocked and still traverses. Finding: the filter is bypassable, illustrating why filtering is not the fix. - The web-server account can read broad areas of the file system. Finding: no least privilege, so a traversal reaches more than it should.
Directory Traversal
Recommendations, in order: map the file parameter to an allow-list of documents so no path is built from input (the class-removing fix); where any path is derived from input, canonicalise and validate it against the intended folder; and run the server with least privilege so a traversal reaches little. The assessor demonstrates the flaw on a planted test file rather than by reading the client's actual sensitive files, which is the minimum needed to prove the finding.
What beginners get wrong
- Thinking filtering
../is the fix. Attackers encode the sequence in forms a naive filter misses; filtering is supplementary, and the robust fix is not to build paths from input. - Building a file path directly from a user parameter. That is the flaw; map inputs to allowed files instead.
- Validating before resolving. A check on the raw input misses encoded traversal; canonicalise first, then validate the resolved path against the intended folder.
- Ignoring least privilege. A web-server account that can read the whole file system turns a traversal into a broad disclosure; restrict it.
- Treating read-only traversal as minor. Disclosure of configuration and source hands over credentials and a map of the application, which is serious.
- Confusing the layer. Traversal can arise from server configuration or application code; note which, and fix at the right level.
Quick revision
- Directory traversal escapes the web root by putting "go up a level" (
../) sequences into input the application turns into a file path, reaching files elsewhere: configuration (credentials), source, system files. - The essence: the server followed a path it should never have been asked to construct from untrusted input.
- Filtering
../is fragile, because the attacker can encode the sequence in forms a naive filter misses but the system still resolves; it is supplementary at best. - Robust fixes, in order: map inputs to a fixed allow-list of files so no path is built from input (removes the class); canonicalise and validate any derived path against the intended folder, checking after resolution; and run the server with least privilege so a traversal reaches little.
Test yourself
- Explain directory traversal at the level of the path.
Directory Traversal
The application builds a file path by combining an intended folder with user-supplied input, and the input contains "go up a level" sequences such as ../. These cause the constructed path to climb out of the intended folder and the web root and back down to a file elsewhere on the server, so the server serves a file outside the area it was meant to, such as a configuration file or source code, because it followed a path built from untrusted input.
- What does a successful directory traversal expose, and why is even read-only traversal serious?
It exposes files outside the intended area: configuration files that may contain database credentials and keys, application source code revealing logic and other flaws, system files such as the account list, and logs. Even read-only, this is serious because disclosing configuration hands over credentials and disclosing source provides a map of the application and its weaknesses, materially advancing an attack.
- Why is filtering the
../sequence a weak defence?
Because an attacker who controls the input can encode the "go up" sequence in forms a naive filter does not recognise, such as alternative encodings, doubled or mixed separators, or platform-specific variants, which the file system still resolves as traversal. Blocking the plain sequence is therefore bypassed by a form the filter did not anticipate, which is why pattern filtering is at most supplementary when the attacker controls the encoding.
- What is the strongest fix for directory traversal, and why?
Not building file paths from user input at all: mapping the user's choice to a fixed set of allowed files, so the user supplies an identifier that the application looks up against a list it controls, and the raw input never becomes part of a path. This removes the class entirely, because with no path constructed from input there is nothing to traverse.
- When a path must be derived from input, how should it be validated, and what backstop limits the damage?
The derived path should be resolved to its canonical absolute form, which collapses all navigation sequences and encodings into a definite path, and then checked to confirm it still lies inside the intended folder, rejecting anything that resolves outside; the validation must occur after resolution so no encoding trick survives. Running the server with least privilege is the backstop, so that even a successful traversal reaches only the little the server's account can read.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.