Search-Engine Reconnaissance and Document Metadata
Chapter Twenty-Two
Syllabus topic Module 1, "Footprinting and Information Gathering Methodology: ... search engine reconnaissance strategies"
Pages 101 to 105 of 578
In one line
Search engines index far more than an organisation intends, and precise queries retrieve the parts that were never meant to be found. Published documents carry hidden metadata, author names, internal paths, software versions, that the author did not know they were sending.
In examination wording: search-engine reconnaissance uses advanced search operators to locate indexed content that an organisation did not intend to expose, such as directory listings, administrative interfaces, configuration files and error messages; document metadata is descriptive information embedded in files by the software that created them, which may disclose user names, internal file paths, software versions and organisational structure.
Why a search engine knows things you never published
The confusion to clear first: organisations think of their website as the pages they linked, and a search engine's index as a copy of those pages. Neither is quite true.
A crawler follows links, but it also finds content in other ways: a link from somebody else's site, a URL that appeared in a public document, a host name discovered from certificate transparency, or a sitemap listing pages nobody links. So a page can be indexed without ever being linked from the site. "Nobody knows the address" is not a security control, and a page reachable by anyone who knows the URL is, in practice, a public page.
Worse, the index persists. Content removed from a site can remain in a search engine's cache for a while, and in web archives indefinitely. Deleting a page does not un-publish what it said.
What search-engine reconnaissance actually does
The technique uses the advanced operators that search engines provide, which are ordinary documented features intended for legitimate use. In general terms they let a searcher restrict results by:
- Site or domain, to see everything indexed for an organisation, including subdomains the searcher did not know existed.
- File type, to find documents rather than pages: spreadsheets, presentations, configuration files, backups, databases.
- Words in the URL or the title, which is how directory listings and administrative interfaces are located, since both produce characteristic titles.
- Words in the body, to find pages containing error messages, credentials or internal terminology.
Combining these is what makes it powerful. Restricting to an organisation's domain, to a document file type, and to a word that appears in sensitive documents will surface material the organisation never intended to publish. The practice of building such queries is widely documented and there are public collections of them.
What it typically finds:
- Directory listings, where a folder with no index page displays its contents, exposing backups, uploads, and files never meant to be reachable.
- Documents that were uploaded for one recipient and left in a public folder: price lists, internal procedures, staff lists, scanned forms.
- Configuration and backup files left in the web root, which may contain database credentials.
- Administrative login pages that were never linked but were indexed.
- Error messages containing software versions and internal file paths.
- Old content that still names systems, staff or suppliers.
Search-Engine Reconnaissance and Document Metadata
All of this is passive by the definition of the earlier chapter: the searcher queries the search engine, not the target, so the target's logs record nothing. That is exactly why it is done first.
Document metadata
The second half of the chapter is the same problem arriving by a different route: information published inside files rather than on pages.
When software creates a document it records descriptive information in the file. Depending on the format and the tool, that can include:
- the author's name and often their user name on the machine, which reveals the organisation's naming convention for accounts;
- the organisation name registered in the software;
- creation and modification times, and the identities of everyone who edited it;
- the software and version used, which is a technology-profile item;
- file paths, which disclose internal server names and folder structures, for example a path showing a departmental file server's name;
- for images, the camera or device, and sometimes location coordinates;
- in some formats, content that was deleted but not removed, such as text hidden behind a redaction drawn as a shape, or earlier revisions retained in the file.
That last one is the most serious and the least understood. A redaction applied as a black rectangle over text in some document formats hides the text visually while leaving it in the file, so it can be recovered by anyone. Documents have been published in this state by organisations that believed they had removed the information.
An attacker who collects an organisation's published documents and extracts the metadata gains: a list of user names, the account naming convention, internal server and path names, and the software versions in use. Tools to do this in bulk are well known, and the whole operation is passive.
The defences
The two halves share a shape: the information was published by accident, so the control is process rather than technology alone.
For search-engine exposure:
- The real fix is that sensitive material must not be reachable. If a document requires authorisation, it must be behind authentication, not merely unlinked. A crawler directive asking search engines not to index a page is a request to well-behaved crawlers, not an access control, and it does not stop anyone who has the URL.
- Disable directory listing, which is a web-server setting and one of the commonest real findings.
- Keep non-public files out of the web root entirely, including backups, exports and configuration.
- Use generic error pages in production.
- Search for your own organisation regularly, using the same operators an attacker would. This is the single most useful action in the chapter, because it finds what is already exposed.
- Request removal of content that is indexed and should not be, and, because the index reflects the site, remove or protect the underlying content first, or it will simply be re-indexed.
Search-Engine Reconnaissance and Document Metadata
A caution worth stating plainly: a crawler-exclusion file that lists the paths you do not want indexed is itself readable, so a file listing your sensitive directories is a map for an attacker. Excluding a path is not protecting it.
For document metadata:
- Strip metadata before publishing. Most office and PDF tools have a document-inspection feature that removes personal and hidden information; for a publishing workflow this should be a required step rather than an optional one.
- Redact properly. Remove the underlying text, do not cover it. Use the software's redaction feature or flatten the document, and verify by searching the published file for the text that should be gone.
- Prefer exported formats for publication, since converting often discards revision history and comments.
- Check images for embedded location and device data.
- Make it a gate. Since this is an accident, the control must be part of the process that publishes, not a reminder to be careful.
A worked example
Anjali audits her college's published material, entirely passively.
- Restricting a search to the college domain and to spreadsheet files returns a document listing student names with enrolment numbers, uploaded two years ago to a folder that was never linked. Finding: personal data is publicly retrievable; it is indexed, so the URL is effectively public.
- A query for the characteristic title of a directory listing, restricted to the domain, returns an open
/uploads/folder containing scanned forms. Finding: directory listing is enabled. - Downloading three published PDFs and inspecting their metadata gives four staff user names in the form
surname_initial, the internal path\\\\fileserver01\\admin\\, and the version of the software used. Finding: account naming convention, an internal server name and a software version are disclosed. - One PDF has a section covered by a black rectangle; selecting the text underneath recovers it. Finding: an improper redaction has published the information it was meant to hide.
Her recommendations: remove the student spreadsheet and treat it as a personal-data incident (not merely a hygiene issue); disable directory listing and move uploads out of the web root; require metadata stripping before publication; re-issue the improperly redacted document and withdraw the original; and add a quarterly self-search using the same queries.
Note that the first and last findings need different responses. Removing the file stops future retrieval but does not undo the disclosure, since it has been indexed and may be cached or archived. Confidentiality, once lost, cannot be restored, which is the point made in the CIA chapter now arriving in a concrete form.
Search-Engine Reconnaissance and Document Metadata
What beginners get wrong
- Believing an unlinked page is private. It can be indexed from other sources, and a URL known to anyone is effectively public. Authentication is the control.
- Using a crawler-exclusion file as protection. It is a request to polite crawlers, it does not restrict access, and the file itself advertises the paths you consider sensitive.
- Thinking removal undoes exposure. Caches and archives persist, and anything already retrieved is gone. Removal stops the bleeding; it does not heal.
- Redacting by drawing over text. In many formats the text remains in the file. Redaction must remove the content, and the result should be verified.
- Treating metadata as trivia. It supplies user names, the account naming convention, internal server paths and software versions, all of which feed later phases.
- Not searching for yourself. The technique is available to defenders at no cost, and it is the fastest way to find what is already exposed.
Quick revision
- Search engines index content that was never linked, so "nobody knows the URL" is not a control; indexes and archives also persist after removal.
- Advanced operators restrict by site, file type, words in the URL or title, and words in the body; combined, they surface directory listings, stray documents, configuration and backup files, unlinked administrative pages and error messages. Querying a search engine is passive.
- Document metadata can disclose author and user names, the organisation, edit history, software versions, internal file paths, image location data, and text that was hidden rather than removed (improper redaction).
- Defences: authenticate what must be protected; disable directory listing; keep non-public files out of the web root; generic error pages; search for yourself regularly; strip metadata and redact properly as a required publishing step, not a reminder.
- A crawler-exclusion file is readable and lists your sensitive paths; excluding is not protecting.
Test yourself
- Why can a page an organisation never linked still appear in a search engine, and what follows for security?
Because crawlers find URLs from other sites, public documents, sitemaps and certificate data, so a page can be indexed without being linked. It follows that obscurity of the address is not a control; anything that must be restricted needs authentication.
- What does a crawler-exclusion file actually do, and why can it make things worse?
It asks well-behaved crawlers not to index the listed paths; it does not prevent access by anyone who has or guesses the URL. It can make things worse because the file is itself publicly readable, so it advertises exactly which paths the organisation considers sensitive.
Search-Engine Reconnaissance and Document Metadata
- Name four kinds of information document metadata can disclose.
The author's name and machine user name (revealing the account naming convention), the organisation and edit history, the software and version used, and internal file paths and server names; images may also carry device and location data.
- Why is drawing a black rectangle over text an unsafe way to redact a document?
Because in many formats the underlying text remains in the file and is merely hidden visually, so it can be selected, copied or extracted by anyone who receives the document. Proper redaction removes the content, and the published file should be verified afterwards.
- An indexed spreadsheet of personal data is discovered and removed. Why is the problem not fully solved?
Because the content has already been indexed and may persist in caches and web archives, and anyone who retrieved it retains it. Removal prevents further retrieval from the source but cannot undo the disclosure, which is why confidentiality loss is treated as irreversible and, for personal data, as an incident.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.