Practical 9 continued: Extracting Every Hyperlink from an HTML File
Chapter Eighteen
Syllabus topic Module 1, practical 9(b), "Write a program to extract all hyperlinks (<a href=\"...\">) from an HTML file"
Pages 106 to 112 of 297
Aim
To extract all the hyperlinks from an HTML file.
The file every listing below uses
<!doctype html>
<html>
<head><title>munotes sample page</title></head>
<body>
<h1>Links of every awkward shape</h1>
<a href="https://www.munotes.in/notes/">Notes, double quoted</a>
<a href='https://www.munotes.in/syllabus/'>Syllabus, single quoted</a>
<a href=https://www.munotes.in/aibe/>AIBE, not quoted at all</a>
<a class="button" href="/for-committees/">href is the SECOND attribute</a>
<a
href="/solved-papers/">the tag is spread over two lines</a>
<A HREF="/blog/">the tag is in capitals</A>
<a name="bookmark">an anchor with no href at all</a>
<a href="#top" title="to the top">a fragment</a>
<img src="picture.png" alt="an image, not a link">
</body>
</html>Nine a tags, and one of them has no href. Before reading on, decide how many hyperlinks that file contains.
The naive pattern, and what it misses
import re
with open("page.html") as handle:
html = handle.read()
naive = re.findall(r'<a href="(.*?)">', html)
print("the naive pattern found", len(naive), "link(s):")
for link in naive:
print(" ", link)the naive pattern found 2 link(s):
https://www.munotes.in/notes/
#top" title="to the topTwo results out of seven links, and look at the second one: it is not an address at all. It is #top" title="to the top, because the pattern insists on "> straight after the address and the nearest "> on that line is the one at the end of the whole tag. So the naive pattern does not only miss links, it also reports a wrong one.
Everything it missed is ordinary HTML that a browser accepts without complaint:
| Missed or mangled | Why |
|---|---|
href='...' | the pattern insists on double quotes |
href=... with no quotes | the same |
<a class="button" href="..."> | the pattern insists href comes first |
| a tag split over two lines | the pattern has a literal space after a |
<A HREF="..."> | the pattern is case sensitive |
href="#top" title="..."> | mangled: it ran on to the "> at the end of the tag |
That is why this exercise is worth doing properly: the pattern that looks right finds a minority of the links and corrupts one of the few it does find.
Building the pattern that works
Take the requirements one at a time.
import re
sample = '''<a href="one.html">1</a> <a href='two.html'>2</a> <a href=three.html>3</a>
<a class="x" href="four.html">4</a> <A HREF="five.html">5</A>'''
print("1 double quotes only :", re.findall(r'href="(.*?)"', sample))
print("2 either quote :", re.findall(r'href=["\'](.*?)["\']', sample))
print("3 or no quote at all :", re.findall(r'href=["\']?([^"\'>\s]+)', sample))
print("4 ignoring case :", re.findall(r'href=["\']?([^"\'>\s]+)', sample, re.IGNORECASE))1 double quotes only : ['one.html', 'four.html']
2 either quote : ['one.html', 'two.html', 'four.html']
3 or no quote at all : ['one.html', 'two.html', 'three.html', 'four.html']
4 ignoring case : ['one.html', 'two.html', 'three.html', 'four.html', 'five.html']Four lines, and each one recovers a family of links the line above it lost.
| Piece | What it does |
|---|---|
href= | the attribute name and the equals sign |
["\']? | an optional double or single quote |
(...) | a capturing group: findall returns what is inside it |
[^"\'>\s]+ | one or more characters that are not a quote, a > or whitespace |
re.IGNORECASE | so HREF is found too |
Practical 9 continued: Extracting Every Hyperlink from an HTML File
The capturing group is why findall returns the address and not the whole tag. With a group in the pattern, findall gives you the group; with no group, it gives the whole match. That is the single most useful thing to know about findall.
Greedy against lazy, which the naive pattern got right by luck
import re
line = '<a href="first.html">one</a> and <a href="second.html">two</a>'
print("greedy .* :", re.findall(r'href="(.*)"', line))
print("lazy .*? :", re.findall(r'href="(.*?)"', line))greedy .* : ['first.html">one</a> and <a href="second.html']
lazy .*? : ['first.html', 'second.html'] is greedy: it takes as much as it can and then gives back only as much as it must, so . ran from the first quote all the way to the last one on the line and swallowed the tag in between. *? is lazy: it takes as little as possible, so it stops at the first closing quote.
Any pattern that matches up to a closing delimiter needs the lazy form, or the first and the last delimiter on the line are paired together. It is the commonest bug in this exercise after the quoting.
The character class [^"\'>\s]+ used in the answer above sidesteps the question entirely, because it cannot cross a quote or a > at all. That is why it is the better pattern: it does not depend on getting greediness right.
The answer
import re
pattern = re.compile(r'''<a\s[^>]*?href\s*=\s*(["']?)([^"'>\s]+)\1''', re.IGNORECASE | re.DOTALL)
with open("page.html") as handle:
html = handle.read()
links = [match.group(2) for match in pattern.finditer(html)]
print(f"{len(links)} hyperlink(s) found:")
for number, link in enumerate(links, 1):
print(f" {number}. {link}")7 hyperlink(s) found:
1. https://www.munotes.in/notes/
2. https://www.munotes.in/syllabus/
3. https://www.munotes.in/aibe/
4. /for-committees/
5. /solved-papers/
6. /blog/
7. #topEvery piece of that pattern earns its place:
| Piece | Why |
|---|---|
<a\s | an a tag, and the \s stops it matching <abbr or <article |
[^>]*? | any other attributes first, but not past the end of the tag |
href\s=\s | spaces are allowed around the equals sign |
(["']?) | group 1: the opening quote, or nothing |
([^"'>\s]+) | group 2: the address itself |
\1 | a back reference: the same quote as group 1 closed it |
re.IGNORECASE | <A HREF= |
re.DOTALL | so [^>]*? may cross the newline in the tag split over two lines |
'''...''' | triple quotes, so both " and ' can appear without escaping |
\1 is the clever part. It means "whatever group 1 matched", so a link opened with a double quote must close with a double quote, one opened with a single quote with a single quote, and one opened with nothing closes with nothing. One pattern, three quoting styles, and no chance of pairing a single quote with a double one.
Practical 9 continued: Extracting Every Hyperlink from an HTML File
Notice what the output does not contain: the <a name="bookmark"> tag, because it has no href, and the <img src=...>, because it is not an a tag. Both are in the file on purpose.
The same job with the standard library parser
A regular expression is what MU's row asks for. A parser is what the job actually deserves, and being able to say why is worth a mark.
from html.parser import HTMLParser
class LinkFinder(HTMLParser):
"""Collects the href of every anchor tag."""
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag == "a":
for name, value in attrs:
if name == "href":
self.links.append(value)
with open("page.html") as handle:
html = handle.read()
finder = LinkFinder()
finder.feed(html)
print(f"the parser found {len(finder.links)} link(s):")
for link in finder.links:
print(" ", link)the parser found 7 link(s):
https://www.munotes.in/notes/
https://www.munotes.in/syllabus/
https://www.munotes.in/aibe/
/for-committees/
/solved-papers/
/blog/
#topHTMLParser is in the standard library, so nothing needs installing. It reads the document as a document: it lower cases the tag and the attribute names itself, it handles every quoting style because that is the specification it implements, and it cannot be fooled by a > inside an attribute value.
Why the parser is right and the pattern is not, in one example
import re
from html.parser import HTMLParser
class LinkFinder(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
if tag == "a":
for name, value in attrs:
if name == "href":
self.links.append(value)
tricky = '''<a title="the > sign in a title" href="real.html">a link</a>
<!-- <a href="commented-out.html">this is inside a comment</a> -->
<a href="also-real.html">another</a>'''
by_pattern = re.findall(r'''<a\s[^>]*?href\s*=\s*(["']?)([^"'>\s]+)\1''',
tricky, re.IGNORECASE | re.DOTALL)
finder = LinkFinder()
finder.feed(tricky)
print("the pattern says:", [pair[1] for pair in by_pattern])
print("the parser says :", finder.links)the pattern says: ['commented-out.html', 'also-real.html']
the parser says : ['real.html', 'also-real.html']Read the two lines. The pattern lost real.html, because the > inside the title attribute looked to [^>]*? like the end of the tag. And it found commented-out.html, which is inside an HTML comment and is not a link on the page at all.
HTML is not a regular language, so no regular expression can parse it. That is the honest answer, and the right thing to write in the journal: the pattern is what the exercise asks for and it works on ordinary pages; a parser is what to reach for when the input is not under your control. Saying that, rather than defending the pattern, is what a good answer looks like.
What to do with the links once you have them
import re
from urllib.parse import urljoin, urlparse
with open("page.html") as handle:
html = handle.read()
pattern = re.compile(r'''<a\s[^>]*?href\s*=\s*(["']?)([^"'>\s]+)\1''',
re.IGNORECASE | re.DOTALL)
links = [m.group(2) for m in pattern.finditer(html)]
absolute = [link for link in links if urlparse(link).scheme]
relative = [link for link in links if not urlparse(link).scheme
and not link.startswith("#")]
fragments = [link for link in links if link.startswith("#")]
print("absolute :", absolute)
print("relative :", relative)
print("fragments :", fragments)
print()
base = "https://www.munotes.in/notes/"
print("resolved against", base)
for link in links:
print(" ", urljoin(base, link))Practical 9 continued: Extracting Every Hyperlink from an HTML File
absolute : ['https://www.munotes.in/notes/', 'https://www.munotes.in/syllabus/', 'https://www.munotes.in/aibe/']
relative : ['/for-committees/', '/solved-papers/', '/blog/']
fragments : ['#top']
resolved against https://www.munotes.in/notes/
https://www.munotes.in/notes/
https://www.munotes.in/syllabus/
https://www.munotes.in/aibe/
https://www.munotes.in/for-committees/
https://www.munotes.in/solved-papers/
https://www.munotes.in/blog/
https://www.munotes.in/notes/#topurlparse splits an address into its parts, and a scheme such as https is what makes an address absolute. urljoin resolves a relative address against the page it was found on, which is what a browser does and what a crawler has to do. Neither needs installing.
The link text as well as the address
import re
with open("page.html") as handle:
html = handle.read()
pattern = re.compile(
r'''<a\s[^>]*?href\s*=\s*(["']?)([^"'>\s]+)\1[^>]*>(.*?)</a\s*>''',
re.IGNORECASE | re.DOTALL)
for match in pattern.finditer(html):
address = match.group(2)
words = re.sub(r"\s+", " ", match.group(3)).strip()
print(f" {address:<36} {words}") https://www.munotes.in/notes/ Notes, double quoted
https://www.munotes.in/syllabus/ Syllabus, single quoted
https://www.munotes.in/aibe/ AIBE, not quoted at all
/for-committees/ href is the SECOND attribute
/solved-papers/ the tag is spread over two lines
/blog/ the tag is in capitals
#top a fragmentGroup 3 is the text between the opening and the closing tag, and re.sub(r"\s+", " ", ...) folds any run of whitespace into one space, which is what a browser does when it displays it. The closing </a\s*> allows a space before the >.
Procedure
- Create
page.htmlwith links using double quotes, single quotes, no quotes,hrefas a later
attribute, a tag split across two lines, a tag in capitals, an a tag with no href, and an img tag. That file is the exercise.
- Save as
practical9b.py. Run the naive patternr'<a href="(.*?)">'first and record how many
of your links it missed.
- Note that the naive pattern also returns one address that is not an address, and say why.
- Build the pattern up in four steps, printing
findallat each: double quotes only, either
quote, quotes optional, then case insensitive.
- Show greedy against lazy on one line holding two links.
- Write the answer with the back reference
\1and both flags, and print every link numbered. - Confirm that the
atag with nohrefand theimgtag are not in the output. - Write the
HTMLParserversion and show it gives the same answer. - Run both over the tricky sample with a
>inside an attribute and a link inside a comment, and
record the two disagreements.
- Split the links into absolute, relative and fragment, and resolve them with
urljoin.
Practical 9 continued: Extracting Every Hyperlink from an HTML File
Result
The naive pattern returned two results out of seven links: it missed the single quoted, unquoted, later-attribute, multi-line and capitalised ones, and it mangled the fragment link into #top" title="to the top by running on to the "> at the end of the tag. The four step build recovered each family in turn. Greedy . swallowed the tag between two links on one line and lazy .? did not. The final pattern, with the back reference and both flags, found every hyperlink in page.html and did not report the a tag with no href or the img tag. HTMLParser gave the same list. On the tricky sample the pattern lost the link whose title attribute contained a > and wrongly reported a link inside an HTML comment, while the parser got both right.
Where marks are lost
- Only handling double quotes. Single quotes and no quotes are both ordinary HTML.
- Assuming
hrefis the first attribute.<a class="..." href="...">is everywhere. - A case sensitive pattern.
<A HREF=is valid. - Greedy
.instead of.?, which pairs the first quote with the last on the line. - No capturing group, so
findallreturns whole tags instead of addresses. - Reporting
<a name="...">as a link. It has nohref. - Reporting
<img src="...">. It is not a hyperlink. - Claiming a regular expression parses HTML. It does not, and saying so is the better answer.
- Not showing the naive pattern failing. The comparison is the evidence.
For the journal
The aim in MU's words, with her <a href="..."> copied exactly. page.html written out as you made it, and one line saying which awkward shape each link is there to test. The naive pattern and its output, with a count of what it missed. The four step build with the output at each step. The greedy against lazy pair. Then the answer, with the table explaining each piece of the pattern and the numbered list of links. One sentence on the back reference: \1 is whatever group 1 matched, so the closing quote is forced to be the same as the opening one. Then the parser version and the tricky sample where the two disagree, and the conclusion: HTML is not a regular language, so the pattern is right for an ordinary page and a parser is right when the input is not under your control.
Quick revision
- The naive
r'<a href="(.*?)">'misses single quotes, no quotes, a laterhref, a multi-line
tag and capitals, and mangles a link followed by another attribute.
- A capturing group
(...)makesfindallreturn the group instead of the whole match. ["\']?is an optional quote of either kind.[^"'>\s]+is an address: no quote, no>, no
Practical 9 continued: Extracting Every Hyperlink from an HTML File
space.
is greedy and?is lazy. Matching up to a delimiter needs the lazy form.\1is a back reference to group 1, which forces the closing quote to match the opening
one.
re.IGNORECASEfor<A HREF=,re.DOTALLso a class may cross a newline.- Use
'''triple quotes'''for a pattern containing both quote characters. <a\snot<a, or the pattern also matches<abbrand<article.html.parser.HTMLParserwithhandle_starttagdoes the job properly, and is in the standard
library.
- HTML is not a regular language. A pattern fails on a
>inside an attribute and on a link
inside a comment. Say so.
urlparse(link).schemetells absolute from relative;urljoin(base, link)resolves it.
Questions you should be able to answer
1. What is wrong with r'<a href="(.*?)">' on an ordinary page? It requires double quotes, requires href to be the first attribute, is case sensitive and cannot cross a newline. It also returns a wrong address when another attribute follows, because it runs on to the "> at the end of the tag.
2. What does a capturing group do to findall? It makes findall return what the group matched rather than the whole match. With two groups it returns tuples.
3. What is the difference between . and .? . is greedy and takes as much as it can; .? is lazy and takes as little as it can. Matching up to a closing quote needs the lazy form.
4. What does \1 mean in a pattern? Whatever group 1 matched. Here it forces the closing quote to be the same character as the opening one.
5. Why <a\s rather than <a? Because <a also matches the start of <abbr and <article. The \s requires whitespace after the tag name.
6. Why is [^"'>\s]+ better than .*? for the address? Because it cannot cross a quote, a > or a space at all, so it does not depend on getting greediness right.
7. What does re.DOTALL change? It lets . and a negated class match a newline, so a tag spread over two lines is still one match.
8. Give two cases where the pattern is wrong and the parser is right. A > inside an attribute value, which ends the tag as far as the pattern is concerned, and a link inside an HTML comment, which the pattern reports and the parser ignores.
9. Why can no regular expression parse HTML in general? Because HTML is not a regular language: it nests, and comments, attribute values and character references can all contain the characters a pattern uses as landmarks.
Practical 9 continued: Extracting Every Hyperlink from an HTML File
10. Which of these is a hyperlink: <a name="top">, <img src="x.png">, <a href="#top">? Only the last. The first has no href and the second is not an a tag.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.