munotes®

Practical 9 continued: Extracting Every Hyperlink from an HTML File

Chapter Eighteen

Syllabus topic Module 1, practical 9(b), "Write a program to extract all hyperlinks (<a href=\"...\">) from an HTML file"

Pages 106 to 112 of 297

Aim

To extract all the hyperlinks from an HTML file.

The file every listing below uses

<!doctype html>
<html>
<head><title>munotes sample page</title></head>
<body>
  <h1>Links of every awkward shape</h1>
  <a href="https://www.munotes.in/notes/">Notes, double quoted</a>
  <a href='https://www.munotes.in/syllabus/'>Syllabus, single quoted</a>
  <a href=https://www.munotes.in/aibe/>AIBE, not quoted at all</a>
  <a class="button" href="/for-committees/">href is the SECOND attribute</a>
  <a
     href="/solved-papers/">the tag is spread over two lines</a>
  <A HREF="/blog/">the tag is in capitals</A>
  <a name="bookmark">an anchor with no href at all</a>
  <a href="#top" title="to the top">a fragment</a>
  <img src="picture.png" alt="an image, not a link">
</body>
</html>

Nine a tags, and one of them has no href. Before reading on, decide how many hyperlinks that file contains.

The naive pattern, and what it misses

import re

with open("page.html") as handle:
    html = handle.read()

naive = re.findall(r'<a href="(.*?)">', html)
print("the naive pattern found", len(naive), "link(s):")
for link in naive:
    print("  ", link)
the naive pattern found 2 link(s):
   https://www.munotes.in/notes/
   #top" title="to the top

Two results out of seven links, and look at the second one: it is not an address at all. It is #top" title="to the top, because the pattern insists on "> straight after the address and the nearest "> on that line is the one at the end of the whole tag. So the naive pattern does not only miss links, it also reports a wrong one.

Everything it missed is ordinary HTML that a browser accepts without complaint:

Missed or mangledWhy
href='...'the pattern insists on double quotes
href=... with no quotesthe same
<a class="button" href="...">the pattern insists href comes first
a tag split over two linesthe pattern has a literal space after a
<A HREF="...">the pattern is case sensitive
href="#top" title="...">mangled: it ran on to the "> at the end of the tag

That is why this exercise is worth doing properly: the pattern that looks right finds a minority of the links and corrupts one of the few it does find.

Building the pattern that works

Take the requirements one at a time.

import re

sample = '''<a href="one.html">1</a> <a href='two.html'>2</a> <a href=three.html>3</a>
<a class="x" href="four.html">4</a> <A HREF="five.html">5</A>'''

print("1 double quotes only  :", re.findall(r'href="(.*?)"', sample))
print("2 either quote        :", re.findall(r'href=["\'](.*?)["\']', sample))
print("3 or no quote at all  :", re.findall(r'href=["\']?([^"\'>\s]+)', sample))
print("4 ignoring case       :", re.findall(r'href=["\']?([^"\'>\s]+)', sample, re.IGNORECASE))
1 double quotes only  : ['one.html', 'four.html']
2 either quote        : ['one.html', 'two.html', 'four.html']
3 or no quote at all  : ['one.html', 'two.html', 'three.html', 'four.html']
4 ignoring case       : ['one.html', 'two.html', 'three.html', 'four.html', 'five.html']

Four lines, and each one recovers a family of links the line above it lost.

PieceWhat it does
href=the attribute name and the equals sign
["\']?an optional double or single quote
(...)a capturing group: findall returns what is inside it
[^"\'>\s]+one or more characters that are not a quote, a > or whitespace
re.IGNORECASEso HREF is found too
munotes.in106

Practical 9 continued: Extracting Every Hyperlink from an HTML File

The capturing group is why findall returns the address and not the whole tag. With a group in the pattern, findall gives you the group; with no group, it gives the whole match. That is the single most useful thing to know about findall.

Greedy against lazy, which the naive pattern got right by luck

import re

line = '<a href="first.html">one</a> and <a href="second.html">two</a>'

print("greedy .*  :", re.findall(r'href="(.*)"', line))
print("lazy   .*? :", re.findall(r'href="(.*?)"', line))
greedy .*  : ['first.html">one</a> and <a href="second.html']
lazy   .*? : ['first.html', 'second.html']

is greedy: it takes as much as it can and then gives back only as much as it must, so . ran from the first quote all the way to the last one on the line and swallowed the tag in between. *? is lazy: it takes as little as possible, so it stops at the first closing quote.

Any pattern that matches up to a closing delimiter needs the lazy form, or the first and the last delimiter on the line are paired together. It is the commonest bug in this exercise after the quoting.

The character class [^"\'>\s]+ used in the answer above sidesteps the question entirely, because it cannot cross a quote or a > at all. That is why it is the better pattern: it does not depend on getting greediness right.

The answer

import re

pattern = re.compile(r'''<a\s[^>]*?href\s*=\s*(["']?)([^"'>\s]+)\1''', re.IGNORECASE | re.DOTALL)

with open("page.html") as handle:
    html = handle.read()

links = [match.group(2) for match in pattern.finditer(html)]

print(f"{len(links)} hyperlink(s) found:")
for number, link in enumerate(links, 1):
    print(f"  {number}. {link}")
7 hyperlink(s) found:
  1. https://www.munotes.in/notes/
  2. https://www.munotes.in/syllabus/
  3. https://www.munotes.in/aibe/
  4. /for-committees/
  5. /solved-papers/
  6. /blog/
  7. #top

Every piece of that pattern earns its place:

PieceWhy
<a\san a tag, and the \s stops it matching <abbr or <article
[^>]*?any other attributes first, but not past the end of the tag
href\s=\sspaces are allowed around the equals sign
(["']?)group 1: the opening quote, or nothing
([^"'>\s]+)group 2: the address itself
\1a back reference: the same quote as group 1 closed it
re.IGNORECASE<A HREF=
re.DOTALLso [^>]*? may cross the newline in the tag split over two lines
'''...'''triple quotes, so both " and ' can appear without escaping

\1 is the clever part. It means "whatever group 1 matched", so a link opened with a double quote must close with a double quote, one opened with a single quote with a single quote, and one opened with nothing closes with nothing. One pattern, three quoting styles, and no chance of pairing a single quote with a double one.

munotes.in107

Practical 9 continued: Extracting Every Hyperlink from an HTML File

Notice what the output does not contain: the <a name="bookmark"> tag, because it has no href, and the <img src=...>, because it is not an a tag. Both are in the file on purpose.

The same job with the standard library parser

A regular expression is what MU's row asks for. A parser is what the job actually deserves, and being able to say why is worth a mark.

from html.parser import HTMLParser


class LinkFinder(HTMLParser):
    """Collects the href of every anchor tag."""

    def __init__(self):
        super().__init__()
        self.links = []

    def handle_starttag(self, tag, attrs):
        if tag == "a":
            for name, value in attrs:
                if name == "href":
                    self.links.append(value)


with open("page.html") as handle:
    html = handle.read()

finder = LinkFinder()
finder.feed(html)

print(f"the parser found {len(finder.links)} link(s):")
for link in finder.links:
    print("  ", link)
the parser found 7 link(s):
   https://www.munotes.in/notes/
   https://www.munotes.in/syllabus/
   https://www.munotes.in/aibe/
   /for-committees/
   /solved-papers/
   /blog/
   #top

HTMLParser is in the standard library, so nothing needs installing. It reads the document as a document: it lower cases the tag and the attribute names itself, it handles every quoting style because that is the specification it implements, and it cannot be fooled by a > inside an attribute value.

Why the parser is right and the pattern is not, in one example

import re
from html.parser import HTMLParser


class LinkFinder(HTMLParser):
    def __init__(self):
        super().__init__()
        self.links = []

    def handle_starttag(self, tag, attrs):
        if tag == "a":
            for name, value in attrs:
                if name == "href":
                    self.links.append(value)


tricky = '''<a title="the > sign in a title" href="real.html">a link</a>
<!-- <a href="commented-out.html">this is inside a comment</a> -->
<a href="also-real.html">another</a>'''

by_pattern = re.findall(r'''<a\s[^>]*?href\s*=\s*(["']?)([^"'>\s]+)\1''',
                        tricky, re.IGNORECASE | re.DOTALL)
finder = LinkFinder()
finder.feed(tricky)

print("the pattern says:", [pair[1] for pair in by_pattern])
print("the parser says :", finder.links)
the pattern says: ['commented-out.html', 'also-real.html']
the parser says : ['real.html', 'also-real.html']

Read the two lines. The pattern lost real.html, because the > inside the title attribute looked to [^>]*? like the end of the tag. And it found commented-out.html, which is inside an HTML comment and is not a link on the page at all.

HTML is not a regular language, so no regular expression can parse it. That is the honest answer, and the right thing to write in the journal: the pattern is what the exercise asks for and it works on ordinary pages; a parser is what to reach for when the input is not under your control. Saying that, rather than defending the pattern, is what a good answer looks like.

What to do with the links once you have them

import re
from urllib.parse import urljoin, urlparse

with open("page.html") as handle:
    html = handle.read()

pattern = re.compile(r'''<a\s[^>]*?href\s*=\s*(["']?)([^"'>\s]+)\1''',
                     re.IGNORECASE | re.DOTALL)
links = [m.group(2) for m in pattern.finditer(html)]

absolute = [link for link in links if urlparse(link).scheme]
relative = [link for link in links if not urlparse(link).scheme
            and not link.startswith("#")]
fragments = [link for link in links if link.startswith("#")]

print("absolute  :", absolute)
print("relative  :", relative)
print("fragments :", fragments)
print()
base = "https://www.munotes.in/notes/"
print("resolved against", base)
for link in links:
    print("  ", urljoin(base, link))
munotes.in108

Practical 9 continued: Extracting Every Hyperlink from an HTML File

absolute  : ['https://www.munotes.in/notes/', 'https://www.munotes.in/syllabus/', 'https://www.munotes.in/aibe/']
relative  : ['/for-committees/', '/solved-papers/', '/blog/']
fragments : ['#top']

resolved against https://www.munotes.in/notes/
   https://www.munotes.in/notes/
   https://www.munotes.in/syllabus/
   https://www.munotes.in/aibe/
   https://www.munotes.in/for-committees/
   https://www.munotes.in/solved-papers/
   https://www.munotes.in/blog/
   https://www.munotes.in/notes/#top

urlparse splits an address into its parts, and a scheme such as https is what makes an address absolute. urljoin resolves a relative address against the page it was found on, which is what a browser does and what a crawler has to do. Neither needs installing.

The link text as well as the address

import re

with open("page.html") as handle:
    html = handle.read()

pattern = re.compile(
    r'''<a\s[^>]*?href\s*=\s*(["']?)([^"'>\s]+)\1[^>]*>(.*?)</a\s*>''',
    re.IGNORECASE | re.DOTALL)

for match in pattern.finditer(html):
    address = match.group(2)
    words = re.sub(r"\s+", " ", match.group(3)).strip()
    print(f"  {address:<36} {words}")
  https://www.munotes.in/notes/        Notes, double quoted
  https://www.munotes.in/syllabus/     Syllabus, single quoted
  https://www.munotes.in/aibe/         AIBE, not quoted at all
  /for-committees/                     href is the SECOND attribute
  /solved-papers/                      the tag is spread over two lines
  /blog/                               the tag is in capitals
  #top                                 a fragment

Group 3 is the text between the opening and the closing tag, and re.sub(r"\s+", " ", ...) folds any run of whitespace into one space, which is what a browser does when it displays it. The closing </a\s*> allows a space before the >.

Procedure

  1. Create page.html with links using double quotes, single quotes, no quotes, href as a later

attribute, a tag split across two lines, a tag in capitals, an a tag with no href, and an img tag. That file is the exercise.

  1. Save as practical9b.py. Run the naive pattern r'<a href="(.*?)">' first and record how many

of your links it missed.

  1. Note that the naive pattern also returns one address that is not an address, and say why.
  2. Build the pattern up in four steps, printing findall at each: double quotes only, either

quote, quotes optional, then case insensitive.

  1. Show greedy against lazy on one line holding two links.
  2. Write the answer with the back reference \1 and both flags, and print every link numbered.
  3. Confirm that the a tag with no href and the img tag are not in the output.
  4. Write the HTMLParser version and show it gives the same answer.
  5. Run both over the tricky sample with a > inside an attribute and a link inside a comment, and

record the two disagreements.

  1. Split the links into absolute, relative and fragment, and resolve them with urljoin.
munotes.in109

Practical 9 continued: Extracting Every Hyperlink from an HTML File

Result

The naive pattern returned two results out of seven links: it missed the single quoted, unquoted, later-attribute, multi-line and capitalised ones, and it mangled the fragment link into #top" title="to the top by running on to the "> at the end of the tag. The four step build recovered each family in turn. Greedy . swallowed the tag between two links on one line and lazy .? did not. The final pattern, with the back reference and both flags, found every hyperlink in page.html and did not report the a tag with no href or the img tag. HTMLParser gave the same list. On the tricky sample the pattern lost the link whose title attribute contained a > and wrongly reported a link inside an HTML comment, while the parser got both right.

Where marks are lost

  • Only handling double quotes. Single quotes and no quotes are both ordinary HTML.
  • Assuming href is the first attribute. <a class="..." href="..."> is everywhere.
  • A case sensitive pattern. <A HREF= is valid.
  • Greedy . instead of .?, which pairs the first quote with the last on the line.
  • No capturing group, so findall returns whole tags instead of addresses.
  • Reporting <a name="..."> as a link. It has no href.
  • Reporting <img src="...">. It is not a hyperlink.
  • Claiming a regular expression parses HTML. It does not, and saying so is the better answer.
  • Not showing the naive pattern failing. The comparison is the evidence.

For the journal

The aim in MU's words, with her <a href="..."> copied exactly. page.html written out as you made it, and one line saying which awkward shape each link is there to test. The naive pattern and its output, with a count of what it missed. The four step build with the output at each step. The greedy against lazy pair. Then the answer, with the table explaining each piece of the pattern and the numbered list of links. One sentence on the back reference: \1 is whatever group 1 matched, so the closing quote is forced to be the same as the opening one. Then the parser version and the tricky sample where the two disagree, and the conclusion: HTML is not a regular language, so the pattern is right for an ordinary page and a parser is right when the input is not under your control.

Quick revision

  • The naive r'<a href="(.*?)">' misses single quotes, no quotes, a later href, a multi-line

tag and capitals, and mangles a link followed by another attribute.

  • A capturing group (...) makes findall return the group instead of the whole match.
  • ["\']? is an optional quote of either kind. [^"'>\s]+ is an address: no quote, no >, no
munotes.in110

Practical 9 continued: Extracting Every Hyperlink from an HTML File

space.

  • is greedy and ? is lazy. Matching up to a delimiter needs the lazy form.
  • \1 is a back reference to group 1, which forces the closing quote to match the opening

one.

  • re.IGNORECASE for <A HREF=, re.DOTALL so a class may cross a newline.
  • Use '''triple quotes''' for a pattern containing both quote characters.
  • <a\s not <a, or the pattern also matches <abbr and <article.
  • html.parser.HTMLParser with handle_starttag does the job properly, and is in the standard

library.

  • HTML is not a regular language. A pattern fails on a > inside an attribute and on a link

inside a comment. Say so.

  • urlparse(link).scheme tells absolute from relative; urljoin(base, link) resolves it.

Questions you should be able to answer

1. What is wrong with r'<a href="(.*?)">' on an ordinary page? It requires double quotes, requires href to be the first attribute, is case sensitive and cannot cross a newline. It also returns a wrong address when another attribute follows, because it runs on to the "> at the end of the tag.

2. What does a capturing group do to findall? It makes findall return what the group matched rather than the whole match. With two groups it returns tuples.

3. What is the difference between . and .? . is greedy and takes as much as it can; .? is lazy and takes as little as it can. Matching up to a closing quote needs the lazy form.

4. What does \1 mean in a pattern? Whatever group 1 matched. Here it forces the closing quote to be the same character as the opening one.

5. Why <a\s rather than <a? Because <a also matches the start of <abbr and <article. The \s requires whitespace after the tag name.

6. Why is [^"'>\s]+ better than .*? for the address? Because it cannot cross a quote, a > or a space at all, so it does not depend on getting greediness right.

7. What does re.DOTALL change? It lets . and a negated class match a newline, so a tag spread over two lines is still one match.

8. Give two cases where the pattern is wrong and the parser is right. A > inside an attribute value, which ends the tag as far as the pattern is concerned, and a link inside an HTML comment, which the pattern reports and the parser ignores.

9. Why can no regular expression parse HTML in general? Because HTML is not a regular language: it nests, and comments, attribute values and character references can all contain the characters a pattern uses as landmarks.

munotes.in111

Practical 9 continued: Extracting Every Hyperlink from an HTML File

10. Which of these is a hyperlink: <a name="top">, <img src="x.png">, <a href="#top">? Only the last. The first has no href and the second is not an a tag.

munotes.in112

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Report or request
Done!